Describe your video in a Telegram chat and get a 16:9 YouTube thumbnail back, with a bold title, an expressive face and colours that stand out in a feed. The bot is an AI thumbnail generator that works like a normal conversation: you write what the thumbnail should show and say, it replies with the image in about a minute. Your first 5 credits are free, and there is nothing to install.
An AI thumbnail maker that lives in a chat
Most thumbnail tools are template editors: you drag a stock photo onto a canvas, type a title, fiddle with outlines and drop shadows. Here you skip the canvas. You write one message that says what the thumbnail shows, what the title reads and how it should feel, and the whole image comes back finished, text included.
That covers the thumbnails people actually need:
- Reaction thumbnails, a big expressive face next to the thing the video is about.
- Title-led thumbnails, where three or four huge words carry the click.
- Your own face, attached as a photo, so every thumbnail on the channel looks like you.
- Quick variants, the same idea in two or three versions to test which one gets clicked.
The bot is not affiliated with YouTube or Google. It makes the image, and you upload it to your video yourself.
How to make a YouTube thumbnail in Telegram
- Open the bot in Telegram and press Start.
- Claim 5 free credits. Subscribe to @zosee_channel, then return to the bot and tap the claim button. The subscription alone grants nothing, the tap is what releases the credits. See the free credits page.
- Choose GPT-Image-2 in the model menu and the 16:9 aspect ratio.
- Send your prompt as a normal message, with the title text in quotes. The thumbnail comes back into the chat.
If something is off, reply to the result with the correction (“make the text bigger”, “change the background to green”) and the bot edits its own output instead of starting over.
Thumbnail prompts to paste
These are the prompts behind the two examples on this page, unchanged.
Cooking video (the ramen):
YouTube thumbnail for a cooking video, 16:9. On the left, a close-up of a surprised young chef with wide eyes and an open-mouth smile, holding a wooden spoon, warm kitchen light. On the right, a steaming bowl of glossy red ramen with a soft-boiled egg. Huge bold white title text with a thick black outline reading “10 MINUTE RAMEN” across the top right. Bright saturated colors, high contrast, clean uncluttered background, readable at small size.
Tech review (the phone):
YouTube thumbnail for a tech review video, 16:9. On the right, a man in his thirties with short dark hair and a skeptical raised eyebrow, looking straight at the camera, lit by a cyan rim light. On the left, a sleek unbranded black smartphone floating at an angle with a glowing screen. Huge bold yellow title text with a thick black outline reading “WORTH IT?” on the left. Dark blue gradient background, high contrast, simple composition, readable at small size.
Both follow the same order: what the video is, who is in the frame and with what expression, the one object the video is about, then the title in quotes with its colour and outline, then the overall look. Swap the subject and the title and keep the rest, and your thumbnails start to look like one channel.
Thumbnails with your own face
A thumbnail with the creator’s face usually beats one with a stranger’s, because viewers learn to recognise you. Attach a sharp, well-lit photo of yourself and put the prompt in the caption of that same message. A photo sent alone, or a prompt sent as a second message, will not be paired.
YouTube thumbnail, 16:9, built around the person in the attached photo with a shocked expression, on the left. On the right, a stack of cash. Huge bold yellow title text with a thick black outline reading “I WAS WRONG”. High contrast, clean background, readable at small size. Keep the face from the attached photo unchanged.
The expression is yours to direct: “shocked”, “laughing”, “skeptical, one eyebrow raised”. The AI avatar generator page has more on keeping a face intact through an edit.
Which model to pick for thumbnails
Text inside the image is the whole job here, so the choice is mostly about lettering:
- GPT-Image-2 is the one to start with. It renders short titles accurately, follows the layout you describe literally, and is the cheapest: 1 credit at 1K, 2 at 2K, 3 at 4K. Both examples above were made with it.
- Qwen3 is the other lettering specialist, and the one to use when the prompt is long and detailed and every part of it has to land. 2 credits at 1K or 2K, Qwen3 Pro 4 to 6 credits.
- NanoBanana Pro is for a busy, photorealistic scene that has to be flawless, 4 to 5 credits.
- NanoBanana 2 handles a single short word cleanly and is a good all-rounder at 2 to 4 credits.
All four offer 16:9.
What it costs
1 to 6 credits per thumbnail depending on the model and resolution, and credits never expire. A 1K GPT-Image-2 thumbnail is 1 credit, so testing three variants of one video costs 3. Payment is Telegram Stars, cryptocurrency, a bank card from any country ($), or a Russian bank card (₽). The pricing page has the packs.
Tips for thumbnails that get clicked
- Keep the title to 3 to 5 words. A thumbnail is seen at the size of a thumbnail. A long sentence turns into grey noise, and short titles are also the ones a model spells right every time.
- Put the exact text in quotes. Write
reading "10 MINUTE RAMEN", not “a title about quick ramen”. The model copies quoted text; it invents unquoted text. - One focal face. One person, one clear expression, looking at the camera. Two faces compete, and a crowd reads as nothing at small size.
- Ask for contrast. Bright text with a thick dark outline on a simple background, and say “high contrast” outright. Busy backgrounds eat the title.
- Check it small. Look at the result zoomed out on your phone before you upload it. If you cannot read the title at a glance, reply “make the text bigger and the background simpler”.
- Read the text before you publish. Check every letter. If a word comes out wrong, reply with the correct spelling in quotes and the bot fixes it.
For more on getting text inside an image right, see how to prompt GPT-Image-2 and how to prompt Qwen3.
Related
- How to prompt GPT-Image-2, the full guide to the model behind these examples.
- AI avatar generator, for a channel profile picture made from your own photo.
- AI photo editor, to fix a screenshot or photo before it goes into a thumbnail.
- Change a photo background, to cut yourself out of a messy room.
- How to use the bot, the walkthrough from Start to your first image.
- Telegram channel avatar generator, for a channel, group or bot avatar that survives the round crop.
- AI images for TikTok, for vertical 9:16 slideshows, covers and trend edits.
- Instagram photo generator, for aesthetic feed posts and Stories from one selfie.