Project ideas for vibe coding: what to build this weekend — apps, bots, mini-sites — and how to ship it.

A live camera with recognition used to mean a server, a video stream, and lag. In July Google brought a fast GPU engine into the browser: your model runs on the video right there, and nothing leaves your device. A weekend build.

On July 20 ByteDance unveiled Seed Audio 1.0: footsteps, rain, hum and voices arrive in one request, already mixed into a single track. A weekend-sized ambience generator rides on it.

On July 8 Cloudflare removed the beginner's biggest wall: the live link comes BEFORE the account. Build an invite page with AI and ship it in a minute.

A talking character used to mean animators, lip-sync, and a game engine. Now a drawing comes alive from a single image and holds a real conversation. Bring your hero to life over a weekend.

AI used to draw someone who merely looked like you. Now it holds your actual face and drops you into the 90s or a cartoon. Build the generator for it this weekend.

Scratches, a yellow haze, blurry faces. This used to take hours in Photoshop — now a model fixes an old photo in one pass. Build the little tool for it this weekend.

Every «now break it down by month» used to reopen the file from scratch. Now the model has a live Python notebook: variables survive between requests, the container lives 30 days, 1,550 hours a month are free.

Rules usually get stuffed into a user message — and the user cancels them with one line. On July 15 the API let the system role sit in the middle of a conversation: that's your voice as the operator, and it outranks. A weekend build.

You used to wait for a picture: hit «generate», stare at a spinner. Now it chases your text live — add a word, the image changes instantly. A weekend build.

A coloring book used to mean drawing every page by hand. Now you type «a space cat firefighter» and a minute later you have eight black-and-white pages to print. In an evening.

AI used to draw a QR-ish pattern — point your camera and it leads nowhere. Now the model writes code and drops a working QR right into the art. With a ready prompt.

A normal bot forgets everything the moment you close the chat. Now memory is a separate service: the bot recalls your name and habits a week later. Free, with a ready prompt.

Before, nudging text on a poster meant regenerating the whole picture. Now it arrives as a stack of layers: drag one, the rest stay put. With a ready-made prompt.

Before, an agent lived only while the tab was open. Close the laptop on the train — it died halfway. Now you set background: true and walk away. With a ready-made prompt.

The whole game world is one stream from a model. Every turn used to stall on 'loading…'. Now the answer pours out in a blink — and the story comes alive. Starter prompt included.

Recorded a voiceover and one word came out wrong? It used to mean re-recording the whole take. Now you change one word and the model resynthesizes only that — in your own voice, seamlessly. Open-source, a weekend build.

Fixing one detail in an AI clip used to mean rebuilding the whole thing — and it came back different. Now you say 'make it night' and only the sky changes. Everything else stays. A weekend build.

A 'swarm of agents' used to be something you wired up by hand, and ten in parallel cost real money. On July 9 OpenAI baked it into the model and a cheap tier made it pennies. Build your own mini-research this weekend.

Running a smart model over every one of 200 emails used to be too expensive — so people filtered by keywords. On July 9 a cheap tier landed at $1 per million tokens. Build your personal 'what matters' screen this weekend.

Voice used to only answer back — you still set the timer yourself. On July 6 OpenAI shipped speech-to-speech that calls your functions on the fly, so now the voice presses the buttons while your hands are busy. A weekend build.

Drop in a document, get a phone number in a couple of minutes — call it and ask by voice about your own file, and it answers. A year ago this was a week of wiring.

You write one sentence — what the game is — and an agent makes the art, code and sound and hands you a browser link. A year ago this wasn't this easy — now it's a weekend build.

On June 4 Krea released Turbo as open weights — a 2K frame in ~2 seconds on ordinary hardware. It breaks the 'make one, wait, tweak' habit. Now you generate twenty options at once and pick. That's a weekend idea-fitting-room for interiors or looks.

On June 13 OpenRouter shipped Fusion: one request fans out to a panel of models, and a judge merges them into a single answer while flagging where they disagree. That's a weekend bot for the questions that matter — with an honest 'the models don't agree here' note.

A job interview, a rejection, an important call — the stuff you're scared to fumble on the first try. On July 1 xAI opened a voice-agent builder: describe it in words, and in two minutes it calls and talks on its own, ~5 cents a minute. A sparring partner you build in an evening.

Automating a website used to mean a brittle script tied to specific buttons — one redesign and it broke. Now the agent just looks at the screen and clicks, like a human. A weekend build.

Take a video you already made and get it back in English — your own delivery, no 'translated… waited… spoke' pauses. A year ago auto-dubbing sounded like a GPS. A weekend build.

Synthesis used to read everything as a flat narrator. Now you drop [whispering] or [excited] into the text and the model performs the scene. A tiny audio play, built in a weekend.

The built-in browser model used to only read text. With Chrome 148 it sees an image and returns structure — a photo of a poster becomes a card. A weekend build.

A translator shows a word's meaning. The fear is saying it out loud. ChatGPT's new feature lets you build a coach that sounds words out, syllable by syllable. A weekend build.

An AI feature used to need a backend, a key, and a token bill. Now the model lives right in the browser — and one file is enough. A weekend build.

Upload one selfie and get eight stickers, all recognizably you: laughing, sad, giving a thumbs up. Faces used to drift from frame to frame. Now they hold.

Record half a minute of yourself, and the app sends a voice message in a language you never learned — still sounding like you. A year ago this didn't work. A weekend build.

Describe a mood in words, get a track you can actually publish. The music API is now trained on licensed data. With a ready brief.

This used to need a vector database and RAG. Now a whole year of chats fits in one prompt — million-token context just got far cheaper.

Not an image, a video: write one line, and a minute later you've got an 8-second vertical clip with sound, ready for Stories. Veo 3.1 in the Gemini API. Starter prompt included.

Big files used to get sliced up and searched (that's RAG). GLM-5.2 just shipped — a million tokens of context, open and cheap. Now the whole document fits in one request. Starter prompt included.

Hand the model an hour-long lecture or tutorial and get back a table of contents: tokens at 02:14, embeddings at 09:40. A year ago this wasn't this easy — now it's a weekend build.

Not a caption under the picture — the same photo, with the words translated and sitting right where they were. This used to be a whole pipeline; now it's one request.

"Where do I turn this on?" — you used to describe what was on your screen. Now the model sees your screen itself and guides you by voice, step by step. A weekend build.

You used to snap a photo and wait for text. Now the model watches a live camera feed and talks back in real time — and you can interrupt it like a person. A weekend build.

AI composers with real vocals used to be pricey and locked in a closed app. Now a full song with voice and lyrics is one request for ~$0.03. That's a weekend greeting-song generator.

Turning a real object into 3D used to mean dozens of photos and special software. Now one snapshot is enough — the model fills in what the photo doesn't show. A weekend build.

Subscriptions, terms, leases — the fine print nobody finishes. Now a cheap model reads it for you and flags where the trap is.

Not «what should I cook» in general, but a recipe from the food in your photo, plus a short list of what to buy. A year ago this wasn't this easy.

You used to just guess whether it'd fit, or wrestle Photoshop. Now the model fuses two of your photos into one scene — the thing, in your room. A weekend build.

It used to redraw your hero from scratch in every frame — and the comic fell apart. Now the face holds from panel to panel. A weekend build.

A restock or ticket watcher used to mean coding a parser, a schedule, a diff. Now the agent does the looking and decides what matters. A weekend build.

"Agent" used to mean you wired up the think→go→check loop by hand. Now it's one API call: the server runs the loop for you. A weekend build.

Narration used to make you wait: half a minute to render, the kid staring at a spinner. Now the model speaks from the first word. A weekend build.

4000 screenshots and you can't find the one you need. Now the model searches the picture itself, not the text on it — and pulls up that exact shot. A weekend build.
The model used to see a single picture. Now it watches the whole video, catches the best moment, and turns it into a cover. A weekend build.

A translator used to mean a cloud API, a key, and per-character fees. Now the browser translates on its own: offline, free, 145 languages. An evening build.

A normal model draws the world as it memorized it during training. The new Seedream searches the web during generation — so the poster comes out with today's facts. That's enough for a weekend project.

Animating a photo used to mean a video editor and a lot of fuss. Now one photo and one sentence is enough — the model makes a short clip with sound on its own. That's a weekend build.

Not 'record — wait — read the reply,' but a real spoken conversation in real time. Last week OpenAI shipped a model that answers with no pause and doesn't break when you interrupt it.

No API key, no per-request bill: the model already lives in your iPhone. Last week Apple added image input and free cloud — building something useful is now a weekend job.

Feed your saved articles and notes to a model and get back an mp3 where two hosts chat about them. A year ago this wasn't this easy — now it's a weekend build.

A model used to glance at a photo and guess 'about twenty'. Now it zooms in and counts exactly — and it's a weekend build.

Image models used to paint a gorgeous background and drunken letters on top. On June 3 Reve 2.0 shipped — it builds the layout first, then the image, so the text on a flyer, menu or sign actually reads.

A year ago a 'smart journal' meant your most private notes flew off to someone else's server. On June 8 Apple opened free access to its model right on the device — now the AI reflects offline, with no key and no API bill.

You mutter your thoughts into the phone while walking, and the app hands back not a wall of text but sorted items: tasks, decisions, ideas. And it doesn't trip over your jargon.

Record 30 seconds of your speech, and any text gets read out in your own voice. A year ago you needed a studio. Now it's a weekend build.

A short list of ideas, each with the skill you'll practice. From a timer to a Telegram bot. Pick one and start today.