๐Ÿ–ผ๏ธ Nano Banana โ€” Create State-of-the-Art AI images from 4ยข cents each โ†’
๐ŸŽฌ XYZ Generator
๐Ÿ“ข
UGC Marketing Guide

What Is UGC Marketing and How to Create UGC Videos with XYZ Generator

Published: 2026-06-20 | Category: Marketing Guide

UGC marketing is one of the most effective ways to drive sales and engagement โ€” and with AI, creating UGC-style videos no longer requires hiring influencers or shooting with a crew. This guide explains what UGC marketing is, what it costs the traditional way, and exactly how to create UGC videos with XYZ Generator.

Creator filming a UGC-style marketing video with a smartphone

What Is UGC Marketing?

For a long time, one of the most effective marketing techniques for driving sales and engagement has been influencer and UGC marketing. The traditional approach involves hiring a human content creator or influencer to promote a product through their social media channels.

This type of User-Generated Content (UGC) is particularly effective because it combines a real person, a natural voice, and an authentic product demonstration. Instead of presenting the product as a traditional advertisement, the creator can demonstrate how it works, share their experience, and interact with the product in real time. This makes the content feel more relatable and engaging to the audience.

The Cost of Traditional UGC

However, traditional UGC has a significant drawback: cost. A decent UGC creator can charge approximately $150โ€“$400 per video. At two videos per week, that can amount to roughly $1,300 per month for a single creator. For a campaign using ten creators, the cost can quickly reach approximately $13,000 per month.

This is before considering the additional time required to find suitable creators, negotiate rates, send products or briefs, coordinate production, and wait for the content to be delivered.

Generate UGC with AI

Fortunately, there is now a faster and more cost-effective way to produce this type of marketing content: AI-generated UGC. Using XYZ Generator and its advanced AI video models, you can create UGC-style videos without hiring and coordinating individual creators for every piece of content.

The fundamental advantage is speed and scalability. Instead of treating every video as a separate production involving a human creator, the process can be handled digitally and repeated whenever new content is required.

A Lower Barrier to Content Production

The barrier to producing UGC-style marketing content is now significantly lower. Just a few months ago, producing a creator-style promotional video could require a substantial budget and a multi-step process:

  • Find a suitable creator.
  • Contact the creator and negotiate an offer.
  • Wait for their response and availability.
  • Provide the product and creative brief.
  • Wait for production and delivery.
  • Review and potentially request revisions.

With AI video generation, much of this workflow can be reduced to a single day. Instead of waiting for a creator to become available, you can generate, review, and iterate on UGC-style videos directly. The result is a marketing workflow that is faster, more scalable, and significantly less expensive to experiment with.

How to Create UGC Videos with AI: A 3-Step Workflow

Step 1: Plan the Visual Environment

Once the research is complete and the script has been written, the next step is to define the visual environment and determine how it should be represented in the video-generation prompts. There are two primary approaches:

  • Prompt the video model directly โ€” provide the script and visual instructions and generate the video.
  • Create a storyboard first โ€” generate individual images representing each scene, review them, and use the approved storyboard as a reference for the final video.

For consistent, high-quality results, the storyboard-first approach is recommended.

Step 2: Build and Review the Storyboard

Generating still images is generally faster and less expensive than generating complete video sequences. This makes the storyboard an efficient environment for testing and refining creative decisions before committing to video generation. Use the storyboard to validate key elements such as:

  • Does the subject look natural, or does she appear overly polished?
  • Is the environment appropriate, or is the room too clean and staged?
  • Is the subject holding and interacting with the correct object?
  • Are the composition, lighting, camera angle, and visual style appropriate?

Iterate on these images until the desired result is achieved. By the time the storyboard is approved, the major creative decisions have already been defined and validated. The video-generation stage can therefore focus on bringing an established visual concept to life rather than experimenting with the underlying creative direction.

AI-generated annotated storyboard for a UGC video advertisement

Step 3: Use the Storyboard as a Reference

The storyboard serves a second important purpose: it becomes a visual reference that is carried forward into the video-generation process. Providing both the prompt and the corresponding reference image gives the video model two complementary sources of information:

  • The prompt describes what should happen.
  • The reference image establishes what the scene should look like.

This reinforces the intended subject, environment, composition, and visual style, increasing consistency between the approved storyboard and the final generated video.

Storyboard Generation Prompt

Here is the prompt used to generate a production storyboard for a 30-second vertical UGC advertisement in XYZ Generator:

Storyboard Prompt

Create ONE single wide 16:9 annotated production STORYBOARD page for a 30 second vertical UGC advert. This output is a director's board document, not six loose photos, not a collage, and not a contact sheet without labels.

PAGE = STORYBOARD. The entire canvas is one working board on plain off-white paper. You must show the board chrome: title bar, frame borders, numbered labels, annotation text under each frame, and a timing strip. If the title bar, annotations or timing strip are missing, the image is wrong โ€” regenerate as a storyboard, not as photos.

LAYOUT: EXACTLY 6 SCENE FRAMES inset into that board in a single horizontal row, numbered 1 to 6, equal size, with generous white paper showing between them. Each inset is a sequential shot keyframe a video model will expand into motion, so every inset has to work on its own as a still โ€” but they live on the board page, they are not separate deliverables.

BOARD DESIGN: a working director's board on plain off-white paper, not a clean grid of product photographs and not a seamless photo montage. A title bar across the top with the concept name "Sona โ€” reframe (older woman)" and the runtime "30s", and the "XYZ Generator" logo small in its top right corner. Thin dark borders around each of the six frames. Under the row of frames, a thin timing strip divided into the six beat lengths (5s, 3s, 6s, 7s, 5s, 4s). Two short hand-drawn arrows between frames, one between 3 and 4 marking the swap from five bowls to one plate, one between 4 and 5 marking the cut from night to daylight. Readable text is allowed on the board chrome and nowhere inside the scene frames.

BOARD ANNOTATIONS, outside the scene frames only, printed under each frame in small readable type: its number, the shot name, whether the beat is spoken to camera or an insert, the camera note, the beat length in seconds, a continuity note, and the spoken line for that beat.

SCENE FRAMES (the six insets only): photographic, not illustrated, 9:16 inside each bordered frame. The artwork inside every frame is text-free: no words, captions, subtitles, speech bubbles, UI copy, wordmarks, labels, packaging text or device text.

INSIDE EACH SCENE FRAME: treat the inset as a frame paused out of a front-camera phone video, not a studio still. Film the talking-head insets like an authentic social-media UGC confessional: phone propped at eye level, her face filling most of the frame, eyes locked on the lens, subtle micro-shake and a slight off-centre tilt, mild wide-angle phone distortion, mild sensor noise. Soft practical face light from a warm lamp near the phone, not a beauty ring light and not portrait retouching.

Her face should show authentic mature skin texture: visible pores, fine lines, natural wrinkles around the eyes and mouth, slight unevenness, and realistic skin texture. Avoid making her look artificially young or heavily retouched. Catch her mid-sentence: mouth open on a consonant, one eyebrow slightly raised, a natural expression, not a held pose.

Frame 1 must stop the scroll on its own โ€” a long shiny olive-oil pour mid-stream into a bowl, with her face filling the shot. Never a beauty portrait, never squared catalogue posing, never a symmetrical food flat-lay. If a face inset would pass as a polished influencer-grid still, it is wrong. The standing-snack insert is the only non-face inset and should look like a paused phone POV, not food photography โ€” one punchy mid-bite, not a cooking sequence.

HANDS: every hand in a frame has an explicit job and an explicit place, and any hand without one is out of frame. Prefer one readable hand on talking-head frames. A visible hand is large enough to read, fully inside the frame, and doing one clear thing. Never finger counting or holding up one, two or three fingers. Never a hand half hidden in a lap, tucked against a body, or cropped by a frame edge.

SHOT LADDER: mild on purpose for the confession beats, like a real front-camera creator who barely moves the phone. Face frames are all eye-level front camera with small height/distance changes only: medium chest-up with oil pour, standing-snack hands insert, tighter face-and-shoulders, medium close-up half a step back, medium chest-up at the counter, medium chest-up at the door. No overhead tabletop hero, no low-angle plate hero, no wide establishing shot, no cinematic orbit. The insert is the only hard cut away from her face.

NO PHONE IS VISIBLE IN ANY FRAME. She films this on her own front camera, so the phone is the camera and cannot also be on her table or in her other hand. No phone, tablet or screen appears anywhere, held, propped, face down, on the counter, or reflected in anything.

REFERENCES: @Image 1 is the "XYZ GENERATOR" logo. Reproduce it once, small, in the top right of the title bar, at the size of a production-house mark. It never appears inside a scene frame, on a prop or anywhere in the kitchen, and it does not define the character, the room or the colour of anything else on the board.

CHARACTER, identical in all six frames: MARGARET, a charismatic 68-year-old woman with a warm natural complexion, expressive dark brown eyes, naturally defined eyebrows, soft mature facial features, and shoulder-length silver-grey hair with subtle natural waves. Her hair is neatly styled but not salon-perfect, with a few natural strands slightly out of place.

She has realistic mature skin with visible pores, fine lines, natural crow's feet, forehead lines, and subtle age-related texture. Her appearance should feel healthy, warm, confident, and completely believable โ€” not glamorous, airbrushed, or artificially youthful.

She wears a simple fitted dark charcoal knit top with three-quarter sleeves, small understated gold earrings, a thin gold necklace, and a simple gold watch on her left wrist when that hand is visible. Short natural nails with a subtle nude finish. The exact same face, silver-grey hair, clothing, jewellery and overall appearance must remain consistent across all six frames.

She is not modelling. She behaves like a real older woman casually talking to a friend through her phone camera at home. Her expressions should feel spontaneous, slightly humorous, confident and conversational rather than performed like a professional advertisement.

SETTING, three setups in one clean modern flat and nowhere else. Frames 1, 3 and 4 are a white quartz kitchen island at night. Frame 2 is a hands-only standing-snack insert at that same island later the same night. Frame 5 is the same kitchen in daylight. Frame 6 is the front door of the flat in open daylight.

The kitchen is clean and contemporary: flat-panel white cabinets, soft warm under-cabinet glow, bare quartz counters, a single matte ceramic vase with stems, no clutter, no dish rack, no fridge, no fridge magnets, no takeaway menus, no crumbs. Behind her in every kitchen frame there is enough clean architecture that the space reads as a real modern apartment, not a void.

FRAME 1, THE VISUAL HOOK, spoken to camera: MEDIUM CHEST-UP, front camera propped at eye level on the clean kitchen island, her face filling the upper two thirds and locked on the lens. She has a slightly conspiratorial, amused expression, one eyebrow lifted, mouth open mid-consonant.

Soft frontal kitchen lamp light, small catch lights in both eyes, pale cabinets a stop darker behind her. In the lower third, large and readable: one dinner taken apart into five small matching matte bowls โ€” chicken, rice, broccoli, sauce and oil โ€” arranged loosely rather than in a perfect grid, with a slim digital kitchen scale beside them.

The scroll-stop is the pour: her right hand holds a clear glass olive oil bottle tilted high over the sauce bowl, a long continuous shiny stream of oil falling mid-air into the bowl that is already glistening โ€” the pour that will not stop, the hidden calories you never count. Left arm out of frame. The oil stream must be the brightest, most readable object in the lower third.

FRAME 2, AN INSERT, she is not in shot: CLOSE CHEST-HEIGHT ANGLE looking slightly down at the clean quartz island later the same night, lit only by the soft warm under-cabinet glow. One punch, not a sequence.

On the bare counter: an open glass jar of peanut butter with the lid beside it. Catch the forgotten standing snack mid-bite โ€” left hand lifts a cracker thick with peanut butter up toward the top of frame, motion blur on the cracker, a knife stuck in the jar. No spreading action, no second cracker, no drawn-out prep. No face, no head, no shoulders.

FRAME 3, spoken to camera: TIGHTER CLOSE-UP, face and shoulders only, same island, same propped front camera pulled a little closer so her eyes dominate. The five bowls are only a soft smear at the very bottom edge or out of frame.

Neither hand visible. She talks straight down the lens, mid-word, with a tiny dismissive tip of the head and an amused, knowing expression. Clean pale cabinets soft behind her.

FRAME 4, THE DEMO LINE, spoken to camera: MEDIUM CLOSE-UP pulled half a step back from frame 3, still eye-level front camera propped on the island. The bowls and scale are gone; one plated dinner sits in the lower foreground โ€” chicken, rice and broccoli on a simple white ceramic plate โ€” readable but secondary to her face.

Left hand taps the plate rim once, fully in frame; right arm out of frame. She talks straight down the lens. She is not holding a phone and she is not photographing anything.

FRAME 5, DAYLIGHT, spoken to camera: MEDIUM CHEST-UP, same clean kitchen later, soft window light as frontal fill, phone propped on the counter at eye level so she faces the lens head-on.

She holds a breakfast bowl from underneath in one hand and a spoon in the other, both hands readable, mid-bite, talking naturally with her mouth slightly full, eyes on the lens. Her expression is warm and amused. Behind her the quartz counter is clear and the cabinets stay clean and modern.

FRAME 6, DAYLIGHT, spoken to camera: MEDIUM CHEST-UP at the front door, front camera at arm's length but still eye-level and close, a clean modern hallway and a strip of street soft behind her rather than a wide establishing shot.

Camera hand out of frame behind the lens; other hand grips a simple tote strap high on her chest, fully readable. She talks to the lens and starts to laugh naturally at her own last line. Slight tilt, a little overexposed by the open door.

OBJECT-STATE CONTINUITY: five bowls, scale and oil pour in frames 1 and 3; single plate from frame 4 onward. Frames 1โ€“4 night, 5โ€“6 daylight. Same face, same silver-grey hair, same clothes and jewellery across every cut. Kitchen stays clean and modern in every kitchen frame. One person only.

NEGATIVE, inside the scene-frame insets only: no studio lighting, no professional camera look, no cinematic grade, no overhead hero food angle, no low tabletop hero angle, no wide establishing shot, no side three-quarter interview angle, no food photography styling, no cluttered counters, no dirty dishes, no dish rack, no fridge, no open fridge, no fridge magnets, no takeaway menus, no crumbs, no plastic leftover tubs, no catalogue posing, no beauty retouching, no artificially young face, no doll skin, no porcelain skin, no airbrushed skin, no skin blur, no glossy plastic finish, no excessively smooth skin, no heavy makeup, no dyed hair, no blonde hair, no second person, no phone or screen, no unassigned second hand, no finger counting, no warped or extra fingers, no long snack-prep sequence in frame 2, no second cracker being spread in frame 2.

FINAL CHECK: this must read as one annotated storyboard page at a glance โ€” off-white paper, title bar, "XYZ GENERATOR" logo, six bordered insets, labels under each, timing strip, two arrows. If it looks like six naked photos or a collage with no board chrome, it is wrong.

Readable text only outside the frames, logo only in the title bar. Six frames, no extras. Face insets should look like the same propped front-camera confession with mild reframes, not six different film setups. Frame 1 must read as a mid-pour oil stream, not a static fork hold. Count the hands in every frame.

BOARD NEGATIVE: no six separate images, no full-bleed photos without a board, no missing title bar, no missing annotations, no missing timing strip, no seamless collage, no contact sheet without labels, no mood board, no poster.

Recommended Structure for Prompting an AI Video Model

A reliable way to prompt an AI video model is to structure the prompt into clearly defined sections. Each section controls a specific aspect of the generation, while repetition across sections reinforces the most important requirements.

1. References

If you attach reference files or images, mention them first. Give each reference its own line and explicitly describe what it controls and what it does not control. An untagged reference may still influence the generation, but the model is free to interpret it however it chooses.

2. Camera

Define how the video is filmed: what camera or device is used, who holds or operates it, the position and framing, and how the camera moves. For UGC-style content, explicitly describe the imperfections that make footage feel authentic: subtle shake, drifting frames, focus hunting, uneven exposure, and small framing changes. Models such as Seedance tend to produce clean and stable footage by default, so these imperfections need to be explicitly requested rather than assumed.

3. Look

Describe the visual characteristics of the footage in concrete terms: lighting, colour, contrast, grain, and overall image quality. Avoid vague descriptions. Skin appearance should also be defined here โ€” realistic texture, visible pores, natural slight unevenness, and no airbrushed or porcelain skin โ€” to prevent the polished look commonly associated with generated imagery.

4. Style

Define how the content is performed, not just how it looks. For a UGC example, the character speaks directly to the camera whenever her face is visible, and voiceover is used only during inserts where she is off-screen. This requirement should be reinforced throughout the prompt in the STYLE, STAGES, and CONSTRAINTS sections.

5. Voice

Treat the voice the same way you treat the character's visual identity. Specify pitch, speaking pace, accent, speech patterns, energy and delivery, and anything to avoid. The objective is to cast the voice deliberately rather than letting the model choose an arbitrary voice.

6. Character

Define the person appearing in the video. At minimum, specify age, build, hair, clothing, and three or four distinctive facial characteristics. These details establish the character's identity and help maintain consistency across multiple shots.

7. Setting

Describe the environment in which the video takes place. Explicitly name every location the advertisement is allowed to visit and exclude locations that should not appear โ€” if you do not define the geography, the model may introduce new environments between shots.

8. Stages

This is the main body of the advertisement. Break the video into individual beats, with each stage describing one meaningful change and its intended end state. For a 30-second video, aim for approximately 110โ€“120 words of spoken content, roughly 15 spoken words per four seconds. For example, one 30-second structure might use 5s โ†’ 7s โ†’ 3s โ†’ 7s โ†’ 4s โ†’ 4s.

9. Brand

Define the spoken brand requirements: what must be spoken, who says it, and how the brand name is pronounced. Do not rely on this section for on-screen text โ€” many models struggle with rendered text, so spoken brand instructions should be treated separately from visual text requirements.

10. Audio

Define the complete audio environment. Describe sound effects as physical events rather than generic labels, specify the voice characteristics, and make an explicit decision about whether music is allowed โ€” for example, "[SOUND] Strictly only naturally occurring sound and foley, no music allowed."

11. Consistency

Define everything that must remain unchanged between cuts: character identity, clothing, hair, props, object counts and states, locations, and spatial relationships. Write numerical quantities explicitly as numbers (for example, 5 bowlsrather than "several bowls") because counts are among the first things that drift between generated shots.

12. Constraints

The final section should contain the complete ban list. Use it for everything the model must notdo. This is the prompt's final quality-control layer: the creative sections tell the model what to create, while the constraints tell it what to avoid. The most effective constraint lists are built from actual generation failures.

๐Ÿ“‹ Recommended Prompt Order

REFERENCES โ†’ CAMERA โ†’ LOOK โ†’ STYLE โ†’ VOICE โ†’ CHARACTER โ†’ SETTING โ†’ STAGES โ†’ BRAND โ†’ AUDIO โ†’ CONSISTENCY โ†’ CONSTRAINTS

This structure separates the major dimensions of video generation while repeatedly reinforcing the requirements that are most likely to drift โ€” particularly character identity, camera behaviour, performance style, audio, object counts, and continuity.

Full Video Prompt Example

Below is the complete 30-second video prompt built with this structure โ€” a UGC-style ad starring "Margaret", an older woman filming herself on her front camera. Paste it into XYZ Generator along with the approved storyboard image.

UGC-style AI video frame of a woman talking to camera for a product ad

Video Prompt (30s UGC Ad)

### REFERENCES

@Image 1 is the approved production storyboard for this ad, containing six frames in order.

Use the six storyboard frames as the authoritative shot plan for the six STAGES below and expand them into one continuous 30-second moving piece.

The storyboard controls the character's appearance, wardrobe, environments, framing, shot progression, object states, lighting changes, and overall visual continuity.

The storyboard is a reference only. Nothing belonging to the storyboard itself appears in the generated video: no panel borders, frame numbers, annotation text, timing strip, arrows, paper background, title bar, or logo from the storyboard.

Do not reinterpret the storyboard as six unrelated images. Treat it as a production plan for one six-shot video.

### CAMERA

Vertical 9:16 smartphone footage, six distinct shots with clean cuts between them.

The video is filmed by Margaret herself using her own front-facing smartphone camera. Talking-head stages are hands-free, with the phone propped at approximately eye level against something on the kitchen counter or table.

The footage should feel like an authentic personal camera-roll recording rather than professionally produced advertising footage.

Maintain subtle imperfections throughout the talking-head shots:

* Small handheld micro-shake.
* Slightly drifting composition.
* Slight off-centre framing.
* Small variations in camera angle and distance.
* Mild wide-angle smartphone distortion.
* Occasional autofocus adjustment.
* Slightly uneven exposure.
* Natural motion blur on moving hands and objects.

The camera movement should remain minimal. This is a casual front-camera confession, not cinematic filmmaking.

Shot progression must follow the storyboard:

1. Medium chest-up at the kitchen island with the oil pour.
2. Close chest-height standing-snack insert.
3. Tighter face-and-shoulders close-up.
4. Medium close-up pulled slightly farther back with the plated dinner.
5. Medium chest-up at the kitchen counter in daylight.
6. Medium chest-up at the front door in daylight.

No overhead tabletop hero shot, no low-angle food shot, no wide establishing shot, no orbit, dolly, crane, tracking shot, or other cinematic camera movement.

The phone is the camera and is therefore never visible. No phone, tablet, monitor, or screen appears anywhere in the video.

### LOOK

The video must look like genuine smartphone footage captured at home.

Use soft practical light rather than professional lighting.

Stages 1, 3 and 4 take place at night in the modern kitchen under a warm practical lamp positioned near the phone. The background falls slightly darker than her face.

Stage 2 is illuminated primarily by the warm under-cabinet kitchen lighting, with the surrounding room falling naturally darker.

Stages 5 and 6 take place in daylight. Stage 5 uses soft window light at the kitchen counter. Stage 6 is slightly overexposed by daylight entering through the open front door.

Warm, natural skin tones. True-to-life colour. Mild smartphone sensor noise. Gentle motion blur on moving hands and the oil stream.

Skin must retain realistic mature texture:

* Visible pores around the nose and cheeks.
* Natural fine lines and wrinkles.
* Natural slight unevenness.
* Subtle texture around the eyes and mouth.
* Slight natural sheen on the nose and forehead.
* No filter-like appearance.

The footage should look like it came directly from someone's actual camera roll.

No cinematic colour grade, studio lighting, beauty lighting, ring light, professional commercial lighting, beauty filter, airbrushed skin, doll skin, porcelain skin, skin blur, or glossy plastic finish.

### STYLE

A charismatic older woman delivering a direct, slightly conspiratorial social-media health and calorie-tracking confession into her front camera.

The performance should feel conversational and spontaneous, as though she is talking to a friend rather than presenting an advertisement.

She is confident, warm, slightly amused, and genuinely convinced that she has discovered something useful.

Whenever her face is visible, she speaks directly to the camera with accurate natural lip sync.

Stage 2 is the only insert. Her face, head and shoulders are completely outside the frame during this shot, while her dialogue continues as voiceover.

Never use voiceover over a shot where her face is visible.

Never show Margaret silently moving her mouth without synchronized dialogue.

Do not make the performance theatrical, overly enthusiastic, polished, or presenter-like.

### VOICE

A warm, mature woman's voice appropriate for a woman in her late sixties.

Natural conversational delivery with a friendly, slightly humorous tone.

Her voice should feel experienced and confident rather than elderly, frail, or artificially aged.

Moderately quick conversational pace, with natural breaths, small pauses, occasional sentence overlap, and subtle changes in emphasis.

She becomes slightly more animated when explaining the problem and the solution, but never sounds like an announcer.

A short natural laugh occurs during the final stage.

The recording should sound intimate and close, like a front-camera smartphone microphone at arm's length.

No announcer tone, narrator delivery, radio voice, commercial voice-over style, exaggerated enthusiasm, over-articulated consonants, or artificial elderly voice.

### CHARACTER

MARGARET, a charismatic 68-year-old woman with a warm natural complexion, expressive dark brown eyes, naturally defined eyebrows, soft mature facial features, and shoulder-length silver-grey hair with subtle natural waves.

Her hair is neatly styled but not salon-perfect, with a few natural strands slightly out of place.

Her face has realistic mature skin: visible pores, fine lines, natural crow's feet, forehead lines, subtle age-related texture, and natural slight unevenness.

She looks healthy, confident, warm and believable. She does not look artificially young and she does not look heavily made-up.

She wears a simple fitted dark charcoal knit top with three-quarter sleeves, small understated gold earrings, a thin gold necklace, and a simple gold watch on her left wrist. Short natural nails with a subtle nude finish.

Same face, same silver-grey hair, same hairstyle, same clothing, same jewellery and same overall appearance in all six shots.

She is not modelling. She sits and stands naturally at home, behaving like someone casually talking to a friend through her phone.

### SETTING

Exactly three locations/setups are permitted:

1. The modern kitchen at night.
2. The same kitchen in daylight.
3. The front door of the same apartment in daylight.

No other locations.

The kitchen has white flat-panel cabinets, a white quartz island, soft warm under-cabinet lighting, clean counters, and a single matte ceramic vase with stems.

The environment should feel like a real modern apartment rather than a commercial set.

Keep the kitchen clean but slightly lived-in and natural. Avoid making it look sterile or artificially staged.

Stage 2 takes place at the same kitchen island as stages 1, 3 and 4.

Stage 5 takes place in the same kitchen in daylight.

Stage 6 takes place at the front door of the same apartment, with a clean modern hallway and a small strip of the outside street visible behind her.

No location changes beyond these three setups.

### STAGES

#### [0โ€“5s] VISUAL HOOK โ€” SPOKEN TO CAMERA

Medium chest-up, front-facing smartphone camera propped at eye level on the kitchen island.

Margaret fills the upper two-thirds of the frame and looks directly into the lens with a slightly conspiratorial, knowing expression. One eyebrow lifts naturally as she begins speaking.

In the lower third, five matching matte bowls contain the components of one dinner: chicken, rice, broccoli, sauce and oil. A slim digital kitchen scale sits beside them.

The key visual hook is the oil pour.

Margaret's right hand holds a clear glass olive-oil bottle high above the sauce bowl and pours a long, continuous stream of oil into it. The stream remains visible for most of the shot.

The oil is already glistening in the bowl.

Her left arm remains out of frame.

She says naturally to camera:

"You're not overweight because you eat too much. You just have no idea how much you're actually eating."

End state: she gradually tips the bottle upright and stops pouring. The oil remains visibly glistening in the bowl. Her eyes stay locked on the lens.

#### [5โ€“8s] STANDING-SNACK INSERT โ€” VOICEOVER

Clean cut.

Margaret's face is completely out of frame.

Close chest-height smartphone angle looking slightly down at the same quartz kitchen island later that night. Only the warm under-cabinet lighting illuminates the counter and her hands.

An open glass jar of peanut butter sits on the counter with the lid beside it and a knife standing in the jar.

Her left hand lifts a cracker already covered with peanut butter toward the top of the frame, where her mouth would be, and takes a bite.

The cracker leaves the upper part of the frame with slight natural motion blur.

This is one quick visual beat, not a food-preparation sequence.

Voiceover:

"Like be honest, you forget the snacks you had standing at the fridge,"

End state: the bitten cracker moves back toward the jar and the shot cuts.

No face, head or shoulders.

#### [8โ€“14s] THE REFRAME โ€” SPOKEN TO CAMERA

Clean cut back to the same kitchen at night.

Tighter face-and-shoulders close-up from the same propped front camera, moved slightly closer so Margaret's eyes dominate the frame.

The five bowls are now only a soft blur at the very bottom of the image or outside the frame.

Neither hand is visible.

She looks directly down the lens, giving a small dismissive movement of her head while speaking.

She says:

"you're guessing your calories, and by the end of the day you're way off. That was literally me, until I found this. It's called Sona."

The word "Sona" is spoken naturally, without advertising emphasis.

End state: she gives one small knowing nod while maintaining eye contact.

#### [14โ€“21s] PRODUCT DEMO โ€” SPOKEN TO CAMERA

Clean cut.

Medium close-up, pulled approximately half a step farther back than stage 3. Still eye-level and filmed from the same front-facing smartphone.

The five bowls and scale are gone.

A single ordinary plated dinner sits in the lower foreground: chicken, rice and broccoli on a simple white ceramic plate.

The plate remains secondary to Margaret's face.

She taps the edge of the plate once with her left hand while continuing to speak directly into the camera.

She says:

"You just take a picture of your food and it works out the calories and the macros in about four seconds. No typing, no guessing, no effort."

End state: she picks up her fork and begins eating.

Do not show a phone or app interface.

#### [21โ€“26s] DAYLIGHT โ€” SPOKEN TO CAMERA

Clean cut to daylight.

Same kitchen, same character, same clothing and same environment.

Medium chest-up framing with the smartphone propped at eye level on the quartz counter.

Soft window light illuminates her face.

Margaret is eating breakfast from a bowl, holding the bowl in one hand and a spoon in the other. Both hands are readable.

She talks naturally between bites, with a relaxed expression and direct eye contact.

She says:

"Now I actually know what I'm eating, and two months in, it shows."

End state: she places the bowl down in the sink.

#### [26โ€“30s] FRONT DOOR โ€” SPOKEN TO CAMERA

Clean cut.

Medium chest-up shot at the front door of the same apartment.

The smartphone is held or positioned at arm's length at approximately eye level, but remains completely invisible.

The modern hallway and a soft strip of the street are visible behind her.

The open doorway creates slightly overexposed daylight around the background.

Margaret has a tote bag over one shoulder and grips the strap high on her chest with one hand.

She looks directly into the lens and delivers the final line naturally.

She says:

"So if you're serious about getting in shape, stop guessing and start tracking the smart way."

She then breaks into a short, genuine laugh at her own delivery.

End state: she holds eye contact with the camera for approximately half a second, then turns and pulls the door shut behind her.

### BRAND

The product name is spoken once in stage 3.

Brand name: Sona.

Pronunciation: "SOH-nuh".

It should be spoken naturally in the sentence:

"It's called Sona."

Never announce the brand name like a commercial.

Do not generate logos, UI, product labels, captions, or readable text inside the video.

### AUDIO

Use natural location sound and realistic Foley throughout.

Relevant physical sounds include:

* Ceramic bowls touching the quartz island.
* A kitchen scale clicking when an object is placed on it.
* Olive oil continuously pouring into a ceramic bowl.
* The glass bottle being tipped upright.
* A knife lightly touching or resting inside the peanut-butter jar.
* A crisp cracker bite.
* A fork touching the ceramic plate.
* A spoon touching the breakfast bowl.
* Quiet residential ambience at the front door.
* Distant birds and traffic.
* The tote strap shifting against clothing.
* The front door opening and closing.

Margaret's voice should sound close, intimate and unamplified, as captured by a smartphone front-camera microphone.

Night scenes have soft indoor room tone. Daylight scenes are slightly brighter and more open. The doorway has subtle exterior ambience.

[SOUND] Strictly naturally occurring sound and Foley only. No background music.

No subtitles.

### CONSISTENCY

The following elements must remain identical across all six stages:

* Same woman: Margaret, 68.
* Same face and facial features.
* Same mature skin texture.
* Same shoulder-length silver-grey hair.
* Same hairstyle.
* Same dark charcoal knit top.
* Same gold earrings.
* Same thin gold necklace.
* Same gold watch.
* Same natural nude nails.
* Same apartment.
* Same kitchen.
* Same front door.
* Exactly 1 person on screen.
* Exactly 5 bowls in stages 1 and 3.
* Exactly 1 kitchen scale in stages 1 and 3.
* Exactly 1 clear olive-oil bottle in stage 1.
* Exactly 1 peanut-butter jar, 1 knife and 1 cracker in stage 2.
* Exactly 1 plated dinner from stage 4 onward.
* Stage 2 remains a short mid-bite insert.
* Stages 1โ€“4 occur at night.
* Stages 5โ€“6 occur in daylight.
* No phone or screen is ever visible.
* Natural mature skin texture remains consistent across every shot.
* Camera style remains smartphone UGC throughout.
* Lighting changes only where specified by the stages.
* Geography remains consistent between cuts.

Any element not explicitly changed by a stage must remain unchanged.

### CONSTRAINTS

No background music.

No on-screen text.

No captions.

No subtitles.

No logos.

No visible phone.

No visible smartphone screen.

No app interface.

No readable scale display.

No voiceover over any shot where Margaret's face is visible.

No face, head or shoulders in stage 2.

No silent talking-head shot.

No unsynchronized mouth movement.

No second person.

No additional locations.

No fridge.

No open fridge.

No cinematic camera movement.

No slow motion.

No overhead food hero shot.

No low-angle tabletop hero shot.

No wide establishing shot.

No side three-quarter interview framing.

No studio lighting.

No ring light.

No professional commercial camera look.

No cinematic colour grade.

No food-photography styling.

No sterile showroom kitchen.

No cluttered counters.

No dirty dishes.

No dish rack.

No fridge magnets.

No plastic leftover containers.

No catalogue posing.

No finger counting.

No counting gestures.

No exaggerated influencer performance.

No beauty filter.

No airbrushed skin.

No doll skin.

No porcelain skin.

No skin blur.

No glossy plastic skin.

No artificially youthful appearance.

No elderly stereotype.

No frail appearance.

No changing hairstyle.

No changing clothes.

No changing jewellery.

No changing character.

No blonde hair.

No dyed hair.

No second cracker in stage 2.

No peanut-butter spreading sequence.

No long snack-preparation sequence.

No warped hands.

No extra fingers.

No missing fingers.

No objects appearing or disappearing without being specified by the stage.

No changes to the kitchen geography.

Nothing from the storyboard board itself appears in the final video.

Start Creating UGC Videos Today

UGC marketing no longer requires a creator budget. With XYZ Generator, you can plan the environment, storyboard the scenes, and generate consistent, authentic-looking UGC videos in a fraction of the time and cost.

Want to go deeper? Read our AI Generated Videos: The Future of Content Creation, explore the Sora and Veo model guide, or compare every model in our full tools list.

Finished UGC-style AI marketing video created with XYZ Generator
๐Ÿ“ข Related Resources โ€” Learn Prompting ยท Blog ยท Pricing

Tags: UGC, Marketing, AI Video, Content Creation, XYZ Generator

Article loaded