Copy complete prompts for the orange booth, hotel lobby, recording studio, street cypher, pets and a lyric-free Beatbox experiment. Each example protects left/right identity, keeps one shared microphone and gives the camera a clear job.
Tested and maintained by the StagePair product team · Updated October 8, 2026 · How we review prompts
01
A reliable prompt has four layers.
Identity: assign Image 1 to the left and Image 2 to the right, then ask the model to preserve both separately.
Stage: describe one coherent environment, its light and exactly one shared microphone.
Performance: give each performer a short role and define how they hand off the rhythm.
Camera and exclusions: specify framing, movement and the failures to avoid.
COPY. ADAPT. TEST.
Six copy-ready Hotel Lobby AI prompts.
Replace durations and aspect ratios only with values your chosen model supports. Attach two separate references rather than a group photo. Prompt wording improves direction; it cannot guarantee a perfect face, mouth or motion in every generation.
01
Classic orange-booth duo
The familiar Hotel Lobby AI setup
Create a 10-second vertical 9:16 rap-duo video in a saturated burnt-orange live-session booth with a seamless matte background and warm studio lighting. Image 1 is the LEFT performer and Image 2 is the RIGHT performer. Preserve both identities, hairstyles and facial features without blending or swapping them. Keep both performers together in a waist-up two-shot from the first frame to the last. Place exactly one silver microphone hanging between them, never covering either face. Start the beat and vocals immediately. The left performer delivers a short line while the right performer nods and reacts, then they trade roles with no silent pause. Use restrained head nods, shoulder rhythm and small gestures below the chin. Add one smooth forward-and-back camera move, no cuts, spins, solo shots, text, logos or watermarks. Generate original clean hip-hop audio; do not use an existing song or artist voice.
02
Couple call-and-response
Warm chemistry without forced romance
Create a 10-second vertical 9:16 couple rap performance in an elegant hotel lobby with warm practical lights, polished stone and soft background depth. Use two separate portrait references: Image 1 remains the LEFT performer and Image 2 remains the RIGHT performer. Preserve each face and keep both people visible throughout. Put one shared vintage microphone between them. They trade short original rap lines in a continuous call-and-response: the listener smiles, nods or gives one small answering gesture, then takes the next line. Include two brief moments of natural eye contact, but return both faces toward the camera. Keep torsos mostly front-facing and communicate affection through expression, not kissing, hugging or forced touching. Use a gentle dolly in and out while maintaining a waist-up two-shot. No long pauses, extra microphones, face obstruction, identity swapping, captions, logos or copyrighted music.
03
Recording-booth lively duo
Larger motion with faces kept clear
Create a 10-second landscape 16:9 rap-duo video inside a professional recording booth with acoustic panels, subtle blue accent lights and a neutral key light on both faces. Keep Image 1 on the LEFT and Image 2 on the RIGHT, preserving both identities. Frame both performers from the waist up around one shared hanging microphone. Begin with both faces toward the lens. Use a playful coordinated groove: one low point toward the camera, a sideways forearm sweep with a shoulder bounce, then a small outward torso pivot and a return to camera. The partners answer each accent with a slight delay instead of mirroring exactly. Keep every hand below face level and never cross the microphone or cover a mouth. Generate an original upbeat hip-hop track and continuous alternating vocals from the first beat to the final beat. Use a smooth, visible camera push and pull with no cuts, full turns, back views, extra limbs, captions or watermarks.
04
Daytime street cypher
A grounded outdoor alternative
Create a 10-second vertical 9:16 street-cypher duet on a photorealistic American city sidewalk in daylight. Include weathered red-brick buildings, black fire escapes, a colorful abstract mural with no readable words and realistic street depth. Image 1 is the LEFT performer and Image 2 is the RIGHT performer; preserve both faces and keep their positions stable. Place exactly one shared microphone on a stand between them. Use a chest-up to waist-up two-shot so both performers remain visible. They trade original rap lines continuously, reacting with head nods, shoulder bounces and small hand accents while staying mostly front-facing. Use natural open-shade frontal light so neither face falls into shadow. Add a controlled forward-back camera move with believable background parallax. No crowds crossing the shot, moving cars, handheld microphones, solo cuts, face blending, readable signs, existing music, captions or watermarks.
05
Anatomy-safe pet duo
Keep paws as paws
Create a 10-second vertical 9:16 duo performance using the two animal reference images. Preserve each animal’s species, face, fur markings, ears, muzzle, collar and natural body proportions. The first pet stays on the LEFT and the second stays on the RIGHT. Keep both naturally seated in a head-and-chest two-shot with one shared microphone suspended between them. Their performance uses gentle head bobs, small natural body sways, muzzle movement and one brief low lift of a naturally shaped front paw. Paws must remain compact and fur-covered. Never add human hands, fingers, thumbs, human arms, human skin, glove-like paws, upright human posture, extra limbs or an off-screen handler. Simplify any gesture that the animal cannot perform naturally. Generate an original playful beat with short alternating vocal sounds and continuous rhythm. Keep the camera smooth and both animals visible. No cuts, identity blending, costumes added beyond the references, captions, logos or watermarks.
06
Two-person Beatbox session
StagePair’s lyric-free experiment
Create a 10-second vertical 9:16 Beatbox duo video in a saturated orange live-session booth. Image 1 is the LEFT performer and Image 2 is the RIGHT performer. Preserve both identities and show both performers together from the waist up around exactly one hanging microphone. Both faces begin toward the camera. Generate one original continuous Beatbox groove at approximately 100 BPM using only mouth-made percussion. The LEFT performer supplies deep B/P kick sounds and rounded bass pulses. The RIGHT performer supplies crisp tss/ka/pf hi-hats and snares. Synchronize each person’s mouth and jaw only with that person’s audible part. Use small head bobs, subtle upper-body sway and brief playful eye contact. Trade one short rhythmic fill near the middle, return immediately to the main groove and finish together on a tight final accent. No spoken words, rap lyrics, singing, instrumental backing track, long solo, silent outro, exaggerated cheek distortion, face obstruction, captions, logos or watermarks.
WHY TRY BEATBOX?
A different test for AI performance.
Most Hotel Lobby AI clips trade rap lines over a generated beat. Beatbox asks the model to build the rhythm from two visible mouths: one performer carries kick and bass, while the other supplies hats and snares. That makes the audio role easier to describe and gives the duo a distinct reason to share one microphone.
It is still experimental. Some generations may soften the percussion, add tonal sounds or miss a mouth cue. Clear front-facing portraits and 10–15 seconds give the exchange more room than a five-second trial. Treat the first result as a test, not a promised studio-grade Beatbox recording.
Inside StagePair
Choose Beatbox under Performance mode. StagePair automatically removes rap lyrics and instrumental backing from the request while preserving the selected stage, model and quality.
No. StagePair turns the important prompt decisions into controls for stage, performance mode, occasion, model, quality, duration and format. These copy-ready prompts are for understanding the format or testing another compatible AI video tool.
Why do the prompts assign LEFT and RIGHT?
Position labels give the model a stable identity map. They reduce ambiguity, but no prompt can guarantee that a generative model will never swap or distort a face. Always review the finished clip.
Can a prompt guarantee exact lyrics or choreography?
No. Prompts guide a probabilistic model. Short actions, clear priorities and one continuous shot are usually easier to follow than a long sequence of precise moves.
What is different about the Beatbox prompt?
It replaces lyrics and instrumental backing with mouth-generated percussion, then assigns complementary rhythmic roles to the two performers. It is experimental, so mouth motion and individual sounds can vary between generations.