Discussions
Aligning Genny narration timing with AI-generated product image sequences
I am building short product explainers where each still scene is created first with https://realisticaiimagegenerator.online/ and then narrated through the Genny API. Because the visual sequence is assembled before synthesis, each scene already has a fixed duration (usually 3–6 seconds).
Is there a recommended way to estimate the spoken duration of a text segment before generating the final audio? I would also like the voice, pacing, and prosody to remain consistent across a sequence of scene-level requests.
Would it be more reliable to send the entire narration in one synthesis request and place explicit breaks between scenes, or to submit one request per scene and join the returned audio? If separate requests are preferable, are there API parameters or metadata that help preserve continuity and make timing predictable?
![Genny API [PROD]](https://files.readme.io/89a130e-small-lovo_logo_blue.png)