Tools for Generative AI
Last year, a consortium consisting of YLE, ITV, Pluxbox, Respeecher, Somersault and Xansr Media, led by RAI and EBU, launched an innovative project that integrated generative AI into the production pipeline. Through rigorous benchmarking, the team identified the most suitable tools for their requirements. They went further by developing a comprehensive workflow. This initiative is not merely about generating text and instantly producing a movie or TV series. The critical factor is control: achieving the desired output necessitates a meticulously structured workflow.
Multiple players and technologies in the market are being announced across territories, and not all of them at time of writing have fully fleshed out models with capabilities that are really useful for a broadcaster or a media company. The challenge of creating a concept from first principles, incorporating new techniques, and reinventing the way that content is produced is an appealing challenge that applies to the entire industry, and there is a need for a well-structured workflow.
Pre-production: Script and Storyline
The team used a combination of Gemini and chatGPT to create the storyline and script with dialogues for the first episode of ‘Tour the World in 80 Days’, revised in a futuristic take.
The first takeaway was that using multiple agents worked best. This approach allowed them to leverage the strengths of each agent, resulting in a more efficient and effective process. However, the team found that relying on fully fledged text-to-script conversion was restrictive. It limited flexibility and adaptability, making it challenging to address dynamic scenarios. To overcome this, a human drive to the re-prompting process was needed. This human intervention ensured that the prompts were contextually relevant and tailored to the specific needs of the situation, enhancing the overall performance and outcome of the project.
The team developed the following storyline:
- It’s 2255 and Phileas Fogg and his companion, Passp2, are two robots who are left in a junkyard.
- Phileas is running out of batteries: only 80 days left.
- They find a hidden portal capable of teleporting them in other places and decide to embark on a journey through it searching for batteries.
Next step, the team addressed the visual storyboard. Auto generated storyboard tools were hard to control and quite restrictive to generate images that match the scene.

Maintaining a consistent style and look for the characters was crucial. The team paid close attention to camera angle positioning to ensure each shot was visually appealing and coherent. However, they encountered some trouble with the side view, which required additional adjustments. To address this, the team had to draw manually some missing images. Moreover, they also found that text on images didn't work as they had hoped. In general, it took many attempts to get the right one, but the effort was worth it to achieve the desired outcome.
Pre-production: Character Design

The character design workflow began with a graphic artist developing original concept art manually in Adobe Photoshop. Subsequently, Adobe Firefly was utilised to generate a stylised interpretation of the concept. The artist then leveraged 3D AIstudio to create a three-dimensional model, which enabled precise character posing across multiple compositions. However, the initial 3D model required professional refinement before it could be suitable for animation.
A specialised 3D artist performed detail modifications in order to correctly apply the skeleton to the 3D model, particularly focusing on improving the hands, the lower arm regions and the head. The topology made by AI was maintained in order to test its behaviour, but some textures were corrected using a 3D texturing software.
Pre-production: Concept Art of the Environments
The team utilised Stable Diffusion to develop the mood board and the concepts art of the environment. Additionally, they employed Photoshop's ‘generative filling’ feature to seamlessly enhance the picture, ensuring a cohesive design.
Pre-production: Style Image

During this critical integration phase, the team consolidated all previously developed visual assets to create the definitive Style Image. The process utilized Blender software for precise character positioning, allowing for accurate poses of both the protagonist and companion character. The final composition, the Style Image, was achieved by incorporating the rendered character models into the environmental backdrop through Photoshop.
Production: Final Video
Starting from the style image, the team used the advanced image-to-video capabilities of Kling.AI and Runway to seamlessly transform static images into dynamic video content. This innovative approach not only overcame the many problems associated with text-to-video, but also significantly reduced the time and effort required for video production.
Key Findings
These bullet items summarize the challenges and learnings encountered during the creation of a video project using AI tools for various stages, from text generation to sound design.
- Text Generation: While AI can generate scripts, a purely automated approach is less effective than an iterative process that allows for control and quality refinement.
- Image Generation: Storyboarding and concept art are possible with AI, reducing the need for dedicated artists. However, consistency and specific elements like camera angles pose challenges.
- Video Generation: Image-to-video yielded better results than text-to-video. Maintaining consistency across shots and addressing animation issues were significant hurdles.
- Sound Generation: Speech-to-speech proved more effective than text-to-speech for capturing emotion and expression, highlighting the limitations of current text-to-speech technology for nuanced vocal performance.
Overall, the project demonstrates the potential of AI tools in video production but also highlights the importance of human intervention for quality control, consistency, and creative direction. The need for iterative refinement and addressing consistency issues across various AI-generated elements emerges as a central theme. Furthermore, the choice of specific AI techniques (e.g., image-to-video vs. text-to-video, speech-to-speech vs. text-to-speech) significantly impacts the outcome and requires careful consideration.
Final Considerations
The project achieved several tangible results, demonstrating the versatility and potential of the created media production workflow. A script, concepts art, and storyboard were developed, laying the groundwork for the pilot episode. Using image-to-video techniques, the team animated one minute of the pilot and created a one-minute title sequence to demonstrate the feasibility of the workflow.
Feedback from industry professionals, including experts at RAI, validated the project's approach. While the pre-production elements - scriptwriting and storyboarding - were praised for their efficiency and quality, the video output was not considered ready for public broadcast. This constructive criticism highlighted both the promise and current limitations of Generative AI in media production.