How we built a scalable travel video engine | Headout Hackin'
What if every iconic place could have a rich, cinematic story built around it? Moonshot uses AI to turn raw footage into high-quality video at scale, making immersive storytelling faster to produce and easier to bring to more guests.
Travel demands video because a photograph can show you what a place looks like, but it cannot capture the feeling of walking through it or the details that make someone stop scrolling and imagine themselves there.
At Headout, we have thousands of experiences that deserve to be presented that way, but producing video at that scale has always been difficult. A single 60-second video can take five to ten hours once you account for the brief, footage, scripting, editing, music, voiceover and review. Multiply that across thousands of experiences and the problem becomes obvious: we are in the age of video, without a scalable way to make it.
Project Moonshot began as an attempt to change that.
It started with a YouTube channel
The team had been watching Chloe in History, where a fictional character films herself dropping into famous moments across time, and asked us a simple question: could we make something like this?
We said yes before we had a plan. At first, the brief was to create one channel, but within a day we realised the more interesting problem was not how to make one clever series of videos, but how to build the system that could make a thousand of them.
That became Moonshot, built by Arvind, Aman and Anish.
What Moonshot does
Moonshot turns a place into a finished short-form video with a single click.
You choose the destination, a duration between 30 and 90 seconds, and a theme, while the orchestration layer coordinates the models, prompts, shots, voiceover and music required to produce the final video.
Three formats work today: a Tour Guide walkthrough, a Time Travel story, and a Cinematic cut built around atmosphere and suspense.
The interface is simple, but the work underneath is not, because generating a video is easy now, while generating one that feels coherent and worth watching remains difficult.
The first night, we did not build the app
We spent the entire first night fighting with prompts. Our earliest outputs looked exactly like the AI slop people are tired of seeing. Faces changed between scenes, framing felt accidental, architecture reinvented itself from one shot to the next, and the soundtrack seemed to belong elsewhere.
Each generation looked impressive on its own, but the finished result still felt wrong, because the models could create striking fragments without understanding how those fragments belonged together. We kept prototyping on Higgsfield until we produced something that made everyone stop. It was not perfect, and the design still whispered “AI”, but the remaining gap finally felt solvable through craft. Only then did we connect the APIs and build the application.
The order mattered: understand the craft first, then automate it.
The real insight was not about the model
Every week, another AI video model launches and the internet moves on to debating which one is now the best. That race is exciting, but it is not a durable product strategy, because if your advantage depends entirely on access to the best model, it may disappear when a better one launches.
The insight was simpler: we may not be able to control the models, but we can control how we use them. We can control how they are prompted, how shots are sequenced, how continuity is maintained, and how several models are combined into one repeatable workflow. That orchestration layer is where we believe the more durable value sits.
The demo and what comes next
For the final demo, we skipped the slides and played a video of Arvind travelling through history and throwing himself into adventure sports despite never filming any of it, followed by a live walkthrough.
Using him as the test subject made the demo more honest, because the question was no longer whether AI could generate an impressive travel video, but whether it could create something convincing featuring someone everyone could recognise.
We are now adding human review gates and deeper controls, because “one click” should not mean handing every decision to the machine; it should mean removing repetitive work while preserving the ability to reject a shot that does not feel right.
The first real test will put Moonshot-generated videos in front of guests after they book, before the same engine expands into ad creative, organic social content and travel vlogs. When the next great model launches, we do not rebuild; we plug it into the orchestration layer, and the entire catalogue gets better.
We started with a hallway question about recreating a YouTube channel and ended the weekend with a working product, a category win and a roadmap for how Headout could create video at scale. The model will keep changing but the system that knows how to use it is what endures.
Hackin’ at Headout is our internal hackathon, where small teams have less than two days to turn real business problems into working products. Interested in building with us? Explore open roles at Headout.