Tuesday, September 1, 2026

Creating Video Templates for rogerbots.com

So the current state of video generation is something like this, 

1. Render everything in HTML, CSS and then render a video out of it, (Hyperframes)
2. Render it in Canvas, and then render HTML into it. (Remotion)

Both of them are great as they provide a vector programming layer behind video generation, and as we all know agents are good at coding, making them good video generation in this method. 

It has massive limitations unfortunately, where the video being used cannot be more directly impacted, where we can remove background, mask and track objects, and replace silhoutte of mask with real footage without losing neighboring noise associated with new replacement / old mask. None of the features which AI video editing is exceptionally good at. Somewhere there needs to exist a harness which combines the programming reliabilty of hyperframes/remotion, imagination of Video Generation models, mixed with real footage of people with right expression, thus giving everyone the power to create more enhanced version of videos with their real footage thats' properly edited without the loss, thus driving better quality videos overall. Let AI do what it's good at - Imagination, and let reliability be handled by discrete math function tested over the years.  

Now, nobody recognises AI with better quality and I agree, AI is good at imagination. 

So, anything that controls the video should be largely part of programming layer on which the video actually rests, 

Zoom Ins -> Programming Layer 

Zoom Outs -> Programming Layer

Camera Motion and Transition where New information is required-> AI Footage 

Lighting -> Programming Layer

Person's face and Eyes -> Real Footage 

Clothes / Low Attention Zones (Background) -> AI Footage

Capturing the Mask -> SAM 3.1 Model (Segmentation) 

Filling the gaps post mask merge -> AI Footage 

Extended Real Footage Effects -> AI Extension for a real Clip

Creating different angle for the same person talking without loss of expression -> Again this is hard but possible, so you use the same reference footage and then turn it into another clip.  

I am little confused with regards to motion capture, text motion capture often requires certain level of layering (for the text in/out effect in the video) and is best served with SAM3.1 over video with templates but the part that's an issue is that advanced text motion graphics are still handled better with models like Minimax H3 Max, 


Video generation and video editing can be handled by AI extremely well when right references are provided. The idea where we can store the periphery, like the background, camera motion and lighting information in a markdown file keeping the person's expression intact is an underexplored approach AI currently, especially with something as simple as launch video. 


My todo list, 
1. Create the video with real footage of you talking, extend it to imaginary hands motion, and dummy clothing without changing my face at all.