MaLiang-Harness: A Programmable Pathto Image and Video Generation

Haoyu Zhao1,2Zihao Zhang2Xudong Wang1Jiaxi Gu3Zuxuan Wu2Shuicheng Yan1

1 National University of Singapore2 Fudan University3 Tencent

Project PageGitHubarXivDaily Papers Coming soon

Video generation

Abstract

MaLiang-Harness organizes MLLM-driven image and video generation as a persistent process of program construction, visual inspection, and revision. It addresses the gap between a program that executes and a visual result that follows the request. Programs, edits, and rendered evidence share a common revision reference.

Hover over an image to inspect local detail. Click to open the full image. On touch screens, tap to enlarge.

Illustrations

Scenes, portraits, and character studies.

Posters and Infographics

Layout, typography, and information-rich compositions. Research figures shown here are illustrative examples, not reported benchmark results.

Reference-based and Editing Examples

Selected outputs from reference-based and revision tasks. Original prompts are available for each result.

Painting and 3D Rendering

Studies of form, materials, and light.

Offline Renderer Demos

Scripted fixtures for renderer testing. These are separate from the model-generated examples above.

Method

Generation proceeds through program construction, visual inspection, and revision, with a shared reference connecting executable state to rendered evidence.

PEG
Persistent executable state.
TGP
Traceable generation process.
REV
Revision-aware editing and verification.

Citation

If you use MaLiang-Harness in your research, please cite our paper.

@misc{zhao2026maliangharness,
  title={MaLiang-Harness: A Programmable Path to Image and Video Generation},
  author={Haoyu Zhao and Zihao Zhang and Xudong Wang and Jiaxi Gu and Zuxuan Wu and Yu-Gang Jiang and Shuicheng Yan},
  year={2026},
  eprint={2609.34309},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2609.34309}
}