Preprints
* indicates equal contribution. Some papers are highlighted.
|
|
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Technical Report, 2026
We introduce an agentic framework that discovers executable world representations for physical reasoning. Structured code models entities, states, events, and relations, allowing agents to simulate physical dynamics and verify their reasoning.
|
|
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
Technical Report, 2026
We introduce an agent-based evaluation framework for visual world models. Specialized agents solve measurable subproblems and organize their findings into an interpretable evidence tree, producing evaluations that closely align with human preferences.
|
|
Physics3D: Learning Physical Properties of 3D Gaussians via Video Diffusion
arXiv preprint, 2024
In this paper, we propose Physics3D, a novel method for learning various physical properties of 3D objects through a video diffusion model. Our approach involves designing a highly generalizable physical simulation system based on a viscoelastic material model, which enables us to simulate a wide range of materials with high-fidelity capabilities.
|
|
High-Fidelity Implicit Text-to-Shape Generation with Voxelized Diffusion
International Journal of Computer Vision (IJCV), 2026
We introduce Diffusion-SDF for text-conditioned 3D shape generation using voxelized signed distance fields. Diffusion-SDF++ further improves geometric detail through octree-based diffusion that adaptively allocates higher-resolution voxels to complex local structures.
|
|
Learning Generalizable Semantic Radiance Fields with Cross-Reprojection Attention
International Journal of Computer Vision (IJCV), 2026
We introduce S-Ray and S-Frustum for generalizable semantic radiance fields. Cross-reprojection attention aggregates multi-view semantic information, while frustum-based rendering improves prediction consistency through efficient attention over a scene tensor.
|
|
CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026 (Highlight)
We propose CFG-Ctrl, a unified framework that reinterprets CFG as a control applied to the first-order continuous-time generative flow. Based on this control-theoretic view, we further introduce SMC-CFG, which uses sliding mode control to drive the flow toward a rapidly convergent sliding manifold.
|
|
ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
IEEE Transactions on Image Processing (TIP), 2025
In this paper, we propose ReconX, a novel 3D scene reconstruction paradigm that reframes the ambiguous reconstruction challenge as a temporal generation task. The key insight is to unleash the strong generative prior of large pre-trained video diffusion models for sparse-view reconstruction.
|
|
DreamReward-X: Boosting High-Quality 3D Generation with Human Preference Alignment
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025
We present a comprehensive framework, coined DreamReward-X, where we introduce a reward-aware noise sampling strategy to unleash text-driven diversity during the generation process while ensuring human preference alignment. Our results demonstrate the great potential for learning from human feedback to improve 3D generation.
|
|
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
IEEE International Conference on Computer Vision (ICCV), 2025
In this paper, we introduce a novel generative framework, coined LangScene-X, to unify and generate 3D consistent multi-modality information for reconstruction and understanding.
|
|
Video-T1: Test-Time Scaling for Video Generation
IEEE International Conference on Computer Vision (ICCV), 2025
We present the generative effects and performance improvements of video generation under test-time scaling (TTS) settings. The videos generated with TTS are of higher quality and more consistent with the prompt than those generated without TTS.
|
|
VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025 (Highlight)
In this paper, we propose VideoScene to distill the video diffusion model to generate 3D scenes in one step, aiming to build an efficient and effective tool to bridge the gap from video to 3D.
|
|
Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image
Conference on Neural Information Processing Systems (NeurIPS), 2024
In this work, we introduce Unique3D, a novel image-to-3D framework for efficiently generating high-quality 3D meshes from single-view images, featuring state-of-the-art generation fidelity and strong generalizability. Unique3D can generate a high-fidelity textured mesh from a single orthogonal RGB image of any object in under 30 seconds.
|
|
Make-Your-3D: Fast and Consistent Subject-Driven 3D Content Generation
Fangfu Liu,
Hanyang Wang,
Weiliang Chen,
Haowen Sun,
Yueqi Duan
European Conference on Computer Vision (ECCV), 2024
We introduce a novel 3D customization method, dubbed Make-Your-3D that can personalize high-fidelity and consistent 3D content from only a single image of a subject with text description within 5
minutes.
|
Honors and Awards
- Outstanding Graduate, Department of Computer Science and Technology, Tsinghua University, 2025
- Tsinghua Excellent Academic Scholarship, Tsinghua University, 2022
|
|