Paper Readings
Papers I presented at lab reading groups based on my research interests. Summaries and analysis reflect my own interpretations, not those of the authors. All figures belong to the cited papers.
Vision-Language Models (VLMs)Layout Generation
Reviewed Jul. 2026
VGGT: Visual Geometry Grounded Transformer
3D ReconstructionFeed-ForwardMulti-Task Model
Reviewed Mar. 2026
GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
Video Diffusion3D ConsistencyCamera Control
Reviewed Feb. 2026
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
Video GenerationLDM
Reviewed Nov. 2025
MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion
Text-to-3D Scene GenerationMulti-View Consistency
Reviewed Sep. 2025
SceneCraft: Layout-Guided 3D Scene Generation
Text-to-3D Scene GenerationSpatial Guidance
Reviewed Jul. 2025
Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models
Text-to-3D Scene GenerationFoundation Model
Reviewed Apr. 2025
Last updated August 2026