Text-driven 3D scene generation is promising for digital content creation, embodied AI simulation, and interactive design, yet practical workflows require refining existing scenes while preserving non-target content. MUSE formulates controllable 3D scene authoring as incremental requirement satisfaction and introduces a memory-grounded multi-agent framework with an Architect, Sculptor, and Inspector. It also introduces AuthorBench for requirement-level controllability and preservation-aware editing.
@article{xu2026muse,title={MUSE: Agentic 3D Scene Authoring via Memory-Grounded Incremental Requirement Satisfaction},author={Xu, Ruijie and Zhu, Xinnan and Ying, Jiayu and Dong, Daoguo and Ji, Yuzhou and Tan, Xin},journal={arXiv preprint arXiv:2606.14168},year={2026},}
JointEdit3D performs feed-forward 3D scene editing in a unified RGB-geometry latent space. It uses asymmetric latent inpainting and a SceneAnchor Branch to propagate an edit from one reference frame while preserving source-scene structure. The work also introduces the SceneEdit3D-15K paired dataset and the SceneEdit3D-Bench evaluation benchmark.
@article{zhu2026jointedit3d,title={JointEdit3D: Feed-Forward 3D Scene Editing in a Unified Latent Space},author={Zhu, Xinnan and Xu, Ruijie and Ying, Jiayu and Dong, Daoguo and Xu, Jiachen and Xie, Yuan and Tan, Xin},journal={arXiv preprint arXiv:2606.13345},year={2026},}
TSHA is a benchmark for trustworthy indoor safety hazard assessment with vision-language models. Its current release contains 66,668 validated question-answer pairs across images, videos, and panoramas, covering 12 parent hazard groups and 64 fine-grained hazard types. Evaluation of 22 popular VLMs shows that current models still lack robust safety hazard assessment capabilities.
@article{yu2026tsha,title={TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios},author={Yu, Qiucheng and Xu, Ruijie and Chen, Mingang and Dong, Jianfeng and Tan, Xin},journal={arXiv preprint arXiv:2603.29759},year={2026},}
OSH-Splat constructs a compact 3D language field for accurate and efficient open-vocabulary queries. It extracts hierarchical semantic information at part, subpart, and whole levels, learns compact semantic features on 3D Gaussians, and optimizes a semantic hyperplane for each text query to improve localization accuracy and robustness.
@article{xu2025oshsplat,title={OSH-Splat: Optimizable Semantic Hyperplanes for Enhanced 3D Language Feature Gaussian Splatting},author={Xu, Ruijie and Ji, Yuzhou and Tan, Xin and Ma, Lizhuang},journal={The Visual Computer},volume={41},number={13},pages={11127--11137},year={2025},doi={10.1007/s00371-025-04091-5},}