A photo to generative AI workflow for rapid 3D heritage representations
Tarih
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Erişim Hakkı
Özet
This paper investigates generative AI workflows that reconstruct 3D architectural heritage models from a single photograph. We compare two pipelines that combine a 2D image model with Tencent’s Hunyuan3D 2.1 image-to-3D system: (1) Gemini Flash 2.5 + Hunyuan3D and (2) Qwen-Image-Edit + Hunyuan3D. Using a dataset of 16 Turkish architectural landmarks photographed from sub-optimal viewpoints, each method first generates an isometric or gently re-angled view and then produces a PBR-textured GLB mesh approximating LOD3 building detail. The pipelines are evaluated in terms of visual fidelity of intermediate images, geometric completeness and sharpness of the 3D meshes, dimensional consistency, processing time, user experience, and downstream compatibility with BIM, GIS, VR and web viewers. Results show that both workflows deliver photorealistic, lightweight 3D assets within minutes, dramatically lowering the cost and expertise barrier compared with conventional photogrammetry or laser scanning. The Qwen-based pipeline better preserves original textures and colors, enriches side-facade information, and offers higher reproducibility thanks to its open-source model and public interfaces. The Gemini-based pipeline provides cleaner, stylized views but is constrained by closed access. We conclude by positioning single-image generative AI as a rapid, complementary tool for cultural heritage visualization, education and preliminary digital twin creation rather than a substitute for metric survey.












