arXiv:2402.18068 Abstract | arXiv Analytics

arXiv:2402.18068 [cs.CV]Abstract References Reviews Resources

SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model

Bin Cao, Jianhao Yuan, Yexin Liu, Jian Li, Shuyang Sun, Jing Liu, Bo Zhao

Published 2024-02-28, updated 2024-03-05Version 2

In the rapidly evolving area of image synthesis, a serious challenge is the presence of complex artifacts that compromise perceptual realism of synthetic images. To alleviate artifacts and improve quality of synthetic images, we fine-tune Vision-Language Model (VLM) as artifact classifier to automatically identify and classify a wide range of artifacts and provide supervision for further optimizing generative models. Specifically, we develop a comprehensive artifact taxonomy and construct a dataset of synthetic images with artifact annotations for fine-tuning VLM, named SynArtifact-1K. The fine-tuned VLM exhibits superior ability of identifying artifacts and outperforms the baseline by 25.66%. To our knowledge, this is the first time such end-to-end artifact classification task and solution have been proposed. Finally, we leverage the output of VLM as feedback to refine the generative model for alleviating artifacts. Visualization results and user study demonstrate that the quality of images synthesized by the refined diffusion model has been obviously improved.

Categories: cs.CV

Keywords: synthetic images, alleviating artifacts, end-to-end artifact classification task, synartifact, compromise perceptual realism

Related articles: Most relevant | Search more

arXiv:1612.07828 [cs.CV] (Published 2016-12-22)

Learning from Simulated and Unsupervised Images through Adversarial Training

Ashish Shrivastava, Tomas Pfister, Oncel Tuzel, Josh Susskind, Wenda Wang, Russ Webb

arXiv:1710.10710 [cs.CV] (Published 2017-10-29)

On Pre-Trained Image Features and Synthetic Images for Deep Learning

Stefan Hinterstoisser, Vincent Lepetit, Paul Wohlhart, Kurt Konolige

arXiv:1712.03904 [cs.CV] (Published 2017-12-11)

Feature Mapping for Learning Fast and Accurate 3D Pose Inference from Synthetic Images

Mahdi Rad, Markus Oberweger, Vincent Lepetit