A Review of Aesthetic Assessment, Preference Alignment, and Explanation in Text-to-Image Generation
DOI:
https://doi.org/10.63593/AS.2709-9830.2026.06.005Keywords:
text-to-image generation, computational aesthetics, aesthetic assessment, preference alignment, explainabilityAbstract
This study presents a bibliometric analysis complemented by a narrative review to systematically examine generative computational aesthetics in text-to-image (T2I) generation from the perspective of a closed-loop optimization framework. Specifically, it investigates the types and functions of aesthetic signals and their interfaces with generative model optimization. Aesthetic signals are categorized according to their output bandwidth into scalar scores, rating distributions, multidimensional attribute labels, natural-language critiques, and saliency maps/spatial attributions. Their respective roles in post hoc reranking, reward modeling and post-training alignment, multi-objective optimization, inference-time control and local editing, as well as personalization and interactive interfaces are further analyzed. Building upon this framework, the study examines the effectiveness of explanation mechanisms within the closed loop by evaluating existing approaches from four complementary dimensions: localizability, verifiability, executability, and causal reliability. The review indicates that although substantial progress has been achieved in individual components of aesthetic assessment, preference alignment, and explainability, a fully integrated aesthetics–alignment–explanation closed-loop paradigm remains largely underexplored. This review therefore provides a unified perspective for understanding and advancing closed-loop optimization in T2I generative systems.
