inproceedings

Rate My Web Design? Using LLMs to Predict Human Subjective Ratings of Generated Web UI Prototypes

Bibliography Reference

Format:
Bakaev, M., & Grigera, J. (2026). Rate My Web Design? Using LLMs to Predict Human Subjective Ratings of Generated Web UI Prototypes. Proceedings of the 2026 IEEE/ACM 48th International Conference on Software Engineering, 456–458. https://doi.org/10.1145/3774748.3795676

Publication Abstract

GenAI services are being increasingly employed in UI/UX engineering, whose automation remains insufficiently advanced compared to other Software Engineering activities. The iterative process involves creating and assessing UI prototypes, but up-to-date GUI Agents lack in evaluating subjective quality dimensions like aesthetics. We had ChatGPT and RateMyImage (RMI) rate 84 web AI-generated UI prototypes, and then benchmarked their ratings with human ones, provided by 39 HCI experts. RMI rated the most similarly to humans (r = 0.546), while GPT-4o showed poor agreement with human participants and poor internal consistency. The best (AIC-optimal) regression model for humans’ subjective ratings that incorporated both pre-LLM UI features and models’ ratings had R2=0.423. We also provide some preliminary guidelines for UI/UX engineers seeking to involve AI in assessing design prototypes.

BibTeX Source Entry

@inproceedings{Bakaev_2026,
  doi = {10.1145/3774748.3795676},
  year = {2026},
  month = {Apr},
  pages = {456–458},
  title = {Rate My Web Design? Using LLMs to Predict Human Subjective Ratings of Generated Web UI Prototypes},
  author = {Bakaev, Maxim and Grigera, Julián},
  series = {ICSE-Companion ’26},
  abstract = {GenAI services are being increasingly employed in UI/UX engineering, whose automation remains insufficiently advanced compared to other Software Engineering activities. The iterative process involves creating and assessing UI prototypes, but up-to-date GUI Agents lack in evaluating subjective quality dimensions like aesthetics. We had ChatGPT and RateMyImage (RMI) rate 84 web AI-generated UI prototypes, and then benchmarked their ratings with human ones, provided by 39 HCI experts. RMI rated the most similarly to humans (r = 0.546), while GPT-4o showed poor agreement with human participants and poor internal consistency. The best (AIC-optimal) regression model for humans’ subjective ratings that incorporated both pre-LLM UI features and models’ ratings had R2=0.423. We also provide some preliminary guidelines for UI/UX engineers seeking to involve AI in assessing design prototypes.},
  booktitle = {Proceedings of the 2026 IEEE/ACM 48th International Conference on Software Engineering},
  publisher = {ACM},
}