How Reliable Are LLM Evaluations of Microservice Architecture Diagrams?
Architecture diagrams play a central role in communicating design intent in microservice systems, but their quality evaluation remains largely manual, subjective, and difficult to scale. This paper investigates the reliability of Large Language Models (LLMs) as automated reviewers of microservice architecture diagrams. We design a guideline-based evaluation framework grounded in five established quality criteria (clarity, consistency, completeness, accuracy, and level of detail) and apply it to assess diagrams using two state-of-the-art LLMs across both visual (PNG) and textual (SVG/XML) representations. To evaluate reliability, we conducted a human audit with 16 software engineers, which yielded human reference scores for 360 guideline-level LLM evaluations and 253 pairwise comparisons of LLM-generated reasoning. Our results show that LLM-generated scores achieve reasonable agreement with human reference scores, with an overall F1-score of 0.761 and fair ordinal agreement (kappa = 0.308). Agreement varies more clearly by quality criterion than by representation, with completeness showing the strongest alignment and clarity the weakest. Although image-based inputs consistently outperform textual inputs descriptively, the representation effect is not statistically robust for score agreement. In contrast, experts significantly prefer image-based reasoning in pairwise comparisons, selecting it in 60.8% of decided cases. These findings suggest that LLMs can effectively support early-stage architectural reviews by identifying potential architecture issues and prioritizing them for human inspection, but their evaluations still require expert validation.
2026/1 - POC2
Orientador: Dr. Eduardo Figueiredo
Palavras-chave: Arquitetura de software, Microserviços, LLM
PDF Disponível