Program interface and UI style anomaly detection method based on large model

Through the UI style anomaly detection method based on the large model, multiple image comparison special models are used to perform structural similarity screening and majority voting strategies, combined with pixel-level comparison, the problems of low efficiency and poor accuracy of UI style detection in the existing technology are solved, and efficient and intelligent UI style anomaly detection is achieved.

CN120472190AActive Publication Date: 2025-08-12CHENGDU KOALA URAN TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510973679.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-08-12
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

The existing UI style anomaly detection methods are inefficient, expensive, error-prone and difficult to deal with complex scenes and text differences. The traditional image comparison algorithm has limited accuracy, poor adaptability and low intelligence.

Method used

Using a large-model-based program interface and UI style anomaly detection method, multiple image comparison-specific large models are deployed, combined with structural similarity filtering and majority voting strategies, pixel-level comparison is performed to locate wrong pixels, identify and draw error boxes.

Benefits of technology

It significantly improves the accuracy and efficiency of UI style abnormal detection, reduces the false alarm rate and omission rate, can intelligently process complex UI and text content, and adapt to different UI design specifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472190A_ABST
    Figure CN120472190A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligence, and discloses a program interface and UI style anomaly detection method based on a large model, which comprises the following steps: deploying a plurality of large models special for image comparison; acquiring a UI style picture set and a program interface screenshot set, and based on structural similarity comparison, screening out sample pairs of which the similarity is lower than a first preset threshold value to form a first sample pair set; performing image comparison on the sample pairs in the first sample pair set by using a plurality of large models, and forming a second sample pair set by using the sample pairs of which the number is inconsistent and exceeds half of the number of the large models according to a comparison result; and positioning error pixel points of the sample pairs in the second sample pair set through pixel-level comparison, and drawing and reporting an error box. According to the method, a plurality of image comparison special large models are adopted for comprehensive judgment. The large model has strong image understanding and feature extraction capabilities, can more accurately identify the subtle difference between the program interface and the UI pattern diagram, and significantly reduces the false alarm rate and the missing report rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent technology, and in particular to a method for detecting program interface and UI style anomalies based on a large model. Background Art

[0002] In today's rapidly iterating software development landscape, the quality of the user interface (UI) has become a key factor in determining a product's success or failure. A beautiful, consistent, and standard-compliant UI not only enhances the user experience but also strengthens brand image and market competitiveness. To ensure that the actual UI presented by the program aligns with the design specifications provided by the UI designer (usually in the form of UI style sheets), development teams must conduct rigorous UI style anomaly detection.

[0003] Traditional UI style anomaly detection relies heavily on manual labor. Testers or developers need to manually compare program interface screenshots with UI style diagrams, checking for discrepancies in color, layout, fonts, spacing, icons, and other aspects. This approach has obvious drawbacks: 1. Inefficiency: Manual inspection is time-consuming and labor-intensive, especially in large projects with numerous UI interfaces and complex elements. The efficiency of manual inspection is extremely low.

[0004] 2. High cost: A lot of human resources are required for inspection, which increases development costs.

[0005] 3. Error-prone: Manual inspection is easily affected by subjective factors, resulting in omissions or misjudgments, leading to inaccurate test results.

[0006] 4. Difficulty in dealing with complex scenarios: Manual inspection is more difficult for complex UI interfaces, such as those containing large amounts of text, dynamic content, and interactive elements.

[0007] To improve the efficiency and accuracy of UI style anomaly detection, the industry has proposed a number of automated detection methods based on image processing technology. These methods typically use traditional image comparison algorithms (such as pixel difference comparison, edge detection, and template matching) to compare program interface screenshots with UI design drawings. However, these methods also have limitations: 1. Limited accuracy: Traditional image comparison algorithms have high requirements on image quality and are easily affected by factors such as image noise, lighting changes, and alignment deviation, resulting in false positives or negative negatives.

[0008] 2. Poor adaptability: It is difficult to handle complex UI layouts and various UI elements, such as different fonts, icons, rounded corners, shadows, etc.

[0009] 3. Low level of intelligence: Various parameters and thresholds usually need to be set manually, which makes it difficult to adapt to different UI design specifications and project requirements.

[0010] 4. Difficulty in processing text: Simple image comparison cannot effectively distinguish the differences in text content, which is a very important part of UI style anomalies. Summary of the Invention

[0011] To address the aforementioned issues with existing UI style anomaly detection methods, the present invention provides a more accurate, efficient, and intelligent solution, specifically including: A method for detecting program interface and UI style anomalies based on a large model includes the following steps: S1. Deploy several large models dedicated to image comparison; S2. Obtain a UI style atlas and a program interface screenshot set, and based on structural similarity comparison, select sample pairs whose similarity is lower than a first preset threshold to form a first sample pair set; S3. Use the plurality of large models to perform image comparison on the sample pairs in the first sample set, and select the sample pairs whose number of inconsistencies exceeds half of the number of the large models to form the second sample set; S4. Locate the erroneous pixel points of the sample pairs in the second sample pair set by pixel-level comparison, draw an error box and report it.

[0012] Preferably, step S2 further includes a preprocessing step: ensuring that the images in the UI style atlas and the program interface screenshot set have the same pixel size and pixel resolution.

[0013] Preferably, the method for comparing structural similarity comprises: S21. Calculate the brightness, contrast and structural similarity between each UI sample in the UI style atlas and each program interface sample in the program interface screenshot set; S22. Calculate a comprehensive structural similarity score based on the brightness, contrast, and structural similarity.

[0014] Preferably, step S4 includes: S42. Converting the UI style diagram and program interface screenshot in the second sample set into an RGB three-channel image; S43. Compare the R, G, and B channels of the UI style diagram and the program interface screenshot respectively by histogram, calculate the difference in the number of pixel values of each channel, and record the coordinates of the pixels where the number of pixel values differs; S44. Draw an error box based on the different pixel coordinates in the program interface screenshot and report it.

[0015] Preferably, step S4 further includes: S41. Perform text area recognition on the UI style diagrams and program interface screenshots in the second sample pair set, record the coordinates of the recognized text areas, and exclude the text areas in steps S42 and S43.

[0016] Preferably, if the text area in the program interface screenshot is missing text, an error box is drawn in the text area.

[0017] Preferably, the method further comprises: floating the recorded text area coordinates upward and downward by a number of pixels in the vertical direction to obtain extended text area coordinates.

[0018] Beneficial effects 1. High Accuracy: This invention uses a large model dedicated to multiple image comparisons for comprehensive judgment. This large model has powerful image understanding and feature extraction capabilities, enabling it to more accurately identify subtle differences between program interfaces and UI style diagrams, significantly reducing false positives and false negatives.

[0019] 2. High Efficiency: This invention uses an automated detection process, eliminating the need for manual comparison. This significantly improves the efficiency of UI style anomaly detection and shortens the development cycle. This reduces reliance on manual labor, saves costs, and improves quality and user experience.

[0020] 3. Accurate positioning: This invention combines pixel-level comparison technology to accurately locate the specific pixel points where there are differences between the UI style image and the program interface screenshot, providing developers with more accurate and detailed error information for convenient and quick repair.

[0021] 4. Intelligent Text Processing: This invention can identify and eliminate interference from text areas, preventing differences in text content from affecting image comparison results. It can also detect issues such as missing text, further improving the comprehensiveness of detection.

[0022] 5. Good adaptability: The present invention does not rely on a specific image comparison algorithm or fixed threshold setting and can adapt to different UI design specifications and project requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 The figure is a flow chart of a method for detecting anomalies in program interface and UI style based on a large model provided in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described below with reference to the accompanying drawings. In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inner", "outer", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings and are only for the convenience of describing the present invention and simplifying the description. They do not indicate or imply that the devices or components referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, they should not be understood as limiting the present invention.

[0025] like Figure 1 As shown, the present invention provides a method for detecting program interface and UI style anomalies based on a large model, comprising the steps of: S1. Deploy several large models dedicated to image comparison.

[0026] The large-scale model specifically designed for image comparison is a large language model specifically applied to or optimized for tasks such as image similarity judgment and difference recognition. Examples include the Vision Transformer (ViT) series based on the Transformer architecture, and ResNet, EfficientNet, and Swin Transformer based on convolutional neural networks (CNNs). These models inherently possess powerful image feature extraction capabilities, and the extracted feature vectors can be used to calculate similarities or differences between images through methods such as metric learning. They can be deployed on a local server or data center, or using a machine learning platform provided by a public or private cloud. The specific model and deployment method can be selected by those skilled in the art based on actual needs and circumstances, and are not further limited by the present invention.

[0027] S2. Obtain a UI style atlas and a program interface screenshot set, and based on structural similarity comparison, screen out sample pairs whose similarity is lower than a first preset threshold to form a first sample pair set.

[0028] The UI style atlas contains interface images provided by UI designers as a benchmark for design specifications. These images are typically derived from exported files from design tools (such as Sketch, Figma, Adobe XD, etc.), and precisely define the appearance, layout, color, size, and other characteristics of each interface element. The program interface screenshot collection contains screenshots of the actual running application in specific states or pages. These screenshots can be captured manually or, more preferably, automatically captured in predetermined scenarios using automated test scripts to ensure coverage of key user interfaces.

[0029] Structural similarity is an indicator that measures the similarity between two images. Compared to simple pixel differences (such as mean square error (MSE)), it better simulates the human visual system's perception of changes in image structural information. In some preferred embodiments, specific steps are provided, including: S21. Calculate the brightness, contrast and structural similarity between each UI sample in the UI style atlas and each program interface sample in the program interface screenshot set respectively.

[0030] S22. Calculate a comprehensive structural similarity score based on the brightness, contrast, and structural similarity. Preferably, the above three components can be weighted and superimposed to calculate a comprehensive structural similarity score.

[0031] The first preset threshold is an empirical value that can be set by those skilled in the art based on actual needs. For example, it can be set to 0.99, 0.98, or adjusted based on actual application scenarios. The basis for setting it is that sample pairs above this threshold are considered to be highly similar in structure, likely to have no UI style anomalies, or the differences are very subtle and can be ignored; sample pairs below this threshold are considered to have significant visual differences.

[0032] This step is essentially a coarse screening process, effectively eliminating a large number of highly visually consistent image pairs and focusing computing resources on samples more likely to harbor issues. This way, only the first sample pairs initially identified as "suspicious" are passed on to subsequent steps, where they are subjected to further, more in-depth analysis and judgment by the more powerful, but also more time-consuming, large-scale image comparison model. This significantly improves the efficiency of the entire anomaly detection process.

[0033] S3. Use the plurality of large models to perform image comparison on the sample pairs in the first sample pair set, and form a second sample pair set with the sample pairs whose comparison results show that the number of inconsistencies exceeds half of the number of the large models.

[0034] Specifically, each large model receives a pair of images (a UI style image and a corresponding program interface screenshot) as input. Based on its internal complex neural network structure and the knowledge learned from large-scale data, each model independently performs deep feature extraction and comparative analysis on the input image pair. Each model ultimately outputs a judgment result on whether the pair of images is consistent. This result is usually a binary label (for example, "consistent" / "inconsistent"), or a confidence score or difference measure that can be converted into a binary label (for example, by comparing with a model-specific internal threshold). After all large models have completed the comparison of the same sample pair, the system collects the judgment results from each model. Then, the following decision steps are taken: S31. The number of large models that statistically judge the sample pair to be "inconsistent" (or "difference exists", "non-match", or other similar conclusions).

[0035] S32. Compare this number with half of the total number of deployed large models. Specifically, determine whether the number of "inconsistent" conclusions exceeds half of the total number of large models.

[0036] S33. Select the sample pairs that meet the majority voting condition (i.e., are judged as "inconsistent" by more than half of the large models) and form the second sample pair set.

[0037] The technical considerations for setting up multiple large models for majority voting in this invention are as follows: 1. A single model may perform poorly or misjudge certain types of images or differences. By combining the judgment results of multiple models, we can effectively leverage the strengths of different models and complement each other, significantly reducing false positives (misclassifying consistent images as inconsistent) and false negatives (misclassifying inconsistent images as consistent) caused by the bias of a single model.

[0038] 2. The majority voting strategy makes the detection results less sensitive to model selection, slight differences in training data, or certain image noise, and the overall performance is more stable and reliable.

[0039] 3. After the rapid screening in step S2 and the in-depth verification in step S3, the second sample pair set contains image pairs that have been determined with high confidence to have UI style differences. This allows the subsequent step S4 to focus resources on accurately locating these truly problematic areas, avoiding unnecessary pixel-level analysis of a large number of samples with no problems or only minor differences.

[0040] S4. Locate the erroneous pixel points of the sample pairs in the second sample pair set by pixel-level comparison, draw an error box and report it.

[0041] It should be understood that pixel-level comparison refers to the process of directly comparing the attribute values (primarily color) of each corresponding pixel between two precisely aligned (i.e., spatially aligned) images, typically of the same size. Existing techniques often employ threshold-based differential image analysis, which calculates the absolute difference between corresponding pixels in the two images. All pixels with a difference greater than a threshold T are marked as "difference points" (e.g., set to white), while pixels with a difference less than or equal to T are marked as "no difference points" (e.g., set to black), thereby generating a binary difference mask image.

[0042] The shortcomings of this method are that it is extremely sensitive to noise and small changes, and is prone to a large number of false positives. For example, it is very sensitive to small jitters in image rendering, anti-aliasing differences, compression artifacts, and even slight changes in lighting. This means that even if there are no substantial functional or visual errors in the UI, the method may mark a large number of differences, resulting in many interfering false positives. In addition, the method lacks intelligent processing of text content and cannot distinguish whether it is a style error of the UI element itself or simply because of normal changes in dynamic text content (such as user names, time, counters, etc.), resulting in a large number of text differences unrelated to style being falsely reported.

[0043] Therefore, in some preferred embodiments, the present invention provides a preferred method for locating erroneous pixels by pixel-level comparison, which specifically includes: S42. Convert the UI style images and program interface screenshots in the second sample pair set into RGB three-channel images. This step ensures that the UI style images and program interface screenshots are compared in the same color space.

[0044] S43. Perform a histogram comparison on the R, G, and B channels of the UI style image and the program interface screenshot respectively, calculate the difference in the number of pixel values in each channel, and record the coordinates of the pixels where the number of pixel values is different. It should be noted that the histogram comparison refers to dividing the image into grids, calculating the color histogram of each channel in each grid, and comparing the similarity of the histograms between corresponding grids (such as using chi-square distance, Bhattacharyya distance, etc.). Grids with significant differences indicate that there is an abnormal color distribution in the area. Calculate the difference in pixel values of each channel for each pixel in the grid with abnormal distribution. When calculating the difference, a small tolerance threshold can be set, for example, the absolute value of the difference is greater than 5, to ignore extremely small color fluctuations that are difficult for the human eye to detect. Preferably, if a difference in pixel value is detected on at least one color channel (exceeding the tolerance), the coordinates (x, y) of the pixel are recorded as an erroneous pixel.

[0045] S44. Draw an error box based on the different pixel coordinates in the program interface screenshot and report it.

[0046] Those skilled in the art will know that after obtaining the coordinates of the pixel points with differences, these discrete points need to be organized into meaningful error areas and visualized. Specifically, algorithms such as connected component analysis in image processing can be used to combine spatially adjacent error pixels into one or more continuous error areas. For each identified error area (or directly based on the minimum enclosing rectangle of all error pixels), a bounding box that can enclose the area is calculated and drawn on the program interface screenshot, usually using a striking color (such as red). This error box intuitively marks the specific location and range of the UI style inconsistency.

[0047] Finally, a screenshot of the program interface containing the error box annotation is reported, along with any additional information (such as the error type, error area coordinates, corresponding UI style diagram, detected difference metrics, etc.). This reporting can be in the form of generating a test report with both pictures and text, or integrating the results into an automated testing platform or continuous integration / continuous deployment process, which is not further limited in this invention.

[0048] Before performing pixel-level comparison, in order to avoid misreporting normal text content differences (such as the difference between the sample icon name and the final version name) as UI style errors, it is necessary to first identify and process the text areas in the image. Specifically, step S4 also includes: S41. Identify text regions on the UI style images and program interface screenshots in the second sample set, record the identified text region coordinates, and exclude these text regions from steps S42 and S43. Text region identification can be achieved using optical character recognition (OCR) technology or a deep learning-based text detection model. Based on the text region coordinates, a binary mask is generated for each image. In this mask, pixels corresponding to the text region are marked as "ignored" or "excluded," while pixels in other regions are marked as "included in comparison."

[0049] In some preferred embodiments, in order to further reduce the interference that may be caused by anti-aliasing of text edges or slight rendering differences, the identified text area coordinates can be expanded. Specifically, the recorded text area coordinates are floated up and down by a number of pixels in the vertical direction (for example, expanded by 1-3 pixels) to obtain the expanded text area coordinates. Preferably, a similar expansion can also be performed in the horizontal direction. Doing so can create an "exclusion area" that is slightly larger than the actual text content, ensuring that subtle changes in the text itself and its edges can be more thoroughly ignored during pixel comparison.

[0050] When identifying text areas, if a text area that clearly exists in the UI style diagram is found to be missing text content in the corresponding program interface screenshot (for example, OCR did not recognize any text in the area, or the recognition result was blank), this usually represents a significant UI defect. In this case, even if the area is excluded in the pixel comparison, it should be considered a specific error type. Specifically, an error box needs to be drawn on the program interface screenshot around the area where text is expected but is actually missing, and it may be accompanied by a specific error label (such as "text missing").

[0051] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for detecting program interface and UI style anomalies based on a large model, characterized in that: Including steps: S1. Deploy several large models dedicated to image comparison; S2. Obtain a UI style atlas and a program interface screenshot set, and based on structural similarity comparison, select sample pairs whose similarity is lower than a first preset threshold to form a first sample pair set; S3. Use the plurality of large models to perform image comparison on the sample pairs in the first sample set, and select the sample pairs whose number of inconsistencies exceeds half of the number of the large models to form the second sample set; S4. Locate the erroneous pixel points of the sample pairs in the second sample pair set by pixel-level comparison, draw an error box and report it.

2. The method for detecting program interface and UI style anomalies based on a large model according to claim 1, wherein: Step S2 also includes a pre-processing step: ensuring that the images in the UI style atlas and the program interface screenshot set have the same pixel size and pixel resolution.

3. The method for detecting program interface and UI style anomalies based on a large model according to claim 1, wherein: The method for comparing structural similarity includes: S21. Calculate the brightness, contrast and structural similarity between each UI sample in the UI style atlas and each program interface sample in the program interface screenshot set; S22. Calculate a comprehensive structural similarity score based on the brightness, contrast, and structural similarity.

4. The method for detecting program interface and UI style anomalies based on a large model according to claim 1, wherein: Step S4 includes: S42. Converting the UI style diagram and program interface screenshot in the second sample set into an RGB three-channel image; S43. Compare the R, G, and B channels of the UI style diagram and the program interface screenshot respectively by histogram, calculate the difference in the number of pixel values of each channel, and record the coordinates of the pixels where the number of pixel values differs; S44. Draw an error box based on the different pixel coordinates in the program interface screenshot and report it.

5. The method for detecting program interface and UI style anomalies based on a large model according to claim 4, wherein: Step S4 further includes: S41. Perform text area recognition on the UI style diagrams and program interface screenshots in the second sample pair set, record the coordinates of the recognized text areas, and exclude the text areas in steps S42 and S43.

6. The method for detecting program interface and UI style anomalies based on a large model according to claim 5, wherein: If text is missing in the text area in the program interface screenshot, an error box is drawn in the text area.

7. The method for detecting program interface and UI style anomalies based on a large model according to claim 5, wherein: Also includes: The recorded text area coordinates are floated vertically upward and downward by a number of pixels to obtain extended text area coordinates.

Citation Information

Patent Citations

  • Image difference detection and model training method, device and program product

    CN113239928A

  • UI picture testing method and device and electronic equipment

    CN114387460A

  • Image comparison method and related device

    CN116883698A

  • Image similarity processing method and device, electronic equipment and storage medium

    CN118447269A

  • Visual restoration detection method and device

    CN119847513A