Functional interface compatibility problem auxiliary detection method and system
By using YOLO model and feature area recognition technology, the missed detection and misdetection problems in functional interface detection are solved, and accurate detection is achieved in different testing environments, especially the identification of imperceptible interface defects.
Patent Information
- Application Number
- CN202510365900.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is prone to missed detection and missed detection when detecting functional interface compatibility problems, and the accuracy of the detection results is low, especially in different testing environments, misjudgment and missed detection are difficult to avoid.
The YOLO model is used to train the functional interface images in different test environments. Through feature area recognition and comparison, functional interface defects in different test environments are identified and detected. The PP-YOLO algorithm and the N-PP-YOLO network are used to accurately locate and compare the feature areas to eliminate the influence of background differences.
It improves the accuracy of functional interface detection, reduces missed and missed detection, and can accurately identify graphical interface defects, especially those that are not easily detectable in different testing environments.
Smart Images

Figure CN120375147A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of software testing, and particularly relates to a method and system for assisting in detecting functional interface compatibility problems. Background Art
[0002] A functional interface refers to a graphical interface for interaction between users and software, mainly involving user operations and function implementation. The functional interface contains various graphical interface elements, such as buttons, text boxes, menus, toolbars, etc. Users interact with the software through these graphical interface elements to complete various tasks. Functional interface compatibility problems include functional interface defects. A functional interface defect refers to misalignment or abnormal display of graphical interface elements, as well as defects such as overlapping or chaotic layout between graphical interface elements in different test environments.
[0003] Testers usually use software testing methods to detect functional interfaces. Due to differences in the backgrounds of graphical interfaces in different test environments, software testing will classify the background differences of graphical interfaces as functional interface defects, resulting in incorrect detection results. For example, in a drawing software, the inability of the background of the graphical interface of some users to display dotted lines properly, the different clarity of solid lines in the backgrounds of the graphical interfaces of different users, or the different clarity of dotted lines in the backgrounds of the graphical interfaces of different users may all be misjudged as functional interface defects. In addition, testers are prone to missing functional interface defects in the corners of the graphical interface, resulting in missed detections. Also, functional interface defects are easy to detect in one test environment but difficult to detect in another test environment, which also leads to low accuracy of the detection results. Summary of the Invention
[0004] One of the purposes of this application is to provide a method and system for assisting in detecting functional interface compatibility problems, which can reduce the occurrence of missed detections and misdetections mentioned in the background art.
[0005] To achieve the above purpose and other related purposes, the first aspect of this application provides a method for assisting in detecting functional interface compatibility problems, which is characterized by including the following steps: obtaining multiple functional interface images in different test environments, and the multiple functional interface images form a training set; establishing a YOLO model; inputting the training set into the YOLO model for training to obtain a trained YOLO model; inputting the functional interface image to be detected into the trained YOLO model to obtain a functional interface image with a feature region in different test environments; comparing the functional interface images with the same feature region in different test environments to obtain a detection result on whether there are functional interface defects in the same feature region in different test environments.
[0006] The second aspect of the present application provides an auxiliary detection system for functional interface compatibility problems, including:
[0007] An acquisition module, configured to acquire multiple functional interface images under different test environments, and the multiple functional interface images form a training set;
[0008] A modeling module, configured to establish a YOLO model;
[0009] A training module, configured to input the training set into the YOLO model for training to obtain a trained YOLO model;
[0010] A generation module, configured to input the functional interface image to be detected into the trained YOLO model to obtain a functional interface image with a feature area under different test environments;
[0011] A comparison module, configured to compare the functional interface images with the same feature area under different test environments to obtain a detection result on whether there are functional interface defects in the same feature area under different test environments.
[0012] The present application has at least the following beneficial effects:
[0013] By using the auxiliary detection method and system for functional interface compatibility problems provided by the present application, multiple functional interface images under different test environments are acquired, and the multiple functional interface images form a training set; a YOLO model is established; the training set is input into the YOLO model for training to obtain a trained YOLO model; the functional interface image to be detected is input into the trained YOLO model to obtain a functional interface image with a feature area under different test environments; the functional interface images with the same feature area under different test environments are compared to obtain data on imperceptible functional interface defects of the application to be detected under different test environments, and further obtain a detection result on whether there are functional interface defects in the same feature area under different test environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 FIG. [ID] shows a working principle diagram of the auxiliary detection system for functional interface compatibility problems according to an embodiment of the present application;
[0016] Figure 2It is a flowchart of an auxiliary detection method for functional interface compatibility problems shown in the embodiments of the present application;
[0017] Figure 3 It is a schematic diagram of the Conv Block module of the auxiliary detection method for functional interface compatibility problems disclosed in the embodiments of the present application;
[0018] Figure 4 It is a schematic diagram of the Identity Block module of the auxiliary detection method for functional interface compatibility problems disclosed in the embodiments of the present application;
[0019] Figure 5 It is a schematic diagram of the backbone network of the auxiliary detection method for functional interface compatibility problems disclosed in the embodiments of the present application;
[0020] Figure 6 It is a schematic diagram of the feature pyramid network of the auxiliary detection method for functional interface compatibility problems disclosed in the embodiments of the present application; Detailed implementation manners
[0021] The following uses specific specific embodiments to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in the present application can also be based on different viewpoints and applications and be modified or changed without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0022] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.
[0023] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units, and a plurality of elements refers to two or more elements, etc.
[0024] The first aspect of the present application provides an auxiliary detection method for functional interface compatibility problems, including the following steps:
[0025] An acquisition step of acquiring a plurality of functional interface images in different test environments, and the plurality of functional interface images form a training set;
[0026] A modeling step of establishing a YOLO model;
[0027] Training step: Input the training set into the YOLO model for training to obtain the trained YOLO model;
[0028] Generation step: Input the functional interface image to be detected into the trained YOLO model to obtain the functional interface image with feature regions under different test environments;
[0029] Comparison step: Compare the functional interface images with the same feature regions under different test environments to obtain the detection results of whether there are functional interface defects in the same feature regions under different test environments.
[0030] Refer to Figure 1 , in the acquisition step, acquire the application operation videos recorded respectively for the application to be detected under different test environments to obtain multiple functional interface images. The multiple functional interface images include the data of functional interface defects that are not easily detected under different test environments.
[0031] In the embodiments of the present application, the application to be detected may be a desktop application with user interaction functions. For example, it may be an industrial design application, etc. The functions of this type of application generally do not have randomness and are definite functional reactions, which are more likely to reflect the performance gaps between different environments. The different test environments may be, for example: different operating systems. Specifically, for example, the different test environments may be a compatible platform and a native operating system platform.
[0032] Further, the different test environments include a first test environment and a second test environment. It is necessary to operate the application to be detected respectively in the first test environment and the second test environment, and record the videos that record the operation process and the interface feedback of the application to be detected for each operation in the test environment, which are recorded as "application operation videos". For example, run the application to be detected on the compatible platform and the native operating system platform respectively, perform the same operations on the application to be detected, and at the same time, record the application operation videos of each platform respectively.
[0033] In the acquisition step, the training set includes the industrial design application software dataset used for design target detection. Consider the classification criteria for the feature extraction regions and the marking positions, sizes, and type labels of the extraction regions.
[0034] For a dataset, generally two factors are mainly considered: one is the classification criterion for the feature extraction region, that is, the classification basis. Taking this method as an example, the main goal is to extract the main interface of the application - the drawing interface. The other is whether the dataset is a GUI dataset for real industrial design applications, and at the same time, the dataset should have sufficiently clear and complete labeled positions, sizes, and type labels for the extraction region. In this study, an industrial design software dataset designed for application compatibility problem detection is used, and the type annotations of the feature extraction region are added through manual annotation as a comparative experiment.
[0035] The training set includes the following:
[0036] Feature region annotation: Manually annotate the regions to be recognized in the key frames, such as the menu bar, text box, and window, etc. Record the position, size, and type of the feature region.
[0037] XML layout file annotation: The XML layout file of each application contains the hierarchical structure and attribute information of the feature region. Parse and analyze the XML file, extract the feature region information, and make corresponding annotations.
[0038] The key frame file format of the industrial design software dataset is PNG, and the attached label file is in JSON format. In the label file, the position, width, and height of the feature extraction region appearing in the key frame are represented by bounds, and the corresponding design semantic type is annotated by component Label.
[0039] Referring to Figure 2 , in the modeling step, the YOLO model belongs to the object detection model, including a backbone network, a feature pyramid network, and a prediction network connected in sequence. The backbone network is used to obtain feature maps of different scales from the images in the training set; the feature pyramid network is used to perform feature fusion on the feature maps of different scales to obtain feature maps of different scales that enhance the image features; the prediction network uses the feature maps of different scales that enhance the image features to detect the functional interface images to be detected.
[0040] This method uses PP-YOLO as the basic research framework for the feature region recognition task. PP-YOLO is a one-stage object detection algorithm based on the basic design idea of the YOLO algorithm. The PP-YOLO network is a one-stage lightweight object detection algorithm optimized on the basis of the YOLOv3 network architecture and combined with a variety of deep learning optimization algorithms. By classifying and regressing the feature information of the entire image, the coordinates of the bounding box of the target object, the confidence level of whether the target object is included, and the task-specific classification to which the target object belongs are obtained. On the basis of the PP-YOLO object detection framework, this method designs a feature region recognition network based on deep learning, named N-PP-YOLO, and the backbone feature extraction network is N_SE-vd that integrates the attention mechanism.
[0041] Through the recognition of the feature region, the non-full-screen desktop area generated during recording can be excluded from the subsequent comparison of application interfaces in different environments, eliminating the additional different elements caused by different backgrounds generated by different desktop operating systems in different environments. When the dataset is large enough, the feature region recognition can directly locate the locations where functional feature differences are likely to occur, greatly increasing the probability of locating defects that cannot be directly observed by the human eye.
[0042] Refer to Figure 5 and Figure 6 The backbone network includes multiple feature extraction modules; the multiple feature extraction modules are used to receive the images of the training set and output feature maps of different scales. The multiple feature extraction modules include a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, a fifth feature extraction module, a sixth feature extraction module, and a seventh feature extraction module connected in series in sequence.
[0043] The first feature extraction module includes a first convolution module. The first convolution module outputs a first feature map and transfers the first feature map to the second feature extraction module. The first convolution module includes a convolutional layer with 16 convolution kernels, a matrix size of 3×3, and a stride of 1. In some embodiments of the present application, the image of the training set is an RGB three-channel image of 576*576, with its width w and height h both being 576. After passing through the first feature extraction module, a first feature map of 576*576 with 16 channels is obtained. The first convolution module also includes a normalization module and a first activation function module. The normalization module can perform batch normalization to accelerate training and improve the stability of the model. The first activation function modules mentioned in this article all include activation units of the Mish function, and the Mish function is used to perform non-linear transformation on each element of the feature map to improve the stability of model training.
[0044] The second feature extraction module receives the first feature map output from the first feature extraction module, outputs the second feature map, and transfers the second feature map to the third feature extraction module. The second feature extraction module includes a ConvBlock module and an Identity Block module connected in sequence. The Conv Block module is used to change the dimensions of the feature map, and the dimensions include the number of channels and the spatial size. The Identity Block module is used to increase the depth of the network and improve the feature learning ability of the model. Both the ConvBlock module and the Identity Block module can be used as Residual blocks in the ResNet residual network. The ResNet residual network is a deep neural network structure that uses the method of residual connection to solve the problem of low accuracy of model object detection caused by the vanishing gradient problem in the training of deep networks.
[0045] Referring to Figure 3 , the Conv Block module includes a second convolutional module, a third convolutional module, a first attention module, a second convolutional module, and a first activation function module connected in sequence.
[0046] The second convolutional module of the Conv Block module includes a convolutional layer with a matrix size of 1×1 and a stride of 1, which is used to reduce the number of channels and the computational complexity. The third convolutional module includes a convolutional layer with a matrix size of 3×3 and a stride of 1, which is used to extract deep features.
[0047] The first attention module of the Conv Block module re-calibrates the weights of each channel of the feature map by compressing in the spatial dimension to remove the influence of spatial positions, and then through the learning and activation of a fully connected network. The first attention module includes a first average pooling module, a first fully connected layer module, and a second fully connected layer module connected in sequence.
[0048] The first average pooling module of the first attention module is used to compress the features in the spatial dimension. For example, a feature map of H*W with C channels outputs a feature map of 1*1 with C channels after global average pooling, that is, it has C different feature values, and each feature value corresponds to the average value of all elements in the channel.
[0049] The first fully connected layer module of the first attention module includes a first fully connected layer and a first activation function module.
[0050] The first fully connected layer is used to reduce the dimension of the feature map passed through the first average pooling module, aiming to reduce the computational load. Specifically, the first fully connected layer compresses the number of channels from C to C / r, where r is the number of nodes in the first fully connected layer. For example, when r is 4, the number of channels of the output feature map after passing through the first fully connected layer is C / 4. The first activation function module performs a non-linear transformation on each element of the feature map to improve the stability of model training.
[0051] The second fully connected layer module of the first attention module includes a second fully connected layer and a second activation function module.
[0052] The second fully connected layer is used to increase the dimension of the feature map passed through the first fully connected layer module, restoring the number of channels of the feature map to C, which is the number of channels before entering the first fully connected layer. For example, the number of nodes in the second fully connected layer is 2C. The second activation function modules mentioned in this article all include activation units of the Sigmoid function. The Sigmoid function is used to increase the attention weight, and the attention weight is used to enhance the important features in the feature map and suppress the unimportant features in the feature map. The attention weight output via the second fully connected layer module is multiplied by the features of the feature map when entering the first attention module channel by channel to obtain a feature map with attention features. The attention features can enable the network to pay more attention to important feature regions and improve the performance of the model. The feature map with attention features undergoes a residual connection after passing through the second convolution module located after the first attention module. The following specifically elaborates on the implementation process of the residual connection.
[0053] The Conv Block module includes a first residual connection path. The first residual connection path includes a second average pooling module and a second convolution module. The feature map input to the Conv Block module is processed by the second average pooling module and the second convolution module of the first residual connection path, and then undergoes a residual connection with the feature map with attention features output by the second convolution module located after the first attention module. The residual connection is used to solve the problem of low accuracy of model object detection caused by the vanishing gradient in the training of deep networks. The feature map after the residual connection is output to the Identity Block module after passing through the first activation function module at the end of the Conv Block module.
[0054] Refer to Figure 4 , the Identity Block module includes a second convolution module, a third convolution module, a second attention module, a second convolution module, and a first activation function module connected in sequence.
[0055] The second convolutional module of the Identity Block module includes a convolutional layer with a matrix size of 1×1 and a stride of 1, which is used to reduce the number of channels and the computational complexity. The third convolutional module includes a convolutional layer with a matrix size of 3×3 and a stride of 1, which is used to extract deep features.
[0056] The second attention module of the Identity Block module removes the influence of spatial positions by compressing in the spatial dimension, and then obtains the weights of each channel of the feature map through the learning and activation of a fully connected network, completing the recalibration of each feature of the feature map in the channel dimension. The second attention module includes a first average pooling module, a first fully connected layer module, and a third fully connected layer module connected in sequence.
[0057] The first average pooling module of the second attention module is used to compress the features in the spatial dimension. For example, a feature map of H*W with C channels is output as a feature map of 1*1 with C channels after global average pooling, that is, it has C different feature values, and each feature value corresponds to the average value of all elements in the channel.
[0058] The first fully connected layer module of the second attention module includes a first fully connected layer and a first activation function module.
[0059] The first fully connected layer is used to reduce the dimension of the feature map transmitted by the first average pooling module to reduce the computational amount. Specifically, the first fully connected layer compresses C channels into C / r channels, where r is the number of nodes in the first fully connected layer. The first activation function module performs a non-linear transformation on each element of the feature map to improve the stability of model training.
[0060] The third fully connected layer module of the second attention module includes a third fully connected layer and a second activation function module. The number of nodes in the third fully connected layer can be C. The second activation function modules mentioned in this article all include activation units of the Sigmoid function. The Sigmoid function is used to increase the attention weights, and the attention weights are used to enhance the important features in the feature map and suppress the unimportant features in the feature map. The attention weights output by the third fully connected layer module are multiplied by the features of the feature map when entering the second attention module channel by channel to obtain a feature map with attention features. The attention features can make the network pay more attention to important feature regions and improve the performance of the model. The feature map with attention features is subjected to residual connection after passing through the second convolutional module located after the second attention module.
[0061] The Identity Block module includes a second residual connection path. The feature map input to the Identity Block module is subjected to residual connection with the feature map with attention features output by the second convolution module located after the second attention module through the second residual connection path. The residual connection is used to solve the problem of low accuracy of model object detection caused by the vanishing gradient problem in the network. The feature map after the residual connection is output to the third feature extraction module after passing through the first activation function module at the end of the Identity Block module.
[0062] Refer to Figure 5 and Figure 6 , the third feature extraction module receives the second feature map output from the second feature extraction module and outputs a third feature map, and transfers the third feature map to the fourth feature extraction module. The third feature extraction module includes a Conv Block module and 2 Identity Block modules connected in sequence.
[0063] Refer to Figure 5 and Figure 6 , the fourth feature extraction module receives the third feature map output from the third feature extraction module and outputs a fourth feature map, and outputs the fourth feature map to the fifth feature extraction module and to the feature pyramid network respectively. The fourth feature extraction module includes a Conv Block module and 8 Identity Block modules connected in sequence.
[0064] Refer to Figure 5 and Figure 6 , the fifth feature extraction module receives the fourth feature map output from the fourth feature extraction module and outputs a fifth feature map, and outputs the fifth feature map to the sixth feature extraction module and to the feature pyramid network respectively. The fifth feature extraction module includes a Conv Block module and 6 Identity Block modules connected in sequence.
[0065] Refer to Figure 5 and Figure 6 , the sixth feature extraction module receives the fifth feature map output from the fifth feature extraction module and outputs a sixth feature map, and outputs the sixth feature map to the seventh feature extraction module and to the feature pyramid network respectively. The sixth feature extraction module includes a Conv Block module and 4 Identity Block modules connected in sequence.
[0066] Refer to Figure 5 and Figure 6, the seventh feature extraction module receives the sixth feature map output from the sixth feature extraction module, outputs the seventh feature map, and outputs the seventh feature map to the feature pyramid network. The seventh feature extraction module includes a Conv Block module and two Identity Block modules connected in sequence. The seventh feature map is used to extract higher-level abstract design features, thereby providing support for the prediction of the feature regions of the overall contour shape of the target.
[0067] The structures and functions of the Conv Block module and the Identity Block module of the third to seventh feature extraction modules are the same as those of the Conv Block module and the Identity Block module of the second feature extraction module, and will not be elaborated here one by one.
[0068] Using the backbone network of this application, compared with ResNet50-vd, the network depth is increased, with a total of 76 convolutional layers. Considering the functional interface compatibility issue, the auxiliary detection method has no requirement for real-time performance. Therefore, the processing speed is sacrificed to enhance the model's ability to extract information from the input key frames and improve the model's recognition ability for feature regions.
[0069] Refer to Figure 6 , the feature pyramid network includes a first feature fusion module and a second feature fusion module. The first feature fusion module is used to improve the recognition ability for small-range targets such as functional interface defects existing in the corners of the graphical interface. The second feature fusion module can improve the positioning accuracy for large-range targets such as graphical interface elements. The specific structures of the first feature fusion module and the second feature fusion module are described in detail below.
[0070] The first feature fusion module includes a 1A feature fusion module, a 1B feature fusion module, and a 1C feature fusion module connected in sequence.
[0071] The 1A feature fusion module makes the deep features of the seventh feature map and the sixth feature map concatenated in the depth direction through an upsampling method. The deep features include deep semantic features. The 1A feature fusion module includes an upsampling module, a feature map merging module, and a Conv module connected in sequence. The Conv module includes a convolutional layer with a size of 3×3, which is used to smooth the feature map.
[0072] The upsampling module of the 1A feature fusion module is used to receive the seventh feature map from the seventh feature extraction module, and after enlarging the seventh feature map to the same scale as the sixth feature map, outputs it to the feature map merging module of the 1A feature fusion module. Using the upsampling module can reduce the computational amount of the model and improve the training speed of the model. The upsampling module includes a convolutional layer with a size of 1×1, which is used to reduce the dimension of the feature map.
[0073] The feature map merging module of the 1A feature fusion module is used to receive the feature map output by the upsampling module of the 1A feature fusion module and the sixth feature map from the sixth feature extraction module that has been processed by the Conv module of the 1A feature fusion module, and after merging the two feature maps, it transmits them to the 1B feature fusion module and the second feature fusion module respectively. During the merging process of the feature map merging module, the deep features of the feature map output by the upsampling module of the 1A feature fusion module and the sixth feature map from the sixth feature extraction module that has been processed by the Conv module of the 1A feature fusion module are concatenated in the depth direction.
[0074] The 1B feature fusion module makes the deep features of the sixth feature map and the fifth feature map concatenated in the depth direction through the upsampling method. The 1B feature fusion module includes an upsampling module, a feature map merging module, and a Conv module connected in sequence. The Conv module includes a convolutional layer with a size of 3×3, which is used to smooth the feature map.
[0075] The upsampling module of the 1B feature fusion module is used to receive the sixth feature map with deep features output by the 1A feature fusion module, and after enlarging the sixth feature map to the same scale as the fifth feature map, it outputs to the feature map merging module of the 1B feature fusion module. The upsampling module includes a convolutional layer with a size of 1×1, which is used to reduce the dimension of the feature map.
[0076] The feature map merging module of the 1B feature fusion module is used to receive the feature map output by the upsampling module of the 1B feature fusion module and the fifth feature map from the fifth feature extraction module that has been processed by the Conv module of the 1B feature fusion module, and after merging the two feature maps, it transmits them to the 1C feature fusion module and the second feature fusion module respectively. During the merging process of the feature map merging module, the deep features of the feature map output by the upsampling module of the 1B feature fusion module and the fifth feature map from the fifth feature extraction module that has been processed by the Conv module of the 1B feature fusion module are concatenated in the depth direction.
[0077] The 1C feature fusion module makes the deep features of the fifth feature map and the fourth feature map concatenated in the depth direction through the upsampling method. The 1C feature fusion module includes an upsampling module, a feature map merging module, and a Conv module connected in sequence. The Conv module includes a convolutional layer with a size of 3×3, which is used to smooth the feature map.
[0078] The upsampling module of the 1C feature fusion module is used to receive the fifth feature map with deep features output by the 1B feature fusion module, and after enlarging the fifth feature map to the same scale as the fourth feature map, it outputs to the feature map merging module of the 1C feature fusion module. The upsampling module includes a convolutional layer with a size of 1×1, which is used to reduce the dimension of the feature map.
[0079] The feature map merging module of the first C feature fusion module is used to receive the feature map output by the upsampling module of the first C feature fusion module and the fourth feature map from the fourth feature extraction module that has been processed by the Conv module of the first C feature fusion module, and after merging the two feature maps, respectively pass them to the second feature fusion module. During the merging process of the feature map merging module, the deep features of the feature map output by the upsampling module of the 1C feature fusion module and the fourth feature map from the fourth feature extraction module that has been processed by the Conv module of the first C feature fusion module are concatenated in the depth direction.
[0080] The second feature fusion module includes a 2A feature fusion module, a 2B feature fusion module, and a 2C feature fusion module connected in sequence.
[0081] The 2A feature fusion module makes the shallow features of the fourth feature map and the fifth feature map concatenated in the depth direction through a downsampling method. The shallow features include shallow edge features. The 2A feature fusion module includes a Conv module, a downsampling module, and a feature map merging module connected in sequence.
[0082] The Conv module of the 2A feature fusion module is used to receive the fourth feature map with deep features output by the feature map merging module of the first C feature fusion module, and output the fourth feature map processed by the Conv module to the prediction network and the downsampling module of the 2A feature fusion module. The Conv module includes a convolutional layer with a size of 3×3 for smoothing the feature map.
[0083] The downsampling module of the 2A feature fusion module is used to receive the fourth feature map output by the Conv module of the 2A feature fusion module, and after shrinking the fourth feature map to the same scale as the fifth feature map, output it to the feature map merging module of the 2A feature fusion module. The downsampling module includes a convolutional layer with a size of 1×1 for dimensionality increase of the feature map.
[0084] The feature map merging module of the 2A feature fusion module is used to receive the feature map output by the downsampling module of the 2A feature fusion module and the fifth feature map output by the feature map merging module of the first B feature fusion module, and after merging the two feature maps, respectively pass them to the 2B feature fusion module and the prediction network. During the merging process of the feature map merging module, the shallow features of the feature map output by the downsampling module of the 2A feature fusion module and the fifth feature map output by the feature map merging module of the first B feature fusion module are concatenated in the depth direction.
[0085] The 2B feature fusion module makes the shallow features of the fifth feature map and the sixth feature map concatenated in the depth direction through a downsampling method. The 2B feature fusion module includes a Conv module, a downsampling module, and a feature map merging module connected in sequence.
[0086] The Conv module of the 2B feature fusion module is used to receive the fifth feature map with shallow features output by the feature map merging module of the 2A feature fusion module, and output the fifth feature map processed by the Conv module to the prediction network and the downsampling module of the 2B feature fusion module.
[0087] The downsampling module of the 2B feature fusion module is used to receive the fifth feature map output by the Conv module of the 2B feature fusion module, and reduce the fifth feature map to the same scale as the sixth feature map and then output it to the feature map merging module of the 2B feature fusion module. The downsampling module includes a 1×1 convolutional layer for dimensionality increase of the feature map.
[0088] The feature map merging module of the 2B feature fusion module is used to receive the feature map output by the downsampling module of the 2B feature fusion module and the sixth feature map output by the feature map merging module of the 1A feature fusion module, and merge the feature maps of the two and then transfer them to the 2C feature fusion module and the prediction network respectively. During the merging process of the feature map merging module, the shallow features of the feature map output by the downsampling module of the 2B feature fusion module and the sixth feature map output by the feature map merging module of the 1A feature fusion module are concatenated in the depth direction.
[0089] The 2C feature fusion module makes the shallow features of the sixth feature map and the seventh feature map concatenated in the depth direction through a downsampling method. The 2C feature fusion module includes a Conv module, a downsampling module, and a feature map merging module connected in sequence.
[0090] The Conv module of the 2C feature fusion module is used to receive the sixth feature map with shallow features output by the feature map merging module of the 2B feature fusion module, and output the sixth feature map processed by the Conv module to the prediction network and the downsampling module of the 2C feature fusion module.
[0091] The downsampling module of the 2C feature fusion module is used to receive the sixth feature map output by the Conv module of the 2C feature fusion module, and reduce the sixth feature map to the same scale as the seventh feature map and then output it to the feature map merging module of the 2C feature fusion module. The downsampling module includes a 1×1 convolutional layer for dimensionality increase of the feature map.
[0092] The feature map merging module of the second C feature fusion module is used to receive the feature map output by the downsampling module of the second C feature fusion module and the seventh feature map output by the seventh feature extraction module, and merge the feature maps of the two and then transfer them to the prediction network. During the merging process of the feature map merging module, the shallow features of the feature map output by the downsampling module of the 2C feature fusion module are concatenated with the seventh feature map output by the seventh feature extraction module in the depth direction.
[0093] The feature pyramid takes advantage of the characteristic that the output size of the backbone network decreases layer by layer through downsampling, and naturally has a pyramid structure. It concatenates the low-resolution feature maps with high-dimensional semantics and the high-resolution feature maps with low-dimensional semantics in the depth direction, minimizing the computational cost and completing feature fusion. The pyramid network of this application is based on the FPN and adopts the feature fusion method of FPN+PANet. On the basis of top-down, it adds a bottom-up feature fusion route, and concatenates the shallow output feature maps containing shallow edge information with the high-level feature maps in the depth direction through downsampling, making the context information contained in each hierarchical level richer. The high-level feature maps responsible for predicting large feature regions are supplemented with rich shallow edge feature information to assist the network in judging large feature regions.
[0094] The prediction network is used to receive the fourth, fifth, sixth, and seventh feature maps with shallow and deep features output by the feature pyramid network, divide the functional interface image to be detected into grids of different sizes according to the sizes of the fourth, fifth, sixth, and seventh feature maps, and each grid is responsible for predicting the category of the target falling into the grid, the coordinates of the predicted bounding box corresponding to the target, and the confidence that the predicted bounding box matches the target. Through the prediction of the above multi-size feature maps, targets of different sizes can be detected. The targets include small-range targets such as functional interface defects and large-range targets such as graphical interface elements.
[0095] The calculation formula for the confidence of the target is where P r (object) indicates whether the center of the target falls into the grid. If it falls, the value is 1; if it does not fall, the value is 0. represents the intersection over union of the predicted bounding box and the ground truth bounding box.
[0096] For example, the size of the fourth feature map is 72×72, the size of the fifth feature map is 36×36, the size of the sixth feature map is 18×18, and the size of the seventh feature map is 9×9, and then four different sizes of grids of 9×9, 18×18, 36×36, and 72×72 are divided.
[0097] The prediction network includes a training stage and a prediction stage.
[0098] The training stage specifically includes the following steps:
[0099] S11. Obtain the coordinate parameters of the true bounding box corresponding to the target to be trained;
[0100] S12. Determine the center of the predicted bounding box according to the coordinate parameters of the true bounding box corresponding to the target to be trained;
[0101] S13. Determine the position of the predicted bounding box according to the center of the predicted bounding box;
[0102] S14. Determine and mark the feature labels of the target to be trained according to the position of the predicted bounding box. The feature labels of the target to be trained include the category of the target to be trained falling into the grid, the coordinates of the predicted bounding box corresponding to the target to be trained, and the confidence that the predicted bounding box matches the target to be trained. The coordinates of the predicted bounding box corresponding to the target to be trained include the upper left coordinates (x, y) of the predicted bounding box, the width w of the predicted bounding box, and the height h of the predicted bounding box.
[0103] The confidence that the predicted bounding box matches the target also includes the confidence in predicting the target category. The calculation formula for the confidence in predicting the target category is where P r (class|object) represents the probability that each grid predicts the target category to which the target to be trained belongs. For example, if there are 3 categories for the target to be trained, then each grid predicts the probabilities of the target belonging to these 3 different categories respectively and obtains a set of probability values, and this set of probability values includes 3 probability values corresponding to 3 different categories.
[0104] The prediction stage specifically includes the following steps:
[0105] S21. Input the functional interface image to be detected into the trained YOLO model;
[0106] S22. Perform grid division on the functional interface image to be detected;
[0107] S23. Each grid is provided with multiple anchor boxes of different predefined sizes;
[0108] S24. Select the anchor box of the corresponding size for prediction according to the sizes of different targets to be predicted, and obtain the category of the target to be predicted, the coordinates of the predicted bounding box corresponding to the target to be predicted, and the confidence that the predicted bounding box matches the target to be predicted. The targets to be predicted include small-range targets such as functional interface defects and large-range targets such as graphical interface elements.
[0109] For example, the number of channels of the functional interface image to be detected is 16, and the size is 576×576. The feature map is divided into four grids of different sizes, namely 9×9, 18×18, 36×36, and 72×72. Each grid has 4 anchor boxes of different predefined sizes. The anchor box size is calculated based on the size of the true bounding box in the training set through dimensional clustering using K-mean++. The center of each anchor box is at the center of the corresponding grid. When the prediction falls into the grid, the intersection of the predicted bounding box and the true bounding box is selected. The anchor box corresponding to the maximum value predicts the target to be predicted.
[0110] In the process of using anchor boxes for prediction, the center coordinates of the anchor boxes are the coordinates of the upper left corner of each grid (P x , P y ) as the starting point. The position parameters of the predicted anchor box for the target to be predicted include the center coordinates of the predicted anchor box (b x , b y ), predict the width of the anchor box b w , predict the height of the anchor box b h The position parameters of the anchor box are related to the network prediction parameters output by the model. The network prediction parameters include the predicted coordinates (t x , t y ), prediction width t w , predicted height t h , the specific calculation formula is as follows: x =σ(t x )+C x ; where t x is the network prediction parameter, σ(t x ) represents the grid prediction value t x Calculate the Sigmoid function so that σ(t x ) is between 0 and 1, ensuring that the center coordinates of the target to be predicted are limited to the current grid and will not deviate from the current grid range responsible for prediction. x Indicates the offset of the grid coordinates of the first row of the feature map from the grid where the center of the target to be predicted is located. y =σ(t y )+C y ; where t y is the network prediction parameter, σ(t y ) represents the grid prediction value t y Calculate the Sigmoid function so that σ(t y ) is between 0 and 1, ensuring that the center coordinates of the target to be predicted are limited to the current grid and will not deviate from the current grid range responsible for prediction. yRepresents the offset of the grid coordinates of the grid where the center of the target to be predicted is located from the first column of the feature map.
[0111] Among them, P w is the preset width of the anchor box, and the preset width of the anchor box is obtained through dimensional clustering calculation of K-mean++ based on the sizes of the true bounding boxes in the training set. t w is the network prediction parameter, and exponential calculation is adopted to improve the convergence speed of the obtained prediction value.
[0112] Among them, P h is the preset height of the anchor box, and the preset height of the anchor box is obtained through dimensional clustering calculation of K-mean++ based on the sizes of the true bounding boxes in the training set. t h is the network prediction parameter, and exponential calculation is adopted, which can improve the convergence speed of the obtained prediction value.
[0113] Compare the functional interface images with the same feature regions under different test environments to obtain the detection results of whether there are functional interface defects in the same feature regions under different test environments, and further obtain the data of imperceptible functional interface defects of the application to be detected under different test environments respectively, and avoid the background technology mentioned that due to the differences in the backgrounds of different user graphical interfaces, software testing will classify the background differences of the graphical interface as functional interface defects, resulting in incorrect detection results.
[0114] Refer to Figure 1 , and the specific comparison of the functional interface images with the same feature regions under different test environments includes:
[0115] S31, different test environments include the first test environment and the second test environment. Obtain the functional interface images with the same feature regions under the first test environment and the second test environment, select one of them as the experimental image, and the other as the control image;
[0116] S32, adopt the Canny edge detection algorithm to extract the edge features of the experimental image and the control image, and obtain the edge feature image corresponding to the experimental image and the edge feature image corresponding to the control image; The edge features include small-range targets such as functional interface defects and large-range targets such as graphical interface elements. The edge feature image includes the contour information and edge information of the edge features.
[0117] The Canny edge detection algorithm uses a two-dimensional zero-mean Gaussian function and performs a convolution operation on the image matrix to achieve noise elimination and smoothing processing of the image. The Canny edge detection algorithm includes the following calculation formulas:
[0118] Gaussian distribution function calculation formula:
[0119] Among them, (i, j) are the coordinates of the image pixel points, and σ is the Gaussian filter parameter, which is used to control the degree of image filtering.
[0120] The calculation formula for obtaining the smoothed image by performing convolution operation on the image matrix: I(i, j) = G(i, j) * f(i, j); where I(i, j) is the smoothed image, and f(i, j) is the experimental image or the control image. Calculate the gradient magnitude and direction of I(i, j), including obtaining the partial derivatives at the point using the difference operation, and using the Euclidean distance and the arctangent function to obtain the gradient and its direction at this point.
[0121] The steps for obtaining the contour information and edge information of the edge features through image noise elimination include:
[0122] S32a, performing non-maximum suppression for edge thinning; the main purpose of non-maximum suppression is to further thin the edges on the basis of edge detection, and only retain those points with the most drastic change in gray value, so as to eliminate the redundant information generated by multiple edge response points. Specifically, for each point in the image, compare the gradient magnitudes of the two adjacent points in the direction where the gray value of this point changes fastest. If the gradient magnitude is greater than the gradient magnitudes of the two adjacent points in the gradient direction, then this point is set as a candidate boundary point, otherwise, it is marked as a non-boundary point.
[0123] S32b, adopting a dual-threshold edge detection and connection edge processing method. Classify and process the candidate boundary points through two preset high and low thresholds. For each candidate boundary point, calculate its edge gradient value. If the edge gradient value of a certain point is greater than the preset high threshold, then this point is marked as a strong edge point. For the candidate boundary points whose edge gradient values are between the high and low thresholds, further processing is required. If these points are spatially connected to the points that have been marked as strong edges, then these points are also classified as edge points. For the candidate boundary points whose edge gradient values are lower than the low threshold, they are directly deleted as noise or background information.
[0124] S33, perform difference calculation on the edge feature image corresponding to the experimental image and the edge feature image corresponding to the control image to obtain a difference map; it is required that the edge feature image corresponding to the experimental image and the edge feature image corresponding to the control image have the same size and number of channels. Traverse each pixel in the two images and calculate the difference between the corresponding pixel values in the two images.
[0125] S34, perform data processing on the difference map, and the data processing includes data processing such as filtering and enhancing contrast; use a Gaussian blur smoothing filter for data processing, and the formula is as follows:
[0126] Among them, (i, j) are the pixel coordinates of the difference map, and σ is the Gaussian filter parameter, which is used to control the degree of image filtering.
[0127] The calculation formula for obtaining the smoothed image by performing convolution operation on the image matrix: I(i, j) = G(i, j) * f(i, j); where I(i, j) is the smoothed image, and f(i, j) is the difference map.
[0128] S35, perform visualization processing on the difference map after data processing. Superimpose the processed difference map on the experimental image and label the difference areas.
[0129] Specifically, it includes the following steps:
[0130] Perform image preprocessing on the experimental image and the processed difference map to ensure that their sizes are the same and convert them into grayscale images;
[0131] Set the transparency; set a transparency value (alpha value) for the difference map, which determines the visibility of the difference map when superimposed.
[0132] Image superposition; realize superimposing the difference map on the experimental image by means of weighted average.
[0133] Label the difference areas; set the threshold of the difference, traverse the difference image, and find the areas that exceed the threshold. These areas represent the significant differences from the experimental image.
[0134] Color mapping; in order to show the differences more clearly, perform color mapping on the difference areas and highlight the differences with red or blue.
[0135] Adopting the method for assisting in detecting functional interface compatibility problems of the present application improves the test efficiency of software testing and lays a foundation for the commercial customization of domestic desktop operating systems.
[0136] Referring to Figure 1 , the second aspect of the present application provides a system for assisting in detecting functional interface compatibility problems, including:
[0137] An acquisition module, which is used to acquire multiple functional interface images in different test environments, and the multiple functional interface images form a training set;
[0138] A modeling module, which is used to establish a YOLO model;
[0139] A training module, which is used to input the training set into the YOLO model for training to obtain a trained YOLO model;
[0140] A generation module, which is used to input the functional interface image to be detected into the trained YOLO model to obtain a functional interface image with characteristic areas in different test environments;
[0141] A comparison module, configured to compare functional interface images with the same feature regions under different test environments, and obtain a detection result on whether there are functional interface defects in the same feature regions under different test environments.
[0142] A generation module, including but not limited to events such as feature region recognition and feature region comparison. The generation module completes the tasks of recognizing the feature regions from the input images and comparing the feature regions. The generation module needs to accurately locate the position and type of the feature regions in the interface by an automated method, and then compare the similarity of the feature regions in different environments through the comparison module, so as to better detect the differences in function availability caused by different displays, and assist manual testing.
[0143] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present application, several improvements and replacements can be made, and these improvements and replacements should also be regarded as the protection scope of the present application.
Claims
1. A method for assisting in detecting functional interface compatibility problems, characterized in that Including the following steps: Obtain multiple functional interface images under different test environments, and the multiple functional interface images form a training set; establish a YOLO model; input the training set into the YOLO model for training to obtain a trained YOLO model; input the functional interface image to be detected into the trained YOLO model to obtain a functional interface image with a feature area under different test environments; compare the functional interface images with the same feature area under different test environments to obtain a detection result on whether there are functional interface defects in the same feature area under different test environments.
2. The auxiliary detection method for functional interface compatibility problems according to claim 1, characterized in that, The YOLO model includes a backbone network, a feature pyramid network, and a prediction network connected in sequence; the backbone network is used to obtain feature maps of different scales from the images in the training set; The feature pyramid network is used to perform feature fusion on the feature maps of different scales to obtain feature maps of different scales that enhance image features; the prediction network uses the feature maps of different scales that enhance image features to detect the functional interface image to be detected.
3. The auxiliary detection method for functional interface compatibility problems according to claim 2, characterized in that The backbone network includes a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, a fifth feature extraction module, a sixth feature extraction module, and a seventh feature extraction module connected in series in sequence; The first feature extraction module includes a first convolution module, and the first convolution module outputs a first feature map and transfers the first feature map to the second feature extraction module; The second feature extraction module outputs a second feature map after receiving the first feature map output from the first feature extraction module and transfers the second feature map to the third feature extraction module; The second feature extraction module includes a Conv Block module and an Identity Block module connected in sequence; The third feature extraction module outputs a third feature map after receiving the second feature map output from the second feature extraction module and transfers the third feature map to the fourth feature extraction module; The third feature extraction module includes a Conv Block module and two Identity Block modules connected in sequence; The fourth feature extraction module outputs a fourth feature map after receiving the third feature map output from the third feature extraction module and outputs the fourth feature map to the fifth feature extraction module and to the feature pyramid network respectively; The fourth feature extraction module includes a Conv Block module and eight Identity Block modules connected in sequence; The fifth feature extraction module outputs a fifth feature map after receiving the fourth feature map output from the fourth feature extraction module and outputs the fifth feature map to the sixth feature extraction module and to the feature pyramid network respectively; The fifth feature extraction module includes a Conv Block module and six Identity Block modules connected in sequence; The sixth feature extraction module outputs a sixth feature map after receiving the fifth feature map output from the fifth feature extraction module and outputs the sixth feature map to the seventh feature extraction module and to the feature pyramid network respectively; The sixth feature extraction module includes a Conv Block module and four Identity Block modules connected in sequence; The seventh feature extraction module receives the sixth feature map output from the sixth feature extraction module, outputs the seventh feature map, and outputs the seventh feature map to the feature pyramid network; The seventh feature extraction module includes a Conv Block module and two Identity Block modules connected in sequence.
4. The functional interface compatibility problem auxiliary detection method according to claim 3, wherein The ConvBlock module includes a second convolution module, a third convolution module, a first attention module, a second convolution module, and a first activation function module connected in sequence.
5. The auxiliary detection method for functional interface compatibility problems according to claim 4, characterized in that The IdentityBlock module includes a second convolution module, a third convolution module, a second attention module, a second convolution module, and a first activation function module connected in sequence.
6. The auxiliary detection method for functional interface compatibility problems according to claim 3, characterized in that The feature pyramid network includes a first feature fusion module and a second feature fusion module; The first feature fusion module includes a 1A feature fusion module, a 1B feature fusion module, and a 1C feature fusion module connected in sequence; the 1A feature fusion module splices the deep features of the seventh feature map with the sixth feature map in the depth direction through an upsampling method; the 1B feature fusion module splices the deep features of the sixth feature map with the fifth feature map in the depth direction through an upsampling method; the 1C feature fusion module splices the deep features of the fifth feature map with the fourth feature map in the depth direction through an upsampling method; The second feature fusion module includes a 2A feature fusion module, a 2B feature fusion module, and a 2C feature fusion module connected in sequence; the 2A feature fusion module splices the shallow features of the fourth feature map with the fifth feature map in the depth direction through a downsampling method; the 2B feature fusion module splices the shallow features of the fifth feature map with the sixth feature map in the depth direction through a downsampling method; the 2C feature fusion module splices the shallow features of the sixth feature map with the seventh feature map in the depth direction through a downsampling method.
7. The auxiliary detection method for functional interface compatibility problems according to claim 1, characterized in that, The prediction network includes a training stage, and the training stage specifically includes the following steps: S11, obtaining the coordinate parameters of the true bounding box corresponding to the target to be trained; S12, determining the center of the predicted bounding box according to the coordinate parameters of the true bounding box corresponding to the target to be trained; S13, determining the position of the predicted bounding box according to the center of the predicted bounding box; S14, determining and marking the feature labels of the target to be trained according to the position of the predicted bounding box.
8. The auxiliary detection method for functional interface compatibility problems according to claim 7, characterized in that The prediction network includes a prediction stage, and the prediction stage specifically includes the following steps: S21, inputting the functional interface image to be detected into the trained YOLO model, S22, performing grid division on the functional interface image to be detected; S23, each grid is provided with multiple anchor boxes of different predefined sizes; S24, selecting the anchor box of the corresponding size according to the size of different targets to be predicted for prediction, and obtaining the category of the target to be predicted, the coordinates of the predicted bounding box corresponding to the target to be predicted, and the confidence level of whether the predicted bounding box matches the target to be predicted.
9. The functional interface compatibility problem assisted detection method according to claim 8, characterized in that Comparing functional interface images with the same feature region in different test environments specifically includes: S31. Different test environments include a first test environment and a second test environment. Obtain functional interface images with the same feature regions in the first test environment and the second test environment, select one of them as the experimental image, and the other as the control image; S32. Use the Canny edge detection algorithm to extract the edge features of the experimental image and the control image, obtaining an edge feature image corresponding to the experimental image and an edge feature image corresponding to the control image; S33. Perform a difference calculation on the edge feature image corresponding to the experimental image and the edge feature image corresponding to the control image to obtain a difference map; S34. Perform data processing on the difference map; the data processing includes filtering and enhancing the contrast; S35. Perform visualization processing on the difference map after data processing; superimpose the processed difference map on the experimental image and label the difference regions.
10. A functional interface compatibility problem assisted detection system, characterized in that, Including: An acquisition module, used to acquire multiple functional interface images in different test environments, and the multiple functional interface images form a training set; A modeling module, used to establish a YOLO model; A training module, used to input the training set into the YOLO model for training to obtain a trained YOLO model; A generation module, used to input the functional interface image to be detected into the trained YOLO model to obtain a functional interface image with a feature region in different test environments; A comparison module, used to compare the functional interface images with the same feature regions in different test environments to obtain a detection result on whether there are functional interface defects in the same feature regions in different test environments.