Abnormality detection method and device, terminal equipment and storage medium

Through multiple detection networks, the display of application interfaces in smart devices is detected, and the display abnormalities are determined in combination with the detection results. The problem of insufficient detection efficiency and accuracy in the prior art is solved, and automated and accurate display abnormalities detection is realized.

CN120029877APending Publication Date: 2025-05-23BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311532819.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect and identify abnormalities displayed on the application interface in smart devices, affecting the user experience.

Method used

By using multiple different detection networks to detect the display interface, and combining the detection results of multiple detection networks, it is determined whether there is a display abnormality in the display interface. These detection networks can include convolutional neural networks, object detection models, etc., and feature extraction and detection are performed through different feature channels and convolution kernel scales.

Benefits of technology

It improves the efficiency and accuracy of detecting abnormal display interfaces, realizes automatic detection, does not require manual operation, has strong versatility, and improves detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029877A_ABST
    Figure CN120029877A_ABST
Patent Text Reader

Abstract

The invention relates to an anomaly detection method and device, terminal equipment and a storage medium. The anomaly detection method comprises the steps of obtaining a to-be-detected display interface; determining whether the display interface is abnormal or not through different detection networks; wherein different detection networks have different detection effects. Whether the display interface is abnormal or not is detected through the multiple detection networks, the detection results of the multiple detection networks are combined, the efficiency and accuracy of detecting the display abnormality of the display interface are improved, automatic detection is achieved, manual operation is not needed, high universality is achieved, and the detection effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of information processing technology, and in particular to an anomaly detection method, apparatus, terminal device and storage medium. Background Art

[0002] With the development of smart devices, smart devices are used in more and more fields and scenarios. For example, mobile smart terminals account for a large proportion of smart devices. These smart devices usually have a display screen and are also installed with applications. Different applications have different functions; the display screen can display various interfaces of the smart device. In some cases, during the operation of the application, some interfaces may display abnormally due to various problems, affecting the display effect and thus affecting the user experience. Summary of the invention

[0003] The present invention provides an anomaly detection method, an apparatus, a terminal device and a storage medium.

[0004] According to a first aspect of an embodiment of the present disclosure, there is provided an abnormality detection method, comprising: acquiring a display interface to be detected; determining whether a display abnormality occurs on the display interface through different detection networks; wherein different detection networks have different detection effects.

[0005] In one embodiment, determining whether the display interface has a display abnormality through different detection networks includes: detecting the display interface through a first detection network to obtain a first detection result; when the first detection result is that the display is normal, detecting the display interface through a second detection network to obtain a second detection result; determining whether the display interface has a display abnormality based on the second detection result; wherein the first detection network is different from the second detection network.

[0006] In one embodiment, determining whether the display interface has a display abnormality through different detection networks also includes: when the first detection result is a display abnormality, detecting the display interface through a third detection network to obtain a third detection result; determining whether the display interface has a display abnormality based on the second detection result and the third detection result; wherein the third detection network is different from the first detection network and the second detection network.

[0007] In one embodiment, the detecting of the display interface by a second detection network to obtain a second detection result at least includes: extracting a first feature map of the display interface; extracting a second feature map and a third feature map from the first feature map by different feature channels; wherein the second feature map and the third feature map have different feature levels; extracting features from the second feature map by convolution kernels of different scales to obtain a fourth feature map; extracting features from the third feature map by convolution kernels of different scales to obtain a fifth feature map; and obtaining the second detection result based on the fourth feature map and the fifth feature map.

[0008] In one embodiment, the detecting of the display interface by a second detection network to obtain a second detection result comprises: detecting the display interface by a trained target detection model to obtain the second detection result; wherein the target detection model comprises at least: a backbone network, comprising a first branch and a second branch, the first branch being used to extract a second feature map from the first feature map, and the second branch being used to extract a third feature map from the first feature map; a selective kernel convolution network, connected to the output end of the backbone network, and used to perform multi-branch convolution on the third feature and the second feature, respectively, to obtain a fourth feature map and a fifth feature map; wherein the sizes of the convolution kernels of different branches are different; a detection network, connected to the output end of the selective kernel convolution network, and used to determine the second detection result based on the fourth feature map and the fifth feature map.

[0009] In one embodiment, the display interface is an interface in an application or a system interface.

[0010] In one embodiment, the display abnormality includes at least: text or picture overlap; image missing; overexposure; patches; white screen; black screen; flowery screen; and / or abnormal control position.

[0011] According to a second aspect of an embodiment of the present disclosure, there is provided an abnormality detection device, comprising: an acquisition module for acquiring a display interface to be detected; and a detection module for determining whether a display abnormality occurs on the display interface through different detection networks; wherein different detection networks have different detection effects.

[0012] According to a third aspect of the embodiments of the present disclosure, a terminal device is provided, including:

[0013] A processor and a memory for storing executable instructions that can be run on the processor, wherein: when the processor is used to run the executable instructions, the executable instructions execute the method described in any one of the above embodiments.

[0014] According to a fourth aspect of the embodiments of the present disclosure, a non-temporary computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the method described in any of the above embodiments is implemented.

[0015] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:

[0016] The solution of the disclosed embodiment detects whether the display interface shows abnormalities through multiple detection networks, and combines the detection results of multiple detection networks, thereby improving the efficiency and accuracy of detecting display interface abnormalities, realizing automated detection without manual operation, having strong versatility, and improving the detection effect.

[0017] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0019] Figure 1 is a schematic diagram showing an abnormality detection method according to an exemplary embodiment;

[0020] Figure 2 is a schematic diagram showing an abnormality detection method according to an exemplary embodiment;

[0021] Figure 3 is a schematic diagram showing another method of detecting anomalies according to an exemplary embodiment;

[0022] Figure 4 is a schematic diagram showing another method of detecting anomalies according to an exemplary embodiment;

[0023] Figure 5 is another detection schematic diagram shown according to an exemplary embodiment;

[0024] Figure 6 is a partial schematic diagram of a target detection model according to an exemplary embodiment;

[0025] Figure 7 is a partial schematic diagram of another target detection model according to an exemplary embodiment;

[0026] Figure 8 is a schematic diagram of an abnormality detection device according to an exemplary embodiment;

[0027] Fig. 9is a schematic diagram of another detection according to an exemplary embodiment;

[0028] Fig.10 The present invention is a block diagram of a terminal device according to an exemplary embodiment. DETAILED DESCRIPTION

[0029] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices consistent with some aspects of the present disclosure as detailed in the appended claims.

[0030] refer to Figure 1 , is a schematic diagram of an anomaly detection method, the method comprising:

[0031] S100: Acquire a display interface to be detected.

[0032] S200: Determine whether a display abnormality occurs on a display interface through different detection networks; wherein different detection networks have different detection effects.

[0033] The method can be applied to terminal devices and servers. The terminal devices may include mobile terminals and fixed terminals. The mobile terminals may be mobile phones, tablet computers, wearable devices, vehicle-mounted central control devices, smart homes, and other devices with display screens and installed applications.

[0034] Terminal devices usually have operating systems and various applications installed. The operating system has a desktop interface, and the desktop has various installed applications. Each application can be displayed through the desktop. Each application has a different display interface, and the interfaces in different applications may be different. The interface in an application usually has multiple controls, and different controls have different functions. Different display interfaces can be switched according to different controls.

[0035] For example, the type of application is not limited, and may be any type of application, for example, instant messaging, game, video playback, map, or other types of applications.

[0036] Exemplarily, any interface can be used as the display interface to be detected.

[0037] Exemplarily, the display interface to be detected may be a screenshot of the interface.

[0038] When the terminal is running an application, the terminal can display the interface in the application and save each interface by taking a screenshot, so that the display interface can be obtained. The display interface to be detected can be one or more or all of them.

[0039] After obtaining the display interface to be detected, the display interface to be detected is detected through different detection networks, and then it is determined whether the display interface to be detected has display abnormalities based on the detection results of these detection networks. These detection networks are all trained display interface abnormality detection networks. The network structures of different detection networks may be different, the training processes may be different, the samples used in training may be different, and the detection effects may also be different. The detection effect may at least include the detected area and detection accuracy, etc., and may also include the method of extracting the features of the display interface and the method of processing the extracted feature map of the display interface.

[0040] Exemplarily, when different detection networks detect the display interface, the display interface may be processed by different detection algorithms.

[0041] For example, the first detection network may process the display interface through a first detection algorithm, the second detection network may process the display interface through a second detection algorithm, and the third detection network may process the display interface through a third detection algorithm.

[0042] The structure of each detection network is not limited, and the detection process of the display interface is not limited. It can be determined according to the actual use requirements, and it is sufficient to detect whether the display interface is abnormal. These detection networks can all be trained by training sample sets, and the training sample sets can include positive samples and labels corresponding to the positive samples, and / or negative samples and labels corresponding to the negative samples. The training methods can include supervised learning, semi-supervised learning, and unsupervised learning.

[0043] Exemplarily, the display abnormality includes at least: text or picture overlap, image missing, overexposure, patches, white screen, black screen, distorted screen and / or abnormal control position, etc. Of course, there may be other display abnormalities.

[0044] Exemplarily, the negative samples in the training sample set of each detection network include at least one of the above display anomalies. Different negative samples have different display anomalies.

[0045] Each detection network can detect the display interface to be detected and obtain a detection result respectively. The detection results of multiple detection networks are combined to determine whether the display of the display interface is abnormal. The display interface is detected by multiple different detection networks. Different detection networks have different focuses on the detection of the display interface, and the detection areas may be different, and the detection accuracy is also different. In this way, the results of the detection of multiple different detection networks can be combined to determine whether the display interface has display abnormalities, thereby improving the efficiency and accuracy of the detection.

[0046] In one embodiment, reference Figure 2 , is a schematic diagram of detecting anomalies, Figure 3 , is another schematic diagram of detecting anomalies, combined with Figure 2 and Figure 3 , S200, detecting whether a display abnormality occurs on the display interface, including:

[0047] S201, detecting a display interface through a first detection network to obtain a first detection result.

[0048] S202, when the first detection result shows that the display is normal, the display interface is detected through a second detection network to obtain a second detection result.

[0049] S203: Determine whether there is a display abnormality on the display interface according to the second detection result.

[0050] The first detection network may be a convolutional neural network, which performs convolution processing on the display interface to extract a feature map of the display interface. The information output by the first detection network is the result of whether the display interface shows abnormality.

[0051] refer to Figure 3 , the first detection network may include Figure 3 The multi-layer convolutional network (Multi-Layer-CNN) in the embodiment may include multiple convolutional layers, and the display interface is detected to see if there is a display abnormality through the multi-layer convolutional network. The multi-layer convolutional network may be a trained multi-layer convolutional network model.

[0052] Of course, the first detection network may also be other network models that can detect whether the display interface is abnormal.

[0053] Exemplarily, the second detection network is different from the first detection network. For example, the structure of the second detection network is different from that of the first detection network, and the second detection network extracts features of the display interface in a different manner.

[0054] The display interface is first detected through the first detection network to obtain a first detection result. If the first detection result shows that the display is normal, the display interface is detected through the second detection network for a second detection to obtain a second detection result, and then it is determined whether the display interface is abnormal according to the second detection result of the second detection. That is, the display interface is detected for a second time to obtain a second detection result, and it is determined whether the display interface is abnormal according to the second detection result.

[0055] Exemplarily, the first detection result may be 1 or 0, where 1 and 0 represent different results, for example, 1 represents normal display, and 0 represents abnormal display.

[0056] Exemplarily, the second detection result may be 1 or 0, where 1 and 0 represent different results, for example, 1 represents normal display, and 0 represents abnormal display.

[0057] Exemplarily, the second detection network may be a target detection network based on deep learning, such as a detection network based on a YOLO model.

[0058] The process of detecting the display interface through the second detection belongs to the process of performing a secondary detection on the second display interface. The second detection result can indicate whether there is a display abnormality on the display interface.

[0059] When the second detection result is that the display is normal, it is determined that the display interface displays normally and there is no abnormality. The first display interface is detected by two different detection networks, the first detection network and the second detection network. When the detection results of both methods are normal, it means that the display interface displays normally. Since the two detection networks have different structures, different detection methods for the display interface, different extracted features, and different detection focuses, there will be differences in the detection results. Combining the two detection results improves the accuracy of the detection, and can more accurately determine whether there is a display abnormality in the display interface.

[0060] In one embodiment, reference Figure 4 , is another schematic diagram of detecting anomalies, combined with Figure 3 and Figure 4 , S200, detecting whether a display abnormality occurs on the display interface, including:

[0061] S204, when the first detection result is that the display is abnormal, the display interface is detected through a third detection network to obtain a third detection result.

[0062] S205: Determine whether a display abnormality occurs on the display interface according to the second detection result and the third detection result.

[0063] Among them, the first detection network, the second detection network and the third detection network are all different.

[0064] In this embodiment, when detecting whether the display interface displays abnormality, the display interface is detected through a variety of different detection networks, thereby improving detection accuracy and reducing missed detection situations.

[0065] The first detection network, the second detection network and the third detection network may be detected in accordance with different target detection algorithms, and different target detection algorithms have different detection accuracy.

[0066] The third detection network can be a detection network based on a convolutional neural network, a detection network obtained by improving the convolutional neural network on the basis of the convolutional neural network, such as a target detection network based on R-CNN, a target detection network based on Fast-RCNN, and a target detection network based on Faster-RCNN.

[0067] The structure of the third detection network is different from that of the second detection network, and the features of the extracted display interface are different from those of the second detection network, and the detection accuracy of display anomalies of target objects of different sizes or colors in the display interface is different.

[0068] Exemplarily, the first detection network, the second detection network and the third detection network are respectively trained target detection network models, which are trained based on abnormal samples and normal samples. The display interface to be detected can be used as the input of these detection network models, and the detection result of whether the display interface shows abnormality can be obtained based on the output of these network models.

[0069] Exemplarily, when the first detection network, the second detection network, and the third detection network detect the display interface, the detection accuracy for target objects of different sizes in the display interface is different.

[0070] Of course, the first detection network, the second detection network and the third detection network may also be other detection networks. In short, the detection networks have different network structures and can detect the display interface based on different target detection algorithms or network models.

[0071] like Figure 3 The process of combining multiple detection networks to determine whether the display interface displays abnormalities belongs to the process of detecting display abnormalities of the display interface by fusing multiple models.

[0072] If the first detection result is abnormal, it means that the display interface may have a display abnormality, and the third detection result of the third detection network is combined to determine whether the display interface is abnormal. For example, the second detection result obtained by the second detection network and the third detection result obtained by the third detection network are used to determine whether the display interface is displayed normally.

[0073] Exemplarily, the third detection result may be 1 or 0, where 1 and 0 represent different results, for example, 1 represents normal display, and 0 represents abnormal display.

[0074] In the case where the first detection result is a display abnormality, the display interface is also detected through the second detection network to obtain a second detection result, and the display interface is also detected through the third detection network to obtain a third detection result, and then the second detection result and the third detection result are combined to determine whether the display interface has a display abnormality. In this way, based on the first detection result, the detection results of the two detection networks can be combined again to determine whether the display interface is abnormal. Since different detection networks may have different detection focuses and different accuracy when detecting the display interface, the accuracy of detecting whether the display interface is abnormal is improved according to the detection results of multiple detection networks.

[0075] Exemplarily, the second detection network includes Figure 6 and Figure 7 The corresponding structure.

[0076] Exemplarily, the second test result and the third test result have respective weights, and the weights can be determined according to actual needs. For example, the weight of the second test result is greater than the weight of the third test result, the weight of the second test result is x=0.6, and the weight of the third test result is y=0.4. Figure 3 shown.

[0077] Exemplarily, the first detection result, the second detection result and the third detection result are numerical values, and the target detection result can be determined based on the second detection result and the third detection result and their corresponding weights. The target detection result is also in numerical form. Whether the display interface displays an abnormality is determined based on the target detection result and a preset threshold.

[0078] Exemplarily, the first detection result, the second detection result, the third detection result and the target detection result may all be values ​​between 0 and 1.

[0079] When the target detection result is greater than the preset threshold, it is determined that the display interface is abnormal. When the target detection result is less than the preset threshold, it is determined that the display interface is normal.

[0080] In one embodiment, the second detection model includes a target detection model, and the structure of the target detection model may include a subsequent Figure 6 and Figure 7 The structure of the corresponding embodiment.

[0081] In one embodiment, reference Figure 5 , is another detection schematic diagram, wherein the display interface is detected by a second detection network to obtain a second detection result, which at least includes:

[0082] S10, extracting a first feature map of the display interface.

[0083] S20, extracting a second feature map and a third feature map from the first feature map through different feature channels; wherein the second feature map and the third feature map have different feature levels.

[0084] S30, extracting features from the second feature map by using convolution kernels of different scales to obtain a fourth feature map.

[0085] S40, extracting features from the third feature map by using convolution kernels of different scales to obtain a fifth feature map.

[0086] S50, determining a second detection result according to the fourth feature map and the fifth feature map.

[0087] For S10, in this embodiment, when detecting whether the display interface is abnormal, the display interface can be first acquired, and then the features of the display interface can be extracted to obtain a first feature map. The method of extracting the first feature map is not limited, and can be extracted by a corresponding feature extraction algorithm, such as a convolution-based feature extraction algorithm.

[0088] For S20, after obtaining the first feature map, extract the second feature map and the third feature map from the first feature map through different feature channels, and the feature levels of the second feature map and the third feature map are different. Extract features from the first feature map through different feature channels to obtain the second feature map and the third feature map, respectively. The second feature map is a high-level feature, and the third feature map is a low-level feature; or, the third feature map is a high-level feature, and the third feature map is a high-level feature. High-level features include high-level semantic features, and low-level features include low-level detail features.

[0089] Exemplarily, the low-level features here may be local feature information, and the high-level features may be global feature information.

[0090] In this way, features at different feature levels can be extracted from the first feature map. While extracting high-level features, more low-level detail features can be retained, thereby improving the receptive field and feature extraction capabilities, thereby improving the feature representation capability, reducing gradient vanishing and information loss, and combining features at different feature levels to facilitate improving the accuracy and efficiency of target detection.

[0091] For S30, after obtaining the second feature map, feature extraction is performed on the second feature map through convolution kernels of different scales to obtain a fourth feature map.

[0092] The second feature map can be subjected to multi-branch convolution, and the size of the convolution kernel of each branch is different, so the corresponding receptive field is different, and the features extracted by each convolution branch are also different. Then, the features obtained by each convolution branch are fused to obtain the fourth feature map. The fusion method can include dimensional changes, such as dimensionality increase and dimensionality reduction, and then normalization and other operations.

[0093] For S40, after the third feature map is obtained, features are extracted from the third feature map using convolution kernels of different scales to obtain a fifth feature map.

[0094] The third feature map can be subjected to multi-branch convolution, and the size of the convolution kernel of each branch is different, so the corresponding receptive field is different, and the features extracted by each convolution branch are also different. Then, the features obtained by each convolution branch are fused to obtain the fifth feature map. The fusion method can include dimensional changes, such as dimensionality increase and dimensionality reduction, and then normalization and other operations.

[0095] The fourth feature map is obtained by selectively applying convolution kernels of different scales to extract features from the second feature map, and the fifth feature map is obtained by applying convolution kernels of different scales to extract features from the third feature map. Due to the different sizes of the convolution kernels, the scales of the extracted features are different, so that multi-scale features are obtained at different levels, the feature representation capability is improved, and thus the accuracy of target detection is improved.

[0096] For S50, after obtaining the fourth characteristic map and the fifth characteristic map, a second detection result is obtained according to the fourth characteristic map and the fifth characteristic map, so as to determine whether a display abnormality occurs on the display interface.

[0097] The method in this embodiment can perform multiple different feature extractions on the feature map of the display interface to obtain feature maps at different levels and scales, and then determine whether the display interface displays abnormalities. While retaining the global features, more local features are also extracted, thereby improving the precision and accuracy of detection and reducing the false detection rate.

[0098] In one embodiment, S203, detecting the display interface through a second detection network to obtain a second detection result includes: detecting the display interface through a trained target detection model to obtain the second detection result.

[0099] refer to Figure 6 , is a partial schematic diagram of a target detection model, the target detection model at least includes:

[0100] The backbone network L includes at least a first branch L1 and a second branch L2. The first branch L1 is used to extract a second feature map from the first feature map, and the second branch L2 is used to extract a third feature map from the first feature map. The backbone network L can implement the function of S20. For example, the first branch L1 can extract high-level features of the first feature map, such as high-level features, and the second branch L2 can extract low-level features of the first feature map, such as local features.

[0101] Exemplarily, the object detection model at least includes: a first backbone network. The first branch L1 in the first backbone network can include multiple parts, such as a convolutional unit (Conv), a residual unit, and a fusion unit (CBL), etc. The second branch L2 can include a convolutional unit (Conv). The number of residual units can be x, where x is a positive integer.

[0102] Exemplarily, refer to Figure 7 , which is a schematic structural diagram of another object detection model. The object detection model at least includes: a second backbone network. The first branch L1 in the second backbone network can include a convolutional unit (Conv) and multiple fusion units (CBL), and the second branch L2 can include a convolutional unit (Conv).

[0103] Exemplarily, the object detection model at least includes: a first backbone network and a second backbone network. The connection relationship between the first backbone network and the second backbone network can be determined according to actual needs. For example, the output end of the first backbone network and the input end of the second backbone network are directly or indirectly connected.

[0104] Exemplarily, the object detection model can include at least one backbone network L.

[0105] Figure 6 and Figure 7 The CSP shown in and is the backbone network, which is used to represent the cross-stage partial connection structure (Cross-Stage-Partial, CSP).

[0106] The selective kernel convolutional network K is connected to the output end of the backbone network L and is used to perform multi-branch convolution on the third feature and the second feature respectively to obtain a fourth feature map and a fifth feature map. Among them, the sizes of the convolutional kernels of different branches are different.

[0107] The selective kernel convolutional network K can implement the processes of S30 and S40 to obtain a fourth feature map and a fifth feature map.

[0108] Exemplarily, the selective kernel convolutional network K can be a Selective Kernel Network, that is, an SKNet convolutional network.

[0109] Exemplarily, the selective kernel convolution network K is connected to the output end of the first branch L1 and the output end of the second branch L2 respectively.

[0110] The detection network W is connected to the output end of the selective kernel convolution network, and is used to output a second detection result indicating whether the display interface has a display abnormality according to the fourth feature map and the fifth feature map. The detection network W can implement the function of S50 and detect display abnormalities on the display interface. The structure of the detection network W is not limited, for example, it includes a convolution unit, a pooling unit and / or a fully connected unit.

[0111] In one embodiment, the target detection network may also include: Figure 6 and Figure 7 The normalization unit (Batch Normalization, BN), correction unit (Leaky relu), feature fusion unit (concat), etc. are shown.

[0112] Figure 6 and Figure 7 The target detection model shown in the figure can be applied to the YOLO-V5 model. In the YOLO-V5 model structure, a selective kernel convolution network K is added to the CSP structure to obtain an updated CSP structure. The updated CSP structure is Figure 6 and Figure 7 The structure except the detection network W is shown in FIG.

[0113] The SKNet network based on the dynamic attention mechanism is introduced into the YOLO-V5 model, and the selective kernel convolution network based on the attention mechanism is added to the CSP structure in the Yolov5 model, thereby improving the accuracy of the target detection model in detecting anomalies.

[0114] Reference Table 1:

[0115] Table 1

[0116]

[0117] The data shown in Table 1 is the data information of the network model used by the first detection network, such as Multi-Layer-CNN. 5 and 3 in column A indicate that the training samples include 3 or 5 different samples showing abnormalities, 7_3 indicates that the ratio of the training set to the validation set (test set) is 7:3, and 8_2 also indicates that the ratio of the training set to the validation set (test set) is 8:2. The total number of samples is 1000, and F1 is an indicator used to comprehensively evaluate the performance of the classification model, which combines precision and recall. It can be concluded from the data in Table 1 that the highest accuracy is 90.3% according to the sample type and ratio in row 7.

[0118] However, the model is prone to a high misjudgment rate for normal samples. Further analysis shows that the abnormal samples have smaller features and the model is not easy to fit and learn completely.

[0119] Refer to Table 2:

[0120] Table 2

[0121]

[0122] The data shown in Table 2 is the data information of the training and testing of the detection model used by the third detection network, for example, the Faster-RCNN model. Column A in Table 2 is the number of training samples. From the data in Table 2, it can be concluded that the model works best when the number of training is about 16,000 times. The second average accuracy of the test after the model training is completed is 90.7. The first average accuracy is the test data during the training process.

[0123] Refer to Table 3:

[0124] Table 3

[0125]

[0126] YOLOV5 still has errors in identifying samples with very small feature information, that is, the detection effect of low-level feature information is poor.

[0127] The data shown in Table 3 is the network model used by the second detection network. The network model includes Figure 5 and Figure 6 The training data and test data of the target detection model shown in Figure 1 are as follows. For example, to improve YOLOV5, the attention mechanism network SKNet is added to the CSP structure in Figure 2. Rows 2-10 indicate the number of training times. According to the data in Table 3, as the number of training times increases, the performance of the model becomes higher and higher, and the precision, recall rate, and average precision will all increase. The accuracy corresponding to row 10 is 92.4%.

[0128] By comparing Table 1, Table 2 and Table 3, the network model used by the second detection network has the best detection effect. The display interface is detected by combining the model used by the first detection network, the model used by the second detection network and the model used by the third detection network, which improves the detection accuracy of whether the display interface shows abnormalities.

[0129] In another embodiment, the first interface is an interface in an application or a system interface.

[0130] In one embodiment, reference Figure 8 , is a schematic diagram of an abnormality detection device, the device comprising:

[0131] Acquisition module 1, acquiring the display interface to be detected;

[0132] The detection module 2 is used to determine whether the display interface has display abnormality through different detection networks; wherein different detection networks have different detection effects.

[0133] In one embodiment, the detection module 2 includes:

[0134] A first detection unit, configured to detect the display interface through a first detection network to obtain a first detection result;

[0135] A second detection unit, configured to detect the display interface through a second detection network to obtain a second detection result when the first detection result is that the display is normal;

[0136] A third detection unit, configured to determine whether the display interface has a display abnormality according to the second detection result;

[0137] The first detection network is different from the second detection network.

[0138] In one embodiment, the detection module 2 further includes:

[0139] a fourth detection unit, configured to detect the display interface through a third detection network to obtain a third detection result when the first detection result is a display abnormality;

[0140] a fifth detection unit, configured to determine whether a display abnormality occurs on the display interface according to the second detection result and the third detection result;

[0141] The third detection network is different from the first detection network and the second detection network.

[0142] In one embodiment, the second detection unit comprises:

[0143] A first extraction subunit, used to extract a first feature map of the display interface;

[0144] A second extraction subunit is used to extract a second feature map and a third feature map from the first feature map through different feature channels; wherein the second feature map and the third feature map have different feature levels;

[0145] a third extraction subunit, configured to perform feature extraction on the second feature map by using convolution kernels of different scales to obtain a fourth feature map;

[0146] a fourth extraction subunit, configured to perform feature extraction on the third feature map by using convolution kernels of different scales to obtain a fifth feature map;

[0147] The detection subunit is used to determine whether a display abnormality occurs on the display interface according to the fourth characteristic map and the fifth characteristic map.

[0148] In one embodiment, the second detection unit is used to: detect whether the display interface has display abnormality through the trained target detection model;

[0149] Wherein, the target detection model at least includes:

[0150] A backbone network, comprising a first branch and a second branch, wherein the first branch is used to extract a second feature map from the first feature map, and the second branch is used to extract a third feature map from the first feature map;

[0151] A selective kernel convolution network, connected to the output end of the backbone network, for performing multi-branch convolution on the third feature and the second feature respectively to obtain a fourth feature map and a fifth feature map; wherein the sizes of the convolution kernels of different branches are different;

[0152] A detection network is connected to an output end of the selective kernel convolutional network, and is used to obtain a second detection result according to the fourth feature map and the fifth feature map.

[0153] In one embodiment, the first interface is an interface in an application or a system interface;

[0154] In one embodiment, the display abnormality at least includes:

[0155] Overlapping text or images;

[0156] Image missing;

[0157] Overexposure;

[0158] Plaque;

[0159] White screen;

[0160] Black screen;

[0161] Flower screen;

[0162] and / or, the controls are located abnormally.

[0163] refer to Fig. 9 , is a schematic diagram of another detection, combined with Figure 3 and Fig. 9 ,

[0164] With the rapid development of mobile APPs, various UI display problems have also frequently occurred, such as abnormal page layout, black and white screen, flowery screen, green screen, etc. Since the UI layout has small feature information and many categories, even the human eye cannot judge it well.

[0165] The detection method provided in this embodiment includes:

[0166] The server receives the display interface to be detected, including the sample frame, and then detects the display interface. First, the display interface is detected by a multi-layer convolutional neural network to obtain a first detection result, which includes "normal" and "abnormal". If an abnormality occurs, continue to use the second detection network and the third detection network to detect the display interface, for example, by combining the improved Yolov5-attention and Fast-RCNN detection networks to determine whether the display interface displays an abnormality, respectively obtain the corresponding second detection result and the third detection result, and then determine whether the display interface displays an abnormality according to the second detection result and the third detection result. It can be concluded from Tables 1 to 3 that the detection effect of the second detection network is better than that of the third detection network, so the weight X of the second detection result is given, and the weight Y of the third detection result is given, X>Y, and finally the target detection result is obtained. The distribution interval of the target detection result is determined according to the target detection result and the preset threshold, so as to determine whether the display interface displays an abnormality. For example, the target detection result is 0 or 1 (1 represents normal, 0 represents abnormal). At the same time, in the above process, if the display interface is detected to be normal through the first detection network, the second detection network, such as the Yolov5-attention network, is directly used to secondarily judge whether the display interface is normal or abnormal and return the result.

[0167] Fig. 9 In the process, the display interface to be detected is read, including sample frames image1, image2, image3, image4, etc., and then the above UI detection method is used to display whether the display interface is abnormal. The detection algorithm may include layout detection and display abnormality detection. The results of the two are used to determine whether an abnormality occurs. If an abnormality occurs, the APK will be notified and the Ui information will be captured in real time. If normal, the model will continue with the next round of detection and recognition.

[0168] In order to verify the optimal effect of our algorithm model, the comparison of the experimental data in Tables 1 to 3 shows that the second detection network model has a higher detection accuracy. The data in Table 1 is the data of the first detection network, for example, the MultilayerCNN detection network model, for abnormal layout recognition, when EPOCH is 1000, the category ratio is 3 categories, and the ratio of training set to validation set is 8:2, the detection network model achieves good experimental results, reaching more than 90% on the training set and validation set. Refer to Table 1.

[0169] Refer to Table 2, which is the data of the third detection network, including the data of the Fast-RCNN detection network. The detection accuracy of the detection network model can reach 90.7%.

[0170] Referring to Table 3, the first detection network model is prone to a high misjudgment rate for normal samples. Therefore, further analysis shows that the abnormal sample features are small, and the model is not easy to fit and learn completely. In order to solve this problem, the Yolov5 target detection algorithm was adopted: first, by marking the abnormal sample data, and then training and predicting the model, the detection network model has significantly improved the recognition rate of positive and negative samples, and can be recognized normally. However, in the actual test process, the model still has errors in the recognition of samples with very small feature information. In order to solve this problem of inaccurate recognition of local feature information, the network structure of the existing Yolov5 algorithm is improved, and the attention mechanism is added to the CSP structure of Yolo. The dynamic attention mechanism SKNet is introduced to improve the local feature information recognition ability of the model. The experimental results show that the recognition detection rate of the improved detection network Yolov5-attention reaches 92.4%, which is significantly improved compared with Multilayer CNN.

[0171] By integrating Yolov5-attention and Fast-RCNN for joint judgment and detection, and improving the existing Yolov5, and adding the attention structure to the CSP structure, the local feature information of the image can be fully learned and the accuracy of the model can be improved.

[0172] Based on the multi-model fusion method of deep learning, the existing Yolov5 algorithm is improved and the attention mechanism is added to solve the problem that the sample frame feature information is small and difficult to identify. The improved algorithm is combined with the classic algorithm Fast-RCNN to make a joint judgment and set different penalty coefficients X and Y, X>Y. Finally, the abnormality is judged based on the experimental results of the fusion of the two and the empirical threshold. Compared with the traditional method, this method has higher recognition efficiency, higher accuracy and lower misjudgment rate. Starting from the sample feature point information, it still has a good recognition effect for the local abnormal features of a small range of abnormalities.

[0173] This detection method can be used for various mobile UI display detection problems: UI layout detection, black and white screen and flower screen detection, etc. It effectively solves the problems of high detection difficulty and many types in the traditional UI detection process. Our method model has a wider coverage, higher accuracy, and supports abnormal information capture after detection.

[0174] The solution of this embodiment focuses on the mobile UI display field. By improving the existing Yolov5 model and proposing a multi-model fusion UI anomaly detection method (Yolov5-attention combined with Fast-RCNN) for the first time, it can efficiently identify various UI problems and complete problem location in cooperation with R&D and testing, reducing costs by 15%, improving efficiency by 30%, and realizing a full-link closed loop.

[0175] It should be noted that the “first” and “second” in the embodiments of the present disclosure are only for the convenience of description and distinction and have no other specific meanings.

[0176] Fig.10 1 is a block diagram of a terminal device according to an exemplary embodiment. For example, the terminal device may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0177] Reference Fig.10 The terminal device may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0178] The processing component 802 generally controls the overall operation of the terminal device, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0179] The memory 804 is configured to store various types of data to support operations on the terminal device. Examples of such data include instructions for any application or method operating on the terminal device, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0180] The power component 806 provides power to various components of the terminal device. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the terminal device.

[0181] The multimedia component 808 includes a screen that provides an output interface between the terminal device and the user, and the display component may include a screen, which may be located on the lens. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the terminal device is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0182] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the terminal device is in an operation mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0183] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, volume buttons, start buttons, and lock buttons.

[0184] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the terminal device. For example, the sensor assembly 814 can detect the open / closed state of the terminal device, the relative positioning of components, such as the display and keypad of the terminal device, and the sensor assembly 814 can also detect the position change of the terminal device or a component of the terminal device, the presence or absence of user contact with the terminal device, the orientation or acceleration / deceleration of the terminal device, and the temperature change of the terminal device. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0185] The communication component 816 is configured to facilitate wired or wireless communication between the terminal device and other devices. The terminal device can access a wireless network based on a communication standard, such as Wi-Fi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0186] In an exemplary embodiment, the terminal device may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.

[0187] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0188] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An anomaly detection method, It is characterized in that include: Get the display interface to be detected; Determine whether the display interface has display abnormality through different detection networks; wherein different detection networks have different detection effects.

2. The method according to claim 1, It is characterized in that The determining whether the display interface has a display abnormality through different detection networks includes: Detecting the display interface through a first detection network to obtain a first detection result; When the first detection result is that the display is normal, detecting the display interface through a second detection network to obtain a second detection result; Determining whether there is a display abnormality on the display interface according to the second detection result; The first detection network is different from the second detection network.

3. The method according to claim 2, It is characterized in that The determining whether the display interface has a display abnormality through different detection networks also includes: When the first detection result is that the display is abnormal, detecting the display interface through a third detection network to obtain a third detection result; Determining whether a display abnormality occurs on the display interface according to the second detection result and the third detection result; The third detection network is different from the first detection network and the second detection network.

4. The method according to claim 2 or 3, It is characterized in that The detecting the display interface through the second detection network to obtain a second detection result at least includes: Extracting a first feature map of the display interface; Extracting a second feature map and a third feature map from the first feature map through different feature channels; wherein the second feature map and the third feature map have different feature levels; Extract features from the second feature map using convolution kernels of different scales to obtain a fourth feature map; Performing feature extraction on the third feature map using convolution kernels of different scales to obtain a fifth feature map; The second detection result is obtained according to the fourth feature map and the fifth feature map.

5. The method according to claim 4, It is characterized in that The detecting the display interface by the second detection network to obtain the second detection result includes: detecting the display interface by the trained target detection model to obtain the second detection result; Wherein, the target detection model at least includes: A backbone network, comprising a first branch and a second branch, wherein the first branch is used to extract a second feature map from the first feature map, and the second branch is used to extract a third feature map from the first feature map; A selective kernel convolution network, connected to the output end of the backbone network, for performing multi-branch convolution on the third feature and the second feature respectively to obtain a fourth feature map and a fifth feature map; wherein the sizes of the convolution kernels of different branches are different; A detection network is connected to an output end of the selective kernel convolutional network, and is used to determine the second detection result according to the fourth feature map and the fifth feature map.

6. The method according to claim 1, It is characterized in that The display interface is an interface in an application or a system interface.

7. The method according to claim 1, It is characterized in that The display abnormality at least includes: Overlapping text or images; Image missing; Overexposure; Plaque; White screen; Black screen; Flower screen; and / or, the controls are located abnormally.

8. An abnormality detection device, It is characterized in that include: An acquisition module is used to acquire the display interface to be detected; The detection module is used to determine whether the display interface has display abnormality through different detection networks; wherein different detection networks have different detection effects.

9. A terminal device, It is characterized in that include: A processor and a memory for storing executable instructions capable of running on the processor, wherein: When the processor is used to run the executable instructions, the executable instructions execute the method described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, It is characterized in that The non-transitory computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method of any one of claims 1 to 7.