Screen detection method, screen detection model training method, device and computer equipment

By segmenting the screen image and fusing feature data, the problem of reduced resolution in screen detection is solved, and highly accurate anomaly detection is achieved.

CN116310472BActive Publication Date: 2026-05-12ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
Filing Date
2022-11-10
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, reducing the size of the device screen image leads to a decrease in resolution, which affects the accuracy of the detection results.

Method used

By segmenting the screen image to obtain multiple image blocks, extracting and fusing the feature data of the image blocks, avoiding downsizing and maintaining resolution, and using a screen detection model for anomaly detection.

Benefits of technology

It improves the accuracy of screen detection, avoids the loss of detailed information, and enhances detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310472B_ABST
    Figure CN116310472B_ABST
Patent Text Reader

Abstract

The embodiment of the present specification discloses a screen detection method, a screen detection model training method, a device and a computer equipment. The screen detection method comprises: acquiring a screen image of a device; segmenting the screen image to obtain a plurality of image blocks; extracting first feature data of the image blocks; fusing the first feature data of the plurality of image blocks to obtain a first feature data sequence; and performing abnormal detection on the screen of the device according to the first feature data sequence to obtain a detection result. The technical scheme of the embodiment of the present specification can improve the accuracy of screen detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of computer technology, and in particular to a screen detection method, a screen detection model training method, an apparatus, and a computer device. Background Technology

[0002] In some application scenarios, it is necessary to detect anomalies on device screens, such as detecting whether the screen is damaged. The screen can include mobile phone screens, tablet screens, laptop screens, etc. In related technologies, screen images of the device can be acquired. Screen images are often large in size; to reduce computational load and improve detection efficiency, the screen image can be reduced in size. The reduced screen image can then be input into a model to obtain the detection results output by the model.

[0003] However, scaling down the screen image reduces its resolution and affects the accuracy of the detection results. Summary of the Invention

[0004] This specification provides a screen detection method, apparatus, and computer device to improve the accuracy of screen detection. The technical solutions of this specification are as follows.

[0005] A first aspect of the embodiments of this specification provides a screen detection method, including:

[0006] Get the device's screen image;

[0007] The screen image is segmented to obtain multiple image blocks;

[0008] Extract the first feature data of the image patch;

[0009] The first feature data of multiple image patches are fused to obtain a first feature data sequence;

[0010] Based on the first feature data sequence, anomaly detection is performed on the device screen to obtain the detection results.

[0011] A second aspect of the embodiments of this specification provides a screen detection model training method, including:

[0012] Obtain multiple image blocks after segmenting the device's screen image;

[0013] The first feature data of the image patch is obtained through the transformation layer in the screen detection model;

[0014] The fusion layer in the screen detection model fuses multiple first feature data to obtain a first feature data sequence.

[0015] The screen detection model uses a classifier to detect anomalies on the device's screen based on the first feature data sequence.

[0016] Based on the detection results, the parameters of the screen detection model are adjusted to obtain the trained screen detection model.

[0017] A third aspect of the embodiments of this specification provides a screen detection device, comprising:

[0018] The acquisition unit is used to acquire the screen image of the device;

[0019] The segmentation unit is used to segment the screen image into multiple image blocks;

[0020] Extraction unit, used to extract the first feature data of image patch;

[0021] The fusion unit is used to fuse the first feature data of multiple image blocks to obtain a first feature data sequence;

[0022] The detection unit is used to perform anomaly detection on the device screen based on the first feature data sequence and obtain the detection result.

[0023] A fourth aspect of the embodiments of this specification provides a screen detection model training apparatus, comprising:

[0024] The first acquisition unit is used to acquire multiple image blocks after the screen image of the device is segmented;

[0025] The second acquisition unit is used to acquire the first feature data of the image patch through the transformation layer in the screen detection model;

[0026] The fusion unit is used to fuse multiple first feature data through the fusion layer in the screen detection model to obtain a first feature data sequence;

[0027] The detection unit is used to perform anomaly detection on the device screen based on the first feature data sequence using the classifier in the screen detection model.

[0028] The adjustment unit is used to adjust the parameters of the screen detection model based on the detection results to obtain the trained screen detection model.

[0029] A fifth aspect of the embodiments of this specification provides a computer device, including:

[0030] At least one processor;

[0031] A memory storing program instructions configured to be executed by the at least one processor, the program instructions including instructions for performing the methods as described in the first or second aspect.

[0032] The technical solution provided in the embodiments of this specification can acquire a device screen image; segment the screen image to obtain multiple image blocks; extract first feature data from the image blocks; fuse the first feature data of multiple image blocks to obtain a first feature data sequence; and perform anomaly detection on the device screen based on the first feature data sequence to obtain a detection result. By segmenting the screen image into blocks, the screen image resolution remains intact, avoiding the need to scale down the screen image during anomaly detection, thus preventing the loss of detailed information and improving the accuracy of screen detection. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. The drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 This is a flowchart illustrating the screen detection method in the embodiments of this specification;

[0035] Figure 2 This is a schematic diagram illustrating the sliding of a sliding window within a screen image in an embodiment of this specification;

[0036] Figure 3 This is a schematic diagram illustrating the determination of a sub-image in a screen image in an embodiment of this specification;

[0037] Figure 4 This is a schematic diagram of the screen detection model in the embodiments of this specification;

[0038] Figure 5 This is a flowchart illustrating the screen detection model training method in the embodiments of this specification;

[0039] Figure 6 This is a schematic diagram of the screen detection device in the embodiments of this specification;

[0040] Figure 7 This is a schematic diagram of the screen detection model training device in the embodiments of this specification;

[0041] Figure 8 This is a schematic diagram of the structure of the computer device in the embodiments of this specification. Detailed Implementation

[0042] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. The specific embodiments described herein are only used to explain this disclosure, and not to limit this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure are within the scope of protection of this disclosure. In addition, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0043] This specification provides a screen detection system. The screen detection system may include a terminal device and a server. The terminal device may include user-facing devices, including but not limited to smartphones, tablets, laptops, desktop computers, and wearable devices. The server may include backend-facing devices. The server may be a single server or a server cluster containing multiple servers.

[0044] The terminal device may have a camera component, which may include a camera or similar device. Using the camera component, the terminal device can capture images of the screen of the device under test, obtaining a screen image; and can send the screen image to a server. The server can receive the screen image and perform anomaly detection on the screen of the device under test based on the screen image, obtaining a detection result. The anomaly detection may include damage detection, display anomaly detection, etc. The damage detection may include at least one of the following: detecting whether the screen has cracks, detecting whether the screen edges are missing, detecting whether the outer screen of the screen is damaged, detecting whether the inner screen of the screen is damaged, etc. The display anomaly detection may include detecting whether the screen's color display is normal, etc. Furthermore, the server may also send the detection results to the terminal device. The terminal device can receive the detection results.

[0045] The following is a scenario example from this specification. It should be noted that this scenario example is only for better illustrating the technical solutions of the embodiments in this specification and does not constitute an improper limitation on the embodiments of this specification.

[0046] Screen breakage insurance, also known as accidental screen breakage insurance, is a type of insurance that emerged with the widespread use of electronic devices. It reimburses consumers for certain expenses when an electronic device's screen malfunctions. When a consumer purchases insurance for the screen of an electronic device, the screen must undergo an malfunction test to prevent fraudulent claims or insurance policies with pre-existing defects.

[0047] Consumers can capture images of the electronic device's screen using a terminal device. The terminal device can send these images to a server established by an insurance institution. The server can receive the screen images and perform anomaly detection on the screen of the device under test based on the images, obtaining a detection result (hereinafter referred to as the first detection result). If the first detection result indicates a screen anomaly, the server can refuse insurance coverage and send a first notification message to the terminal device. This first notification message indicates insurance coverage failure. The terminal device can receive the first notification message. If the first detection result indicates a normal screen, the server can provide the screen image to an auditor for further anomaly detection using manual review, obtaining a detection result (hereinafter referred to as the second detection result). If the second detection result indicates a screen anomaly, the server can refuse insurance coverage and send a first notification message to the terminal device. The terminal device can receive the first notification message. If the second detection result indicates a normal screen, the server can execute the pre-set insurance coverage steps and send a second notification message to the terminal device. This second notification message indicates successful insurance coverage. The terminal device can receive the second notification message. Auditors can directly input the second detection result into the server. The server can receive the second detection result input by the auditor. Alternatively, the server can send screen images to a designated device. The designated device can receive the screen images; provide the screen images to reviewers; receive the second detection result input by the reviewers; and send the second detection result to the server. The server can receive the second detection result. The designated device may include a device oriented towards the reviewers. By using screen images, fraudulent or pre-existing insurance claims can be filtered out, reducing manual review workload and saving labor costs.

[0048] This specification provides a screen detection method. The screen detection method can be applied to a server in the screen detection system. Please refer to... Figure 1 The screen detection method may include the following steps.

[0049] Step S11: Acquire a screen image of the device.

[0050] In some embodiments, the terminal device can capture screen images of the device under test and send the screen images to a server. The server can receive the screen images. Alternatively, the terminal device can also capture video of the device under test; extract the video frame containing the screen as a screen image from the video; and send the screen image to a server. The server can receive the screen images. Alternatively, the terminal device can also capture video of the device under test and send the video to a server. The server can receive the video; and extract the video frame containing the screen as a screen image from the video. The device under test may include electronic devices with screens such as smartphones, tablets, laptops, televisions, and smart wearable devices. The terminal device can be a terminal device in the screen detection system. The terminal device may have a camera component. The camera component may include a camera, etc. The terminal device can capture screen images and / or video through the camera component.

[0051] Step S13: Segment the screen image to obtain multiple image blocks.

[0052] In some embodiments, by segmenting the screen image, a large image can be reduced to smaller images, resulting in multiple image patches for anomaly detection. This avoids downsampling processing such as downsizing the screen image, thus preventing the loss of detailed information and improving the accuracy of the detection results. Furthermore, since the size of the image patches is smaller than the size of the screen image, the computational load for anomaly detection is reduced, thereby improving the efficiency of anomaly detection.

[0053] The image blocks can be rectangular or circular, etc. The multiple image blocks can have the same size. For example, the image blocks can be rectangular and have the same side length. Or, the image blocks can be circular and have the same radius. Of course, the multiple image blocks can also have different sizes.

[0054] In some embodiments, a screen image can be segmented according to set segmentation parameters to obtain multiple image blocks. The segmentation parameters may include the number of horizontal segments and the number of vertical segments. The number of horizontal segments and the number of vertical segments may be the same or different. For example, the shape of the screen image may be rectangular. The side length of one side of the screen image can be divided by the number of horizontal segments to obtain the side length of one image block; the side length of the other side of the screen image can be divided by the number of vertical segments to obtain the side length of the other image block; the screen image can be segmented according to the side lengths of the image blocks to obtain multiple image blocks.

[0055] In some embodiments, the sliding window can also slide across the screen image in predetermined steps to obtain multiple image blocks. In practical applications, the sliding window can be slid across the screen image in predetermined steps according to a certain sliding strategy. For example, please refer to... Figure 2 The sliding window can be rectangular. The set step size can include a horizontal step size and a vertical step size. Starting from the top left corner of the screen image, the sliding window can be moved across the screen image according to the horizontal and vertical step sizes, moving from left to right horizontally and from top to bottom vertically. Each slide covers an area on the screen image, which can be designated as the area containing an image block. The area containing the image block can be segmented from the screen image to obtain the corresponding image block.

[0056] The step size can be flexibly set according to actual needs. The step size can be equal to the side length of the sliding window. For example, the step size can include a horizontal step size and a vertical step size. The horizontal step size can be equal to the side length of one side of the sliding window, and the vertical step size can be equal to the side length of the other side of the sliding window. This ensures that there is no overlapping area between adjacent image patches. Alternatively, the step size can be less than the side length of the sliding window. For example, the step size can include a horizontal step size and a vertical step size. The horizontal step size can be less than the side length of one side of the sliding window, and the vertical step size can be equal to the side length of the other side of the sliding window. Or, the horizontal step size can be equal to the side length of one side of the sliding window, and the vertical step size can be less than the side length of the other side of the sliding window. Or, the horizontal step size can be less than the side length of one side of the sliding window, and the vertical step size can be less than the side length of the other side of the sliding window. This ensures that there is an overlapping area between adjacent image patches. Of course, the step size can also be greater than the side length of the sliding window.

[0057] In some embodiments, the screen image can be directly segmented to obtain multiple image blocks. Alternatively, please refer to [link to relevant documentation]. Figure 3 Furthermore, it is possible to determine the sub-image containing the screen region within the screen image; the sub-image can be segmented to obtain multiple image blocks. Specifically, the sub-image containing the screen region can be detected based on a border detection model. The border detection model may include neural network models, etc. By determining and segmenting the sub-image within the screen image, interference from other regions outside the screen area can be avoided in anomaly detection. The process of segmenting the sub-image is similar to the process of segmenting the screen image.

[0058] Step S15: Extract the first feature data of the image patch.

[0059] In some embodiments, the first feature data is used to reflect the features of an image patch. The first feature data can be a feature vector or a feature matrix, etc. The dimension of the first feature data can be N. The dimension refers to the number of data elements in the first feature data. N can be 1, 300, 512, or 600, etc. Specifically, the image patch can be flattened to obtain a feature vector as the first feature data. The data elements in the feature vector can include the pixel values ​​of the pixels in the image patch under multiple color channels. For example, the width of the image patch can be represented as width, the height of the image patch can be represented as height, and the number of color channels of the image patch can be represented as channel. Then the dimension of the feature vector can be represented as width × height × channel. Flattening can flatten the image patch, avoiding downsampling of the screen image, thereby avoiding resolution loss and image detail loss caused by downsampling. Of course, other methods can also be used to extract the first feature data of the image patch. For example, the first feature data can also be obtained according to a feature extraction algorithm. The feature extraction algorithm can include the SIFT algorithm, the HOG algorithm, etc. In addition, in practical applications, the first feature data of multiple image patches can be extracted one by one. Alternatively, a parallel approach can be used to extract the first feature data of multiple image patches simultaneously, thereby improving efficiency.

[0060] Step S17: Fuse the first feature data of multiple image blocks to obtain a first feature data sequence.

[0061] In some embodiments, specific mathematical operations such as addition, subtraction, multiplication, and division can be performed on the first feature data of multiple image blocks to obtain a first feature data sequence. The first feature data sequence can be used to determine the detection result.

[0062] In some embodiments, a feature data sequence (hereinafter referred to as the second feature data sequence) can be constructed based on the first feature data of multiple image blocks; the second feature data sequence can be encoded to obtain another feature data sequence (hereinafter referred to as the first feature data sequence). The second feature data sequence may include multiple ordered first feature data. Specifically, the position of each first feature data in the second feature data sequence can be determined; the second feature data sequence can be constructed based on the position of the first feature data. The position of the first feature data can be random. Alternatively, the position of the corresponding first feature data can be determined based on the position of the image block in the screen image. For example, the position of an image block in the screen image can be represented as (m, n). (m, n) indicates that the image block is the image block in the m-th row and n-th column of the screen image. Then the position of the corresponding first feature data can be represented as (m-1)×N+n, where N represents the number of image blocks in each row of the screen image. (m-1)×N+n indicates that the first feature data is the (m-1)×N+n-th first feature data in the second feature data sequence. Specifically, a machine learning model can be used to encode the second feature data sequence. The machine learning model may include a sequence model. The sequence model is suitable for processing data sequences. The machine learning model includes RNN (Recurrent Neural Networks), LSTM (Long Short-Term Memory), GRU (Gate Recurrent Unit), Transformer, etc. By encoding the second feature data sequence, multiple first feature data in the second feature data sequence can be fused to obtain the first feature data sequence. The first feature data sequence is used to determine the detection result. The first feature data sequence may include several ordered second feature data. The second feature data can be feature vectors or feature matrices, etc. The dimension of the second feature data can be the same as or different from that of the first feature data. Furthermore, the number of second feature data in the first feature data sequence can be the same as the number of first feature data in the second feature data sequence. Of course, the number of second feature data in the first feature data sequence can also be different from the number of first feature data in the second feature data sequence. The above process of obtaining the first feature data sequence from the screen image does not perform downsampling, thus avoiding loss of detailed information from the screen image and improving the accuracy of the detection result.

[0063] Step S19: Based on the first feature data sequence, perform anomaly detection on the device screen to obtain the detection result.

[0064] In some embodiments, anomaly detection of the device screen can be performed based on a first feature data sequence to obtain a detection result. The first feature data sequence may include the first feature data sequence obtained through specific mathematical operations in the aforementioned embodiments, or it may also include the first feature data sequence obtained by encoding a second feature data sequence in the aforementioned embodiments. The anomaly detection may include damage detection, display anomaly detection, etc. Damage detection may include at least one of the following: detecting whether the screen has cracks, detecting whether the screen edge is missing, detecting whether the outer screen is damaged, detecting whether the inner screen is damaged, etc. Display anomaly detection may include detecting whether the screen's color display is normal, etc. The detection result can be used to indicate whether the screen is abnormal. For example, the detection result can be a specific type (abnormal, normal, etc.). Alternatively, the detection result may also include result data. The result data may include numerical values, which represent the probability of screen abnormality. Of course, the detection result can also be used to represent anomaly types. For example, the detection result may be a specific type (screen has cracks, screen edge is missing, outer screen is damaged, inner screen is damaged, screen color display is abnormal, etc.). Alternatively, the detection result may also include result data. The result data may include a vector. The vector may include multiple numerical values. The multiple values ​​are used to represent the probability that the screen belongs to multiple exception types.

[0065] In some embodiments, anomaly detection of the device screen may include a variety of implementation methods.

[0066] In the first implementation, a second feature data at a preset position can be obtained from the first feature data sequence; the obtained second feature data can be input into a classifier to obtain a detection result. The preset position can be an empirical value. Alternatively, the preset position can also be obtained through machine learning. For example, the preset position can include the position of the second feature data input to the classifier within the first feature data sequence during the training process (e.g., the training process of a screen detection model). For example, if the preset position is 1, then the first second feature data in the first feature data sequence can be input into the classifier. Or, for instance, if the preset position is 2, then the second second feature data in the first feature data sequence can be input into the classifier. The classifier includes support vector machines, convolutional neural networks (CNNs), multi-layer perceptrons (MLPs), etc.

[0067] In the second implementation, considering that determining the detection result based on a single second feature data might lead to low accuracy, multiple preset positions of second feature data can be obtained from the first feature data sequence. These multiple second feature data can be fused, and the fused result can be input into a classifier to obtain the detection result. The preset positions can be empirical values. Alternatively, the preset positions can be obtained through machine learning. For example, the preset positions can include the positions of the second feature data input to the classifier in the first feature data sequence during the training process (e.g., the training process of a screen detection model). The number of preset positions can be 2, 3, 5, etc. This embodiment does not specifically limit the method used to fuse the multiple second feature data. For example, the average, median, weighted average, etc., of the multiple second feature data can be calculated as the fusion result. Furthermore, if the multiple preset positions can include 1 and 2, then the first second feature data x1 and the second second feature data x2 can be obtained from the first feature data sequence; the result can be obtained according to the formula... To perform fusion, preset parameters for τ and π are used.

[0068] In the third implementation, second feature data at multiple preset positions can be obtained from the first feature data sequence; the obtained second feature data can be input into multiple classifiers to obtain multiple sub-detection results; the multiple sub-detection results can be fused to obtain a detection result. The preset positions can be empirical values. Alternatively, the preset positions can also be obtained through machine learning. For example, the preset positions can include the positions of the second feature data input to the classifier in the first feature data sequence during the training process (e.g., the training process of a screen detection model). The number of preset positions can be 2, 3, 5, etc. Each obtained second feature data can be input into a classifier to obtain a corresponding sub-detection result. The sub-detection result is used to indicate whether the screen is abnormal and / or the type of abnormality. This embodiment does not specifically limit the method used to fuse the multiple sub-detection results. For example, the sub-detection result can include result data. The average, median, weighted average, etc., of the multiple sub-detection results can be calculated as the final detection result. Alternatively, the sub-detection result can be a specific type. A vote can be performed on the multiple sub-detection results, and the sub-detection result with the most votes can be used as the final detection result.

[0069] In some embodiments, anomaly detection can also be performed on the device screen using a screen detection model to obtain detection results. The screen detection model may include a deep neural network model, a vision transformer, etc.

[0070] The screen detection model does not have a downsampling structure, so that the process of obtaining the first feature data sequence from the screen image does not involve downsampling, thus avoiding the loss of detailed information of the screen image and improving the accuracy of the detection results.

[0071] Please see Figure 4 The screen detection model may include a transformation layer, a fusion layer, and a classifier. The transformation layer is used to acquire first feature data of image patches. The fusion layer is used to fuse the first feature data of multiple image patches to obtain a first feature data sequence. The fusion layer includes multiple sequentially stacked encoding blocks (EncoderBlocks) with identical structures. The number of encoding blocks is P. P can be 8, 9, 12, etc. Each encoding block includes a normalization layer (Norm), a multi-head attention layer, and a multilayer perceptron (MLP), etc. The normalization layer is used for normalization. The multi-head attention layer includes a linear layer (also called a fully connected layer, hereinafter referred to as the first linear layer) and multiple attention heads. The first linear layer is used to summarize the outputs of the multiple attention heads. The number of attention heads in the multi-head attention layer is H. H can be 3, 4, 6, etc. The multiple attention heads have the same structure. Each attention head may include a linear layer (hereinafter referred to as the second linear layer) and a self-attention layer. The second linear layer is used to perform linear transformations. The self-attention layer is used to perform attention-based computations. The classifier is used to detect anomalies on the screen based on the first feature data sequence, obtaining the detection result. The classifier may include support vector machines, convolutional neural networks (CNNs), MLPs (Multi-Layer Perception), etc. Figure 4 In the middle, symbol It indicates fusion.

[0072] In practical applications, a transformation layer can be used to extract the first feature data of an image patch. A fusion layer can be used to fuse the first feature data of multiple image patches to obtain a first feature data sequence. A classifier can then be used to perform anomaly detection on the screen based on the first feature data sequence to obtain the detection result. In some scenario examples, the image patch can be input into the transformation layer to flatten it and obtain the first feature data. A second feature data sequence can be constructed based on the first feature data of multiple image patches. The second feature data sequence can be input into the fusion layer to obtain the first feature data sequence. The fusion layer may include multiple sequentially stacked encoding modules with identical structures. The second feature data sequence can be input to multiple encoding modules. The input to the first encoding module can be the second feature data sequence. The input to the intermediate encoding modules can be the output of the previous encoding module. The input to the last encoding module can be the output of the previous encoding module, and the output can be the first feature data sequence. A classifier can then be used to perform anomaly detection on the screen based on the first feature data sequence to obtain the detection result.

[0073] The screen detection model can have one classifier. Second feature data at a preset position in the first feature data sequence can be input into the classifier to obtain a detection result. Alternatively, second feature data at multiple preset positions in the first feature data sequence can be fused; the fused result can then be input into the classifier to obtain a detection result. Of course, the screen detection model can also have multiple classifiers. Second feature data at multiple preset positions in the first feature data sequence can be input into multiple classifiers to obtain multiple sub-detection results; these sub-detection results can then be fused to obtain the final detection result.

[0074] The method described in this specification can acquire a device screen image; segment the screen image to obtain multiple image blocks; extract first feature data from the image blocks; fuse the first feature data of the multiple image blocks to obtain a first feature data sequence; and perform anomaly detection on the device screen based on the first feature data sequence to obtain a detection result. This segmentation method preserves the screen image resolution, avoids downscaling during anomaly detection, prevents loss of detail, and improves the accuracy of screen detection.

[0075] This specification also provides a method for training a screen detection model. The screen detection model may include a transformation layer, a fusion layer, and a classifier. The method can be applied to a server in the screen detection system, or it can be applied to other devices. Please refer to... Figure 5 The method may include the following steps.

[0076] Step S21: Obtain multiple image blocks after segmenting the screen image of the device.

[0077] Step S23: Obtain the first feature data of the image patch through the transformation layer in the screen detection model;

[0078] Step S25: The multiple first feature data are fused through the fusion layer in the screen detection model to obtain the first feature data sequence.

[0079] Step S27: Using the classifier in the screen detection model, perform anomaly detection on the device screen based on the first feature data sequence.

[0080] Step S29: Adjust the parameters of the screen detection model based on the detection results to obtain the trained screen detection model.

[0081] In some embodiments, the specific architecture of the screen detection model can be found in [reference needed]. Figure 1 The description in the corresponding embodiments.

[0082] In some embodiments, a training sample set can be obtained. The training sample set includes one or more training samples. Each training sample includes a screen image and its corresponding label. The label indicates whether the screen corresponding to the screen image is abnormal and / or the type of abnormality. The screen images in the training samples can be segmented to obtain multiple corresponding image patches.

[0083] In some embodiments, image patches can be input into a transformation layer to flatten them, obtaining first feature data. A second feature data sequence can be constructed based on the first feature data of multiple image patches. The second feature data sequence can be input into a fusion layer to obtain a first feature data sequence. An anomaly detection of the screen can be performed using a classifier based on the first feature data sequence to obtain a detection result. In practical applications, the screen detection model can have one classifier. Second feature data at one location in the first feature data sequence can be input into the classifier to obtain a detection result. Alternatively, second feature data at multiple locations in the first feature data sequence can be fused; the fused result can be input into a classifier to obtain a detection result. Of course, the screen detection model can also have multiple classifiers. Second feature data at multiple locations in the first feature data sequence can be input into multiple classifiers to obtain multiple detection results. These locations can be used as preset positions when the screen detection model is used online, for example... Figure 1 The preset position in the corresponding embodiment.

[0084] In some embodiments, loss information can be determined based on the detection results and labels using a loss function; the parameters of the screen detection model can be adjusted based on the loss information to obtain the trained screen detection model. In practical applications, the screen detection model can have one classifier. The number of detection results can be one. Loss information can be determined based on the detection results and labels output by the classifier using a loss function. The loss function can include cross-entropy loss, maximum likelihood loss (MLE), etc. Alternatively, the screen detection model can have multiple classifiers, and the number of detection results can be multiple. The loss function can include multiple function terms. For example, the loss function can be obtained by adding the multiple function terms. The function terms can include cross-entropy loss, maximum likelihood loss (MLE), etc. Each function term corresponds to a classifier. Sub-loss information can be determined based on the detection results and labels output by the classifier using the function term corresponding to that classifier. Mathematical operations such as addition, weighted addition, subtraction, multiplication, etc., can be performed on multiple sub-loss information to obtain the loss information. The model parameters can be adjusted using the backpropagation mechanism based on the loss information.

[0085] The method described in this specification can acquire multiple image blocks after segmenting a device's screen image; it can acquire first feature data of the image blocks through a transformation layer in a screen detection model; it can fuse multiple first feature data into a first feature data sequence through a fusion layer in the screen detection model; it can perform anomaly detection on the device's screen based on the first feature data sequence through a classifier in the screen detection model; and it can adjust the parameters of the screen detection model based on the detection results to obtain a trained screen detection model. This facilitates anomaly detection on the screen using a screen detection model.

[0086] This specification also provides a screen detection device. The screen detection device can be applied to a server in the screen detection system. Please refer to... Figure 6 The screen detection device may include the following units.

[0087] Acquisition unit 31 is used to acquire the screen image of the device;

[0088] The segmentation unit 33 is used to segment the screen image to obtain multiple image blocks;

[0089] Extraction unit 35 is used to extract the first feature data of the image patch;

[0090] Fusion unit 37 is used to fuse the first feature data of multiple image blocks to obtain a first feature data sequence;

[0091] The detection unit 39 is used to perform anomaly detection on the screen of the device based on the first feature data sequence and obtain the detection result.

[0092] This specification also provides a training apparatus in its embodiments. The apparatus is used to train a screen detection model. The screen detection model may include a conversion layer, a fusion layer, and a classifier, etc. The apparatus can be applied to a server in the screen detection system, or it can also be applied to other devices. Please refer to... Figure 7 The device may include the following units.

[0093] The first acquisition unit 41 is used to acquire multiple image blocks after the screen image of the device is segmented;

[0094] The second acquisition unit 43 is used to acquire the first feature data of the image patch through the transformation layer in the screen detection model;

[0095] The fusion unit 45 is used to fuse multiple first feature data through the fusion layer in the screen detection model to obtain a first feature data sequence;

[0096] The detection unit 47 is used to perform anomaly detection on the screen of the device based on the first feature data sequence through the classifier in the screen detection model.

[0097] Adjustment unit 49 is used to adjust the parameters of the screen detection model according to the detection results to obtain the trained screen detection model.

[0098] The following describes an embodiment of the computer device described in this manual. Figure 8 This is a schematic diagram of the hardware structure of the computer device in this embodiment. For example... Figure 8 As shown, the computer device may include one or more (only one is shown in the figure) processors, memory, and transmission modules. Of course, those skilled in the art will understand that... Figure 8 The hardware structure shown is for illustrative purposes only and does not limit the hardware structure of the computer device described above. In practice, the computer device may also include more... Figure 8 Showing more or fewer component units; or, having the same as Figure 8 The different configurations shown.

[0099] The memory may include high-speed random access memory; or it may include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. Of course, the memory may also include remotely accessible network memory. The memory can be used to store program instructions or modules of application software, such as those described in this specification. Figure 1 or Figure 5The program instructions or modules corresponding to the embodiments.

[0100] The processor can be implemented in any suitable manner. For example, the processor can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers, etc. The processor can read and execute program instructions or modules in the memory.

[0101] The transmission module can be used to transmit data via a network, such as the Internet, corporate intranet, local area network, or mobile communication network.

[0102] This specification also provides an embodiment of a computer storage medium. The computer storage medium includes, but is not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), cache, hard disk drive (HDD), memory card, etc. The computer storage medium stores computer program instructions. When the computer program instructions are executed, they implement: this specification. Figure 1 or Figure 5 The program instructions or modules corresponding to the embodiments.

[0103] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0104] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. A computer can be a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0105] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0106] Those skilled in the art will understand that the descriptions of the various embodiments have different focuses, and parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. Furthermore, it is understood that those skilled in the art, after reading this specification, can conceive of any combination of some or all of the embodiments listed in this specification without creative effort, and such combinations are also within the scope of disclosure and protection of this specification.

[0107] Although this specification has been described by way of examples, those skilled in the art will recognize that many variations and modifications are possible with respect to this specification, and it is intended that the appended claims cover such variations and modifications without departing from the spirit of this specification.

Claims

1. A screen detection method, comprising: Get the device's screen image; The screen image is segmented to obtain multiple image blocks. The segmentation of the screen image includes: segmenting the sub-image containing the screen area within the screen image to obtain multiple image blocks. The multiple image blocks are used to perform anomaly detection on the screen of the device within the screen area. Extract the first feature data of the image patch; The first feature data of multiple image patches are fused to obtain a first feature data sequence; Anomaly detection is performed on the device screen based on a first feature data sequence to obtain a detection result. The first feature data sequence includes several ordered second feature data. The anomaly detection on the device screen includes: inputting the second feature data at a preset position in the first feature data sequence into a classifier to obtain a detection result; or, fusing the second feature data at multiple preset positions in the first feature data sequence and inputting the fusion result into a classifier to obtain a detection result; or, inputting the second feature data at multiple preset positions in the first feature data sequence into multiple classifiers to obtain multiple sub-detection results and fusing the multiple sub-detection results to obtain a detection result.

2. The method according to claim 1, wherein segmenting the screen image comprises: The sliding window slides across the screen image in a set step size to obtain multiple image blocks; The set step size is less than the side length of the sliding window.

3. The method according to claim 1, wherein extracting the first feature data of the image patch comprises: The image patch is flattened to obtain the first feature data; The fusion of the first feature data of multiple image patches includes: The position of the corresponding first feature data is determined based on the position of the image block in the screen image; Construct a second feature data sequence based on the position of the first feature data; The second feature data sequence is encoded to obtain the first feature data sequence.

4. The method according to claim 1, wherein extracting the first feature data of the image patch comprises: The first feature data of the image patch is extracted through the transformation layer in the screen detection model; The fusion of the first feature data of multiple image patches includes: The first feature data of multiple image patches are fused through the fusion layer in the screen detection model. The abnormal detection of the device screen includes: The screen detection model uses a classifier to detect anomalies on the device's screen based on the first feature data sequence.

5. A screen detection model training method, comprising: The screen image of the device is segmented into multiple image blocks. The segmentation of the screen image of the device includes: segmenting the sub-image containing the screen area within the screen image. The multiple image blocks are used to perform anomaly detection on the screen of the device within the screen area. The first feature data of the image patch is obtained through the transformation layer in the screen detection model; The fusion layer in the screen detection model fuses multiple first feature data to obtain a first feature data sequence. Anomaly detection of the device screen is performed using a classifier in the screen detection model based on a first feature data sequence, wherein the first feature data sequence includes several ordered second feature data. The anomaly detection of the device screen includes: if the screen detection model includes one classifier, inputting the second feature data at a preset position in the first feature data sequence into the classifier to obtain a detection result; or, fusing the second feature data at multiple preset positions in the first feature data sequence and inputting the fused result into the classifier to obtain a detection result; or, if the screen detection model includes multiple classifiers, inputting the second feature data at multiple preset positions in the first feature data sequence into multiple classifiers to obtain multiple sub-detection results, and fusing the multiple sub-detection results to obtain a final detection result. Based on the detection results, the parameters of the screen detection model are adjusted to obtain the trained screen detection model.

6. A screen detection device, comprising: The acquisition unit is used to acquire the screen image of the device; The segmentation unit is used to segment the screen image to obtain multiple image blocks. The segmentation of the screen image includes: segmenting the sub-image containing the screen area within the screen image to obtain multiple image blocks. The multiple image blocks are used to perform anomaly detection on the screen of the device within the screen area. Extraction unit, used to extract the first feature data of image patch; The fusion unit is used to fuse the first feature data of multiple image blocks to obtain a first feature data sequence; The detection unit is used to perform anomaly detection on the screen of a device based on a first feature data sequence and obtain a detection result. The first feature data sequence includes several ordered second feature data. The anomaly detection on the screen of the device includes: inputting the second feature data at a preset position in the first feature data sequence into a classifier to obtain a detection result; or, fusing the second feature data at multiple preset positions in the first feature data sequence and inputting the fusion result into a classifier to obtain a detection result; or, inputting the second feature data at multiple preset positions in the first feature data sequence into multiple classifiers to obtain multiple sub-detection results and fusing the multiple sub-detection results to obtain a detection result.

7. A screen detection model training device, comprising: The first acquisition unit is used to acquire multiple image blocks after segmenting the screen image of the device. The segmentation of the screen image of the device includes: segmenting the sub-image where the screen area is located within the screen image. The multiple image blocks are used to perform anomaly detection on the screen of the device within the screen area. The second acquisition unit is used to acquire the first feature data of the image patch through the transformation layer in the screen detection model; The fusion unit is used to fuse multiple first feature data through the fusion layer in the screen detection model to obtain a first feature data sequence; The detection unit is used to perform anomaly detection on the device screen based on a first feature data sequence using a classifier in a screen detection model. The first feature data sequence includes several ordered second feature data. The anomaly detection on the device screen includes: when the screen detection model includes a classifier, inputting the second feature data at a preset position in the first feature data sequence into the classifier to obtain a detection result; or, fusing the second feature data at multiple preset positions in the first feature data sequence and inputting the fused result into the classifier to obtain a detection result; or, when the screen detection model includes multiple classifiers, inputting the second feature data at multiple preset positions in the first feature data sequence into multiple classifiers to obtain multiple sub-detection results, and fusing the multiple sub-detection results to obtain a detection result. The adjustment unit is used to adjust the parameters of the screen detection model based on the detection results to obtain the trained screen detection model.

8. A computer device, comprising: At least one processor; A memory storing program instructions configured to be executed by the at least one processor, the program instructions including instructions for performing the method according to any one of claims 1-5.