Display screen anomaly detection method and device, electronic equipment, medium and product
By analyzing the video frame feature map and region feature map of the display screen, and combining temporal changes and structural anomaly intensity, the feature is fused using a nonlinear model, which solves the problem of missing detection of minor anomalies in vehicle display screens, and improves detection accuracy and user experience.
Patent Information
- Application Number
- CN202511555900.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-10-29
AI Technical Summary
Existing technologies struggle to effectively identify minute anomalies such as flickering corners and slight display misalignments on in-vehicle displays, resulting in a high rate of missed detections, impacting user experience and posing potential driving safety hazards.
By acquiring monitoring videos of the display screen, extracting video frame feature maps and dividing them into multiple regional feature maps, analyzing temporal change characteristics and spatial structure characteristics, and using a nonlinear model to fuse the intensity of temporal change and structural anomaly, the recognition rate of screen corner areas and minor display anomalies is improved.
It improves the recognition rate of screen corner areas and minor display anomalies, reduces the false negative rate, enhances user experience, and strengthens driving safety.
Smart Images

Figure CN121033036A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically to methods, devices, electronic devices, media, and products for detecting abnormalities in display screens. Background Technology
[0002] With the rapid increase in the penetration rate of intelligent connected vehicles, the complexity of intelligent cockpit systems has further intensified. Frequent hardware resource allocation during screen rendering and the exponential increase in the amount of information interacting between screens can easily lead to crashes and screen malfunctions in in-vehicle screen systems. These malfunctions not only affect the user experience of the intelligent cockpit but also pose potential driving safety hazards due to display failure.
[0003] Traditional detection methods locate screen anomalies through hardware diagnostics, but rely on manual operation. Existing technologies rely on automated scripts for anomaly detection, but require the development of separate scripts for different vehicle infotainment systems, resulting in poor compatibility. Some virtual machine-based detection methods monitor in-vehicle screen malfunctions, but these have high hardware resource requirements. To address these issues, existing technologies also offer image recognition-based detection methods, using image recognition models to identify screen anomalies. However, these methods often focus on overall screen anomalies, resulting in a high rate of missed detections for minor anomalies such as flickering at screen corners or slight display misalignments. Summary of the Invention
[0004] This invention provides a method, apparatus, electronic device, medium, and product for detecting display screen anomalies, in order to solve the problem that existing technologies easily miss faults or anomalies in the corners or small display areas of the display screen.
[0005] In a first aspect, the present invention provides a method for detecting display screen anomalies, the method comprising: Acquire monitoring video of the display screen and extract video frame feature maps from the monitoring video; Based on the video frame feature map, extract the temporal variation features between consecutive video frames; The video frame feature map is divided into multiple region feature maps, and the spatial structure features of each region feature map are extracted; wherein, each region feature map corresponds one-to-one with a different display area of the display screen. Determine the temporal variation intensity of the temporal variation characteristics, obtain the structural anomaly intensity of each display area based on the spatial structural characteristics, and determine the target display area where the structural anomaly intensity is greater than the anomaly intensity threshold. Based on the intensity of temporal variation and the intensity of structural anomalies in the target display area, the temporal variation features and the spatial structural features corresponding to the target display area are fused. Based on the obtained fused features, the anomaly detection results of the display screen are determined.
[0006] This invention utilizes video frame feature maps from monitored video feeds on display screens to extract temporal variation features between consecutive video frames and spatial structural features corresponding to different display areas of the screen, enabling separate analysis of the spatial structure of different display areas. Next, it analyzes the temporal variation intensity of the temporal variation features and the structural anomaly intensity of each display area to measure the impact of these features on screen anomaly detection. Furthermore, it fuses the spatial structural features of target display areas with high temporal variation and structural anomaly intensities to improve the recognition rate of screen corner areas and minor display anomalies, thus obtaining the final anomaly detection results for the display screen.
[0007] In one optional implementation, the temporal variation features and the spatial structural features corresponding to the target display area are fused based on the intensity of temporal variation and the intensity of structural anomalies in the target display area, including: The intensity of temporal variation and the intensity of structural anomalies in the target display area are input into a pre-trained nonlinear model to obtain fusion coefficients; whereby the fusion coefficients are used to characterize the importance of the temporal variation features. Based on the fusion coefficient, the temporal variation features and the spatial structural features corresponding to the target display area are fused to obtain the fused features.
[0008] This invention inputs the intensity of temporal variation and the intensity of structural anomalies in the target display area into a pre-trained nonlinear model. The nonlinear model then determines the degree of influence of temporal variation features and spatial structural features of the target display area on screen anomaly detection, thereby outputting fusion coefficients. These fusion coefficients allow for the adjustment of the weighting of temporal variation features and spatial structural features of the target display area during feature fusion, improving the adaptability of screen anomaly detection in various scenarios.
[0009] In one optional implementation, the structural anomaly intensity of each display area is obtained based on spatial structural characteristics, including: Based on the brightness and texture features in the spatial structure features, the brightness variance features and texture variance features of each display area are obtained; Determine the brightness variance characteristics of each display area and the brightness difference characteristics between the preset brightness variance, and determine the texture variance characteristics of each display area and the texture difference characteristics between the preset texture variance; The structural anomaly intensity of each display area is calculated based on the brightness difference characteristics and texture difference characteristics.
[0010] This invention utilizes brightness features for anomaly detection, covering core fault types, while texture features are a key dimension describing the structural patterns of screen content. By calculating the variance of brightness and texture features, and then using the resulting brightness and texture difference features, the strength of structural anomalies in each display area is measured. This achieves full-scene coverage while supporting mutual verification, with low computational cost, making it more suitable for the lightweight requirements of screen anomaly detection.
[0011] In one optional implementation, determining the temporal variation intensity of the temporal variation characteristics includes: Determine the root mean square error characteristics of time-series variations; Based on the comparison between the mean squared error characteristics and the preset mean squared error range, the temporal variation intensity of the temporal variation characteristics is calculated.
[0012] This invention captures the fluctuation characteristics of temporal variation features by calculating the mean squared error (MSE) of these features, suppressing isolated noise in individual video frames, avoiding local misjudgments, and with extremely low computational overhead. Furthermore, by comparing the MSE features with a preset MSE range, the intensity of temporal variation features is measured, thereby improving the accuracy of screen anomaly identification by combining the intensity of temporal variation.
[0013] In one optional implementation, based on the video frame feature map, temporal variation features between consecutive video frames are extracted, including: Based on the video frame feature map, the difference features between every two consecutive video frames are calculated; Based on the difference characteristics, the temporal variation characteristics between consecutive video frames are obtained.
[0014] This invention calculates the difference features between every two consecutive video frames, uses these difference features as input for temporal variation features, filters out a large amount of static redundant information in consecutive video frames, reduces the input of irrelevant static redundant information, and improves the robustness of screen anomaly detection.
[0015] In one optional implementation, the anomaly detection result of the display screen is determined based on the obtained fusion features, including: The fused features are compressed to obtain a global feature vector of fixed length; The global feature vector is input into a pre-trained anomaly classification model to obtain anomaly detection results for the display screen; the anomaly detection results include screen flickering, black screen, distorted screen, white screen, and normal screen.
[0016] This invention compresses fused features into a fixed-length global feature vector, aggregates the global information of the fused features, and uses the global feature vector to determine whether the display screen has anomalies such as screen flickering, black screen, distorted screen, or white screen. Moreover, the global feature vector can effectively reduce the feature dimension and avoid overfitting of the pre-trained anomaly classification model when performing anomaly detection.
[0017] In one alternative implementation, the training process of the anomaly classification model includes: Construct a screen image dataset and a pre-defined model structure; Data augmentation operations are performed on the screen image dataset to obtain an augmented image dataset. The data augmentation operations include image occlusion, geometric transformation, lighting adjustment, screen anomaly simulation, and resolution adjustment. Add a category label to each image in the augmented image dataset; the category label corresponds to the anomaly type displayed on the screen. The pre-trained anomaly classification model is obtained by training a pre-defined model structure using an augmented image dataset.
[0018] This invention is based on a constructed screen image dataset. Through data augmentation operations such as image occlusion, geometric transformation, lighting adjustment, screen anomaly simulation, and resolution adjustment, an enhanced image dataset is constructed and category labels are added. This enhanced image dataset is then used for model training to improve the robustness, generalization ability, scene adaptability, and detection accuracy of the anomaly classification model.
[0019] In one optional implementation, the pre-trained anomaly classification model is deployed in the cloud after pruning and quantization; the global feature vector is input into the pre-trained anomaly classification model to obtain the anomaly detection results displayed on the screen, including: Encapsulate the global feature vector into a classification request; By using a pre-defined API, a classification request is sent to the cloud, which then calls a pre-trained anomaly classification model based on the classification request to obtain the anomaly classification result. Receive the anomaly classification results returned from the cloud and obtain the anomaly detection results displayed on the screen.
[0020] This invention achieves lightweight deployment by pruning and quantizing the anomaly classification model and then deploying it in the cloud. This allows the vehicle's infotainment system to obtain the anomaly classification results returned from the cloud by calling the anomaly classification model deployed in the cloud, thereby reducing the latency of screen anomaly detection.
[0021] In a second aspect, the present invention provides a display screen anomaly detection device, the device comprising: The first processing module is used to acquire the monitoring video of the display screen and extract the video frame feature map of the monitoring video; The second processing module is used to extract the temporal variation features between consecutive video frames based on the video frame feature map; The third processing module is used to divide the video frame feature map into multiple region feature maps and extract the spatial structure features of each region feature map; wherein, the region feature map corresponds one-to-one with different display areas of the display screen; The fourth processing module is used to determine the intensity of temporal change characteristics, obtain the structural anomaly intensity of each display area based on spatial structural characteristics, and determine the target display area where the structural anomaly intensity is greater than the anomaly intensity threshold. The fifth processing module is used to fuse the temporal variation features and the spatial structural features corresponding to the target display area based on the intensity of temporal variation and the intensity of structural anomalies in the target display area, and to determine the anomaly detection result of the display screen based on the obtained fused features.
[0022] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the display screen anomaly detection method of the first aspect or any corresponding embodiment described above.
[0023] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the display screen anomaly detection method of the first aspect or any corresponding embodiment described above.
[0024] Fifthly, the present invention provides a computer program product, including computer instructions, which are used to cause a computer to execute the display screen anomaly detection method of the first aspect or any corresponding embodiment described above.
[0025] The beneficial effects of this invention are as follows: This invention utilizes video frame feature maps from monitored video feeds on display screens to extract temporal variation features between consecutive video frames and spatial structural features corresponding to different display areas of the screen, enabling separate analysis of the spatial structure of different display areas. Next, it analyzes the temporal variation intensity of the temporal variation features and the structural anomaly intensity of each display area to measure the impact of these features on screen anomaly detection. Furthermore, it fuses the spatial structural features of target display areas with high temporal variation and structural anomaly intensities to improve the recognition rate of screen corner areas and minor display anomalies, thus obtaining the final anomaly detection results for the display screen. Attached Figure Description
[0026] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention; Figure 2 This is a schematic flowchart of a first method for detecting display screen anomalies according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a second process for a display screen anomaly detection method according to an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the process of constructing an enhanced image dataset according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the training process of an anomaly classification model according to an embodiment of the present invention; Figure 6 This is a model framework diagram of a display screen anomaly detection system according to an embodiment of the present invention; Figure 7 This is a structural block diagram of a display screen anomaly detection device according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0030] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0031] As an optional application scenario of this invention, such as Figure 1 As shown, the display screen anomaly detection system may include at least one terminal device and at least one server. Figure 1 The system is illustrated in the example, which includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through a network 110.
[0032] Specifically, the terminal device can be a smartphone, tablet, laptop, PDA, desktop computer, game console, smart TV, smart wearable device, in-vehicle terminal, VR (Virtual Reality) device, AR (Augmented Reality) device, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranet, local area network, wide area network, mobile communication network, and combinations thereof.
[0033] With the rapid increase in the penetration rate of intelligent connected vehicles, the complexity of intelligent cockpit systems has further intensified. Industry data shows that in-vehicle infotainment system issues have become a focal point of user complaints. In recent years, in-vehicle infotainment system-related malfunctions (including black screens, freezes, and blurry images) have accounted for 9.56% to 15.89% of total vehicle complaints. The reasons for this include the frequent scheduling of the central processing unit (CPU) / graphics processing unit (GPU) during 3D rendering, and the exponential increase in the amount of information exchanged between screens, all of which can easily lead to system crashes. In some models, the black screen failure rate is as high as 24%. Such anomalies not only disrupt the user experience of the intelligent cockpit but also pose potential driving safety hazards due to display failures.
[0034] In response to the aforementioned potential screen malfunctions and the need for automakers to invest in improving the reliability of in-vehicle infotainment systems, the commonly used methods for detecting screen malfunctions currently include the following: (1) Traditional detection method: This method covers multiple aspects such as hardware inspection, power / signal line testing, system and software diagnosis, and uses tools to measure screen voltage, check signals, and read fault codes to locate abnormal problems. However, this method relies on manual operation, is inefficient, is difficult to troubleshoot complex hidden software or compatibility faults, and is not robust enough to scenarios such as complex lighting, screen resolution changes, and dynamic updates of display content, resulting in a high rate of missed detection and false detection.
[0035] (2) Detection method based on automated scripts: The core principle of this method is to interact with the vehicle's native services based on the C-layer interface (i.e., the interface connecting the application layer and the vehicle's native services), draw the specified color by running the script, and then judge the fault based on whether the actual display effect on the screen meets the expectations. However, the developed automated scripts need to be adapted separately for different vehicle systems, resulting in poor compatibility. Moreover, it can only detect abnormal display screens and is difficult to identify occasional software defects.
[0036] (3) Virtual machine-based detection method: This method adopts a dual virtual machine collaborative mode, in which the monitor driver monitors the abnormal pin signals of the screen and then transmits the monitoring information to the screen monitor for display. However, this method has specific requirements for environment and hardware. On the one hand, it relies on the virtual machine environment to build, which has a high operating threshold and high requirements for hardware resources; on the other hand, its detection range is narrow, focusing only on abnormalities of motor control unit, and is suitable for relatively simple fixed detection scenarios.
[0037] (4) Image recognition-based detection method: This method first captures the real-time image of the vehicle screen using a camera, and then inputs the image into a pre-trained image recognition model, such as a Convolutional Neural Network (CNN) or YOLO. The image recognition model extracts and analyzes the image features to identify abnormal problems such as black screen, flickering screen, distorted screen, color distortion, and display misalignment. However, general image recognition models often struggle to balance accuracy and speed, and have a high rate of missing detection for minor anomalies such as flickering at the screen corners and slight display misalignment. They also have weak overall anti-interference capabilities and are easily affected by the external environment.
[0038] This invention provides a method for detecting display screen anomalies. Based on the video frame feature map of the monitored video of the display screen, it extracts the temporal variation features between consecutive video frames and the spatial structure features corresponding to different display areas of the display screen, so as to analyze the spatial structure of different display areas separately. Next, it analyzes the temporal variation intensity of the temporal variation features and the structural anomaly intensity of each display area, and determines the target display area with a high structural anomaly intensity. Then, it uses the temporal variation intensity and the structural anomaly intensity of the target display area to perform feature fusion of the temporal variation features and the spatial structure features of the target display area, thereby improving the recognition rate of screen corner areas and minor display anomalies, and obtaining the display screen anomaly detection result.
[0039] According to an embodiment of the present invention, a method for detecting display screen anomalies is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0040] This embodiment provides a method for detecting abnormal display screens, which can be used in the aforementioned mobile terminals, such as in-vehicle terminals, mobile phones, tablet computers, etc. Figure 2 This is a flowchart of a display screen anomaly detection method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Acquire the monitoring video of the display screen and extract the video frame feature map of the monitoring video.
[0041] Specifically, taking a vehicle-mounted screen as an example, the vehicle-mounted camera monitors the screen to obtain a monitoring video, which consists of a series of continuous video frames. Then, a pre-defined neural network model is used to extract features from each video frame of the monitoring video, extracting the RGB feature map of each video frame to obtain the corresponding video frame feature map.
[0042] In this embodiment, when extracting features from the monitoring video, the preset neural network model used can be an efficient convolutional neural network (EfficientNet), such as the EfficientNet-B3 model. Efficient convolutional neural networks have the advantages of small number of parameters, small amount of computation, and high feature extraction accuracy, and can effectively avoid hardware resource overload, which is consistent with the accuracy and efficiency balance requirements of display screen anomaly detection.
[0043] Step S202: Extract the temporal variation features between consecutive video frames based on the video frame feature map.
[0044] Specifically, efficient convolutional neural network models include multiple channels, each representing a different dimension of color or brightness perceived by the human eye. For example, an RGB feature map has three channels (corresponding to the red (R) channel, green (G) channel, and blue (B) channel), while a video frame feature map includes the video frame feature value corresponding to each spatial location (i.e., pixel coordinates) and each channel. Efficient convolutional neural network models typically utilize the Squeeze-and-Excitation (SE) attention mechanism to address the problem of uneven feature importance across different channels. However, SE attention is a single-frame independent attention mechanism, processing only the static features of a single frame. It cannot distinguish the feature importance of different spatiotemporal regions within the same channel and is highly sensitive to noise, making it susceptible to the influence of noise (irrelevant interference information) when processing data.
[0045] This embodiment calculates the temporal changes of the feature map of continuous video frames at each spatial location (i.e., pixel coordinates) and each channel, extracts the temporal change features between continuous video frames, and utilizes the dynamic correlation between continuous frames to better identify dynamic anomalies of the display screen and improve the detection rate of anomalies such as periodic screen flickering (e.g., sudden brightness changes 3 times per second, which appear normal in a single frame but are only noticeable when comparing consecutive frames).
[0046] Step S203: Divide the video frame feature map into multiple region feature maps and extract the spatial structure features of each region feature map; wherein, the region feature map corresponds one-to-one with different display areas of the display screen.
[0047] Specifically, the display area of a screen includes a border area (the outer area of the screen), a central area (the main display area in the center of the screen), and partitioned display areas (functional display areas such as sidebars, notification bars, navigation bars, and split-screen windows). When a vehicle camera monitors a screen, it inevitably includes physical areas outside the screen (such as the steering wheel and dashboard surrounding the infotainment system), leading to video frame feature maps containing invalid, redundant spatial features. Furthermore, the types of anomalies vary significantly across different display areas; for example, the border area often experiences light leakage / damage, the central area often has screen flickering / dead pixels, and the partitioned display areas often have display misalignment.
[0048] Therefore, in this embodiment, based on the display area of the display screen, the video frames of the monitored video are divided into multiple non-overlapping pixel regions. Each pixel region can correspond one-to-one with the physical area outside the display screen, the border area, the center area, and the partitioned display area within the display screen, and a region mask is generated for each pixel region. The region mask is used to mark and distinguish different pixel regions of the video frame. For example, the region mask corresponding to the physical area outside the display screen is "00", the region mask corresponding to the border area is "01", the region mask corresponding to the center area is "10", and the region mask corresponding to the partitioned display area is "11", etc. The process of generating the region mask can be found in the description of related technologies, and will not be repeated here.
[0049] Furthermore, weights are assigned based on region masks, and the extracted video feature maps are divided into multiple non-overlapping region feature maps. For example, the feature map portion with a region mask of "00" is filtered out (i.e., the weight is 0), so that only the region feature maps corresponding one-to-one with different display areas of the display screen (border area, middle area, partition display area) are extracted, and the background pixel areas (such as dashboard, steering wheel, etc.) in the video feature maps are filtered out.
[0050] In step S203, for each region feature map, the pixel brightness value and texture gradient value at each spatial location (i.e., pixel coordinate) in the region feature map are extracted to obtain the brightness features and texture features of the region feature map, and thus obtain the spatial structure features of the region feature map.
[0051] In some embodiments, a weighted average method can be used to calculate the average brightness of a pixel at each spatial location in the region feature map across the three RGB channels, thus obtaining the pixel brightness value. The texture gradient value can be extracted based on an image edge detection algorithm (such as the Sobel operator). For details, please refer to the description of the relevant technology, which will not be elaborated here.
[0052] This embodiment extracts spatial structure features from the feature maps corresponding to each display area, enhancing the detection adaptability to different display areas. By utilizing the weight allocation of region masks, targeted processing can be applied to different pixel regions of the video frame feature maps. For example, using region masks can directly block background pixels, increasing the model's attention to and calculation of the pixel regions corresponding to the display screen, thereby reducing the amount of invalid pixels processed by more than 50% and improving inference speed by approximately 30%. Thus, through region partitioning and region mask structures, while improving the anomaly detection accuracy of different display areas, computational power can be allocated differentially to avoid high computational consumption across the entire image.
[0053] Step S204: Determine the temporal change intensity of the temporal change characteristics, obtain the structural anomaly intensity of each display area based on the spatial structure characteristics, and determine the target display area where the structural anomaly intensity is greater than the anomaly intensity threshold.
[0054] Specifically, anomaly analysis is performed on temporal variation characteristics and spatial structure characteristics to obtain the temporal variation intensity of temporal variation characteristics and the structural anomaly intensity of each display area.
[0055] Specifically, because related technologies focus on detecting anomalies in the overall image, they have a high rate of missing detection for minor anomalies such as screen corner flickering and slight display misalignment. This embodiment compares the structural anomaly intensity of each display area with an anomaly intensity threshold (which can be set according to actual detection accuracy requirements; the higher the accuracy requirement, the lower the anomaly intensity threshold). It then filters out target display areas with higher structural anomaly intensity, allowing for anomaly detection using the spatial structural features of these target display areas. This ignores feature data that has little impact on the detection target and improves the detection accuracy for minor display areas, avoiding missed detections of screen anomalies and enabling timely detection of minor screen anomalies, thus improving the user experience.
[0056] Step S205: Based on the intensity of temporal change and the intensity of structural anomaly in the target display area, the temporal change features and the spatial structural features corresponding to the target display area are fused, and the anomaly detection result of the display screen is determined according to the obtained fused features.
[0057] Specifically, by utilizing the intensity of temporal changes and the intensity of structural anomalies in the target display area, the influence of temporal change features and the spatial structural features corresponding to the target display area on screen anomaly detection is measured. Thus, during feature fusion, the fusion ratio of temporal change features and spatial structural features corresponding to the target display area is dynamically adjusted to obtain fusion features that are more adapted to the actual scene. In this way, the fusion features are used to detect anomalies on the display screen and obtain anomaly detection results.
[0058] The display screen anomaly detection method provided in this embodiment extracts temporal variation features between consecutive video frames and spatial structure features corresponding to different display areas of the display screen based on the video frame feature map of the monitored video of the display screen, so as to analyze the spatial structure of different display areas separately. Next, it analyzes the temporal variation intensity of the temporal variation features and the structural anomaly intensity of each display area to measure the influence of temporal variation features and spatial structure features on screen anomaly detection. Then, it performs feature fusion on the spatial structure features of target display areas with high temporal variation features and high structural anomaly intensity to improve the recognition rate of screen corner areas and minor display anomalies, thus obtaining the display screen anomaly detection result.
[0059] This embodiment provides a method for detecting abnormal display screens, which can be used in the aforementioned mobile terminals, such as in-vehicle terminals, mobile phones, tablet computers, etc. Figure 3 This is a flowchart of a display screen anomaly detection method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps: Step S301: Acquire the monitoring video of the display screen and extract the video frame feature map of the monitoring video. For details, please refer to [link to relevant documentation]. Figure 2 Step S201 of the illustrated embodiment will not be described again here.
[0060] Step S302: Extract the temporal variation features between consecutive video frames based on the video frame feature map.
[0061] In step S302, the differential features between every two consecutive video frames are calculated based on the video frame feature map, and the temporal variation features between consecutive video frames are obtained based on the differential features.
[0062] Specifically, to determine the feature maps of every two consecutive corresponding video frames, the difference features between every two consecutive video frames can be extracted using the following formula:
[0063] in, Indicates the first The spatial location of feature maps of each video frame The video frame feature value at channel c, No. The spatial location of feature maps of each video frame The video frame feature value at channel c, The number of feature maps in the video frame. For differential features. The feature map of the k-th video frame to the... The temporal variation features between the feature maps of each video frame are .
[0064] Existing technologies employing SE attention only calculate weights for spatial domain channels within a single frame, failing to perceive the temporal changes in channel features and hindering the capture of dynamic information between video frames. This embodiment, however, calculates differential features between consecutive video frames to collect core temporal information—frame-to-frame changes—providing input for temporal attention. By calculating differential features, a large amount of static redundant information in consecutive video frames is filtered out. For example, the differential feature result for a fixed background in screen anomaly detection is close to 0, effectively reducing the input of static redundant information ineffective for screen anomaly detection. This allows the model to accurately focus on anomaly areas on the screen, significantly reducing robustness issues caused by environmental interference.
[0065] This embodiment calculates the difference features between every two consecutive video frames, uses these difference features as input for temporal change features, filters out a large amount of static redundant information in consecutive video frames, reduces the input of irrelevant static redundant information, and improves the robustness of screen anomaly detection.
[0066] Step S303 involves dividing the video frame feature map into multiple region feature maps and extracting the spatial structure features of each region feature map; wherein, each region feature map corresponds one-to-one with a different display area of the display screen. For details, please refer to [link to details]. Figure 2 Step S203 of the illustrated embodiment will not be described again here.
[0067] Step S304: Determine the temporal change intensity of the temporal change characteristics, obtain the structural anomaly intensity of each display area based on the spatial structure characteristics, and determine the target display area where the structural anomaly intensity is greater than the anomaly intensity threshold.
[0068] Specifically, step S304 includes: Step S3041: Determine the mean squared error characteristics of the time series variation features, and calculate the time series variation intensity of the time series variation features based on the comparison relationship between the mean squared error characteristics and the preset mean squared error range.
[0069] Specifically, the mean squared error of the time series variation characteristics can be calculated using the following formula:
[0070] in, The mean squared error characteristic is at channel c. Let be the mean of all difference features of channel c. For the high-resolution feature map of video frames, The width of the feature map of the video frame.
[0071] Furthermore, the temporal variation intensity of the temporal variation characteristics at each channel is calculated based on the mean square error characteristics. The calculation formula is as follows:
[0072] in, The temporal variation intensity represents the temporal variation characteristics at channel c. Then, the temporal variation intensity... , This represents the minimum value within the preset mean square error range. This represents the maximum value of the preset mean squared error range. The preset mean squared error range can be set by obtaining the mean squared error range corresponding to the temporal variation characteristics of various screens.
[0073] In this embodiment, the mean squared error (MSE) feature is used to measure the degree of deviation of data from the mean. By using the MSE feature to calculate the temporal variation intensity of the difference features in continuous video frames, the fluctuation characteristics of the difference features (the dynamic fluctuation of data over time) can be captured, while suppressing isolated noise in individual video frames (interference signals that exist alone in the data, are unrelated to surrounding effective information, have a small range, and affect local areas). This accurately quantifies the persistence of temporal variations while avoiding local misjudgments. Furthermore, the calculation of the MSE only involves basic operations such as the difference features and the mean of the difference features, without the need for complex convolutions, matrix multiplications, or activation functions. The computational complexity of the MSE for a single channel is [insert value here]. (N is the number of video frames), the overall computational overhead is extremely low, and it is fully adaptable to scenarios with limited computing power, such as embedded devices and real-time detection.
[0074] This embodiment captures the fluctuation characteristics of temporal variation features by calculating the mean squared error (MSE) of these features, suppressing isolated noise in individual video frames, avoiding local misjudgments, and with extremely low computational overhead. Furthermore, by comparing the MSE features with a preset MSE range, the intensity of temporal variation features is measured, thereby improving the accuracy of screen anomaly detection by combining the intensity of temporal variation.
[0075] Step S3042: Based on the brightness and texture features in the spatial structure features, obtain the brightness variance features and texture variance features of each display area.
[0076] In some embodiments, for each region feature map, the brightness variance feature of that region feature map can be extracted according to the following formula:
[0077] in, Indicates the characteristics of luminance variance. In spatial location The pixel brightness value at channel c. This represents the average brightness value of all pixels in channel c of the region feature map. For the height of the region feature map, The width of the region feature map.
[0078] In some embodiments, for each region feature map, the texture variance feature of that region feature map can be extracted according to the following formula:
[0079] in, Represents texture variance features. In spatial location Texture gradient value at channel c This is the mean of all texture gradient values in channel c of the region feature map.
[0080] Step S3043: Determine the brightness difference characteristics between the brightness variance characteristics of each display area and the preset brightness variance, and determine the texture difference characteristics between the texture variance characteristics of each display area and the preset texture variance.
[0081] In some embodiments, for each display area, the brightness difference characteristics corresponding to that display area can be determined according to the following formula:
[0082] in, Indicates brightness difference characteristics. This is the preset brightness variance. It should be noted that the preset brightness variance can be set based on the baseline value of the brightness variance of a normal screen.
[0083] In some embodiments, for each display area, the texture difference features corresponding to that display area can be determined according to the following formula:
[0084] in, Indicates texture difference features, This is the preset texture variance. It should be noted that the preset texture variance can be set based on the texture variance baseline value of a normal screen.
[0085] Step S3044: Calculate the structural anomaly intensity of each display area based on the brightness difference characteristics and texture difference characteristics.
[0086]
[0087] in, Indicates structural anomalous strength. The weights corresponding to the brightness difference features. The settings can be adjusted based on the importance of brightness difference characteristics.
[0088] In this embodiment, since brightness features are a direct indicator of the presence of screen displays, anomaly detection based on brightness features can cover core fault types, while texture features are a key dimension describing the structural patterns of screen content. Variance calculation is performed using brightness and texture features, and the resulting brightness and texture difference features are used to measure the intensity of structural anomalies in each display area. This achieves full-scene coverage while supporting mutual verification, with low computational cost, making it more suitable for the lightweight requirements of screen anomaly detection.
[0089] Specifically, the display screen displays continuous images that change over time in response to user actions and real-time information updates. Furthermore, the information and content elements displayed in fixed areas of the interface are arranged in an orderly manner, thus exhibiting a significant spatiotemporal distribution pattern. Because the original data features have complex correlations, especially nonlinear correlations and multi-factor interactions, they are often limited by spatial dimensions. Low-dimensional processing alone is generally insufficient to effectively characterize relationships in complex spaces. For example, two types of data may completely overlap on a two-dimensional plane and cannot be separated by a straight line, but when mapped to three-dimensional space, they may exhibit a separable spatial distribution. In high-dimensional space, previously mixed key feature information, such as dynamic dependencies in spatiotemporal data and abstract textures in images, is amplified and decoupled. Complex relationships that are difficult to fit linearly, such as nonlinear dependencies and dynamic interaction patterns between data, are also transformed into more easily processed forms in high-dimensional space.
[0090] Therefore, this embodiment extracts brightness and texture features to map low-dimensional data to a high-dimensional space, uncovering patterns that cannot be found in the low-dimensional space from a high-dimensional perspective. After extracting these features, the core parts of the high-dimensional feature data are compressed by calculating the intensity of temporal changes and structural anomalies, in order to capture the most important regular characteristics in the data, while ignoring feature data with less impact on the target. Through layer-by-layer processing, the original data is transformed into structured and generalized abstract relationships, giving the extracted data spatiotemporal characteristics.
[0091] Step S305: Based on the intensity of temporal change and the intensity of structural anomaly in the target display area, the temporal change features and the spatial structural features corresponding to the target display area are fused, and the anomaly detection result of the display screen is determined according to the obtained fused features.
[0092] Specifically, step S305 includes: Step S3051: Input the intensity of temporal change and the intensity of structural anomaly in the target display area into the pre-trained nonlinear model to obtain the fusion coefficient; wherein, the fusion coefficient is used to characterize the importance of the temporal change features.
[0093] Specifically, due to the intensity of temporal changes The intensity of structural anomalies reflects the temporal dynamics (such as screen flickering and stuttering) of consecutive video frames. The spatial structure anomalies reflected in a single frame (such as black screen or distorted screen) are not simply linearly related; therefore, splicing features will be considered. Input a pre-trained nonlinear model to uncover the nonlinear relationship between temporal anomalies and spatial structure anomalies, and determine the intensity of temporal changes. and structural anomaly strength The target association pattern is used to obtain fusion coefficients that characterize the importance of features corresponding to temporal changes.
[0094] It should be noted that nonlinear models can be constructed based on multi-layer perceptrons (MLPs). By training the nonlinear model, it can autonomously learn the intensity of temporal changes under different anomalous scenarios. and structural anomaly strength The optimal correlation pattern is obtained, thereby improving the model's adaptability. Nonlinear models are... The activation function yields the fusion coefficients. :
[0095] in, express The weight matrix corresponding to the function, yes The bias term corresponding to the function, Represents a linear transformation. The function can be a nonlinear function used to filter irrelevant or redundant features. express The weight matrix corresponding to the activation function, yes The bias term corresponding to the activation function. , , as well as Adjustments can be made during the training process of the nonlinear model.
[0096] This embodiment further explores the higher-order correlation between the intensity of temporal variation and the intensity of structural anomalies by nesting linear and nonlinear transformations, demonstrating their impact on feature fusion and achieving quantitative output of the dynamic strategy. Compared to traditional fusion methods such as fixed weights and linear weighting, this embodiment generates fusion coefficients through nonlinear models and activation functions. It can dynamically adapt to different scenarios such as time-driven screen anomalies (e.g., flickering, lag), space-driven screen anomalies (e.g., black screen, partial screen distortion), and mixed anomalies (e.g., flickering + screen distortion), achieving non-linear correlation of spatiotemporal features and matching the model's expressive power. Experimental data shows that the detection accuracy of the above method for flickering, black screen, and mixed anomalies is significantly higher than that of a fixed weight (fusion coefficient). =0.5) Increased by 18%, 12%, and 23%.
[0097] Step S3052: Based on the fusion coefficient, the temporal variation features and the spatial structure features corresponding to the target display area are fused to obtain the fused features.
[0098] Specifically, the temporal variation features and the spatial structure features corresponding to the target display area are fused according to the following formula:
[0099] in, Indicates spatial location ,aisle Fusion characteristics at the location, This represents the temporal change characteristics corresponding to the current time t (i.e., the temporal change characteristics between the video frame feature map extracted at time t-1 and the video frame feature map extracted at time t). It captures dynamic changes between consecutive frames; It represents the spatial structural features (including brightness and texture features) of the target display area, capturing the spatial layout, texture and other structural information within a single frame.
[0100] It should be noted that before fusing the temporal variation features and the spatial structure features corresponding to the target display area, the temporal variation features need to be cropped according to the pixel region corresponding to the target display area to ensure... , as well as The corresponding pixel regions are consistent, that is, the spatial structural features and temporal change features corresponding to the target display area are fused to obtain the fused features.
[0101] It should be noted that the fusion coefficient of the time-series variation feature part Coefficients of spatial structural features The sum of the two is 1, which conforms to basic mathematical rules and reflects the complementary relationship of their inverse relationship. When the fusion coefficient... When the fusion coefficient approaches 1, it indicates that the temporal variation characteristics have a very significant impact on the current anomaly detection; therefore, a higher value is assigned to represent the degree of this impact. Conversely, when the fusion coefficient... When the coefficient approaches 0, spatial structural features (such as black screen or distorted screen) become more important and therefore have higher weights. This demonstrates that the fusion coefficients learned from the nonlinear model... It can perform weighted fusion processing based on the characteristics of the input data, and dynamically allocate attention resources according to the real-time scene, thereby achieving the effect of adaptively adjusting the fusion ratio of temporal change features and spatial structure features.
[0102] Compared to models that use only single-channel feature dimensions, this embodiment can more comprehensively and accurately represent the features of the input data, achieving a higher recognition rate and improving the model's performance in anomaly detection, as well as its adaptability to different scenarios and task requirements. On the other hand, the fusion coefficient... The value of intuitively reflects the proportion of temporal variation features and spatial structure features in the final fused features, which improves the interpretability of the model and helps the model understand which type of information it relies on more when processing different inputs, thus providing a basis for model analysis and optimization.
[0103] This embodiment inputs the intensity of temporal variation and the intensity of structural anomalies in the target display area into a pre-trained nonlinear model. The nonlinear model determines the degree of influence of temporal variation features and spatial structural features of the target display area on screen anomaly detection, thereby outputting fusion coefficients. These fusion coefficients can be used to adjust the weighting of temporal variation features and spatial structural features of the target display area during feature fusion, improving the adaptability of screen anomaly detection in various scenarios.
[0104] Step S3053: Determine the anomaly detection result of the display screen based on the obtained fusion features.
[0105] In some optional implementations, step S3053 above includes: Step a1: Compress the fused features to obtain a global feature vector of fixed length.
[0106] Specifically, after obtaining the fused features, a Global Average Pooling (GAP) layer is used to compress the fused features into a fixed-length global feature vector. The GAP layer can perform a global averaging operation on the fused features, compressing the two-dimensional fused features into a fixed-length vector. This can aggregate the global information of the feature map, eliminate differences in spatial dimensions, and greatly reduce the feature dimensionality, thereby reducing the number of parameters in the subsequent anomaly classification model and avoiding overfitting.
[0107] Step a2: Input the global feature vector into the pre-trained anomaly classification model to obtain the anomaly detection results of the display screen; wherein, the anomaly detection results include screen flickering, black screen, distorted screen, white screen and normal screen.
[0108] This embodiment compresses the fused features into a fixed-length global feature vector, aggregates the global information of the fused features, and then uses the global feature vector to determine whether the display screen has anomalies such as screen flickering, black screen, distorted screen, or white screen. Moreover, the global feature vector can effectively reduce the feature dimension and avoid overfitting of the pre-trained anomaly classification model when performing anomaly detection.
[0109] In some embodiments, the training process of the anomaly classification model includes: Step c1: Construct the screen image dataset and the preset model structure.
[0110] Specifically, such as Figure 4As shown, screen images in various states, including normal, distorted, black, white, and flickering screens, are captured using an in-vehicle camera or screen recording device, covering scenarios such as strong light, low light, and dynamic interfaces. The screen images are then resized to a standard 512*512 pixel size to construct a screen image dataset. Furthermore, a pre-defined model architecture is built based on a fully connected layer and a Softmax classifier.
[0111] Step c2 involves performing data augmentation operations on the screen image dataset to obtain an augmented image dataset. The data augmentation operations include image occlusion, geometric transformation, lighting adjustment, screen anomaly simulation, and resolution adjustment.
[0112] For example, see again Figure 4 The model performs local image occlusion on screen images in the screen image dataset; performs geometric transformations on screen images such as random rotation (±15°), horizontal flipping, and scaling (0.8~1.2 times); adjusts the lighting of screen images such as brightness increase / decrease (±20%), contrast change (±15%), and Gaussian noise addition (Gaussian noise is 0.1); performs color dithering or Gaussian blurring on local areas of screen images to simulate abnormal screen states such as screen flickering / white screen transition; and adjusts the size of screen images to 224×224, 380×380, 456×456, etc., and trains the model with screen images of various resolutions.
[0113] This embodiment uses one or more of the above data augmentation operations in combination to augment the screen image dataset, thereby enhancing the robustness of the anomaly classification model.
[0114] Step c3: Add a category label to each image in the augmented image dataset; the category label corresponds to the anomaly type displayed on the screen.
[0115] Specifically, see again Figure 4 To enhance the annotation of each image in the image dataset, category labels were added, including flickering, distorted, black, white, and normal screens, corresponding to the anomaly types of the displayed screen. All image samples were labeled using a manual annotation system to form a high-quality multimodal training dataset, effectively improving the generalization ability of the anomaly classification model.
[0116] Step c4: Train the pre-set model structure using the enhanced image dataset to obtain a pre-trained anomaly classification model.
[0117] Specifically, such as Figure 5As shown, an iterative optimization training strategy is employed to improve the performance of the pre-trained model structure. After initializing the model using pre-trained model weights, the loss between the predicted class and the class label is calculated using the cross-entropy loss function, and the parameters are updated via backpropagation using the AdamW optimizer. During training, the model weights of the pre-trained model structure are periodically saved, and model pruning and quantization operations are performed after convergence. Finally, a pre-trained anomaly classification model is output, which achieves a compression rate improvement of over 30% and an accuracy loss of less than 1%, facilitating lightweight deployment.
[0118] In some embodiments, during the training of the anomaly classification model, the training framework can be PyTorch, the optimizer can be AdamW, and the initial learning rate can be 1e-4. Furthermore, the cross-entropy loss function is used to optimize the classification task, and the pre-trained model weights from EfficientNet are used for transfer learning to accelerate model convergence. Model pruning and INT8 quantization are used to compress the model size, achieving lightweight deployment and improving model inference speed (reaching over 30 FPS).
[0119] In some embodiments, during the training of the anomaly classification model, images of the vehicle's display screen are acquired in real time. Temporal change features and spatial structure features of different display areas are extracted, and the intensity of temporal change and structural anomaly are calculated. The temporal change features and the spatial structure features corresponding to the target display area are then fused to obtain fused features. The fused features are input into the anomaly classification model to obtain the probability distribution of anomaly types of the display screen output by the anomaly classification model. Anomaly detection results are obtained, and an alarm is triggered based on the target anomaly type corresponding to the highest probability, prompting the user to repair it in time.
[0120] In some embodiments, the pre-trained anomaly classification model is deployed in the cloud after pruning and quantization. The vehicle-mounted system encapsulates the global feature vector into a classification request and sends the request to the cloud via a preset API. This allows the cloud to invoke the pre-trained anomaly classification model based on the classification request and obtain the anomaly classification result. The vehicle-mounted system receives the anomaly classification result returned by the cloud and displays the anomaly detection result on the screen.
[0121] In some embodiments, taking an in-vehicle infotainment system as an example, after the anomaly classification model is trained, the pre-trained anomaly classification model is pruned and quantized to construct a lightweight feature extraction network. Subsequently, an inference engineering framework (including a preset calling interface (e.g., an API interface) is built and deployed on the in-vehicle infotainment system. This framework includes a preset neural network model used for video frame feature extraction, such as the EfficientNet-B3 model, a model framework for temporal variation features and spatial structure features, a computation framework for temporal variation intensity and structural anomaly intensity, and a feature fusion framework) and PyTorch is used to deploy the pruned and quantized anomaly classification model to the cloud. This allows the in-vehicle infotainment system to remotely call the preset calling interface, ensuring low-latency, high-concurrency real-time inference service capabilities.
[0122] Furthermore, the vehicle's infotainment system collects monitoring video from the display screen using the vehicle's camera, calls the deployed inference service for real-time analysis and processing, and obtains a global feature vector. The system then encapsulates the global feature vector into a classification request and sends the request to the cloud via an API call to obtain the abnormal classification results returned by the cloud (the probabilities of normal screen / distorted screen / black screen / white screen / flickering screen respectively). Based on the abnormal type with the highest probability in the abnormal classification results, the system triggers the corresponding alarm to remind the user to perform maintenance.
[0123] The following section uses a vehicle-mounted screen as an example to illustrate the abnormal detection scheme for the display screen of the present invention in detail with a specific application example.
[0124] Related technologies either rely on manual intervention and specific adaptations or focus on single scenarios, and are mostly based on static features of a single frame for detection, making it difficult to effectively identify dynamic display problems. Even relatively superior detection methods based on image recognition models struggle to balance accuracy and speed simultaneously. Furthermore, the attention mechanisms employed by these technologies use fixed weight allocation, lacking adaptive processing for temporal changes (such as dynamic content updates) and spatial structural characteristics (such as screen corner details) in in-vehicle scenarios. This results in a high rate of missed detections for small-sized anomalies and weak anti-interference capabilities.
[0125] like Figure 6 As shown in the diagram, this embodiment provides a model architecture for display screen anomaly detection, employing a hierarchical image classification architecture to handle the screen detection task. The input image is processed by the EfficientNet-B3 backbone network to extract basic video frame feature maps. Then, a fusion attention module performs channel-dimensional compression, followed by dual channel and spatial attention enhancement to obtain fused features. These fused features are compressed into a fixed-length global feature vector using global average pooling. The global feature vector is then mapped through a fully connected layer, and the Softmax classifier outputs the display screen anomaly detection result.
[0126] This invention proposes a display screen anomaly detection method based on temporal channel and spatial structure coupled attention. It introduces a dual modeling mechanism of temporal attention and structural attention, and achieves adaptive fusion of the two attention features through a fusion coefficient, ultimately realizing accurate detection of anomalies in vehicle-mounted screens. Verification shows that this method can achieve an accuracy of approximately 99% in detecting anomalies in vehicle-mounted screens, with an average single-frame detection time of ≤30ms, achieving a balance between detection accuracy and efficiency.
[0127] Specifically, the dynamic correlation between consecutive frames is captured by calculating the difference features of adjacent video frames, and the mean squared error feature is used to stabilize data fluctuations, thereby constructing an efficient temporal attention structure. The model processed by the temporal attention module can filter static noise interference, reduce redundant information, reduce the amount of feature data processing, and significantly improve the robustness and efficiency of dynamic anomaly recognition, thus solving the problem of SE attention failing to detect anomalies in the temporal dimension.
[0128] Furthermore, based on the abnormal characteristics of different display areas of the vehicle's infotainment screen, feature segmentation and region masking are performed. By calculating luminance variance and texture variance features, core fault scenarios are covered, and a stable spatial attention structure is established to capture spatially abnormal features. This structure is used to distinguish the feature priorities of different display areas, solving the problem of related technologies missing detection of local spatial anomalies.
[0129] The fusion attention mechanism of this invention first introduces a temporal attention mechanism by calculating the differential features of consecutive frames and the intensity of temporal changes, and then introduces a structural attention mechanism by dividing spatial features and calculating the intensity of structural anomalies. Next, fusion coefficients are used to achieve adaptive fusion of the two mechanisms, enhancing the interpretability of the model. Finally, this mechanism is inserted into the EfficientNet model. Through the fusion attention module, the ability to extract spatiotemporal features from abnormal screen regions is enhanced, improving the model's adaptability to various abnormal scenarios.
[0130] Specifically, a nonlinear model is used to explore the nonlinear relationship between temporal attention and structural attention, calculate the fusion coefficient, and dynamically achieve adaptive fusion of the two. By collaboratively identifying spatiotemporal features, the core anomalies of the display screen are pinpointed, avoiding the omission of continuous frame anomalies and spatial differences by single-channel dimensional attention, and effectively addressing vehicle screen detection in real-world multi-dimensional mixed anomaly scenarios.
[0131] This invention establishes a spatiotemporal attention and dynamic fusion mechanism to achieve adaptive fusion of spatiotemporal features for screen anomaly detection in different scenarios, exhibiting excellent adaptability and interpretability. Furthermore, it improves anomaly detection accuracy across various scenarios, reduces computational overhead by combining a lightweight classification architecture (global average pooling and fully connected layer design), and enhances the model's generalization ability for anomalies such as distorted screens, black screens, and white screens through multi-scenario data augmentation strategies. This allows for more accurate and comprehensive feature extraction while supporting lightweight and efficient multi-classification detection tasks, demonstrating significant advantages in automotive screen anomaly detection.
[0132] Compared with related technologies, this invention improves the average accuracy of vehicle screen anomaly detection by 14.9% while still maintaining an average single-frame detection time of less than 30ms. It achieves the construction of a high-precision, high-efficiency and interpretable vehicle screen anomaly detection model, enabling 24-hour uninterrupted detection of vehicle screen anomalies, thus improving the accuracy of vehicle screen anomaly detection and the level of intelligence in automotive cabin testing.
[0133] This embodiment also provides a display screen anomaly detection device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0134] This embodiment provides a display screen anomaly detection device, such as... Figure 7 As shown, it includes: The first processing module 701 is used to acquire the monitoring video of the display screen and extract the video frame feature map of the monitoring video; The second processing module 702 is used to extract the temporal change features between consecutive video frames based on the video frame feature map; The third processing module 703 is used to divide the video frame feature map into multiple region feature maps and extract the spatial structure features of each region feature map; wherein, the region feature map corresponds one-to-one with different display areas of the display screen; The fourth processing module 704 is used to determine the temporal change intensity of the temporal change characteristics, obtain the structural anomaly intensity of each display area based on the spatial structure characteristics, and determine the target display area where the structural anomaly intensity is greater than the anomaly intensity threshold. The fifth processing module 705 is used to fuse the temporal change features and the spatial structural features corresponding to the target display area based on the intensity of temporal change and the intensity of structural anomalies in the target display area, and to determine the anomaly detection result of the display screen based on the obtained fused features.
[0135] In some optional implementations, the second processing module 702 is further configured to: Based on the video frame feature map, the difference features between every two consecutive video frames are calculated; Based on the difference characteristics, the temporal variation characteristics between consecutive video frames are obtained.
[0136] In some alternative implementations, the fourth processing module 704 is further configured to: Determine the root mean square error characteristics of time-series variations; Based on the comparison between the mean squared error characteristics and the preset mean squared error range, the temporal variation intensity of the temporal variation characteristics is calculated.
[0137] In some alternative implementations, the fourth processing module 704 is further configured to: Based on the brightness and texture features in the spatial structure features, the brightness variance features and texture variance features of each display area are obtained; Determine the brightness variance characteristics of each display area and the brightness difference characteristics between the preset brightness variance, and determine the texture variance characteristics of each display area and the texture difference characteristics between the preset texture variance; The structural anomaly intensity of each display area is calculated based on the brightness difference characteristics and texture difference characteristics.
[0138] In some optional implementations, the fifth processing module 705 is further configured to: The intensity of temporal variation and the intensity of structural anomalies in the target display area are input into a pre-trained nonlinear model to obtain fusion coefficients; whereby the fusion coefficients are used to characterize the importance of the temporal variation features. Based on the fusion coefficient, the temporal variation features and the spatial structural features corresponding to the target display area are fused to obtain the fused features.
[0139] In some optional implementations, the fifth processing module 705 is further configured to: The fused features are compressed to obtain a global feature vector of fixed length; The global feature vector is input into a pre-trained anomaly classification model to obtain anomaly detection results for the display screen; the anomaly detection results include screen flickering, black screen, distorted screen, white screen, and normal screen.
[0140] In some alternative implementations, the training process for the anomaly classification model includes: Construct a screen image dataset and a pre-defined model structure; Data augmentation operations are performed on the screen image dataset to obtain an augmented image dataset. The data augmentation operations include image occlusion, geometric transformation, lighting adjustment, screen anomaly simulation, and resolution adjustment. Add a category label to each image in the augmented image dataset; the category label corresponds to the anomaly type displayed on the screen. The pre-trained anomaly classification model is obtained by training a pre-defined model structure using an augmented image dataset.
[0141] In some optional implementations, the pre-trained anomaly classification model is deployed in the cloud after pruning and quantization; the fifth processing module 705 is also used for: Encapsulate the global feature vector into a classification request; By using a pre-defined API, a classification request is sent to the cloud, which then calls a pre-trained anomaly classification model based on the classification request to obtain the anomaly classification result. Receive the anomaly classification results returned from the cloud and obtain the anomaly detection results displayed on the screen.
[0142] The display screen anomaly detection device provided in this embodiment of the invention can execute the display screen anomaly detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.
[0143] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0144] The following is a detailed reference. Figure 8 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from memory 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device. The processor 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0145] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0146] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a memory 808, or installed from a ROM 802. When the computer program is executed by the processor 801, it performs the functions defined in the display screen anomaly detection method of the embodiments of the present invention.
[0147] Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention.
[0148] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the display screen anomaly detection method shown in the above embodiments is implemented.
[0149] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0150] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for detecting display screen anomalies, characterized in that, The method includes: Acquire monitoring video of the display screen and extract video frame feature maps from the monitoring video; Based on the video frame feature map, extract the temporal variation features between consecutive video frames; The video frame feature map is divided into multiple region feature maps, and the spatial structure features of each region feature map are extracted; wherein, each region feature map corresponds one-to-one with a different display area of the display screen; Determine the temporal change intensity of the temporal change characteristics, obtain the structural anomaly intensity of each display area based on the spatial structure characteristics, and determine the target display area where the structural anomaly intensity is greater than the anomaly intensity threshold. Based on the intensity of temporal change and the intensity of structural anomaly in the target display area, the temporal change features and the spatial structural features corresponding to the target display area are fused, and the anomaly detection result of the display screen is determined according to the obtained fused features.
2. The display screen anomaly detection method according to claim 1, characterized in that, The process of fusing the temporal variation features and the spatial structural features corresponding to the target display area based on the temporal variation intensity and the structural anomaly intensity of the target display area includes: The intensity of the temporal variation and the intensity of the structural anomaly in the target display area are input into a pre-trained nonlinear model to obtain fusion coefficients; wherein, the fusion coefficients are used to characterize the importance of the temporal variation features. Based on the fusion coefficient, the temporal variation features and the spatial structure features corresponding to the target display area are fused to obtain the fused features.
3. The display screen anomaly detection method according to claim 2, characterized in that, The step of obtaining the structural anomaly intensity of each display area based on the spatial structural characteristics includes: Based on the brightness and texture features in the spatial structure features, the brightness variance features and texture variance features of each display area are obtained; Determine the brightness variance characteristics of each display area and the brightness difference characteristics between the preset brightness variance, and determine the texture variance characteristics of each display area and the texture difference characteristics between the preset texture variance; The structural anomaly intensity of each display area is calculated based on the brightness difference characteristics and the texture difference characteristics.
4. The display screen anomaly detection method according to claim 2, characterized in that, Determining the temporal variation intensity of the temporal variation characteristics includes: Determine the mean square error characteristics of the time-series variation features; The temporal variation intensity of the temporal variation feature is calculated based on the comparison relationship between the mean squared error feature and the preset mean squared error range.
5. The display screen anomaly detection method according to any one of claims 1-4, characterized in that, The step of extracting temporal variation features between consecutive video frames based on the video frame feature map includes: Based on the video frame feature map, the difference features between every two consecutive video frames are calculated; Based on the differential features, the temporal variation features between consecutive video frames are obtained.
6. The display screen anomaly detection method according to any one of claims 1-4, characterized in that, The step of determining the anomaly detection result of the display screen based on the obtained fusion features includes: The fused features are compressed to obtain a global feature vector of fixed length; The global feature vector is input into a pre-trained anomaly classification model to obtain anomaly detection results for the display screen; wherein, the anomaly detection results include screen flickering, black screen, distorted screen, white screen, and normal screen.
7. The display screen anomaly detection method according to claim 6, characterized in that, The training process of the anomaly classification model includes: Construct a screen image dataset and a pre-defined model structure; Data augmentation operations are performed on the screen image dataset to obtain an enhanced image dataset; the data augmentation operations include image occlusion, geometric transformation, lighting adjustment, screen anomaly simulation, and resolution adjustment. Add a category label to each image in the enhanced image dataset; the category label corresponds to the anomaly type displayed on the screen. The pre-trained anomaly classification model is obtained by training the pre-set model structure using the enhanced image dataset.
8. The display screen anomaly detection method according to claim 7, characterized in that, The pre-trained anomaly classification model is deployed in the cloud after pruning and quantization; the step of inputting the global feature vector into the pre-trained anomaly classification model to obtain the anomaly detection results displayed on the screen includes: The global feature vector is encapsulated into a classification request; The classification request is sent to the cloud through a preset calling interface, so that the cloud can call the pre-trained anomaly classification model based on the classification request to obtain the anomaly classification result; Receive the anomaly classification results returned from the cloud and obtain the anomaly detection results displayed on the screen.
9. A display screen anomaly detection device, characterized in that, The device includes: The first processing module is used to acquire the monitoring video of the display screen and extract the video frame feature map of the monitoring video; The second processing module is used to extract the temporal variation features between consecutive video frames based on the video frame feature map; The third processing module is used to divide the video frame feature map into multiple region feature maps and extract the spatial structure features of each region feature map; wherein, the region feature map corresponds one-to-one with different display areas of the display screen; The fourth processing module is used to determine the temporal change intensity of the temporal change characteristics, obtain the structural anomaly intensity of each display area based on the spatial structure characteristics, and determine the target display area where the structural anomaly intensity is greater than the anomaly intensity threshold. The fifth processing module is used to fuse the temporal change features and the spatial structural features corresponding to the target display area based on the temporal change intensity and the structural anomaly intensity of the target display area, and determine the anomaly detection result of the display screen according to the obtained fused features.
10. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the display screen anomaly detection method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the display screen anomaly detection method according to any one of claims 1 to 8.
12. A computer program product, characterized in that, Includes computer instructions, said computer instructions being used to cause a computer to perform the display screen anomaly detection method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method for detecting images with brightness abnormality and LED display screen uniformity correction method
CN104900178A
LED display screen fault detection method and system
CN119478431A
Performance detection method and system for liquid crystal display module
CN120141807A
Display screen dead pixel fault prediction detection method and system
CN120235879A
LED display screen fault detection method, device, equipment and medium
CN120579027A