Rearview mirror intelligent detection method and system based on machine vision

By using a machine vision-based intelligent detection method for rearview mirrors, and leveraging an improved neural network model and control signals from the drive module, comprehensive, accurate, and rapid detection of rearview mirrors is achieved, solving the problems of low efficiency and incompleteness in traditional detection methods.

CN121937355APending Publication Date: 2026-04-28WUXI VOCATIONAL INSTITUTE OF COMMERCE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUXI VOCATIONAL INSTITUTE OF COMMERCE
Filing Date
2025-11-25
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Traditional rearview mirror inspection methods, whether manual or automated, suffer from low efficiency and incomplete inspection.

Method used

A machine vision-based intelligent inspection method for rearview mirrors is adopted. By acquiring video data of the rearview mirror to be inspected, preprocessing it, and then using an improved neural network model to perform frame-by-frame reasoning, combined with the control signal triggering timing of the rearview mirror drive module, the functional and appearance defects of the rearview mirror are analyzed to achieve integrated automated inspection.

Benefits of technology

It enables comprehensive, accurate, and rapid inspection of rearview mirrors, avoiding the subjectivity and missed detection problems of manual inspection, and improving inspection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937355A_ABST
    Figure CN121937355A_ABST
Patent Text Reader

Abstract

The invention provides a rearview mirror intelligent detection method and system based on machine vision, and relates to the technical field of image processing, and the method comprises the steps: obtaining video data of a rearview mirror to be detected; preprocessing each image frame of the video data to obtain a standardized image sequence; performing frame-by-frame reasoning on each standardized image frame in the standardized image sequence through a rearview mirror detection model based on a neural network, and detecting and positioning a target position of the rearview mirror from the standardized image frames; a time sequence is triggered by combining a control signal of a rearview mirror driving module, the target position is analyzed, and whether all functions of the rearview mirror are normal or not is detected; key image frames are selected from the standardized image sequence; reasoning the key image frame through an appearance defect detection module based on a neural network, and detecting whether the rearview mirror has an appearance defect and a corresponding defect position; and outputting various function detection results and appearance defect detection results of the rearview mirror.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a machine vision-based intelligent detection method and system for rearview mirrors. Background Technology

[0002] As a key component for safe driving, rearview mirrors directly affect a driver's ability to observe road conditions behind and to the sides, making them a crucial part of ensuring driving safety. The quality and performance of rearview mirrors not only affect the driver's field of vision and blind spot size, but also involve multiple technical indicators such as mirror reflectivity, weather resistance, and shock resistance. Rearview mirror inspection is an indispensable part of automobile manufacturing, after-sales maintenance, and annual vehicle inspections. With the development of intelligent vehicles, some rearview mirrors have integrated cameras and blind spot monitoring functions, further increasing the complexity and technical requirements of inspection. Therefore, comprehensive and professional rearview mirror inspection is of great significance for ensuring driving safety and improving the driving experience.

[0003] Traditional rearview mirror inspection methods often rely on manual inspection. Due to the subjectivity of manual inspection, missed inspections are common, and manual inspection is time-consuming, labor-intensive, and the quality is not always satisfactory. Therefore, automated inspection equipment has been developed. However, when using automated equipment, tools are needed to fix the rearview mirror in place, and the mirror needs to be rotated during the inspection process for a comprehensive view. The limitations imposed by these fixing tools typically affect the convenience of the inspection process.

[0004] Therefore, traditional rearview mirror inspection methods, whether manual or automated, suffer from low inspection efficiency and incomplete inspection. Summary of the Invention

[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a machine vision-based intelligent detection method for rearview mirrors, which can solve the technical problems of low detection efficiency and incomplete detection in traditional rearview mirror detection methods, whether manual or automatic.

[0006] A first aspect of this invention proposes a machine vision-based intelligent detection method for rearview mirrors, comprising:

[0007] S1: Acquire video data of the rearview mirror to be tested;

[0008] S2: Preprocess each image frame of the video data to obtain a standardized image sequence;

[0009] S3: Using a neural network-based rearview mirror detection model, perform frame-by-frame reasoning on each standardized image frame in the standardized image sequence to detect and locate the target position of the rearview mirror from the standardized image frames.

[0010] S4: Combine the control signal triggering timing of the rearview mirror drive module to analyze the target position and detect whether the various functions of the rearview mirror are normal;

[0011] S5: Select key image frames from the standardized image sequence;

[0012] S6: By using a neural network-based appearance defect detection module, reasoning is performed on the key image frames to detect whether the rearview mirror has appearance defects and the corresponding defect locations;

[0013] S7: Output the test results of various functions of the rearview mirror and the test results of appearance defects.

[0014] A second aspect of this invention provides a machine vision-based intelligent rearview mirror detection system, comprising: a processor and a memory;

[0015] The memory stores programs or instructions that can run on the processor, which, when executed by the processor, implement the steps of the machine vision-based intelligent rearview mirror detection method as described in the first aspect.

[0016] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0017] In this embodiment of the invention, deep neural network recognition and functional synchronous verification are used to realize an integrated automated process for detecting appearance defects and functional status of rearview mirrors. The detection is comprehensive, accurate, and fast, and can avoid the subjectivity and missed detection problems of manual detection. Attached Figure Description

[0018] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0019] Figure 1 This is a flowchart illustrating a machine vision-based intelligent detection method for rearview mirrors provided in an embodiment of the present invention.

[0020] Figure 2 This is a schematic diagram of the structure of a rearview mirror detection model based on a neural network provided in an embodiment of the present invention.

[0021] Figure 3 This is a schematic diagram of the structure of a rearview mirror intelligent detection platform provided in an embodiment of the present invention.

[0022] Figure 4 This is a schematic diagram of the structure of a machine vision-based intelligent rearview mirror detection system provided in an embodiment of the present invention. Detailed Implementation

[0023] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] The intelligent rearview mirror detection method based on machine vision provided by the present invention will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0025] Reference manual attached Figure 1 The diagram shows a flowchart of a machine vision-based intelligent detection method for rearview mirrors provided by an embodiment of the present invention.

[0026] This invention provides a machine vision-based intelligent detection method for rearview mirrors, which may include the following steps:

[0027] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0028] S1: Acquire video data of the rearview mirror to be tested.

[0029] Specifically, a camera can be used to collect video data of the rearview mirror to be tested.

[0030] S2: Preprocess each image frame of the video data to obtain a standardized image sequence.

[0031] Specifically, the original video is first decoded frame by frame, and each frame undergoes denoising, brightness equalization, and color correction to eliminate noise and lighting deviations caused by the shooting environment. Then, distortion correction, rotation correction, or cropping correction is applied to the images to ensure the rearview mirror maintains a consistent posture within the frame. Next, the image size is uniformly scaled to adjust all frames to the fixed resolution required for neural network input, while normalization is performed to map pixel values ​​to a standard numerical range. If necessary, channel arrangement adjustment, histogram equalization, and edge enhancement are also performed to enable the subsequent detection model to perform inference under consistent, clear, and stable visual input conditions. Through these preprocessing steps, the original video can be converted into a standardized image sequence with uniform structure and stable quality, laying a reliable data foundation for subsequent deep learning-based functional detection and appearance defect recognition.

[0032] S3: Using a neural network-based rearview mirror detection model, the system performs frame-by-frame reasoning on each standardized image frame in the standardized image sequence to detect and locate the target position of the rearview mirror from the standardized image frames.

[0033] Optionally, the target locations include: the outer contour edge location, the turn signal light location, the blind spot monitoring signal light location, and the lens contour edge location.

[0034] Alternatively, neural networks such as YOLO and U-Net can be used for rearview mirror detection.

[0035] Optionally, the bounding boxes and corresponding confidence levels of each target location can be identified from standardized image frames using a neural network.

[0036] Reference manual attached Figure 2 The diagram shows a schematic of the structure of a rearview mirror detection model based on a neural network provided in an embodiment of the present invention.

[0037] In one possible implementation, the neural network-based rearview mirror detection model uses YOLOv8 as a framework, which has a backbone network, a neck network, and a head.

[0038] It should be noted that rearview mirror images have strong reflective properties, complex curved structures, large differences in the scale of appearance details, concentrated functional areas, and diverse and random defect morphologies. The default convolutional structure and static detection head of the traditional YOLOv8 neural network struggle to simultaneously capture both detailed texture and overall semantics. To address this issue, this invention makes the following improvements to the YOLOv8 neural network:

[0039] In the backbone network, multi-scale grouped dilated convolutional modules are used to replace conventional convolutional modules to extract multi-scale features. Convolutions with different dilation rates are used in parallel to extract multi-scale features such as small cracks, light area contours, and the overall shell, overcoming the limitation of the fixed receptive field of conventional 3×3 convolutions.

[0040] In the neck network, a cross-level perceptual self-attention module is used to integrate high-level semantic context while preserving details of low-level semantic features, and an adaptive detail enhancement module is used for multi-scale feature fusion. This allows low-level edge and texture details to be preserved while fusing with high-level semantic structure, effectively suppressing interference from pseudo-features such as specular reflections and bright spots, and achieving cross-scale semantic consistency.

[0041] The detection head uses a dynamic head. It can adaptively adjust the feature fusion method according to the local characteristics of the input image, thereby enhancing the ability to identify appearance defects and functional components with different shapes and scales.

[0042] Specifically, the architecture of the neural network-based rearview mirror detection model is as follows:

[0043] The input terminal connects to the CBS module. The input terminal is used to input normalized image frames.

[0044] The CBS module sequentially connects to the first multi-scale grouped dilated convolutional module (MSGDC), the first C2f module, the second multi-scale grouped dilated convolutional module, the second C2f module, the third multi-scale grouped dilated convolutional module, the third C2f module, and the SPPF module. The multi-scale grouped dilated convolutional modules are used to extract image features. The SPPF module performs multi-scale pooling on the last layer of features, further aggregating the global semantics from different receptive fields, which serves as the "high-level semantic input" for subsequent attention modules.

[0045] The CBS (Convolution-Batch Normalization-Swish) module, composed of convolutional layers, batch normalization layers, and activation functions, is the most basic feature extraction unit in the network.

[0046] Among them, the C2f (Cross-Stage Partial with Feature Fusion) module is an improved cross-stage feature fusion structure. By dividing the input features into multiple branches and performing cross-level connections in a deep layer, it achieves efficient feature reuse and multi-channel semantic fusion.

[0047] Among them, the multi-scale grouped dilated convolution module introduces multiple grouped dilated convolutions with different dilation rates in parallel in the same layer, enabling the network to obtain feature representation capabilities of small, medium, and large receptive fields simultaneously without significantly increasing the computational cost.

[0048] Among them, the SPPF (Spatial Pyramid Pooling - Fast) module is a lightweight and accelerated version of Spatial Pyramid Pooling (SPP). It forms a multi-scale receptive field through multi-level max pooling operations, enabling the network to capture local details and global context information simultaneously on a single feature layer.

[0049] It should be noted that the combination of MSGDC + C2f + SPPF fully extracts the multi-scale structural features, contour information and key region features of the rearview mirror, and compresses them into multi-level feature maps, laying a good foundation for subsequent cross-level attention and multi-scale fusion.

[0050] The SPPF module is connected to the first cross-level perceptual self-attention module (CPSA). The second C2f module and the first cross-level perceptual self-attention module are both connected to the second cross-level perceptual self-attention module. The first C2f module and the second cross-level perceptual self-attention module are both connected to the third cross-level perceptual self-attention module. In the cross-level perceptual self-attention module, low-level semantic features are integrated into the high-level semantic context while preserving details. The first cross-level perceptual self-attention module, the second cross-level perceptual self-attention module, and the third cross-level perceptual self-attention module are all connected to the adaptive detail enhancement module (ADE). In the adaptive detail enhancement module, multi-scale feature fusion is performed.

[0051] It should be noted that the cross-level perception self-attention module is essentially a process in which high-level semantics guides low-level details from top to bottom. High-level semantics are injected into the middle and shallow layers step by step, so that the edges and texture details of the shallow layers "know which semantic region they belong to" (such as the light area, the edge of the mirror shell, and the mirror area) while being preserved.

[0052] Furthermore, the adaptive detail enhancement module addresses the question of "which layer of detail is more useful at the current scale," rather than simply piecing together multi-scale features in a crude manner, but rather performing a weighted and refined fusion. This is particularly crucial for targets like rearview mirrors, which have strong reflectivity, complex edges, and small defects.

[0053] The adaptive detail enhancement module is connected to the dynamic detection head. Based on the fused features, the dynamic detection head detects and locates the target position of the rearview mirror from the standardized image frame.

[0054] The logic of the entire neural network-based rearview mirror detection model is as follows: First, a multi-scale grouped dilated convolutional module + C2f + SPPF is used to fully extract the multi-scale structural features of the rearview mirror in the backbone network. Then, high-level semantics are injected into low-level details through cross-level perceptual self-attention. In the adaptive detail enhancement module, the multi-scale features are masked and weighted and fused. Finally, the dynamic detection head completes the accurate detection and localization of the target position of the rearview mirror on the fused features, thereby significantly improving the detection accuracy and robustness in rearview mirror scenarios with strong reflection, variable scale, and subtle defects.

[0055] This invention innovatively proposes a novel cross-level perceptual self-attention mechanism to be used in conjunction with a multi-scale grouped dilated convolution module, and also innovatively introduces an adaptive detail enhancement module.

[0056] In one possible implementation, image features are extracted using a multi-scale grouped dilated convolution module, specifically including:

[0057] The input features are subjected to multiple grouped dilated convolutions with different dilation rates, followed by residual connections and nonlinear activation.

[0058]

[0059] in, This represents the output feature map of the nth grouped convolution in layer (l-1), RPReLU represents the RPReLU activation function, B_Cov represents the dilated convolution, Ba represents the binarization function, and X represents the output feature map of the nth grouped convolution in layer (l-1). l-1 denoted as the input feature map of layer l-1, dil represents the inflation rate, the inflation rates are 2n-1, and n represents the total number of groups, n=1, 2 or 3.

[0060] The binarization function is as follows:

[0061]

[0062] Where sign represents the sign function, X represents the input feature map, b represents the scaling factor, and a represents the bias.

[0063] Feature fusion is performed on the output feature maps of each group of dilated convolutions to extract image features:

[0064]

[0065] Among them, H l-1 This represents the image feature map output by the (l-1)th layer. This represents the output feature map of the first group convolution in the (l-1)th layer. This represents the output feature map of the second grouped convolution in layer (l-1). This represents the output feature map of the third grouped convolution in the (l-1)th layer.

[0066] In this embodiment of the invention, the advantage of using a multi-scale grouped dilated convolution module is that it can significantly improve the network's ability to represent multi-scale structures with almost no increase in computational cost. By applying grouped dilated convolutions with different dilation rates to the same input feature, the network can simultaneously capture fine-grained textures, small defect features (from the small dilation rate branch), local region structures (from the medium dilation rate branch), and overall contours and large-scale semantics (from the large dilation rate branch), effectively solving the problem of traditional convolution having a fixed receptive field and difficulty in taking into account information at different scales. The binarization processing function makes the feature distribution more stable, which is beneficial for dilated convolution to maintain feature consistency under complex lighting and reflective scenes. Residual connections and RPReLU activation further improve gradient propagation efficiency and nonlinear expression capability. Finally, by fusing and normalizing the features of each branch, the module can generate high-quality multi-scale feature maps containing rich hierarchical information, significantly enhancing the robustness and accuracy of the network in target detection tasks such as rearview mirrors that are complex in shape, have a large scale span, and are sensitive to details.

[0067] Despite the numerous advantages of the multi-scale grouped dilated convolution module (MSGDC), it can only construct features based on local neighborhood features and is insensitive to global semantic relationships. Furthermore, dilated convolution is prone to texture discontinuities and semantic disconnects when processing images with strong reflections, complex backgrounds, or long-range dependent structures, making it difficult to determine global relationships between features using only local convolution. In addition, while multi-scale branches offer rich features, they lack the ability to dynamically model the semantic importance of different levels, easily leading to an imbalance in representation, such as "overly strong low-level details" or "insufficient high-level semantics." To address this issue, this invention innovatively proposes introducing a cross-level perceptual self-attention module. By achieving feature interaction and adaptive weight allocation between different semantic levels, high-level global semantics are injected into low-level detailed features. Simultaneously, dual channel and spatial attention is used to selectively strengthen key structural regions, thereby compensating for the shortcomings of the multi-scale grouped dilated convolution module, such as weak global information modeling ability, susceptibility to reflection interference, and insufficient hierarchical semantic fusion, achieving a more stable, comprehensive, and semantically consistent feature representation.

[0068] In one possible implementation, within the cross-level perceptual self-attention module, low-level semantic features are integrated into the high-level semantic context while preserving details, specifically including:

[0069] Using the semantic features output from the previous layer as a guide, the semantic features output from the previous layer are added to and fused with the semantic features of the current layer to obtain preliminary fused features:

[0070]

[0071] Where, x r Indicates preliminary fusion characteristics, x l Represents low-level semantic features, x h It represents high-level semantic features.

[0072] An attention map along the channel dimension is applied to the initial fused features, and a channel weight map is generated through linear transformation and activation function:

[0073]

[0074] Where, m c This represents the channel weight map, σ represents the activation function, and W... c This represents the channel attention matrix.

[0075] It should be noted that channel attention is used to selectively enhance key semantic channels and suppress reflective noise or background textures that are irrelevant to the target.

[0076] After applying the channel weight map to the initial fused features, spatial dimension attention is then applied to generate the spatial weight map:

[0077]

[0078] Where, m s Represents a spatial weighted graph, W s Represents the spatial attention matrix. This represents element-wise product.

[0079] It should be noted that spatial attention further determines the actual key pixel locations in the spatial region, making the model more focused on key locations such as the edges of the rearview mirror, the light area, and the mirror surface.

[0080] The initial fused features are double-weighted using channel and spatial weight maps to output multi-level semantically enhanced features:

[0081]

[0082] Where, x b This represents multi-level semantic enhancement features.

[0083] In this embodiment of the invention, by adding and fusing low-level semantic features with high-level semantic features, and applying channel attention and spatial attention in sequence, the cross-level perception self-attention module can introduce deep global semantics without destroying shallow texture details. This enables the network to simultaneously possess fine-grained structure recognition capabilities and long-range dependency modeling capabilities, making up for the shortcomings of the multi-scale grouped dilated convolution module, such as weak global information modeling capabilities, susceptibility to reflection interference, and insufficient hierarchical semantic fusion, and achieving a more stable, comprehensive, and semantically consistent feature expression.

[0084] Furthermore, while the multi-level semantic enhancement features obtained through the cross-level perceptual self-attention module are comprehensive, how to effectively select and fuse them across different scales remains an unsolved problem. Traditional fusion methods often simply and crudely piece together multi-scale features. This static and non-selective fusion leads to interference between fine-grained textures in high-resolution features and high-level semantics in low-resolution features, smoothing out important details while unintentionally amplifying irrelevant background or reflective areas. Therefore, this invention innovatively proposes an adaptive detail enhancement module.

[0085] In one possible implementation, multi-scale feature fusion is performed in the adaptive detail enhancement module, specifically including:

[0086] Convolutional operations are applied to multi-level semantic enhancement features at various scales to complete basic feature preprocessing, providing a unified feature representation for subsequent cross-scale alignment and mask generation.

[0087] Select a semantic enhancement feature at a certain scale as the anchor scale, and the other scales as non-anchor scales. Through upsampling or downsampling operations, adjust the features of each non-anchor scale to the same spatial size as the anchor scale to obtain the aligned cross-scale feature set.

[0088] Based on the alignment, for each non-anchored scale, a corresponding detail mask is generated through convolutional mapping and concatenation of activation functions:

[0089]

[0090] Among them, S i W represents the detail mask for the i-th non-anchored scale, Sigmoid represents the Sigmoid activation function, and W... i This represents the convolution kernel with the i-th non-anchored scale. This represents the i-th non-anchored scale feature after alignment.

[0091] It should be noted that this detail mask is essentially a weight map with values ​​ranging from 0 to 1, which can dynamically characterize the contribution of features at the current scale in different spatial locations. Through this learnable masking mechanism, the network can automatically and selectively retain high-frequency details that are sensitive to appearance defects, while suppressing irrelevant noise or background features, thus achieving adaptive weighted control of cross-scale features.

[0092] The anchored scale features processed by convolution are fused element-wise with all non-anchored scale detail masks to obtain the final fused features:

[0093]

[0094] Where Y represents the fusion feature, Wa X represents the convolution kernel with the anchoring scale. a This indicates the anchoring scale characteristics.

[0095] It should be noted that the anchor scale features after convolution are fused element-wise with all non-anchor scale detail masks to achieve fine-grained modulation of the anchor features by multi-scale information. Through this dot-multiplication fusion method, the anchor scale features are simultaneously constrained and reinforced by the detail masks in both spatial and scale dimensions, enabling them to introduce key details from other scales while maintaining overall semantic stability.

[0096] In this embodiment of the invention, the adaptive detail enhancement module eliminates the need for simple concatenation or element-wise addition when fusing multi-scale features. Instead, it first performs convolutional preprocessing on semantic enhancement features at different scales, then aligns the features at each scale through upsampling or downsampling to ensure feature space consistency. Subsequently, convolutional mapping and sigmoid activation are used to generate detail masks for each non-anchored scale, enabling the network to adaptively determine the importance of the feature at that scale in detail representation based on the current image content. Finally, anchored scale features are explicitly enhanced with effective details and irrelevant noise is suppressed through element-wise modulation of multiple detail masks, achieving a "selective, multi-level, and dynamically controllable" cross-scale fusion method. This method avoids the feature conflicts, detail overload, and semantic imbalance problems that occur in traditional multi-scale fusion, enabling the network to extract key structural information more accurately in scenarios such as rearview mirrors with strong reflections, fine-grained defects, and significant scale differences, thereby significantly improving detection robustness.

[0097] S4: Combine the control signal triggering timing of the rearview mirror drive module to analyze the target position and check whether the various functions of the rearview mirror are normal.

[0098] In one possible implementation, S4 specifically includes sub-steps S401 to S404:

[0099] S401: Determine whether the brightness at the turn signal position is above the first brightness threshold. If yes, confirm that the turn signal function of the rearview mirror is normal. Otherwise, confirm that the turn signal function of the rearview mirror is abnormal.

[0100] Those skilled in the art can set the size of the first brightness threshold according to the actual situation, and the present invention does not limit it.

[0101] S402: Determine whether the brightness at the location of the blind spot monitoring signal light is above the second brightness threshold. If yes, confirm that the blind spot monitoring signal light of the rearview mirror is functioning normally. Otherwise, confirm that the blind spot monitoring signal light of the rearview mirror is malfunctioning.

[0102] Those skilled in the art can set the size of the second brightness threshold according to the actual situation, and the present invention does not limit it.

[0103] S403: Calculate the rearview mirror folding angle based on the changes in the outer contour edge between consecutive frames, and determine whether the rearview mirror folding angle is above a first angle threshold. If yes, determine that the rearview mirror folding function is normal. Otherwise, determine that the rearview mirror folding function is abnormal.

[0104] Those skilled in the art can set the size of the first angle threshold according to the actual situation, and the present invention does not limit it.

[0105] Optionally, S403 specifically includes:

[0106] S4031: Calculate the change in outer contour between consecutive frames.

[0107] Among them, the change in the outer contour can quantify the actual displacement of the rearview mirror during the folding process.

[0108] S4032: When the change in the outer contour between consecutive frames is less than the first change threshold, the rearview mirror folding angle is normally calculated based on the change in the outer contour edge.

[0109] It should be noted that when the change in the outer contour is less than the first change threshold, it usually means that the folding action is smooth and the change pattern between frames is consistent. In this case, directly calculating the folding angle based on the change in the outer contour edge can obtain a reliable and accurate attitude estimate. By using a fast processing method for stable frames, unnecessary calculations can be reduced, the real-time performance of folding angle detection can be improved, and additional filtering can be avoided to prevent delays, thus keeping the system efficient and accurate during the stable action phase.

[0110] S4033: When the change in the outer contour between consecutive frames is between the first change threshold and the second change threshold, the change in the outer contour between consecutive frames is smoothed by moving average, and then the rearview mirror folding angle is calculated based on the change in the outer contour edge. If the second change threshold is greater than the first change threshold, the unstable situation is recorded.

[0111] It should be noted that when the inter-frame outer contour change falls between the first and second change thresholds, it indicates slight jitter or local instability in the folding action. Directly using the original change values ​​in this case may lead to fluctuations or errors in the folding angle conversion. Applying a moving average smoothing process to the change values ​​effectively filters short-term noise, preserves the true folding trend, and improves the continuity and stability of angle estimation. Simultaneously recording instability can serve as an auxiliary indicator for monitoring the health status of the rearview mirror folding mechanism, aiding in subsequent maintenance and diagnosis.

[0112] S4034: When the change in the outer contour between consecutive frames is greater than the second change threshold, the frame is regarded as an unstable frame and discarded. The rearview mirror folding angle is converted using adjacent stable frames, and the unstable situation is recorded.

[0113] It should be noted that when the change in the outer contour exceeds the second change threshold, it indicates that the frame has been severely jittered, obstructed, blurred, or mistakenly captured, and is no longer suitable as a basis for calculating the folding angle. Marking this frame as an unstable frame and removing it, and using an adjacent stable frame for angle conversion, can prevent extreme noise from causing huge deviations in the folding angle estimation, significantly improving the robustness and accuracy of the overall detection. Simultaneously recording the instability helps identify mechanical jamming or abnormal movements during the folding process, providing additional evidence for judging the reliability of the rearview mirror function.

[0114] Those skilled in the art can set the magnitudes of the first and second change thresholds according to actual circumstances; this invention does not impose any limitations.

[0115] S404: Calculate the lens folding angle based on the lens contour edge changes between consecutive frames, and determine whether the lens folding angle is above a second angle threshold. If yes, determine that the rearview mirror's lens adjustment function is normal. Otherwise, determine that the rearview mirror's lens adjustment function is abnormal.

[0116] Those skilled in the art can set the size of the second angle threshold according to the actual situation, and the present invention does not limit it.

[0117] Alternatively, the conversion process for the lens folding angle can also refer to the conversion process for the rearview mirror folding angle, introducing contour changes to obtain a more stable result.

[0118] S5: Select key image frames from the standardized image sequence.

[0119] In one possible implementation, S5 specifically includes sub-steps S501 to S505:

[0120] S501: Based on the target location detection results in S3, select image frames from the standardized image sequence in which all target locations have been detected.

[0121] It should be noted that selecting image frames in which all key parts (such as the outer contour, mirror surface, and lighting area) are successfully detected ensures that the rearview mirror target in subsequent images used for appearance defect detection is sufficiently visible and structurally complete, preventing errors in subsequent appearance judgment due to occlusion, angular deviation, or detection failure. This step effectively eliminates incomplete, misaligned, or partially invisible frames, providing a reliable starting point for key frame selection.

[0122] S502: Calculate the proportion of the total area of ​​the bounding boxes of the target locations in the selected image frames to the entire image to determine the complete visibility of the rearview mirror appearance of the image frame.

[0123] It should be noted that by calculating the proportion of the target bounding box area to the entire image, the visible range and completeness of the rearview mirror in the image can be quantified, solving the problem of inconsistent criteria for judging whether the rearview mirror's appearance is fully displayed. This indicator can effectively filter frames that are too far away, have cropped mirror edges, or where the rearview mirror only partially enters the frame, thus ensuring that the rearview mirror's shape is complete and its appearance area is available for accurate analysis in candidate frames.

[0124] S503: Select image frames whose rearview mirror appearance is complete and visible to a greater than a threshold, and then determine the sharpness and illumination uniformity of the selected image frames.

[0125] Those skilled in the art can set the threshold value according to the actual situation, and the present invention does not impose any limitations.

[0126] It should be noted that after selecting frames with complete and visible appearance, further judgment on sharpness and illumination uniformity can prevent blurry frames, overexposed frames, underexposed frames, or frames with uneven illumination distribution from entering the appearance defect detection process. This step strengthens the quality constraints of key frames, ensuring that the final candidate frames not only have complete targets but also possess good image quality with clear details and reasonable brightness, which is especially important for capturing minor defects such as cracks and scratches on rearview mirrors.

[0127] S504: The rearview mirror appearance integrity visibility, sharpness, and illumination uniformity of the selected image frames are weighted and summed to calculate the criticality index of each image frame.

[0128] It should be noted that by calculating the criticality index through weighted summation of multiple evaluation indicators, a comprehensive score can be achieved for candidate frames based on multi-dimensional quality factors. This ensures that the keyframes selected by the system not only meet a single indicator but also possess an optimal balance across various features. This method avoids biases caused by excessively strong single parameters, making the keyframe selection mechanism more objective, comprehensive, and robust, thus improving the overall reliability of appearance defect detection.

[0129] S505: The image frame with the highest criticality index is identified as the critical image frame.

[0130] It should be noted that selecting the image frame with the highest criticality index as the key image frame ensures that the image frame used for appearance defect detection is in optimal condition in terms of target integrity, clarity, and lighting conditions, thus providing the best input for the neural network's defect recognition. This step significantly improves the accuracy of rearview mirror appearance defect recognition, avoids false positives and false negatives caused by poor frame quality, and provides crucial assurance for the overall performance of the detection system.

[0131] S6: Through the appearance defect detection module based on neural network, reasoning is performed on key image frames to detect whether there are appearance defects in the rearview mirror and the corresponding defect locations.

[0132] Alternatively, mature neural networks such as YOLO and U-Net can be used for appearance defect detection. There are many mature algorithms for visual defect detection, and this invention does not limit which specific neural network is used for appearance defect detection.

[0133] S7: Outputs the test results of various functions of the rearview mirror and the test results of appearance defects.

[0134] In this embodiment of the invention, deep neural network recognition and functional synchronous verification are used to realize an integrated automated process for detecting appearance defects and functional status of rearview mirrors. The detection is comprehensive, accurate, and fast, and can avoid the subjectivity and missed detection problems of manual detection.

[0135] Reference manual attached Figure 3 The diagram shows a structural schematic of a rearview mirror intelligent detection platform provided in an embodiment of the present invention.

[0136] To construct a rearview mirror intelligent inspection platform, a main control unit 100 and a rearview mirror drive unit 200 and an image acquisition unit 300 connected to the main control unit 100 can be used. The rearview mirror drive unit 200 includes a rearview mirror drive module 210, a turntable 220, and a rearview mirror mounting base 230. The rearview mirror drive module 210 controls the folding and resetting, turn signal light activation / deactivation, blind spot monitoring signal light activation / deactivation, lens angle adjustment, and lens heating functions of the rearview mirror 400 to be inspected according to the inspection commands from the main control unit 100. The rearview mirror 400 to be inspected is fixed to the rearview mirror mounting base 230 by bolts. The rearview mirror mounting base 230 is fixed to the turntable 220 by bolts and rotates synchronously with the turntable 220. The rearview mirror drive module 210 is fixed to the rearview mirror mounting base 230 by bolts and connected to... The terminal block is electrically connected to the rearview mirror 400 to be tested; the image acquisition unit 300 includes a camera 310 and a robotic arm 320. The robotic arm 320 adjusts the position of the camera 310 according to the detection command of the main control unit 100. The camera 310 is fixed to the base at the top of the robotic arm 320 by bolts and is used to acquire video data of the rearview mirror 400 under the drive of the rearview mirror drive unit 200; the main control unit 100 is used to send detection commands to the rearview mirror drive unit 200 and the image acquisition unit 300, control the coordinated operation of the control robotic arm 320, the turntable 220 and the rearview mirror drive module 210, and, based on the above-mentioned intelligent rearview mirror detection algorithm based on machine vision, detect the rearview mirror 400 to be tested according to the video data, obtain the detection result, and display the detection result on the working interface in real time.

[0137] Reference manual attached Figure 4 The diagram shows a schematic of the structure of a machine vision-based intelligent rearview mirror detection system provided in an embodiment of the present invention.

[0138] This invention provides a machine vision-based intelligent rearview mirror detection system 20, comprising: a processor 201 and a memory 202;

[0139] The memory 202 stores programs or instructions that can run on the processor 201. When the program or instructions are executed by the processor 201, they implement the steps of the above-described intelligent rearview mirror detection method based on machine vision and achieve the same technical effect. To avoid repetition, the present invention will not elaborate further.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the protection scope of the present invention.

Claims

1. A machine vision-based intelligent detection method for rearview mirrors, characterized in that, include: S1: Acquire video data of the rearview mirror to be tested; S2: Preprocess each image frame of the video data to obtain a standardized image sequence; S3: Using a neural network-based rearview mirror detection model, perform frame-by-frame reasoning on each standardized image frame in the standardized image sequence to detect and locate the target position of the rearview mirror from the standardized image frames. S4: Combine the control signal triggering timing of the rearview mirror drive module to analyze the target position and detect whether the various functions of the rearview mirror are normal; S5: Select key image frames from the standardized image sequence; S6: By using a neural network-based appearance defect detection module, reasoning is performed on the key image frames to detect whether the rearview mirror has appearance defects and the corresponding defect locations; S7: Output the test results of various functions of the rearview mirror and the test results of appearance defects.

2. The intelligent rearview mirror detection method based on machine vision according to claim 1, characterized in that, The target locations include: the outer contour edge location, the turn signal light location, the blind spot monitoring signal light location, and the lens contour edge location.

3. The intelligent rearview mirror detection method based on machine vision according to claim 1, characterized in that, The neural network-based rearview mirror detection model uses YOLOv8 as its framework, which includes a backbone network, a neck network, and a detection head. In the backbone network, multi-scale grouped dilated convolutional modules are used instead of conventional convolutional modules to extract multi-scale features; In the neck network, a cross-level perceptual self-attention module is used to incorporate high-level semantic context while preserving details of low-level semantic features, and an adaptive detail enhancement module is used for multi-scale feature fusion. The detection head uses a dynamic detection head.

4. The intelligent rearview mirror detection method based on machine vision according to claim 3, characterized in that, The specific architecture of the neural network-based rearview mirror detection model is as follows: The input terminal is connected to the CBS module, and the input terminal is used to input the standardized image frame; The CBS module is sequentially connected to the first multi-scale grouped dilated convolution module, the first C2f module, the second multi-scale grouped dilated convolution module, the second C2f module, the third multi-scale grouped dilated convolution module, the third C2f module, and the SPPF module, and uses the multi-scale grouped dilated convolution module to extract image features. The SPPF module is connected to the first cross-level perception self-attention module. The second C2f module and the first cross-level perception self-attention module are both connected to the second cross-level perception self-attention module. The first C2f module and the second cross-level perception self-attention module are both connected to the third cross-level perception self-attention module. In the cross-level perception self-attention module, low-level semantic features are integrated into the high-level semantic context while preserving details. The first cross-level perception self-attention module, the second cross-level perception self-attention module, and the third cross-level perception self-attention module are all connected to the adaptive detail enhancement module. In the adaptive detail enhancement module, multi-scale feature fusion is performed. The adaptive detail enhancement module is connected to the dynamic detection head, and based on the fusion features, the dynamic detection head detects and locates the target position of the rearview mirror from the standardized image frame.

5. The intelligent rearview mirror detection method based on machine vision according to claim 4, characterized in that, The extraction of image features using a multi-scale grouped dilated convolution module specifically includes: The input features are subjected to multiple grouped dilated convolutions with different dilation rates, followed by residual connections and nonlinear activation. Feature fusion is performed on the output feature maps of each group of dilated convolutions to extract image features.

6. The intelligent rearview mirror detection method based on machine vision according to claim 4, characterized in that, In the cross-level perception self-attention module, the low-level semantic features are integrated into the high-level semantic context while preserving details, specifically including: Using the semantic features output from the previous layer as a guide, the semantic features output from the previous layer are added and fused with the semantic features of the current layer to obtain preliminary fused features; An attention mapping along the channel dimension is applied to the preliminary fusion features, and a channel weight map is generated through linear transformation and activation function; After applying the channel weight map to the initial fusion feature, spatial dimension attention is then applied to generate a spatial weight map. The preliminary fusion features are double-weighted using the channel weight map and the spatial weight map to output multi-level semantic enhancement features.

7. The intelligent rearview mirror detection method based on machine vision according to claim 6, characterized in that, The adaptive detail enhancement module performs multi-scale feature fusion, specifically including: Convolutional operations are applied to the multi-level semantic enhancement features at each scale to complete the basic feature preprocessing, providing a unified feature representation for subsequent cross-scale alignment and mask generation. Select a semantic enhancement feature at a certain scale as the anchor scale, and the other scales as non-anchor scales. Through upsampling or downsampling operations, adjust the features of each of the non-anchor scales to the same spatial size as the anchor scale to obtain an aligned cross-scale feature set. Based on the alignment, for each non-anchored scale, a corresponding detail mask is generated by convolutional mapping and concatenation activation functions; The anchored scale features after convolution are fused element-wise with all the detail masks of the non-anchored scales to obtain the final fused features.

8. The intelligent rearview mirror detection method based on machine vision according to claim 2, characterized in that, S4 specifically includes: S401: Determine whether the brightness at the position of the turn signal light is above the first brightness threshold; if yes, determine that the turn signal light of the rearview mirror is functioning normally; otherwise, determine that the turn signal light of the rearview mirror is malfunctioning. S402: Determine whether the brightness at the location of the blind spot monitoring signal light is above the second brightness threshold; if yes, determine that the blind spot monitoring signal light of the rearview mirror is functioning normally; otherwise, determine that the blind spot monitoring signal light of the rearview mirror is functioning abnormally. S403: Calculate the rearview mirror folding angle based on the change of the outer contour edge between consecutive frames, and determine whether the rearview mirror folding angle is above a first angle threshold; if so, determine that the folding function of the rearview mirror is normal; otherwise, determine that the folding function of the rearview mirror is abnormal. S404: Calculate the lens folding angle based on the lens contour edge change between consecutive frames, and determine whether the lens folding angle is above the second angle threshold; if so, determine that the lens adjustment function of the rearview mirror is normal; otherwise, determine that the lens adjustment function of the rearview mirror is abnormal.

9. The intelligent rearview mirror detection method based on machine vision according to claim 2, characterized in that, S5 specifically includes: S501: Based on the target location detection results in S3, select image frames from the standardized image sequence in which all target locations have been detected; S502: Calculate the proportion of the total area of ​​the bounding boxes of the target locations in the selected image frames to the entire image to determine the complete visibility of the rearview mirror appearance of the image frame. S503: Select image frames whose appearance integrity and visibility are greater than a threshold, and then determine the clarity and illumination uniformity of the selected image frames. S504: The rearview mirror appearance integrity visibility, sharpness and illumination uniformity of the selected image frames are weighted and summed to calculate the criticality index of each image frame. S505: The image frame with the highest criticality index is determined as the critical image frame.

10. A machine vision-based intelligent rearview mirror detection system, characterized in that, include: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the machine vision-based intelligent detection method for rearview mirrors as described in any one of claims 1 to 9.