Camera shielding state detection method
By using multi-branch structure and C2f-Star module and other technical means in camera occlusion detection, the existing methods have been solved in terms of robustness and real-timeness, and more efficient and accurate occlusion detection is achieved.
Patent Information
- Application Number
- CN202510077616.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing camera occlusion detection methods are less robust when dealing with lighting changes, noise interference and complex backgrounds in dynamic scenes, and edge detection and SSIM calculations may be time-consuming to process each frame, affecting real-time.
The feature extraction module with a multi-branch structure is adopted, and the fusion feature is combined with the C2f-Star module and attention module, and the feature level is fully utilized to improve the accuracy and real-time detection through 3×3 convolution, 1×1 convolution and identity transformation.
On the premise of ensuring detection accuracy, the model is lightweight and real-time, and the accuracy and generalization ability of camera occlusion state detection are improved.
Smart Images

Figure CN120014264A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of camera occlusion detection, and in particular to a camera occlusion state detection method. Background Art
[0002] With the continuous development of computer vision technology, image acquisition equipment has begun to enter all aspects of life and production. Including in security monitoring, vehicle advanced driving assistance system (ADAS), driver monitoring system (DMS), driving recorder and other systems, the camera, as a key visual sensor, is responsible for providing real-time environmental images to support the normal operation and decision-making of the system. However, in actual use, the camera is often blocked by external objects (for example, dust, hands, leaves, clothing, etc.), resulting in the device being unable to obtain the real scene, thus affecting the normal use of the function and bringing safety hazards. Therefore, camera occlusion detection has important practical significance.
[0003] At present, the method for camera occlusion detection mainly performs camera occlusion detection through color histogram or texture analysis of the similarity between the previous and next frames. For example, in patent application number CN201911114424X, patent name A method for detecting occlusion in driving recorder video, the video frame image is divided into multiple sub-blocks, and the Marr-Hildreth edge operator is used to calculate the number of texture edges of each sub-block. When the number of edges is lower than the preset threshold, the sub-block is determined to be occluded, and the similarity between consecutive frames is determined by SSIM to determine whether the video is continuously occluded. This type of method often performs poorly in dealing with lighting changes, noise interference and complex backgrounds in dynamic scenes, has low robustness, and edge detection and SSIM calculations may be time-consuming for processing each frame, especially in high-resolution and high-frame rate videos, which may affect real-time performance. In addition, some camera occlusion detection methods mainly use deep learning methods. In order to improve the detection effect, more complex models and strategies are designed. For example, in patent application number CN2021107937477, patent name A surveillance video image occlusion detection method based on deep learning, the proportion of the obstacle color in the color histogram is first used to make a judgment. If it exceeds a certain threshold, yolov5 is used to detect the occluded part by target detection. Patent application number CN2022117187838, patent name A camera occlusion detection method, device, storage medium and electronic device converts occlusion detection into a classification + semantic segmentation task. The accuracy of this type of algorithm is better than that of traditional methods, but most of them are more complex, and lack real-time and lightweight.
[0004] In addition, many current works adopt an end-to-end training method and only supervise the output of the last layer of the model. In this way, the model relies on the feedback information of the final layer to adjust the previous feature extraction layer layer by layer. It does not fully utilize the features of different levels in the model, ignores the contribution of the intermediate layer to specific targets (such as local details and edge information), and may cause false detection.
[0005] In practical applications, there are many types of occlusions and the occlusion scenes are diverse. Some non-occlusion scenes can be easily confused with the features of occlusion scenes (for example, the sky, open roads, etc.), which is a challenge for data collection and processing. Summary of the invention
[0006] The purpose of the present invention is to provide a camera occlusion state detection method that can fully utilize different levels of features in the model and has real-time performance, so as to improve the accuracy of detection.
[0007] The camera occlusion state detection method of the present invention comprises the following steps:
[0008] S1. Training phase: The image features with occluders and non-occluders are divided into three paths in the feature extraction Rep-Conv module. One path is passed through the 3×3 convolution and batch normalization layer through the weights and bias parameters to obtain the convolution output features corresponding to the branch. Another path is passed through the 1×1 convolution and batch normalization layer through the weights and bias parameters to obtain the convolution output features corresponding to the branch. Another path is obtained by the identity transformation to obtain the output features corresponding to the branch. The equivalent convolution kernel with the recognition target edge color texture and original image features is obtained by fusing the above three branches.
[0009] S2. After completing the training of the Rep-Conv module, pass its output features to the C2f-Star module, and train the C2f-Star module to obtain a C2f-Star module that has the ability to identify occluders and non-occluders and combine information from different dimensions;
[0010] S3. After completing the training of the C2f-Star module, pass its output features to the attention module, and train the attention module to obtain an attention module capable of identifying key feature information of occluders;
[0011] S4. After completing the training of the attention module, the output features thereof are passed to the final classifier, and the final classifier is trained to obtain a final classifier capable of identifying occluders and non-occluders;
[0012] S5, detection stage: input each acquired frame of image into the trained equivalent convolution kernel for convolution calculation, and obtain the feature map of the original image by combining the edge color texture of the occluder and non-occluder, information exchange between channels;
[0013] S6, the edge color texture, information exchange between channels, and the feature map of the original image features are divided into two parts according to the number of feature channels in the trained C2f-Star module, one part of which is subjected to multiple feature transformations by multiple Faster-Star modules to obtain a feature map reflecting the essence of the occluder, and the other part is subjected to the residual path to obtain the initial edge color texture of the occluder and non-occluder, information exchange between channels, and the feature map of the original image;
[0014] S7, the feature map reflecting the essence of the occluder and the initial edge color texture of the occluder and non-occluder are exchanged between channels, the feature map of the original image is integrated and spliced through the channels, and then the number of channels is adjusted to output the feature map of the occluder and non-occluder combining different dimensional information;
[0015] S8, inputting the feature map that integrates the information of the occluder and the non-occluder in different dimensions into the attention module, highlighting the key features of the occluder by weighted adjustment of each channel, and obtaining a feature map with the key feature information of the occluder;
[0016] S9. The final classifier calculates the occlusion score of each frame image through the feature map of the key feature information of the occlusion object, and determines whether the occlusion score is greater than the occlusion threshold. If so, the camera in the frame image is in an occluded state; otherwise, the camera in the frame image is in a non-occluded state.
[0017] The camera occlusion state detection method described in the present invention expands the 3×3 convolution into a multi-branch structure in the training stage to increase the nonlinearity and feature expression ability of the model, and in the detection stage, the multi-branch structure is fused into an equivalent convolution layer in the reasoning stage through re-parameterization, thereby achieving lightweight and real-time performance of the model while ensuring detection accuracy, thereby improving the accuracy of camera occlusion state detection; in addition, the feature map of the original image combined with edge color texture, information exchange between channels, is divided into two parts according to the number of feature channels in the trained C2f-Star module, one part of which is subjected to multiple feature extraction by multiple Faster-Star modules. The feature map that reflects the nature of the occlusion is transformed, and the other part is obtained through the residual path to obtain the initial edge color texture. The information is exchanged between channels, and the feature map of the original image is used to calculate the occlusion score of each frame of the image, so as to judge the occlusion state of the camera, so as to further make full use of the different levels of features in the model and improve the accuracy of the camera occlusion state detection; in addition, by using the feature map that combines information of different dimensions to weight the feature parameters of the channel through the attention module, a feature map of the key feature information of the occlusion is obtained, which can improve the model's attention to key features and suppress irrelevant information, thereby improving the discrimination and generalization capabilities of the camera occlusion state detection.
[0018] As a preferred solution of the present invention, in step S1, the convolution output feature corresponding to the branch is obtained through the 3×3 convolution and batch normalization layers through weights and bias parameters. The specific design is as follows: 3×3 =W 3×3 *x, where W 3×3 is the weight of the 3x3 convolution, x is the input image feature with occluders and non-occluders, y 3×3 is the feature map obtained after convolution; then, the convolutional features are adjusted by the batch normalization layer, and the output is: in, is the feature map of the output edge color texture, μ 3×3 , b 3×3 is the mean, variance, and bias parameter of the current feature map, γ 3×3 and β 3×3 are the learned parameters of the batch normalization layer.
[0019] As a preferred solution of the present invention, in step S1, the convolution output feature corresponding to the branch is obtained through the 1×1 convolution and batch normalization layer through weight and bias parameters, and the specific design is as follows: the 1x1 convolution is padded and expanded into a 3x3 feature form; is the weight of the 1×1 convolution; the calculation of the padded feature is as follows: Among them, x is the input image feature with occlusions and non-occlusions, y1×1 is the output feature map obtained by 1x1 convolution; then, the convolutional features are adjusted by the batch normalization layer, and the output is: in, Yes 1×1 The feature map with information exchange between feature channels after adjustment by the normalization layer, μ 1×1 , b 1×1 , are the mean, variance, and bias parameters of the current feature map, γ 1×1 and β 1×1 are the learned parameters of the batch normalization layer.
[0020] As a preferred solution of the present invention, in step S1, the output feature corresponding to the branch is obtained by identity transformation, and the specific design is as follows: the convolution kernel of a unit matrix is padded and expanded into a 3x3 feature form; is the weight of the identity transformation; the feature after padding is calculated as follows: Among them, x is the input image feature with occlusions and non-occlusions, y identity is the feature map output after the identity transformation; then, after adjustment by the batch normalization layer, the output is: in, is the feature map of the original image feature output, μ identity , b identity , are the mean, variance, and bias parameters of the current feature map, γ identity and β identity are the learned parameters of the batch normalization layer.
[0021] As a preferred solution of the present invention, there are multiple Rep-Conv modules and C2f-Star modules, two Rep-Conv modules are arranged near the input end of the image features, and the remaining Rep-Conv modules and C2f-Star modules are arranged at intervals and connected in series, and the C2f-Star module is used as the final output of the image features, and the Rep-Conv modules and the C2f-Star modules form feature extraction modules of different levels;
[0022] When the Rep-Conv module and the C2f-Star module are trained, the first C2f-Star module, the second C2f-Star module, and the third C2f-Star module behind the Rep-Conv module are respectively connected to the first auxiliary classifier, the second auxiliary classifier, and the third auxiliary classifier;
[0023] Calculate a first difference value between the occlusion score and non-occlusion probability score calculated by the final classifier and the occlusion label and non-occlusion label corresponding to the original image through a loss function;
[0024] The first auxiliary classifier calculates the occlusion score and non-occlusion score of each frame image through the key feature information feature map of the occluder at this stage;
[0025] The second auxiliary classifier calculates the occlusion score and non-occlusion score of each frame image through the key feature information feature map of the occluder at this stage;
[0026] The third auxiliary classifier calculates the occlusion score and non-occlusion score of each frame image through the key feature information feature map of the occluder at this stage;
[0027] The occlusion scores and non-occlusion probability scores calculated by the first auxiliary classifier, the second auxiliary classifier, and the third auxiliary classifier are respectively compared with the occlusion labels and non-occlusion labels corresponding to the original image to calculate the corresponding difference values through the loss function;
[0028] The occlusion scores and non-occlusion scores calculated by the first auxiliary classifier, the second auxiliary classifier, and the third auxiliary classifier are respectively compared with the occlusion scores and non-occlusion scores calculated by the final classifier to calculate the corresponding difference values through the loss function;
[0029] All difference values are weighted and combined to calculate the final difference value;
[0030] Based on the final difference value, gradient back propagation is performed to update the Rep-Conv module, C2f-Star module, attention module, weights and bias parameters of each classifier. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a structural schematic diagram of the training phase of the camera occlusion state detection method of the present invention;
[0032] Figure 2 Schematic diagram of the Rep-Conv module divided into three training paths;
[0033] Figure 3 It is a structural schematic diagram of the detection phase of the camera occlusion state detection method of the present invention;
[0034] Figure 4 Schematic diagram of the process of the camera occlusion state detection method of the present invention;
[0035] Figure 5 Schematic diagram of the process of performing multiple feature transformations for the C2f-Star module;
[0036] Figure 6 Schematic diagram of using various types of occluders. DETAILED DESCRIPTION
[0037] The following will be combined with the accompanying drawings to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0038] The structural diagram of the training phase of camera occlusion state detection is as follows: Figure 1 As shown, the image features with occluders and non-occluders are mainly processed by the Rep-Conv (heavy parameter) module and divided into three-way processing. The feature map of edge color texture, exchange between channels, and original image features is mainly processed by the C2f-Star module according to the number of feature channels. There are multiple Rep-Conv modules and C2f-Star modules. Two Rep-Conv modules are set near the input end of the image features. The remaining Rep-Conv modules and C2f-Star modules are set at intervals and connected in series. The C2f-Star module is used as the final output of the image features. The Rep-Conv module and the C2f-Star module form feature extraction modules of different levels. The output of the feature extraction module is connected to the attention module, and the output of the attention module is connected to the final classifier (fully connected layer). In addition, when the Rep-Conv module and the C2f-Star module are trained, the first C2f-Star module, the second C2f-Star module, and the third C2f-Star module after the Rep-Conv module are respectively connected to the first auxiliary classifier, the second auxiliary classifier, and the third auxiliary classifier.
[0039] A camera occlusion state detection method, such as Figure 3 As shown, S1, training stage: the image features with occlusions and non-occlusions are divided into three paths in the feature extraction Rep-Conv module, one path passes through the 3×3 convolution and batch normalization layer through the weight and bias parameters to obtain the convolution output features corresponding to the branch, the other path passes through the 1×1 convolution and batch normalization layer through the weight and bias parameters to obtain the convolution output features corresponding to the branch, and the other path obtains the output features corresponding to the branch through the identity transformation. The equivalent convolution kernel with the ability to identify the edge color texture and original image features of the target is obtained by fusing the above three branches.
[0040] like Figure 2 As shown in Figure 1, the convolution output features corresponding to this branch are obtained through the weight and bias parameters of the 3×3 convolution and batch normalization layers. The specific design is as follows: 3×3 =W 3×3 *x, where W 3×3is the weight of the 3x3 convolution, x is the input image feature with occluders and non-occluders, y 3×3 is the feature map obtained after convolution; then, the convolutional features are adjusted by the batch normalization layer, and the output is: in, is the feature map of the output edge color texture, μ 3×3 , b 3×3 is the mean, variance, and bias parameter of the current feature map, γ 3×3 and β 3×3 are the learned parameters of the batch normalization layer.
[0041] After the 1×1 convolution and batch normalization layers, the convolution output features corresponding to the branch are obtained through weight and bias parameters. The specific design is as follows: the 1x1 convolution is padded and expanded into a 3x3 feature form; is the weight of the 1×1 convolution; the calculation of the padded feature is as follows: Among them, x is the input image feature with occlusions and non-occlusions, y 1×1 is the output feature map obtained by 1x1 convolution; then, the convolutional features are adjusted by the batch normalization layer, and the output is:
[0042] in, Yes 1×1 The feature map with information exchange between feature channels after adjustment by the normalization layer, μ 1×1 , b 1×1 , are the mean, variance, and bias parameters of the current feature map, γ 1×1 and β 1×1 are the learned parameters of the batch normalization layer.
[0043] The output features corresponding to this branch are obtained through identity transformation. The specific design is as follows: the convolution kernel of a unit matrix is padded and expanded into a 3x3 feature form; is the weight of the identity transformation; the feature after padding is calculated as follows: Among them, x is the input image feature with occlusions and non-occlusions, y identity is the feature map output after the identity transformation; then, after adjustment by the batch normalization layer, the output is: in, is the feature map of the original image feature output, μ identity , b identity , are the mean, variance, and bias parameters of the current feature map, γ identity and β identityare the learned parameters of the batch normalization layer.
[0044] The weights and bias parameters of the three branches are added together to obtain the fused convolution kernel and bias parameters, so as to fuse the three branches of 3x3 convolution, 1x1 convolution and identity transformation in the training phase into an equivalent convolution kernel with weights and bias parameters for identifying the edge color texture features of occluders and non-occluders and the original image features. In this way, while maintaining the multi-branch structure's good learning ability for complex feature distribution, the model structure is kept lightweight. The three-branch fusion formula is as follows: The weight and bias parameters of the equivalent convolution kernel can be expressed as follows:
[0045]
[0046] The output of the equivalent convolution kernel after fusion is: fused =W fused *x+b fused , where W fused is the weight of the equivalent convolution kernel, b fused is the bias parameter of the equivalent convolution kernel.
[0047] S2. After completing the training of the Rep-Conv module, pass its output features to the C2f-Star module, and train the C2f-Star module to obtain a C2f-Star module that has the ability to identify occluders and non-occluders and combine information of different dimensions.
[0048] S3. After completing the training of the C2f-Star module, pass its output features to the attention module, and train the attention module to obtain an attention module capable of identifying key feature information of occluders.
[0049] S4. After completing the training of the attention module, the features of its output are passed to the final classifier, and the final classifier is trained to obtain a final classifier capable of identifying occluders and non-occluders.
[0050] When training the Rep-Conv module and the C2f-Star module, the first C2f-Star module, the second C2f-Star module and the third C2f-Star module after the Rep-Conv module are connected to the first auxiliary classifier, the second auxiliary classifier and the third auxiliary classifier respectively; the occlusion score and the non-occlusion probability score calculated by the final classifier are used to calculate the first difference value with the occlusion label and the non-occlusion label corresponding to the original image through the loss function; the first auxiliary classifier calculates the occlusion score and the non-occlusion score of each frame of the image through the feature map of the key feature information of the occlusion object at this stage; the second auxiliary classifier calculates the occlusion score and the non-occlusion score of each frame of the image through the feature map of the key feature information of the occlusion object at this stage; the third auxiliary classifier calculates the occlusion score and the non-occlusion score of each frame of the image through the feature map of the key feature information of the occlusion object at this stage; The occlusion score and non-occlusion score of each frame image are calculated from the feature map of the key feature information of the stage occlusion object; the occlusion scores and non-occlusion scores calculated by the first auxiliary classifier, the second auxiliary classifier and the third auxiliary classifier are respectively compared with the occlusion labels and non-occlusion labels corresponding to the original image through the loss function to calculate the corresponding difference values; the occlusion scores and non-occlusion scores calculated by the first auxiliary classifier, the second auxiliary classifier and the third auxiliary classifier are respectively compared with the occlusion scores and non-occlusion scores calculated by the final classifier through the loss function to calculate the corresponding difference values; all the difference values are weighted and combined to calculate the final difference value; according to the final difference value, gradient backpropagation is performed to update the weights and bias parameters of the Rep-Conv module, the C2f-Star module, the attention module, and each classifier.
[0051] The loss function consists of two parts: the supervised loss of the final classifier and the supervised loss of the auxiliary classifier. The loss function is defined as follows: Among them, L final is the supervision loss between the final classifier and the true label (the occluded label and the non-occluded label corresponding to the original image); is the supervision loss of the i auxiliary classifiers, which is used to guide the feature learning of the middle layer of the feature extraction module; α is the adjustment parameter to control the proportion of the auxiliary supervision loss in the total loss; N is the number of auxiliary classification heads. The loss function formula of the final classifier is expressed as follows:
[0052]
[0053] Among them, L ce is the cross entropy, p final is the final classification output of the model (occlusion score and non-occlusion probability score), and y is the true label. By minimizing this loss, the model learns the correct classification decision.
[0054] The loss function of the auxiliary classifier consists of two parts, one is the supervision of the real label, and the other is the predicted distribution of the classifier in S4 as the supervision of the soft label. The formula is as follows: Among them, p i is the predicted probability distribution of the output of the i-th auxiliary classifier (there are three in total); y is the true label, which is in the form of a one-hot distribution; is the soft label output by the final classifier, expressed in the form of probability distribution; β is the adjustment parameter that balances the weight of real label supervision and soft label supervision; L ce (p i ,y) is the cross entropy loss, which is used to constrain the consistency between the auxiliary classifier prediction distribution and the true label. is the KL divergence, which measures the difference between the predicted distribution and the soft label. The formula for KL divergence is as follows, In the actual prediction stage of the model, no auxiliary classifier is used. Instead, the end-to-end model structure is used directly to output the prediction results, thus ensuring the lightweight and real-time performance of the model.
[0055] By adding an auxiliary classifier after the output of the three C2f-Star modules in the backbone of the model (corresponding to 8x, 16x, and 32x downsampling), the prediction results of the features with different downsampling ratios are output, and the real labels are used as supervision. In this way, the middle layer of the model can extract more detailed and hierarchical features during training. In particular, the 8x and 16x downsampling layers can usually capture the edges and occlusion features of the occluded targets, while the 32x downsampling layers often contain richer global information and are more suitable for capturing the overall contour and scene features. This hierarchical supervision design improves the model's adaptability to different occlusion levels and different scale features, and can effectively enhance the accuracy of occlusion detection. In addition, this training method can enhance the flow of intermediate layer gradients and effectively alleviate the problem of gradient vanishing when the number of network layers is deep.
[0056] In addition, by introducing the prediction distribution of the final classifier as the output of the teacher model to guide the prediction distribution of the auxiliary classifier, the auxiliary classifier can better learn features that meet the final task objectives. This approach not only balances the optimization direction of the auxiliary classifier and the final classifier, but also enhances the coordination of global features and further improves the generalization ability of the model. In addition, the final classifier of the model is still optimized using the real label, which guarantees the performance of the final task as a whole.
[0057] For the collection of positive samples (occluders), the present invention uses various types of occluders, including towels and rags, leaves, paper, plastic bags, water bottles, limbs, small pendants in cars, etc. Figure 6As shown. The collection process is as follows: 1. Use various occluders (including direct occlusion with the palm of your hand) to block the lens from all directions (up, down, left, and right). Block three times in each direction, and each occlusion lasts about 3 seconds. It is considered an occlusion only when 80% to 90% of the picture is blocked, but be careful not to block the picture to the point of complete blackness, otherwise there will be no features to learn. 2. After collecting fixed occlusions, you can use occluders for moving occlusions (simulating the relative movement of occluders and the lens). Specifically, use the occluder to slowly translate left and right and up and down in front of the lens. The picture is considered as a translation from unobstructed state → completely obstructed → unobstructed. A total of two translations are performed, one from left to right and one from top to bottom, and each translation lasts about 2 to 3 seconds.
[0058] For the collection of negative samples (non-occluded objects), the present invention is divided into two types: conventional negative samples and difficult negative samples. Conventional negative samples refer to normal, unobstructed or partially obstructed images taken by the lens while driving and in a stationary state. In actual applications, the texture of some non-occluded scenes is very close to the occlusion, for example, large tracts of sky or large tracts of open road appear in the picture; scenes with very low illumination or even close to darkness; scenes with more motion blur during driving; scenes where the lens is covered with mud spots and water droplets, etc. The present invention performs special sampling for the above-mentioned scenes, thereby greatly improving the stability of the present invention in various usage scenarios.
[0059] S5, such as Figure 4 As shown, in the detection stage, each acquired frame of the image is input into the trained equivalent convolution kernel for convolution calculation to obtain a feature map of the original image by combining the color texture of the edges of the occluders and non-occluders, exchanging information between channels, and obtaining the feature map of the original image.
[0060] S6. The feature map of the original image combined with edge color texture, information exchange between channels, is divided into two parts according to the number of feature channels in the trained C2f-Star module. One part is transformed multiple times by multiple Faster-Star modules to obtain a feature map reflecting the essence of the occluder, and the other part is obtained through the residual path to obtain the initial edge color texture of the occluder and non-occluder, information exchange between channels, and the feature map of the original image.
[0061] like Figure 5 As shown in the figure, the Faster-Star module is optimized for feature extraction efficiency and nonlinear enhancement within the C2f-Star module. The structure and functions of the Faster-Star module are as follows:
[0062] 1. Partial channel convolution (PwConv)
[0063] The Faster-Star module first uses a PwConv (partial channel convolution) to downsample the input features. PwConv is a variant of DwConv (separate convolution). Compared with DwConv, PwConv only convolves part of the channels and reconcatenates the convolved features with the channel features that have not been convolved to maintain the integrity of the information.
[0064] 2. Channel integration and splicing
[0065] The concatenated features are integrated through two parallel 1×1 convolutions. One branch is activated by a ReLU function after the 1×1 convolution to introduce nonlinear expression capabilities. The other branch is only 1×1 convolution without activation to retain linear information.
[0066] 3. Dot multiplication operation
[0067] The output of the two branches will perform a dot product operation, and the dot product formula is as follows:
[0068]
[0069] Among them, W 1 is the weight of the convolution kernel of the first branch, W 2 is the weight of the convolution kernel of the second branch, x is the input feature, and d represents the dimension of the feature.
[0070] Through the dot multiplication operation, the model can obtain a kernel function that maps low-dimensional features to high-dimensional nonlinear features, learn the correlation between feature channels, improve the model's understanding of the overall information of the feature map, and enhance the generalization of the model.
[0071] 4. Dimension Adjustment
[0072] The number of channels is adjusted through a 1×1 convolution in the point product result to adapt to the subsequent network structure requirements.
[0073] S7, the feature map reflecting the essence of the occluder and the initial edge color texture of the occluder and non-occluder, the information exchange between channels, the feature map of the original image is integrated and spliced through the channels, and then the number of channels is adjusted to output the feature map of the occluder and non-occluder combining different dimensional information.
[0074] S8. Input the feature map that integrates the information of occluders and non-occluders in different dimensions into the attention module, and highlight the key features of the occluders by weighted adjustment of each channel to obtain a feature map with the key feature information of the occluders.
[0075] S9. The final classifier calculates the occlusion score of each frame image through the feature map of the key feature information of the occlusion object, and determines whether the occlusion score is greater than the occlusion threshold. If so, the camera in the frame image is in an occluded state; otherwise, the camera in the frame image is in a non-occluded state.
[0076] The definition of whether there is occlusion in a scene is: if the occlusion area (occlusion score) accounts for more than 80% (occlusion threshold) of the picture, it is defined as occlusion, otherwise it is not considered as occlusion.
[0077] The above embodiments are only used to illustrate the detailed scheme of the present invention, and the present invention is not limited to the above detailed scheme, that is, it does not mean that the present invention must rely on the above detailed scheme to be implemented. Those skilled in the art should understand that any improvement of the present invention, equivalent replacement of the raw materials of the product of the present invention, addition of auxiliary components, selection of specific methods, etc., are all within the protection scope and disclosure scope of the present invention.
Claims
1. A camera occlusion state detection method, comprising the following steps: S1. Training phase: The image features with occluders and non-occluders are divided into three paths in the feature extraction Rep-Conv module. One path is passed through the 3×3 convolution and batch normalization layer through the weights and bias parameters to obtain the convolution output features corresponding to the branch. Another path is passed through the 1×1 convolution and batch normalization layer through the weights and bias parameters to obtain the convolution output features corresponding to the branch. Another path is obtained by the identity transformation to obtain the output features corresponding to the branch. The equivalent convolution kernel with the recognition target edge color texture and original image features is obtained by fusing the above three branches. S2. After completing the training of the Rep-Conv module, pass its output features to the C2f-Star module, and train the C2f-Star module to obtain a C2f-Star module that has the ability to identify occluders and non-occluders and combine information from different dimensions; S3. After completing the training of the C2f-Star module, pass its output features to the attention module, and train the attention module to obtain an attention module capable of identifying key feature information of occluders; S4. After completing the training of the attention module, the output features thereof are passed to the final classifier, and the final classifier is trained to obtain a final classifier capable of identifying occluders and non-occluders; S5, detection stage: input each acquired frame of image into the trained equivalent convolution kernel for convolution calculation, and obtain the feature map of the original image by combining the edge color texture of the occluder and non-occluder, information exchange between channels; S6. The feature map of the original image is divided into two parts according to the number of feature channels in the trained C2f-Star module by combining edge color texture, information exchange between channels, and one part is subjected to multiple feature transformations by multiple Faster-Star modules to obtain a feature map reflecting the essence of the occluder, and the other part is subjected to the residual path to obtain the initial edge color texture of the occluder and non-occluder, information exchange between channels, and the feature map of the original image; S7, the feature map reflecting the essence of the occluder and the initial edge color texture of the occluder and non-occluder are exchanged between channels, the feature map of the original image is integrated and spliced through channels, and then the number of channels is adjusted to output the feature map that integrates the information of the occluder and non-occluder in different dimensions; S8, inputting the feature map that integrates the information of the occluder and the non-occluder in different dimensions into the attention module, highlighting the key features of the occluder by weighted adjustment of each channel, and obtaining a feature map with the key feature information of the occluder; S9. The final classifier calculates the occlusion score of each frame image through the feature map of the key feature information of the occlusion object, and determines whether the occlusion score is greater than the occlusion threshold. If so, the camera in the frame image is in an occluded state; otherwise, the camera in the frame image is in a non-occluded state.
2. The camera occlusion state detection method according to claim 1, characterized in that: In step S1, the convolution output features corresponding to the branch are obtained through the 3×3 convolution and batch normalization layers through weight and bias parameters. The specific design is as follows: 3×3 =W 3×3 *x, where W 3×3 is the weight of the 3x3 convolution, x is the input image feature with occluders and non-occluders, y 3×3 is the feature map obtained after convolution; then, the convolutional features are adjusted by the batch normalization layer, and the output is: in, is the feature map of the output edge color texture, μ 3×3 , b 3×3 is the mean, variance, and bias parameter of the current feature map, γ 3×3 and β 3×3 are the learned parameters of the batch normalization layer.
3. The camera occlusion state detection method according to claim 1, characterized in that: In step S1, the convolution output features corresponding to the branch are obtained through the weight and bias parameters of the 1×1 convolution and batch normalization layers. The specific design is as follows: the 1x1 convolution is padded and expanded into a 3x3 feature form; is the weight of the 1×1 convolution; The calculation of the features after padding is as follows: Among them, x is the input image feature with occlusions and non-occlusions, y 1×1 is the output feature map obtained by 1x1 convolution; then, the convolutional features are adjusted by the batch normalization layer, and the output is: in, Yes 1×1 The feature map with information exchange between feature channels after adjustment by the normalization layer, μ 1×1 , b 1×1 , are the mean, variance, and bias parameters of the current feature map, γ 1×1 and β 1×1 are the learned parameters of the batch normalization layer.
4. The camera occlusion state detection method according to claim 1, characterized in that: In step S1, the output feature corresponding to the branch is obtained by identity transformation, and the specific design is as follows: the convolution kernel of a unit matrix is padded and expanded into a 3x3 feature form; is the weight of the identity transformation; the feature after padding is calculated as follows: Among them, x is the input image feature with occlusions and non-occlusions, y identity is the feature map output after the identity transformation; then, after adjustment by the batch normalization layer, the output is: in, is the feature map of the original image feature output, μ identity , b identity , are the mean, variance, and bias parameters of the current feature map, γ identity and β identity are the learned parameters of the batch normalization layer.
5. The camera occlusion state detection method according to claim 1, characterized in that: There are multiple Rep-Conv modules and C2f-Star modules. Two Rep-Conv modules are set near the input end of the image features. The remaining Rep-Conv modules and C2f-Star modules are set at intervals and connected in series. The C2f-Star module is used as the final output of the image features. The Rep-Conv modules and the C2f-Star modules form feature extraction modules of different levels. When the Rep-Conv module and the C2f-Star module are trained, the first C2f-Star module, the second C2f-Star module, and the third C2f-Star module behind the Rep-Conv module are respectively connected to the first auxiliary classifier, the second auxiliary classifier, and the third auxiliary classifier; Calculate a first difference value between the occlusion score and non-occlusion score calculated by the final classifier and the occlusion label and non-occlusion label corresponding to the original image through a loss function; The first auxiliary classifier calculates the occlusion score and non-occlusion score of each frame image through the key feature information feature map of the occluder at this stage; The second auxiliary classifier calculates the occlusion score and non-occlusion score of each frame image through the key feature information feature map of the occluder at this stage; The third auxiliary classifier calculates the occlusion score and non-occlusion score of each frame image through the key feature information feature map of the occluder at this stage; The occlusion scores and non-occlusion scores calculated by the first auxiliary classifier, the second auxiliary classifier, and the third auxiliary classifier are respectively compared with the occlusion labels and non-occlusion labels corresponding to the original image to calculate the corresponding difference values through the loss function; The occlusion scores and non-occlusion scores calculated by the first auxiliary classifier, the second auxiliary classifier, and the third auxiliary classifier are respectively compared with the occlusion scores and non-occlusion scores calculated by the final classifier to calculate the corresponding difference values through the loss function; All difference values are weighted and combined to calculate the final difference value; Based on the final difference value, gradient back propagation is performed to update the Rep-Conv module, C2f-Star module, attention module, weights and bias parameters of each classifier.
Citation Information
Patent Citations
Surface defect detection method for blue film coated lithium battery
CN118247246A
Model training method and device, target detection method and device and electronic equipment
CN118521856A