Vehicle blind area monitoring method

By constructing the CWF-SSD model, using the cascading attention mechanism and weighted fusion module, the problem of inaccurate identification of small obstacles in blind spot monitoring technology is solved, and high-precision identification of small targets in vehicle blind spots is achieved.

CN120014561APending Publication Date: 2025-05-16ANHUI POLYTECHNIC UNIV MECHANICAL & ELECTRICAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411868231.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing blind spot monitoring technology can easily filter out smaller obstacles and cannot accurately identify obstacles.

Method used

The CWF-SSD model is constructed, including the VGG16 network, a cascading attention mechanism and a weighted fusion module, and the model is trained and tested by preprocessing and marking the vehicle image data, and finally used to monitor the vehicle blind spots in real time.

Benefits of technology

It improves the recognition accuracy of small targets within the vehicle's blind spot and enhances the accurate recognition ability of obstacles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014561A_ABST
    Figure CN120014561A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle blind area monitoring, in particular to a vehicle blind area monitoring method which comprises the steps that a cascade weighted fusion CWF-SSD model is constructed, and the CWF-SSD model comprises a VGG16 network, a cascade attention mechanism and a weighted fusion module. Wherein the VGG16 network is used for feature extraction; the cascaded attention machine model structure can enhance target feature extraction and feature representation, and target features are processed and grouped in a cascaded manner; the weighted fusion module is used for integrating feature information from different scales; and processing real-time image data acquired by a high-definition camera on the vehicle based on the final CWF-SSD model to monitor the blind area of the vehicle. According to the invention, the cascaded attention mechanism and the weighted fusion module are introduced, so that the recognition precision of the small target in the vehicle blind area range can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle blind spot monitoring, and in particular to a vehicle blind spot monitoring method. Background Art

[0002] Blind Spot Detection (BSD) is an important feature in automotive safety systems, designed to help drivers detect areas around the vehicle that are out of sight, known as "blind spots". This technology uses sensors to monitor the vehicle's surroundings. When other vehicles, pedestrians or objects enter the blind spot, the system warns the driver visually or audibly to reduce the risk of collision.

[0003] In traditional blind spot monitoring technology, the mainstream sensor is millimeter wave radar. Millimeter wave radar has a long detection distance, strong anti-interference ability, and will not be affected by extreme weather. However, the electromagnetic waves emitted by millimeter wave radar can easily filter out smaller obstacles and cannot accurately identify obstacles. Summary of the invention

[0004] In view of this, the purpose of the present invention is to provide a vehicle blind spot monitoring method to solve the problem that the current blind spot monitoring technology easily filters out smaller obstacles and cannot accurately identify obstacles.

[0005] Based on the above purpose, the present invention provides a method for monitoring a vehicle blind spot, comprising:

[0006] S1: Acquire vehicle image data of various scenes, preprocess the vehicle image data, use a bounding box to mark the position of the vehicle on the preprocessed vehicle image data, and divide the marked vehicle image data into a training set and a test set;

[0007] S2: Construct a CWF-SSD model, which includes a VGG16 network, a cascaded attention mechanism and a weighted fusion module, wherein the VGG16 network is used for feature extraction; the cascaded attention machine model structure can enhance target feature extraction and feature representation, and process and group target features in a cascade form, so that the network can selectively focus on different parts of the input features to obtain more discriminative feature representations; the weighted fusion module is used to integrate feature information from different scales;

[0008] S3: training the CWF-SSD model using the training set to obtain a trained CWF-SSD model;

[0009] S4: Testing the trained CWF-SSD model using the test set, returning to step S3 if the test fails, and using the CWF-SSD model that passes the test as the final CWF-SSD model if the test fails.

[0010] S5: Based on the final CWF-SSD model, the real-time image data acquired by the high-definition camera on the vehicle is processed to monitor the vehicle blind spot.

[0011] Optionally, the preprocessing of the vehicle image data includes: performing preprocessing operations such as normalization and scaling on the vehicle image data.

[0012] Optionally, the weighted fusion module integrates feature information from different scales through the following function:

[0013] W c1 ,W c2 =Conv1(S c ),Conv2(S c )

[0014]

[0015] where x = c, s

[0016] Output=(W c1 +W s1 )*T1+(W c2 +W s2 )*T2

[0017] Where S c represents cascaded attention features, Conv1(), Conv2() represent one-dimensional convolution, W c ,W s Represent spatial features and channel features respectively; W x1 ,W x2 It represents the important feature information of spatial dimension and channel dimension calculated by softmax method; T1 and T2 represent the features of cascaded submodules, output is the fused features, and e represents the base of natural logarithm.

[0018] Optionally, the convolution operation is performed through a convolution module, and the cascade attention mechanism is composed of a grouping mechanism, a cascade processing unit and a feature combination module, which weakens background features and captures important information.

[0019] Optionally, the VGG16 network is a CNN architecture for feature extraction and classification.

[0020] Optionally, the CWF-SSD model also includes feature selection and multi-scale information capture to suppress interference features.

[0021] In this method, a cascaded attention mechanism and a weighted fusion module are added by constructing a CWF-SSD model. First, the cascaded attention mechanism enhances the diversity of the attention map by assigning a unique feature subset to each head. Similar to grouped convolution, the cascaded grouped attention mechanism can reduce the amount of computation and model parameters by h times, because the number of input and output channels in the convolution layer is reduced by h times. Second, adding cascaded attention heads can deepen the network structure, thereby increasing the capacity of the model without adding any additional parameters. Since the attention map in each head is calculated based on a smaller channel dimension, it only produces a relatively small delay. Feature information from different scales of spatial and channel enhancements enters the weighted fusion module for adaptive fusion, which determines the importance of features at different scales by learning content-related weight maps, thereby weakening the impact of low-quality depth maps. At the same time, the weighted fusion module enhances the nonlinear representation ability of the network through nonlinear feature enhancement units to ensure expressiveness during the fusion process. The final CWF-SSD model is obtained by training and testing the CWF-SSD model, and then deployed and applied to a real vehicle. The model monitors the blind spots of the vehicle through a high-definition camera.

[0022] From the above, it can be seen that by introducing the cascaded attention mechanism and weighted fusion, the optimized final CWF-SSD model is applied to the blind spot monitoring technology, which can improve the recognition accuracy of small targets within the vehicle blind spot. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0024] Figure 1 It is a structural block diagram of the CWF-SSD model of an embodiment of the present invention;

[0025] Figure 2 This is a detection effect diagram of an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with specific embodiments.

[0027] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should be understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0028] like Figure 1 As shown, a method for monitoring a vehicle blind spot includes:

[0029] S1: Acquire vehicle image data of various scenes, preprocess the vehicle image data, use a bounding box to mark the position of the vehicle on the preprocessed vehicle image data, and divide the marked vehicle image data into a training set and a test set;

[0030] S2: Construct a CWF-SSD model, which includes a VGG16 network, a cascade attention mechanism (CAM) and weighted fusion (WF). The VGG16 network is used for feature extraction. The cascade attention mechanism model structure can enhance target feature extraction and feature representation, and process and group target features in a cascade form, so that the network can selectively focus on different parts of the input features, so that it can obtain a more discriminative feature representation; the weighted fusion module is used to integrate feature information from different scales;

[0031] S3: training the CWF-SSD model using the training set to obtain a trained CWF-SSD model;

[0032] S4: Testing the trained CWF-SSD model using the test set, returning to step S3 if the test fails, and using the CWF-SSD model that passes the test as the final CWF-SSD model if the test fails.

[0033] S5: Based on the final CWF-SSD model, the real-time image data acquired by the high-definition camera on the vehicle is processed to monitor the vehicle blind spot.

[0034] In this method, a cascaded attention mechanism and a weighted fusion module are added by constructing a CWF-SSD model. First, the cascaded attention mechanism enhances the diversity of the attention map by assigning a unique feature subset to each head. Similar to grouped convolution, the cascaded grouped attention mechanism can reduce the amount of computation and model parameters by h times, because the number of input and output channels in the convolution layer is reduced by h times. Second, adding cascaded attention heads can deepen the network structure, thereby increasing the capacity of the model without adding any additional parameters. Since the attention map in each head is calculated based on a smaller channel dimension, it only produces a relatively small delay. Feature information from different scales of spatial and channel enhancements enters the weighted fusion module for adaptive fusion, which determines the importance of features at different scales by learning content-related weight maps, thereby weakening the impact of low-quality depth maps. At the same time, the weighted fusion module enhances the nonlinear representation ability of the network through nonlinear feature enhancement units to ensure expressiveness during the fusion process. The final CWF-SSD model is obtained by training and testing the CWF-SSD model, and then deployed and applied to a real vehicle. The model monitors the blind spots of the vehicle through a high-definition camera.

[0035] From the above, it can be seen that by introducing the cascaded attention mechanism and weighted fusion, the optimized final CWF-SSD model is applied to the blind spot monitoring technology, which can improve the recognition accuracy of small targets within the vehicle blind spot.

[0036] In some embodiments, the preprocessing of the vehicle image data includes: performing preprocessing operations of normalizing, flipping, rotating, and scaling the vehicle image data.

[0037] In some embodiments, the weighted fusion module integrates feature information from different scales through the following function:

[0038] W c1 ,W c2 =Conv1(S c ),Conv2(S c )

[0039]

[0040] where x = c, s

[0041] Output=(W c1 +W s1 )*T1+(W c2 +W s2 )*T2

[0042] Where S crepresents cascaded attention features, Conv1(), Conv2() represent one-dimensional convolution, W c ,W s Represent spatial features and channel features respectively; W x1 ,W x2 It represents the important feature information of spatial dimension and channel dimension calculated by softmax method; T1 and T2 represent the features of cascaded submodules, output is the fused features, and e represents the base of natural logarithm.

[0043] By redesigning the basic structure of the CWF-SSD model, the network model can maximize feature extraction of the hidden layers in the network and fully share contextual information.

[0044] In some embodiments, the VGG16 network is a CNN architecture for feature extraction and classification.

[0045] In some embodiments, the CWF-SSD model also includes a feature selection and multi-scale information capture module for suppressing interfering features.

[0046] like Figure 2 As shown in the figure, in order to clearly demonstrate the effectiveness of the CWF-SSD algorithm, the experimental data of the reproduced SSD algorithm and the CWF-SSD algorithm are compared, and some images are randomly selected from the test set and recognized using the SSD algorithm and the CWF-SSD algorithm respectively. Figure 2 The detection results are shown, where Figures (a) and (c) represent the detection results of the SSD algorithm, while Figures (b) and (d) show the detection results of the CWF-SSD algorithm.

[0047] observe Figure 2 It can be found that in each set of comparison images, targets of the same category are marked by boxes of the same color. Overall, in scenes with simpler backgrounds, both the SSD algorithm and the CWF-SSD algorithm can accurately locate targets and identify their categories, but the CWF-SSD algorithm has a higher confidence score and more accurate target positioning. In scenes with more complex backgrounds, as shown in Image1 and Image2, the original SSD algorithm misses buses and cars, and has poor recognition effects on occluded targets. In contrast, the CWF-SSD algorithm can not only effectively identify vehicles in images, but also has a higher accuracy rate in detecting people. Therefore, applying the optimized SSD model to blind spot monitoring technology can improve the recognition accuracy of small targets within the vehicle blind spot.

[0048] It should be understood by those skilled in the art that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present invention (including the claims) is limited to these examples. Within the scope of the present invention, the technical features in the above embodiments or in different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the above aspects of the present invention, which are not provided in detail for the sake of simplicity.

[0049] The present invention is intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for monitoring a vehicle blind spot, characterized in that: include: S1: Acquire vehicle image data of various scenes, preprocess the vehicle image data, use a bounding box to mark the position of the vehicle on the preprocessed vehicle image data, and divide the marked vehicle image data into a training set and a test set; S2: Construct a CWF-SSD model, which includes a VGG16 network, a cascaded attention mechanism and a weighted fusion module, wherein the VGG16 network is used for feature extraction, the cascaded attention machine model structure is used to enhance target feature extraction and feature representation, and the cascaded features are processed and grouped so that the network can selectively focus on different parts of the input features to obtain discriminative feature representations, and the weighted fusion module is used to integrate feature information from different scales; S3: training the CWF-SSD model using the training set to obtain a trained CWF-SSD model; S4: Testing the trained CWF-SSD model using the test set, returning to step S3 if the test fails, and using the CWF-SSD model that passes the test as the final CWF-SSD model if the test fails. S5: Based on the final CWF-SSD model, the real-time image data acquired by the high-definition camera on the vehicle is processed to monitor the vehicle blind spot.

2. The method for monitoring a vehicle blind spot according to claim 1, characterized in that: The preprocessing of the vehicle image data includes: performing normalization and scaling preprocessing operations on the collected vehicle blind spot image data.

3. The method for monitoring a vehicle blind spot according to claim 1, characterized in that: The weighted fusion module integrates feature information from different scales through the following function: W c1 ,W c2 =Conv1(S c ),Conv2(S c ) where x = c, s Output=(W c1 +W s1 )*T1+(W c2 +W s2 )*T2 Where S c represents cascaded attention features, Conv1(), Conv2() represent one-dimensional convolution, W c ,W s Represent spatial features and channel features respectively; W x1 ,W x2 It represents the important feature information of spatial dimension and channel dimension calculated by softmax method; T1 and T2 represent the features of cascaded submodules, output is the fused features, and e represents the base of natural logarithm.

4. The method for monitoring a vehicle blind spot according to claim 3, characterized in that: The cascade attention mechanism is composed of a grouping mechanism, a cascade processing unit and a feature combination module.

5. The method for monitoring a vehicle blind spot according to claim 1, characterized in that: The VGG16 network is a CNN architecture used for feature extraction and classification.

6. The method for monitoring a vehicle blind spot according to claim 1, characterized in that: The CWF-SSD model also includes feature selection and multi-scale information capture modules for suppressing interfering features.