Unmanned multi-sensor fusion method and system based on attention mechanism
By adopting a multi-sensor fusion method based on attention mechanism in unmanned driving technology, the fuzzy algorithm is used to improve the cross attention mechanism, and the problem of limited three-dimensional object detection performance under a single sensor data input is solved, achieving more accurate and comprehensive feature fusion, and improving the efficiency and accuracy of target recognition.
Patent Information
- Application Number
- CN202510036816.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-06-10
AI Technical Summary
In the existing unmanned driving technology, under the input of a single sensor data, the three-dimensional object detection performance is limited, and the multi-sensor feature fusion is not sufficient, making it difficult to obtain stronger semantic features at the instance level, resulting in uneven fusion of features such as long-distance objects and small targets.
The unmanned driving multi-sensor fusion method based on attention mechanism is adopted, and the cross-attention mechanism is improved through the fuzzy algorithm, combined with the advantages of the fuzzy algorithm and attention mechanism, an improved fuzzy attention mechanism algorithm is formed to realize the feature fusion of point cloud data and image data.
A more accurate and comprehensive feature fusion is achieved, improving the efficiency and accuracy of target recognition, especially in the recognition of long-distance objects and small targets.
Smart Images

Figure CN120125939A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of driverless technology, and particularly relates to a driverless multi-sensor fusion method and system based on an attention mechanism. Background Technique
[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] 3D object detection is a key technology in the perception system. It detects objects by using sensor data such as lidar and cameras, helping intelligent devices to perceive other vehicles, pedestrians, obstacles, or construction objects during the driverless process. Although the point cloud data collected by lidar is rich in spatial information, its data is sparse, the field of view is limited, and it lacks color and texture information. In contrast, camera images provide rich color and semantic information but are difficult to perform accurate depth perception. The deficiencies of these two types of data respectively restrict the performance of 3D object detection under the input of single-moment and single-modal data. Moreover, existing methods do not pay enough attention to spatially-positioned guided fusion in multi-sensor feature fusion, making it difficult to obtain stronger semantic features at the instance level and resulting in unbalanced feature fusion for distant objects, small targets, etc. Summary of the Invention
[0004] To solve the above problems, the present invention proposes a driverless multi-sensor fusion method and system based on an attention mechanism. The present invention improves the cross-attention mechanism with a fuzzy algorithm, combines the respective advantages of the fuzzy algorithm and the attention mechanism to form an improved fuzzy attention mechanism algorithm, and realizes more accurate and comprehensive feature fusion.
[0005] According to some embodiments, the present invention adopts the following technical solutions:
[0006] A driverless multi-sensor fusion method based on an attention mechanism, comprising the following steps:
[0007] Obtain point cloud data and image data, and respectively use corresponding feature extractors for the two to extract pixel features of the image data and point cloud features of the point cloud data;
[0008] For the extracted pixel features and point cloud features, use a fuzzy-improved cross-attention mechanism for feature fusion;
[0009] The process of feature fusion by the fuzzy improved cross-attention mechanism is as follows: construct a membership function according to the density of the point cloud, calculate the membership degree of the points corresponding to the pixel features to obtain the fuzzy distance weight, use each pixel feature in the image as a query, generate keys and values based on the point cloud features, perform dot product operations on the fuzzy distance weight with the query and the key, and map and transform them into fuzzy semantic attention weights. Apply the fuzzy semantic attention weights to the values, and obtain the final output vector through weighted summation to achieve the fuzzy weighted fusion of pixel features and point cloud features and generate the fused feature representation.
[0010] As an alternative implementation, the process of constructing the membership function according to the density of the point cloud is as follows: determine a search radius for each point in the point cloud, count the number of other points within the search radius to which the point belongs, and perform logarithmic processing on the counted number iteratively;
[0011] Use the data after logarithmic processing to determine the parameters of the skew normal distribution;
[0012] According to the obtained skew normal distribution membership function, calculate the membership degree of the points corresponding to the image pixel features to obtain the fuzzy distance weight.
[0013] As a further step, the formula for logarithmic transformation is where m is the number of points within the set search radius for each point, regarded as the local density of the point. Among them, the distance between the origin O and the other two points P 1 、P 2 is d 1 、d 2 , the angle between OP 1 and OP 2 for adjacent light beams is α, β th is the angle threshold parameter, R 1 is the distance coefficient, R 1 = max(d th / |d 1 -d 2 |, 0.5), d th is the distance threshold parameter.
[0014] As a further step, the skew normal distribution membership function is:
[0015]
[0016] where φ is the density function of the standard normal distribution, Φ is the distribution function of the standard normal distribution, and the parameters μ, σ, and α represent the location parameter, scale parameter, and skewness parameter respectively.
[0017] As an alternative implementation, use the softmax function to map and transform it into fuzzy semantic attention weights.
[0018] An unmanned driving method applies the above-mentioned unmanned multi-sensor fusion method based on the attention mechanism to identify the fusion feature representation of at least two features among obstacles, unmanned driving devices, driving / operation routes, environmental signs, and operation objects in the driving scenario.
[0019] According to the fusion feature representation, perform motion control of the unmanned driving device.
[0020] An unmanned multi-sensor fusion system based on the attention mechanism includes:
[0021] A feature extraction module is configured to obtain point cloud data and image data, and respectively use corresponding feature extractors for both to extract pixel features of the image data and point cloud features of the point cloud data.
[0022] A feature fusion module is configured to perform feature fusion on the extracted pixel features and point cloud features using a fuzzy improved cross-attention mechanism. The process of performing feature fusion by the fuzzy improved cross-attention mechanism is as follows: construct a membership function according to the density of the point cloud, calculate the membership degree of the points corresponding to the pixel features to obtain a fuzzy distance weight, use each pixel feature in the image as a query, and generate keys and values based on the point cloud features. Perform a dot product operation on the fuzzy distance weight, query, and key, and map and transform it into a fuzzy semantic attention weight. Apply the fuzzy semantic attention weight to the values, and obtain the final output vector through weighted summation to achieve fuzzy weighted fusion of the pixel features and point cloud features, and generate a fused feature representation.
[0023] An unmanned driving system applies the above-mentioned unmanned multi-sensor fusion system based on the attention mechanism to identify the fusion feature representation of at least two features among obstacles, unmanned driving devices, driving / operation routes, environmental signs, and operation objects in the driving scenario.
[0024] According to the fusion feature representation, perform motion control of the unmanned driving device.
[0025] According to the different driving scenarios, the identified object features are different. For example, in road driving, pedestrians, vehicles, and signal signs can be identified; in construction driving, construction machinery and construction targets can be identified; in agricultural operation scenarios, agricultural machinery, farmland, and crops can be identified, etc. These can all be adjusted according to the specific scenario.
[0026] An unmanned driving device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the above method are completed;
[0027] Or it includes the above system.
[0028] Driverless devices include but are not limited to driverless vehicles, drones, unmanned boats, etc.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] 1. The driverless multi-sensor fusion method and system based on the attention mechanism of the present invention improve the cross-attention mechanism with a fuzzy algorithm, combine the respective advantages of the fuzzy algorithm and the attention mechanism, form an improved fuzzy attention mechanism algorithm, and achieve more accurate and comprehensive feature fusion.
[0031] 2. The driverless multi-sensor fusion method and system based on the attention mechanism of the present invention use the improved fuzzy attention mechanism algorithm to perform multi-sensor feature fusion on images and point clouds, obtain a better fusion effect, and improve the efficiency and accuracy of feature fusion in target recognition.
[0032] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0034] Figure 1 It is a flowchart showing a driverless multi-sensor fusion method based on the attention mechanism provided by an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The present invention will be further described below in conjunction with the drawings and embodiments.
[0036] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0037] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0038] In the case of no conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0039] Example 1:
[0040] As Figure 1 shown, Example 1 of the present invention provides an unmanned multi-sensor fusion method based on an attention mechanism, including the following steps:
[0041] Step 1: Obtain image and point cloud data.
[0042] Among them, the image can be captured by devices such as cameras, and the point cloud can be collected by devices such as lidar; the recognition targets can be pedestrians, vehicles, and signal signs in the road; construction machinery and construction targets in the construction site; agricultural machinery, crops, and farmland during agricultural operations, etc.
[0043] Step 2: Extract point cloud features in the point cloud data and pixel features in the image data.
[0044] In this embodiment, an appropriate encoder, model, etc. can be used to extract the pixel features and point cloud features of the image. The encoder and model can use existing technologies and will not be elaborated here.
[0045] Step 3: Use a fuzzy improved cross-attention mechanism to perform feature fusion on the extracted pixel features and point cloud features;
[0046] The fuzzy improved cross-attention mechanism is as follows: First, a membership function is constructed according to the Gaussian function based on the density of the point cloud, and the membership degree of the points corresponding to the photo pixel features is calculated to obtain the fuzzy distance weight. A query Q is generated from the pixel features, and keys K and values V are generated from the point cloud feature set. The dot product of the fuzzy distance weight and the query Q and the key K is calculated, and the fuzzy semantic attention weight A is obtained through the softmax function. The fuzzy semantic attention weight is calculated through the fuzzy distance weight, and the fuzzy semantic attention weight is applied to the value V, and then the fused feature representation is obtained through weighted summation, so as to perform fuzzy weighted fusion on the pixel features and point cloud features and generate the fused feature representation.
[0047] The process of constructing the membership function is as follows: For each point in the point cloud, a search radius is determined. Within this radius, the number of other points within the distance from this point is counted, and the logarithm of the counted number is iteratively processed. The formula for logarithmic transformation is where m is the number of points within the specified distance for each point, regarded as the local density of this point, where the origin O and any other two points P 1 、P 2 The distance between them is d 1 、d 2 , OP 1 The angle between OP 2 and the adjacent light beams of OP th is α, β 1is the distance coefficient, R 1 = max(d th / |d 1 -d 2 |, 0.5), d th is the distance threshold parameter.
[0048] Use the data after logarithmic transformation to determine the parameters of the skew normal distribution. The skew normal distribution has three parameters: the location parameter μ, the scale parameter σ, and the skewness parameter α. These parameters are determined by analyzing the point cloud data through big data tools. The probability density function of the skew normal distribution usually has the following form:
[0049]
[0050] where φ is the density function of the standard normal distribution, Φ is the distribution function of the standard normal distribution, and the parameters α, σ, μ represent the location parameter, the scale parameter, and the skewness parameter respectively.
[0051] Based on the membership function of the skew normal distribution obtained from the above process, calculate the membership degree of the points corresponding to the pixel features of the image to obtain the fuzzy distance weight.
[0052] The target information collected in real time is introduced into the obtained multi-sensor fusion method for driverless based on the attention mechanism to identify pedestrians, vehicles, signal signs, etc. in the corresponding road; construction machinery and construction targets, etc. in the construction site; and the characteristics of agricultural machinery, crops, farmland, etc. in the agricultural operation process, to obtain the multi-sensor fusion method and system for driverless based on the attention mechanism.
[0053] Example 2:
[0054] Embodiment 2 of the present invention provides a multi-sensor fusion system for driverless based on the attention mechanism, including:
[0055] A feature extraction module, configured to obtain point cloud data and image data, and respectively use corresponding feature extractors for both to extract the pixel features of the image data and the point cloud features of the point cloud data;
[0056] A feature fusion module, configured to perform feature fusion on the extracted pixel features and point cloud features using a fuzzy improved cross-attention mechanism; the process of performing feature fusion by the fuzzy improved cross-attention mechanism is to construct a membership function according to the density of the point cloud, calculate the membership degree of the points corresponding to the pixel features, obtain the fuzzy distance weight, use each pixel feature in the image as a query, and generate keys and values based on the point cloud features, perform a dot product operation on the fuzzy distance weight with the query and the key, and map and transform it into a fuzzy semantic attention weight, apply the fuzzy semantic attention weight to the values, and obtain the final output vector through weighted summation to achieve fuzzy weighted fusion of the pixel features and point cloud features and generate a fused feature representation.
[0057] The working method of the system is the same as the attention mechanism-based multi-sensor fusion method for unmanned driving provided in Embodiment 1, and will not be elaborated here.
[0058] Embodiment 3:
[0059] Embodiment 3 of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the attention mechanism-based multi-sensor fusion method for unmanned driving as described in Embodiment 1 of the present invention.
[0060] Embodiment 4:
[0061] Embodiment 4 of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the attention mechanism-based multi-sensor fusion method for unmanned driving as described in Embodiment 1 of the present invention.
[0062] Embodiment 5:
[0063] An unmanned driving method applies the attention mechanism-based multi-sensor fusion method provided in Embodiment 1 to identify the fused feature representations of at least two features among obstacles, unmanned devices, driving / operation routes, environmental signs, and operation objects in a driving scenario;
[0064] According to the fused feature representation, perform motion control of the unmanned device.
[0065] According to different driving scenarios, the identified object features are different. For example, in road driving, pedestrians, vehicles, and signal signs can be identified; in construction driving, construction machinery and construction targets can be identified; in agricultural operation scenarios, agricultural machinery, farmland, and crops can be identified, etc. These can all be adjusted according to the specific scenario.
[0066] Of course, it is not limited to the above scenarios, and can also include scenarios such as the ocean and the sky.
[0067] Example 6:
[0068] An unmanned driving system applies the unmanned multi-sensor fusion system based on the attention mechanism in Example 2 to identify the fusion feature representation of at least two features among obstacles, unmanned driving devices, driving / operation routes, environmental signs, and operation objects in a driving scenario;
[0069] Based on the fusion feature representation, motion control of the unmanned driving device is performed.
[0070] Example 7:
[0071] An unmanned driving device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the unmanned multi-sensor fusion method based on the attention mechanism in Example 1 are completed;
[0072] Or it includes the unmanned multi-sensor fusion system based on the attention mechanism in Example 2.
[0073] The unmanned driving device includes but is not limited to unmanned vehicles, drones, unmanned ships, etc.
[0074] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0075] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0076] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0077] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more processes and / or blocks Figure 1 specified in the block or blocks.
[0078] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made by those skilled in the art without creative efforts within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An unmanned driving multi-sensor fusion method based on attention mechanism, characterized in that: The following steps are involved: Obtain point cloud data and image data, and use corresponding feature extractors to extract pixel features of the image data and point cloud features of the point cloud data respectively; The extracted pixel features and point cloud features are fused using the fuzzy improved cross attention mechanism; The process of feature fusion by the fuzzy improved cross attention mechanism is as follows: construct a membership function according to the density of the point cloud, calculate the membership of the point corresponding to the pixel feature, obtain the fuzzy distance weight, take each pixel feature in the image as a query, and generate a key and value based on the point cloud feature, perform a dot product operation on the fuzzy distance weight, the query and the key, and map them into a fuzzy semantic attention weight, apply the fuzzy semantic attention weight to the value, obtain the final output vector by weighted summation, realize the fuzzy weighted fusion of pixel features and point cloud features, and generate a fused feature representation.
2. The unmanned driving multi-sensor fusion method based on the attention mechanism as claimed in claim 1, characterized in that: The process of constructing the membership function according to the density of the point cloud is to determine a search radius for each point in the point cloud, count the number of other points within the search radius to which the point belongs, and iterate the logarithmic processing of the counted number; Determine the parameters of the skewed normal distribution using the logarithmically processed data; According to the obtained skew normal distribution membership function, the membership of the points corresponding to the image pixel features is calculated to obtain the fuzzy distance weight.
3. The unmanned driving multi-sensor fusion method based on the attention mechanism as claimed in claim 2, characterized in that: The formula for logarithmic transformation is Where m is the number of points within the set search radius for each point, which is regarded as the local density of the point. The distance between the origin O and the other two P1 and P2 is d1 and d2, and the angle between the adjacent beams of OP1 and OP2 is α and β. th is the angle threshold parameter, R1 is the distance coefficient, R1=max(d th / |d1-d2|,0.5), d th is the distance threshold parameter.
4. The unmanned driving multi-sensor fusion method based on the attention mechanism as claimed in claim 3, characterized in that: The skewed normal distribution membership function is: Among them, φ is the density function of the standard normal distribution, Φ is the distribution function of the standard normal distribution, and the parameters μ , σ , α represent the location parameter, scale parameter, and skewness parameter respectively.
5. The unmanned driving multi-sensor fusion method based on the attention mechanism as claimed in claim 1, characterized in that: Use the softmax function mapping to transform into fuzzy semantic attention weights.
6. An unmanned driving method, characterized in that: Applying the unmanned driving multi-sensor fusion method based on the attention mechanism described in any one of claims 1 to 5 to identify obstacles, unmanned driving equipment, driving / operating routes, environmental signs and fusion feature representations of at least two features of operating objects in driving scenes; The motion control of the unmanned driving device is performed according to the fused feature representation.
7. An unmanned driving multi-sensor fusion system based on attention mechanism, characterized by: include: A feature extraction module is configured to obtain point cloud data and image data, and use corresponding feature extractors to extract pixel features of the image data and point cloud features of the point cloud data respectively; The feature fusion module is configured to use a fuzzy improved cross attention mechanism to perform feature fusion on the extracted pixel features and point cloud features; the process of feature fusion using the fuzzy improved cross attention mechanism is to construct a membership function according to the density of the point cloud, calculate the membership of the point corresponding to the pixel feature, obtain the fuzzy distance weight, use each pixel feature in the image as a query, and generate a key and value based on the point cloud feature, perform a dot product operation on the fuzzy distance weight, the query and the key, and map them into a fuzzy semantic attention weight, apply the fuzzy semantic attention weight to the value, and obtain the final output vector by weighted summation, so as to realize the fuzzy weighted fusion of the pixel features and the point cloud features and generate a fused feature representation.
8. An unmanned driving system, characterized in that: The unmanned driving multi-sensor fusion system based on the attention mechanism as described in claim 7 can identify obstacles, unmanned driving equipment, driving / operating routes, environmental signs and fusion feature representations of at least two features of operating objects in driving scenes; The motion control of the unmanned driving device is performed according to the fused feature representation.
9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the steps in the method as claimed in claims 1 to 6 are implemented.
10. An unmanned driving device, characterized in that: The method comprises a memory and a processor and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the steps of the method according to any one of claims 1 to 6 are completed; Or comprising the system of claim 7 or 8.