A multi-source sensor fusion recognition method based on networked three-dimensional reconstruction

CN118445757BActive Publication Date: 2026-09-25TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410623510.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2026-09-25
Estimated Expiration
2044-05-20

AI Technical Summary

Technical Problem

这将显著限制现有智能汽车对障碍物的理解,导致识别精度上限受限

Benefits of technology

[0046]本发明首先对来自于一辆智能汽车的单一来源信息进行处理以获取障碍物信息,然后利用多辆智能汽车的单一来源信息构成多源信息,以重构对障碍物的三维视角的观测,得到信息融合图,并利用结合了图卷积网络和卷积神经网络的模型对信息融合图进行融合识别,以实现来自不同智能汽车的传感器信息巧妙的融合,从多维角度观测同一物体,显著提高智能汽车的识别精度上限,提高智能汽车在复杂环境中的适应性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118445757B_ABST
    Figure CN118445757B_ABST
Patent Text Reader

Abstract

The application relates to a multi-source sensor fusion recognition method based on networked three-dimensional reconstruction, which comprises the following steps: acquiring single-source information collected by a sensor of an intelligent automobile and performing pretreatment; constructing an information fusion graph according to multi-source information, wherein the multi-source information is composed of the pretreated single-source information of multiple intelligent automobiles around the same obstacle; taking the information fusion graph as the input of an information fusion model, performing multi-source information fusion and recognition by using the information fusion model, and outputting a recognition result. Compared with the prior art, the application can significantly improve the upper limit of the recognition precision of the intelligent automobile and improve the adaptability of the intelligent automobile in a complex environment by skillfully fusing the sensor information from different intelligent automobiles and observing the same object from a multi-dimensional angle based on the existing Internet of Things technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent driving vehicles, and in particular to a multi-source sensor fusion recognition method based on networked 3D reconstruction. Background Technology

[0002] Intelligent vehicles have garnered attention for their ability to reduce the frequency of accidents. However, because intelligent vehicles are required to operate in open and complex driving environments, improving their adaptability in high-latitude and numerous traffic scenarios has become a challenge. Furthermore, the rapid development of the Internet of Things (IoT) has enabled high-speed communication between intelligent vehicles, providing a novel technological means to observe the same object from multiple spatial perspectives and improve object recognition accuracy.

[0003] Existing perception algorithms are mainly divided into two categories: single-sensor recognition and multi-source sensor recognition methods. Due to the differentiated characteristics of different sensor types, multi-source sensor recognition methods are significantly more adaptable than single-sensor recognition methods and have become the mainstream intelligent vehicle recognition approach. Although multi-source sensors have achieved good results, they utilize the output information from sensors deployed within the same intelligent vehicle. Therefore, even though the information comes from multiple sensors, the position of the intelligent vehicle relative to the observed object is fixed—that is, the angle of the observed object is within a fixed and narrow range. This significantly limits the current intelligent vehicle's understanding of obstacles, resulting in a limited upper limit to recognition accuracy. Summary of the Invention

[0004] The purpose of this invention is to provide a multi-source sensor fusion recognition method based on networked 3D reconstruction. Here, "multi-source" refers to data from different intelligent vehicles. By observing the same object from multiple angles, the upper limit of the recognition accuracy of intelligent vehicles can be significantly improved, thereby enhancing the adaptability of intelligent vehicles in complex environments.

[0005] The objective of this invention can be achieved through the following technical solutions:

[0006] A multi-source sensor fusion identification method based on networked 3D reconstruction includes the following steps:

[0007] S1, acquire single-source information collected by sensors of a smart car and preprocess it;

[0008] S2, construct an information fusion graph based on multi-source information, wherein the multi-source information consists of pre-processed single-source information from multiple intelligent vehicles around the same obstacle;

[0009] S3 uses the information fusion graph as input to the information fusion model, performs multi-source information fusion and recognition using the information fusion model, and outputs the recognition results.

[0010] The sensors include millimeter-wave radar, cameras, and a global positioning system. The single-source information collected includes the relative distance between the current intelligent vehicle and obstacles collected by millimeter-wave radar, image information collected by cameras, and the world coordinates of the intelligent vehicle collected by the global positioning system.

[0011] The preprocessing includes location information processing and image information compression processing based on graph convolutional neural networks.

[0012] The location information processing specifically involves:

[0013] The relative distance (x) between the intelligent vehicle and obstacles is obtained using millimeter-wave radar. r ,y r The world coordinates (x, y) of the intelligent vehicle are obtained through the Global Positioning System. v ,y v The relative distance is added to the world coordinates of the intelligent vehicle to obtain the position (x, y) of the obstacle in the world coordinate system. d ,y d The calculation formula is as follows:

[0014] x d =x r +x v ;y d =y r +y v

[0015] The image information compression processing based on graph convolutional neural networks includes the following steps:

[0016] S121, Image grayscale processing

[0017] The color image information captured by the camera is converted to grayscale by averaging the pixel values ​​of the three RGB channels of each pixel in the color image, as shown in the following formula:

[0018]

[0019] In the formula u G For each pixel, R, G, and B represent the pixel values ​​of the red, green, and blue channels, respectively.

[0020] S122, Construct the basic graph model

[0021] The basic graph model consists of nodes U s and the edge E connecting the nodes s Composition, i.e. G s =(U s E s ), where node U s The grayscale value u of a pixel GThe edges connecting nodes are represented by Euclidean metrics in the pixel image coordinate system, and the specific calculation formula is as follows:

[0022]

[0023] In the formula, e ij This represents the edge connecting the i-th pixel and the j-th pixel in the pixel coordinate system; c ix and c iy c represents the coordinates of the i-th pixel in the pixel coordinate system; jx and c jy This represents the coordinates of the j-th pixel in the pixel coordinate system; e sij Let E be the edge matrix s The elements, where E s It is an N×N square matrix that represents the connection relationship between two pixels, where N represents the number of pixels on an edge of a grayscale image;

[0024] S123, Information Compression Based on Graph Convolutional Neural Networks

[0025] Graph convolutional networks are used to predict the basic graph model, using A=E s +I calculates the adjacency matrix A of the basic graph model, where I is the identity matrix;

[0026] via d mm =∑A mn Calculate the element d of the metric matrix mm , forming a metric matrix D, where D is a diagonal matrix;

[0027] Iteration is performed using the following formula:

[0028]

[0029] In the formula, W represents the network parameters to be learned, and l G Z represents the number of graph convolution iterations, and Z is the iteration element, i.e., Zi. 0 =U s ;

[0030] Through each iteration of the graph convolutional network, nodes perceive information about their neighbors, ultimately deriving this information from the hidden state H of the graph convolutional neural network. s Perceive all information in the image and compress it.

[0031] The preprocessed single-source information includes the hidden state H of the graph convolutional neural network. s Obstacle location information (x d ,y d ) and the coordinate information of the intelligent vehicle relative to the obstacle (x r ,y r(), representing multi-source information as X i =[H si ,x ri ,y ri x d ,y d ], where i is the intelligent car number around the same obstacle.

[0032] The construction of the information fusion graph based on multi-source information specifically involves:

[0033] Define the information fusion graph as G M =(U M E M ), where the node information matrix U M It consists of the hidden states of a graph convolutional neural network based on information from a single source, i.e., U M element u in Mi =H si Edge matrix E M E is a Q×Q square matrix used to express the difference in viewing angles of the same obstacle, where Q represents the number of intelligent vehicles observing the same obstacle. The connection relationship between nodes is determined by the angle between the objects observed by two intelligent vehicles, forming the edge matrix E. M element e in Mij The calculation formula is as follows:

[0034]

[0035] In the formula and Let i and j be the distance vectors of the two smart cars i and j relative to the obstacle.

[0036] The number of intelligent vehicles observing the same obstacle, Q, is a fixed value determined according to the computing power of the computing device. When the actual number of intelligent vehicles around the same obstacle is greater than Q, they are eliminated by ranking the distance vectors of the two intelligent vehicles relative to the obstacle from closest to farthest. When the actual number of intelligent vehicles around the same obstacle is less than Q, they are completed by using unit vectors.

[0037] The information fusion model includes an information fusion module based on graph convolutional neural networks and a recognition module based on convolutional neural networks, wherein,

[0038] The information fusion module uses A′=E M +I calculates the adjacency matrix A′ of the information fusion graph, where I is the identity matrix; through d′ mm =∑A′ mn Calculate the element d′ of the metric matrix mm This forms the metric matrix D′, which is a diagonal matrix; and it is iterated using the following formula:

[0039]

[0040] In the formula, W′ represents the network parameters to be learned, l M Z' represents the number of graph convolution iterations in the information fusion module, where Z′ is the iteration element, i.e., Z 0 =U M After iteration, the information fusion module outputs Z. M ;

[0041] The recognition module uses the output Z of the graph convolutional neural network of the information fusion module. M As input, the data is processed by a convolutional neural network with a softmax layer, using the following convolution formula, and the output is the recognition result through the softmax layer:

[0042] H = Conv(Z) M )

[0043] In the formula, Conv(·) is the convolution calculation formula, and H is the output of the convolutional neural network.

[0044] The loss function of the information fusion model is determined based on the difference between the labeling of the highest probability object output by the softmax layer and the labeling of the real object.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] This invention first processes single-source information from a single intelligent vehicle to obtain obstacle information. Then, it uses single-source information from multiple intelligent vehicles to construct multi-source information, reconstructing a three-dimensional perspective observation of the obstacle to obtain an information fusion map. Finally, it uses a model combining graph convolutional networks and convolutional neural networks to perform fusion recognition on the information fusion map, thereby achieving a clever fusion of sensor information from different intelligent vehicles. This allows for observation of the same object from multiple dimensions, significantly improving the upper limit of the recognition accuracy of intelligent vehicles and enhancing their adaptability in complex environments. Attached Figure Description

[0047] Figure 1 This is a flowchart of the method of the present invention;

[0048] Figure 2 This is a schematic diagram of the information perception process of a convolutional neural network.

[0049] Figure 3 This is a schematic diagram illustrating the differences in perspective when observing obstacles from a three-dimensional viewpoint. Detailed Implementation

[0050] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0051] To improve the adaptability of intelligent vehicles in open environments, this invention proposes a multi-source sensor fusion recognition method based on networked 3D reconstruction. Utilizing existing IoT technology, it cleverly fuses sensor information from different intelligent vehicles, observing the same object from multiple perspectives. This significantly improves the upper limit of recognition accuracy for intelligent vehicles and enhances their adaptability in complex environments. Specifically, for example... Figure 1 As shown, it includes the following steps:

[0052] S1: Acquire single-source information collected by sensors in a smart car and preprocess it.

[0053] This step will process information from the sensors of a smart car. It's important to note that "single source" here refers to information from the entire smart car, not just a single sensor. In this embodiment, the information from the smart car includes data from millimeter-wave radar, cameras, and GPS outputs.

[0054] Preprocessing includes the following steps:

[0055] S11, Location Information Processing

[0056] The relative distance (x) between the intelligent vehicle and obstacles is obtained using millimeter-wave radar. r ,y r The world coordinates (x, y) of the intelligent vehicle are obtained through the Global Positioning System. v ,y v The relative distance is added to the world coordinates of the intelligent vehicle to obtain the position (x, y) of the obstacle in the world coordinate system. d ,y d The calculation formula is as follows:

[0057] x d =x r +x v ;y d =y r +y v

[0058] S12, Image information compression processing based on graph convolutional neural networks

[0059] S121, Image grayscale processing

[0060] To convert the color image information captured by the camera to grayscale, this embodiment uses the averaging method, which involves taking the average of the pixel values ​​of the RGB (red, green, and blue) channels of each pixel in the color image to obtain the corresponding grayscale value. The formula is as follows:

[0061]

[0062] In the formula u G R, G, and B represent the grayscale value of each pixel, respectively, representing the pixel values ​​of the red, green, and blue channels.

[0063] S122, Construct the basic graph model

[0064] The basic graph model consists of nodes U s and the edge E connecting the nodes s Composition, i.e. G s =(U s E s ), where node U s The grayscale value u of a pixel G The edges connecting nodes are represented by Euclidean metrics in the pixel image coordinate system, and the specific calculation formula is as follows:

[0065]

[0066] In the formula, e ij This represents the edge connecting the i-th pixel and the j-th pixel in the pixel coordinate system; c ix and c iy c represents the coordinates of the i-th pixel in the pixel coordinate system; jx and c jy This represents the coordinates of the j-th pixel in the pixel coordinate system; e sij Let E be the edge matrix s The elements, where E s It is an N×N square matrix that represents the connection relationship between two pixels, where N represents the number of pixels on an edge of a grayscale image.

[0067] S123, Information Compression Based on Graph Convolutional Neural Networks

[0068] To reduce the burden of information transmission, graph convolutional networks are used to predict the basic graph model.

[0069] via A=E s +I calculates the adjacency matrix A of the basic graph model, where I is the identity matrix;

[0070] via d mm =∑A mn Calculate the metric matrix element d mm , forming a metric matrix D, where D is a diagonal matrix;

[0071] Iteration is performed using the following formula:

[0072]

[0073] In the formula, W represents the network parameters to be learned, and l G Z represents the number of graph convolution iterations, and Z is the iteration element, i.e., Zi. 0 =U s .

[0074] As can be seen from the above calculation process, in each iteration, the node will perceive neighbor information, such as... Figure 2 As shown, the hidden state H of the graph convolutional neural network is finally obtained. s Perceive all information in the graph and perform information compression. This is to ensure the final hidden state H... s It can effectively perceive information from the entire image, where l G It can be defined as N-1. The number of iterations is related to the computing power of the computing platform of the ride algorithm; it can be reduced, but the corresponding accuracy will be reduced.

[0075] S2, construct an information fusion graph based on multi-source information, wherein the multi-source information consists of pre-processed single-source information from multiple intelligent vehicles around the same obstacle.

[0076] As can be seen from the above single-source information preprocessing process, the preprocessed single-source information includes the hidden state H of the graph convolutional neural network. s Obstacle location information (x d ,y d ) and the coordinate information of the intelligent vehicle relative to the obstacle (x r ,y r This embodiment represents multi-source information as X. i =[H si ,x ri ,y ri x d ,y d [i] represents the intelligent vehicle ID surrounding the same obstacle. Information is filtered only by the world coordinates of the obstacle to ensure that all information describes the same obstacle.

[0077] Multi-source information fusion graphs are used to fuse camera information from different smart cars to reconstruct a 3D perspective observation of obstacles. The information fusion graph is defined as G. M =(U M E M ), where the node information matrix U M It consists of the hidden states of a graph convolutional neural network based on information from a single source, i.e., U M element u inMi =H si .

[0078] To demonstrate that the information output by a single intelligent vehicle represents observations of a 3D image of an obstacle from different angles, the connection relationships between nodes are determined here by the angle between the observations of the object by two intelligent vehicles. Specifically, as... Figure 3 As shown, when the angle between the distance vectors of two intelligent vehicles relative to an obstacle is zero, the information increment is 0, meaning the object is observed from the same 3D perspective. Conversely, when the angle between the distance vectors of two intelligent vehicles relative to an obstacle is 180°, the information increment is 1, meaning the object is observed from two completely different 3D perspectives. That is, the edge matrix E... M E is a Q×Q square matrix used to express the difference in viewing angles of the same obstacle, where Q represents the number of intelligent vehicles observing the same obstacle. The connection relationship between nodes is determined by the angle between the objects observed by two intelligent vehicles, forming the edge matrix E. M element e in Mij The calculation formula is as follows:

[0079]

[0080] In the formula and Let i and j be the distance vectors of the two smart cars i and j relative to the obstacle.

[0081] The number of intelligent vehicles observing the same obstacle, Q, is a fixed value determined based on the computing power of the computing device. When the actual number of intelligent vehicles around the same obstacle is greater than Q, they are eliminated by ranking the distance vectors of the two intelligent vehicles relative to the obstacle from closest to farthest. When the actual number of intelligent vehicles around the same obstacle is less than Q, they are completed by using unit vectors.

[0082] S3 uses the information fusion graph as input to the information fusion model, performs multi-source information fusion and recognition using the information fusion model, and outputs the recognition results.

[0083] In this embodiment, the information fusion model is used to fuse information from multiple sources of sensors and to achieve the recognition task. It includes an information fusion module based on graph convolutional neural networks and a recognition module based on convolutional neural networks.

[0084] The information fusion module is similar to the information compression based on graph convolutional neural networks in step S123, specifically: through A′=E M +I calculates the adjacency matrix A′ of the information fusion graph, where I is the identity matrix; through d′ mm =∑A′ mn Calculate the element d′ of the metric matrix mmThis forms the metric matrix D′, which is a diagonal matrix; and it is iterated using the following formula:

[0085]

[0086] In the formula, W′ represents the network parameters to be learned, l M Z' represents the number of graph convolution iterations in the information fusion module, where Z′ is the iteration element, i.e., Z 0 =U M After iteration, the information fusion module outputs Z. M However, it's important to note that, in order to comprehensively understand the obstacles, the number of iterations for graph convolution is l. M It is defined as Q-1. The number of iterations is related to the computing power of the computing platform of the ride algorithm, and can be reduced, but the corresponding accuracy will be reduced.

[0087] The recognition module uses the output Z of the graph convolutional neural network of the information fusion module. M As input, the data is processed by a convolutional neural network with a softmax layer, using the following convolution formula, and the output is the recognition result through the softmax layer:

[0088] H = Conv(Z) M )

[0089] In the formula, Conv(·) is the convolution calculation formula, and H is the output of the convolutional neural network.

[0090] The ultimate goal of the information fusion model is to identify the category of objects. Therefore, the loss function is expressed based on the difference between the label of the object with the highest probability output by the softmax layer and the label of the real object. For example, if the object name is correct, the loss is recorded as 0, and if it is incorrect, it is recorded as 1.

[0091] The method of this invention has four steps when deployed in a real vehicle, as detailed below:

[0092] 1. Determine the model of the camera on the deployed vehicle and the dimension N of the input features in the network.

[0093] 2. Prepare a dataset from extreme environments as the training dataset for the model, which contains accurate results with human annotations.

[0094] 3. Based on the existing computing platform, set the iteration parameters l of the network. G and l M That is, the number of iterations of the information compression graph convolutional network and the information fusion graph convolutional network, and then the model is trained until the loss converges.

[0095] 4. Download the converged model to the intelligent vehicle recognition and control system to complete the model deployment.

[0096] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A multi-source sensor fusion identification method based on networked 3D reconstruction, characterized in that, Includes the following steps: S1, acquire single-source information collected by sensors of a smart car and preprocess it; the sensors include millimeter-wave radar, camera and global positioning system, the single-source information collected includes the relative distance between the current smart car and obstacles collected by millimeter-wave radar, image information collected by camera, and world coordinates of the smart car collected by global positioning system; the preprocessing includes location information processing and image information compression processing based on graph convolutional neural network. S2, construct an information fusion graph based on multi-source information, wherein the multi-source information consists of preprocessed single-source information from multiple intelligent vehicles around the same obstacle; the preprocessed single-source information includes the hidden states of the graph convolutional neural network. Obstacle location information And the coordinate information of the intelligent vehicle relative to the obstacle. Representing multi-source information as , i Number the intelligent vehicles around the same obstacle; The construction of the information fusion graph based on multi-source information specifically involves: Define the information fusion graph as The node information matrix It consists of the hidden states of a graph convolutional neural network derived from information from a single source, i.e. elements in edge matrix It is A square matrix is ​​used to express the degree of difference in the viewing angles of observing the same obstacle, where... This represents the number of intelligent cars observing the same obstacle. The connection relationship between nodes is determined by the angle between the objects observed by two intelligent cars, forming an edge matrix. elements in The calculation formula is as follows: , In the formula and Two smart cars i and j Distance vector relative to the obstacle; S3 uses the information fusion graph as input to the information fusion model, performs multi-source information fusion and recognition using the information fusion model, and outputs the recognition result; The information fusion model includes an information fusion module based on graph convolutional neural networks and a recognition module based on convolutional neural networks, wherein, Information fusion module through Calculate the adjacency matrix of the information fusion graph ,in, It is the identity matrix; through Calculate the elements of the metric matrix ,in Adjacency matrix of the basic graph model The Middle Line number Column elements, ; Construct a metric matrix , It is a diagonal matrix; and it is iterated using the following formula: In the formula The network parameters to be learned, This indicates the number of graph convolution iterations in the information fusion module. For the iterable element, i.e. After iteration, the information fusion module outputs... ; The recognition module uses the output of the graph convolutional neural network of the information fusion module. As input, the data is processed by a convolutional neural network with a softmax layer, using the following convolution formula, and the output is the recognition result through the softmax layer: In the formula Here is the formula for calculating convolution. This is the output of the convolutional neural network.

2. The multi-source sensor fusion identification method based on networked 3D reconstruction according to claim 1, characterized in that, The location information processing specifically involves: The relative distance between the intelligent vehicle and obstacles is obtained using millimeter-wave radar. Obtain the world coordinates of intelligent vehicles through the Global Positioning System The relative distance is added to the world coordinates of the intelligent vehicle to obtain the position of the obstacle in the world coordinate system. The calculation formula is as follows: 。 3. The multi-source sensor fusion identification method based on networked 3D reconstruction according to claim 1, characterized in that, The image information compression processing based on graph convolutional neural networks includes the following steps: S121, Image grayscale processing The color image information captured by the camera is converted to grayscale by averaging the pixel values ​​of the three RGB channels of each pixel in the color image, as shown in the following formula: In the formula For each pixel, R, G, and B represent the pixel values ​​of the red, green, and blue channels, respectively. S122, Construct the basic graph model The basic graph model consists of nodes and the edges connecting nodes Composition, that is , among which, nodes grayscale value of a pixel The edges connecting nodes are represented by Euclidean metrics in the pixel image coordinate system, and the specific calculation formula is as follows: In the formula, In the pixel coordinate system, the first... i The pixel and the j Connecting edges of 1 pixel; and Indicates the first i The coordinates of each pixel in the pixel coordinate system; and Indicates the first j The coordinates of each pixel in the pixel coordinate system; edge matrix The elements, among which It is A square matrix represents the connection relationship between two pixels. N This indicates the number of pixels on each edge of a grayscale image. S123, Information Compression Based on Graph Convolutional Neural Networks Graph convolutional networks are used to predict the basic graph model, through Calculate the adjacency matrix of the basic graph model A ,in, It is the identity matrix; pass Calculate the elements of the metric matrix ,in Adjacency matrix of the basic graph model A The Middle Line number Column elements, ; Construct a metric matrix , D It is a diagonal matrix; Iteration is performed using the following formula: In the formula The network parameters to be learned, This indicates the number of graph convolution iterations. For the iterable element, i.e. ; Through each iteration of the graph convolutional network, nodes perceive information about their neighbors, ultimately deriving this information from the hidden states of the graph convolutional neural network. Perceive all information in the image and compress it.

4. The multi-source sensor fusion identification method based on networked 3D reconstruction according to claim 1, characterized in that, The number of intelligent cars observing the same obstacle It is a fixed value, determined based on the computing power of the computing device; when the actual number of intelligent vehicles around the same obstacle is greater than... At that time, the number of intelligent vehicles around the obstacle is eliminated by ranking them from closest to farthest using the distance vectors between the two intelligent vehicles. When the actual number of intelligent vehicles around the same obstacle is less than [number missing], [the elimination process is repeated]. When the time is right, it is completed using a unit vector.

5. The multi-source sensor fusion identification method based on networked 3D reconstruction according to claim 1, characterized in that, The loss function of the information fusion model is determined based on the difference between the labeling of the highest probability object output by the softmax layer and the labeling of the real object.