A distributed multi-level perception fusion method, device and equipment based on vehicle-road cooperation and a storage medium
By using a perception fusion method combining convolutional neural networks and vehicle-road cooperative systems, local perception results from vehicle-side and roadside units are fused. This solves the problems of massive roadside perception information and heavy transmission and computational burden in vehicle-road cooperative systems, achieving efficient perception result fusion and field of view gain, and optimizing the information processing efficiency of autonomous driving.
Patent Information
- Application Number
- CN202210884720.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-26
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-07-26
AI Technical Summary
In existing vehicle-road cooperative systems, the amount of roadside perception information is enormous, resulting in a heavy burden on network transmission and computing. Vehicles struggle to efficiently acquire the perception gains of surrounding roadside units, and the fusion method fails to effectively consider the vehicle's driving direction, leading to insufficient auxiliary information.
A perception fusion method based on convolutional neural networks and vehicle-road cooperation is adopted. The vehicle-side and roadside units fuse local perception results and adjust the confidence level through convolutional neural networks. Combined with vehicle-road cooperative computing, the fusion of distributed multi-level perception results is realized, which reduces the amount of 5G data transmission and computation on the vehicle side, while improving object recognition capability and field of view.
Perception fusion is achieved through distributed computing on the roadside, allowing the vehicle to obtain a field of vision gain that meets the needs of autonomous driving with less network transmission and computing power. This improves object recognition capabilities, reduces blind spots, and optimizes the information transmission and processing efficiency of the vehicle-road cooperative system.
Smart Images

Figure CN115438711B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of vehicle-road cooperation, and in particular to a distributed multi-level perception fusion method and device based on vehicle-road cooperation, equipment and a storage medium. BACKGROUND
[0002] There are four possible fusion modes in the current vehicle-road cooperation: 1. The vehicle end receives the perception results of each road side unit of the road end and performs fusion operation; 2. A central node is added to the road side, and the results of the entire road side perception system are fused and sent to all vehicle ends in the vehicle-road cooperation system; 3. A central node is added to the road side to calculate the distance between the vehicle and the road side unit in real time, and the perception results of the nearest road side unit are sent to the vehicle end; 4. A central node is added to the road side to calculate the distance between the vehicle and the road side unit in real time, and the data of the road side units within a certain range near the vehicle are sent to the vehicle end.
[0003] Assuming that there are n road side units and k vehicles in the vehicle-road cooperation system, and each road side unit has at most four other road side units around it. The disadvantage of fusion mode 1 is that the vehicle end needs to receive all the road end perception results, and each vehicle needs to receive them. The entire vehicle-road cooperation system needs to transmit the road end perception results n*k times per unit time through 5G, and perform k times of fusion operation. Compared with fusion mode 1, fusion mode 2 completes the fusion operation at the road side, transmits n times on the road side optical fiber broadband, transmits k times on 5G, performs 1 time of fusion operation at the road side, and performs 1 time of fusion operation at the vehicle end. The advantages of fusion modes 1 and 2 are that they obtain the view of all RSUs on the road, and the help to the vehicle end is the largest. Fusion mode 3 abandons the idea of sending all road side perception results to the vehicle end, and selects to send the perception results of the nearest road side unit to the vehicle end. It transmits k times on 5G, performs 1 time of fusion operation at the vehicle end, and performs distance calculation at the road end to determine which road side unit data the vehicle receives. Although the transmission times are the same as those of fusion mode 2, the single frame data amount transmitted on 5G in fusion mode 3 is much smaller than that in fusion mode 2. The disadvantage of fusion mode 3 is that the fusion content is less, and the gain is smaller than that of fusion modes 1 and 2. Fusion mode 4 seeks a compromise, and only sends the road end data within a certain range near the vehicle to the vehicle end. Assuming that the number of road side units within 100 meters near the vehicle is at most m (m is much smaller than n), the vehicle-road cooperation system needs to transmit at most m*k times of perception data on 5G. In addition, fusion modes 3 and 4 match the vehicle and the road side unit based on the distance between the vehicle and the road side unit, and do not consider the driving direction of the vehicle. It is possible that the nearest road side unit matched to the vehicle is behind the vehicle, and in this case, the auxiliary information provided by the RSU is greatly discounted for driving.
[0004] It can be seen that the road end perception information is huge. How to use less network transmission, road end and vehicle end calculation to make the driving vehicle obtain more gain from the road end information is a main problem to be solved at present. SUMMARY
[0005] The technical problem solved by the present application is to overcome the shortcomings of the prior art and provide a distributed multi-level perception fusion method, device and equipment based on vehicle-road cooperation and a storage medium.
[0006] To solve the above technical problems, the technical solutions of the present application are as follows:
[0007] A distributed multi-level perception fusion method based on vehicle-road cooperation, comprising,
[0008] The vehicle end and the road side unit fuse the perception results of the laser radar and the camera themselves;
[0009] The perception fusion method based on convolutional neural network and vehicle-road cooperation is used to obtain the road side unit fusion perception results obtained by each road side unit taking itself as the main observation and a plurality of road side units within a certain range around itself as the auxiliary observation, and the vehicle end fusion perception results obtained by each vehicle end taking itself as the main observation and a plurality of vehicle ends within a certain range around itself as the auxiliary observation;
[0010] The perception fusion method based on convolutional neural network and vehicle-road cooperation is used to obtain the fusion perception results obtained by the vehicle end taking itself as the main observation and the nearest road side unit in front of the vehicle end as the auxiliary observation.
[0011] As a preferred scheme of the distributed multi-level perception fusion method based on vehicle-road cooperation, wherein the fusion of the perception results of the laser radar and the camera of the vehicle end and the road side unit comprises,
[0012] The camera information in the vehicle end and the road side unit is respectively used to assist in adjusting the confidence of the laser radar detection model output in the vehicle end and the road side unit by the CLOCs method.
[0013] As a preferred scheme of the distributed multi-level perception fusion method based on vehicle-road cooperation, wherein the perception fusion method based on convolutional neural network and vehicle-road cooperation comprises,
[0014] Obtain the detection frame of the main observation and the auxiliary observation;
[0015] Calculate the first tensor of the main observation and the auxiliary observation detection frame, and create an empty sparse matrix of the first tensor;
[0016] The first tensor extracts the combination of the primary observation detection frame and the auxiliary observation detection frame based on the primary observation and the auxiliary observation detection frame, records the index of each extracted detection frame combination in the first tensor, and groups the extracted detection frame combinations into a second tensor as an input tensor of the convolutional neural network;
[0017] The second tensor is convolved using a 1*1 convolution to obtain one-dimensional features of each detection frame combination in the second tensor, which is used as the confidence of the corresponding primary observation after adjustment by the auxiliary observation;
[0018] According to the index of each detection frame combination in the first tensor, the confidence of the detection frame combination is placed in the empty sparse matrix of the first tensor;
[0019] The maximum pooling is used to select the maximum confidence from several confidences of each primary observation detection frame after adjustment by the auxiliary observation as the confidence of the primary observation.
[0020] As a preferred scheme of the distributed multi-level perception fusion method based on vehicle-road cooperation, wherein the first tensor of the primary observation and the auxiliary observation detection frame comprises,
[0021] The first tensor T i,j ={t i,j , IoU i,j , s i , s j , d i , d j},
[0022] Wherein, i represents the i-th object in the primary observation, and j represents the j-th object in the auxiliary observation;
[0023] t i,j represents the time difference normalization weight of the i-th object in the primary observation and the j-th object in the auxiliary observation, and the calculation method of t i,j , wherein, represents the time difference, represents the maximum delay, is a coefficient inversely proportional to the curvature of the curve;
[0024] IoU i,j represents the intersection over union of the i-th object in the primary observation detection frame and the j-th object detection frame in the auxiliary observation detection frame, and the calculation method of IoU i,j The calculation method is that: first, the intersection area S of the main observation detection frame and the auxiliary observation detection frame projected on the x-y plane is calculated, then the intersection length L of the main observation detection frame and the auxiliary observation detection frame projected on the z axis is calculated, and the intersection volume V1 is obtained by multiplying the intersection area S and the intersection length L, then the union volume V2 is obtained by subtracting the intersection volume V1 from the sum of the volume of the main observation detection frame and the volume of the auxiliary observation detection frame, and finally, the IoU is obtained by dividing the intersection volume V1 by the union volume V2;
[0025] s represents the confidence of the model output about the main observation detection frame;
[0026] d represents the normalized distance of the detected object from the observation center.
[0027] As a preferred scheme of the distributed multi-level perception fusion method based on vehicle-road cooperation, wherein: the 1*1 convolution is used to convolve the second tensor to obtain one-dimensional features of each group of detection frame combinations in the second tensor, which is used as the confidence of the corresponding main observation after adjustment by the auxiliary observation,
[0028] The 1*1 convolution is used to linearly transform the feature space constructed by the second tensor to another feature space, and the dimension of the feature is increased to eighteen, and the RELU activation function is used to increase the non-linear excitation after the dimension increase;
[0029] The 1*1 convolution is used to linearly transform the feature space obtained in the previous step to another feature space, and the dimension of the feature is increased to thirty-six, and the RELU activation function is used to increase the non-linear excitation after the dimension increase;
[0030] The 1*1 convolution is used to linearly transform the feature space obtained in the previous step to another feature space, and the RELU activation function is used to increase the non-linear excitation of the new feature space result;
[0031] The 1*1 convolution is used to linearly transform the feature space obtained in the previous step to another feature space, and the dimension of the feature is reduced to one dimension, and one-dimensional features of each group of detection frame combinations in the second tensor are obtained, which are used as the confidence of the main observation after adjustment by the auxiliary observation.
[0032] As a preferred scheme of the distributed multi-level perception fusion method based on vehicle-road cooperation, wherein: the perception fusion method based on convolutional neural network and vehicle-road cooperation is adopted to obtain the road side unit fusion perception result obtained by taking each road side unit as the main observation and a plurality of road side units within a certain range around the road side unit as the auxiliary observation, and the vehicle end fusion perception result obtained by taking each vehicle end as the main observation and a plurality of vehicle ends within a certain range around the vehicle end as the auxiliary observation,
[0033] The perception fusion method based on the convolutional neural network and the vehicle-road cooperation is used to obtain a plurality of fusion perception results of each roadside unit, the roadside unit taking itself as the main observation and a plurality of roadside units within a certain range around itself as the auxiliary observation, and a fusion perception result of each vehicle end, the vehicle end taking itself as the main observation and a plurality of vehicle ends within a certain range around itself as the auxiliary observation.
[0034] For each roadside unit, the plurality of fusion perception results are integrated by taking the average value or the maximum value to obtain the integrated fusion perception result of each roadside unit, and for each vehicle end, the plurality of fusion perception results are integrated by taking the average value or the maximum value to obtain the integrated fusion perception result of each vehicle end.
[0035] For each roadside unit, the integrated fusion perception results of the plurality of roadside units within a certain range around itself are added and combined, and then the suboptimal solution is filtered through the NMS algorithm to obtain the expanded fusion perception result corresponding to each roadside unit, which is used as the roadside unit fusion perception result of each roadside unit, and for each vehicle end, the integrated fusion perception results of the plurality of vehicle ends within a certain range around itself are added and combined, and then the suboptimal solution is filtered through the NMS algorithm to obtain the expanded fusion perception result corresponding to each vehicle end, which is used as the vehicle end fusion perception result of each vehicle end.
[0036] As a preferred scheme of the distributed multi-level perception fusion method based on the vehicle-road cooperation, wherein the roadside unit fusion perception result obtained by each roadside unit taking itself as the main observation and a plurality of roadside units within a certain range around itself as the auxiliary observation and the vehicle end fusion perception result obtained by each vehicle end taking itself as the main observation and a plurality of vehicle ends within a certain range around itself as the auxiliary observation include,
[0037] For each roadside unit, the perception results of the plurality of roadside units within a certain range around itself are combined to form the overall auxiliary observation corresponding to itself, and for each vehicle end, the perception results of the plurality of vehicle ends within a certain range around itself are combined to form the overall auxiliary observation corresponding to itself.
[0038] The perception fusion method based on the convolutional neural network and the vehicle-road cooperation is used to obtain a plurality of fusion perception results of each roadside unit, the roadside unit taking itself as the main observation and the overall auxiliary observation corresponding to itself as the auxiliary observation, and a plurality of fusion perception results of each vehicle end, the vehicle end taking itself as the main observation and the overall auxiliary observation corresponding to itself as the auxiliary observation.
[0039] For each road side unit, the fusion perception results of several road side units within a certain range around the road side unit are added and combined, and then a suboptimal solution is filtered through an NMS algorithm to obtain an extended fusion perception result corresponding to each road side unit, which is used as the road side unit fusion perception result of each road side unit.
[0040] The application further provides a distributed multi-level perception fusion device based on vehicle-road cooperation, comprising,
[0041] The self fusion module is used for fusing the perception results of the vehicle end and the road side unit.
[0042] The same-end local fusion module is used for acquiring the road side unit fusion perception result obtained by taking each road side unit as a main observation and several road side units within a certain range around the road side unit as auxiliary observations and the vehicle end fusion perception result obtained by taking each vehicle end as a main observation and several vehicle ends within a certain range around the vehicle end as auxiliary observations by using the perception fusion method based on a convolutional neural network and vehicle-road cooperation.
[0043] The vehicle-road cooperation fusion module is used for acquiring the fusion perception result obtained by taking the vehicle end as a main observation and the road side unit closest to the vehicle end as an auxiliary observation by using the perception fusion method based on a convolutional neural network and vehicle-road cooperation.
[0044] The application further discloses a computer device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the method according to any one of the above-mentioned distributed multi-level perception fusion methods based on vehicle-road cooperation when executing the program.
[0045] The application further discloses a computer readable storage medium, which stores a computer program, and the program is executable on a processor to implement the method according to any one of the above-mentioned distributed multi-level perception fusion methods based on vehicle-road cooperation.
[0046] The application has the following beneficial effects:
[0047] The application completes the road side unit perception fusion in a distributed manner at the road end, and obtains the visual field gain meeting the driving demand of the automatic driving vehicle in the system with less 5G transmission data amount and vehicle end calculation amount. Thus, the demand of the vehicle end for obtaining the perception result of the surrounding road side unit is ingeniously transferred to the road side unit closest to it, and the corresponding computing power and network transmission demand are also transferred to the road side unit with higher computing power and smaller delay. Meanwhile, the perception fusion method based on the convolutional neural network and the vehicle-road cooperation is added in the fusion architecture, which not only improves the object recognition capability, but also retains the ability to increase the visual field range and reduce the blind area. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can also be obtained according to these drawings without creative labor for those skilled in the art.
[0049] Figure 1 The flowchart of the distributed multi-level perception fusion method based on vehicle-road cooperation provided by the application is shown in the figure.
[0050] Figure 2 The flowchart of the perception fusion method based on the convolutional neural network and the vehicle-road cooperation is shown in the figure.
[0051] Figure 3 The relationship between the normalized time difference and the subjective observation and auxiliary observation time difference is shown in the figure.
[0052] Figure 4 The specific flowchart of step S102 in the distributed multi-level perception fusion method based on vehicle-road cooperation provided by the application is shown in the figure.
[0053] Figure 5 Another specific flowchart of step S102 in the distributed multi-level perception fusion method based on vehicle-road cooperation provided by the application is shown in the figure.
[0054] Figure 6 The specific flowchart of step S102d in the distributed multi-level perception fusion method based on vehicle-road cooperation provided by the application is shown in the figure.
[0055] Figure 7 The schematic diagram of a typical vehicle-road cooperation intersection scene is shown in the figure.
[0056] Figure 8 The multi-level fusion architecture diagram for b car as the current vehicle to obtain gain from other vehicles and road ends is shown in the figure.
[0057] Figure 9Another multi-level fusion architecture diagram for obtaining gains from other vehicles and road ends by taking the b vehicle as the current vehicle;
[0058] Figure 10 A structure schematic diagram of the distributed multi-level perception fusion device based on vehicle-road cooperation provided by the application is provided.
[0059] Figure 11 A schematic diagram of the computer device provided by the application is provided. DETAILED DESCRIPTION
[0060] In order to make the content of the application more easily understood, the application will be further described in detail below according to the specific embodiments and in combination with the accompanying drawings.
[0061] Figure 1 A flowchart of the distributed multi-level perception fusion method based on vehicle-road cooperation provided by the embodiment of the application is provided. The method comprises steps S101-S103, and the specific steps are described as follows:
[0062] Step S101: The vehicle end and the road side unit fuse the perception results of the laser radar and the camera thereof.
[0063] Specifically, each perception module for observation in the automatic driving vehicle and the road side unit comprises at least one laser radar and one camera. Inside each automatic driving vehicle, the camera information is used to assist in adjusting the confidence of the laser radar detection model output by using the CLOCs method. Similarly, in each road side unit, the camera information is used to assist in adjusting the confidence of the laser radar detection model output by using the CLOCs method.
[0064] It should be noted that CLOCs is a method for obtaining better results by fusing the candidate boxes detected from the camera and the laser radar based on the convolutional neural network, which belongs to the prior art and will not be described here.
[0065] Step S102: The perception fusion method based on the convolutional neural network and the vehicle-road cooperation is used to obtain the road side unit fusion perception result obtained by each road side unit taking itself as the main observation and a plurality of road side units within a certain range around itself as the auxiliary observation, and the vehicle end fusion perception result obtained by each vehicle end taking itself as the main observation and a plurality of vehicle ends within a certain range around itself as the auxiliary observation.
[0066] Specifically, after fusing the cameras and LiDAR within the roadside units and the vehicle, a perception fusion method combining convolutional neural networks and vehicle-to-infrastructure (V2I) communication is used to adjust the confidence level of each roadside unit within a certain range around it. This achieves perception fusion between roadside units, where each roadside unit primarily observes itself, with several surrounding roadside units serving as auxiliary observations, thus realizing fused perception between roadside units. Simultaneously, a perception fusion method combining convolutional neural networks and V2I is used to adjust the confidence level of each autonomous vehicle within a certain range around it, achieving perception fusion between vehicles. This means each vehicle primarily observes itself, with several surrounding vehicles serving as auxiliary observations, thus realizing fused perception between vehicles.
[0067] The flowchart of the above perception fusion method based on convolutional neural networks and vehicle-road cooperation can be found in [link to flowchart]. Figure 2 Specifically, this includes:
[0068] Step a: Obtain the detection boxes for the main observation and auxiliary observation.
[0069] Specifically, the detection boxes for the main observation and auxiliary observation can be directly obtained from the current roadside unit and surrounding roadside units. It should be noted that the detection box output by each roadside unit is a 3D bounding box. The 3D bounding box is expressed as {x, y, z, l, w, h, r}, where the unit is meters, x, y, and z are the coordinates of the center point of the 3D bounding box, l is the length of the 3D bounding box, w is the width of the 3D bounding box, h is the height of the 3D bounding box, and r is the angle of rotation of the 3D bounding box around the z-axis.
[0070] Step b: Calculate the first tensor of the main observation and auxiliary observation detection boxes, and create an empty sparse matrix of the first tensor.
[0071] Specifically, the first quantity T i,j ={t i,j IoU i,j s i s j d i d j}
[0072] Where i represents the i-th object in the primary observation and j represents the j-th object in the auxiliary observation.
[0073] t i,j t represents the normalized weight of the time difference between the i-th object in the primary observation and the j-th object in the auxiliary observation. i,j The calculation method is as follows: ,in, D represents the time difference, and D represents the maximum delay. is a coefficient inversely proportional to the curvature of the curve.
[0074] IoU i,j represents the intersection over union of the i-th object in the primary observation bounding box and the j-th object bounding box in the auxiliary observation bounding box. The IoU i,j The calculation method of IoU is as follows: first, calculate the intersection area S of the primary observation bounding box and the auxiliary observation bounding box projected onto the x-y plane, then calculate the intersection length L of the primary observation bounding box and the auxiliary observation bounding box projected onto the z-axis, and multiply the intersection area S and the intersection length L to obtain the intersection volume V1, then subtract the intersection volume V1 from the sum of the volume of the primary observation bounding box and the volume of the auxiliary observation bounding box to obtain the union volume V2, and finally divide the intersection volume V1 by the union volume V2 to obtain the IoU.
[0075] s represents the confidence of the model output with respect to the primary observation bounding box.
[0076] d represents the normalized distance of the detected object from the observation center. The maximum distance is the diagonal of the three-dimensional bounding box detection range, and the corresponding normalized distance is 1.
[0077] After the calculation of the first tensor is completed, an empty sparse matrix of the first tensor is created.
[0078] It should be noted that each point in the point cloud data has its own timestamp, and the timestamp of the object is obtained by averaging the time of the points in the three-dimensional bounding box of the object. Considering that the time difference between different observations has a great influence on the credibility of the IoU, the embodiment uses a normalized weight to measure the credibility. Referring to Figure 3 , the x-axis is x and the y-axis is t. From left to right, alpha = 1 / 3, alpha = 0.5, and alpha = 1. The smaller the alpha, the smaller the delay when the weight is equal to 0.5, i.e. the objects with small delay occupy a larger proportion of credibility. The alpha is the is a coefficient inversely proportional to the curvature of the curve.
[0079] Step c: based on the first tensor of the primary observation and auxiliary observation bounding boxes, extract the combination of the primary observation bounding box and the auxiliary observation bounding box with intersection, record the index of each group of extracted bounding box combinations in the first tensor, and form a second tensor with the extracted bounding box combinations as the input tensor of the convolutional neural network.
[0080] Specifically, the combination of the primary observation bounding box and the auxiliary observation bounding box with intersection is the combination of the primary observation bounding box and the auxiliary observation bounding box with IoU>0. Therefore, the combination of the primary observation bounding box and the auxiliary observation bounding box with IoU>0 is extracted, and the extracted bounding box combination is grouped to form a second tensor as an input tensor of the convolutional neural network. At the same time, when the bounding box combination is extracted, the index of each extracted bounding box combination in the first tensor is recorded.
[0081] Step d: using 1*1 convolution to convolve the second tensor to obtain one-dimensional features of each bounding box combination in the second tensor, which is used as the confidence of the primary observation after adjustment by the auxiliary observation.
[0082] Specifically, referring to Figure 4 , the step specifically includes the following steps:
[0083] Step d-1: using 1*1 convolution to linearly transform the feature space constructed by the second tensor to another feature space, and increasing the dimension of the feature to eighteen dimensions, and using the RELU activation function to increase the nonlinear excitation of the result after dimension increasing.
[0084] Step d-2: using 1*1 convolution to linearly transform the feature space obtained in the previous step to another feature space, and increasing the dimension of the feature to thirty-six dimensions, and using the RELU activation function to increase the nonlinear excitation of the result after dimension increasing.
[0085] Step d-3: using 1*1 convolution to linearly transform the feature space obtained in the previous step to another feature space, and using the RELU activation function to increase the nonlinear excitation of the result of the new feature space.
[0086] Step d-4: using 1*1 convolution to linearly transform the feature space obtained in the previous step to another feature space, and reducing the dimension of the feature to one dimension, to obtain one-dimensional features of each bounding box combination in the second tensor, which is used as the confidence of the primary observation after adjustment by the auxiliary observation.
[0087] Step e: according to the index of each bounding box combination in the first tensor, the confidence of the bounding box combination is put into the empty sparse matrix of the first tensor. That is, according to the index of each bounding box combination recorded in the first tensor when extracting each bounding box combination in step c, the calculated confidence is put into the empty sparse matrix of the first tensor to realize the homing.
[0088] It should be noted that there may be multiple auxiliary observation bounding boxes corresponding to each primary observation bounding box. After homing the calculated confidence according to the index of each bounding box combination in the first tensor, it can be known that each calculated confidence is obtained after which auxiliary observation bounding box assists which primary observation bounding box.
[0089] Step f: using maximum pooling to select the maximum confidence from the several confidences of the auxiliary observation adjusted for each subjective observation detection frame as the confidence of the subjective observation.
[0090] The above is a flowchart of the perception fusion method based on convolutional neural network and vehicle-road cooperation. Step S102 describes obtaining the road side unit fusion perception result obtained by each road side unit taking itself as the subjective observation and several road side units within a certain range around itself as the auxiliary observation, and the vehicle end fusion perception result obtained by each vehicle end taking itself as the subjective observation and several vehicle ends within a certain range around itself as the auxiliary observation, using the perception fusion method based on convolutional neural network and vehicle-road cooperation. See Figure 5 , which specifically includes the following steps:
[0091] Step S102a: obtaining several fusion perception results obtained by each road side unit taking itself as the subjective observation and several road side units within a certain range around itself as the auxiliary observation, and fusion perception results obtained by each vehicle end taking itself as the subjective observation and several vehicle ends within a certain range around itself as the auxiliary observation, using the perception fusion method based on convolutional neural network and vehicle-road cooperation.
[0092] Specifically, for each road side unit, several other road side units are arranged within a certain range around it. For each road side unit, taking itself as the subjective observation and any one of the other road side units within a certain range around it as the auxiliary observation, a fusion perception result can be obtained by the perception fusion method based on convolutional neural network and vehicle-road cooperation. Therefore, in this step, for each road side unit, the number of other road side units arranged within a certain range around it is N, and the number of fusion perception results obtained is also N.
[0093] Similarly, for each vehicle end, taking itself as the subjective observation and any one of the other vehicle ends within a certain range around it as the auxiliary observation, a fusion perception result can be obtained by the perception fusion method based on convolutional neural network and vehicle-road cooperation. For each vehicle end, the number of other vehicle ends within a certain range around it is M, and the number of fusion perception results obtained is also M.
[0094] It should be noted that the confidence of each subjective observation detection frame is calculated by the above method. The expression form corresponding to the subjective observation detection frame is {x, y, z, l, w, h, r}, and thus the fusion perception result of each subjective observation is obtained.
[0095] Step S102b: for each road side unit, the obtained several fusion perception results are averaged or maximized to integrate, to obtain the integrated fusion perception result of each road side unit, and for each vehicle end, the obtained several fusion perception results are averaged or maximized to integrate, to obtain the integrated fusion perception result of each vehicle end.
[0096] Specifically, for each road side unit, the obtained N fusion perception results are averaged or maximized to integrate, to obtain the perception result of the current road side unit as the main observation and the other road side units within a certain range around the current road side unit as the auxiliary observation, which is the corresponding integrated fusion perception result of each road side unit.
[0097] Similarly, for each vehicle end, the obtained M fusion perception results are averaged or maximized to integrate, to obtain the perception result of the current vehicle end as the main observation and the other vehicle ends within a certain range around the current vehicle end as the auxiliary observation, which is the corresponding integrated fusion perception result of each vehicle end.
[0098] Step S102c: for each road side unit, the integrated fusion perception results of several road side units within a certain range around the road side unit are added and combined, and then the NMS algorithm is used to filter suboptimal solutions, to obtain the expanded fusion perception result corresponding to each road side unit, which is used as the road side unit fusion perception result of each road side unit, and for each vehicle end, the integrated fusion perception results of several vehicle ends within a certain range around the vehicle end are added and combined, and then the NMS algorithm is used to filter suboptimal solutions, to obtain the expanded fusion perception result corresponding to each vehicle end, which is used as the vehicle end fusion perception result of each vehicle end.
[0099] Specifically, since the perception fusion method based on convolutional neural network and vehicle-road cooperation cannot solve the blind area problem, the fusion perception results of the current road side unit and the surrounding road side units are combined together, and then the NMS algorithm is used to filter suboptimal solutions, to obtain the expanded fusion perception result of the current road side unit. The same is true for the vehicle end.
[0100] In addition, another way of obtaining the road side unit fusion perception result of each road side unit and the vehicle end fusion perception result of each vehicle end by using the perception fusion method based on convolutional neural network and vehicle-road cooperation is provided in the embodiment, which is shown in Figure 6 , and specifically includes the following steps:
[0101] Step S102a': for each road side unit, the perception results of several road side units within a certain range around itself are merged to form the overall auxiliary observation corresponding to itself, and for each vehicle end, the perception results of several vehicle ends within a certain range around itself are merged to form the overall auxiliary observation corresponding to itself.
[0102] Step S102b': the fusion perception results of each road side unit obtained by taking itself as the primary observation and the overall auxiliary observation corresponding to itself as the auxiliary observation, and the fusion perception results of each vehicle end obtained by taking itself as the primary observation and the overall auxiliary observation corresponding to itself as the auxiliary observation are obtained by using the perception fusion method based on convolutional neural network and vehicle-road cooperation.
[0103] Step S102c': for each road side unit, the fusion perception results of several road side units within a certain range around itself are added and combined, and then the suboptimal solution is filtered by the NMS algorithm to obtain the expanded fusion perception results corresponding to each road side unit, which are used as the road side unit fusion perception results of each road side unit, and for each vehicle end, the fusion perception results of several vehicle ends within a certain range around itself are added and combined, and then the suboptimal solution is filtered by the NMS algorithm to obtain the expanded fusion perception results corresponding to each vehicle end, which are used as the vehicle end fusion perception results of each vehicle end.
[0104] The above steps are to first combine other auxiliary observations around each road side unit or vehicle end to form an overall auxiliary observation, and then obtain the fusion perception results by using the perception fusion method based on convolutional neural network and vehicle-road cooperation.
[0105] Step S103: the fusion perception results obtained by taking the vehicle end as the primary observation and the nearest road side unit of the vehicle end as the auxiliary observation are obtained by using the perception fusion method based on convolutional neural network and vehicle-road cooperation.
[0106] Specifically, in step S102, each road side unit and each autonomous vehicle obtains its own fusion perception results, i.e. its own local field of view. However, the road side unit has a higher and farther observation, while the vehicle end has a closer and more accurate observation. Therefore, by taking the vehicle end as the primary observation and the nearest road side unit of the vehicle end as the auxiliary observation, and by using the perception fusion method based on convolutional neural network and vehicle-road cooperation to complete the vehicle-road cooperation fusion, the field of view gain of the driving demand of the autonomous vehicle can be obtained.
[0107] It should be noted that the vehicle end only receives the fusion perception results of the road side unit closest to itself. When performing fusion perception, if the driving direction of the autonomous vehicle is not considered, the closest road side unit matched with the vehicle may be located behind the vehicle, in which case the assistance information provided by the road side unit is greatly discounted in terms of driving assistance. Therefore, in the present embodiment, the closest road side unit to the vehicle needs to be determined first. The specific determination method is as follows:
[0108] Project the three-dimensional coordinates of the vehicle and the road side unit to the x-y plane.
[0109] Solve the equation of the straight line passing through the vehicle position coordinate point and perpendicular to the driving direction of the vehicle by using the characteristic that the data product is 0 when the vectors are perpendicular.
[0110] Substitute the coordinate point of the road side unit into the equation to screen out the road side unit in front of the vehicle.
[0111] Iterate through the screened road side units, calculate the distance between them and the current vehicle, and further screen out the closest road side unit, which is used as the auxiliary observation.
[0112] After determining the road side unit closest to the vehicle end, the vehicle end is used as the main observation, the road side unit closest to the vehicle end in front of the vehicle end is used as the auxiliary observation, and the fusion perception result of the vehicle end is obtained by using the above-mentioned fusion perception method based on convolutional neural network and vehicle-road cooperation.
[0113] Thus, the above-mentioned method ingeniously transfers the demand of the vehicle end for the perception result of the surrounding road side unit to the road side unit closest to the vehicle end, and the corresponding computing power and network transmission demand are also transferred to the road side unit with higher computing power and smaller delay. In addition, the fusion architecture is added with the fusion perception method based on convolutional neural network and vehicle-road cooperation, which not only improves the object recognition capability, but also retains the ability to reduce the blind area.
[0114] Figure 7 A typical vehicle-road cooperation intersection scene is shown in FIG. 1, in which the ellipse represents an autonomous vehicle and the rectangle represents a road side unit. Figure 8 FIG. 2 is a multi-level fusion architecture diagram in which vehicle b is the current vehicle and obtains gains from other vehicles and road ends.
[0115] The first layer of the left part of the multi-level fusion architecture diagram is the fusion of the perception results of the roadside units 1-5. Then, taking roadside unit 2 as the main observation and roadside units 1, 3, 4, and 5 as auxiliary observations, the perception fusion method based on convolutional neural network and vehicle-road cooperation (i.e., Fuse) is used to fuse the four fusion perception results, respectively, 21, 23, 24, and 25. Then, the average or maximum value of the four fusion perception results is taken to integrate, and the integrated fusion perception result 2' of roadside unit 2 is obtained. Similarly, the integrated fusion perception results 1', 3', 4', and 5' of roadside units 1, 3, 4, and 5 are obtained. The integrated fusion perception results of roadside units 1-5 are added together, and then the NMS algorithm is used to filter the suboptimal solution, and the extended perception result of roadside unit 2, i.e., the large-range and high-quality fusion perception result, is obtained. Similarly, referring to Figure 1 , the roadside units that can assist in observing roadside unit 1 within a certain range around it are roadside units 2, 3, 4, and 6, and the extended perception result of roadside unit 1 can be obtained. Similarly, all roadside units will have a large-range and high-quality fusion perception result.
[0116] The first layer of the right part of the multi-level fusion architecture diagram is the fusion of the perception results of the autonomous vehicles a, b, c, and d. Then, taking vehicle b as the main observation and vehicles a, c, and d as auxiliary observations, the perception fusion method based on convolutional neural network and vehicle-road cooperation is used to fuse the three fusion perception results, respectively, b a , b c , and b d . Then, the average or maximum value of the three fusion perception results is taken to integrate, and the integrated fusion perception result b' of vehicle b is obtained. Similarly, the fusion perception results a', c', and d' of vehicles a, c, and d are obtained. The fusion perception results of autonomous vehicles a, b, c, and d are added together, and then the NMS algorithm is used to filter the suboptimal solution, and the extended fusion perception result of vehicle b is obtained.
[0117] Since the nearest roadside unit in front of vehicle b is roadside unit 2, the extended perception result of roadside unit 2 and the extended perception result of vehicle b are fused using the perception fusion method based on convolutional neural network and vehicle-road cooperation, and the fusion perception result of vehicle b is obtained, which is used as the final perception result to guide the driving of vehicle b.
[0118] Figure 9Another multi-level fusion architecture diagram for obtaining gains from other vehicles and road ends by taking b vehicle as the current vehicle. The first layer of the left part of the multi-level fusion architecture diagram is the fusion of the perception results of the road side units 1-5. With the road side unit 2 as the main observation and the road side units 1, 3, 4, and 5 as auxiliary observations, the perception results of the road side units 1, 3, 4, and 5 are first merged (i.e., the merge step in Figure 7 ), to form the overall auxiliary observation (i.e., the auxiliary in Figure 7 ) corresponding to the road side unit 2. Then, the perception fusion method based on convolutional neural network and vehicle-road cooperation (i.e., Fuse) is adopted, with the road side unit 2 as the main observation and the overall auxiliary observation of the road side units 1, 3, 4, and 5 after merging as the auxiliary observation, to obtain the fused perception result 2' of the road side unit 2. Similarly, the fused perception results 1', 3', 4', and 5' of the road side units 1, 3, 4, and 5 can be obtained. The fused perception results of the road side units 1-5 are added and combined together, and then the NMS algorithm is used to filter the suboptimal solution, to obtain the extended perception result 2'' of the road side unit 2, i.e., the large-range and high-quality fused perception result. Similarly, referring to Figure 1 , the road side units that can perform auxiliary observations for the road side unit 1 within a certain range around the road side unit 1 are the road side units 2, 3, 4, and 6, and thus the extended perception result 1'' of the road side unit 1 can be obtained. Similarly, all road side units will have large-range and high-quality fused perception results.
[0119] The first layer of the right part of the multi-level fusion architecture diagram is the fusion of the perception results of the autonomous driving vehicles a, b, c, and d. With the b vehicle as the main observation and the a vehicle, c vehicle, and d vehicle as auxiliary observations, the perception results of the a vehicle, c vehicle, and d vehicle are first merged (i.e., the merge step in Figure 7 ), to form the overall auxiliary observation (i.e., the auxiliary in Figure 7 ) corresponding to the b vehicle. Then, the perception fusion method based on convolutional neural network and vehicle-road cooperation (i.e., Fuse) is adopted, with the b vehicle as the main observation and the overall auxiliary observation formed by the a vehicle, c vehicle, and d vehicle after merging as the auxiliary observation, to obtain the fused perception result b' of the b vehicle. Similarly, the fused perception results a', c', and d' of the a vehicle, c vehicle, and d vehicle can be obtained. The fused perception results of the autonomous driving vehicles a, b, c, and d are added and combined together, and then the NMS algorithm is used to filter the suboptimal solution, to obtain the extended fused perception result b'' of the b vehicle.
[0120] Since the nearest road side unit in front of the b vehicle is the road side unit 2, finally, the perception result of the road side unit 2 after extension and the perception result of the b vehicle after extension are fused by using the perception fusion method based on convolutional neural network and vehicle-road cooperation, to obtain the fused perception result of the b vehicle, which is used as the final perception result for guiding the driving of the b vehicle.
[0121] Figure 10 A structure schematic diagram of a distributed multi-level perception fusion device based on vehicle-road cooperation is provided for an embodiment of the present application. The device comprises a self fusion module, a same-end local fusion module, and a vehicle-road cooperation fusion module.
[0122] Specifically, the self fusion module is configured to fuse the perception results of the laser radar and the camera of the vehicle end and the road side unit.
[0123] The same-end local fusion module is configured to obtain the road side unit fusion perception result obtained by each road side unit taking itself as the main observation and a plurality of road side units within a certain range around itself as the auxiliary observation, and the vehicle end fusion perception result obtained by each vehicle end taking itself as the main observation and a plurality of vehicle ends within a certain range around itself as the auxiliary observation, by using a perception fusion method based on a convolutional neural network and vehicle-road cooperation.
[0124] The vehicle-road cooperation fusion module is configured to obtain the fusion perception result obtained by taking the vehicle end as the main observation and the closest road side unit of the vehicle end as the auxiliary observation, by using a perception fusion method based on a convolutional neural network and vehicle-road cooperation.
[0125] Referring to Figure 11 The embodiment also provides a computer device, and components of the computer device can include but are not limited to one or more processors or processing units, system memory, and a bus connecting different system components including the system memory and the processing unit.
[0126] The bus represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor or local bus using any of a variety of bus architectures. For example, these architectures include but are not limited to an industry standard architecture (ISA) bus, a microchannel architecture (MAC) bus, an enhanced ISA bus, a video electronics standards association (VESA) local bus, and a peripheral component interconnect (PCI) bus.
[0127] The computer system / server typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer system / server and includes both volatile and nonvolatile media, removable and non-removable media.
[0128] The system memory can include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The computer device can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system suitable for storing physical memory can be used for reading from and writing to non-removable, non-volatile magnetic media (e.g., a "hard drive"). A magnetic disk drive can also be provided for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive can be provided for reading from or writing to a removable, non-volatile optical disk (e.g., a CD-ROM, DVD-ROM or other optical media). In these instances, each can be connected to the bus by one or more data media interfaces. The memory can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the application.
[0129] Program / utility, having a set (at least one) of program modules, can be stored in, for example, memory by way of example, such program modules include an operating system, one or more application programs, other program modules, and program data, each of or some combination of which can include an implementation of a network environment. Program modules are generally executed by processing unit and / or provided data to other program modules in memory, as desired.
[0130] The computer device can also communicate with one or more external devices such as a keyboard, a pointing device, a display, etc. through input / output (I / O) interface(s). Additionally, the computer device can communicate with one or more networks such as a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet, through network adapter.
[0131] The processing unit executes the program modules stored in the system memory to perform the functions and / or methods of embodiments of the application.
[0132] The computer program mentioned above can be provided in a computer storage medium, i.e. the computer storage medium is encoded with the computer program, which, when executed by one or more computers, causes the one or more computers to perform the method flow and / or device operation shown in the above embodiments of the application.
[0133] With the development of time and technology, the medium meaning is more and more extensive, and the computer program propagation path is no longer limited to a tangible medium, but can also be directly downloaded from a network, etc. Any combination of one or more computer-readable media can be used. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.
[0134] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such propagated data signals can take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can be used to carry or store program code for use by or in connection with an instruction execution system, apparatus or device.
[0135] The program code contained on the computer-readable medium can be transmitted using any suitable medium, including, but not limited to, wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0136] Computer program code for carrying out operations of the present application can be written in one or more programming languages or combinations of languages including object-oriented, such as Java, Smalltalk, C++, conventional procedural programming languages, such as the "C" programming language, or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0137] Besides the above-mentioned embodiments, the present application can also have other embodiments; any technical solution formed by equivalent replacement or equivalent transformation falls within the protection scope required by the present application.
Claims
1. A distributed multi-level perception fusion method based on vehicle-road cooperation, characterized in that: The application relates to a vehicle-end and roadside-unit fusion perception method based on a convolutional neural network and vehicle-road cooperation. The vehicle-end and roadside-unit fusion perception method based on the convolutional neural network and vehicle-road cooperation comprises the following steps: The vehicle-end and roadside-unit fusion perception method based on the convolutional neural network and vehicle-road cooperation comprises the following steps: The vehicle-end and roadside-unit fusion perception method based on the convolutional neural network and vehicle-road cooperation comprises the following steps: The method for determining the nearest roadside unit in front of the vehicle-end comprises the following steps: The three-dimensional coordinates of the vehicle and the roadside unit are projected onto an x-y plane; The equation of a straight line perpendicular to the driving direction of the vehicle is solved by using the characteristic that the data product is 0 when vectors are perpendicular; The roadside unit coordinate points in front of the vehicle are screened out by substituting the roadside unit coordinate points into the equation; The distance between the screened roadside unit and the current vehicle is calculated, and the nearest roadside unit is further screened out to serve as the auxiliary observation.
2. The method of claim 1, wherein: The vehicle-end and roadside-unit fusion perception method based on the convolutional neural network and vehicle-road cooperation comprises the following steps: The confidence of the laser radar detection model output of the vehicle-end and roadside unit is adjusted by using the camera information in the vehicle-end and roadside unit as auxiliary observation through the CLOCs method.
3. The method of claim 1, wherein: The vehicle-end and roadside-unit fusion perception method based on the convolutional neural network and vehicle-road cooperation comprises the following steps: The detection frame of the main observation and the auxiliary observation is obtained; The first tensor of the main observation and the auxiliary observation detection frame is calculated, and the empty sparse matrix of the first tensor is created; The combination of the main observation detection frame and the auxiliary observation detection frame with an intersection is extracted based on the first tensor of the main observation and the auxiliary observation detection frame, the index of each group of extracted detection frame combinations in the first tensor is recorded, and the extracted detection frame combinations are composed into a second tensor as the input tensor of the convolutional neural network; The second tensor is convolved by using a 1*1 convolution, and the one-dimensional feature of each group of detection frame combinations in the second tensor is obtained as the confidence of the main observation adjusted by the auxiliary observation; The confidence of the detection frame combination is put into the empty sparse matrix of the first tensor according to the index of each group of detection frame combinations in the first tensor; The maximum confidence is selected as the confidence of the main observation by using the maximum pooling from the several confidences of the main observation detection frame adjusted by the auxiliary observation.
4. The method of claim 3, wherein: The vehicle-end and roadside-unit fusion perception method based on the convolutional neural network and vehicle-road cooperation comprises the following steps: The first tensor T i,j = {t i,j , IoU i,j , s i , s j , d i , d j}, Wherein, i represents the ith object in the main observation, and j represents the jth object in the auxiliary observation; t i,j denotes the time difference normalized weight of the ith object in the main observation and the jth object in the auxiliary observation, and the t i,j The calculation method is as follows: where where, denotes the time difference, D denotes the maximum delay, and a is a coefficient inversely proportional to the curvature of the curve. IoU i,j represents the intersection over union of the i-th object detection box in the primary observation detection box and the j-th object detection box in the auxiliary observation detection box, and the IoU i,j The calculation method of the IoU is as follows: first, the intersection area S of the primary observation detection box and the auxiliary observation detection box projected onto the x-y plane is calculated, then the intersection length L of the primary observation detection box and the auxiliary observation detection box projected onto the z axis is calculated, and the intersection volume V1 is obtained by multiplying the intersection area S and the intersection length L; the union volume V2 is obtained by subtracting the intersection volume V1 from the sum of the volume of the primary observation detection box and the volume of the auxiliary observation detection box; and finally, the IoU is obtained by dividing the intersection volume V1 by the union volume V2. S represents the confidence of the main observation detection frame output by the model; D represents the normalized distance of the detected object from the observation center.
5. The method of claim 3, wherein: The vehicle-end and roadside-unit fusion perception method based on the convolutional neural network and vehicle-road cooperation comprises the following steps: The second tensor is linearly transformed into another feature space using a 1*1 convolution, the dimension of the feature is increased to eighteen, and a RELU activation function is used on the result after dimension increase to increase the non-linear excitation; The feature space obtained in the previous step is linearly transformed into another feature space using a 1*1 convolution, the dimension of the feature is increased to thirty-six, and a RELU activation function is used on the result after dimension increase to increase the non-linear excitation; The feature space obtained in the previous step is linearly transformed into another feature space using a 1*1 convolution, and a RELU activation function is used on the result of the new feature space to increase the non-linear excitation; The feature space obtained in the previous step is linearly transformed into another feature space using a 1*1 convolution, the dimension of the feature is reduced to one, and the one-dimensional feature of each group of bounding box combination in the second tensor is obtained as the confidence of the subjective observation after adjustment of the auxiliary observation.
6. The method of claim 5, wherein: The perception fusion method based on the convolutional neural network and the vehicle-road cooperation obtains the road side unit fusion perception result obtained by each road side unit taking itself as a main observation and a plurality of road side units within a certain range around itself as auxiliary observations, and the vehicle end fusion perception result obtained by each vehicle end taking itself as a main observation and a plurality of vehicle ends within a certain range around itself as auxiliary observations, and the perception fusion method based on the convolutional neural network and the vehicle-road cooperation comprises the following steps of: The perception fusion method based on the convolutional neural network and the vehicle-road cooperation obtains the road side unit fusion perception result obtained by each road side unit taking itself as a main observation and a plurality of road side units within a certain range around itself as auxiliary observations, and the vehicle end fusion perception result obtained by each vehicle end taking itself as a main observation and a plurality of vehicle ends within a certain range around itself as auxiliary observations; For each road side unit, the plurality of fusion perception results are integrated by taking the average value or the maximum value, to obtain the integrated fusion perception result of each road side unit, and for each vehicle end, the plurality of fusion perception results are integrated by taking the average value or the maximum value, to obtain the integrated fusion perception result of each vehicle end; For each road side unit, the integrated fusion perception results of the plurality of road side units within a certain range around itself are added and combined, and then a NMS algorithm is used to filter suboptimal solutions, to obtain the expanded fusion perception result corresponding to each road side unit, which is used as the road side unit fusion perception result of each road side unit, and for each vehicle end, the integrated fusion perception results of the plurality of vehicle ends within a certain range around itself are added and combined, and then a NMS algorithm is used to filter suboptimal solutions, to obtain the expanded fusion perception result corresponding to each vehicle end, which is used as the vehicle end fusion perception result of each vehicle end.
7. The method of claim 5, wherein: The perception fusion method based on the convolutional neural network and the vehicle-road cooperation obtains the road side unit fusion perception result obtained by each road side unit taking itself as a main observation and a plurality of road side units within a certain range around itself as auxiliary observations, and the vehicle end fusion perception result obtained by each vehicle end taking itself as a main observation and a plurality of vehicle ends within a certain range around itself as auxiliary observations, and the perception fusion method based on the convolutional neural network and the vehicle-road cooperation comprises the following steps of: For each roadside unit, the perception results of several roadside units within a certain range around itself are combined to form the overall auxiliary observation corresponding to itself, and for each vehicle end, the perception results of several vehicle ends within a certain range around itself are combined to form the overall auxiliary observation corresponding to itself; The fusion perception result of each roadside unit is obtained by taking the roadside unit itself as the primary observation and the overall auxiliary observation corresponding to the roadside unit itself as the auxiliary observation, and the fusion perception result of each vehicle end is obtained by taking the vehicle end itself as the primary observation and the overall auxiliary observation corresponding to the vehicle end itself as the auxiliary observation, by using the perception fusion method based on the convolutional neural network and the vehicle-road cooperation; For each roadside unit, the fusion perception results of several roadside units within a certain range around itself are added and combined, and then the suboptimal solution is filtered by using the NMS algorithm to obtain the expanded fusion perception result corresponding to each roadside unit, which is taken as the roadside unit fusion perception result of each roadside unit, and for each vehicle end, the fusion perception results of several vehicle ends within a certain range around itself are added and combined, and then the suboptimal solution is filtered by using the NMS algorithm to obtain the expanded fusion perception result corresponding to each vehicle end, which is taken as the vehicle end fusion perception result of each vehicle end.
8. A distributed multi-level perception fusion device based on vehicle-road cooperation, characterized in that: The method comprises the following steps: The self fusion module is used for fusing the perception results of the vehicle end and the roadside unit; The same-end local fusion module is used for obtaining the roadside unit fusion perception result of each roadside unit by taking the roadside unit itself as the primary observation and several roadside units within a certain range around the roadside unit as the auxiliary observation, and obtaining the vehicle end fusion perception result of each vehicle end by taking the vehicle end itself as the primary observation and several vehicle ends within a certain range around the vehicle end as the auxiliary observation, by using the perception fusion method based on the convolutional neural network and the vehicle-road cooperation; The vehicle-road cooperation fusion module is used for obtaining the fusion perception result by taking the vehicle end as the primary observation and the roadside unit closest to the vehicle end as the auxiliary observation, by using the perception fusion method based on the convolutional neural network and the vehicle-road cooperation. The method for determining the roadside unit closest to the front of the vehicle end comprises the following steps: The three-dimensional coordinates of the vehicle and the roadside unit are projected onto the x-y plane; The equation of the straight line perpendicular to the driving direction of the vehicle is solved by using the characteristic that the data product is 0 when the vectors are perpendicular; The roadside unit coordinate point is substituted into the equation to screen out the roadside unit in front of the vehicle; The distance between the screened roadside unit and the current vehicle is calculated, and the roadside unit closest to the vehicle is further screened out and taken as the auxiliary observation.
9. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: The processor implements the method of any one of claims 1-7 when executing the program.
10. A computer readable storage medium having stored thereon a computer program, characterized in that: The program is executed by the processor to implement the method of any one of claims 1-7.
Citation Information
Patent Citations
Obstacle perception fusion method and device, equipment and storage medium
CN112418092A
Perception fusion method, device and equipment based on vehicle infrastructure cooperation and medium
CN113537362A
Vehicle sensing information fusion method, device and equipment and storage medium
CN114386481A
Perception fusion method, device and equipment based on convolutional neural network and vehicle-road cooperation, and storage medium
CN115438712A