Autonomous driving vehicle-road collaborative perception method, device and computer-readable storage medium
By using the vehicle re-identification network and hashing algorithm for feature processing in autonomous vehicle-road collaborative perception, and combining the ICP algorithm for three-dimensional information fusion, the problems of large bandwidth consumption and poor real-time performance in the existing technology are solved, and more efficient and real-time vehicle-road collaborative perception is achieved.
Patent Information
- Application Number
- CN202211357496.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-01
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-11-01
AI Technical Summary
The existing autonomous driving vehicle-road collaborative perception methods consume a lot of bandwidth and are not very real-time.
Feature extraction and matching are performed through the vehicle re-identification network, and feature conversion is performed using the hash algorithm of random hyperplane projection, and three-dimensional information is fusion with the ICP algorithm to reduce the amount of data transmission and improve real-time.
It reduces bandwidth consumption, improves the real-time nature of collaborative perception, and expands the field of view of the perception device.
Smart Images

Figure CN115743171B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving, and specifically to an autonomous driving vehicle-road collaborative perception method, device, and computer-readable storage medium. Background Art
[0002] In the field of autonomous driving vehicle-road collaboration, more research is focused on more efficient network communication protocols, embedded neural network acceleration and other technologies. The perception information interaction between vehicles and roads relies heavily on the positioning information of the Global Positioning System (GPS), resulting in a low degree of accuracy in collaborative intelligence. However, under the current limited actual communication bandwidth, there is less research on how to reduce the amount of collaborative perception data transmission and efficiently complete collaborative perception. More research is needed to supplement this. Especially in the actual scenario of vehicle-road collaboration, which is time-sensitive, a large amount of vehicle-side and road-side perception data interact during peak hours. If the vehicle-road collaborative perception information is not specially processed, network bandwidth congestion will inevitably occur, further affecting the efficiency and real-time performance of vehicle-road collaborative perception. Summary of the invention
[0003] The main technical problem that this application solves is that the existing autonomous driving vehicle-road collaborative perception methods consume a lot of bandwidth and are not very real-time.
[0004] According to the first aspect, an embodiment provides an autonomous driving vehicle-road collaborative perception method, including: a first perception object obtains image information and three-dimensional information of at least one vehicle in a first field of view, and a second perception object obtains image information and three-dimensional information of at least one vehicle in a second field of view, and the image information and three-dimensional information of at least one vehicle in the first field of view and the image information and three-dimensional information of at least one vehicle in the second field of view include the same vehicle; the first perception object extracts features of the image information of at least one vehicle in the first field of view based on a vehicle re-identification network to obtain a first vector feature, and the second perception object extracts features of the image information of at least one vehicle in the second field of view based on a vehicle re-identification network to obtain a second vector feature; the second perception object determines whether to perform feature conversion based on the sum of the number of vehicles contained in the first field of view and the second field of view, and if the If the total number of vehicles is greater than a first threshold, the first perception object converts the first vector feature based on a hash algorithm of random hyperplane projection to obtain a first binary feature, and the second perception object converts the second vector feature based on a hash algorithm of random hyperplane projection to obtain a second binary feature; the first perception object sends the first binary feature and the three-dimensional information of at least one vehicle in the first field of view to the second perception object, and the second perception object matches the first binary feature and the second binary feature to obtain a matching result, which includes at least one pair of successfully matched vehicles; the second perception object calculates the three-dimensional information corresponding to the at least one pair of successfully matched vehicles based on the ICP algorithm to obtain a coordinate transformation matrix; the second perception object fuses the three-dimensional information of the first field of view and the three-dimensional information of the second field of view based on the coordinate transformation matrix to obtain fused three-dimensional information.
[0005] In one embodiment, another autonomous driving vehicle-road collaborative perception method is also provided, including: obtaining image information and three-dimensional information of at least one vehicle in a second field of view, the image information and three-dimensional information of at least one vehicle in the second field of view and the image information and three-dimensional information of at least one vehicle in the first field of view include the same vehicle, and the image information and three-dimensional information of at least one vehicle in the first field of view are obtained by a first perception object; performing feature extraction on the image information of at least one vehicle in the second field of view based on a vehicle re-identification network to obtain a second vector feature; determining whether to perform feature conversion based on the sum of the number of vehicles included in the first field of view and the second field of view, if the sum of the number of vehicles is greater than a first threshold, converting the second vector feature based on a hash algorithm of random hyperplane projection to obtain a second binary feature; receiving a first binary feature sent by the first perception object and the three-dimensional information of at least one vehicle in the first field of view, and then matching the first binary feature and the second binary feature to obtain a matching result, the matching result including at least one pair of successfully matched vehicles; calculating the three-dimensional information corresponding to the at least one pair of successfully matched vehicles based on an ICP algorithm to obtain a coordinate transformation matrix; fusing the three-dimensional information of the first field of view and the three-dimensional information of the second field of view based on the coordinate transformation matrix to obtain fused three-dimensional information.
[0006] The fused three-dimensional information is fused three-dimensional information, which combines the information of the first field of view and the second field of view, thereby expanding the field of view of the perception device. The fused three-dimensional information can be presented on the first perception object or the second perception object.
[0007] According to the second aspect, an embodiment provides an electronic device, comprising: a memory; a processor; and a computer program; wherein the computer program is stored in the memory and configured to be executed by the processor to implement the method as described in the first aspect.
[0008] According to the third aspect, an embodiment provides a computer-readable storage medium, on which a program is stored, and the program can be executed by a processor to implement the method as described in any one of the above-mentioned first aspects.
[0009] According to the autonomous driving vehicle-road collaborative perception method, device and computer-readable storage medium of the above-mentioned embodiments, the bandwidth consumption is reduced through the vehicle re-identification network, the real-time performance of collaborative perception is improved, and the field of view of the perception device is expanded. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 A scene diagram of an autonomous driving vehicle-road collaborative perception method provided in an embodiment of the present application;
[0011] Figure 2 A schematic diagram of a vehicle-road collaborative perception method for autonomous driving provided in an embodiment of the present application;
[0012] Figure 3 A schematic diagram of a vehicle re-identification network provided in an embodiment of the present application;
[0013] Figure 4 A schematic diagram of vehicle feature conversion provided in an embodiment of the present application;
[0014] Figure 5 A schematic diagram of the fusion of an ICP algorithm provided in an embodiment of the present application;
[0015] Figure 6 A schematic diagram after fusion provided in an embodiment of the present application;
[0016] Figure 7 A schematic diagram of a fusion effect after vehicle matching provided in an embodiment of the present application;
[0017] Figure 8 A schematic diagram of feature matching and feature fusion provided in an embodiment of the present application;
[0018] Fig. 9 A schematic diagram of three-dimensional information fusion on both sides of a vehicle and a road provided in an embodiment of the present application;
[0019] Fig.10 A schematic diagram of an electronic device provided in this application. DETAILED DESCRIPTION
[0020] The present application is further described in detail below by specific embodiments in conjunction with the accompanying drawings. Wherein similar elements in different embodiments adopt associated similar element numbers. In the following embodiments, many detailed descriptions are intended to enable the present application to be better understood. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, and methods. In some cases, some operations related to the present application are not shown or described in the specification, in order to avoid the core part of the present application being overwhelmed by too much description, and for those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the description in the specification and the general technical knowledge in the art.
[0021] In addition, the features, operations or characteristics described in the specification can be combined in any appropriate manner to form various implementations. At the same time, the steps or actions in the method description can also be interchanged or adjusted in a manner that is obvious to those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for the purpose of clearly describing a certain embodiment and are not meant to be a required sequence, unless otherwise specified that a certain sequence must be followed.
[0022] The serial numbers of the components in this document, such as "first", "second", etc., are only used to distinguish the objects described and do not have any order or technical meaning. The "connection" and "coupling" mentioned in this application, unless otherwise specified, include direct and indirect connections (couplings).
[0023] Real-time perception interaction between vehicle terminals and road terminals can minimize driving blind spots. The method proposed in this application is an efficient vehicle-road collaborative perception based on vehicle re-identification, which involves the most core aspect of vehicle-road collaborative perception - perception fusion. The algorithms mainly involved in vehicle-road collaborative perception fusion are two categories: re-identification and perception fusion methods.
[0024] In terms of vehicle-road collaboration, more research is done on technologies such as more efficient network communication protocols and neural network acceleration on the embedded side. The perception information interaction between vehicles and roads relies heavily on GPS positioning information, resulting in a low degree of accuracy in collaborative intelligence. However, under the current limited actual communication bandwidth, there is less research on the important issue of how to reduce the amount of collaborative perception data transmission and efficiently complete collaborative perception, and more research is needed to supplement it. Especially in the actual scenario of vehicle-road collaboration, which is time-sensitive, a large amount of vehicle-side and road-side perception data interacts during peak vehicle hours. If the vehicle-road collaborative perception information is not specially processed, network bandwidth congestion will inevitably occur, further affecting the efficiency and real-time performance of vehicle-road collaborative perception. In order to solve this communication pain point and to reduce the size of transmitted data as much as possible while ensuring the accuracy and speed of collaborative perception, this application has launched a study on a method of using lightweight re-identification technology to reduce the amount of perception data on both sides of the vehicle and road, reduce transmission delays, and efficiently perform perception fusion.
[0025] A search of existing literature found that in existing vehicle-road cooperative system research, the perception information interaction between vehicles and roads relies heavily on GPS positioning information, resulting in a low degree of accuracy in cooperative intelligence. In addition, the amount of data transmitted by vehicle-road cooperation is large, resulting in insufficient transmission and feature matching speed.
[0026] As an advanced stage of autonomous driving, vehicle-road collaboration has only become a research focus in recent years. An important part of vehicle-road collaboration is the need for rapid interaction and fusion of vehicle-road information, and completing this task requires rapid feature recognition, transmission, matching and fusion.
[0027] The original purpose of vehicle re-identification was to better solve the problems of traffic supervision and criminal investigation. It has evolved from early sensor-based methods and artificial feature-based methods to current methods mainly based on deep learning. Since the focus of vehicle re-identification is feature extraction and comparison, and deep learning is good at extracting target features, deep learning has become a hot method for vehicle re-identification. In the scenario of vehicle-road collaboration, vehicle re-identification combined with vehicle-road collaboration can bring more possibilities, and can associate roadside and vehicle-side targets to truly achieve the unification of vehicle-road perception results. However, it also faces some more serious problems. For example, the current vehicle re-identification deep learning model has many parameters and is not suitable for deployment in computing devices with low computing power on the car side and road side. In addition, the bandwidth on the car side and road side is limited. During peak hours and high-speed operation, the delay and communication volume of vehicle-road collaboration interaction need to be shorter and smaller. Most of the current model methods are not enough to meet these requirements, which is also one of the factors restricting the development of vehicle-road collaboration.
[0028] With the further development of sensor technology, deep learning has also begun to shine in the field of 3D target detection, including 3D target detection based on LiDAR point cloud data, 3D target detection based on camera and LiDAR data fusion, monocular 3D target detection method and binocular 3D target detection method, etc. These methods are based on vehicle-side data for perception, and it is difficult to obtain more comprehensive vehicle driving environment perception data. Since the process of vehicle-road collaborative perception requires the interaction of perception data on the vehicle side and the road side, how to perceive and fuse the 3D perception data on both sides in the shortest time is related to the final implementation effect of vehicle-road collaboration. However, the commonly used point cloud registration methods such as iterative closest point (ICP) (registration methods based on point primitives, registration methods based on geometric primitives, registration methods based on voxel primitives, etc.) are not suitable for direct application in time-sensitive scenarios such as vehicle-road collaboration. The fusion method based on GPS positioning is not suitable for the fusion of collaborative perception information because it is not intelligent enough and accurate enough.
[0029] The embodiments of the present application include multiple usage scenarios. Figure 1 A scene diagram of a vehicle-road collaborative perception method for autonomous driving in one embodiment.
[0030] In the present application, the following examples are provided: Figure 2 The method for autonomous driving vehicle-road cooperative perception shown in the figure includes steps S201 to S206:
[0031] Step S201, the first sensing object obtains image information and three-dimensional information of at least one vehicle in a first field of view, and the second sensing object obtains image information and three-dimensional information of at least one vehicle in a second field of view, and the image information and three-dimensional information of the at least one vehicle in the first field of view and the image information and three-dimensional information of the at least one vehicle in the second field of view include the same vehicle;
[0032] In one embodiment, the first sensing object may be a sensing device on the road side or a sensing device on the vehicle side. In one embodiment, the second sensing object may be a sensing device on the road side or a sensing device on the vehicle side.
[0033] In one embodiment, the vehicle in the first field of view or the second field of view can also be replaced by other objects in the first field of view or the second field of view, which can also achieve the effect of expanding the field of view. The image information and three-dimensional information of other objects in the first field of view or the second field of view correspond one to one. Other objects may include flowers, grass, people, trees, street lights, buildings, road signs, signs, etc., and this application does not limit this. The present application can generate fused three-dimensional information based on other objects in the first field of view or the second field of view to expand the field of view. In one embodiment, the image information and three-dimensional information of at least one vehicle correspond one to one.
[0034] The three-dimensional information includes information about the position, size, and direction of the vehicle. In one embodiment, the position information may be the x, y, and z coordinates of the center point of the three-dimensional bounding box of the vehicle with the sensing device as the origin of the coordinate system. x, y, and z are letters representing spatial coordinates, x represents the horizontal coordinate, y represents the vertical coordinate, and z represents the height. The sensing device is a vehicle-side sensor or a roadside sensor. The size information is the length, width, and height of the vehicle. The direction can be represented as the yaw angle information of the vehicle. The eight vertices of the three-dimensional bounding box of the vehicle can be calculated based on the x, y, and z coordinates of the center point of the three-dimensional bounding box of the vehicle and the length, width, height of the vehicle, and the yaw angle information of the vehicle. Therefore, the three-dimensional information may be the x, y, and z coordinates of the center point of the three-dimensional bounding box of the vehicle, the length, width, height of the vehicle, and the yaw angle of the vehicle.
[0035] In one embodiment, the three-dimensional information may also be three-dimensional point cloud information of a vehicle target with the sensing device as the origin of the coordinate system, and the sensing device may be a vehicle-side sensor or a road-side sensor. For example, the three-dimensional information may be the original three-dimensional point cloud coordinate value of the vehicle or a simplified partial three-dimensional point cloud coordinate value.
[0036] The three-dimensional information may also be other types of information, which is not limited here.
[0037] The first field of view and the second field of view include the same vehicle. For example, the first field of view includes vehicles 1, 2, and 3, and the second field of view includes vehicles 3, 4, and 5. The first field of view and the second field of view include the same vehicle 3. Based on the same vehicle 3, the information of the first field of view and the second field of view can be integrated to achieve an expanded field of view.
[0038] Step S202: the first sensing object extracts features from the image information of at least one vehicle in the first field of view based on the vehicle re-identification network to obtain a first vector feature, and the second sensing object extracts features from the image information of at least one vehicle in the second field of view based on the vehicle re-identification network to obtain a second vector feature;
[0039] The model of the vehicle re-identification network consists of a Conv1 layer, a MaxPool layer (maximum pooling layer), a Stage2 layer, a Stage3 layer, a Stage4 layer, a Conv5 layer, an AdaptiveAvgPool2d layer (adaptive average pooling layer), a BNNeck layer, and a fully connected layer. The Conv1 layer includes a convolutional layer, a BatchNorm layer, and a SiLU activation function layer. The Stage2 layer includes an InvertedResidual structure (inverted residual structure), a convolutional layer, a BatchNorm2d layer, a SiLU activation function layer, a maximum pooling layer, and a SENet attention layer (Squeeze-and-Excitation Module), Linear layer (linear layer), the Stage3 layer includes InvertedResidual structure (inverted residual structure), convolution layer, BatchNorm2d layer, SiLU activation function layer, maximum pooling layer, SENet attention layer, Linear layer (linear layer), the Stage4 layer includes InvertedResidual structure (inverted residual structure), convolution layer, BatchNorm2d layer, SiLU activation function layer, maximum pooling layer, SENet attention layer, Linear layer (linear layer), the Conv5 layer includes convolution layer, BatchNorm layer, SiLU activation function layer, the BNNeck layer includes a BatchNorm1d layer, and the fully connected layer includes a Linear layer (linear layer); the activation function of the model of the vehicle re-identification network is SiLU (Sigmoid Linear Unit), and half-precision floating point parameters are used for training. In some embodiments, Figure 3 Schematic diagram of the vehicle re-identification network model.
[0040] In one embodiment, the re-identification network model is modified based on ShuffleNetV2, and the modification includes: 1) using the pre-trained weights of the feature extraction backbone network of shufflenetV2 2.0 as the pre-trained weights of our re-identification model, 2) replacing the original ReLU (Rectified Linear Unit) activation function in the network with the SiLU activation function to prevent the occurrence of gradient saturation in the original model training, 3) adding the SENet (Squeeze-and-Excitation Module) attention mechanism to the feature extraction backbone part, 4) adding BNNeck before the fully connected layer of the network, and 5) using half-precision floating point parameters for model training to further reduce the model size.
[0041] In the deep feature extraction stage of the re-identification network model, the loss function is the key to network training. In the existing vehicle re-identification algorithms, most of them use the joint loss function of Cross-entropy loss (cross entropy loss function) and Triplet Loss (triplet loss function) to induce training of the re-identification network. This training method can make the vehicle image features extracted by the model have the advantages of small intra-class distance and large inter-class distance in the feature metric space. However, in subsequent studies, in order to maximize the free compression of features, a hash algorithm of random hyperplane projection is used to divide the feature space to achieve feature data dimensionality reduction. Therefore, the target features extracted by the trained re-identification model must also have excellent clustering performance in the metric space, and then the features of the same vehicle are more likely to be divided into the same space after dimensionality reduction. Therefore, the present invention designs the application of ArcFace loss function to constrain vehicle re-identification feature extraction network training, maps vehicle image features to a hypersphere, and then constrains the angular distance of image features in vector space, optimizing the spatial separability of features. Therefore, the present invention combines Cross-entropy loss, Triplet Loss and ArcFace Loss to jointly perform constrained training on the vehicle re-identification feature extraction network, effectively improving the feature extraction performance of the network and maintaining good re-identification accuracy after feature dimensionality reduction, which can more intelligently narrow the target three-dimensional information feature matching range for our subsequent vehicle-road collaborative perception fusion.
[0042] The loss function during the training of the re-identification network model is L final , L Final =λ1L CE +λ2L T +λ3L A , where λ1, λ2, λ3 are weighting coefficients L CE is the cross entropy loss function, L Tis the triple loss function, L A is the ArcFace loss function; where y i is the predicted classification value, y i is the actual classification value; L T =(d a,p -d a,n +m), where m is the margin parameter; d a,p is the feature distance between positive sample pairs; d a,n is the feature distance between negative sample pairs;
[0043]
[0044] Where N is the batch size; n is the number of categories; s is the scaling factor; m is the angle interval parameter; θ j W j and x i The angle between (the weight of the last fully connected layer (FC, Fully connected layer) in the deep learning network is W, x i is the feature of the deep learning network input fully connected layer) (the FC layer weight of the j∈{1, 2, .., yi, .., n} category is W j , T represents the transpose in matrix multiplication;), the angle θ j Available through Calculated; y i Represents the feature x during training i The real category, For prediction y i The target weight of the category, for and x i The angle between Available through Calculated.
[0045] If the vehicle re-identification network model is trained only with the commonly used cross entropy loss function (L CE ) and triplet loss function (L T), after training, the binary bit features converted by the hash function (LSH) based on the random projection method are matched with the target, and the matching accuracy will be greatly reduced compared with that before the feature conversion. For example, when the feature is increased from 2048bit to 1024bit and then to 512bit, a large rank1 will appear every time it decreases by half, causing the accuracy to decrease. This is extremely unfavorable for target re-identification in vehicle-road collaboration. Rank indicates the probability that the n images with the highest confidence in the search results have the correct result. Rank 1 indicates the first hit, and rank k indicates the hit within the kth time. For this reason, in addition to the commonly used cross entropy loss function and triple loss function, the present invention also adds the Arcface loss function (L A ).
[0046] Table 1 shows the experimental verification table before and after adding the Arcface loss function. As shown in Table 1, through experimental verification on the VeRi776 dataset, after adding the Arcface loss function, the feature matching accuracy loss after the floating-point feature is converted to the bit feature through LSH can be reduced. After testing, the re-identification feature lengths of 2048bit, 1024bit, and 512bit are almost unchanged compared with the feature accuracy before compression. After adding the Arcface loss function, while significantly improving the model detection accuracy, it also greatly reduces the accuracy loss caused by feature compression conversion. For example, from the Rank1 accuracy loss of 35.7% from the 2048bit to 64bit feature conversion before adding the Arcface loss function, to the Rank1 accuracy loss of 19.9% from the feature conversion, the accuracy loss rate is reduced by 44.3%, and the overall accuracy is also greatly improved. When the dimension is reduced to 128bit, rank1 can still be greater than 85%. At the same time, the mAP (mean average precision) indicator also has similar performance.
[0047]
[0048] Table 1
[0049] Step S203, the second perception object determines whether to perform feature conversion based on the sum of the number of vehicles included in the first field of view and the second field of view. If the sum of the number of vehicles is greater than the first threshold, the first perception object converts the first vector feature based on the hash algorithm of random hyperplane projection to obtain a first binary feature, and the second perception object converts the second vector feature based on the hash algorithm of random hyperplane projection to obtain a second binary feature;
[0050] The sum of the number of included vehicles represents the total number of vehicles included in the first field of view and the second field of view. For example, if the first field of view includes 10 vehicles and the second field of view includes 20 vehicles, the sum of the number of included vehicles is 30 vehicles.
[0051] The hash function based on random projection is defined as shown in formula (1): Randomly select a non-zero vector X in the n-dimensional feature space = [x1, x2, ..., x n ] is the normal vector of the hyperplane passing through the origin of the coordinate system. Since the n-dimensional space is divided into positive and negative spaces by the hyperplane, the space where the normal vector is located is the positive space. Let the hash values of the vectors in the positive and negative spaces in the set S be 1 and 0 respectively. The vector hash value can be calculated by judging whether the angle between the vector V and the normal vector X is greater than 90 degrees. Therefore, the definition formula of the hash function is as follows:
[0052]
[0053] Assume that the angle between vectors V and U is θ, and the probability that the two vectors are divided into the same space by the hyperplane determined by the randomly selected normal vector X is p(θ) = 1-θ÷π. Assume that the similarity between vectors V and U is s, according to θ = arccos(s), we can get p(s) = 1-arccos(s)÷π, and the similarity s and probability p are monotonically increasing. The hash function of formula (1) can only generate two-bit hash values. In order to increase the length of the hash value and improve the representation ability of the hash value, multiple independent hash functions are required. The hash values are the same only when two feature vectors have enough equal h(V) values. Therefore, the hash function H(V) is defined as shown in formula (2):
[0054] H(V)=(h b (V),h b-1 (V),...,h1(V))2 (2)
[0055] Where b is the binary number of the output; h i (V) is the ith independent h(V) function.
[0056] If the floating-point or half-precision floating-point features extracted by the re-identification network are transmitted directly, the amount of data will still be large, and the speed of feature transmission on both sides of the vehicle and road in the vehicle-road collaborative scenario will still be greatly affected. If the vehicle side and the road side need to transmit 20 perceived moving vehicle targets respectively in the vehicle-road collaborative environment, in the case of floating-point features, the feature size extracted by ShuffleBNLSH for each target is 2048×32=65536 bits, and the data volume of 20 vehicles will reach 1310720 bits. This amount of data is difficult to be satisfactory in time-sensitive communication scenarios such as vehicle-road collaboration that requires real-time interaction.
[0057] Therefore, in the vehicle-road cooperative scenario, when there are fewer vehicles and the vehicles are moving slowly, the demand for bandwidth and speed is not particularly tight, while during peak hours or when the speed is faster, the bandwidth and cooperative speed will become more critical. The bit features of vehicle-road cooperative perception will have a greater impact on the accuracy of re-identification when they are reduced to a certain length. Therefore, we designed a dynamic adaptive cooperative transmission method. When there are fewer vehicles in the cooperative field of view and the vehicle speed is slow, we do not perform bit conversion of features, and use half-precision floating point vehicle-road cooperative perception to extract features. When there are more vehicles, we make decisions on the bit conversion of vehicle re-identification features based on the total number of vehicles in the field of view (vehicle number information can be easily obtained when the vehicle and the road communicate). When the number of vehicles is greater than a certain number, we use shorter bit perception features for transmission and comparison. When the number of vehicles is medium, we use medium-length bit features for transmission and comparison. When there are fewer vehicles, we use longer bit features for transmission. This dynamic feature transmission method can better balance the relationship between vehicle-road re-identification accuracy and transmission bandwidth, making the communication channel between vehicles and roads more efficient. Figure 4 It is a schematic diagram of vehicle feature conversion in one embodiment of the present application.
[0058] In one embodiment, if the total number of vehicles is greater than the first threshold and less than the second threshold, the size of the converted first binary feature and the second binary feature is 1024 bits; the first threshold can be 10 vehicles, and the second threshold can be 20 vehicles.
[0059] In one embodiment, if the total number of vehicles is greater than the second threshold and less than the third threshold, the size of the converted first binary feature and the second binary feature is 512 bits; the second threshold can be 20 vehicles, and the third threshold can be 40 vehicles.
[0060] In one embodiment, if the total number of vehicles is greater than a third threshold, the size of the converted first binary feature and the second binary feature is 256 bits; the third threshold may be 40 vehicles.
[0061] In one embodiment, if the sum of the number of vehicles is less than a first threshold, the first vector feature and the second vector feature are not subjected to feature conversion. The first threshold may be 10 vehicles.
[0062] After processing by a hash function based on the random projection method, the amount of feature data for each vehicle is preset to between 128 bits and 1024 bits. The amount of floating-point feature data before compression can be reduced to 1 / 128 to 1 / 32. Compared with the amount of data before compression, the communication pressure of vehicle-road cooperative perception data can be greatly reduced.
[0063] Step S204, the first sensing object sends the first binary feature and the three-dimensional information of at least one vehicle in the first field of view to the second sensing object, and the second sensing object matches the first binary feature and the second binary feature to obtain a matching result, wherein the matching result includes at least one pair of successfully matched vehicles;
[0064] The speed of feature matching is also one of the keys to vehicle-road collaboration. Euclidean distance is a commonly used feature distance measurement method in re-identification. Assume that there are two points x(x1, x2, ..., x n ) and y(y1,y2,...,y n ), then the Euclidean distance d(x,y) between x and y is calculated as shown in formula (3). The vehicle-road collaboration process requires frequent calculation of the similarity of the target features between the vehicle and the road. However, the Euclidean distance of high-dimensional features requires frequent multiplication and square root calculations, and this calculation overhead is not small. In order to speed up the feature matching speed, we abandon the commonly used method based on Euclidean distance for feature matching and use the method based on Hamming distance as shown in formula (4) to complete the matching. The Hamming distance can be used to calculate the similarity of the vehicles contained in the first binary feature and the second binary feature, and obtain at least one pair of successfully matched vehicles above the similarity threshold.
[0065]
[0066]
[0067] In formula (4), x and y are binary data, and d h represents the Hamming distance, Represents XOR calculation. The computational complexity of the Hamming distance is much lower than that of the Euclidean distance, which can reduce the feature matching calculation time and further improve the efficiency of vehicle-road cooperative perception fusion.
[0068] Using the Hamming distance to calculate the similarity between the first binary feature and the second binary feature will save time in calculating the Euclidean distance compared with the half-precision floating-point features. For example, in the VeRi776 dataset, the binary features and half-precision floating-point features of 1678 query images and 9901 images to be matched proposed by the feature extraction network were compared and tested. The test results show that compared with the common 2048-dimensional half-precision floating-point features for Euclidean distance calculation, the Hamming distance can be accelerated by 72.32%; compared with the common 512-dimensional half-precision floating-point features for Euclidean distance calculation, the Hamming distance can be accelerated by 13.94%.
[0069] Step S205, the second sensing object calculates the coordinate transformation matrix based on the ICP algorithm for the three-dimensional information corresponding to at least one pair of successfully matched vehicles;
[0070] In one embodiment, one or more pairs can be randomly selected from at least one pair of successfully matched vehicles, and a coordinate transformation matrix can be calculated; for example, if a pair of vehicles is selected, the coordinate transformation matrix is calculated based on the pair of vehicles; if multiple pairs of vehicles are selected, at least one initial transformation matrix is calculated based on the multiple pairs of vehicles; if there is a transformation matrix with numerical abnormalities in at least one initial transformation matrix, the transformation matrix with numerical abnormalities is eliminated, and finally a transformation matrix is randomly selected from the at least one initial transformation matrix after elimination as the coordinate transformation matrix; if there is no transformation matrix with numerical abnormalities in at least one initial transformation matrix, a transformation matrix is randomly selected as the coordinate transformation matrix.
[0071] The three-dimensional information corresponding to at least one pair of successfully matched vehicles indicates that the three-dimensional information of a vehicle in the first field of view has been matched with the three-dimensional information of a vehicle in the second field of view. For example, the first field of view includes vehicle A, and the eight vertices of a group of three-dimensional bounding boxes detected by vehicle A in the first field of view are A, B, C, D, E, F, G, and H. The second field of view includes vehicle A, and the eight vertices of a group of three-dimensional bounding boxes detected by vehicle A in the second field of view are 1, 2, 3, 4, 5, 6, 7, and 8. The eight vertex information of vehicle A in the first field of view corresponds one-to-one with the eight vertex information of vehicle A in the second field of view, so that the eight pairs of vertex information corresponding to the vehicle in the first field of view and the second field of view are input into the ICP algorithm to calculate the coordinate transformation matrix. As an example, the process of calculating the coordinate transformation matrix after the points in the eight pairs of vertices are matched one-to-one, and the two groups of vertices can be regarded as two groups of point clouds, including: (1) solving the center of mass of the vehicle in the two groups of point clouds: (2) finding the coordinates of the two groups of point clouds without the center of mass: (3) finding the average of each point matrix and solving to obtain the coordinate transformation matrix. The schematic diagram of using the ICP algorithm for feature fusion when the point clouds do not match is as follows: Figure 5 (The yellow point cloud is the point cloud collected from the vehicle side, and the blue point cloud is the point cloud collected from the road side). This application calculates the matching degree of features through Hamming distance, and then performs coordinate transformation based on the ICP algorithm, and then obtains the fusion of the three-dimensional information of the first field of view and the second field of view. The schematic diagram after fusion is as follows: Figure 6 (The yellow point cloud is the point cloud collected from the vehicle side, and the blue point cloud is the point cloud collected from the road side).
[0072] Step S206: The second sensing object fuses the three-dimensional information of the first field of view and the three-dimensional information of the second field of view based on the coordinate transformation matrix to obtain fused three-dimensional information.
[0073] In one embodiment, the three-dimensional bounding box of each target converted by the coordinate transformation matrix is subjected to an IoU calculation. If there is any overlapping IoU in the converted three-dimensional bounding boxes, one of the bounding boxes with an IoU exceeding 30% (the IoU is not limited and can take a value from 0% to 100%) is removed (the removal can be performed in different ways, the ways are not limited, such as based on the confidence level and preference of the positions on both sides of the road) to obtain the final fusion result.
[0074] The fused three-dimensional information can be displayed on the first perception object or the second perception object.
[0075] In one embodiment, the fusion effect diagram after vehicle matching is as follows: Figure 7 (The yellow bounding box is the target detected on the car side, and the blue bounding box is the target detected on the road side).
[0076] In one embodiment, Figure 8 This is a schematic diagram of feature matching and feature fusion shown in steps S204, S205, and S206 provided in an embodiment of the present application.
[0077] In one embodiment, the three-dimensional information fusion effect diagram on both sides of the road is as follows: Fig. 9 (The yellow bounding box is the target detected on the car side, and the blue bounding box is the target detected on the road side).
[0078] This fusion method does not require GPS positioning information access. It only requires a small number of common vehicle targets on both sides of the road to enable ICP to adaptively select better source point clouds and target point clouds. It can not only speed up the fusion of target information perceived on both sides of the road, but also achieve the goal of reducing vehicle-side blind spots and helping to improve traffic safety.
[0079] Based on the same inventive concept, this embodiment also provides an autonomous driving vehicle-road collaborative perception method. The explanation of the method can be referred to the above and will not be repeated here. The method includes:
[0080] Acquire image information and three-dimensional information of at least one vehicle in a second field of view, wherein the image information and three-dimensional information of the at least one vehicle in the second field of view and the image information and three-dimensional information of the at least one vehicle in the first field of view include the same vehicle, and the image information and three-dimensional information of the at least one vehicle in the first field of view are acquired by the first sensing object;
[0081] Extracting features from image information of at least one vehicle in the second field of view based on the vehicle re-identification network to obtain a second vector feature;
[0082] Determine whether to perform feature conversion based on the sum of the number of vehicles included in the first field of view and the second field of view, and if the sum of the number of vehicles is greater than the first threshold, convert the second vector feature based on the hash algorithm of random hyperplane projection to obtain a second binary feature;
[0083] Receiving a first binary feature sent by a first sensing object and three-dimensional information of at least one vehicle in a first field of view, and then matching the first binary feature with the second binary feature to obtain a matching result, wherein the matching result includes at least one pair of successfully matched vehicles;
[0084] Based on the ICP algorithm, the coordinate transformation matrix is calculated for the three-dimensional information corresponding to at least one pair of successfully matched vehicles;
[0085] The three-dimensional information of the first field of view and the three-dimensional information of the second field of view are fused based on the coordinate transformation matrix to obtain fused three-dimensional information.
[0086] Based on the same inventive concept, the embodiment of the present application provides an electronic device, such as Fig.10 As shown, including:
[0087] Processor 41; memory 42 for storing executable instructions of processor 41; wherein processor 41 is configured to execute to implement an autonomous driving vehicle-road collaborative perception method as provided above: a first perception object obtains image information and three-dimensional information of at least one vehicle in a first field of view, and a second perception object obtains image information and three-dimensional information of at least one vehicle in a second field of view, and the image information and three-dimensional information of at least one vehicle in the first field of view and the image information and three-dimensional information of at least one vehicle in the second field of view include the same vehicle; the first perception object extracts features from the image information of at least one vehicle in the first field of view based on a vehicle re-identification network to obtain a first vector feature, and the second perception object extracts features from the image information of at least one vehicle in the second field of view based on a vehicle re-identification network to obtain a second vector feature; the second perception object extracts features from the image information of at least one vehicle in the second field of view based on the vehicle re-identification network to obtain a second vector feature; the second perception object extracts features from the first field of view and the second field of view based on the number of vehicles contained in the first field of view and the second field of view The sum of determines whether to perform feature conversion. If the sum of the number of vehicles is greater than the first threshold, the first perception object converts the first vector feature based on the hash algorithm of random hyperplane projection to obtain the first binary feature, and the second perception object converts the second vector feature based on the hash algorithm of random hyperplane projection to obtain the second binary feature; the first perception object sends the first binary feature and the three-dimensional information of at least one vehicle in the first field of view to the second perception object, and the second perception object matches the first binary feature and the second binary feature to obtain a matching result, which includes at least one pair of successfully matched vehicles; the second perception object calculates the three-dimensional information corresponding to at least one pair of successfully matched vehicles based on the ICP algorithm to obtain a coordinate transformation matrix; the second perception object fuses the three-dimensional information of the first field of view and the three-dimensional information of the second field of view based on the coordinate transformation matrix to obtain fused three-dimensional information. Based on the same inventive concept, this embodiment provides a non-temporary computer-readable storage medium. When the instructions in the storage medium are executed by the processor 41 of the electronic device, the electronic device is able to execute and implement an autonomous driving vehicle-road collaborative perception method as provided above.
[0088] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only for example and does not constitute a limitation of this specification. Although not explicitly stated here, those skilled in the art may make various modifications, improvements and corrections to this specification. Such modifications, improvements and corrections are suggested in this specification, so such modifications, improvements and corrections still belong to the spirit and scope of the exemplary embodiments of this specification.
[0089] At the same time, this specification uses specific words to describe the embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" refer to a certain feature, structure or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more in different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures or characteristics in one or more embodiments of this specification can be appropriately combined.
[0090] In addition, unless explicitly stated in the claims, the order of processing elements and sequences, the use of alphanumeric characters, or the use of other names in this specification are not intended to limit the order of the processes and methods of this specification. Although the above disclosure discusses some invention embodiments that are currently considered useful through various examples, it should be understood that such details are only for illustrative purposes, and the attached claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the essence and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.
[0091] Similarly, it should be noted that in order to simplify the description disclosed in this specification and thus help understand one or more embodiments of the invention, in the above description of the embodiments of this specification, multiple features are sometimes combined into one embodiment, figure or description thereof. However, this disclosure method does not mean that the features required by the subject matter of this specification are more than the features mentioned in the claims. In fact, the features of the embodiments are less than all the features of the single embodiment disclosed above.
[0092] In some embodiments, numbers describing the number of components and attributes are used. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise specified, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may change according to the required features of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of this specification are approximate values, in specific embodiments, the setting of such numerical values is as accurate as possible within the feasible range.
[0093] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, documents, etc., cited in this specification is hereby incorporated by reference in its entirety. Except for application history documents that are inconsistent with or conflicting with the content of this specification, documents that limit the broadest scope of the claims of this specification (currently or later attached to this specification) are also excluded. It should be noted that if the descriptions, definitions, and / or use of terms in the materials attached to this specification are inconsistent or conflicting with the content of this specification, the descriptions, definitions, and / or use of terms in this specification shall prevail.
[0094] Finally, it should be understood that the embodiments in this specification are only used to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, as an example and not a limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly introduced and described in this specification.
Claims
1. A method for autonomous driving vehicle-road cooperative perception, characterized in that: include: The first sensing object acquires image information and three-dimensional information of at least one vehicle in a first field of view, and the second sensing object acquires image information and three-dimensional information of at least one vehicle in a second field of view, wherein the image information and three-dimensional information of the at least one vehicle in the first field of view and the image information and three-dimensional information of the at least one vehicle in the second field of view include the same vehicle; The first sensing object extracts features from the image information of at least one vehicle in the first field of view based on the vehicle re-identification network to obtain a first vector feature, and the second sensing object extracts features from the image information of at least one vehicle in the second field of view based on the vehicle re-identification network to obtain a second vector feature; The second perception object determines whether to perform feature conversion based on the sum of the number of vehicles included in the first field of view and the second field of view. If the sum of the number of vehicles is greater than a first threshold, the first perception object converts the first vector feature based on a hash algorithm of random hyperplane projection to obtain a first binary feature, and the second perception object converts the second vector feature based on a hash algorithm of random hyperplane projection to obtain a second binary feature; The first sensing object sends the first binary feature and the three-dimensional information of at least one vehicle in the first field of view to the second sensing object, and the second sensing object matches the first binary feature and the second binary feature to obtain a matching result, wherein the matching result includes at least one pair of successfully matched vehicles; The second sensing object calculates the coordinate transformation matrix based on the ICP algorithm for the three-dimensional information corresponding to the at least one pair of successfully matched vehicles; The second sensing object fuses the three-dimensional information of the first field of view and the three-dimensional information of the second field of view based on the coordinate transformation matrix to obtain fused three-dimensional information; Among them, the coordinate transformation matrix is calculated based on the ICP algorithm for the three-dimensional information corresponding to the at least one pair of successfully matched vehicles, including: randomly selecting one or more pairs from the at least one pair of successfully matched vehicles, and calculating the coordinate transformation matrix; the random selection of one or more pairs from the at least one pair of successfully matched vehicles, and calculating at least one initial transformation matrix, including: if a pair of vehicles is selected, the coordinate transformation matrix is calculated based on the pair of vehicles; if multiple pairs of vehicles are selected, at least one initial transformation matrix is calculated based on the multiple pairs of vehicles; if there is a transformation matrix with numerical abnormalities in the at least one initial transformation matrix, the transformation matrix with numerical abnormalities is eliminated, and finally a transformation matrix is randomly selected from the at least one initial transformation matrix after elimination as the coordinate transformation matrix; if there is no transformation matrix with numerical abnormalities in the at least one initial transformation matrix, a transformation matrix is randomly selected as the coordinate transformation matrix.
2. The autonomous driving vehicle-road collaborative perception method according to claim 1, characterized in that: The model of the vehicle re-identification network consists of a Conv1 layer, a MaxPool layer, a Stage2 layer, a Stage3 layer, a Stage4 layer, a Conv5 layer, an AdaptiveAvgPool2d layer, a BNNeck layer, and a fully connected layer; the activation function of the model of the vehicle re-identification network is SiLU, and half-precision floating-point parameters are used for training.
3. The autonomous driving vehicle-road collaborative perception method according to claim 1, characterized in that: The loss function of the vehicle re-identification network model during training is: L Final , L Final = λ 1 L CE + λ 2 L T + λ 3 L A ,in λ 1. λ 2. λ 3 is the weighting coefficient, L CE is the cross entropy loss function, L T is the triple loss function, L A is the ArcFace loss function; ,in y i is the predicted classification value, is the actual classification value; L T = ( d a,p - d a,n + m ),in m is the margin parameter; d a,p is the feature distance between positive sample pairs; d a,n is the feature distance between negative sample pairs; ,in N is the batch size, n is the number of categories, s is the scaling factor, m is the angle interval parameter, θ j for W j and x i The angle between W j is the weight of the fully connected layer, x i is the feature of the deep learning network input fully connected layer, y i Represents the features during training x i The real category, θ yi for W yi and x i The angle between W yi For prediction y i The target weight for the category.
4. The autonomous driving vehicle-road collaborative perception method according to claim 1, characterized in that: Also includes: If the sum of the number of vehicles is less than the first threshold, no feature conversion is performed on the first vector feature and the second vector feature.
5. The autonomous driving vehicle-road collaborative perception method according to claim 1, characterized in that: If the sum of the number of vehicles is greater than a first threshold, the first vector feature is converted to obtain a first binary feature based on a hash algorithm of random hyperplane projection, and the second vector feature is converted to obtain a second binary feature based on a hash algorithm of random hyperplane projection, including: If the sum of the number of vehicles is greater than the first threshold value and less than the second threshold value, the size of the converted first binary feature and the second binary feature is 1024 bits; If the sum of the number of vehicles is greater than the second threshold value and less than the third threshold value, the size of the converted first binary feature and the second binary feature is 512 bits; If the sum of the number of vehicles is greater than the third threshold, the sizes of the converted first binary feature and the second binary feature are 256 bits.
6. The autonomous driving vehicle-road collaborative perception method according to claim 1, characterized in that: The matching result obtained by matching the first binary feature and the second binary feature includes at least one pair of successfully matched vehicles, including: using the Hamming distance to calculate the similarity of the vehicles included in the first binary feature and the second binary feature, and obtaining at least one pair of successfully matched vehicles above a similarity threshold.
7. An electronic device, characterized in that: include: Memory; Processor; And a computer program; wherein the computer program is stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps corresponding to the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Fast hash vehicle retrieval method based on multi-task deep learning
CN107885764A
Multi-vehicle cooperatively rapid mapping method
CN109100730A