Obstacle recognition method and system for multi-unit vehicle, device, product, and medium
By deploying cameras on each carriage of a multi-car articulated vehicle and processing and fusing bird's-eye view features, the problem of multi-car articulated vehicles having difficulty accurately identifying obstacles has been solved, achieving low-cost, high-reliability obstacle perception and autonomous driving support.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2026-04-02
AI Technical Summary
Multi-unit articulated vehicles struggle to accurately identify obstacles around them in a low-cost manner, and existing lidar solutions are expensive and have short lifespans.
Multiple cameras are deployed on each carriage. The sub-circular view images are processed by a single-carriage bird's-eye view converter, and the bird's-eye view features of each carriage are fused to generate multi-carriage fused features to identify obstacle information, including 3D coordinates and categories.
It reduces the cost of obstacle perception, improves the reliability and comprehensiveness of obstacle perception, reduces blind spots, enhances environmental perception capabilities, and supports autonomous driving decision-making.
Smart Images

Figure CN2024130118_02042026_PF_FP_ABST
Abstract
Description
Obstacle identification method, system, device, product and medium for multi-unit vehicle
[0001] The present application claims priority to the Chinese patent application No. 202411352104.9, filed on September 26, 2024, and titled "Obstacle identification method, system, device, product and medium for multi-unit vehicle", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of transportation, in particular to an obstacle identification method, system, device, product and medium for multi-unit vehicle. BACKGROUND
[0003] Multi-unit articulated vehicle is a new type of urban passenger bus, which moves through rubber wheels without fixed tracks. Compared with ordinary buses, it can transport more passengers through articulated discs connecting multiple carriages, relieving the pressure of the public transportation system. Since the multi-unit articulated vehicle uses articulated discs to connect multiple carriages, the vehicle body is relatively long, and it is generally difficult for the driver to detect obstacles around the vehicle body with the naked eye. The current multi-unit articulated vehicle mainly uses laser radar for obstacle perception, but laser radar has problems such as high price and short service life.
[0004] Therefore, how to provide a solution to the above technical problems is a problem that those skilled in the art need to solve at present.
[0005] SUMMARY
[0006] The purpose of the present application is to provide an obstacle identification method, system, device, product and medium for multi-unit vehicle, which can reduce the cost of obstacle perception and improve the reliability and comprehensiveness of obstacle perception.
[0007] In one aspect, the present application provides an obstacle identification method for a multi-unit vehicle, the multi-unit vehicle comprising multiple articulated carriages, each carriage being provided with multiple cameras for capturing sub-panoramic images in its own capture area, the obstacle identification method comprising:
[0008] For each carriage, the sub-panoramic images captured by the multiple cameras on the carriage are obtained, two-dimensional image features are extracted from all the sub-panoramic images, and the camera parameters, bird's eye view encoding and two-dimensional image features corresponding to the cameras are processed by a single-carriage bird's eye view converter to obtain the current bird's eye view features of the carriage;
[0009] The current bird's eye view features of multiple carriages are fused to obtain multi-carriage fusion features;
[0010] Obtain the obstacle information of the multi-vehicle by using the multi-carriage fusion feature;
[0011] The processing includes:
[0012] Calculate the first key vector and the first value vector by using the camera parameters and the first two-dimensional image feature;
[0013] Obtain the first to-be-processed bird's eye view feature based on the bird's eye view encoding, and obtain the second to-be-processed bird's eye view feature by using the first key vector, the first value vector and the first to-be-processed bird's eye view feature;
[0014] Calculate the second key vector and the second value vector by using the camera parameters and the second two-dimensional image feature; the feature down-sampling rate corresponding to the second two-dimensional image feature is greater than the feature down-sampling rate corresponding to the first two-dimensional image feature;
[0015] Process the second key vector, the second value vector and the second to-be-processed bird's eye view feature to obtain the second intermediate feature;
[0016] Process the second intermediate feature to obtain the current bird's eye view feature of the carriage.
[0017] Optionally, the single-carriage bird's eye view converter includes a first fusion attention module, a first convolution module, a first full connection network module, a second full connection network module, a second fusion attention module and a second convolution module;
[0018] Calculate the first key vector and the first value vector by using the first full connection network module and the camera parameters and the first two-dimensional image feature;
[0019] Process the first key vector, the first value vector and the first to-be-processed bird's eye view feature by using the first fusion attention module to obtain the first intermediate feature;
[0020] Process the first intermediate feature by using the first convolution module to obtain the second to-be-processed bird's eye view feature;
[0021] Calculate the second key vector and the second value vector by using the second full connection network module and the camera parameters and the second two-dimensional image feature; the feature down-sampling rate corresponding to the second two-dimensional image feature is greater than the feature down-sampling rate corresponding to the first two-dimensional image feature;
[0022] Process the second key vector, the second value vector and the second to-be-processed bird's eye view feature by using the second fusion attention module to obtain the second intermediate feature;
[0023] The second intermediate feature is processed by the second convolution module to obtain a current bird's eye view feature of the car.
[0024] Optionally, the first fusion attention module and the second fusion attention module each include a first deformable attention unit and a second deformable attention unit.
[0025] Further comprising:
[0026] The first to-be-processed bird's eye view feature is divided into a grid of P*P size, and features in each of the grids are extracted as inputs of the first deformable attention unit.
[0027] The grid of P*P size is divided into a grid of G*G size, and features in multiple related grids are extracted as inputs of the second deformable attention unit.
[0028] P and G are each an integer greater than 1.
[0029] Optionally, a process of fusing current bird's eye view features of multiple cars to obtain multi-car fusion features includes:
[0030] An articulation chain angle between every two adjacent cars is obtained.
[0031] Based on all the articulation chain angles and all the current bird's eye view features, splicing features of the multiple cars are obtained.
[0032] The splicing features are fused to obtain multi-car fusion features.
[0033] Optionally, a process of obtaining splicing features of the multiple cars based on all the articulation chain angles and all the current bird's eye view features includes:
[0034] According to an articulation chain angle between an i-th car and an i-1-th car, a current bird's eye view feature of the i-th car is converted into a bird's eye view space coordinate system of a car head car to obtain a to-be-spliced feature of the i-th car; i = 2, 3…, n, n is a total number of cars of the multiple-unit vehicle.
[0035] The current bird's eye view feature of the first car is determined as the to-be-spliced feature of the first car.
[0036] The to-be-spliced features of the multiple cars are spliced to obtain splicing features.
[0037] Optionally, a process of fusing the splicing features to obtain multi-car fusion features includes:
[0038] The splicing features are input into a bird's eye view fusion converter to obtain multi-car fusion features.
[0039] Optionally, the process of obtaining the obstacle information of the multi-vehicle combination vehicle by using the multi-vehicle combination fusion feature comprises:
[0040] inputting the multi-vehicle combination fusion feature into an aerial view decoder to obtain the obstacle information of the multi-vehicle combination vehicle, wherein the obstacle information comprises three-dimensional coordinates of the obstacle in the whole-vehicle coordinate system of the multi-vehicle combination vehicle, length, width, height and heading angle of the obstacle.
[0041] Optionally, the obstacle information further comprises a category of the obstacle.
[0042] Optionally, the process further comprises:
[0043] automatically controlling the multi-vehicle combination vehicle based on the obstacle information.
[0044] In another aspect, the present application further provides a multi-vehicle combination obstacle identification system, wherein the multi-vehicle combination vehicle comprises multiple articulated vehicle compartments, each vehicle compartment is provided with multiple cameras, the cameras are used to collect sub-surround view images in their own collection areas, and the multi-vehicle combination obstacle identification system comprises:
[0045] a processing module, configured to, for each vehicle compartment, acquire the sub-surround view images collected by the multiple cameras on the vehicle compartment, extract two-dimensional image features from all the sub-surround view images, and process camera parameters corresponding to the cameras, aerial view encoding and the two-dimensional image features by using a single-vehicle compartment aerial view converter to obtain current aerial view features of the vehicle compartment;
[0046] a fusion module, configured to fuse the current aerial view features of multiple vehicle compartments to obtain multi-vehicle combination fusion features;
[0047] a decoding module, configured to obtain obstacle information of the multi-vehicle combination vehicle by using the multi-vehicle combination fusion features;
[0048] the processing comprises:
[0049] calculating the camera parameters and the first two-dimensional image features to obtain a first key vector and a first value vector;
[0050] obtaining a first to-be-processed aerial view feature based on the aerial view encoding, and obtaining a second to-be-processed aerial view feature based on the first key vector, the first value vector and the first to-be-processed aerial view feature;
[0051] calculating the camera parameters and the second two-dimensional image features to obtain a second key vector and a second value vector; the feature down-sampling rate corresponding to the second two-dimensional image features is greater than the feature down-sampling rate corresponding to the first two-dimensional image features;
[0052] processing the second key vector, the second value vector and the second to-be-processed bird's eye view feature to obtain a second intermediate feature;
[0053] processing the second intermediate feature to obtain the current bird's eye view feature of the car.
[0054] In another aspect, the present application also provides a computer program product comprising computer programs / instructions which, when executed by a processor, implement the steps of the obstacle identification method of the multi-formation vehicle according to any one of the above.
[0055] In another aspect, the present application also provides an electronic device comprising:
[0056] a memory for storing computer programs;
[0057] a processor for implementing the steps of the obstacle identification method of the multi-formation vehicle according to any one of the above when executing the computer programs.
[0058] In another aspect, the present application also provides a computer readable storage medium having computer programs stored thereon, the computer programs being executed by a processor to implement the steps of the obstacle identification method of the multi-formation vehicle according to any one of the above.
[0059] The present application provides an obstacle identification method of a multi-formation vehicle, which realizes three-dimensional perception of obstacles around the multi-formation articulated vehicle through a camera arranged on the multi-formation vehicle. Compared with a laser radar, the camera has lower cost and longer service life. The current bird's eye view feature of each car is obtained based on the sub-surround view image collected by the camera of the car, the current bird's eye view features of the multi-section cars are fused to obtain a multi-car fusion feature, the non-overlapping fields of view from different cameras are fused into a complete scene representation, the blind area of a single camera is filled, and subsequent global understanding and analysis are facilitated, so that the obstacle information around the vehicle can be accurately identified, and the reliability and comprehensiveness of obstacle perception are improved. The present application also provides an obstacle identification system of a multi-formation vehicle, an electronic device, a computer program product and a computer readable storage medium, which have the same beneficial effects as the obstacle identification method of the multi-formation vehicle. BRIEF DESCRIPTION OF DRAWINGS
[0060] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0061] Fig. 1 is a flow chart of a multi-formation vehicle obstacle identification method provided by the present application;
[0062] Fig. 2 is a camera layout schematic diagram provided by the present application;
[0063] Fig. 3 is a schematic diagram of a single-carriage bird's-eye view converter provided by the present application;
[0064] Fig. 4 is a schematic diagram of a fusion attention module provided by the present application;
[0065] Fig. 5 is a local extraction schematic diagram of a first deformable attention unit provided by the present application;
[0066] Fig. 6 is a global extraction schematic diagram of a second deformable attention unit provided by the present application;
[0067] Fig. 7 is a schematic diagram of a bird's-eye view fusion converter provided by the present application;
[0068] Fig. 8 is a multi-formation vehicle obstacle identification flow schematic diagram provided by the present application;
[0069] Fig. 9 is a structural schematic diagram of a multi-formation vehicle obstacle identification system provided by the present application;
[0070] Fig. 10 is a structural schematic diagram of an electronic device provided by the present application;
[0071] Fig. 11 is a structural schematic diagram of a computer-readable storage medium provided by the present application. DETAILED DESCRIPTION
[0072] The core of the present application is to provide a multi-formation vehicle obstacle identification method, system, device, product and medium, which can reduce obstacle perception cost and improve the reliability and comprehensiveness of obstacle perception.
[0073] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0074] In a first aspect, referring to Fig. 1, the present application provides a multi-formation vehicle obstacle identification method, the multi-formation vehicle comprising multiple articulated carriages, each carriage being provided with multiple cameras, the cameras being used to collect sub-surround view images in their own collection areas, and the multi-formation vehicle obstacle identification method comprising:
[0075] S101: For each carriage, obtain the sub-surround view images collected by the plurality of cameras on the carriage, extract two-dimensional image features from all sub-surround view images, process the camera parameters, bird's eye view encoding and two-dimensional image features corresponding to the cameras by a single-carriage bird's eye view converter to obtain the current bird's eye view features of the carriage;
[0076] The processing includes:
[0077] The camera parameters and the first two-dimensional image features are calculated to obtain a first key vector and a first value vector;
[0078] The first key vector, the first value vector and the first to-be-processed bird's eye view features are obtained based on the bird's eye view encoding to obtain a second to-be-processed bird's eye view feature;
[0079] The camera parameters and the second two-dimensional image features are calculated to obtain a second key vector and a second value vector; the feature down-sampling rate corresponding to the second two-dimensional image features is greater than the feature down-sampling rate corresponding to the first two-dimensional image features;
[0080] The second key vector, the second value vector and the second to-be-processed bird's eye view feature are processed to obtain a second intermediate feature;
[0081] The second intermediate feature is processed to obtain the current bird's eye view features of the carriage.
[0082] In this embodiment, the multi-formation vehicle can be a multi-formation articulated vehicle, i.e., a vehicle formed by articulating multiple carriages through an articulation device. In order to reduce the cost of obstacle perception, this embodiment uses multiple cameras arranged on the multi-formation vehicle to collect sub-surround view images around the vehicle body, so as to subsequently obtain a complete surround view image of the multi-formation vehicle based on the sub-surround view images collected by the multiple cameras. As an optional embodiment, at least one camera can be arranged on the left side and the right side of the carriage respectively to collect sub-surround view images of the side of the carriage, one camera is arranged on the front side of the head carriage to collect sub-surround view images of the front view angle, and one camera is arranged on the rear side of the tail carriage to collect sub-surround view images of the rear view angle. Referring to FIG. 2, the multi-formation articulated vehicle includes three carriages, i.e., a head carriage, a tail carriage and a middle carriage. The head carriage and the middle carriage are connected by a hinge, and the tail carriage and the middle carriage are connected by a hinge. Five cameras are arranged on the head carriage and the tail carriage respectively, and two cameras are arranged on the left side and the right side of the middle carriage respectively to compensate for the blind area of a single camera.
[0083] Each camera can collect sub-surround view images of its corresponding collection area at a preset sampling frequency. The sampling frequency can be set according to actual engineering needs, which is not limited in this embodiment.
[0084] In order to make the position and size of the obstacle in space more easily and accurately described, and make the spatial relationship between the obstacle and the multi-vehicle more intuitive, the embodiment converts the image captured by the camera from the original view (such as front view, side view) to the bird's eye view feature, so as to improve the accuracy of obstacle detection, facilitate more accurate positioning of the position of the obstacle, and facilitate path planning and obstacle avoidance in the automatic driving system. It can be understood that a plurality of cameras are arranged on each car, and a plurality of sub-loop view images captured by the cameras can be obtained at the same time. The two-dimensional image features are extracted from the sub-loop view images captured by the plurality of cameras. Feature extraction helps to remove noise and irrelevant information in the original sub-loop view image, better aligns the features under different perspectives, ensures the consistency of spatial information, and thus improves the accuracy of obstacle detection. In addition, considering that directly performing BEV (Bird's Eye View) conversion on the original image may involve a large number of geometric transformations and image processing operations, extracting two-dimensional image features and performing BEV conversion on the two-dimensional image features can reduce the amount of calculation.
[0085] Specifically, in the embodiment, the single-carriage feature extractor can be used to extract features from the sub-loop view images captured by the plurality of cameras of the single carriage. The input of the single-carriage feature extractor is the sub-loop view images captured by the plurality of cameras of the single carriage, which is assumed to be N×H×W×C, where N is the number of cameras on the carriage, H is the height of the sub-loop view image, W is the width of the sub-loop view image, and C is the channel of the sub-loop view image. The output of the single-carriage feature extractor is two-dimensional image features, and the feature size is N×H'×W'×C', where N is the number of cameras on the carriage, H' is the height of the two-dimensional image features, W' is the width of the two-dimensional image features, and C' is the channel of the two-dimensional image features. ResNet (Residual Network), EfficientNet (a kind of convolutional neural network), VoVNet (Volumetric Overlapped Voting Network) and the like can be used to extract two-dimensional image features of the sub-loop view image.
[0086] In an exemplary embodiment, the process of converting the two-dimensional image features into the current bird's eye view features of the carriage includes:
[0087] Obtaining camera parameters corresponding to the camera;
[0088] Converting the two-dimensional image features into the current bird's eye view features of the carriage based on the camera parameters.
[0089] The camera parameters in the embodiment include intrinsic parameters and extrinsic parameters of the camera. The intrinsic parameters of the camera describe inherent properties of the camera itself, which are irrelevant to the external environment of the camera. These parameters include but are not limited to focal length, principal point coordinates, radial distortion coefficient, tangential distortion coefficient, image scale factor, etc. The extrinsic parameters of the camera describe the position and direction of the camera relative to a certain global coordinate system. These parameters include but are not limited to rotation parameters, translation parameters, etc. The camera parameters are used to project two-dimensional image features to the bird's eye view perspective, ensuring accurate mapping from two-dimensional image features to the current BEV features, so that the features maintain the correct spatial relationship and scale during the conversion. In addition, in a multi-camera perception system, as an optional embodiment, the intrinsic and extrinsic parameters of each camera need to be accurately calibrated in order to fuse image features captured by different cameras into a unified BEV representation.
[0090] In an exemplary embodiment, the process of converting two-dimensional image features into the current bird's eye view features of the vehicle cabin based on the camera parameters includes:
[0091] Inputting the camera parameters, the bird's eye view encoding and the two-dimensional image features into the single-vehicle-cabin bird's eye view converter corresponding to the vehicle cabin to obtain the current bird's eye view features of the vehicle cabin.
[0092] In the embodiment, an independent single-vehicle-cabin bird's eye view converter is configured for each vehicle cabin to realize the conversion of two-dimensional image features into current bird's eye view features, thereby improving the conversion efficiency.
[0093] The bird's eye view encoding represents the features in the single-vehicle-cabin bird's eye view (BEV) perspective, which is a learnable vector.
[0094] In an exemplary embodiment, referring to FIG. 3, the single-vehicle-cabin bird's eye view converter includes a first fusion attention module, a first convolution module, a first fully connected network module, a second fully connected network module, a second fusion attention module and a second convolution module.
[0095] The process of inputting the camera parameters, the bird's eye view encoding and the two-dimensional image features into the single-vehicle-cabin bird's eye view converter corresponding to the vehicle cabin to obtain the current bird's eye view features of the vehicle cabin includes:
[0096] The first fully connected network module is used to calculate the camera parameters and the first two-dimensional image features to obtain a first key vector and a first value vector;
[0097] The first to-be-processed bird's eye view features are obtained based on the bird's eye view encoding;
[0098] The first fusion attention module is used to process the first key vector, the first value vector and the first to-be-processed bird's eye view features to obtain a first intermediate feature;
[0099] The first intermediate feature is processed by using the first convolution module to obtain a second to-be-processed bird's eye view feature;
[0100] The camera parameters and the second two-dimensional image feature are calculated by using the second full connection network module to obtain a second key vector and a second value vector; the feature down-sampling rate corresponding to the second two-dimensional image feature is greater than the feature down-sampling rate corresponding to the first two-dimensional image feature;
[0101] The second key vector, the second value vector and the second to-be-processed bird's eye view feature are processed by using the second fusion attention module to obtain a second intermediate feature;
[0102] The second intermediate feature is processed by using the second convolution module to obtain the current bird's eye view feature of the carriage.
[0103] In this embodiment, the structure of the single-carriage bird's eye view converter is shown in FIG. 3, and the input is bird's eye view encoding, camera parameters and two-dimensional image features, and the output is the current bird's eye view feature of the single carriage. The purpose is to convert the two-dimensional image feature to the BEV perspective, and in FIG. 3, the bird's eye view encoding is used to calculate the correlation with the image feature. The key vector K and the value vector V are calculated by the two-dimensional image feature and the camera internal and external parameters through the full connection network, the key vector K and the value vector V are calculated by the bird's eye view encoding, and the three are input into the fusion attention module. In FIG. 3, the down-sampling rate of 4 times represents the feature down-sampling rate corresponding to the first two-dimensional image feature, the down-sampling rate of 8 times represents the feature down-sampling rate corresponding to the second two-dimensional image feature, and Q1, K1 and V1 represent the first to-be-processed bird's eye view feature, the first key vector and the first value vector in sequence, and Q2, K2 and V2 represent the second to-be-processed bird's eye view feature, the second key vector and the second value vector in sequence.
[0104] S102: The current bird's eye view features of the multi-section carriages are fused to obtain a multi-carriage fusion feature;
[0105] In this embodiment, after obtaining the current bird's eye view feature of the single carriage, the current bird's eye view features of the multi-section carriages are fused, which can provide more comprehensive coverage around the vehicle, reduce the blind area, enhance the environmental perception ability, and consider that the cameras of different carriages may have their own advantages and limitations due to different positions and perspectives. The current bird's eye view features of each carriage are fused to complement these differences. The multi-carriage fusion feature after fusion is represented in a unified coordinate system, which improves the robustness of perception and more accurately detects the surrounding obstacles. At the same time, it is helpful to realize the coordinated control between carriages and improve the stability and efficiency of driving.
[0106] S103: The obstacle information of the multi-formation vehicle is obtained by using the multi-carriage fusion feature.
[0107] In the embodiment, the multi-carriage fusion features can be used for global understanding and analysis to obtain the obstacle information of the multi-formation vehicle, and the obstacle information includes but is not limited to three-dimensional coordinates, length, width, height and heading angle of the obstacle in the vehicle coordinate system. Then, the obstacle information is sent out through the local area network in the multi-formation vehicle, so that the vehicle control unit in the multi-formation vehicle can avoid the obstacle according to the obstacle information, and automatic driving is realized.
[0108] As can be seen, in the embodiment, the three-dimensional perception of the obstacles around the multi-formation articulated vehicle is realized by the camera arranged on the multi-formation vehicle. Compared with the laser radar, the camera has lower cost and longer service life. The current bird's eye view features of each carriage are obtained based on the sub-surrounding view images collected by the camera of each carriage. The current bird's eye view features of the multiple carriages are fused to obtain the multi-carriage fusion features. The non-overlapping fields of view from different cameras are fused into a complete scene representation to fill in the blind area of a single camera. This is convenient for subsequent global understanding and analysis, so that the obstacle information around the vehicle can be accurately identified, and the reliability and comprehensiveness of obstacle perception are improved.
[0109] On the basis of the above embodiment:
[0110] In an exemplary embodiment, the first fusion attention module and the second fusion attention module each include a first deformable attention unit and a second deformable attention unit.
[0111] Further comprising:
[0112] The first to-be-processed bird's eye view features are divided into P*P size grids, and the features in each grid are extracted as the input of the first deformable attention unit.
[0113] The P*P size grid is divided into G*G number of grids, and the features in the multiple related grids are extracted as the input of the second deformable attention unit, and P and G are both integers greater than 1.
[0114] Referring to FIG. 4, FIG. 4 shows the calculation process of the l-1th layer vector z l-1 to the l+1th layer vector z l+1 of the fusion attention module (including the first fusion attention module and the second fusion attention module), and the fusion attention module is composed of layer normalization (LayerNorm), deformable attention unit and multilayer perceptron (MLP, Multilayer Perceptron). The first deformable attention unit and the second deformable attention unit focus on local and global bird's eye view features, respectively. The specific extraction positions are shown in FIG. 5 and FIG. 6.
[0115] Referring to FIG. 5, the first deformable attention unit local extraction divides the first to-be-processed bird's eye view feature into a P x P fixed-size grid, and extracts the features in each grid as the input of the first deformable attention unit. In this way, the features of the same grid position of different cameras in a single carriage can be effectively focused on, which is conducive to solving the occlusion problem of a single camera.
[0116] Referring to FIG. 6, the second deformable attention unit global extraction divides the P x P fixed grid into G x G grids, and extracts the features in multiple related grids as the input of the second deformable attention unit. As shown in FIG. 6, p11, p21, p31, p41, p51, p61, p71, and p81 are a group of related grids, and the remaining groups of related networks are similar. In this way, the neural network can focus on the global information of the current scene, reduce the amount of calculation, help the neural network understand the position distribution of each camera in a single carriage, and better fuse the information of multiple cameras.
[0117] The calculation method of Q, K, and V is as follows:
[0118] Where Q is the to-be-processed bird's eye view feature, K is the key vector, V is the value vector, softmax is the normalization exponential function, d k is the vector dimension.
[0119] In an exemplary embodiment, the process of fusing the current bird's eye view features of the multi-section carriages to obtain the multi-carriage fusion features includes:
[0120] Obtaining the articulation chain angle between every two adjacent carriages;
[0121] Obtaining the splicing features of the multi-section carriages based on all articulation chain angles and all current bird's eye view features;
[0122] Fusing the splicing features to obtain the multi-carriage fusion features;
[0123] Wherein, the process of obtaining the splicing features of the multi-section carriages based on all articulation chain angles and all current bird's eye view features includes:
[0124] According to the articulation chain angle between the i-th carriage and the i-1-th carriage, the current bird's eye view feature of the i-th carriage is converted to the bird's eye view space coordinate system of the car head carriage to obtain the to-be-spliced feature of the i-th carriage; i = 2, 3…, n, n is the total number of carriages of the multi-section vehicle;
[0125] The current bird's eye view feature of the first carriage is determined as the to-be-spliced feature of the first carriage;
[0126] Splicing the to-be-spliced features of the multi-section carriages to obtain the splicing features.
[0127] In this embodiment, the current bird's eye view feature of the current car body is converted to the bird's eye view space coordinate system of the front car body according to the hinge chain angle of the current car body and the previous car body, which can eliminate the coordinate deviation caused by the position and direction difference between different car bodies, ensure that the sub-surrounding view information of the whole vehicle is in the same reference system, and splice the converted to-be-spliced feature to obtain a spliced feature. The spliced bird's eye view feature can provide a complete panoramic view around the vehicle, thereby enhancing the perception ability of the vehicle to the surrounding environment, reducing the blind area, and improving the driving safety.
[0128] It can be understood that taking the front car body as the reference point can more intuitively assist the automatic driving system to make decisions such as steering, acceleration and braking.
[0129] In an example embodiment, the process of fusing the spliced feature to obtain the multi-carriage fusion feature includes:
[0130] The spliced feature is input into the bird's eye view fusion converter to obtain the multi-carriage fusion feature.
[0131] The structure of the bird's eye view fusion converter is shown in FIG. 7. The bird's eye view fusion converter is similar to the single-carriage bird's eye view converter and is also composed of layer normalization, a deformable attention unit and a multi-layer perception mechanism. The difference is that the input of the bird's eye view fusion converter is the spliced feature, and the query vector, key vector and value vector input into the local attention unit and global attention unit are obtained by different fully connected layers from the BEV feature. Since the field of view of each car body is not the same, focusing on the same local area of the BEV can make up for the possible field of view blind area of a single car body. In addition, the global attention unit can be used for global interaction through sparse sampling to obtain the context understanding of the map semantics, and finally obtain the multi-carriage fusion feature.
[0132] In an example embodiment, the process of obtaining obstacle information of a multi-formation vehicle using the multi-carriage fusion feature includes:
[0133] The multi-carriage fusion feature is input into the bird's eye view decoder to obtain the obstacle information of the multi-formation vehicle. The obstacle information includes the three-dimensional coordinates of the obstacle in the whole vehicle coordinate system of the multi-formation vehicle, the length, width and height of the obstacle, and the heading angle.
[0134] In an example embodiment, the obstacle information further includes the category of the obstacle.
[0135] After obtaining the multi-carriage fusion features, the multi-carriage fusion features are sent into a BEV decoder to obtain the surrounding obstacle information of the multi-formation vehicle. The BEV decoder is composed of two 1x1 convolutional layers, which are used for frame regression and classification respectively. The regression output is (x, y, z, w, d, h, theta), which respectively represents the 3D position, size (length, width and height) and yaw angle. The classification output is the confidence score of each anchor box as an object or background, which is used to distinguish the categories of obstacles. Finally, the obtained obstacle information is sent to the vehicle control unit of the multi-formation vehicle to help the multi-formation vehicle avoid obstacles and realize automatic driving. Referring to FIG. 8, FIG. 8 is a schematic diagram of an obstacle recognition process for the multi-formation vehicle shown in FIG. 2.
[0136] In summary, the present application provides a surround view perception scheme for a multi-formation articulated vehicle. The 3D perception of the surrounding obstacles of the multi-formation articulated vehicle is realized through a surround view camera system including multiple cameras. The perception of obstacles is realized through a low-cost camera scheme, which can effectively reduce the cost compared to the laser radar scheme. To meet the demand for 3D obstacle perception for automatic driving of the multi-formation articulated vehicle, a 3D obstacle perception scheme based on a transformer (a model architecture based on self-attention mechanism) network structure is provided. The images collected by the cameras of each carriage are used as input, and the 3D obstacle information around the vehicle is output to provide guidance for the planning and control module of the automatic driving, and to provide protection for the safe driving of the multi-formation articulated train. To meet the demand for multi-carriage fusion perception of the multi-formation articulated train, a multi-carriage fusion perception scheme is provided. The overall surround view features of the multi-formation articulated train and the surround view features of the individual carriages are focused on to complete the perception of the obstacles around the entire vehicle body. This perception scheme can effectively fuse multiple cameras and the perception range of the cameras of multiple carriages, fill in the blind area of a single camera, and better perceive the surrounding obstacle information of the multi-formation articulated vehicle by combining the global position information of multiple cameras.
[0137] In the second aspect, referring to FIG. 9, the present application further provides an obstacle recognition system for a multi-formation vehicle. The multi-formation vehicle includes multiple articulated carriages. Each carriage is provided with multiple cameras. The cameras are used to collect sub-surround view images in their own collection areas. The obstacle recognition system for the multi-formation vehicle includes:
[0138] The processing module 11 is configured to obtain, for each carriage, the sub-surround view images collected by the multiple cameras on the carriage, extract two-dimensional image features from all the sub-surround view images, and process the camera parameters, the bird's eye view encoding and the two-dimensional image features corresponding to the cameras by the single-carriage bird's eye view converter to obtain the current bird's eye view features of the carriage.
[0139] The fusion module 12 is configured to fuse the current bird's eye view features of the multiple carriages to obtain multi-carriage fusion features.
[0140] A decoding module 13 is configured to obtain obstacle information of the multi-vehicle by using the multi-carriage fusion feature;
[0141] The processing includes:
[0142] The camera parameters and the first two-dimensional image features are calculated to obtain a first key vector and a first value vector;
[0143] The first key vector, the first value vector, and the first to-be-processed aerial view feature are processed by using the first fusion attention module to obtain a first intermediate feature;
[0144] The camera parameters and the second two-dimensional image features are calculated to obtain a second key vector and a second value vector; the feature down-sampling rate corresponding to the second two-dimensional image features is greater than the feature down-sampling rate corresponding to the first two-dimensional image features;
[0145] The second key vector, the second value vector, and the second to-be-processed aerial view feature are processed to obtain a second intermediate feature;
[0146] The second intermediate feature is processed to obtain the current aerial view feature of the carriage.
[0147] In an exemplary embodiment, the single-carriage aerial view converter includes a first fusion attention module, a first convolution module, a first fully connected network module, a second fully connected network module, a second fusion attention module, and a second convolution module;
[0148] The camera parameters, the aerial view encoding, and the two-dimensional image features are input into the single-carriage aerial view converter corresponding to the carriage to obtain the current aerial view feature of the carriage, and the process includes:
[0149] The camera parameters and the first two-dimensional image features are calculated by using the first fully connected network module to obtain a first key vector and a first value vector;
[0150] The first to-be-processed aerial view feature is obtained based on the aerial view encoding;
[0151] The first key vector, the first value vector, and the first to-be-processed aerial view feature are processed by using the first fusion attention module to obtain a first intermediate feature;
[0152] The first intermediate feature is processed by using the first convolution module to obtain a second to-be-processed aerial view feature;
[0153] The camera parameters and the second two-dimensional image features are calculated by using the second fully connected network module to obtain a second key vector and a second value vector; the feature down-sampling rate corresponding to the second two-dimensional image features is greater than the feature down-sampling rate corresponding to the first two-dimensional image features;
[0154] The second key vector, the second value vector and the second to-be-processed bird's eye view feature are processed by using a second fusion attention module to obtain a second intermediate feature;
[0155] The second intermediate feature is processed by using a second convolution module to obtain the current bird's eye view feature of the carriage.
[0156] In an exemplary embodiment, the first fusion attention module and the second fusion attention module each include a first deformable attention unit and a second deformable attention unit;
[0157] Further comprising:
[0158] The first division module is configured to divide the first to-be-processed bird's eye view feature into a grid of P*P size, and extract the features in each grid as the input of the first deformable attention unit;
[0159] The second division module is configured to divide the grid of P*P size into a grid of G*G size, and extract the features in the multiple related grids as the input of the second deformable attention unit, P and G are both integers greater than 1.
[0160] In an exemplary embodiment, the process of fusing the current bird's eye view features of the multi-carriage to obtain the multi-carriage fusion feature includes:
[0161] Obtaining the articulation chain angle between every two adjacent carriages;
[0162] Obtaining the splicing feature of the multi-carriage based on all articulation chain angles and all current bird's eye view features;
[0163] Fusing the splicing feature to obtain the multi-carriage fusion feature.
[0164] In an exemplary embodiment, the process of obtaining the splicing feature of the multi-carriage based on all articulation chain angles and all current bird's eye view features includes:
[0165] According to the articulation chain angle between the ith carriage and the i-1th carriage, the current bird's eye view feature of the ith carriage is converted to the bird's eye view space coordinate system of the head carriage to obtain the to-be-spliced feature of the ith carriage; i=2, 3…, n, n is the total number of carriages of the multi-formation vehicle;
[0166] The current bird's eye view feature of the first carriage is determined as the to-be-spliced feature of the first carriage;
[0167] Splicing the to-be-spliced features of the multi-carriage to obtain the splicing feature.
[0168] In an exemplary embodiment, the process of fusing the splicing feature to obtain the multi-carriage fusion feature includes:
[0169] input the splicing feature into the bird's eye view fusion converter to obtain a multi-compartment fusion feature.
[0170] In an example embodiment, a process of obtaining obstacle information of the multi-compartment vehicle by using the multi-compartment fusion feature comprises:
[0171] inputting the multi-compartment fusion feature into a bird's eye view decoder to obtain the obstacle information of the multi-compartment vehicle, the obstacle information comprising three-dimensional coordinates of the obstacle in a whole-vehicle coordinate system of the multi-compartment vehicle, length, width and height of the obstacle, and a heading angle.
[0172] In an example embodiment, the obstacle information further comprises a category of the obstacle.
[0173] In an example embodiment, the method further comprises:
[0174] controlling the multi-compartment vehicle for automatic driving control based on the obstacle information.
[0175] In a third aspect, the present application further provides a computer program product comprising computer programs / instructions which, when executed by a processor, implement the steps of the obstacle identification method of the multi-compartment vehicle as described in any one of the above embodiments.
[0176] For the computer program product provided by the present application, please refer to the above embodiments, and the present application will not be repeated here.
[0177] The computer program product provided by the present application has the same beneficial effects as the obstacle identification method of the multi-compartment vehicle described above.
[0178] In a fourth aspect, referring to FIG. 10, the present application further provides an electronic device comprising:
[0179] a memory 21 for storing computer programs;
[0180] a processor 22 for executing the computer programs to implement the steps of the obstacle identification method of the multi-compartment vehicle as described in any one of the above embodiments.
[0181] The electronic device further comprises:
[0182] an input interface 23 connected to the processor 22 via a communication bus 26, for obtaining externally imported computer programs, parameters and instructions, and saving them to the memory 21 under the control of the processor 22. The input interface can be connected to an input device to receive parameters or instructions manually input by a user. The input device can be a touch layer overlaid on a display screen, or a key, trackball or touchpad provided on a terminal housing.
[0183] The display unit 24 is connected to the processor 22 through the communication bus 26, and is configured to display data sent by the processor 22. The display unit can be a liquid crystal display screen or an electronic ink display screen, etc.
[0184] The network port 25 is connected to the processor 22 through the communication bus 26, and is configured to be connected to external terminal devices for communication. The communication technology used in the communication connection can be wired communication technology or wireless communication technology, such as mobile high-definition link technology, universal serial bus, high-definition multimedia interface, wireless fidelity technology, Bluetooth communication technology, low-power Bluetooth communication technology, IEEE 802.11s-based communication technology, etc.
[0185] For the electronic device provided by the present application, please refer to the above-mentioned embodiments, and the present application will not be described here.
[0186] The electronic device provided by the present application has the same beneficial effects as the above-mentioned obstacle identification method for the multi-formation vehicle.
[0187] In the fifth aspect, referring to FIG. 11, the present application further provides a computer readable storage medium 30, and the computer readable storage medium 30 stores a computer program 31. When the computer program 31 is executed by a processor, the steps of the obstacle identification method for the multi-formation vehicle described in any one of the above-mentioned embodiments are implemented.
[0188] The computer readable storage medium 30 can include a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc. various media that can store program codes.
[0189] For the computer readable storage medium provided by the present application, please refer to the above-mentioned embodiments, and the present application will not be described here.
[0190] The computer readable storage medium provided by the present application has the same beneficial effects as the above-mentioned obstacle identification method for the multi-formation vehicle.
[0191] It is also noted that, in this disclosure, relational terms such as first and second, and the like, can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises," "comprising," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0192] The above description of disclosed embodiments provides enabling concepts for practicing or using the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for obstacle recognition in multi-car trains, wherein the multi-car trains comprise multiple articulated carriages, characterized in that, Each section of the carriages is provided with multiple cameras for collecting sub-360° images in their respective collection areas, and the obstacle recognition method for the multiple-unit vehicle comprises: For each carriage, the sub-360° images collected by the multiple cameras on the carriage are obtained, two-dimensional image features are extracted from all the sub-360° images, camera parameters, bird's eye view encoding and the two-dimensional image features corresponding to the cameras are processed by a single-carriage bird's eye view converter to obtain current bird's eye view features of the carriage; The current bird's eye view features of multiple carriages are fused to obtain multi-carriage fusion features; The multi-carriage fusion features are used to obtain obstacle information of the multiple-unit vehicle; The processing comprises: The camera parameters and first two-dimensional image features are calculated to obtain first key vectors and first value vectors; The first key vectors, the first value vectors and the first to-be-processed bird's eye view features are obtained based on the bird's eye view encoding to obtain second to-be-processed bird's eye view features; The camera parameters and second two-dimensional image features are calculated to obtain second key vectors and second value vectors; the feature down-sampling rate corresponding to the second two-dimensional image features is greater than the feature down-sampling rate corresponding to the first two-dimensional image features; The second key vectors, the second value vectors and the second to-be-processed bird's eye view features are processed to obtain second intermediate features; The second intermediate features are processed to obtain the current bird's eye view features of the carriage.
2. The obstacle recognition method for a multiple unit vehicle according to claim 1, characterized by, The single-carriage bird's eye view converter comprises a first fusion attention module, a first convolution module, a first full connection network module, a second full connection network module, a second fusion attention module and a second convolution module; The first key vectors and the first value vectors are calculated by using the first full connection network module to calculate the camera parameters and the first two-dimensional image features; The first key vectors, the first value vectors and the first to-be-processed bird's eye view features are processed by using the first fusion attention module to obtain first intermediate features; The first intermediate features are processed by using the first convolution module to obtain second to-be-processed bird's eye view features; The second key vectors and the second value vectors are calculated by using the second full connection network module to calculate the camera parameters and the second two-dimensional image features; the feature down-sampling rate corresponding to the second two-dimensional image features is greater than the feature down-sampling rate corresponding to the first two-dimensional image features; The second key vectors, the second value vectors and the second to-be-processed bird's eye view features are processed by using the second fusion attention module to obtain second intermediate features; The second intermediate features are processed by using the second convolution module to obtain the current bird's eye view features of the carriage. The first fusion attention module and the second fusion attention module each comprise a first deformable attention unit and a second deformable attention unit; 3. The obstacle recognition method for a multiple unit vehicle according to claim 2, characterized by, Further comprising: The first to-be-processed bird's eye view features are divided into P×P grids, and the features in each grid are extracted as the input of the first deformable attention unit; averaging a P×P size grid to divide a G×G number of grids, and extracting features in multiple related grids as inputs of the second deformable attention unit; P and G are both integers greater than 1.
4. The method of obstacle recognition for a multiple unit vehicle of claim 1, wherein, The process of fusing the current bird's eye view features of multiple sections of the carriages to obtain the multi-carriage fusion features includes: obtaining the hinging chain angles between every two adjacent carriages; obtaining the splicing features of the multiple sections of the carriages based on all the hinging chain angles and all the current bird's eye view features; fusing the splicing features to obtain the multi-carriage fusion features.
5. The method of obstacle recognition for a multiple unit vehicle of claim 4, wherein, The process of obtaining the splicing features of the multiple sections of the carriages based on all the hinging chain angles and all the current bird's eye view features includes: converting the current bird's eye view features of the ith carriage to the bird's eye view space coordinate system of the head carriage according to the hinging chain angle between the ith carriage and the i-1th carriage, to obtain the to-be-spliced features of the ith carriage; i = 2, 3…, n, n is the total number of carriages of the multiple-unit vehicle; determining the current bird's eye view features of the first carriage as the to-be-spliced features of the first carriage; splicing the to-be-spliced features of the multiple sections of the carriages to obtain the splicing features.
6. The method of obstacle recognition for a multiple unit vehicle of claim 4, wherein, The process of fusing the splicing features to obtain the multi-carriage fusion features includes: inputting the splicing features into a bird's eye view fusion converter to obtain the multi-carriage fusion features.
7. The method of obstacle recognition for a multi -unit vehicle according to claim 1, characterized in that, The process of obtaining the obstacle information of the multiple-unit vehicle by using the multi-carriage fusion features includes: inputting the multi-carriage fusion features into a bird's eye view decoder to obtain the obstacle information of the multiple-unit vehicle, the obstacle information including three-dimensional coordinates of the obstacle in the whole-vehicle coordinate system of the multiple-unit vehicle, length, width, height and heading angle of the obstacle.
8. The method of obstacle recognition for a multiple unit vehicle of claim 7, wherein, The obstacle information further includes the category of the obstacle.
9. The method of obstacle recognition for a multiple unit vehicle of any one of claims 1-8, wherein, Further including: automatically controlling the multiple-unit vehicle based on the obstacle information.
10. An obstacle recognition system for a multiple unit vehicle, the multiple unit vehicle comprising a plurality of articulated carriages, characterized in that, Each of the carriages is provided with multiple cameras, the cameras are used to collect sub-surround view images in their own collection areas, and the obstacle recognition system of the multiple-unit vehicle includes: a processing module, configured to, for each of the carriages, obtain the sub-surround view images collected by the multiple cameras on the carriage, extract two-dimensional image features from all the sub-surround view images, process the camera parameters corresponding to the cameras, bird's eye view encoding and the two-dimensional image features by a single-carriage bird's eye view converter to obtain current bird's eye view features of the carriage; a fusion module, configured to fuse the current bird's eye view features of multiple sections of the carriages to obtain multi-carriage fusion features; a decoding module, configured to obtain the obstacle information of the multiple-unit vehicle by using the multi-carriage fusion features; the processing includes: calculating the camera parameters and the first two-dimensional image features to obtain a first key vector and a first value vector; obtaining a first to-be-processed bird's eye view feature based on the bird's eye view encoding, and obtaining a second to-be-processed bird's eye view feature based on the first key vector, the first value vector and the first to-be-processed bird's eye view feature; The camera parameters and the second two-dimensional image features are calculated to obtain a second key vector and a second value vector; the second two-dimensional image features correspond to a feature down-sampling rate greater than a feature down-sampling rate corresponding to the first two-dimensional image features; The second key vector, the second value vector, and the second to-be-processed bird's-eye view feature are processed to obtain a second intermediate feature; The second intermediate feature is processed to obtain a current bird's-eye view feature of the vehicle compartment.
11. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by the processor to implement the steps of the obstacle identification method of the multi-formation vehicle according to any one of claims 1-9.
12. An electronic device, comprising: The computer program / instruction is executed by the processor to implement the steps of the obstacle identification method of the multi-formation vehicle according to any one of claims 1-9. The computer program / instruction is executed by the processor to implement the steps of the obstacle identification method of the multi-formation vehicle according to any one of claims 1-9. 13. A computer-readable storage medium, characterized in that,
Citation Information
Patent Citations
Train operation environment obstacle sensing method based on multi-mode fusion recognition
CN115953662A
Aerial view feature determination method, image processing method, device and equipment
CN116863153A
Vehicle aerial view generation method and device, storage medium and electronic device
CN117830526A
Roadside aerial view target detection method and device, computing equipment and storage medium
CN117994748A
Obstacle identification method, system and device for multi-marshalling vehicle, product and medium
CN118865332A