An environment perception method and device, computer equipment and medium
By combining 4D millimeter-wave radar and vision devices, the high cost and low efficiency problems caused by relying on deep learning in existing technologies are solved, achieving more accurate and robust environmental perception. By matching and fusing 3D target information from radar and vision data, the computational requirements are reduced.
Patent Information
- Application Number
- CN202311147056.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-05
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2043-09-05
AI Technical Summary
Existing autonomous driving environmental perception technologies rely on deep learning algorithms, which requires a large amount of labeled data and computing resources, resulting in high costs and low efficiency, and do not fully utilize the 3D target information of 4D millimeter-wave radar.
A method combining 4D millimeter-wave radar and vision equipment is adopted. 3D point cloud data and environmental image data are acquired by radar and vision equipment respectively to determine the first and second 3D target boxes. Matching and fusion are performed at the alternating time to generate a new fused target box. Tracking is performed using a fusion strategy with memory and Kalman filtering.
It does not rely on deep learning, reducing the demand for labeled data and computing resources, improving the accuracy and robustness of environmental perception, providing more accurate perception results by combining radar and visual data, and simplifying the calculation process by considering the position and shape of the bounding box.
Smart Images

Figure CN117315616B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and in particular to an environmental perception method, device, computer equipment, and medium. Background Technology
[0002] Current autonomous driving and advanced driver assistance systems (ADAS) widely utilize various types of sensors, such as radar, cameras, and LiDAR, for environmental perception. Each of these sensors has its own advantages and disadvantages. For example, radar is highly robust to changes in weather conditions and lighting, while cameras can provide rich color and texture information. Therefore, the development of sensor fusion technology aims to leverage the strengths of different sensor types to improve the accuracy and reliability of environmental perception.
[0003] Existing sensor fusion methods using 4D millimeter-wave radar and visual cameras typically rely on techniques such as deep learning or feature matching to fuse the data. For example, existing ADAS systems use 4D millimeter-wave radar and cameras for data fusion, but primarily rely on deep learning methods. The advantage of this approach is its ability to automatically learn features from the data, but the disadvantages are the need for large amounts of data and computational resources, and the difficulty in interpreting the results. Related technical details can be found in Aptiv's technical white paper. Another example is a fusion system that also uses 4D millimeter-wave radar and visual cameras, but it mainly relies on manually designed features and rules for data matching and fusion, requiring complex parameter design and failing to fully utilize the rich 3D target information provided by 4D millimeter-wave radar.
[0004] Existing technology discloses a 4D millimeter-wave radar and vision fusion perception method, including the following steps: data acquisition and preprocessing; extracting radar backbone features from millimeter-wave radar point cloud data in the form of points to obtain radar feature data; extracting image backbone features from image data to obtain image feature data, then performing 3D image detection to obtain the image 3D bounding box of the localized object and the feature data of the corresponding box, and transforming the detected image 3D bounding box to the vehicle coordinate system; transforming the radar feature data and image 3D bounding box to polar coordinate form, then associating the radar feature data and image 3D bounding box, and then performing attention encoding fusion on the associative radar feature data and the feature data corresponding to the image 3D bounding box; result output and post-processing. This algorithm relies on deep learning and requires the use of chips with high computing power.
[0005] Existing technology discloses a parking space detection method based on the fusion of 4D millimeter-wave radar and image recognition. This method includes: acquiring radar images collected by a time-synchronized 4D millimeter-wave radar and environmental images collected by an onboard camera; projecting the 4D millimeter-wave radar reflection points in the radar image onto image points in the environmental image to obtain radar projection points; in the image coordinate system, associating and matching the radar projection points with target pixels obtained by semantic segmentation; in the vehicle coordinate system, fusing the associating and matching radar projection points and target pixels to obtain fused points; clustering the fused points in the vehicle coordinate system and outputting the clustering results as obstacles; if it is determined that the space formed between any two obstacles can accommodate the vehicle, then the space is identified as the vehicle's parking space. This patent's fusion method is primarily applied in the parking field.
[0006] While the above methods improve environmental perception performance to some extent, they still cannot meet all driving situations and environmental conditions due to their reliance on deep learning or failure to fully utilize the advantages of 4D millimeter-wave radar. Summary of the Invention
[0007] In view of this, embodiments of the present invention provide an environmental perception method, apparatus, computer device, and medium, solving the technical problem in the prior art where environmental perception relies on deep learning algorithms, requiring a large amount of labeled data and computing resources, resulting in high cost and low efficiency. The method includes:
[0008] The radar is used to acquire 3D point cloud data of the vehicle's environment, and the vision device is used to acquire environmental image data of the vehicle's environment. The radar is a 4D millimeter-wave radar, and the radar and vision device are installed on the vehicle.
[0009] Determine the first 3D bounding box of the object to be perceived in the environment where the vehicle is located based on 3D point cloud data.
[0010] Determine the second 3D bounding box of the object to be perceived in the environment where the vehicle is located based on environmental image data;
[0011] In the first and second 3D target boxes, target boxes that match the fusion target box are alternately determined, and the determined target boxes are merged with the fusion target box to generate a new fusion target box. The new fusion target box is tracked to generate a continuous fusion track.
[0012] This invention also provides an environmental sensing device that solves the technical problem in the prior art where environmental sensing relies on deep learning algorithms, requiring a large amount of labeled data and computing resources, resulting in high cost and low efficiency. The device includes:
[0013] The data acquisition module is used to acquire 3D point cloud data of the vehicle's environment using radar and to acquire environmental image data of the vehicle's environment using vision equipment. The radar is a 4D millimeter-wave radar, and the radar and vision equipment are installed on the vehicle.
[0014] The first 3D target bounding box determination module is used to determine the first 3D target bounding box of the object to be perceived in the environment where the vehicle is located based on 3D point cloud data.
[0015] The second 3D target bounding box determination module is used to determine the second 3D target bounding box of the object to be perceived in the environment where the vehicle is located based on the environmental image data.
[0016] The association matching and fusion tracking module is used to alternately determine target boxes that match the fusion target box in the first 3D target box and the second 3D target box, merge the determined target boxes with the fusion target box to generate a new fusion target box, track the new fusion target box, and generate a continuous fusion track.
[0017] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-described environmental perception methods, solving the technical problem in the prior art where environmental perception relies on deep learning algorithms, requiring a large amount of labeled data and computing resources, resulting in high cost and low efficiency.
[0018] This invention also provides a computer-readable storage medium storing a computer program that executes any of the above-described environmental perception methods, thereby solving the technical problem in the prior art where environmental perception relies on deep learning algorithms, requiring a large amount of labeled data and computing resources, resulting in high cost and low efficiency.
[0019] Compared with the prior art, the beneficial effects that at least one technical solution adopted in the embodiments of this specification can achieve include at least:
[0020] The fusion method of 4D millimeter-wave radar and vision device data proposed in this invention does not rely on deep learning. Therefore, it does not require a large amount of labeled data and computing resources, which can reduce costs and improve efficiency. It makes full use of the rich 3D target information provided by 4D millimeter-wave radar. By combining radar and vision device data, the two data sources can be used simultaneously to provide more accurate and robust perception results. The 3D target box matching method takes into account the position and shape of the box, which can provide more accurate matching results. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of an environmental perception method provided in an embodiment of the present invention;
[0023] Figure 2 This is a flowchart illustrating an implementation of the above-described environmental perception method according to an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of a memory-based fusion strategy framework for an environmental perception method provided in an embodiment of the present invention;
[0025] Figure 4 This is a structural block diagram of a computer device provided in an embodiment of the present invention;
[0026] Figure 5 This is a structural block diagram of an environmental sensing device provided in an embodiment of the present invention. Detailed Implementation
[0027] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0028] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0029] In this embodiment of the invention, an environmental perception method is provided, such as... Figure 1 As shown, the method includes:
[0030] Step S101: Use radar to acquire 3D point cloud data of the vehicle's environment and use vision equipment to acquire environmental image data of the vehicle's environment. The radar is a 4D millimeter-wave radar, and the radar and vision equipment are installed on the vehicle.
[0031] Step S102: Determine the first 3D bounding box of the object to be perceived in the environment where the vehicle is located based on the 3D point cloud data.
[0032] Step S103: Determine the second 3D bounding box of the object to be perceived in the environment where the vehicle is located based on the environmental image data;
[0033] Step S104: Alternately determine target boxes that match the fusion target box in the first 3D target box and the second 3D target box, and merge the determined target boxes with the fusion target box to generate a new fusion target box. Track the new fusion target box to generate a continuous fusion track.
[0034] like Figure 2 As shown, in specific environmental perception, 4D millimeter-wave radar and vision devices are used to acquire environmental data respectively. The 4D millimeter-wave radar acquires 3D point cloud data of the environment around the vehicle, including distance, speed, angle, and height information for each point. The vision device can acquire high-resolution environmental image data. The 4D millimeter-wave point cloud acquired by the radar is post-processed to obtain the first 3D target bounding box. The image acquired by the vision camera is processed, and traditional 3D target detection algorithms are used to analyze the image and perform appropriate post-processing to extract the second 3D target bounding box. The 3D target bounding boxes output by the radar and camera are matched and associated with the fused target. The matched fused target is tracked, and the state (such as position, speed, and direction) of the fused target is tracked through continuous frames to output the fused target bounding box.
[0035] Specifically, a memory-based fusion strategy is employed, where the 3D target bounding boxes output by the radar (first 3D target bounding box) and the 3D target bounding boxes output by the vision device (second 3D target bounding box) are matched and associated with the fused target bounding box after matching, respectively. Instead of directly matching the targets output by the vision device and the radar, these strategies are used to match the fused target bounding boxes separately. This fusion strategy avoids problems such as discontinuous fused targets due to frame drops from a single sensor.
[0036] like Figure 3 As shown, each small box represents the target bounding box information. The first row, "Fusion," represents the fused target bounding box information. "Global" on the left indicates the globally fused target. "Sensori" represents the i-th sensor (radar or vision device), and "sensorj" represents the j-th sensor (radar or vision device). In the figure, 't' represents the time axis, and k below the time axis... x This represents a relative time. Due to the different sensor frame rates and transmission delays, we set up a global fusion target bounding box. Each time a new sensor sends target information, it is directly fused with the global fusion target bounding box. Figure 3 Chinese K i-1At time k, sensor i sends target information, which is directly fused with the target bounding box. G-3 The fusion is complete at any given moment. This framework retains information about the global track; when sensor data arrives, it predicts the global track and immediately fuses the sensor data. The fusion end is at k... G-3 At time sensori in k i-1 When data is reported in real time, it is immediately fused. Then, the global track data is used for prediction, at k... G-2 When sensor data arrives, it is fused again, and so on. Compared to the memoryless strategy, the global track is retained and maintained after creation. The memory-based fusion strategy retains the information of the global track. When sensor data arrives, it predicts the global track and immediately fuses the sensor data. This strategy is highly flexible and has a "plug-and-play" advantage because different sensor data are previously isolated and do not directly interact; they only interact indirectly through the global track. Therefore, adding or removing data from a particular sensor frame will not significantly affect the fusion process.
[0037] Specifically, radar-acquired data requires post-processing, including noise reduction and spatial point cloud clustering. First, the RANSAC algorithm is used to remove ground noise from the radar data. Then, the DBSCAN clustering algorithm is used to perform spatial point cloud clustering on the filtered 3D point cloud data to generate target boxes (the first target box). The purpose of post-processing is to aggregate the point cloud into multiple independent 3D target boxes based on the spatial distribution of the radar echoes.
[0038] In practice, to obtain 3D bounding boxes from environmental image data, the following steps are taken to determine the second 3D bounding box of the object to be perceived in the vehicle's environment:
[0039] The non-maximum suppression algorithm is used to remove detection boxes that overlap with maxima in the environmental image data to generate denoised environmental image data; the denoised environmental image data is then filtered and tracked to generate at least one second 3D target box.
[0040] Specifically, a deep learning-based 3D object detection algorithm is used to parse the image to obtain 3D bounding boxes (second 3D bounding boxes). Appropriate post-processing, such as non-maximum suppression (NMS), is then performed to eliminate redundant 3D bounding boxes. Finally, Kalman filtering is used for tracking to obtain the final 3D bounding boxes containing information such as distance, velocity, height, and object type.
[0041] In practice, to effectively fuse the target boxes, the following steps are taken to alternately determine target boxes that match the fusion target box within the first and second 3D target boxes, and then fuse the determined target boxes with the fusion target box to generate a new fusion target box:
[0042] At different times, the first 3D target box and the second 3D target box are alternately set as matching boxes, and the fused target box of the previous time is set as the matching box of the next time. At each time, the matching box that matches the matching box is determined and fused to generate a new fused target box for each time.
[0043] In practice, to find the target box pair with the highest matching degree and then merge the target box pairs, the following steps are taken to determine and merge the matching boxes that match the matched box at each time step, generating the merged target box at each time step:
[0044] At the current moment, the bounding box to be matched is paired with each matching box to form a target box pair, resulting in multiple target box pairs. The matching algorithm is used to calculate the total matching cost of each target box pair, resulting in multiple total matching costs. The total matching cost value represents the matching degree between the bounding box and the matching box. The smaller the total matching cost value, the higher the matching degree. A cost matrix is constructed using all the total matching costs. The element with the smallest total matching cost value is selected from the cost matrix as the optimal element. The matching box corresponding to the optimal element is merged with the bounding box to obtain the merged target box at the current moment.
[0045] In practice, to quickly and effectively calculate the total matching cost, the following steps are taken to calculate the total matching cost for each pair of target boxes using a matching algorithm:
[0046] For the matching and matched boxes in the target box pair, calculate the Mahalanobis distance and the intersection-union ratio (IUR). Normalize the Mahalanobis distance to generate an adjusted Mahalanobis distance, ensuring the adjusted distance is less than 1. Set weights for the Mahalanobis distance and IUR, where the sum of the Mahalanobis distance weights and IUR weights is 1. Use the formula cost = w d *d′+w i *i calculates the total matching cost, where cost is the total matching cost, and w d Here, d′ represents the Mahalanobis distance weights, and w represents the adjusted Mahalanobis distance. t Let i be the crossover-union ratio weight, and i be the crossover-union ratio.
[0047] In one embodiment, a greedy algorithm is used to construct a cost matrix and select the optimal element to match the matching box (either the first or second 3D target box) with the matched box (the target box being tracked after matching at the previous time step). Specifically, first, the matching cost of each pair of target boxes is calculated. A cost matrix is constructed using the matching costs, where each element represents the matching cost of a pair of target boxes. Second, the element with the lowest cost is found from the cost matrix, and the corresponding matching box (either the first or second 3D target box) is matched with the matched box (the target box being tracked after matching at the previous time step) (the greedy algorithm always selects the best matching pair at the moment, i.e., the matching pair with the lowest cost). Third, all elements related to the already matched boxes are removed from the cost matrix to ensure that each target box can only be matched once. Finally, the above steps are repeated until all target boxes have been matched, or no matching pair that meets the cost threshold can be found. If no matching pair that meets the cost threshold can be found, it will be considered unmatched. The advantage of the greedy algorithm is its simplicity in implementation and its ability to obtain near-optimal results (the element with the lowest cost). However, since the greedy algorithm only considers the current best match and does not take into account the overall matching effect, it may not be able to obtain the globally optimal matching result in some cases.
[0048] In another embodiment, the Hungarian algorithm is used to construct a cost matrix and select the optimal element to match the matching box (either the first or second 3D target box) with the matched box (the target box tracked after the previous time step). The Hungarian algorithm is also known as the optimal assignment algorithm. In the Hungarian algorithm, a cost matrix is first constructed, where each element is the weighted average of the normalized Mahalanobis distance and normalized IoU between the two boxes. Then, the Hungarian algorithm is used to find the optimal one-to-one match in this cost matrix. The Hungarian algorithm guarantees a globally optimal match (minimizing the total cost), but its computational complexity is high, especially when the number of target boxes is large.
[0049] Specifically, the Mahalanobis distance between the target bounding boxes is calculated using the following steps:
[0050] Suppose we have two 3D bounding boxes A (the matching box) and B (the bounding box to be matched), whose characteristics can be represented by vectors a and b, respectively. The elements of these vectors include the x, y, and z coordinates of the center point, as well as the height, width, and length of the bounding box. Where a = [x...] a ,y a ,z a ,h a ,w a ,l a ], [x b ,y b ,zb ,h b ,w b ,l b Under the above settings, the Mahalanobis distance formula can be defined as: D = Where T is the transpose of a vector, * represents matrix multiplication, and S -1 This represents the inverse of the covariance matrix. The covariance matrix S is derived statistically from the data. The Mahalanobis distance formula calculates the square of the Mahalanobis distance because square root operations can introduce additional complexity. The matching method in this embodiment of the invention focuses on the relative magnitude of the Mahalanobis distance; therefore, square root operations are ignored.
[0051] Specifically, the intersection-over-union (IoU) ratio of two 3D bounding boxes (the matching box and the matched box) is calculated through the following steps:
[0052] Calculate the intersection volume and the joint volume, then divide them. Assume there are two 3D bounding boxes A (the matching box) and B (the matched box), each defined by a pair of diagonal points, or by the coordinates of its center point (x, y, z) and dimensions (width w, height h, length l). Taking diagonal point definition as an example, the two diagonal points of A are A1(x1, y1, z1) and A2(x2, y2, z2), and the two diagonal points of B are B1(x3, y3, z3) and B2(x4, y4, z4); calculate the intersection volume. Find the diagonal points of the intersection box. The coordinates of one diagonal point are max(A1, B1), and the coordinates of the other diagonal point are min(A2, B2). If these two points form a valid 3D box (i.e., all coordinates satisfy diagonal point 1 < diagonal point 2), then the intersection volume can be calculated as length multiplied by width multiplied by height. If these two points do not form a valid 3D bounding box (i.e., there is some coordinate that satisfies diagonal point 1 >= diagonal point 2), then the intersection volume is 0; calculate the union volume. The union volume is the sum of the volumes of the two bounding boxes minus their intersection volume; calculate the IoU, which is the intersection volume divided by the union volume.
[0053] Specifically, Mahalanobis distance is a way to measure the similarity between data points, taking into account the weights of each feature and the relationships between them. In our application, features can be the position, size, shape, etc., of the bounding boxes. IoU is a way to measure the degree of overlap between two bounding boxes. Both can provide useful information for the matching process. Therefore, we normalize Mahalanobis distance and IoU to the same range (e.g., 0 to 1), and then take a weighted average of them, with the weights summing to 1, to obtain the overall similarity. The Mahalanobis distance is normalized using the following formula: Where, d maxThis is the preset maximum Mahalanobis distance. Since a smaller Mahalanobis distance indicates a better match, we use 1 minus the exponent term for normalization. After processing, d' will also be a value between 0 and 1, and a smaller value indicates a better match. For IoU, it is already a value between 0 and 1, so no additional normalization is needed.
[0054] Specifically, using the formula cost = w d *d′+w i *i calculates the total matching cost, where w d and w i These are the preset weights for Mahalanobis distance and intersection-union ratio, respectively, and they satisfy w d +w i =1. Using the total matching cost formula to calculate the total matching cost ensures that the total matching cost is a value between 0 and 1, with a smaller value indicating a better match. Furthermore, the relative importance of Mahalanobis distance and IoU in calculating the total matching cost can be controlled by adjusting the weights.
[0055] In practice, to improve matching accuracy and reduce matching errors, the following steps are taken to perform post-processing after the target matching is completed:
[0056] After fusing the matching box included in the optimal element with the matched box to obtain the fused target box at the current time, if it is determined that the Mahalanobis distance corresponding to the optimal element is greater than the first threshold and / or the intersection-union ratio is less than the second threshold, then the fused target box at the current time is not used for tracking.
[0057] Specifically, after the target matching is completed, post-processing is required, such as threshold filtering. If the Mahalanobis distance between two target boxes is too large (greater than the first threshold) or the IoU is too small (less than the second threshold), they should be considered a mismatch even if they are matched during the matching process. This usually means that one of the 3D target boxes may be an incorrect detection result, or that the two target boxes correspond to different real targets.
[0058] In practice, to obtain a continuously fused track, the following steps are performed to track the new fused target box and generate a continuously fused track:
[0059] Determine the state variables of the new fusion target box; set the initial state and initial parameters for the Kalman filter of the filtering algorithm; execute the following iterative steps to generate a continuous fusion track by tracking the new fusion target box until the tracking ends or the new fusion target box is lost, and then end the iterative steps: In each time step, predict the state variables of the new fusion target box at the next time step based on the state variables of the new fusion target box at the current time and the parameters of the Kalman filter, update the predicted state variables through the observation equation, and readjust the parameters of the Kalman filter according to the observation noise and system noise.
[0060] Specifically, filtering algorithms (such as Kalman filtering) are used to track the matched and fused targets. The purpose of tracking is to track the state of the fused target bounding box, such as position, velocity, and orientation, through consecutive frames. Filtering mainly includes several steps: initialization, prediction, and update.
[0061] At the start of tracking, the Kalman filter is initialized. The initial frame can generate a fused bounding box after matching using the successful results of radar and vision device matching. Initial values are set for other parameters of the Kalman filter, such as the state transition matrix, observation matrix, process noise covariance matrix, and observation noise covariance matrix. If the state vector of the fused bounding box is set to include the 3D position (x, y, z) and velocity (v) of the bounding box... x ,v y ,v z ), where x, y, and z represent the x-axis, y-axis, and z-axis directions, respectively. At this point, the state vector of the fused target bounding box is set to X = [x, y, z, v]. x ,v y ,v z ] T , where T is the transpose of the vector;
[0062] Predicting the next state. At each time step, the Kalman filter first makes a prediction. Based on the state of the previous time step and knowledge of the system behavior (usually represented by the state transition matrix), the prediction yields a predicted state and a prediction error covariance. The formula for the predicted state is X. pre =FX + Bu, predicting covariance P pre =FPF T +Q, where X is the current state, F is the state transition matrix (used to describe the dynamics of the system), P is the state covariance matrix (used to describe the uncertainty of the state), B is the control input matrix, u is the control input, and Q is the process noise covariance matrix (used to describe the uncertainty of the model).
[0063] The updated state is output. Each frame receives a new measurement (in this embodiment, the new position of the fused target box). The Kalman filter updates its state estimate based on this measurement and the predicted state. During the update process, the Kalman gain formula K = P is used. pre H T HP pre H T +R) -1 Calculate the Kalman gain, then use the Kalman gain to adjust the predicted state, and utilize P = (IK*H)*P pre Update the prediction error covariance, and finally output the updated state X = X pre +K*(zH*X pre ), where H is the observation matrix (used to describe how the observations are obtained from the state), R is the observation noise covariance matrix (used to describe the uncertainty of the measurement), z is the actual observation value, and I is the identity matrix.
[0064] In this embodiment, a computer device is provided, such as... Figure 4 As shown, it includes a memory 401, a processor 402, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-described environmental perception methods.
[0065] Specifically, the computer device can be a computer terminal, a server, or a similar computing device.
[0066] In this embodiment, a computer-readable storage medium is provided, which stores a computer program that performs any of the above-described environmental perception methods.
[0067] Specifically, computer-readable storage media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media does not include transient media, such as modulated data signals and carrier waves.
[0068] Based on the same inventive concept, this invention also provides an environmental sensing device, as described in the following embodiments. Since the principle by which the environmental sensing device solves the problem is similar to that of the environmental sensing method, the implementation of the environmental sensing device can refer to the implementation of the environmental sensing method, and repeated details will not be elaborated further. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0069] Figure 5 This is a structural block diagram of an environmental sensing device according to an embodiment of the present invention, such as... Figure 5 As shown, it includes: a data acquisition module 501, a first 3D target bounding box determination module 502, a second 3D target bounding box determination module 503, and an association matching and fusion tracking module 504. The structure is described below.
[0070] The data acquisition module 501 is used to acquire 3D point cloud data of the vehicle's environment using radar and to acquire environmental image data of the vehicle's environment using vision equipment. The radar is a 4D millimeter-wave radar, and the radar and vision equipment are installed on the vehicle.
[0071] The first 3D target bounding box determination module 502 is used to determine the first 3D target bounding box of the object to be perceived in the environment where the vehicle is located based on 3D point cloud data.
[0072] The second 3D target bounding box determination module 503 is used to determine the second 3D target bounding box of the object to be perceived in the environment where the vehicle is located based on the environmental image data.
[0073] The association matching and fusion tracking module 504 is used to alternately determine the target box that matches the fusion target box in the first 3D target box and the second 3D target box, merge the determined target box with the fusion target box to generate a new fusion target box, track the new fusion target box, and generate a continuous fusion track.
[0074] In one embodiment, the second 3D target bounding box determination module includes:
[0075] The denoising unit is used to remove detection boxes that overlap with maxima in the environmental image data using a non-maximum suppression algorithm, and generate denoised environmental image data.
[0076] The target bounding box generation unit is used to filter and track the denoised environmental image data to generate at least one second 3D target bounding box.
[0077] In one embodiment, the association matching and fusion tracking module includes:
[0078] The matching box fusion unit is used to alternately set the first 3D target box and the second 3D target box as matching boxes according to different time periods, set the fusion target box of the previous time period as the matching box of the next time period, and at each time period, determine the matching box that matches the matching box and perform fusion to generate a new fusion target box for each time period.
[0079] The fusion tracking unit is used to determine the state variables of the new fusion target box; set the initial state and initial parameters for the Kalman filter of the filtering algorithm; and execute the following loop steps to generate a continuous fusion track by tracking the new fusion target box until the tracking ends or the new fusion target box is lost, at which point the loop steps end: within each time step, the state variables of the new fusion target box at the next time step are estimated based on the state variables of the new fusion target box at the current time step and the parameters of the Kalman filter, the estimated state variables are updated through the observation equation, and the parameters of the Kalman filter are readjusted according to the observation noise and system noise.
[0080] In one embodiment, the matching box fusion unit is configured to, at the current moment, form target box pairs with each matching box to obtain multiple target box pairs, calculate the total matching cost of each target box pair using a matching algorithm, and obtain multiple total matching costs, wherein the value of the total matching cost represents the matching degree between the matching box and the matching box, and the smaller the value of the total matching cost, the higher the matching degree; construct a cost matrix using all the total matching costs, select the element with the smallest total matching cost value from the cost matrix as the optimal element; and fuse the matching box corresponding to the optimal element with the matching box to obtain the fused target box at the current moment.
[0081] In one embodiment, the matching box fusion unit is further configured to calculate the Mahalanobis distance and the intersection-union ratio (IUR) for the matching boxes and matched boxes in the target box pair; normalize the Mahalanobis distance to generate an adjusted Mahalanobis distance, such that the adjusted Mahalanobis distance is less than 1; set the weights for the Mahalanobis distance and the IUR respectively, wherein the sum of the weights for the Mahalanobis distance and the IUR is 1; and use the formula cost = w d *d′+w i *i calculates the total matching cost, where cost is the total matching cost, and w d Here, d′ represents the Mahalanobis distance weights, and w represents the adjusted Mahalanobis distance. i Let i be the crossover-union ratio weight, and i be the crossover-union ratio.
[0082] In one embodiment, the environmental sensing device further includes a threshold filtering module, which is used to fuse the matching box included in the optimal element with the matched box to obtain the fused target box at the current time. If it is determined that the Mahalanobis distance corresponding to the optimal element is greater than a first threshold and / or the intersection-union ratio is less than a second threshold, then the fused target box at the current time is not used for tracking.
[0083] The embodiments of the present invention achieve the following technical effects:
[0084] The environmental perception method proposed in this invention does not rely on deep learning. Therefore, it does not require a large amount of labeled data and computational resources, thus reducing costs and improving efficiency. The environmental perception method in this invention uses explicit algorithms for data processing and fusion, making it easier to understand and interpret the system's behavior. It fully utilizes the rich 3D target information provided by 4D millimeter-wave radar. By combining data from radar and vision devices, both data sources can be used simultaneously to provide more accurate and robust perception results. By using weighted Mahalanobis distance and IoU as matching costs, it can more accurately match the 3D target boxes of radar and vision devices. Furthermore, the matching method considers the position and shape of the boxes, providing more accurate matching results. It uses a greedy algorithm for data association, which is simple and computationally efficient, and in most cases can provide results comparable to more complex matching algorithms. The Kalman filter can effectively track the state of each target. As a linear predictive filter, the Kalman filter can estimate the system state from noisy data, making the tracking results more stable and accurate.
[0085] Obviously, those skilled in the art should understand that the modules or steps of the above-described embodiments of the present invention can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the embodiments of the present invention are not limited to any particular hardware and software combination.
[0086] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An environmental perception method, characterized in that, include: The vehicle uses radar to acquire 3D point cloud data of its environment and a vision device to acquire environmental image data of the environment. The radar is a 4D millimeter-wave radar and the radar and the vision device are mounted on the vehicle. Based on the 3D point cloud data, determine the first 3D target bounding box of the object to be perceived in the environment where the vehicle is located. Based on the environmental image data, a second 3D target bounding box of the object to be perceived in the environment where the vehicle is located is determined; The process involves alternately determining target boxes that match the fusion target box within the first 3D target box and the second 3D target box, and then fusing the determined target boxes with the fusion target box to generate a new fusion target box. This includes: alternately setting the first 3D target box and the second 3D target box as matching boxes at different times, setting the fusion target box from the previous time as the matching box for the next time, and at each time, determining and fusing the matching box that matches the matching box to generate a new fusion target box for each time. The fusion target box is a global fusion target box and is used to retain information about the global track. When the 3D point cloud data or the environmental image data arrives, the global track is predicted, and then the environmental image data or the 3D point cloud data is fused so that the first 3D target box or the second 3D target box indirectly contacts each other through the global fusion target box. The new fused target bounding box is tracked to generate a continuous fused track.
2. The environmental perception method as described in claim 1, characterized in that, At each time step, the matching box that matches the matched box is determined and merged to generate the fused target box at each time step, including: At the current moment, the matched box is paired with each of the matched boxes to form a target box pair, resulting in multiple target box pairs. The total matching cost of each target box pair is calculated using a matching algorithm, resulting in multiple total matching costs. The value of the total matching cost represents the matching degree between the matched box and the matched box. The smaller the value of the total matching cost, the higher the matching degree. Construct a cost matrix using all the total matching costs, and select the element with the smallest total matching cost from the cost matrix as the optimal element; The matching box corresponding to the optimal element is merged with the matched box to obtain the fused target box at the current time.
3. The environmental perception method as described in claim 2, characterized in that, The total matching cost for each of the target box pairs is calculated using a matching algorithm, including: For the matching box and the matched box in the target box pair, calculate the Mahalanobis distance and the intersection-union ratio; The Mahalanobis distance is normalized to generate an adjusted Mahalanobis distance, such that the adjusted Mahalanobis distance is less than 1. Set Mahalanobis distance weight and intersection-union ratio weight respectively, wherein the sum of the Mahalanobis distance weight and the intersection-union ratio weight is 1; Using formula Calculate the total matching cost, where cost is the total matching cost. The Mahalanobis distance weights are... The adjusted Mahalanobis distance, The intersection-union ratio weights are... Let be the crossover-union ratio.
4. The environmental perception method as described in claim 3, characterized in that, Also includes: After fusing the matching box and the matched box included in the optimal element to obtain the fused target box at the current time, if it is determined that the Mahalanobis distance corresponding to the optimal element is greater than the first threshold and / or the intersection-union ratio is less than the second threshold, then the fused target box at the current time is not used for tracking.
5. The environmental perception method as described in any one of claims 1 to 4, characterized in that, Based on the environmental image data, a second 3D bounding box is determined for the objects to be perceived in the vehicle's environment, including: The non-maximum suppression algorithm is used to remove the detection boxes that overlap with the maxima in the environmental image data to generate denoised environmental image data. The denoised environmental image data is filtered and tracked to generate at least one second 3D target bounding box.
6. The environmental perception method as described in any one of claims 1 to 4, characterized in that, Tracking the new fused target bounding box to generate a continuous fused track includes: Determine the state variables of the new fusion target box; Set the initial state and initial parameters for the Kalman filter of the filtering algorithm; The following iterative steps are performed to generate a continuous fused track by tracking the new fused target box until tracking ends or the new fused target box is lost, at which point the iterative steps end: Within each time step, the state variables of the new fused target box at the next time step are estimated based on the state variables of the new fused target box at the current time and the parameters of the Kalman filter. The estimated state variables are updated through the observation equation, and the parameters of the Kalman filter are readjusted according to the observation noise and system noise.
7. An environmental sensing device, characterized in that, include: The data acquisition module is used to acquire 3D point cloud data of the vehicle's environment using radar and to acquire environmental image data of the vehicle's environment using a vision device. The radar is a 4D millimeter-wave radar, and the radar and the vision device are mounted on the vehicle. The first 3D target bounding box determination module is used to determine the first 3D target bounding box of the object to be sensed in the environment where the vehicle is located based on the 3D point cloud data. The second 3D target bounding box determination module is used to determine the second 3D target bounding box of the object to be sensed in the environment where the vehicle is located based on the environmental image data. The association matching and fusion tracking module is used to alternately determine target boxes that match the fusion target box in the first 3D target box and the second 3D target box, and merge the determined target boxes with the fusion target box to generate a new fusion target box, track the new fusion target box, and generate a continuous fusion track. The correlation matching and fusion tracking module includes: The matching box fusion unit is used to alternately set the first 3D target box and the second 3D target box as matching boxes at different times, and set the fusion target box of the previous time as the matching box of the next time. At each time, the matching box that matches the matching box is determined and fused to generate a new fusion target box for each time. The fusion target box is a global fusion target box and is used to retain the information of the global track. When the 3D point cloud data or the environmental image data arrives, the global track is predicted and then the environmental image data or the 3D point cloud data is fused so that the first 3D target box or the second 3D target box indirectly contacts each other through the global fusion target box.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the environmental perception method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that performs the environmental perception method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Data fusion method and device, electronic equipment and storage medium
CN116310674A
Highway thrown object detection method and device based on background model and tracking
CN116434160A