Target detection method and device based on vision and radar, vehicle and medium
By fusing camera and radar data, the target state is dynamically corrected, solving the problem of low target object perception accuracy when the vehicle turns right, and achieving higher perception accuracy and robustness.
Patent Information
- Application Number
- CN202511590925.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-02-17
AI Technical Summary
In existing technologies, the accuracy of target object perception is low when a vehicle turns right due to reliance on a single sensor. This is especially prone to errors and misjudgments in complex environments, and it is impossible to effectively obtain information about the inner wheel difference blind spot.
By synchronously acquiring target object data from cameras and radar, performing intra-frame and inter-frame data fusion, and combining the advantages of different sensors, the target state is dynamically corrected to generate more reliable target object data.
It significantly improves the accuracy and robustness of target object perception, overcomes the problems of single sensor error and visual misjudgment, and ensures accurate tracking and stability of target objects.
Smart Images

Figure CN121541185A_ABST
Abstract
Description
Technical Field
[0001] This application relates to vehicles, and more particularly to a target detection method, device, vehicle, and medium based on vision and radar. Background Technology
[0002] For some larger vehicles, such as trucks, a significant inner wheel difference blind spot is created when turning right due to their wide turning radius. If pedestrians, non-motorized vehicles, or small vehicles cross this blind spot, the driver's view is obstructed, making it difficult to obtain specific information about these targets in time, greatly increasing the risk of collisions and endangering personal safety. Therefore, it is necessary to obtain real-time and comprehensive information about the inner wheel difference blind spot through reliable data collection methods to provide accurate information for driving decisions and driver assistance systems.
[0003] To meet the precise perception requirements for truck right turns, the current approach primarily utilizes machine vision and millimeter-wave radar. Specifically, observation data (such as longitudinal distance and speed) from the millimeter-wave radar system and observation data (such as lateral position and target type) from the machine vision system are acquired separately. Then, the two types of data are matched by coordinate unification and correlation between point traces and flight paths. Subsequently, Kalman filtering is used to track the target and update its lifecycle status. Finally, the output signals from both types of sensors are fused using a preset formula, combining the advantages of both to improve the perception effect.
[0004] However, existing technologies still suffer from the technical problem of low accuracy in perceiving target objects. Summary of the Invention
[0005] This application provides a target detection method, device, vehicle, and medium based on vision and radar to improve the accuracy of target object perception.
[0006] In a first aspect, embodiments of this application provide a target detection method based on vision and radar, including:
[0007] The system acquires first target object data and second target object data of the vehicle's driving environment at the current moment. The first target object data is collected by a camera, and the second target object data is collected by radar.
[0008] The first target object data and the second target object data are merged to generate the third target object data;
[0009] Based on the third target object data and the historical third target object data generated at the previous moment, the fourth target object data is determined.
[0010] In this technical solution, the first target object data and the second target object data are acquired simultaneously and fused to generate the third target object data, effectively combining the advantages of different sensors; then, the fourth target object data is determined based on the third target object data and the historical third target object generated at the previous moment, realizing dynamic state correction during target tracking, overcoming the perception reliability problem caused by single sensor error or visual misjudgment, and significantly improving the overall target object perception accuracy and robustness.
[0011] In one possible implementation, the target object data includes target feature values of the target object and the confidence level corresponding to each target feature value. The target features include at least one of the following: location, type, size, velocity, and orientation.
[0012] In this implementation, the target object data includes the confidence level of feature values such as location, type, size, speed, and orientation. This ensures that each feature value has assessable credibility, laying the foundation for subsequent data processing and decision-making processes, and further enhancing the reliability of the perception results.
[0013] In one possible implementation, fusing the first target object data and the second target object data to generate the third target object data includes:
[0014] Construct a first set based on the first target object contained in the first target object data;
[0015] Construct a second set based on the second target object contained in the second target object data;
[0016] For any first target object in the first set, calculate the value of the target association index between the first target object and each second target object in the second set, wherein the target association index includes at least one of distance difference, IOU, size difference and speed difference;
[0017] The second target object whose value of the target association index in the second set meets the corresponding preset condition is determined as the candidate target object;
[0018] Associate the first target object with the candidate target object;
[0019] Update the third set and the fourth set, wherein the third set is used to store target object pairs that have a relationship, and the fourth set is used to store first target objects that do not have a relationship, and / or second target objects that do not have a relationship;
[0020] The third target object data is generated based on the third set and the fourth set.
[0021] In this implementation, by constructing a first set and a second set, and calculating the target correlation index between target objects, candidate target objects that meet preset conditions are associated with the first target object. This process then updates the third and fourth sets, ultimately generating the third target object data. This mechanism achieves accurate association and data integration of the same target object from different sources, effectively improving the accuracy and data integrity of target matching in multi-sensor perception systems.
[0022] In one possible implementation, associating the first target object with the candidate target object includes:
[0023] If the candidate target object is already associated with other first target objects, then determine the first IOU between the candidate target object and the first target object, and the second IOU between the candidate target object and the other first target objects;
[0024] If the first IOU is greater than the second IOU, then the association between the candidate target object and the other first target objects is terminated, and the candidate target object is associated with the first target object.
[0025] In this implementation, by determining whether the candidate target object already has an association during the association process, and updating the association when the first IOU is greater than the second IOU, the many-to-one target contention problem is effectively solved, ensuring a stable association between the target object and the best candidate, thereby improving the accuracy and reliability of the third target object data.
[0026] In one possible implementation, generating the third target object data based on the third set and the fourth set includes:
[0027] Based on the target object pairs with related relationships contained in the third set, determine the target feature value pairs of the target object pairs from the first target object data and the second target object data;
[0028] For each pair of target objects, the target feature value pairs are fused to generate the third target object data;
[0029] Based on the first target objects in the fourth set that have no association relationship, the target feature values of the first target objects that have no association relationship are obtained from the first target object data to generate the third target object data;
[0030] Based on the second target objects that have no association in the fourth set, the target feature values of the second target objects that have no association are obtained from the second target object data, and the third target object data is generated.
[0031] In this implementation, by performing data fusion on target object pairs with related relationships based on the third set and retaining unmatched independent target objects based on the fourth set, it is ensured that all perceived target objects can be effectively processed and output, avoiding the loss of effective targets and significantly improving the integrity of the third target object data.
[0032] In one possible implementation, determining the fourth target object data based on the third target object data and the historical third target object data generated at the previous moment includes:
[0033] Based on the historical third target object data, determine the predicted location of the target object;
[0034] The fourth target object data is determined based on the predicted location of the target object, the historical third target object data, and the third target object data.
[0035] In this implementation, the third target object data is further corrected based on the historical third target object data generated in the previous moment. This enables dynamic correction of the target state during the tracking process, effectively smoothing the trajectory and suppressing measurement noise, thereby improving the stability and accuracy of the tracking.
[0036] In one possible implementation, acquiring the first target object data and the second target object data of the vehicle driving environment at the current moment includes:
[0037] Acquire initial target object data of the vehicle driving environment at the current moment. The initial target object data includes initial first target object data and / or initial second target object data. The initial first target object data is acquired by the camera, and the initial second target object data is acquired by the radar.
[0038] For any target object in the initial target object data, if the target object in the initial target object data is of the first type and the target object in the historical initial target object data of the previous time is of the second type, then the type continuity number of the target object in the previous time is determined according to the historical initial target object data of the previous time. The type continuity number is used to represent the number of times the target object was continuously identified as the second type in the previous time. The previous time includes the previous time.
[0039] If the number of consecutive types is greater than a preset number, then the first type in the initial target object data is corrected to the second type, and target object data is generated;
[0040] If the number of consecutive occurrences of the type is less than or equal to the preset number, then the initial target object data is determined as the target object data. Specifically, the target object data corresponding to the initial first target object data is the first target object data, and the target object data corresponding to the initial second target object data is the second target object data.
[0041] In this implementation, by statistically analyzing the number of consecutive types of the target object in historical frames, the target object in the initial target object data is corrected when the historical types are stable. This suppresses instantaneous type jumps caused by environmental interference, improving the continuity and stability of the output type. When the historical types are unstable, the perception result of the current frame is trusted to ensure responsiveness to changes in the actual type. This technical solution significantly improves the reliability of the output type of a single sensor in complex scenarios such as occlusion and sudden changes in lighting, providing higher-quality and more reliable input data for subsequent multi-sensor data fusion.
[0042] In one possible implementation, the method further includes:
[0043] Within a time period of a preset continuous duration preceding the current moment, determine the number of type changes of the target object;
[0044] Based on the number of type changes of the target object, the number of consecutive types, and the type corresponding to the target object within the preset continuous time period, the initial confidence level in the initial target object data is updated to generate target object data.
[0045] In this implementation, the initial confidence level in the initial target object data is updated by comprehensively considering the number of type changes and the number of type continuities of the target object. When the number of type continuities is high and the number of type changes is low, the updated confidence level is correspondingly improved; when type changes are frequent, the updated confidence level is correspondingly reduced, thereby improving the accuracy of the confidence level and laying the foundation for subsequent data fusion.
[0046] Secondly, embodiments of this application provide a target detection device based on vision and radar, comprising:
[0047] The acquisition module is used to acquire first target object data and second target object data of the vehicle driving environment at the current moment. The first target object data is collected by a camera, and the second target object data is collected by radar.
[0048] The fusion module is used to fuse the first target object data and the second target object data to generate the third target object data;
[0049] The determination module is used to determine the fourth target object data based on the third target object data and the historical third target object data generated at the previous moment.
[0050] In one possible implementation, the target object data includes target feature values of the target object and the confidence level corresponding to each target feature value. The target features include at least one of the following: location, type, size, velocity, and orientation.
[0051] In one possible implementation, the fusion module is specifically used for:
[0052] Construct a first set based on the first target object contained in the first target object data;
[0053] Construct a second set based on the second target object contained in the second target object data;
[0054] For any first target object in the first set, calculate the value of the target association index between the first target object and each second target object in the second set, wherein the target association index includes at least one of distance difference, IOU, size difference and speed difference;
[0055] The second target object whose value of the target association index in the second set meets the corresponding preset condition is determined as the candidate target object;
[0056] Associate the first target object with the candidate target object;
[0057] Update the third set and the fourth set, wherein the third set is used to store target object pairs that have a relationship, and the fourth set is used to store first target objects that do not have a relationship, and / or second target objects that do not have a relationship;
[0058] The third target object data is generated based on the third set and the fourth set.
[0059] In one possible implementation, the fusion module is specifically used for:
[0060] If the candidate target object is already associated with other first target objects, then determine the first IOU between the candidate target object and the first target object, and the second IOU between the candidate target object and the other first target objects;
[0061] If the first IOU is greater than the second IOU, then the association between the candidate target object and the other first target objects is terminated, and the candidate target object is associated with the first target object.
[0062] In one possible implementation, the fusion module is specifically used for:
[0063] Based on the target object pairs with related relationships contained in the third set, determine the target feature value pairs of the target object pairs from the first target object data and the second target object data;
[0064] For each pair of target objects, the target feature value pairs are fused to generate the third target object data;
[0065] Based on the first target objects in the fourth set that have no association relationship, the target feature values of the first target objects that have no association relationship are obtained from the first target object data to generate the third target object data;
[0066] Based on the second target objects that have no association in the fourth set, the target feature values of the second target objects that have no association are obtained from the second target object data, and the third target object data is generated.
[0067] In one possible implementation, a module is defined, specifically for:
[0068] Based on the historical third target object data, determine the predicted location of the target object;
[0069] The fourth target object data is determined based on the predicted location of the target object, the historical third target object data, and the third target object data.
[0070] In one possible implementation, the acquisition module is specifically used for:
[0071] Acquire initial target object data of the vehicle driving environment at the current moment. The initial target object data includes initial first target object data and / or initial second target object data. The initial first target object data is acquired by the camera, and the initial second target object data is acquired by the radar.
[0072] For any target object in the initial target object data, if the target object in the initial target object data is of the first type and the target object in the historical initial target object data of the previous time is of the second type, then the type continuity number of the target object in the previous time is determined according to the historical initial target object data of the previous time. The type continuity number is used to represent the number of times the target object was continuously identified as the second type in the previous time. The previous time includes the previous time.
[0073] If the number of consecutive types is greater than a preset number, then the first type in the initial target object data is corrected to the second type, and target object data is generated;
[0074] If the number of consecutive occurrences of the type is less than or equal to the preset number, then the initial target object data is determined as the target object data. Specifically, the target object data corresponding to the initial first target object data is the first target object data, and the target object data corresponding to the initial second target object data is the second target object data.
[0075] In one possible implementation, the vision- and radar-based target detection device further includes an update module for:
[0076] Within a time period of a preset continuous duration preceding the current moment, determine the number of type changes of the target object;
[0077] Based on the number of type changes of the target object, the number of consecutive types, and the type corresponding to the target object within the preset continuous time period, the initial confidence level in the initial target object data is updated to generate target object data.
[0078] Thirdly, embodiments of this application provide a vehicle, including: a vehicle body, a memory, and a processor;
[0079] The memory stores computer-executed instructions;
[0080] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0081] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0082] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0083] The target detection method, device, vehicle, and medium based on vision and radar provided in this application simultaneously acquire first target object data and second target object data, and fuse them to generate third target object data, effectively combining the advantages of different sensors; then, based on the third target object data and the historical third target object data generated at the previous moment, the fourth target object data is determined, and dynamic state correction is achieved during target tracking, overcoming the perception reliability problem caused by single sensor error or visual misjudgment, and significantly improving the target object perception accuracy and robustness. Attached Figure Description
[0084] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0085] Figure 1 A flowchart illustrating the vision- and radar-based target detection method provided in this application. Figure 1 ;
[0086] Figure 2 A flowchart illustrating the vision- and radar-based target detection method provided in this application. Figure 2 ;
[0087] Figure 3 A flowchart illustrating the vision- and radar-based target detection method provided in this application. Figure 3 ;
[0088] Figure 4 A flowchart illustrating the vision- and radar-based target detection method provided in this application. Figure 4 ;
[0089] Figure 5 A flowchart illustrating the vision- and radar-based target detection method provided in this application. Figure 5 ;
[0090] Figure 6 A schematic diagram of the target detection device based on vision and radar provided in this application;
[0091] Figure 7 This is a structural diagram of the vehicle provided in this application.
[0092] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0093] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0094] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0095] The main reason for the low accuracy of target object perception in existing technologies is:
[0096] 1. In the data fusion stage, the lateral and longitudinal position information relies entirely on a single sensor source, without considering the potentially large errors that may exist in a single source under actual road conditions. For example, in rainy or bright light conditions, the lateral position detection of machine vision is easily interfered with, and the longitudinal distance measurement of millimeter-wave radar in complex multi-target scenarios may also be inaccurate.
[0097] 2. Regarding target type judgment, the solution completely trusts the machine vision module, ignoring the common false detection and false detection problems of machine vision in situations such as occlusion and sudden changes in light. It may misjudge pedestrians as obstacles or miss small non-motorized vehicles.
[0098] In summary, existing technologies suffer from low accuracy in perceiving target objects.
[0099] Based on the aforementioned technical problems, the technical concept of this application is as follows: In analyzing the problems existing in the prior art, the inventors recognized that relying solely on data from a single sensor in a truck right-turn scenario is prone to significant errors due to environmental interference, and that completely trusting visual perception in target type judgment can easily lead to false detections and missed detections. Based on this, the inventors conceived of first simultaneously acquiring multi-source target object data from cameras and radar, and then generating more reliable third target object data through intra-frame fusion to complement the advantages of the sensors; subsequently, in inter-frame fusion, combining the historical third target object data generated at the previous moment, the fourth target object data is determined, thereby dynamically correcting the target state during tracking, reducing the unreliability of perception caused by single sensor errors or visual misjudgments, and improving the accuracy and robustness of target object perception.
[0100] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments.
[0101] The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0102] Figure 1 A flowchart illustrating the vision- and radar-based target detection method provided in this application. Figure 1 .like Figure 1 As shown, the method includes:
[0103] S11. Obtain the first target object data and the second target object data of the vehicle driving environment at the current moment.
[0104] In this step, it is necessary to simultaneously acquire the perception results from different sensors to provide a foundation for subsequent data fusion.
[0105] The data for the first target object was acquired by a camera, while the data for the second target object was acquired by radar.
[0106] In a specific scenario, the vehicle can be a truck, and the driving environment can be the area to the right of the vehicle.
[0107] In one possible implementation, multi-frame video data acquired by a camera, multi-frame radar data acquired by radar, and the extrinsic parameter matrix between the camera and radar are acquired. Then, the extrinsic parameter matrix is used to uniformly transform the multi-frame video data and multi-frame radar data into the vehicle coordinate system, and further timestamp alignment and synchronization are performed to obtain target multi-frame video data and target multi-frame radar data. Subsequently, for the current moment, the first target object data is identified from the target multi-frame video data, and the second target object data is identified from the target multi-frame radar data.
[0108] Specifically, the method for aligning and synchronizing timestamps of multi-frame video data and multi-frame radar data is as follows: the two frames with the closest time distance in the multi-frame video data and multi-frame radar data are matched. For the two successfully matched frames, the position of the target object in the other frame is corrected according to the speed of the target object, based on one of the frames.
[0109] Furthermore, historical frame data (data preceding the current frame) of the target multi-frame video data can be used to correct the initially identified target object data, and historical frame data (data preceding the current frame) of the target multi-frame radar data can also be used to correct the initially identified target object data, to ensure the accuracy of the final first and second target object data. It should be understood that the specific correction process in this method will be explained later. Figure 4 The embodiments shown are described in detail, and will not be repeated here.
[0110] Optionally, the target object data includes the target feature values of the target object and the confidence level corresponding to each target feature value. The target features include at least one of the following: location, type, size, speed, and orientation.
[0111] It should be understood that the confidence level corresponding to each target feature value in the first target object data can be the confidence level associated with target identification of the video frame corresponding to the current moment from multiple frames of video data using existing technology; similarly, the confidence level corresponding to each target feature value in the second target object data can be the confidence level associated with target identification of the video frame corresponding to the current moment from multiple frames of radar data using existing technology.
[0112] The target object data includes confidence levels for core features such as location, type, size, speed, and orientation, making each feature value have assessable credibility. This lays the foundation for subsequent data processing and decision-making processes, further enhancing the reliability of the perception results.
[0113] It is understandable that the first target object data, the second target object data, the third target object data, and the fourth target object data all contain the target feature values of the target objects and the confidence levels corresponding to each target feature value.
[0114] For example, the target object can be a moving traffic participant, such as pedestrians, two-wheeled vehicles (e.g., bicycles, electric vehicles, and motorcycles), small vehicles (e.g., cars, vans), and large vehicles (e.g., trucks, construction vehicles, tricycles, buses, and coaches).
[0115] S12. Merge the first target object data and the second target object data to generate the third target object data;
[0116] In this step, to achieve complementary advantages between sensors, the first target object data and the second target object data are fused within the frame to generate a set of more accurate and reliable third target object data.
[0117] In one possible implementation, the first and second target object data are first associated to match the same target object from both the camera and radar. Target objects from the two sources at the current moment are defined as the first set and the second set, respectively. The distance, intersection-over-union (IOU), size difference, and velocity difference between the two source target objects are used as association indicators, and corresponding thresholds are set. Through traversal searching and judgment, target objects that meet all association threshold conditions and have the smallest distance are associated to form matching pairs; unmatched target objects are retained as independent target objects. Subsequently, the associated target object pairs are fused, and the confidence levels of each target feature value are updated to finally generate the third target object data.
[0118] It should be understood that the specific implementation method of this step can be found in [reference needed]. Figure 2 The embodiments shown are not described in detail here.
[0119] S13. Determine the fourth target object data based on the third target object data and the historical third target object data generated in the previous moment.
[0120] In this step, to achieve continuous and stable tracking of the target object, the third target object data calculated at the current moment can be corrected by using historical third target object data to determine the fourth target object data, thereby ensuring the continuity of the target object in time sequence.
[0121] Among them, the historical third target object data generated in the previous moment is the third target object data generated in the previous moment by executing S11-S12.
[0122] In one possible implementation, the predicted location of the target object can be determined based on historical third-target object data. Then, the fourth target object data is determined based on the predicted location of the target object, historical third-target object data, and the third-target object data.
[0123] It should be understood that this implementation method will Figure 3 The embodiments shown are described in detail, and will not be repeated here.
[0124] The target detection method based on vision and radar provided in this application first acquires first target object data and second target object data of the vehicle driving environment at the current moment. The first and second target object data are then fused to generate third target object data. Finally, fourth target object data is determined based on the third target object data and historical third target object data generated at the previous moment. The first target object data is acquired by a camera, and the second target object data is acquired by radar. This technical solution simultaneously acquires the first and second target object data and fuses them to generate the third target object data, effectively combining the advantages of different sensors. Furthermore, the fourth target object data is determined based on the third target object data and historical third target object data generated at the previous moment. This enables dynamic state correction during target tracking, overcoming the perception reliability problem caused by single sensor errors or visual misjudgments, and significantly improving the overall target object perception accuracy and robustness.
[0125] Figure 2 A flowchart illustrating the vision- and radar-based target detection method provided in this application. Figure 2 ,like Figure 2 As shown, S12 includes:
[0126] S21. Construct a first set based on the first target object contained in the first target object data.
[0127] For example, the first set can be set A.
[0128] S22. Construct a second set based on the second target object contained in the second target object data.
[0129] For example, the second set can be set B.
[0130] S23. For any first target object in the first set, calculate the value of the target association index between the first target object and each second target object in the second set.
[0131] The target correlation index includes at least one of the following: distance difference, intersection-union ratio (IOU), size difference, and velocity difference.
[0132] In one possible implementation, for each first target object in set A ( Search in set B for... The distance difference between them is less than the preset distance ( The second target object forms a set to be matched. Then, the matching sets are processed in ascending order of distance difference. The second target objects contained therein are arranged. Then, the set to be matched is... The second target object in the calculation and The IOU, size difference, and speed difference between them.
[0133] S24. The second target object whose value of the target association index in the second set meets the corresponding preset conditions is determined as the candidate target object.
[0134] In one possible implementation, according to the set to be matched The arrangement order (from smallest to largest distance difference) is based on the preset IOU (Interval of Union). ), preset size difference ( ) and preset speed difference ( ), the set to be matched The second target object, whose IOU is greater than the preset IOU, whose size difference is less than the preset size difference, and whose speed difference is less than the preset speed difference, is determined as the candidate target object.
[0135] Where i represents the i-th first target object in set A, and j represents The j-th second target object in the process.
[0136] Furthermore, the preset condition is that among the second set of second target objects whose distance difference from the first target object is less than a preset distance difference, whose IOU is greater than a preset IOU, whose size difference is less than a preset size difference, and whose speed difference is less than a preset speed difference, the first target object has the smallest distance difference.
[0137] S25. Associate the first target object with the candidate target object.
[0138] In one implementation, the first target object can be directly associated with the candidate target object.
[0139] In another possible implementation, before associating the first target object with the candidate target object, it is first determined whether the candidate target object already has an association with other first target objects. If the candidate target object already has an association with other first target objects, then the first IOU between the candidate target object and the first target object, and the second IOU between the candidate target object and the other first target object are determined. If the first IOU is greater than the second IOU, then the association between the candidate target object and the other first target object is terminated, and the candidate target object is then associated with the first target object. Conversely, if the first IOU is less than or equal to the second IOU, then the association between the candidate target object and the other first target object is maintained, and no association is established between candidate target objects.
[0140] In this implementation, by determining whether the candidate target object already has an association during the association process, and updating the association when the first IOU is greater than the second IOU, the many-to-one target contention problem is effectively solved, ensuring a stable association between the target object and the best candidate, thereby improving the accuracy and reliability of the third target object data.
[0141] Optionally, determining whether the candidate target object already has an association with other first target objects can be done by determining whether the candidate target object already exists in the set. Implemented in [the context]. Among them, the set... Used to store the second target object in set B that has established a relationship with the first target object in set A, its initial value is empty.
[0142] Furthermore, after traversing all the first target objects in set A, for the first target objects that do not have any association relationship, we can look for the second target objects that do not have any association relationship (those that do not exist in the set A). The second target object (in the first target object) is selected as the candidate target object of the first target object, and the two are associated.
[0143] S26. Update the third and fourth sets.
[0144] The third set is used to store target object pairs that have a relationship, and the fourth set is used to store the first target object that does not have a relationship, and / or the second target object that does not have a relationship.
[0145] After establishing the relationship between the first target object and the second target object, the target object pair with the relationship is added to the third set. In order to update the third set; after traversing all the first target objects, add all first target objects that have no relationship and / or all second target objects that have no relationship to the fourth set. Stored in ).
[0146] S27. Generate the third target object data based on the third set and the fourth set.
[0147] For the third set, based on the related target object pairs contained in the third set, target feature value pairs of the target object pairs are determined from the first target object data and the second target object data. For each target object pair, the target feature value pairs are fused to generate the third target object data.
[0148] Next, we will explain the specific fusion method for each feature in terms of location, type, size, speed, and orientation.
[0149] 1. Location
[0150] The vehicle driving environment is divided into three zones: Zone 1, Zone 2, and Zone 3. For cameras, the distance between the camera and the vehicle is less than 10 meters in Zone 1, 10-15 meters in Zone 2, and greater than 15 meters in Zone 3. For radar, the distance is less than 10 meters in Zone 1, 10-20 meters in Zone 2, and greater than 20 meters in Zone 3.
[0151] Determine whether the target object pair meets the fusion condition, where the fusion condition is: the distance difference is less than the preset fusion distance difference. The ratio of the maximum length to the minimum length in the target object pair is less than the first ratio (e.g., 2:1), and the ratio of the maximum width to the minimum width in the target object pair is less than the second ratio (e.g., 2:1). After determining that the target object pair meets the fusion conditions, the position pairs in the target object pair are weighted and summed according to the weights pre-assigned to the camera and radar to obtain the fused position, which is the position in the third target object data.
[0152] In this process, if the target object pair does not meet the fusion conditions, the position of the first target object in the target object pair is taken as the fused position, which is the position in the third target object data.
[0153] Specifically, if the type of the fused target object pair is a non-four-wheeled vehicle, the preset fusion distance difference is the diagonal length of the preset size corresponding to the non-four-wheeled vehicle; if the type of the fused target object pair is a four-wheeled vehicle, the preset fusion distance difference is half the diagonal length of the preset size corresponding to the non-four-wheeled vehicle.
[0154] 2. Type
[0155] The type of the second target object in the target object pair is determined as the fused type when the following conditions are met; otherwise, the type of the first target object in the target object pair is determined as the fused type.
[0156] (1) When the target object is in the first region or the second region, the type of the first target object is different from the type of the previous time, and the number of consecutive types of the second target object is greater than the first preset number, and the type of the second target object is the same as the type of the first target object at the current time or the previous time.
[0157] In other words, when the target object is in the first or second region, if the type of the first target object captured by the camera at the current moment is different from the type of the first target object captured at the previous moment, and the type of the second target object captured by the radar is consistent for a long time, and the type of the second target object captured by the radar is the same as the type of the first target object captured by the camera at the current or previous moment, then the radar result is more trusted.
[0158] (2) When the target object pair is in the third region, when the number of type changes of the first target object in the target object pair is greater than the second preset number, and the number of consecutive types of the second target object in the target object pair is greater than the first preset number, and the preset size corresponding to the type of the second target object in the target object pair is consistent with the identified size.
[0159] In other words, when the target object is in the third region, if the type of the first target object captured by the camera changes frequently over a period of time, but the type of the second target object captured by the radar remains consistent for a long time, and the size of the second target object captured by the radar is consistent with the size corresponding to that type, then the radar's identification result is trusted.
[0160] (3) When the number of type changes of the first target object in the target object pair is greater than the second preset number and the confidence level is less than the first preset confidence level, and the number of consecutive types of the second target object in the target object pair is greater than the first preset number, and the confidence level of the second target object in the target object pair is greater than the second preset confidence level.
[0161] In other words, if the type of the first target object captured by the camera changes frequently over a period of time and the reliability of the type is not high, then if the type of the second target object captured by the radar remains consistent over a long period of time and the reliability of the type is high, then the radar's identification result can be trusted.
[0162] 3. Size
[0163] The corresponding preset size is determined based on the type after fusion, and the size closest to the preset size in the target object pair is determined as the size after fusion.
[0164] 4. Speed
[0165] The velocity of the second target object in the target object pair is determined as the fused velocity.
[0166] 5. Orientation
[0167] The absolute velocity of the target object pair is determined. If the absolute velocity is greater than a preset velocity, the velocity direction value of the target object pair is calculated and determined as the orientation after fusion. If the absolute velocity is less than or equal to the preset velocity, the orientation of the first target object in the target object pair is determined as the orientation after fusion. It should be understood that the absolute velocity of the target object pair can be determined based on the velocity pair of the target object pair and the speed of the vehicle, which will not be elaborated here.
[0168] For the fourth set, based on the first target objects in the fourth set that have no relation to each other, the target feature values of the first target objects that have no relation to each other are obtained from the first target object data, and the third target object data is generated.
[0169] Based on the second target objects in the fourth set that have no relation to each other, the target feature values of the second target objects that have no relation to each other are obtained from the second target object data, and the third target object data is generated.
[0170] That is, for the first and second target objects in the fourth set, their own target feature values are determined as the fused target feature values.
[0171] Regarding the above fusion process, by performing data fusion on target object pairs with related relationships based on the third set, and retaining unmatched independent target objects based on the fourth set, it is ensured that all perceived target objects can be effectively processed and output, avoiding the loss of effective targets and significantly improving the integrity of the third target object data.
[0172] against Figure 2 The illustrated embodiment constructs a first set and a second set, calculates target correlation indicators between target objects, associates candidate target objects that meet preset conditions with the first target object, updates the third set and the fourth set, and finally generates third target object data. This mechanism achieves accurate association and data integration of the same target object from different sources, effectively improving the accuracy and data integrity of target matching in multi-sensor perception systems.
[0173] Figure 3 A flowchart illustrating the vision- and radar-based target detection method provided in this application. Figure 3 ,like Figure 3 As shown, S13 includes:
[0174] S31. Determine the predicted location of the target object based on the historical third target object data.
[0175] In this step, the position of the target object at the current time can be predicted based on the historical third target object data calculated at the previous time step, providing a benchmark for subsequent inter-frame association and fusion.
[0176] The predicted position of the target object at the current moment can be calculated using the following formula:
[0177]
[0178] in, The predicted position of the target object at the current moment; The merged position of the target object at the previous moment is the position contained in the historical third target object data; It is the time difference between the previous moment and the current moment. The velocity of the target object after fusion at the previous moment is the velocity contained in the historical third target object data.
[0179] S32. Determine the fourth target object data based on the predicted location of the target object, historical third target object data, and third target object data.
[0180] Next, we will explain the specific processing methods for each feature from five aspects: location, type, size, speed, and orientation.
[0181] 1. Location
[0182] A fusion weight is assigned to the predicted position and the fused position (the position in the third target object data). Then, a weighted sum is calculated based on the fusion weight, and the resulting value is determined as the final target position in the fourth target object data.
[0183] It should be understood that the merged location belongs to the third target object data.
[0184] Among them, the farther the target object is from the vehicle, the higher the fusion weight of the predicted location ( The larger the value, the greater the fusion weight of the merged position. The smaller the value of the fusion weight, the closer the target object is to the vehicle; conversely, the closer the target object is to the vehicle, the smaller the fusion weight of the predicted position, and the larger the fusion weight of the fused position. Specifically, the fusion weight can be determined in the following ways:
[0185] a) When the distance between the target object and the vehicle is less than or equal to 5 meters, , ;
[0186] b) When the distance between the target object and the vehicle is greater than 5 meters and less than or equal to 15 meters, , ;
[0187] c) When the distance between the target object and the vehicle is greater than 15 meters, , .
[0188] Regarding the confidence level of a location, the confidence level corresponding to the area where the target object is located can be determined as the final target confidence level.
[0189] For example, if the target object is in region 1, the confidence level of the location is set to 0.9; if the target object is in region 2, the confidence level of the location is set to 0.8; and if the target object is in region 3, the confidence level of the location is set to 0.6.
[0190] Furthermore, before determining the confidence level corresponding to the region where the target object is located as the final target confidence level, the difference between the predicted position and the fused position can be confirmed. If the difference is greater than the preset size corresponding to the target object, the confidence level corresponding to the region where the target object is located is multiplied by a preset coefficient (e.g., 0.8), and the result of the multiplication is determined as the final target confidence level; conversely, if the difference is less than or equal to the preset size corresponding to the target object, the confidence level corresponding to the region where the target object is located is directly determined as the final target confidence level.
[0191] 2. Type
[0192] The fused type of the target object at the previous time step is determined from the historical third target object data, and it is then determined whether the fused type at the previous time step is consistent with the fused type at the current time step. If they are inconsistent, the number of consecutive types and the number of type changes of the target object are further determined from the historical third target object data at previous time steps. If the number of consecutive types is greater than 10, and the fused type at the previous time step is consistent with the type corresponding to the historical first target object data at the previous time step (or the type corresponding to the historical second target object data), then the fused type at the previous time step is determined as the target type, and the target confidence score is calculated according to the following formula:
[0193]
[0194] in, The target confidence level is less than or equal to 1. The confidence level corresponds to the fused type from the previous time step. If the fused type from the previous time step is consistent with the historical first target object data from the previous time step, then the historical first target object data from the previous time step is... Determined as Conversely, if the type of the fused data from the previous time step is consistent with the historical second target object data from the previous time step, then the historical second target object data from the previous time step will be... Determined as .
[0195] In cases other than those described above, the fused type at the current moment is determined as the final target type, and the target confidence corresponding to the type is calculated using the following formula:
[0196]
[0197] in, This represents the confidence level of the merged type. It should be understood that if the merged type at the current moment is consistent with the data of the first target object at the current moment, then the data of the first target object at the current moment will be... Determined as Conversely, if the type of the fused object at the current moment is consistent with the data of the second target object at the current moment, then the data of the second target object at the current moment will be... Determined as .
[0198] 3. Size
[0199] The current fused size is determined as the target size, and a corresponding preset size is determined based on the target type. The fused size from the previous time step is determined from the historical third target object data, and a first difference between it and the preset size corresponding to the target type is calculated. A second difference between the current fused size and the preset size corresponding to the target type is also calculated. If both the first and second differences are less than the preset difference, the target confidence level of the size is calculated as follows:
[0200]
[0201] in, The target confidence level is less than or equal to 1. This represents the confidence level of the fused size at the current time step. Similar to the type, if the fused size at the current time step is consistent with the data of the first target object at the current time step, then the confidence level of the first target object data at the current time step is determined as... Conversely, if the fused size at the current moment is consistent with the data of the second target object at the current moment, then the confidence level of the data of the second target object at the current moment is determined as... .
[0202] Conversely, if the first difference and / or the second difference are not less than the preset difference, the target confidence level of the size is calculated as follows:
[0203]
[0204] 4. Speed
[0205] The fused velocity at the current moment is determined as the target velocity, and the target confidence of the position is determined as the target confidence of the velocity.
[0206] 5. Orientation
[0207] Based on the historical third target object data, determine the fused orientation of the previous moment, calculate the angle difference between the fused orientation of the previous moment and the fused orientation of the current moment. If the angle difference is less than the preset angle difference, then the fused orientation of the current moment is determined as the target orientation, and the target confidence of the position is determined as the target confidence of the velocity.
[0208] Conversely, if the angle difference is greater than or equal to the preset angle difference, weights are assigned to the fused orientation of the previous moment and the fused orientation of the current moment (e.g., both are 0.5), and the weighted sum is determined as the target orientation. The preset ratio (0.7) is multiplied by the target confidence of the position, and the result is determined as the target confidence of the target orientation.
[0209] It should be understood that the fourth target object data includes target location, target type, target size, target speed, and target orientation, as well as the target confidence level corresponding to each of the target location, target type, target size, target speed, and target orientation.
[0210] In the above embodiments, by further correcting the third target object data based on the historical third target object generated in the previous moment, the target state can be dynamically corrected during the tracking process, effectively smoothing the trajectory and suppressing measurement noise, thereby improving the stability and accuracy of tracking.
[0211] Figure 4 A flowchart illustrating the vision- and radar-based target detection method provided in this application. Figure 4 ,like Figure 4 As shown, S11 includes:
[0212] S41. Obtain the initial target object data of the vehicle driving environment at the current moment.
[0213] In this step, the initial target object data includes initial first target object data and / or initial second target object data. The initial first target object data is acquired by a camera, and the initial second target object data is acquired by radar.
[0214] Referring to the relevant content in S11, for the current moment, the initial first target object data is identified from the target multi-frame video data, and the initial second target object data is identified from the target multi-frame radar data.
[0215] S42. For any target object in the initial target object data, if the target object in the initial target object data is of type 1 and the target object in the historical initial target object data of the previous time step is of type 2, then determine the number of consecutive types of the target object in the previous time step based on the historical initial target object data of the previous time step.
[0216] In this step, a type jump occurs when the current output type (first type) of the target object is detected to be inconsistent with the historical type (second type) of the previous time step. To determine whether this jump is a misjudgment caused by noise interference, it is necessary to determine the type stability of the target object in previous time steps.
[0217] It should be understood that during vehicle operation, the data for the first and second target objects at the current moment are generated in real time. Therefore, the historical initial target object data from previous moments are the first and second target object data processed and stored at those previous moments.
[0218] In one possible implementation, starting from the previous moment and tracing backwards, if the target object at that moment is of type 2, the type continuity number is initialized to 1, and tracing continues backwards; if the target object at the previous moment is still of type 2, the continuity number is incremented by 1; and so on, until the target object at a previous moment is no longer of type 2. The accumulated value obtained at this point is the final type continuity number.
[0219] The consecutive type number represents the number of times the target object has been continuously identified as the second type in the previous time step, where the previous time step includes the last time step.
[0220] S43. If the number of consecutive types is greater than the preset number of times, then the first type in the initial target object data is modified to the second type, and the target object data is generated.
[0221] In this step, if the target object has been stable in the second type for a long time in the previous moment, then the sudden change to the first type in the current frame is considered to be unreliable sensor noise and should be corrected.
[0222] When the initial target object data is the initial first target object data, if the number of consecutive types is greater than 10, the first type in the initial target object data is modified to the second type, and the target object data is generated.
[0223] Furthermore, if the number of consecutive types is greater than 10 and less than 30, the type in the initial target object data will be corrected to the second type in the next 5 frames; if the number of consecutive types is greater than or equal to 30, the type in the initial target object data will be corrected to the second type in the next 10 frames.
[0224] When the initial target object data is the initial second target object data, if the number of consecutive types is greater than 20, the first type in the initial target object data will be corrected to the second type to generate the target object data.
[0225] Furthermore, if the number of consecutive types is greater than 20, the type in the initial target object data will be corrected to the second type in the following 5 frames.
[0226] For example, the preset number of times can be set according to the sensor type; for instance, it can be set to 10 for a camera and 20 for radar. It should be understood that the preset number of times can be preset based on empirical or experimental values, and this application embodiment does not impose specific limitations on this.
[0227] It should be understood that the target object data corresponding to the initial first target object data is the first target object data, and the target object data corresponding to the initial second target object data is the second target object data.
[0228] S44. If the number of consecutive types is less than or equal to the preset number of times, then the initial target object data is determined as the target object data.
[0229] In this step, if the target object is only briefly identified as the second type, it indicates that the historical type stability of the target object is insufficient. The type change in the current frame may be a valid change or a transient disturbance. Therefore, it is necessary to maintain the current perception result. Thus, the initial target object data can be directly determined as the target object data.
[0230] In the above technical solution, by statistically analyzing the number of consecutive types of the target object in historical frames, when the historical types are stable, the target object in the initial target object data is corrected to suppress instantaneous type jumps caused by environmental interference, thereby improving the continuity and stability of the output type. When the historical types are unstable, the perception result of the current frame is trusted to ensure the responsiveness to changes in the actual type. This technical solution significantly improves the reliability of the output type of a single sensor in complex scenarios such as occlusion and sudden changes in lighting, providing higher quality and more reliable input data for subsequent multi-sensor data fusion.
[0231] In practical applications, the number of type changes of the target object can be counted within a preset continuous historical time period starting from the current moment. Then, based on the number of type changes, the number of consecutive types, and the type of the target object within the preset continuous time period, the initial confidence level in the initial target object data is updated to generate the target object data.
[0232] It should be understood that the preset continuous duration can be pre-set based on empirical or experimental values, and this application embodiment does not impose specific limitations on this.
[0233] Alternatively, the confidence level of the type in the initial target object data can be corrected using the following formula:
[0234]
[0235]
[0236] in, Update coefficients for type confidence. For type confidence update factor, For consecutive numbers of type, For type variation number, The confidence level of the target object in the target object data. The initial confidence level for the target object.
[0237] Furthermore, It is determined based on the type of the target object within a preset continuous time period. Specifically, if the type of the target object within the preset continuous time period is all four-wheeled vehicles (small vehicles + large vehicles) or all non-four-wheeled vehicles (pedestrians + two-wheeled vehicles), then it is determined... The value is 1.1; conversely, if the type of the target object within the preset continuous time period includes both four-wheeled vehicles and non-four-wheeled vehicles, then it is determined that... It is 1.05.
[0238] In the above embodiments, the initial confidence level in the initial target object data is updated by comprehensively considering the number of type changes and the number of type continuities of the target object. When the number of type continuities is high and the number of type changes is low, the updated confidence level is correspondingly improved; when type changes are frequent, the updated confidence level is correspondingly reduced, thereby improving the accuracy of the confidence level and laying the foundation for subsequent data fusion.
[0239] Optionally, in some embodiments, for a pair of target objects that were associated in the previous time, if their distance at the current time is less than a preset split distance threshold, the association is maintained; if their distance is greater than or equal to the split distance threshold, the association between the first target object and the second target object in the target object pair is canceled so that they can be associated with other target objects.
[0240] Based on the target detection methods based on vision and radar provided in any of the above embodiments, the following will explain their implementation through several examples.
[0241] Figure 5 A flowchart illustrating the vision- and radar-based target detection method provided in this application. Figure 5 ,like Figure 5 As shown, the method includes:
[0242] S51: Acquire multi-frame video data using a camera.
[0243] S52: Collects multiple frames of radar data via radar.
[0244] S53. Preprocess multi-frame video data and multi-frame radar data.
[0245] The preprocessing includes timestamp alignment, synchronization, and category correction. Category correction can be found in [reference needed]. Figure 4 The embodiments shown are not described in detail here.
[0246] S54, intra-frame correlation and fusion.
[0247] This step includes intra-frame target object association and intra-frame feature value fusion.
[0248] It should be understood that the specific implementation method in this step can be referred to the relevant content in S12, and will not be repeated here.
[0249] S55, inter-frame tracking and fusion.
[0250] This step includes inter-frame target object location prediction, inter-frame target object association, and inter-frame feature value fusion.
[0251] It should be understood that the specific implementation method in this step can be referred to the relevant content in S13, and will not be repeated here.
[0252] S56, Lifecycle Management.
[0253] For the first target object that has no association at the current moment, and / or the second target object that has no association, its state is continuously saved for the next 5 consecutive frames. If no association is established for it within the next 5 consecutive frames, the target object will no longer be output after the 5th frame.
[0254] Based on any of the above embodiments, the target detection method based on vision and radar provided in this application has the following technical effects:
[0255] (1) In the data preprocessing stage, by calculating the number of consecutive categories and the number of category changes of the target, and correcting the category that jumps based on the number of consecutive categories, the problem of unstable category output caused by environmental interference of a single sensor is effectively suppressed.
[0256] (2) In the process of associating target objects, the distance between target objects is used as the main association index, and IOU, size difference and speed difference are combined as constraints. Different association thresholds are set for different types of targets, so as to achieve accurate matching of the same target object among multiple sensors.
[0257] (3) During the intra-frame data fusion process, by dividing the region and allocating fusion weights according to the region where the target is located, and setting fusion distance conditions and fusion size conditions, the radar position is weighted and fused only when the conditions are met; otherwise, the camera position is selected first, thus avoiding errors caused by the mismatch between the radar target and the category size. In the category fusion process, the continuity of the region and category and the confidence level are combined to improve the accuracy of category judgment. The orientation fusion determines the orientation based on the absolute velocity of the target, which effectively enhances the stability of the orientation output.
[0258] (4) In the inter-frame target object association, the predicted position is matched with the fused position at the current time. Based on the association index constraint, the target that has been fused in the previous frame is further split. If the distance of its source target in the current frame does not exceed the split threshold, the fusion state is maintained to improve the stability of target ID in complex scenarios.
[0259] (5) During the inter-frame data fusion process, based on the differences in attributes between regions and sensors and historical attribute values, the final output of each target feature value is determined in a rule-based manner, and the confidence level reflecting the accuracy and stability of the attributes is output synchronously, so as to provide downstream modules with more complete and reliable perception results.
[0260] Figure 6 A schematic diagram of the target detection device based on vision and radar provided in this application is shown below. Figure 6 As shown, the target detection device 60 based on vision and radar provided in this embodiment includes:
[0261] The acquisition module 61 is used to acquire the first target object data and the second target object data of the vehicle driving environment at the current moment. The first target object data is collected by the camera and the second target object data is collected by the radar.
[0262] The fusion module 62 is used to fuse the first target object data and the second target object data to generate the third target object data;
[0263] The determination module 63 is used to determine the fourth target object data based on the third target object data and the historical third target object data generated in the previous moment.
[0264] In one possible implementation, the target object data includes target feature values of the target object and the confidence level corresponding to each target feature value. The target features include at least one of the following: location, type, size, velocity, and orientation.
[0265] In one possible implementation, the fusion module 62 is specifically used for:
[0266] Construct a first set based on the first target object contained in the first target object data;
[0267] Construct a second set based on the second target object contained in the second target object data;
[0268] For any first target object in the first set, calculate the value of the target association index between the first target object and each second target object in the second set. The target association index includes at least one of distance difference, IOU, size difference and velocity difference.
[0269] The second target object whose value of the target association index in the second set meets the corresponding preset condition is determined as the candidate target object;
[0270] Associate the first target object with the candidate target objects;
[0271] Update the third set and the fourth set. The third set is used to store target object pairs that have a relationship, and the fourth set is used to store the first target object that does not have a relationship, and / or the second target object that does not have a relationship.
[0272] Generate the third target object data based on the third and fourth sets.
[0273] In one possible implementation, the fusion module 62 is specifically used for:
[0274] If the candidate target object is already associated with other first target objects, then determine the first IOU between the candidate target object and the first target object, and the second IOU between the candidate target object and other first target objects;
[0275] If the first IOU is greater than the second IOU, then the association between the candidate target object and other first target objects is removed, and the candidate target object is associated with the first target object.
[0276] In one possible implementation, the fusion module 62 is specifically used for:
[0277] Based on the target object pairs with related relationships contained in the third set, determine the target feature value pairs of the target object pairs from the first target object data and the second target object data;
[0278] For each pair of target objects, the target feature values are fused to generate third target object data;
[0279] Based on the first target objects in the fourth set that have no relation to each other, the target feature values of the first target objects that have no relation to each other are obtained from the first target object data, and the third target object data is generated.
[0280] Based on the second target objects in the fourth set that have no relation to each other, the target feature values of the second target objects that have no relation to each other are obtained from the second target object data, and the third target object data is generated.
[0281] In one possible implementation, module 63 is specifically used for:
[0282] Based on historical data of the third target object, the predicted location of the target object is determined;
[0283] Based on the predicted location of the target object, historical data of the third target object, and data of the third target object, determine the data of the fourth target object.
[0284] In one possible implementation, module 61 is specifically used for:
[0285] Acquire initial target object data of the vehicle driving environment at the current moment. The initial target object data includes initial first target object data and / or initial second target object data. The initial first target object data is acquired by a camera, and the initial second target object data is acquired by radar.
[0286] For any target object in the initial target object data, if the target object in the initial target object data is of type 1 and the target object in the historical initial target object data of the previous time is of type 2, then the type continuity number of the target object in the previous time is determined based on the historical initial target object data of the previous time. The type continuity number is used to represent the number of times the target object was continuously identified as type 2 in the previous time. The previous time includes the previous time.
[0287] If the number of consecutive types exceeds the preset number, the first type in the initial target object data will be corrected to the second type, and the target object data will be generated.
[0288] If the number of consecutive types is less than or equal to the preset number of times, then the initial target object data is determined as the target object data. Specifically, the target object data corresponding to the initial first target object data is the first target object data, and the target object data corresponding to the initial second target object data is the second target object data.
[0289] In one possible implementation, the vision- and radar-based target detection device 60 further includes an update module for:
[0290] Within a time period of a preset continuous duration preceding the current moment, determine the number of type changes of the target object;
[0291] Based on the number of type changes, the number of consecutive types of the target object, and the type of the target object within a preset continuous time period, the initial confidence level in the initial target object data is updated to generate target object data.
[0292] The target detection device based on vision and radar provided in this embodiment can execute the methods provided in the above method embodiments. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0293] Figure 7 This is a structural diagram of the vehicle provided in this application. Figure 7 As shown, the vehicle 70 provided in this embodiment includes: a vehicle body 71, at least one processor 72, and a memory 73. Optionally, the vehicle 70 also includes a communication component 74. The processor 72, memory 73, and communication component 74 are connected via a bus 75.
[0294] In a specific implementation, at least one processor 72 executes computer execution instructions stored in memory 73, causing at least one processor 72 to perform the above-described method.
[0295] The specific implementation process of processor 72 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0296] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0297] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0298] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0299] This application also provides an electronic device, which can be a terminal device, such as a laptop, desktop computer, tablet computer, or vehicle terminal, or a server. In practical applications, whether the electronic device is a terminal device or a server can be determined according to the actual situation, and no specific limitation is imposed.
[0300] Specifically, the electronic device includes a processor and a memory, which are connected via a bus.
[0301] In practice, the processor executes computer execution instructions stored in memory, causing at least one processor to perform the above-described method.
[0302] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0303] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0304] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0305] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an application-specific integrated circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0306] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0307] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0308] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0309] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0310] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0311] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A vision and radar based target detection method, characterized in that, The method comprises: acquiring first target object data and second target object data of a driving environment of a vehicle at a current time, the first target object data being acquired by a camera, and the second target object data being acquired by a radar; fusing the first target object data and the second target object data to generate third target object data; determining fourth target object data according to the third target object data and historical third target object data generated at a previous time.
2. The method of claim 1, wherein, The target object data comprises target feature values of target objects and confidence degrees corresponding to the target feature values, and the target features comprise at least one of the following: position, type, size, speed and orientation.
3. The method according to claim 1 or 2, characterized in that, The fusing of the first target object data and the second target object data to generate third target object data comprises: constructing a first set according to first target objects contained in the first target object data; constructing a second set according to second target objects contained in the second target object data; calculating a value of a target association index between any first target object in the first set and each second target object in the second set, the target association index comprising at least one of a distance difference, an intersection over union (IOU), a size difference and a speed difference; determining a second target object in the second set as a candidate target object if the value of the target association index of the second target object satisfies a corresponding preset condition; associating the first target object with the candidate target object; updating a third set and a fourth set, the third set being used to store target object pairs in an association relationship, and the fourth set storing first target objects not in an association relationship and / or second target objects not in an association relationship; generating the third target object data according to the third set and the fourth set.
4. The method of claim 3, wherein, The associating of the first target object with the candidate target object comprises: if the candidate target object is already in an association relationship with other first target objects, determining a first IOU between the candidate target object and the first target object and a second IOU between the candidate target object and the other first target objects; if the first IOU is greater than the second IOU, disassociating the candidate target object from the other first target objects and associating the candidate target object with the first target object.
5. The method of claim 3, wherein, The generating of the third target object data according to the third set and the fourth set comprises: determining target feature value pairs of target object pairs in an association relationship from the first target object data and the second target object data according to the target object pairs contained in the third set; fusing the target feature value pairs for each target object pair to generate the third target object data; acquiring target feature values of first target objects not in an association relationship from the first target object data according to the first target objects not in an association relationship contained in the fourth set to generate the third target object data; According to the second target object without the existing correlation relationship in the fourth set, a target feature value of the second target object without the existing correlation relationship is obtained from the second target object data, and the third target object data is generated.
6. The method of any one of claims 1, 2, 4, or 5, wherein, The fourth target object data is determined according to the third target object data and historical third target object data generated at a previous moment. A predicted position of the target object is determined according to the historical third target object data. The fourth target object data is determined according to the predicted position of the target object, the historical third target object data and the third target object data.
7. The method of any one of claims 1, 2, 4, or 5, wherein, The first target object data and the second target object data of the vehicle driving environment at the current moment are obtained, including: The initial target object data of the vehicle driving environment at the current moment is obtained, the initial target object data including initial first target object data and / or initial second target object data, the initial first target object data being collected by the camera and the initial second target object data being collected by the radar; For any target object in the initial target object data, if the target object in the initial target object data is of a first type and the target object in historical initial target object data at a previous moment is of a second type, a type continuous number of the target object at the previous moment is determined according to historical initial target object data at the previous moment, the type continuous number representing a number of times that the target object is continuously identified as the second type at the previous moment, the previous moment including the previous moment; If the type continuous number is greater than a preset number of times, the first type in the initial target object data is corrected to the second type to generate target object data; If the type continuous number is less than or equal to the preset number of times, the initial target object data is determined as target object data; wherein the target object data corresponding to the initial first target object data is first target object data, and the target object data corresponding to the initial second target object data is second target object data.
8. The method of claim 7, wherein, The method further includes: A type change number of the target object is determined within a time period of a preset continuous duration from the current moment; An initial confidence in the initial target object data is updated according to the type change number of the target object, the type continuous number and a type corresponding to the target object within the preset continuous duration to generate target object data.
9. A vision and radar based target detection apparatus, characterized by, It includes: An acquisition module is configured to obtain first target object data and second target object data of a vehicle driving environment at a current moment, the first target object data being collected by a camera and the second target object data being collected by a radar; A fusion module is configured to fuse the first target object data and the second target object data to generate third target object data; A determination module is configured to determine fourth target object data according to the third target object data and historical third target object data generated at a previous moment.
10. A vehicle characterized by comprising: It includes: A vehicle body, a memory and a processor; The memory stores computer execution instructions; The processor executes computer-executable instructions stored in the memory such that the processor performs the method of any of claims 1-8.
11. A computer readable storage medium, characterized in that, The computer-readable storage medium has stored therein computer-executable instructions that, when executed by a processor, perform the method of any of claims 1-8.
12. A computer program product, characterised in that, A computer program that, when executed by a processor, performs the method of any of claims 1-8.