A method, apparatus, device and medium for extracting driving scenes
By filtering and matching candidate targets in keyframes during autonomous driving scene generation and utilizing threshold features, the problem of abrupt changes in multi-target tracking results is solved, improving the accuracy and efficiency of scene extraction while reducing computational complexity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-11
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies suffer from jumps in multi-target tracking results during autonomous driving scenario generation, which affects the progress of scenario library construction and the efficiency of scenario generation. Furthermore, they are computationally complex and have high false detection and false negative rates.
By identifying candidate targets in keyframes, and utilizing vehicle driving data frames and relative driving data frames of environmental targets, combined with first and second threshold features, target frames are filtered and matched to reduce the likelihood of target identification ID jumps and improve the accuracy of scene extraction.
It achieves accurate tracking of scene targets, reduces the false detection rate and false negative rate of driving scene extraction, and reduces computational complexity.
Smart Images

Figure CN115601685B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technology, and in particular to a method, apparatus, device and medium for extracting driving scenes. Background Technology
[0002] As the functions of intelligent driving vehicles become more diverse and their design and operational domains expand, traditional testing methods based on single-function testing can no longer meet the testing needs of intelligent driving vehicles. Scenario-based testing has gradually become an essential path to achieving high-level autonomous driving. By digitally recreating autonomous driving application scenarios through mathematical modeling and establishing scenario models that are as close to the real world as possible, and then analyzing and studying them through simulation testing, the goal of testing and verifying autonomous driving algorithms and systems can be achieved.
[0003] In existing technologies, scene generation schemes based on autonomous driving data collection often experience tracking result jumps when performing multi-target tracking. This affects the trajectory extraction results of the target in consecutive frames, and the jump phenomenon impacts the construction progress of the scene library and the generation efficiency of the collected scenes. Summary of the Invention
[0004] This application provides a driving scene extraction method, apparatus, device, and medium, which can reduce the false negative rate of driving scene extraction, reduce the computational complexity of driving scene extraction, and improve the accuracy of driving scene extraction.
[0005] According to one aspect of this application, a driving scene extraction method is provided, the method comprising:
[0006] Candidate targets in keyframes are determined based on the vehicle's driving data frames, the relative driving data frames of environmental targets relative to the vehicle, and a first threshold feature of the target driving scene; wherein, the driving data frames include at least one of the vehicle's data acquisition timestamp information, speed information, and orientation angle information; the relative driving data frames include at least one of the environmental target's data acquisition timestamp information, point cloud data information, orientation angle information, relative position information, and relative speed information;
[0007] The tracking result of the candidate target is determined based on the candidate target in the key frame, a preset number of target frames before the key frame, and a preset number of target frames after the key frame;
[0008] Based on the tracking results of the candidate targets and the second threshold features of the target driving scene, the target driving scene extraction result is determined.
[0009] According to another aspect of this application, a driving scene extraction device is provided, comprising:
[0010] The target determination module is used to determine candidate targets in key frames based on the vehicle's driving data frames, the relative driving data frames of environmental targets relative to the vehicle, and a first threshold feature of the target driving scene; wherein, the driving data frames include at least one of the vehicle's data acquisition timestamp information, speed information, and direction angle information; the relative driving data frames include at least one of the environmental target's data acquisition timestamp information, point cloud data information, direction angle information, relative position information, and relative speed information;
[0011] The target tracking module is used to determine the tracking result of the candidate target based on the candidate target in the key frame, a preset number of target frames before the key frame, and a preset number of target frames after the key frame;
[0012] The scene extraction module is used to determine the target driving scene extraction result based on the tracking result of the candidate target and the second threshold feature of the target driving scene.
[0013] According to another aspect of this application, a driving scene extraction device is provided, the device comprising:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the driving scene extraction method described in any embodiment of this application.
[0017] According to another aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the driving scene extraction method described in any embodiment of this application.
[0018] The technical solution of this application embodiment determines candidate targets in keyframes based on vehicle driving data frames, relative driving data frames of environmental targets to the vehicle, and a first threshold feature of the target driving scene; determines the tracking result of the candidate targets based on the candidate targets in the keyframes, a preset number of target frames before the keyframes, and a preset number of target frames after the keyframes; and determines the target driving scene extraction result based on the tracking result of the candidate targets and a second threshold feature of the target driving scene. This technical solution aims to achieve accurate tracking of scene targets and reduce the false detection rate and false negative rate of driving scene extraction.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart of a driving scene extraction method provided in Embodiment 1 of this application;
[0022] Figure 2 This is a schematic diagram of driving scene extraction provided in Embodiment 1 of this application;
[0023] Figure 3 A flowchart of a driving scene extraction method provided in Embodiment 2 of this application;
[0024] Figure 4 This is a schematic diagram of the structure of a driving scene extraction device provided in Embodiment 3 of this application;
[0025] Figure 5 This is a schematic diagram of the structure of a device for implementing a driving scene extraction method according to an embodiment of this application. Detailed Implementation
[0026] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0027] It should be noted that the terms "first," "second," "target," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] Example 1
[0029] Figure 1 This is a flowchart illustrating a driving scene extraction method provided in Embodiment 1 of this application. This embodiment is applicable to situations involving the extraction of driving scenes. The method can be executed by a driving scene extraction device, which can be implemented in hardware and / or software and can be configured in a device with data processing capabilities. Figure 1 As shown, the method includes:
[0030] S110. Based on the vehicle's driving data frame, the relative driving data frame of the environmental target relative to the vehicle, and the first threshold feature of the target driving scene, determine the candidate target in the key frame; wherein, the driving data frame includes at least one of the vehicle's data acquisition timestamp information, speed information, and direction angle information; the relative driving data frame includes at least one of the environmental target's data acquisition timestamp information, point cloud data information, direction angle information, relative position information, and relative speed information.
[0031] The vehicle's driving data frame consists of driving data collected during preset driving conditions. This data can be acquired via an industrial control computer from the vehicle's CAN (Controller Area Network) line, including the vehicle's data acquisition timestamp, speed, and steering angle. The data acquisition timestamp can correspond to the timestamps for the vehicle's speed and steering angle information.
[0032] The environmental targets refer to objects present in the surrounding environment when the vehicle is driving according to a preset driving state, such as vehicles and pedestrians. Relative driving data frames can be collected by the vehicle's intelligent sensing unit from sensors such as LiDAR, cameras, IMU (Inertial Measurement Unit), and GNSS (Global Navigation Satellite System). This data includes environmental target data acquisition timestamps, point cloud data, orientation angles, relative positions, and relative speeds. The point cloud data is a collection of massive points representing environmental target characteristics collected by the LiDAR sensor, and may include the environmental target's three-dimensional coordinates, color, and reflection intensity.
[0033] It should be noted that the vehicle's driving data frames and the relative driving data frames of environmental targets relative to the vehicle can be collected by the industrial control computer and intelligent sensing unit deployed on the vehicle, respectively, or by drones, or by roadside equipment. This embodiment of the invention does not limit the method of data frame collection.
[0034] The target driving scenario can be a scenario that meets the testing requirements of the vehicle to be tested, such as overtaking scenario, avoidance scenario, turning through intersection scenario, lane changing scenario, parking scenario, etc.
[0035] Automatically extracting the target driving scene from target information collected by the vehicle's intelligent sensing unit is computationally complex and consumes significant computing power. Furthermore, the target ID is prone to changing. Manual screening and correction, on the other hand, incurs substantial costs and is inefficient. To avoid the impact of easily changing target IDs in target tracking algorithms, this invention identifies candidate targets in keyframes and then matches these candidate targets against other data frames, thus mitigating the issue of easily changing target IDs.
[0036] Keyframes can be relative driving data frames containing the target driving scenario from which the target needs to be extracted. It is understood that a keyframe may contain one or more targets from the target driving scenario. Candidate targets can be environmental targets in the keyframe that meet a first threshold feature. The first threshold feature can be a feature that initially defines the target driving scenario, in order to accurately filter and locate keyframes containing the target driving scenario from which the target needs to be extracted.
[0037] For example, the target driving scenario is set as a lane-changing and cutting-in scenario. The first threshold feature is the characteristics of the vehicle and the preceding vehicle when the preceding vehicle is at the lane divider in the same direction. Specifically, this can be: the vehicle's speed threshold is greater than or equal to 5 km / h; the vehicle's steering wheel angle threshold is ±20°; the lateral relative distance range between the environmental vehicles and the vehicle is (-1.5 m, -1.2 m) and (1.2 m, 1.5 m); the longitudinal relative distance range between the environmental vehicles and the vehicle is [0 m, 50 m]; and the relative orientation range between the environmental vehicles and the vehicle is (-1 rad, 1 rad). Based on the collected vehicle's driving data frames and the relative driving data frames of environmental vehicles relative to the vehicle, relative driving data frames that simultaneously satisfy the above first threshold features are selected as keyframes. Environmental targets in the keyframes that meet the above first threshold features are candidate targets.
[0038] As an optional but non-limiting implementation, candidate targets in keyframes are determined based on the vehicle's driving data frames, the relative driving data frames of environmental targets to the vehicle, and a first threshold feature of the target driving scene, including but not limited to the following steps A1 to A3:
[0039] Step A1: Based on the vehicle's driving data frames and the first threshold features of the target driving scenario, determine the first timestamp corresponding to the driving data frames that meet the first threshold features.
[0040] In this embodiment of the invention, at least one of the speed information and direction angle information in the vehicle's driving data frame can be compared with a first threshold feature of the target driving scenario to determine the driving data frame that meets the first threshold feature, and the data collection timestamp information of the vehicle can be extracted based on the driving data frame to obtain the first timestamp.
[0041] Step A2: Based on the relative driving data frame of the environmental target relative to the vehicle and the first threshold feature of the target driving scene, determine the second timestamp corresponding to the relative driving data frame that meets the first threshold feature.
[0042] In this embodiment of the invention, at least one of the direction angle information, relative position information, and relative speed information in the relative driving data frame of the environmental target relative to the vehicle can be compared with the first threshold feature of the target driving scene to determine the relative driving data frame that meets the first threshold feature, and the data acquisition timestamp information of the environmental target can be extracted based on the relative driving data frame to obtain the second timestamp.
[0043] Step A3: Determine the target timestamp that is the same as the first timestamp and the second timestamp, and determine the environmental target in the relative driving data frame corresponding to the first target timestamp in the consecutive target timestamps as the candidate target in the key frame.
[0044] It should be noted that the target timestamps that have the same first and second timestamps can be consecutive timestamps. In this case, the first target timestamp is used as the key timestamp, and the environmental targets in the relative driving data frames corresponding to the key timestamps are determined as candidate targets in the key frames.
[0045] As an optional but non-limiting implementation, before determining the first timestamp corresponding to the driving data frame that meets the first threshold feature based on the vehicle's driving data frame and the first threshold feature of the target driving scenario, the method further includes: synchronizing the data acquisition timestamp information in the vehicle's driving data frame and the data acquisition timestamp information in the relative driving data frame of the environmental target relative to the vehicle.
[0046] If the vehicle's driving data frame is obtained by collecting the vehicle's CAN line information through an industrial control computer, and the relative driving data frame of the environmental target relative to the vehicle is obtained by collecting sensor data such as lidar, camera, IMU, and GNSS from the vehicle's intelligent sensing unit, the data acquisition timestamp information of the driving data frame and the relative driving data frame may be out of sync due to the asynchronous acquisition frequencies of the various data acquisition devices.
[0047] In this embodiment of the invention, the data acquisition timestamp information of the driving data frame and the data acquisition timestamp information in the relative driving data frame are synchronized. Specifically, the synchronization process adjusts the data acquisition timestamp information of the driving data frame and the data acquisition timestamp information in the relative driving data frame using a reference clock.
[0048] S120. Determine the tracking result of the candidate target based on the candidate target in the key frame, a preset number of target frames before the key frame, and a preset number of target frames after the key frame.
[0049] The preset number of target frames before and after the keyframe can be used as the tracking range for candidate targets, and the preset number can be determined according to actual needs. The tracking result can be the matching result of the candidate targets.
[0050] In this embodiment of the invention, candidate targets in keyframes are matched with environmental targets in target frames, and environmental targets in the target frame that belong to the same target as the candidate targets are determined as the tracking results of the candidate targets. The tracking results may include at least one of the following: data acquisition timestamp information of the candidate target frame, point cloud data information of the target in the target frame, orientation angle information, relative position information, and relative velocity information.
[0051] S130. Based on the tracking results of the candidate targets and the second threshold features of the target driving scene, determine the target driving scene extraction result.
[0052] It should be noted that the features of candidate targets obtained at keyframes may be exactly the same as the features of the objects to be extracted in the target driving scene, but they are not necessarily the objects to be extracted in the target driving scene. For example, if the target driving scene is a lane change scene, the keyframe includes two candidate targets, both of which have features that are present when they are at the lane divider. However, after analyzing a preset number of target frames before and after the keyframe, one candidate target is found to be changing from the lane where the vehicle is located to another lane, and the other candidate target is changing from another lane to the lane where the vehicle is located.
[0053] In view of the above problems, embodiments of the present invention further filter the tracking results of candidate targets by using a second threshold feature of the target driving scene to obtain accurate target driving scene extraction results. The second threshold feature can be a feature that precisely defines the target driving scene, thereby accurately filtering the objects to be extracted from the driving scene. The target driving scene extraction result can be a continuous driving data frame of the objects to be extracted from the driving scene, and may include at least one of the following: data acquisition timestamp information, point cloud data information, direction angle information, relative position information, and relative speed information of the objects to be extracted from the driving scene.
[0054] For example, the above-mentioned target driving scenario is the scenario where the vehicle in front changes lanes and cuts in, which will be used as an example for explanation. Figure 2 This is a schematic diagram illustrating a driving scene extraction method provided in Embodiment 1 of this application. Figure 2 As shown, vehicle M travels in a straight line along the middle lane. Candidate vehicle 1 and candidate vehicle 2 are obtained at keyframes. The target driving scene extraction result is determined based on the tracking results of candidate vehicle 1 and candidate vehicle 2 within 50 frames before and after the keyframe. The second threshold feature is set as the trajectory of the preceding vehicle cutting into the lane of the vehicle from another lane. Specifically, this can be defined as follows: the lateral relative distance between the environmental vehicles and the vehicle is greater than 2.2 meters in the first 50 frames of the keyframe, and less than 1.2 meters in the last 50 frames of the keyframe. Based on the tracking results of each candidate target in the keyframe, the candidate target information that meets the second threshold feature is taken as the target driving scene extraction result.
[0055] As an optional but non-limiting implementation, the target driving scene extraction result is determined based on the tracking result of the candidate target and the second threshold feature of the target driving scene, including but not limited to the following steps B1 to B2:
[0056] Step B1: Determine the scene target based on the tracking results of the candidate target and the second threshold feature of the target driving scene; wherein the scene target is the target extraction object of the target driving scene.
[0057] In this embodiment of the invention, the target tracking result that meets the second threshold feature is selected from the tracking results of the candidate targets, and the candidate target corresponding to the target tracking result is taken as the scene target.
[0058] Step B2: Based on the tracking results of the scene target, determine the driving trajectory and collect the timestamp information of the scene target.
[0059] In this embodiment of the invention, the driving trajectory of the scene target is generated based on the direction angle information, relative position information and relative speed information in the tracking results of the scene target; then, continuous data frames about the scene target are obtained based on the driving trajectory of the scene target and the collection timestamp information, so as to realize the extraction of the target driving scene.
[0060] This invention discloses a driving scene extraction method. The method determines candidate targets in keyframes based on vehicle driving data frames, relative driving data frames of environmental targets relative to the vehicle, and a first threshold feature of the target driving scene. The driving data frames include at least one of the vehicle's data acquisition timestamp information, speed information, and direction angle information. The relative driving data frames include at least one of the environmental target's data acquisition timestamp information, point cloud data information, direction angle information, relative position information, and relative speed information. The method determines the tracking results of the candidate targets based on the candidate targets in the keyframes, a preset number of target frames before the keyframes, and a preset number of target frames after the keyframes. Finally, the method determines the target driving scene extraction result based on the tracking results of the candidate targets and a second threshold feature of the target driving scene. This technical solution aims to achieve accurate tracking of scene targets and reduce the false detection rate and false negative rate in driving scene extraction.
[0061] Example 2
[0062] Figure 3 This is a flowchart of a driving scene extraction method provided in Embodiment 2 of this application. This embodiment is an optimized solution of the above-mentioned embodiment. The method includes:
[0063] S310. Based on the vehicle's driving data frame, the relative driving data frame of the environmental target relative to the vehicle, and the first threshold feature of the target driving scene, determine the candidate target in the key frame; wherein, the driving data frame includes at least one of the vehicle's data acquisition timestamp information, speed information, and direction angle information; the relative driving data frame includes at least one of the environmental target's data acquisition timestamp information, point cloud data information, direction angle information, relative position information, and relative speed information.
[0064] S320. Based on the candidate targets in the key frame, the environmental targets in a preset number of target frames before the key frame, and the environmental targets in a preset number of target frames after the key frame, determine the tracking target in the target frame that corresponds to the candidate target.
[0065] In this embodiment of the invention, the environmental targets included in the target frame are screened, and the environmental targets in the target frame that match the candidate targets are used as the tracking targets of the candidate targets. The matching criterion is that the candidate target and the environmental target belong to the same target. For example, the similarity between the candidate target and each environmental target can be determined separately, and the environmental target with the highest similarity in each target frame can be used as the tracking target of the candidate target. The similarity can be determined by comparing the pixel value histograms of the keyframe corresponding image and the target frame corresponding image, or by comparing the hash values of the keyframe corresponding image and the target frame corresponding image using a hash algorithm. Of course, this embodiment of the invention does not limit the method for matching candidate targets.
[0066] As an optional but non-limiting implementation, based on the candidate targets in the keyframe, the environmental targets in a preset number of target frames before the keyframe, and the environmental targets in a preset number of target frames after the keyframe, the tracking target corresponding to the candidate target in the target frame is determined, including but not limited to steps C1-C2:
[0067] Step C1: Based on the point cloud data information of the candidate target, the point cloud data information of each environmental target in a preset number of target frames before the key frame, and the point cloud data information of each environmental target in a preset number of target frames after the key frame, determine the overlap between each environmental target in each target frame and the candidate target.
[0068] The point cloud data information may include the bounding boxes of environmental targets. In this embodiment of the invention, the overlap between each environmental target and the candidate target in each target frame can be determined by calculating the intersection-union ratio of the bounding boxes of the candidate targets and the environmental targets.
[0069] Step C2: Determine the environmental target with the maximum overlap in each target frame as the tracking target corresponding to the candidate target.
[0070] The maximum overlap indicates that the environmental target and the candidate target are most similar and can be identified as belonging to the same target. Therefore, the environmental target corresponding to the maximum overlap in each target frame and the candidate target in the key frame belong to the same target.
[0071] S330. Based on the information of the tracked target in each target frame, determine the tracking result of the candidate target.
[0072] In this embodiment of the invention, based on the information of the tracked target in each target frame, the ID of the tracked target in each target frame and the ID of the candidate target in the key frame are unified. For example, the ID of the candidate target can be used to replace the ID of the tracked target in each target frame.
[0073] The advantage of this setting is that it can avoid the situation where the tracking target identification ID corresponding to the candidate target is prone to change in the target tracking algorithm.
[0074] S340. Based on the tracking results of the candidate targets and the second threshold features of the target driving scene, determine the target driving scene extraction result.
[0075] This invention provides a driving scene extraction method. The method determines candidate targets in keyframes based on vehicle driving data frames, relative driving data frames of environmental targets relative to the vehicle, and a first threshold feature of the target driving scene. The driving data frames include at least one of the vehicle's data acquisition timestamp information, speed information, and direction angle information. The relative driving data frames include at least one of the environmental target's data acquisition timestamp information, point cloud data information, direction angle information, relative position information, and relative speed information. Based on the candidate targets in the keyframes, environmental targets in a predetermined number of target frames before the keyframe, and environmental targets in a predetermined number of target frames after the keyframe, the method determines the tracking targets corresponding to the candidate targets in the target frames. Based on the tracking target information in each target frame, the method determines the tracking result of the candidate targets. Based on the tracking result of the candidate targets and a second threshold feature of the target driving scene, the method determines the target driving scene extraction result. This technical solution aims to achieve accurate tracking of scene targets, avoid easily changing target identification IDs, and reduce the false detection rate and false negative rate of driving scene extraction.
[0076] Based on the above embodiments, optionally, before determining the candidate targets in the keyframes according to the vehicle's driving data frames, the relative driving data frames of the environmental targets to the vehicle, and the first threshold features of the target driving scene, the method further includes, but is not limited to, steps D1-D2:
[0077] Step D1: Detect environmental targets based on the target detection model and generate a single-frame recognition list of environmental targets.
[0078] The target detection model can fuse and identify targets based on data collected by the vehicle's intelligent sensing unit to obtain environmental target information that meets the recognition requirements. Based on the environmental target information detected by the target detection model, a single-frame environmental target recognition list is generated. This list may include the environmental target's data acquisition timestamp, point cloud data, orientation angle, position, and speed information.
[0079] Step D2: Based on the environmental target single-frame identification list and the vehicle's driving data frame, determine the relative driving data frame of the environmental target relative to the vehicle.
[0080] In this embodiment of the invention, the relative position and relative speed information of the environmental target relative to the vehicle can be obtained based on the position and speed information in the environmental target single-frame identification list and the vehicle's driving data frames. For example, the relative position information can be the longitudinal distance and the lateral distance of the environmental target relative to the vehicle.
[0081] The advantage of this setting is that it facilitates the comparison of feature thresholds between the vehicle environment and the target driving scene, providing convenience for determining the extraction results of the target driving scene.
[0082] Example 3
[0083] Figure 4 This is a schematic diagram of a driving scene extraction device provided in Embodiment 3 of this application. This device can execute the driving scene extraction method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. Figure 4 As shown, the device includes:
[0084] The target determination module 410 is used to determine candidate targets in key frames based on the vehicle's driving data frames, the relative driving data frames of environmental targets relative to the vehicle, and a first threshold feature of the target driving scene; wherein, the driving data frames include at least one of the vehicle's data acquisition timestamp information, speed information, and direction angle information; the relative driving data frames include at least one of the environmental target's data acquisition timestamp information, point cloud data information, direction angle information, relative position information, and relative speed information;
[0085] The target tracking module 420 is used to determine the tracking result of the candidate target based on the candidate target in the key frame, a preset number of target frames before the key frame, and a preset number of target frames after the key frame;
[0086] The scene extraction module 430 is used to determine the target driving scene extraction result based on the tracking result of the candidate target and the second threshold feature of the target driving scene.
[0087] This invention provides a driving scene extraction device. The device determines candidate targets in keyframes based on vehicle driving data frames, relative driving data frames of environmental targets to the vehicle, and a first threshold feature of the target driving scene. It then determines the tracking result of the candidate targets based on the candidate targets in the keyframes, a preset number of target frames before the keyframes, and a preset number of target frames after the keyframes. Finally, it determines the target driving scene extraction result based on the tracking result of the candidate targets and a second threshold feature of the target driving scene. This technical solution aims to achieve accurate tracking of scene targets and reduce the false detection rate and false negative rate in driving scene extraction.
[0088] Furthermore, the target determination module 410 includes:
[0089] The first timestamp determination unit is used to determine the first timestamp corresponding to the driving data frame that meets the first threshold feature based on the vehicle's driving data frame and the first threshold feature of the target driving scenario.
[0090] The second timestamp determination unit is used to determine the second timestamp corresponding to the relative driving data frame that meets the first threshold feature based on the relative driving data frame of the environmental target relative to the vehicle and the first threshold feature of the target driving scenario.
[0091] The candidate target determination unit is used to determine the target timestamp that is the same as the first timestamp and the second timestamp, and to determine the environmental target in the relative driving data frame corresponding to the first target timestamp in the consecutive target timestamps as the candidate target in the key frame.
[0092] Furthermore, the target determination module 410 also includes:
[0093] The synchronization processing unit is used to synchronize the data acquisition timestamp information in the vehicle's driving data frame and the data acquisition timestamp information in the relative driving data frame of the environmental target relative to the vehicle before determining the first timestamp corresponding to the driving data frame that meets the first threshold feature based on the vehicle's driving data frame and the first threshold feature of the target driving scenario.
[0094] Furthermore, the target tracking module 420 includes:
[0095] The tracking target determination unit is used to determine the tracking target corresponding to the candidate target in the target frame based on the candidate target in the key frame, each environmental target in a preset number of target frames before the key frame, and each environmental target in a preset number of target frames after the key frame;
[0096] The tracking result determination unit is used to determine the tracking result of the candidate target based on the information of the tracked target in each target frame.
[0097] Furthermore, the target identification unit includes:
[0098] The target overlap determination subunit is used to determine the overlap between each environmental target in each target frame and the candidate target based on the point cloud data information of the candidate target, the point cloud data information of each environmental target in a preset number of target frames before the key frame, and the point cloud data information of each environmental target in a preset number of target frames after the key frame.
[0099] The target determination subunit is used to determine the environmental target with the maximum overlap in each target frame as the tracking target corresponding to the candidate target.
[0100] Furthermore, the scene extraction module 430 includes:
[0101] A scene target determination unit is used to determine the tracking result of the scene target based on the tracking result of the candidate target and the second threshold feature of the target driving scene; wherein, the scene target is the target extraction object of the target driving scene;
[0102] The driving scene extraction unit is used to determine the driving trajectory of the scene target and collect timestamp information based on the tracking results of the scene target.
[0103] Furthermore, the device also includes:
[0104] The target recognition list generation module is used to detect environmental targets based on a target detection model and generate a single-frame recognition list of environmental targets before determining candidate targets in key frames based on the vehicle's driving data frames, the relative driving data frames of environmental targets to the vehicle, and the first threshold features of the target driving scene.
[0105] The relative driving data frame determination module is used to determine the relative driving data frame of the environmental target relative to the vehicle based on the environmental target single frame identification list and the vehicle's driving data frame.
[0106] The driving scene extraction device provided in this application embodiment can execute a driving scene extraction method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the method execution.
[0107] Example 4
[0108] Figure 5A schematic diagram of the structure of a device 10 that can be used to implement embodiments of this application is shown. The device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0109] like Figure 5 As shown, device 10 includes at least one processor 11 and a memory, such as read-only memory (ROM) 12, random access memory (RAM) 13, etc., communicatively connected to at least one processor 11. The memory stores computer programs executable by at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of device 10. The processor 11, ROM 12, and RAM 13 are interconnected via bus 14. Input / output (I / O) interface 15 is also connected to bus 14.
[0110] Multiple components in device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0111] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as driving scene extraction methods.
[0112] In some embodiments, the driving scene extraction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the driving scene extraction method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the driving scene extraction method by any other suitable means (e.g., by means of firmware).
[0113] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0114] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0115] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0116] To provide interaction with a user, the systems and techniques described herein can be implemented on a device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback); and input from the user can be received in any form (including sound input, voice input, or haptic input).
[0117] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0118] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0119] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this application can be achieved, and this is not limited herein.
[0120] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A driving scene extraction method characterized by, The method comprises: According to the driving data frame of the vehicle, the relative driving data frame of the environmental target relative to the vehicle, and the first threshold feature of the target driving scene, the relative driving data frame that meets the first threshold feature is determined as the candidate target in the key frame; wherein the driving data frame comprises at least one of the data acquisition timestamp information, speed information and direction angle information of the vehicle; the relative driving data frame comprises at least one of the data acquisition timestamp information, point cloud data information, direction angle information, relative position information and relative speed information of the environmental target; the first threshold feature comprises: self-vehicle speed threshold, self-vehicle steering wheel angle threshold, lateral relative distance range of environmental vehicle and self-vehicle, longitudinal relative distance range of environmental vehicle and self-vehicle, and relative orientation range of environmental vehicle and self-vehicle; According to the candidate target in the key frame, the target frame before the key frame and the target frame after the key frame, the tracking result of the candidate target is determined; According to the tracking result of the candidate target, and the second threshold feature of the target driving scene, the target driving scene extraction result is determined; the second threshold feature comprises: the lateral relative distance of the environmental vehicle and the self-vehicle in the key frame before the frame and the lateral relative distance of the environmental vehicle and the self-vehicle in the key frame after the frame; According to the driving data frame of the vehicle, the relative driving data frame of the environmental target relative to the vehicle, and the first threshold feature of the target driving scene, the candidate target in the key frame is determined, comprising: According to the driving data frame of the vehicle and the first threshold feature of the target driving scene, the first timestamp corresponding to the driving data frame meeting the first threshold feature is determined; According to the relative driving data frame of the environmental target relative to the vehicle and the first threshold feature of the target driving scene, the second timestamp corresponding to the relative driving data frame meeting the first threshold feature is determined; The target timestamp of the same first timestamp and second timestamp is determined, and the environmental target in the relative driving data frame corresponding to the first target timestamp in the continuous target timestamp is determined as the candidate target in the key frame; Before determining the first timestamp corresponding to the driving data frame meeting the first threshold feature according to the driving data frame of the vehicle and the first threshold feature of the target driving scene, the method further comprises: The data acquisition timestamp information in the driving data frame of the vehicle and the data acquisition timestamp information in the relative driving data frame of the environmental target relative to the vehicle are synchronously processed; wherein the synchronous processing is to uniformly adjust the data acquisition timestamp information in the driving data frame and the data acquisition timestamp information in the relative driving data frame by reference clock.
2. The method of claim 1, wherein, According to the candidate target in the key frame, the target frame before the key frame and the target frame after the key frame, the tracking result of the candidate target is determined, comprising: The tracking target corresponding to the candidate target in the target frame is determined according to the candidate target in the key frame, each environment target in a preset number of target frames before the key frame, and each environment target in a preset number of target frames after the key frame. The tracking result of the candidate target is determined according to the information of the tracking target in each target frame.
3. The method of claim 2, wherein, The tracking target corresponding to the candidate target in the target frame is determined according to the candidate target in the key frame, each environment target in a preset number of target frames before the key frame, and each environment target in a preset number of target frames after the key frame, including: The coincidence degrees of each environment target and the candidate target in each target frame are respectively determined according to the point cloud data information of the candidate target, the point cloud data information of each environment target in a preset number of target frames before the key frame, and the point cloud data information of each environment target in a preset number of target frames after the key frame. The environment target corresponding to the maximum coincidence degree in each target frame is determined as the tracking target corresponding to the candidate target.
4. The method of claim 1, wherein, The target driving scene extraction result is determined according to the tracking result of the candidate target and the second threshold feature of the target driving scene, including: The scene target is determined according to the tracking result of the candidate target and the second threshold feature of the target driving scene; wherein the scene target is a target extraction object of the target driving scene; The driving trajectory and the collection timestamp information of the scene target are determined according to the tracking result of the scene target.
5. The method of claim 1, wherein, Before the candidate target in the key frame is determined according to the driving data frame of the vehicle, the relative driving data frame of the environment target relative to the vehicle, and the first threshold feature of the target driving scene, the method further includes: The environment target is detected based on a target detection model to generate an environment target single-frame recognition list; The relative driving data frame of the environment target relative to the vehicle is determined according to the environment target single-frame recognition list and the driving data frame of the vehicle.
6. A driving scene extraction apparatus characterized by comprising: The device includes: The target determination module is configured to filter the relative driving data frame that meets the first threshold feature as a key frame according to the driving data frame of the vehicle, the relative driving data frame of the environment target relative to the vehicle, and the first threshold feature of the target driving scene, and the environment target in the key frame that meets the first threshold feature as a candidate target; wherein the driving data frame includes at least one of data collection timestamp information, speed information, and direction angle information of the vehicle; the relative driving data frame includes at least one of data collection timestamp information, point cloud data information, direction angle information, relative position information, and relative speed information of the environment target; and the first threshold feature includes a vehicle speed threshold, a vehicle steering wheel angle threshold, a lateral relative distance range of the environment vehicle and the vehicle, a longitudinal relative distance range of the environment vehicle and the vehicle, and a relative orientation range of the environment vehicle and the vehicle. The target tracking module is configured to determine the tracking result of the candidate target according to the candidate target in the key frame, a preset number of target frames before the key frame, and a preset number of target frames after the key frame. The scene extraction module is configured to determine a target driving scene extraction result according to the tracking result of the candidate target and second threshold features of the target driving scene, wherein the second threshold features include lateral relative distances between the environment vehicle and the ego vehicle in the key frame front frame and the key frame rear frame. The target determination module includes: The first timestamp determination unit is configured to determine a first timestamp corresponding to a driving data frame that meets the first threshold features of the target driving scene according to the driving data frame of the vehicle and the first threshold features. The second timestamp determination unit is configured to determine a second timestamp corresponding to a relative driving data frame that meets the first threshold features of the target driving scene according to the relative driving data frame of the environment target relative to the vehicle and the first threshold features. The candidate target determination unit is configured to determine a target timestamp that is the same as the first timestamp and the second timestamp, and determine the environment target in the relative driving data frame corresponding to the first target timestamp in the continuous target timestamps as the candidate target in the key frame. The target determination module further includes: The synchronization processing unit is configured to perform synchronization processing on data acquisition timestamp information in the driving data frame of the vehicle and data acquisition timestamp information in the relative driving data frame of the environment target relative to the vehicle before determining the first timestamp corresponding to the driving data frame that meets the first threshold features of the target driving scene according to the driving data frame of the vehicle and the first threshold features.
7. A driving scene extraction device characterized by comprising: The device includes: at least one processor; and a memory connected with the at least one processor in communication; wherein The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the driving scene extraction method of any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing the processor to execute the driving scene extraction method of any one of claims 1-5 when executed.
Citation Information
Patent Citations
Passage support device and passage support method for vehicle
JP2012252407A
Obstacle tracking method, storage medium, and electronic device
US20220319189A1
Vehicle accident recording method and apparatus, and vehicle
WO2021004380A1