Scene data acquisition method and apparatus, device, and storage medium

CN115862321BActive Publication Date: 2026-08-21NEUSOFT REACH AUTOMOBILE TECH (SHENYANG) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211462477.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2026-08-21
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

[0002]近年来,随着自动车辆的数量不断增加,导致的车辆事故也不断增加;但是,由于事故数据较难获得,自车车主或其它车主难以直观地了解到事故的原因,也无法根据已经发生的事故学习如何规避风险,并导致了大量的事故数据浪费;因此,自车或其它车主在后续的驾驶过程中仍然无法对已经导致事故的事故场景进行规避,所以,往往会因为相同的原因,继续产生多起交通事故

Benefits of technology

[0043] The above-mentioned technical solution of this application 1) realizes that after obtaining the original video data corresponding to the accident scene, the vehicle is located based on the original video data in different ways. Specifically, the vehicle is first located based on the lane lines to obtain the first pose of the vehicle at each moment; the vehicle is second located based on VSLAM to obtain the second pose of the vehicle at each moment; then the first pose is corrected based on the second pose at the same moment to obtain the target pose of the vehicle at each moment; and the third pose of each target vehicle at each moment is obtained based on 3D recognition. In addition, the embodiments of this application also obtain the dynamic parameters of the vehicle and each target vehicle at each moment, thereby realizing the accurate acquisition of scene data corresponding to the accident scene. The method is simple and the scene data is easy to obtain, which solves the problem of difficult acquisition of scene data in related technologies, greatly improves the utilization rate of accident data, and facilitates the autonomous driving system to learn from the accident scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862321B_ABST
    Figure CN115862321B_ABST
Patent Text Reader

Abstract

The application discloses a scene data acquisition method and device, equipment and a storage medium. The method comprises the following steps: acquiring original video data based on a target vehicle recorder, performing lane line extraction on the original video data, obtaining initial lane lines at each time, and obtaining a first pose of a vehicle at each time based on each initial lane line; calculating a second pose of the vehicle at each time based on the original video data; obtaining a target pose of the vehicle at each time based on the first pose and the second pose of the vehicle at each time; performing target vehicle identification based on the original video data, and calculating a third pose of each target vehicle at each time; and obtaining scene data at each time based on the initial lane line at each time, the target pose of the vehicle and the third pose of each target vehicle. The technical scheme of the application is easy to obtain scene data of an accident scene, and the accident scene can be played back based on the scene data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of autonomous driving technology, and in particular relates to a method, apparatus, device and storage medium for acquiring scene data. Background Technology

[0002] In recent years, with the continuous increase in the number of autonomous vehicles, the number of vehicle accidents has also been increasing. However, due to the difficulty in obtaining accident data, it is difficult for autonomous vehicle owners or other vehicle owners to intuitively understand the causes of accidents, and they are unable to learn how to avoid risks based on the accidents that have already occurred, resulting in a large waste of accident data. Therefore, autonomous vehicle owners or other vehicle owners are still unable to avoid accident scenarios that have already caused accidents in subsequent driving processes, so multiple traffic accidents often continue to occur for the same reasons. Summary of the Invention

[0003] This application aims to at least partially address one of the technical problems in the related art. Therefore, one objective of this application is to provide a method, apparatus, device, and storage medium for acquiring scene data.

[0004] To address the aforementioned technical problems, embodiments of this application provide the following technical solutions:

[0005] A method for acquiring scene data, comprising:

[0006] Based on the raw video data acquired by the target dashcam, lane lines are extracted from the raw video data to obtain the initial lane lines at each moment, and based on each initial lane line, the first pose of the vehicle at each moment is obtained.

[0007] Based on the original video data, the second pose of the vehicle at each moment is calculated;

[0008] Based on the first pose and the second pose of the vehicle at each moment, the target pose of the vehicle at each moment is obtained.

[0009] Based on the original video data, target vehicles are identified, and the third pose of each target vehicle at each time step is calculated.

[0010] Based on the initial lane line, the target pose of the vehicle, and the third pose of each target vehicle at each time moment, scene data is obtained at each time moment; wherein, the scene data is used to replay the accident scene that matches the scene data.

[0011] Optionally, obtaining the vehicle's first pose at each moment based on each initial lane line includes:

[0012] Perform coordinate transformation on the initial lane lines at each time step to obtain the transformed lane lines at each time step;

[0013] Perform a top-down view transformation on the lane change lines at each time point to obtain the top-down lane lines;

[0014] Based on the overhead view of the lane lines and the initial lane lines, the first coordinate transformation parameters are calculated.

[0015] Based on the first coordinate transformation parameters and the original video data at each time moment, the first pose of the vehicle at each time moment is obtained.

[0016] Optionally, calculating the second pose of the vehicle at each moment based on the original video data includes:

[0017] VSLAM calculates the original pose of the vehicle at each time step based on the original video data at each time step;

[0018] Based on the original pose of the vehicle at each time step, the second pose of the vehicle at each time step is calculated.

[0019] Optionally, obtaining the target pose of the vehicle at each time step based on the first pose and the second pose at each time step includes:

[0020] Based on the overhead lane lines and the VSLAM, the second coordinate transformation parameters are obtained;

[0021] Based on the second coordinate transformation parameters, the first pose at each time moment is transformed to obtain the first transformed pose at each time moment;

[0022] Match the first transformed pose with the second pose at each time step;

[0023] If the first transformed pose and the second pose are successfully matched, the first transformed pose is determined as the target pose of the vehicle; or, if the first transformed pose and the second pose are not successfully matched, the second pose is determined as the target pose of the vehicle.

[0024] Optionally, the step of identifying target vehicles based on the original video data and calculating the third pose of each target vehicle includes:

[0025] Based on the original video data, target vehicle identification is performed to obtain identification results; wherein, the identification results include at least one of the target vehicles;

[0026] The tracking algorithm is used to track each target vehicle to obtain the third pose of each target vehicle at each time step.

[0027] Optionally, after identifying target vehicles based on the original video data and calculating the third pose of each target vehicle, the method further includes:

[0028] Based on the third pose of the target vehicle at time t and the third pose of the target vehicle at time t+n, the first running speed of the target vehicle within n time moments is calculated; where t and n are positive integers.

[0029] Based on the first running speed of the target vehicle within n time periods, the first running acceleration corresponding to each of the n time periods is calculated.

[0030] Optionally, the process of obtaining scene data at each time step based on the initial lane line, the target pose of the vehicle, and the third pose of each target vehicle includes:

[0031] Based on the overhead lane line, the target pose of the vehicle, and the third pose of each target vehicle at each time moment, the first sub-scene data at each time moment is obtained;

[0032] Based on the first operating speed of each target vehicle, the first operating acceleration of each target vehicle, the second operating acceleration of the vehicle, and the second operating speed of the vehicle at each time moment, the second sub-scene data at each time moment is obtained;

[0033] The scene data at each time step is obtained based on the first sub-scene data and the second sub-scene data at each time step.

[0034] Embodiments of this application also provide a scene data acquisition device, including:

[0035] The first acquisition module is used to acquire raw video data based on the target dashcam, extract lane lines from the raw video data, obtain the initial lane lines at each moment, and obtain the first pose of the vehicle at each moment based on each initial lane line.

[0036] The calculation module is used to calculate the second pose of the vehicle at each moment based on the original video data;

[0037] The second acquisition module is used to obtain the target pose of the vehicle at each time step based on the first pose and the second pose of the vehicle at each time step.

[0038] The recognition module is used to identify target vehicles based on the original video data and calculate the third pose of each target vehicle at each time step.

[0039] The simulation module is used to obtain scene data at each time step based on the initial lane line, the target pose of the vehicle, and the third pose of each target vehicle at each time step; wherein the scene data is used to replay the accident scene that matches the scene data.

[0040] Embodiments of this application also provide an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the method described above.

[0041] Embodiments of this application also provide a computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the method described above.

[0042] The embodiments of this application have the following technical effects:

[0043] The above-mentioned technical solution of this application 1) realizes that after obtaining the original video data corresponding to the accident scene, the vehicle is located based on the original video data in different ways. Specifically, the vehicle is first located based on the lane lines to obtain the first pose of the vehicle at each moment; the vehicle is second located based on VSLAM to obtain the second pose of the vehicle at each moment; then the first pose is corrected based on the second pose at the same moment to obtain the target pose of the vehicle at each moment; and the third pose of each target vehicle at each moment is obtained based on 3D recognition. In addition, the embodiments of this application also obtain the dynamic parameters of the vehicle and each target vehicle at each moment, thereby realizing the accurate acquisition of scene data corresponding to the accident scene. The method is simple and the scene data is easy to obtain, which solves the problem of difficult acquisition of scene data in related technologies, greatly improves the utilization rate of accident data, and facilitates the autonomous driving system to learn from the accident scene.

[0044] 2) It can obtain scene data corresponding to any accident scenario, and then simulate or replay the scene data corresponding to any accident scenario to test the performance of the autonomous driving system in the accident scenario or determine whether the autonomous driving system will cause an accident in the accident scenario, thereby reducing the probability of the autonomous driving system causing an accident and ultimately improving the safety of the autonomous driving system.

[0045] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0046] Figure 1 This is a flowchart illustrating a scene data acquisition method provided in an embodiment of this application;

[0047] Figure 2 This is an accident scenario obtained based on scene data playback, as provided in the embodiments of this application;

[0048] Figure 3 This is a schematic diagram of the structure of a scene data acquisition device provided in an embodiment of this application. Detailed Implementation

[0049] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0050] To facilitate understanding of the embodiments by those skilled in the art, some terms are explained below:

[0051] (1) VSLAM: Visual Simultaneous Localization and Mapping.

[0052] (2) GPS: Global Positioning System.

[0053] (2) 3D Recognition: 3D Behavior Recognition is a human behavior recognition technology based on Riemannian geometry and deep image learning. This technology can detect the actual position of an object in spatial coordinates, and at the same time, combined with behavior recognition algorithms, it can achieve high-precision and rapid recognition and capture of target behavior.

[0054] (3) focs3D: A revolutionary, high-performance 3D laser scanner with intuitive touchscreen control.

[0055] (4) CenterNet: An anchor-free model that models the center point of an object's bounding box, uses keypoint estimation to find the center point, and regresses it to all other object attributes, such as size and 3D position.

[0056] like Figure 1As shown, an embodiment of this application provides a method for acquiring scene data, including:

[0057] Step S11: Based on the target dashcam, acquire the original video data, extract lane lines from the original video data to obtain the initial lane lines at each moment, and based on each initial lane line, obtain the first pose of the vehicle at each moment.

[0058] In an optional embodiment of this application, obtaining the vehicle's first pose at each moment based on each initial lane line includes:

[0059] Perform coordinate transformation on the initial lane lines at each time step to obtain the transformed lane lines at each time step;

[0060] Perform a top-down view transformation on the lane change lines at each time point to obtain the top-down lane lines;

[0061] Based on the overhead view of the lane lines and the initial lane lines, the first coordinate transformation parameters are calculated.

[0062] Based on the first coordinate transformation parameters and the original video data at each time moment, the first pose of the vehicle at each time moment is obtained.

[0063] In an optional embodiment of this application, the start time of the original video data is determined based on the time T when the accident occurs or a time T close to the time when the accident occurs; wherein, T > 0;

[0064] First, the length of the original video data (preset time period) is preset, and then the end time of the original video data is determined; for example, 1 minute, etc., but the embodiments of this application do not specifically limit this.

[0065] Regarding the source of the original video data, some car owners will upload the videos obtained by their dashcams to the server (e.g., Bilibili, Youku, or other public datasets). Therefore, when obtaining the original video data, it can be downloaded from the server; that is, it is obtained based on the network. Correspondingly, the dashcam that provides the original video data is identified as the target dashcam.

[0066] Furthermore, embodiments of this application can also obtain raw video data from other car owners based on the target dashcam by cooperating with buses, taxis, and car clubs.

[0067] Furthermore, when obtaining raw video data through cooperation with buses, taxis, and car clubs, GPS data at the corresponding time can also be obtained to make the positioning of the vehicle and the target vehicle more accurate.

[0068] In an optional embodiment of this application, based on the original video data, the pose change and positioning of the vehicle in the dashcam coordinate system corresponding to the target dashcam are first calculated; then, based on the coordinate transformation extrinsic parameters between the target dashcam coordinate system and the vehicle dashcam coordinate system, the pose change and positioning in the target dashcam coordinate system are converted into the first pose of the vehicle; wherein, the coordinate transformation extrinsic parameters can be preset or labeled according to actual needs, or calculated based on the positioning data of the target dashcam and the vehicle dashcam;

[0069] Furthermore, based on the original video data, the original lane lines in the target dashcam's coordinate system are first obtained; then, based on the aforementioned coordinate transformation extrinsic parameters, the original lane lines are transformed to obtain the initial lane lines in the vehicle's dashcam's coordinate system; wherein, the initial lane lines also include dimensional information such as the width of the lane lines.

[0070] Furthermore, in order to facilitate determining the position of the vehicle's first pose relative to the initial lane line at each moment, embodiments of this application convert the initial lane line at each moment to obtain a top-view lane line, that is, convert the initial lane line in the dashcam coordinate system of the vehicle dashcam into a top-view lane line in top-view format; then match the vehicle's first pose at each moment with the top-view lane line at the same moment to determine the position of the vehicle relative to the top-view lane line at each moment; for example, determine the distance of the vehicle relative to the edge of the top-view lane line at each moment.

[0071] It should be noted that if the coordinate transformation extrinsic parameters are not preset or labeled, or if the coordinate transformation extrinsic parameters cannot be obtained, they can be preset according to the vehicle type of the vehicle.

[0072] The embodiments of this application realize the acquisition of original video data corresponding to the accident scene. The method is simple and easy to operate.

[0073] Step S12: Based on the original video data, calculate the second pose of the vehicle at each moment;

[0074] In an optional embodiment of this application, calculating the second pose of the vehicle at each moment based on the original video data includes:

[0075] VSLAM calculates the original pose of the vehicle at each time step based on the original video data at each time step;

[0076] Based on the original pose of the vehicle at each time step, the second pose of the vehicle at each time step is calculated.

[0077] Specifically, the VSLAM-based coordinate system generally starts from the first frame of the original video data. Based on each frame, the second pose of the vehicle at each time step is obtained. Then, by concatenating the trajectory points of adjacent frames, the motion trajectory of the vehicle can be obtained, thereby achieving continuous localization of the vehicle.

[0078] Step S13: Based on the first pose and the second pose of the vehicle at each time step, obtain the target pose of the vehicle at each time step.

[0079] In an optional embodiment of this application, obtaining the target pose of the vehicle at each time step based on the first pose and the second pose of the vehicle at each time step includes:

[0080] Based on the overhead lane lines and the VSLAM, the second coordinate transformation parameters are obtained;

[0081] Based on the second coordinate transformation parameters, the first pose at each time moment is transformed to obtain the first transformed pose at each time moment;

[0082] Match the first transformed pose with the second pose at each time step;

[0083] If the first transformed pose and the second pose are successfully matched, the first transformed pose is determined as the target pose of the vehicle; or, if the first transformed pose and the second pose are not successfully matched, the second pose is determined as the target pose of the vehicle.

[0084] The first pose involved in the embodiments of this application also corresponds to the coordinate system of the vehicle's dashcam; while the second pose is obtained based on VSLAM, that is, the second pose corresponds to the VSLAM coordinate system of the vehicle; in order to compare the first pose and the second pose of the vehicle at the same time, the first pose and the second pose are unified under the same coordinate system (the coordinate system of VSLAM) and then compared.

[0085] Specifically, since the top-view lane lines are obtained based on the dashcam coordinate system of the vehicle, in order to obtain the matching relationship between the first pose and the second pose of the vehicle at each moment, the second coordinate transformation parameter between the dashcam coordinate system and the VSLAM coordinate system is obtained. Based on the second coordinate transformation parameter, the top-view lane lines are converted into target lane lines in the VSLAM coordinate system. Then, the target lane lines obtained at each moment are saved as the target lane lines corresponding to the scene data at each moment.

[0086] Furthermore, based on the second coordinate transformation parameters, the first pose of the vehicle at each moment is transformed into the first transformed pose, thereby achieving position matching between the first transformed pose and the second pose in the same coordinate system (the coordinate system of VSLAM).

[0087] Furthermore, after obtaining the first transformed pose of the first pose in the VSLAM coordinate system:

[0088] The first and second poses of the vehicle at each moment are compared;

[0089] If the first transition pose and the second pose of the vehicle at a certain moment cannot be matched or are inconsistent, then the second pose at that moment is determined as the target pose of the vehicle at that moment; or, if the first transition pose and the second pose of the vehicle at a certain moment can be matched or are consistent, then the first transition pose at that moment is determined as the target pose of the vehicle at that moment.

[0090] In the embodiments of this application, the first transformed pose of the vehicle at each time step is corrected based on the second pose of the vehicle obtained by VSLAM at each time step, which is based on the positioning of the target lane line. This achieves accurate positioning of the vehicle by combining VSLAM and lane lines, improves the positioning accuracy and the accuracy of scene simulation, and is beneficial to the accurate testing and learning of accident scenarios by the autonomous driving system.

[0091] Step S14: Based on the original video data, identify the target vehicles and calculate the third pose of each target vehicle at each time step;

[0092] In an optional embodiment of this application, the step of identifying target vehicles based on the original video data and calculating the third pose of each target vehicle includes:

[0093] Based on the original video data, target vehicle identification is performed to obtain identification results; wherein, the identification results include at least one of the target vehicles;

[0094] The tracking algorithm is used to track each target vehicle to obtain the third pose of each target vehicle at each time step.

[0095] In the embodiments of this application, based on 3D recognition, each frame of the original video data is identified to determine each target vehicle used to simulate each accident scenario, and the third pose of each target vehicle at each moment is obtained; wherein, 3D recognition can be implemented based on FOCS3D or CenterNet, etc.

[0096] Specifically, firstly, based on the dashcam coordinate system of the target dashcam corresponding to the original video data, the pose change and positioning of each target vehicle at each moment are obtained; then, based on the above coordinate transformation extrinsic parameters, the pose change and positioning of each target vehicle at each moment are transformed to the dashcam coordinate system of the vehicle, and correspondingly, the third pose of each target vehicle at each moment is obtained.

[0097] Furthermore, the third pose at each moment is tracked using a tracking algorithm to obtain the third pose corresponding to the next moment. By analogy, the trajectory points of each target vehicle in each frame of the image can be obtained, that is, the coordinates at each moment. By connecting multiple trajectory points within a preset time period in chronological order, the running trajectory of each target vehicle within the preset time period can be obtained, thereby realizing continuous positioning of each target vehicle within the preset time period.

[0098] Furthermore, in order to determine the position of each target vehicle relative to the vehicle and the target lane line at each time step, the third pose of each target vehicle at each time step is transformed into the VSLAM coordinate system of the vehicle based on the second coordinate transformation parameters mentioned above, and the third transformed pose of each target vehicle at each time step is obtained.

[0099] In an optional embodiment of this application, during the actual target vehicle identification process, there may be many other vehicle participants appearing in the original video data. Some of these other vehicle participants have no impact on the accident. Therefore, when identifying the target vehicle, the other vehicle participants actually present in the original video data can be filtered according to their relative positions to the vehicle itself. This filters out unnecessary data, reduces the waste of system computing power, and ultimately obtains at least one target vehicle. The relative position can be preset to 100 meters away from the vehicle, etc. The embodiments of this application do not specifically limit this. In the actual filtering process, the filtering parameters (e.g., the relative position can be preset to 100 meters away from the vehicle) can be preset according to actual needs.

[0100] In an optional embodiment of this application, each frame of the original video data can be identified based on 3D recognition to obtain traffic light information; specifically, the traffic light information includes, but is not limited to, the position and color of the traffic lights.

[0101] Similarly, based on the coordinate system transformation method described above, the traffic light information is also transformed into the VSLAM coordinate system of the vehicle, and the target traffic light information is obtained.

[0102] In an optional embodiment of this application, if traffic light information cannot be obtained from the original video data based on 3D recognition, it can be extracted from the map using GPS data at the corresponding time.

[0103] In the embodiments of this application, each target vehicle is identified and its third transition pose is obtained. Then, the third transition pose of each target vehicle, the target pose of the vehicle, the target lane line, and the target traffic light information at the same time are matched to determine the state of the lane line, the vehicle, the target vehicle, and the traffic light at the time of the accident or the time close to the accident, thus realizing the preliminary simulation of the accident scenario at each time.

[0104] An optional embodiment of this application, after identifying target vehicles based on the original video data and calculating the third pose of each target vehicle, further includes:

[0105] Based on the third pose of the target vehicle at time t and the third pose of the target vehicle at time t+n, the first running speed of the target vehicle within n time moments is calculated; where t and n are positive integers.

[0106] Based on the first running speed of the target vehicle within n time periods, the first running acceleration corresponding to each of the n time periods is calculated.

[0107] In an optional embodiment of this application, in order to more accurately simulate or replay each accident scenario, not only are the positions of each target vehicle, the vehicle itself, and the traffic light relative to the target lane line obtained at each moment, but the dynamic parameters of each target vehicle, the vehicle itself, and the traffic light are also calculated to more accurately reconstruct each accident scenario.

[0108] Specifically, taking the calculation of the first running speed of each target vehicle as an example, the first running speed of each target vehicle at each moment can be calculated based on the position difference between frames and the corresponding time difference between frames. Specifically, the coordinates of each target vehicle at each moment can be obtained based on 3D recognition. Based on the coordinates of two adjacent moments, the position difference of each target vehicle can be calculated. Furthermore, the position difference of each target vehicle can also be calculated based on several adjacent frames of images corresponding to each target vehicle (i.e., within a preset time period).

[0109] Furthermore, since the first running speed of each target vehicle at each moment is calculated solely based on the inter-frame position difference and the corresponding time difference, the error is relatively large. Therefore, in order to improve the accuracy of the calculation result of the first running speed, windowing optimization is performed. For example, a preset distance or preset time period (preset consecutive adjacent frames) can be taken to calculate the first running speed of each target vehicle at each moment, and then differential is performed to calculate the first running acceleration and behavior of each target vehicle at each moment. Among them, for the behavior of each target vehicle, for example, if the behavior of the target vehicle at a certain moment is to cross, it is calculated after reorganizing the basic running information such as the third transformation pose and the first running speed of the target vehicle at that moment. The embodiments of this application do not specifically limit this calculation process, and can be calculated by selecting appropriate formulas or other methods according to actual needs.

[0110] Similarly, the second running speed and second running acceleration of the vehicle at each moment can be calculated based on the above method, which will not be described in detail in the embodiments of this application.

[0111] In the embodiments of this application, after achieving the preliminary simulation of the accident scenario at each moment, the dynamic parameters of the vehicle, the target vehicle, and the traffic lights are further determined, such as the second running speed and second running acceleration of the vehicle, the first running speed and first running acceleration of the target vehicle, and the color of the traffic lights.

[0112] The embodiments of this application realize the localization of the vehicle based on the original video data corresponding to the accident scene through different methods after obtaining the original video data. Specifically, the vehicle is first localized based on lane lines to obtain the first pose of the vehicle at each moment; the vehicle is second localized based on VSLAM to obtain the second pose of the vehicle at each moment; then the first pose is corrected based on the second pose at the same moment to obtain the target pose of the vehicle at each moment; and the third pose of each target vehicle at each moment is obtained based on 3D recognition. In addition, the embodiments of this application also obtain the dynamic parameters of the vehicle and each target vehicle at each moment, thereby realizing the accurate acquisition of scene data corresponding to the accident scene. The method is simple and the scene data is easy to obtain, which solves the problem of difficult acquisition of scene data in related technologies, greatly improves the utilization rate of accident data, and facilitates the autonomous driving system to learn from the accident scene.

[0113] Step S15: Based on the initial lane line, the target pose of the vehicle, and the third pose of each target vehicle at each time moment, obtain scene data for each time moment; wherein, the scene data is used to replay historical scenes that match the scene data.

[0114] In an optional embodiment of this application, obtaining scene data at each time step based on the initial lane line, the target pose of the vehicle, and the third pose of each target vehicle includes:

[0115] Based on the overhead lane line, the target pose of the vehicle, and the third pose of each target vehicle at each time moment, the first sub-scene data at each time moment is obtained;

[0116] Based on the first operating speed of each target vehicle, the first operating acceleration of each target vehicle, the second operating acceleration of the vehicle, and the second operating speed of the vehicle at each time moment, the second sub-scene data at each time moment is obtained;

[0117] The scene data at each time step is obtained based on the first sub-scene data and the second sub-scene data at each time step.

[0118] In the embodiments of this application, based on the first running speed, first running acceleration, and third pose of each target vehicle at each moment, the second running speed, second running acceleration, and first transformed pose of the vehicle, and the position and color of the traffic light, scene data corresponding to each moment of the accident scenario is obtained. Based on the scene data, information such as the position of the vehicle and each target vehicle relative to the target lane line at each moment within a preset time period can be regressed. The regressed information such as the position of the vehicle and each target vehicle relative to the target vehicle at each moment within the preset time period is abstracted and stored, thereby enabling the acquisition of scene data corresponding to any accident scenario based on the above method.

[0119] Furthermore, in the embodiments of this application, after obtaining the scene data at each moment corresponding to the accident scenario, the scene data at each moment is input into any simulator to simulate or replay the accident scenario corresponding to the previous accident, so as to test and learn the accident scenario, thereby reducing the probability of accidents occurring in autonomous driving and improving the safety of autonomous driving.

[0120] like Figure 2 The image shows an accident scenario obtained based on scene data playback:

[0121] Specifically, within a target lane of width xxx, the vehicle is moving at a constant second speed of X km / h. 55m ahead, a car of size xxx, with a first speed of xxx and a first acceleration of xxx, is crossing the lane. The traffic light is green.

[0122] In an optional embodiment of this application, after obtaining scene data of an accident scenario, the scene data can be abstracted to describe the accident scenario in a script-like manner.

[0123] For example, an accident scenario can be described in the following way:

[0124] 1) Within a target lane of width xxx, when the vehicle is moving at a constant second speed of X km / h, a car of size xxx, with a first speed of xxx and a first acceleration of xxx crosses the lane 50m ahead.

[0125] 2) Within a target lane of width xxx, a car is moving at a constant second speed of X km / h. A car of size xxx is 50m behind it, moving at a first speed of xxx and a first acceleration of xxx.

[0126] 3) Within the target lane with a width of xxx, when the vehicle is moving at a constant second speed of X km / h, a car with a size of xxx, a second speed of xxx, and a second acceleration of xxx crosses the lane 60m ahead. The traffic light is green.

[0127] 4) Within a target lane of width xxx, when the vehicle travels at a second speed of X km / h and a speed of Y km / h... 2 The second acceleration accelerates forward. A car of size xxx is 80m ahead. The first speed is xxx. The first acceleration is xxx. The traffic light is green.

[0128] According to the embodiments of this application, based on the above method, it is possible to obtain scene data corresponding to any accident scenario, and then simulate or replay the obtained scene data corresponding to any accident scenario to test the performance of the autonomous driving system in the accident scenario or determine whether the autonomous driving system will cause an accident in the accident scenario, thereby reducing the probability of the autonomous driving system causing an accident and ultimately improving the safety of the autonomous driving system.

[0129] like Figure 3 As shown, embodiments of this application also provide a scene data acquisition device 20, including:

[0130] The first acquisition module 31 is used to acquire raw video data based on the target dashcam, extract lane lines from the raw video data, obtain the initial lane lines at each moment, and obtain the first pose of the vehicle at each moment based on each initial lane line.

[0131] Calculation module 32 is used to calculate the second pose of the vehicle at each moment based on the original video data;

[0132] The second acquisition module 33 is used to obtain the target pose of the vehicle at each time based on the first pose and the second pose of the vehicle at each time.

[0133] The recognition module 34 is used to identify target vehicles based on the original video data and calculate the third pose of each target vehicle at each time step.

[0134] The simulation module 35 is used to obtain scene data at each time step based on the initial lane line, the target pose of the vehicle, and the third pose of each target vehicle at each time step; wherein the scene data is used to replay the accident scene that matches the scene data.

[0135] Optionally, obtaining the vehicle's first pose at each moment based on each initial lane line includes:

[0136] Perform coordinate transformation on the initial lane lines at each time step to obtain the transformed lane lines at each time step;

[0137] Perform a top-down view transformation on the lane change lines at each time point to obtain the top-down lane lines;

[0138] Based on the overhead view of the lane lines and the initial lane lines, the first coordinate transformation parameters are calculated.

[0139] Based on the first coordinate transformation parameters and the original video data at each time moment, the first pose of the vehicle at each time moment is obtained.

[0140] Optionally, calculating the second pose of the vehicle at each moment based on the original video data includes:

[0141] VSLAM calculates the original pose of the vehicle at each time step based on the original video data at each time step;

[0142] Based on the original pose of the vehicle at each time step, the second pose of the vehicle at each time step is calculated.

[0143] Optionally, obtaining the target pose of the vehicle at each time step based on the first pose and the second pose at each time step includes:

[0144] Based on the overhead lane lines and the VSLAM, the second coordinate transformation parameters are obtained;

[0145] Based on the second coordinate transformation parameters, the first pose at each time moment is transformed to obtain the first transformed pose at each time moment;

[0146] Match the first transformed pose with the second pose at each time step;

[0147] If the first transformed pose and the second pose are successfully matched, the first transformed pose is determined as the target pose of the vehicle; or, if the first transformed pose and the second pose are not successfully matched, the second pose is determined as the target pose of the vehicle.

[0148] Optionally, the step of identifying target vehicles based on the original video data and calculating the third pose of each target vehicle includes:

[0149] Based on the original video data, target vehicle identification is performed to obtain identification results; wherein, the identification results include at least one of the target vehicles;

[0150] The tracking algorithm is used to track each target vehicle to obtain the third pose of each target vehicle at each time step.

[0151] Optionally, after identifying target vehicles based on the original video data and calculating the third pose of each target vehicle, the method further includes:

[0152] Based on the third pose of the target vehicle at time t and the third pose of the target vehicle at time t+n, the first running speed of the target vehicle within n time moments is calculated; where t and n are positive integers.

[0153] Based on the first running speed of the target vehicle within n time periods, the first running acceleration corresponding to each of the n time periods is calculated.

[0154] Optionally, the process of obtaining scene data at each time step based on the initial lane line, the target pose of the vehicle, and the third pose of each target vehicle includes:

[0155] Based on the overhead lane line, the target pose of the vehicle, and the third pose of each target vehicle at each time moment, the first sub-scene data at each time moment is obtained;

[0156] Based on the first operating speed of each target vehicle, the first operating acceleration of each target vehicle, the second operating acceleration of the vehicle, and the second operating speed of the vehicle at each time moment, the second sub-scene data at each time moment is obtained;

[0157] The scene data at each time step is obtained based on the first sub-scene data and the second sub-scene data at each time step.

[0158] Embodiments of this application also provide an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the method described above.

[0159] Embodiments of this application also provide a computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the method described above.

[0160] Furthermore, other configurations and functions of the apparatus in the embodiments of this application are known to those skilled in the art, and will not be described in detail here to reduce redundancy.

[0161] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0162] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0163] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0164] In the description of this application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc., indicating the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0165] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0166] In this application, unless otherwise expressly specified and limited, the terms "installation," "connection," "joining," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise expressly limited. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0167] In this application, unless otherwise expressly specified and limited, "above" or "below" the second feature can mean that the first feature is in direct contact with the second feature, or that the first feature is in indirect contact with the second feature through an intermediate medium. Furthermore, "above," "on top of," and "over" the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0168] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A method for acquiring scene data, characterized in that, include: Based on the raw video data acquired by the target dashcam, lane lines are extracted from the raw video data to obtain the initial lane lines at each moment, and based on each initial lane line, the first pose of the vehicle at each moment is obtained. Based on the original video data, the second pose of the vehicle at each moment is calculated; Based on the first pose and the second pose of the vehicle at each moment, the target pose of the vehicle at each moment is obtained. Based on the original video data, target vehicles are identified, and the third pose of each target vehicle at each time step is calculated. Based on the initial lane lines, the target pose of the vehicle, and the third pose of each target vehicle at each time moment, scene data is obtained at each time moment; wherein, the scene data is used to replay the accident scene that matches the scene data; The process of obtaining the vehicle's first pose at each moment based on each initial lane line includes: Perform coordinate transformation on the initial lane lines at each time step to obtain the transformed lane lines at each time step; Perform a top-down view transformation on the lane change lines at each time point to obtain the top-down lane lines; Based on the overhead view of the lane lines and the initial lane lines, the first coordinate transformation parameters are calculated. Based on the first coordinate transformation parameters and the original video data at each time moment, the first pose of the vehicle at each time moment is obtained; The step of obtaining the target pose of the vehicle at each time step based on the first pose and the second pose of the vehicle at each time step includes: Based on the overhead lane lines and VSLAM, the second coordinate transformation parameters are obtained; Based on the second coordinate transformation parameters, the first pose at each time moment is transformed to obtain the first transformed pose at each time moment; Match the first transformed pose with the second pose at each time step; If the first transformed pose and the second pose are successfully matched, the first transformed pose is determined as the target pose of the vehicle; or, if the first transformed pose and the second pose are not successfully matched, the second pose is determined as the target pose of the vehicle.

2. The method according to claim 1, characterized in that, The step of calculating the second pose of the vehicle at each moment based on the original video data includes: The VSLAM calculates the original pose of the vehicle at each time step based on the original video data at each time step. Based on the original pose of the vehicle at each time step, the second pose of the vehicle at each time step is calculated.

3. The method according to claim 1, characterized in that, The step of identifying target vehicles based on the original video data and calculating the third pose of each target vehicle includes: Based on the original video data, target vehicle identification is performed to obtain identification results; wherein, the identification results include at least one of the target vehicles; The tracking algorithm is used to track each target vehicle to obtain the third pose of each target vehicle at each time step.

4. The method according to claim 1, characterized in that, After identifying target vehicles based on the original video data and calculating the third pose of each target vehicle, the method further includes: Based on the third pose of the target vehicle at time t and the third pose of the target vehicle at time t+n, the first running speed of the target vehicle within n time moments is calculated; where t and n are positive integers. Based on the first running speed of the target vehicle within n time periods, the first running acceleration corresponding to each of the n time periods is calculated.

5. The method according to claim 4, characterized in that, The scene data for each time moment is obtained based on the initial lane line, the target pose of the vehicle, and the third pose of each target vehicle, including: Based on the overhead lane line, the target pose of the vehicle, and the third pose of each target vehicle at each time moment, the first sub-scene data at each time moment is obtained; Based on the first operating speed of each target vehicle, the first operating acceleration of each target vehicle, the second operating acceleration of the vehicle, and the second operating speed of the vehicle at each time moment, the second sub-scene data at each time moment is obtained; The scene data at each time step is obtained based on the first sub-scene data and the second sub-scene data at each time step.

6. A scene data acquisition device, characterized in that, include: The first acquisition module is used to acquire raw video data based on the target dashcam, extract lane lines from the raw video data, obtain the initial lane lines at each moment, and obtain the first pose of the vehicle at each moment based on each initial lane line. The calculation module is used to calculate the second pose of the vehicle at each moment based on the original video data; The second acquisition module is used to obtain the target pose of the vehicle at each time step based on the first pose and the second pose of the vehicle at each time step. The recognition module is used to identify target vehicles based on the original video data and calculate the third pose of each target vehicle at each time step. The simulation module is used to obtain scene data at each time step based on the initial lane line, the target pose of the vehicle, and the third pose of each target vehicle at each time step; wherein the scene data is used to replay the accident scene that matches the scene data. The first obtaining module is specifically used for: Perform coordinate transformation on the initial lane lines at each time step to obtain the transformed lane lines at each time step; Perform a top-down view transformation on the lane change lines at each time point to obtain the top-down lane lines; Based on the overhead view of the lane lines and the initial lane lines, the first coordinate transformation parameters are calculated. Based on the first coordinate transformation parameters and the original video data at each time moment, the first pose of the vehicle at each time moment is obtained; The second obtaining module is specifically used for: Based on the overhead lane lines and VSLAM, the second coordinate transformation parameters are obtained; Based on the second coordinate transformation parameters, the first pose at each time moment is transformed to obtain the first transformed pose at each time moment; Match the first transformed pose with the second pose at each time step; If the first transformed pose and the second pose are successfully matched, the first transformed pose is determined as the target pose of the vehicle; or, if the first transformed pose and the second pose are not successfully matched, the second pose is determined as the target pose of the vehicle.

7. An electronic device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the method as claimed in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Vehicle positioning method and system

    CN113869203A

  • Method and device for generating traffic scene

    CN114721963A

  • Pose determination method and device, computer equipment and storage medium

    CN114913500A