Information generation method and device for test scene of vehicle, processor and vehicle

By using a multi-UAV collaborative data acquisition and target detection model, vehicle test scenario information is generated, solving the accuracy problem of single-UAV aerial photography and achieving high-precision and comprehensive test scenario information generation, supporting key scenario recognition for autonomous driving systems.

CN121366366APending Publication Date: 2026-01-20CHINA FAW CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511240241.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of information generation in vehicle testing scenarios is low. Single-drone aerial photography is limited by field of view and flight altitude, making it impossible to achieve large-scale, high-precision data collection simultaneously. Furthermore, it relies on manual annotation or simple rule extraction, which cannot effectively identify dangerous interactive features during the movement process.

Method used

Multiple drones are used to collaboratively collect initial image information of the vehicle. Target image information is generated through frame-level synchronous processing and target detection models. Combined with movement trajectory and location information, risk is assessed using a driving risk field model to generate information about the vehicle test scenario.

Benefits of technology

It achieves highly accurate generation of vehicle test scenario information, avoids information loss and blind spots from a single perspective, improves the completeness and accuracy of test scenario information, and supports the identification of key scenarios in autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121366366A_ABST
    Figure CN121366366A_ABST
Patent Text Reader

Abstract

The invention discloses an information generation method and device for a test scene of a vehicle, a processor and the vehicle. The method comprises the steps that a plurality of unmanned aerial vehicles are used for obtaining initial image information of a road where a vehicle is located, and the initial image information comprises road scene data of the road where the vehicle is located; frame-level synchronization processing is carried out on the initial image information to obtain target image information, and the target image information is used for representing timestamps and position information corresponding to all frames in the initial image information; a target detection model is used for analyzing the target image information to obtain the moving track of the vehicle, and the target detection model is used for positioning and classifying the vehicle in the target image information; and generating information of a target test scene where the vehicle is located based on the moving track and the position information. The technical problem of low accuracy of information generation of the test scene of the vehicle is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicles, in particular to a vehicle test scene information generation method, device, processor and vehicle. BACKGROUND

[0002] At present, the initial image information of the road where the vehicle is located is mainly collected in the form of single unmanned aerial vehicle aerial photography. Due to the limitations of the field of view of the unmanned aerial vehicle camera, the performance of the camera, the flight height and other aspects, the current aerial photography data can only achieve small-range high-precision collection. If large-range collection is pursued, the single unmanned aerial vehicle needs to fly to a very high altitude, resulting in poor display effect of the vehicle in the initial image information, and the high-precision vehicle moving track cannot be accurately extracted; if high-precision collection is pursued, the area needs to be flown at a lower altitude, resulting in limited collection range and inability to effectively capture long-time continuous information of the vehicle during driving. In addition, the generation of the information of the test scene of the vehicle in the prior art mainly relies on manual labeling or simple rule extraction, which cannot effectively identify the dangerous interaction features in the motion process, resulting in insufficient accuracy of the extraction of key test scene information. Therefore, there is still a technical problem of low accuracy of the generation of the test scene of the vehicle.

[0003] At present, there is no effective solution to the technical problem of low accuracy of the generation of the information of the test scene of the vehicle. SUMMARY

[0004] The embodiments of the present application provide a vehicle test scene information generation method, device, processor and vehicle to at least solve the technical problem of low accuracy of the generation of the information of the test scene of the vehicle.

[0005] According to an aspect of the embodiments of the present application, a vehicle test scene information generation method is provided. The method can include: acquiring initial image information of a road where a vehicle is located by using a plurality of unmanned aerial vehicles, wherein the initial image information contains road scene data of the road where the vehicle is located; performing frame-level synchronization processing on the initial image information to obtain target image information, wherein the target image information is used to represent the time stamp and position information corresponding to each frame in the initial image information; analyzing the target image information by using a target detection model to obtain a moving track of the vehicle, wherein the target detection model is used to locate and classify the vehicle in the target image information; and generating information of a target test scene where the vehicle is located based on the moving track and the position information.

[0006] Optionally, based on the movement trajectory and the position information, information of a target test scene in which the vehicle is located is generated, including: determining, according to the movement trajectory and the position information, information of an initial test scene of the vehicle, wherein the information of the initial test scene is used to represent information of a corresponding test scene of the vehicle in the initial test scene; performing risk assessment on the vehicle in the initial test scene by using a driving risk field model to obtain a risk score, wherein the risk score is used to represent a risk degree of the vehicle in the initial test scene; in response to the risk score being greater than a risk score threshold, determining that the information of the initial test scene is the information of the target test scene; and the method further includes: storing target image information corresponding to the information of the target test scene.

[0007] Optionally, the target image information is analyzed by using a target detection model to obtain the movement trajectory of the vehicle, including: performing identification processing on the target image information by using the target detection model, and frame selecting and marking the vehicle in the target image information; tracking attribute information of the vehicle with the frame selection and marking by using a target tracking model, wherein the target tracking model is used at least for tracking the same vehicle in the target image information across multiple frames, and the attribute information is used to represent the speed and direction of the vehicle; and performing interpolation processing and smoothing processing on the attribute information to obtain the movement trajectory of the vehicle in a three-dimensional space.

[0008] Optionally, the initial image information of the road in which the vehicle is located is obtained by using a plurality of unmanned aerial vehicles, including: dividing the road in which the vehicle is located into a plurality of sub-regions, wherein each sub-region can contain one or more unmanned aerial vehicles; deploying the plurality of unmanned aerial vehicles in the plurality of sub-regions and determining setting parameters of the unmanned aerial vehicles, wherein the setting parameters are used to represent preset initial flight heights, waypoint coordinates and shooting angles of the unmanned aerial vehicles; and in response to the setting parameters satisfying a setting parameter threshold, obtaining the initial image information of the road in which the vehicle is located.

[0009] Optionally, the initial image information is subjected to frame-level synchronization processing to obtain the target image information, including: performing denoising processing and distortion correction processing on the initial image information to obtain processed initial image information, wherein the image quality of the processed initial image information is greater than that of the initial image information before processing; and performing frame-level synchronization processing on the processed initial image information to obtain the target image information.

[0010] Optionally, the method further includes: obtaining state information of the plurality of unmanned aerial vehicles, wherein the state information is used to represent at least the running states of the plurality of unmanned aerial vehicles; determining, based on the state information, an action space and a reward function of the unmanned aerial vehicles, wherein the action space is used to represent action options taken by the unmanned aerial vehicles when performing a task, and the reward function is used to quantify the action effect of the unmanned aerial vehicles; evaluating the state information of the unmanned aerial vehicles according to the action space and the reward function; and adjusting a flight path of the unmanned aerial vehicles according to the evaluation result.

[0011] Optionally, the method further includes: classifying the information of the target test scene to obtain a test scene type; adding a semantic label and a risk level to the information of the target test scene corresponding to the test scene type; determining a similarity between the information of the target test scene and information of a test scene in the test scene library according to the semantic label and the risk level; and saving the information of the target test scene to the test scene library in response to the similarity being less than or equal to a similarity threshold.

[0012] According to another aspect of the embodiments of the present application, a device for generating information of a test scene of a vehicle is further provided. The device can include: an acquisition unit configured to acquire initial image information of a road where the vehicle is located by using a plurality of unmanned aerial vehicles, wherein the initial image information comprises road scene data of the road where the vehicle is located; a processing unit configured to perform frame-level synchronization processing on the initial image information to obtain target image information, wherein the target image information is used to represent time stamps and position information corresponding to each frame in the initial image information; an identification unit configured to analyze the target image information by using a target detection model to obtain a moving track of the vehicle, wherein the target detection model is used to locate and classify the vehicle in the target image information; and a generation unit configured to generate information of a target test scene where the vehicle is located based on the moving track and the position information.

[0013] According to another aspect of the embodiments of the present application, a computer readable storage medium is further provided. The computer readable storage medium comprises a stored program, wherein the program, when executed, controls a device where the computer readable storage medium is located to perform the method of the embodiments of the present application.

[0014] According to another aspect of the embodiments of the present application, a processor is further provided. The processor is configured to execute a program, wherein the program, when executed, performs the method of the embodiments of the present application.

[0015] According to another aspect of the embodiments of the present application, a vehicle is further provided. The vehicle is configured to perform the method of the embodiments of the present application.

[0016] In the embodiment of the present application, if information of a target test scene where the vehicle is located needs to be generated, a plurality of unmanned aerial vehicles can be used to obtain initial image information of a road where the vehicle is located, wherein the initial image information contains road scene data of the road where the vehicle is located; frame-level synchronization processing is performed on the initial image information to obtain target image information, wherein the target image information is used to represent time stamps and position information corresponding to each frame in the initial image information; a target detection model is used to analyze the target image information to obtain a moving track of the vehicle, wherein the target detection model is used to locate and classify the vehicle in the target image information; and information of the target test scene where the vehicle is located is generated based on the moving track and the position information. In the embodiment of the present application, the initial image information of the vehicle and the road where the vehicle is located is captured from different heights, angles and positions by the plurality of unmanned aerial vehicles, which provides more comprehensive and three-dimensional road scene data and avoids the information loss and blind area problems that may be caused by a single perspective. Through frame-level synchronization processing, the initial image information captured by each unmanned aerial vehicle is accurately assigned with time stamps and position information, which ensures the spatio-temporal consistency between the initial image information collected by different unmanned aerial vehicles. The target detection model is used to analyze the target image information, which can quickly and accurately locate and classify the vehicle in the target image information, so as to generate the information of the target test scene where the vehicle is located based on the moving track and the position information, thereby solving the technical problem of low accuracy of the information generation of the test scene of the vehicle and achieving the technical effect of improving the accuracy of the information generation of the test scene of the vehicle. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and serve to explain the principles of the application. In the drawings:

[0018] Figure 1 FIG. 1 is a flowchart of a method for generating information of a test scene of a vehicle according to an embodiment of the present application;

[0019] Figure 2 FIG. 3 is a schematic diagram of an intelligent extraction system of a natural driving key test scene based on cooperation of a plurality of unmanned aerial vehicles according to an embodiment of the present application;

[0020] Figure 3 FIG. 4 is a schematic diagram of a key test scene recognition system according to an embodiment of the present application;

[0021] Figure 4 FIG. 5 is a schematic diagram of a device for generating information of a test scene of a vehicle according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] In the following, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative work should belong to the protection scope of the present application.

[0023] It should be noted that the terms "first", "second" and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in other than the order illustrated or described herein. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a list of steps or units need not be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to such processes, methods, products or devices.

[0024] According to the embodiments of the present application, an embodiment of a method for generating information of a test scene of a vehicle is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0025] Figure 1 is a flowchart of a method for generating information of a test scene of a vehicle according to the embodiments of the present application, as shown in Figure 1 The method can include the following steps:

[0026] In step S102, initial image information of a road where the vehicle is located is obtained by using a plurality of unmanned aerial vehicles.

[0027] In the technical solution provided in the above step S102 of the present application, the initial image information contains road scene data of the road where the vehicle is located.

[0028] In this embodiment, due to the limited field of view of the camera of a single UAV, especially when pursuing high-precision shooting, it is necessary to fly at a low height, which will greatly limit its coverage and make it difficult to capture the road scene data of the road where the vehicle is located during the driving process of the vehicle. If the UAV flies at a high altitude in order to expand the coverage, the vehicle in the initial image information will become very small, affecting the accuracy of vehicle behavior recognition and trajectory extraction. Therefore, in order to avoid missing key information due to single perspective, the initial image information of the road where the vehicle is located can be collected by multiple UAVs.

[0029] Optionally, multiple UAVs are used to cooperatively collect initial image information of the road where the vehicle is located from different directions and heights. The initial image information can include road conditions, activities of traffic participants, and surrounding environment, etc. The road conditions can include road surface conditions, road signs, traffic signal lights, etc. The activities of traffic participants can include the motion state of vehicles, pedestrians, bicycles, etc. The surrounding environment can include weather conditions, lighting conditions, building layout, etc.

[0030] It should be noted that the specific content of the initial image information described above is only for illustration, and is not specifically limited here. As long as the image information can be used to capture panoramic information of the road where the vehicle is located and provide a basis for generating information of the target test scene where the vehicle is located in the subsequent process, it is within the protection scope of the embodiments of the present application.

[0031] In step S104, frame-level synchronization processing is performed on the initial image information to obtain target image information.

[0032] In the technical solution provided in step S104 of the present application, after obtaining the initial image information of the road where the vehicle is located by using multiple UAVs, frame-level synchronization processing can be performed on the initial image information to obtain target image information. The target image information is used to represent the time stamp and position information corresponding to each frame of the initial image information.

[0033] In this embodiment, in order to ensure that the initial image information obtained from different UAVs is accurately synchronized in time and space, frame-level synchronization processing can be performed on the initial image information to obtain target image information. The target image information can not only include the initial image information itself, but also carry the time stamp and accurate three-dimensional position information (position information) corresponding to each frame of the initial image information.

[0034] Optionally, the initial image information can be subjected to frame-level synchronization processing by using a network time protocol or a dedicated time synchronization server, so that the timestamps of the initial image information recorded by each unmanned aerial vehicle can be kept consistent even if the unmanned aerial vehicles are dispersed in different locations. Then, three-dimensional coordinate information of each frame of initial image information can be recorded by using an integrated global positioning system (GPS) and an inertial navigation system (INS) to obtain target image information. During the synchronization processing, the position information can be further calibrated to ensure that the relative positions between the target image information can be accurately determined even in the case of rapid movement or complex environmental conditions.

[0035] Optionally, the accurate timestamps and position information of the target image information are crucial for three-dimensional scene reconstruction, which can accurately restore the movement trajectory of the vehicle and the relative positions of the traffic participants, and provide a basis for subsequent target detection and trajectory tracking.

[0036] In step S106, a target detection model is used to analyze the target image information to obtain the movement trajectory of the vehicle.

[0037] In the technical solution of step S106 described above, after the frame-level synchronization processing of the initial image information is completed to obtain the target image information, a target detection model can be used to analyze the target image information to obtain the movement trajectory of the vehicle. The target detection model is used to locate and classify the vehicle in the target image information.

[0038] In this embodiment, after the frame-level synchronization processing is completed, a target detection model, such as a You Only Look Once (YOLO) target detection algorithm or a Faster Region-based Convolutional Neural Network (Faster R-CNN) target detection algorithm, can be used to process each frame of the target image information to automatically identify the position of the vehicle in the target image information, classify the vehicle (such as a small car, a truck, a bus, etc.), and output the bounding box coordinates of the vehicle to provide an accurate starting point for subsequent trajectory tracking.

[0039] Optionally, the target image information subjected to the synchronization processing can be subjected to deep analysis by using the target detection model, which not only accurately identifies and locates the vehicle in the target image information, but also constructs a complete movement trajectory of the vehicle. The target detection model can automatically identify and locate the vehicle, which significantly improves the detection accuracy and efficiency and reduces the data processing time and labor cost compared with manual annotation or traditional rule-based methods.

[0040] Step S108, based on the movement trajectory and position information, generate information of the target test scene where the vehicle is located.

[0041] In the technical solution of step S108 of the present application described above, after analyzing the target image information using the target detection model to obtain the movement trajectory of the vehicle, the information of the target test scene where the vehicle is located can be generated based on the movement trajectory and position information.

[0042] In this embodiment, according to the three-dimensional position information, combined with the movement trajectory of the vehicle, the dynamic model of the vehicle and its surrounding environment in the target test scene can be reconstructed through three-dimensional modeling technology, such as the motion state of the vehicle, the distribution of pedestrians, the road conditions, etc. The environmental factors (such as light, weather, obstacles, etc.) captured in the target image information can be combined with the movement trajectory of the vehicle to create a test scene description containing rich details. For example, the specific behavior pattern of the vehicle on a wet and slippery road in rainy weather. Key features such as the relative position between vehicles, speed differences, changes in driving direction, etc. are extracted from the reconstructed test scene description.

[0043] Optionally, since the target test scene information contains rich environmental and traffic participant behavior information, not only limited to vehicle trajectory, but also including other factors that may affect the decision-making of the autonomous driving system, such as pedestrian flow, obstacle existence, traffic signals, etc., the authenticity and integrity of the information of the target test scene can be ensured.

[0044] The steps S102 to S108 of the present application, if the information of the target test scene where the vehicle is located needs to be generated, a plurality of unmanned aerial vehicles can be used to obtain initial image information of the road where the vehicle is located, wherein the initial image information contains road scene data of the road where the vehicle is located; frame-level synchronization processing is performed on the initial image information to obtain target image information, wherein the target image information is used to represent the time stamp and position information corresponding to each frame in the initial image information; a target detection model is used to analyze the target image information to obtain the moving track of the vehicle, wherein the target detection model is used to locate and classify the vehicle in the target image information; and the information of the target test scene where the vehicle is located is generated based on the moving track and the position information. In the embodiment of the present application, the initial image information of the vehicle and the road where the vehicle is located is captured from different heights, angles and positions by the plurality of unmanned aerial vehicles, which provides more comprehensive and three-dimensional road scene data and avoids the information loss and blind area problem caused by a single perspective; through frame-level synchronization processing, the initial image information captured by each unmanned aerial vehicle is accurately assigned with a time stamp and position information, which ensures the spatio-temporal consistency between the initial image information collected by different unmanned aerial vehicles; the target detection model is used to analyze the target image information, which can quickly and accurately locate and classify the vehicle in the target image information, so as to generate the information of the target test scene where the vehicle is located based on the moving track and the position information, solve the technical problem of low accuracy of the information generation of the test scene of the vehicle, and achieve the technical effect of improving the accuracy of the information generation of the test scene of the vehicle.

[0045] The above method of the embodiment will be further introduced below.

[0046] As an optional embodiment, in step S108, the information of the target test scene where the vehicle is located is generated based on the moving track and the position information, including: determining the initial test scene information of the vehicle according to the moving track and the position information, wherein the initial test scene information is used to represent the test scene information corresponding to the vehicle in the initial test scene; using a driving risk scene model to perform risk assessment on the vehicle in the initial test scene to obtain a risk score, wherein the risk score is used to represent the risk degree of the vehicle in the initial test scene; and in response to the risk score being greater than a risk score threshold, determining that the initial test scene information is the target test scene information; the method further includes: storing the target image information corresponding to the target test scene information.

[0047] In this embodiment, in the process of generating the information of the target test scene where the vehicle is located based on the movement trajectory and the location information, the movement trajectory of the vehicle can be analyzed, and in combination with the specific geographic location (location information) where the vehicle is located, a series of possible test scene information (initial test scene information) can be determined. For example, when the vehicle frequently changes lanes or performs complex driving operations at intersections, these test scene information can be initially identified as initial test scene information.

[0048] Optionally, after determining the initial test scene information of the vehicle, the driving risk field model can be used to perform risk assessment on the vehicle in the initial test scene. The driving risk field model can define a risk function for quantitatively modeling the potential conflict relationship between traffic participants.

[0049] The relative distance d between vehicle i and vehicle j ij (t) can be represented by the following formula:

[0050]

[0051] Wherein, x i (t), y i (t) can be used to represent the position coordinates of vehicle i, x j (t), y j (t) can be used to represent the position coordinates of vehicle j.

[0052] The relative speed v ij (t) can be represented by the following formula:

[0053] v ij (t) = v i (t) - v j (t)

[0054] Wherein, v i (t) can be used to represent the speed of vehicle i, v j (t) can be used to represent the speed of vehicle j.

[0055] The local risk value R ij (t) can be represented by the following formula:

[0056]

[0057] Wherein, a and β can be used to represent the experience weight coefficient, v max may be used to represent the maximum relative speed threshold, t c may be used to represent the time to collision (Time to Collision). The time to collision can be represented by the following formula:

[0058]

[0059] The risk matrix R(t) can be represented by the following formula:

[0060]

[0061] wherein, may be used to represent a set of n*n order matrices.

[0062] Alternatively, each vehicle is regarded as a node, and the inter-vehicle risk value is regarded as an edge weight, a vehicle interaction graph can be established, and the vehicle interaction graph can be represented by the following formula:

[0063]

[0064] wherein, θ can be used to represent a set threshold.

[0065] Alternatively, a graph convolutional network (GCN) is used to extract the inter-vehicle interaction features; then, a long short-term memory (LSTM) module can be accessed to model the time evolution of traffic behavior, so as to output the risk score at each time.

[0066] The graph convolutional layer update can be represented by the following formula:

[0067]

[0068] wherein, may be used to represent an adjacency matrix with a self-loop, and I can be used to represent a unit matrix; may be used to represent a corresponding degree matrix, and W (l) may be used to represent the learnable parameters of the lth layer, and σ can be used to represent an activation function.

[0069] Through the above steps, when the risk score is greater than the risk score threshold, the information (information of the target test scene) of the potential key scene can be marked.

[0070] As an optional embodiment, in step S106, the target image information is analyzed by using the target detection model to obtain the moving track of the vehicle, including: the target image information is recognized and processed by using the target detection model, and the vehicle in the target image information is marked by framing; the attribute information of the vehicle with the framing mark is tracked by using the target tracking model, wherein the target tracking model is used at least for tracking the same vehicle in the target image information across multiple frames, and the attribute information is used to represent the speed and direction of the vehicle; the attribute information is processed by interpolation and smoothing to obtain the moving track of the vehicle in the three-dimensional space.

[0071] In this embodiment, in the process of analyzing the target image information by using the target detection model to obtain the moving track of the vehicle, the target detection model (such as YOLO or Faster R-CNN) can be applied to process each frame of target image information to identify and frame each vehicle in the target image information. Then, the target tracking model can be used to continuously track the framed vehicles, which not only limits the position tracking in the visual level, but also involves the continuous monitoring of the vehicle attribute information (such as speed, direction, etc.). The target tracking model ensures that the same vehicle can be continuously identified and tracked even in complex environments and occlusion conditions through multi-frame processing, so as to collect a continuous sequence of vehicle attribute information.

[0072] Optionally, due to the noise and discontinuity that may exist in the actual image acquisition process, the attribute information may contain errors. In order to improve the accuracy and continuity of the moving track, the collected attribute information can be interpolated to fill in possible missing values, and smoothed to eliminate abnormal points and ensure the smoothness and continuity of the data.

[0073] Optionally, after the attribute information is interpolated and smoothed, the vehicle attribute information after interpolation and smoothing is converted into three-dimensional space (x, y, z) coordinates, wherein the z coordinate can be further refined by the vertical shooting angle of the unmanned aerial vehicle and the GPS data, so as to combine the time stamp of each frame to construct the continuous moving track of the vehicle in the three-dimensional space.

[0074] As an optional embodiment, in step S102, the initial image information of the road where the vehicle is located is obtained by using a plurality of unmanned aerial vehicles, including: dividing the road where the vehicle is located into a plurality of sub-regions, wherein each sub-region can contain one or more unmanned aerial vehicles; deploying a plurality of unmanned aerial vehicles in the plurality of sub-regions and determining the setting parameters of the unmanned aerial vehicles, wherein the setting parameters are used to represent the preset initial flight height, waypoint coordinates and shooting angle of the unmanned aerial vehicles; and in response to the setting parameters satisfying the setting parameter threshold, obtaining the initial image information of the road where the vehicle is located.

[0075] In this embodiment, in the process of obtaining the initial image information of the road where the vehicle is located by using a plurality of unmanned aerial vehicles, a multi-unmanned aerial vehicle cooperative acquisition system with spatial coverage capability can be constructed to ensure the high precision and wide coverage of the traffic scene data. The area of the road where the vehicle is located can be divided into a plurality of sub-regions according to the length, width and traffic density of the target test section, and each sub-region is responsible for acquisition by one or more unmanned aerial vehicles.

[0076] Optionally, multiple UAVs can be deployed in multiple sub-regions, and the UAVs can be equipped with high-definition cameras, GPS positioning, inertial navigation INS, and wireless communication modules. Each UAV can be preset with an initial flight height (80-150 meters), waypoint coordinates, and a shooting angle to obtain initial image information of the road where the vehicle is located.

[0077] Optionally, a communication and synchronization mechanism can be established, and all UAVs exchange state information through ad hoc network communication. A unified timestamp server can be set up to realize time synchronization of image frames and trajectory data. A cooperative acquisition strategy can also be initialized to set a minimum safe distance between UAVs to prevent flight conflicts. A UAV cluster formation control logic can be started to ensure that the overall acquisition range has no blind area.

[0078] As an optional embodiment, in step S104, the initial image information is subjected to frame-level synchronization processing to obtain target image information, including: performing denoising processing and distortion correction processing on the initial image information to obtain processed initial image information, wherein the image quality of the processed initial image information is greater than that of the initial image information before processing; and performing frame-level synchronization processing on the processed initial image information to obtain the target image information.

[0079] In this embodiment, during the frame-level synchronization processing of the initial image information to obtain the target image information, each frame of initial image information can be subjected to filtering denoising, color correction, and lens distortion correction processing. For example, the above processing can be completed using an open source computer vision library (such as OpenCV).

[0080] Optionally, after the above processing is completed, the initial image information collected by all UAVs can be subjected to frame-level synchronization using a network time protocol (NTP) or a dedicated synchronization device to match the timestamps and GPS coordinates (position information) corresponding to each frame to obtain the target image information.

[0081] As an optional embodiment, the method further includes: obtaining state information of the multiple UAVs, wherein the state information is used to at least represent running states of the multiple UAVs; determining an action space and a reward function of the UAVs based on the state information, wherein the action space is used to represent action options taken by the UAVs when performing tasks, and the reward function is used to quantify action effects of the UAVs; evaluating the state information of the UAVs according to the action space and the reward function; and adjusting a flight path of the UAVs according to an evaluation result.

[0082] In this embodiment, the agent state space can be defined, each UAV is taken as an independent agent, and the state information of multiple UAVs is obtained, such as the current position, speed, attitude, battery capacity, field of view coverage area, etc. Then, the action space and the reward function can be designed.

[0083] Optionally, the action space can include flight direction adjustment, height adjustment, camera pitch angle change. The reward function design can be +1, indicating successful capture of key vehicle behavior; the reward function design can be -1, indicating entering the occlusion area or collecting blur; the reward function design can be -0.5, indicating too close to other UAVs.

[0084] Optionally, according to the current state information of the UAV and the pre-defined action space and reward function, the current behavior of the UAV can be evaluated in real time by using a reinforcement learning algorithm, such as the Proximal Policy Optimization (PPO) algorithm, and the expected reward value of each action option under the current state can be calculated. According to the evaluation result, the next action of the UAV can be decided. Through continuous iteration, the appropriate action strategy under different states can be learned, so as to dynamically adjust the flight path and data collection mode of the UAV, so as to maximize the completion quality and efficiency of the task.

[0085] As an optional embodiment, the method further includes: classifying the information of the target test scene to obtain a test scene type; adding a semantic label and a risk level to the information of the target test scene corresponding to the test scene type; determining a similarity between the information of the target test scene and the information of the test scene in the test scene library according to the semantic label and the risk level; and saving the information of the target test scene to the test scene library in response to the similarity being less than or equal to a similarity threshold.

[0086] In this embodiment, after obtaining the information of the target test scene, the information of the target test scene can be classified, such as classified according to the scene type (such as intersection, lane change, emergency braking, etc.); the semantic label and the risk level can be automatically added based on the visual large language model.

[0087] Optionally, a vector similarity matching algorithm (cosine similarity) is used to determine whether the information of the target test scene already exists; if the similarity is less than or equal to a similarity threshold, the information of the target test scene is saved to the test scene library.

[0088] Optionally, the information of the target test scene can be stored in a concise data format (JavaScript Object Notation, referred to as JSON) format, including video clip links, vehicle trajectory data, risk scores, scene type labels, timestamp and geographic coordinate information, etc. Incremental learning and online updating can be supported, and the test scene library can be dynamically adjusted and optimized with the obtained information of the target test scene, so that the test scene library can adapt to the changing traffic environment and the needs of autonomous driving technology.

[0089] In the embodiment of the application, if it is necessary to generate information of a target test scene in which a vehicle is located, a plurality of unmanned aerial vehicles can be used to obtain initial image information of a road in which the vehicle is located, wherein the initial image information contains road scene data of the road in which the vehicle is located; frame-level synchronization processing is performed on the initial image information to obtain target image information, wherein the target image information is used to represent timestamp and position information corresponding to each frame in the initial image information; a target detection model is used to analyze the target image information to obtain a moving trajectory of the vehicle, wherein the target detection model is used to locate and classify the vehicle in the target image information; and information of the target test scene in which the vehicle is located is generated based on the moving trajectory and the position information. In the embodiment of the application, the initial image information of the vehicle and the road in which the vehicle is located is captured from different heights, angles and positions by the plurality of unmanned aerial vehicles, so that more comprehensive and three-dimensional road scene data is provided, and the problem of information loss and blind area caused by a single perspective is avoided. Through frame-level synchronization processing, the initial image information captured by each unmanned aerial vehicle is accurately assigned with timestamp and position information, so that the spatio-temporal consistency between the initial image information captured by different unmanned aerial vehicles is ensured. The target detection model is used to analyze the target image information, so that the vehicle in the target image information can be quickly and accurately located and classified, and then the information of the target test scene in which the vehicle is located is generated based on the moving trajectory and the position information, thereby solving the technical problem of low accuracy of the information of the test scene of the vehicle, and achieving the technical effect of improving the accuracy of the information of the test scene of the vehicle.

[0090] The technical solutions of the embodiments of the application will be described below in conjunction with preferred embodiments.

[0091] Currently, the driving scene data collection means often adopts a single unmanned aerial vehicle for aerial photography to obtain vehicle motion trajectory and behavior data of other traffic participants. The above-mentioned method can provide real-time traffic scene information to a certain extent, but is severely restricted by the performance and operation of the unmanned aerial vehicle. On the one hand, the single unmanned aerial vehicle is limited by the flight height and camera field of view, and cannot simultaneously consider large-scale coverage and high-precision image collection. If high-altitude photography is selected to expand the field of view, a large amount of image details will be lost, and the accurate identification of the vehicle trajectory becomes a problem; on the contrary, low-altitude photography can ensure image definition, but can only cover a limited area, and it is difficult to capture continuous motion information of the vehicle in a long period and a large range, especially in a complex traffic scene, the limitation is more prominent. On the other hand, after the single unmanned aerial vehicle collects data, the identification and extraction of key test scenes mainly rely on manual annotation or simple rule setting, and lack in-depth analysis of the dynamic interaction characteristics of traffic participants, thereby leading to low scene extraction efficiency, excessive redundant data, and inability to accurately identify the key test scenes challenging the automatic driving system. Therefore, there is still a technical problem of low accuracy of the generation of the test scene of the vehicle.

[0092] The embodiment of the present application proposes a natural driving key test scene intelligent extraction method based on multi-unmanned aerial vehicle cooperation, constructs a large-scale data collection network through a multi-unmanned aerial vehicle cooperative collection scheme, realizes scene reconstruction in a three-dimensional coordinate system, constructs an unmanned aerial vehicle scheduling and path planning scheme based on a PPO algorithm, realizes optimal state allocation of multi-unmanned aerial vehicles, and further designs a key scene evaluation system integrating driving risk field theory and graph convolution time series analysis, guides unmanned aerial vehicle state control based on the risk field, improves key scene identification, collection and extraction efficiency based on convolution time series analysis, achieves cooperative optimization of unmanned aerial vehicle aerial photography control and automatic driving test scene extraction, solves the technical problem of low accuracy of the information generation of the test scene of the vehicle, and realizes the technical effect of improving the accuracy of the information generation of the test scene of the vehicle.

[0093] The embodiment of the present application will be further introduced below.

[0094] Figure 2 It is a schematic diagram of a natural driving key test scene intelligent extraction system based on multi-unmanned aerial vehicle cooperation according to the embodiment of the present application. The natural driving key test scene intelligent extraction system 200 based on multi-unmanned aerial vehicle cooperation comprises a multi-unmanned aerial vehicle cooperative collection network deployment module 201, a unmanned aerial vehicle path planning module 202 based on a PPO algorithm, a three-dimensional scene reconstruction and data preprocessing module 203, a key test scene identification and extraction module 204, and a test scene library construction and update module 205.

[0095] The multi-UAV cooperative collection network deployment module 201, in which a multi-UAV cooperative collection system with spatial coverage capability can be constructed to ensure high-precision and wide-area coverage of traffic scene data. The entire region can be divided into several sub-regions according to the target test road length, width and traffic density, and each sub-region is responsible for collection by one or more UAVs; a UAV platform can be deployed, and a UAV with a high-definition camera, GPS positioning, inertial navigation (INS) and wireless communication module can be used.

[0096] Optionally, each UAV can be preset with an initial flight height (80-150 meters), waypoint coordinates and shooting angle; a communication and synchronization mechanism can be established, and all UAVs exchange state information through ad hoc network communication; a unified timestamp server can be set up to realize time synchronization of image frames and trajectory data; a cooperative collection strategy can be initialized to set the minimum safe distance between UAVs to prevent flight conflicts; and UAV cluster formation control logic can be started to ensure that the overall collection range has no blind area.

[0097] The UAV path planning module 202 based on the PPO algorithm, in which the agent state space can be defined, and each UAV is an independent agent, whose state includes: current position, speed, attitude, battery capacity, field of view coverage area, etc.; the action space and reward function can be designed. The action space includes: flight direction adjustment, height adjustment, camera pitch angle change. The reward function is designed as: +1: successfully capturing key vehicle behavior; -1: entering an occluded area or collecting a blur; -0.5: too close to other UAVs; the PPO model can be trained, and preliminary training can be performed in a simulated environment, and training samples can be generated using simulated traffic flow data; a pre-trained model can be loaded before actual deployment to support online fine-tuning. Path planning and task scheduling can be performed. Real-time reception of UAV state information updates the global environment state; the corresponding action is output through the PPO policy network to control the UAV flight path; multi-UAV task switching is supported, and the UAV automatically changes positions when the battery is low, thereby realizing dynamic path optimization using the PPO algorithm in reinforcement learning to improve collection efficiency and quality.

[0098] The three-dimensional scene reconstruction and data preprocessing module 203, in which image denoising and distortion correction can be performed. Each frame of image is filtered and denoised, color corrected and lens distortion corrected; image preprocessing is completed using OpenCV tools. Temporal and spatial synchronization and frame alignment can be performed. NTP protocol or special synchronization equipment is used for frame-level synchronization of all images collected by the unmanned aerial vehicle; the time stamp and GPS coordinate corresponding to each frame are matched; target detection and tracking can be performed. YOLO, Faster R-CNN and other target detection algorithms are applied to identify vehicles, pedestrians and other road users; combined with online real-time multi-target tracking algorithm (Deeply Learned Online and Realtime Tracking, referred to as DeepSORT) for cross-frame tracking; the identifier (Identifier, referred to as ID), category and trajectory point sequence of each target are output.

[0099] The key test scene identification and extraction module 204, in which a driving risk field model can be constructed, and a risk function can be defined for quantitatively modeling the potential conflict relationship between road users.

[0100] The relative distance d between vehicle i and vehicle j ij (t) can be represented by the following formula:

[0101]

[0102] wherein x i (t), y i (t) can be used to represent the position coordinates of vehicle i, x j (t), y j (t) can be used to represent the position coordinates of vehicle j.

[0103] The relative speed v ij (t) can be represented by the following formula:

[0104] v ij (t) = v i (t) - v j (t)

[0105] wherein v i (t) can be used to represent the speed of vehicle i, and v j (t) can be used to represent the speed of vehicle j.

[0106] The local risk value R ij (t) can be represented by the following formula:

[0107]

[0108] wherein, a and b can be used to represent the experience weight coefficient, v max may be used to represent the maximum relative speed threshold, t c may be used to represent the time to collision. The time to collision can be represented by the following formula:

[0109]

[0110] The risk matrix R(t) can be represented by the following formula:

[0111]

[0112] wherein, may be used to represent a set of n*n order matrices.

[0113] Optionally, each vehicle is regarded as a node, and the inter-vehicle risk value is regarded as an edge weight, and a vehicle interaction graph can be established, and the vehicle interaction graph can be represented by the following formula:

[0114]

[0115] wherein, θ can be used to represent a set threshold.

[0116] Optionally, a graph convolutional network (GCN) is used to extract the inter-vehicle interaction features; then, a long short-term memory (LSTM) module can be accessed to model the time sequence evolution of the traffic behavior, so as to output the risk score at each time.

[0117] The graph convolutional layer update can be represented by the following formula:

[0118]

[0119] wherein, may be used to represent an adjacency matrix with a self-loop, and I can be used to represent a unit matrix; may be used to represent a corresponding degree matrix, W (l) may be used to represent the learnable parameters of the lth layer, and σ can be used to represent an activation function.

[0120] When the scene risk score exceeds the set threshold, it is marked as a potential key scene; the images, trajectories, and semantic labels in the corresponding time period are packaged and saved.

[0121] The test scene library construction and update module 205, in which scene classification and labeling can be performed. Classification is performed according to scene types (such as intersections, lane changes, emergency braking, etc.); semantic labels and risk levels are automatically added based on a visual large language model. Scene deduplication and categorization can be performed. A vector similarity matching algorithm (cosine similarity) is used to determine whether a new scene already exists; if the similarity is higher than a set threshold, it is not duplicated in the library; a structured database can be constructed. Scene information is stored in JSON format, including: video segment link, vehicle trajectory data, risk score, scene type label, timestamp, and geographic coordinate information; incremental learning and online updating can be supported. Newly collected data is regularly uploaded to the server; the graph convolution model and risk prediction model are updated using an incremental learning mechanism.

[0122] Figure 3 is a schematic diagram of a key test scene recognition system according to an embodiment of the application, as shown, comprising: an input layer 301, a preprocessing layer 302, a target detection layer 303, a trajectory extraction layer 304, a risk modeling layer 305, a graph structure layer 306, a graph convolution + time series modeling layer 307, and an output layer 308. Figure 3

[0123] The input layer 301 can input video images, GPS / INS data.

[0124] The preprocessing layer 302 can perform denoising, distortion correction, and frame synchronization.

[0125] The target detection layer 303 is used to box vehicles using YOLO / Faster R-CNN.

[0126] The trajectory extraction layer 304 is used to track using DeepSORT, outputting ID and trajectory.

[0127] The risk modeling layer 305 is used to calculate collision time and construct a risk matrix.

[0128] The graph structure layer 306 is used to construct a vehicle interaction graph.

[0129] The graph convolution + time series modeling layer 307 is used to extract features using GCN and model evolution using LSTM.

[0130] The output layer 308 scores the scene, and when the score exceeds a score threshold, it is extracted as a key scene.

[0131] A typical intersection of a city trunk road and a branch road is a typical crossroads, with dense traffic and frequent high-risk behaviors such as lane changes and emergency braking. The following example of field deployment in this scenario verifies its ability to recognize key test scenes.

[0132] ​Table 1 is a multi-UAV deployment table, as shown in Table 1, different main tasks under different flight altitudes and different coverage directions of different UAVs are shown.

[0133] Table 1 is a multi-UAV deployment table

[0134] UAV number Flight height Covering direction Main task UAV-01 100m Straight ahead of the main road Capture straight vehicle trajectory UAV-02 120m Left side of the branch road Monitor left turn and merging vehicles UAV-03 120m Right side of the branch road Capture right turn vehicle behavior UAV-04 80m Center of the intersection Take pictures of global interaction

[0135] Optionally, data acquisition and synchronization: all UAV video frame rates can be uniformly set to 30fps; frame-level synchronization can be performed using an NTP time server; the acquisition time can be, for example, 30 minutes; the total data volume can be about 60GB of video + corresponding GPS / INS data.

[0136] Optionally, scene modeling and trajectory extraction: YOLOv5 can be used for vehicle detection; DeepSORT can be used for cross-frame tracking; vehicle ID, coordinates, speed, and acceleration can be extracted; and complete trajectory information of each vehicle can be output.

[0137] Optionally, risk field and vehicle interaction atlas construction: the collision time between each pair of vehicles can be calculated; a risk threshold θ = 3s can be set to construct an adjacency matrix; each vehicle can be a node in the graph; node features can include speed, acceleration, direction angle, lane offset, etc.

[0138] Optionally, graph convolution and time series modeling: GCN can be used to extract graph structure features; two layers of LSTM can be used to model time series evolution; an overall scene risk score s(t) at each time can be output; a threshold τ = 0.75 can be set, and if it is exceeded, it is marked as a potential key scene.

[0139] According to an embodiment of the present application, a vehicle test scene information generation device is also provided. It should be noted that the vehicle test scene information generation device can be used to execute the vehicle test scene information generation method in the embodiment.

[0140] Figure 4 is a schematic diagram of a vehicle test scene information generation device according to an embodiment of the present application. As shown in Figure 4 The vehicle test scene information generation device 400 can include an acquisition unit 402, a processing unit 404, an identification unit 406, and a generation unit 408.

[0141] The acquisition unit 402 is configured to acquire initial image information of a road where a vehicle is located by using a plurality of UAVs, wherein the initial image information includes road scene data of the road where the vehicle is located.

[0142] The processing unit 404 is configured to perform frame-level synchronization processing on the initial image information to obtain target image information, wherein the target image information is used to represent time stamps and position information corresponding to each frame in the initial image information.

[0143] The recognition unit 406 is configured to analyze the target image information by using a target detection model to obtain a moving track of the vehicle, wherein the target detection model is used to locate and classify the vehicle in the target image information.

[0144] The generation unit 408 is configured to generate information of a target test scene in which the vehicle is located based on the moving track and the position information.

[0145] Optionally, the generation unit 408 comprises: a first determination module configured to determine information of an initial test scene of the vehicle according to the moving track and the position information, wherein the information of the initial test scene is used to represent information of a test scene corresponding to the vehicle in the initial test scene; an evaluation module configured to perform risk evaluation on the vehicle in the initial test scene by using a driving risk scene model to obtain a risk score, wherein the risk score is used to represent a risk degree of the vehicle in the initial test scene; and a second determination module configured to determine that the information of the initial test scene is the information of the target test scene in response to the risk score being greater than a risk score threshold. The method further comprises: storing target image information corresponding to the information of the target test scene.

[0146] Optionally, the recognition unit 406 comprises: a recognition module configured to perform recognition processing on the target image information by using the target detection model, and frame the vehicle in the target image information; a tracking module configured to track attribute information of the framed vehicle by using a target tracking model, wherein the target tracking model is used to track the same vehicle in the target image information across multiple frames, and the attribute information is used to represent a speed and a direction of the vehicle; and a first processing module configured to perform interpolation processing and smoothing processing on the attribute information to obtain a moving track of the vehicle in a three-dimensional space.

[0147] Optionally, the acquisition unit 402 comprises: a division module configured to divide a road in which the vehicle is located into a plurality of sub-regions, wherein each sub-region can contain one or more unmanned aerial vehicles; a third determination module configured to deploy a plurality of unmanned aerial vehicles in the plurality of sub-regions, and determine setting parameters of the unmanned aerial vehicles, wherein the setting parameters are used to represent preset initial flight heights, waypoint coordinates and shooting angles of the unmanned aerial vehicles; and a first acquisition module configured to acquire initial image information of the road in which the vehicle is located in response to the setting parameters satisfying a setting parameter threshold.

[0148] Optionally, the processing unit 404 comprises: a second processing module, configured to perform denoising processing and distortion correction processing on the initial image information to obtain processed initial image information, wherein the image quality of the processed initial image information is greater than the image quality of the initial image information before processing; and a third processing module, configured to perform frame-level synchronization processing on the processed initial image information to obtain the target image information.

[0149] Optionally, the apparatus further comprises: a second acquisition module, configured to acquire state information of the plurality of unmanned aerial vehicles, wherein the state information is used to at least represent running states of the plurality of unmanned aerial vehicles; a third determination module, configured to determine, based on the state information, an action space and a reward function of the unmanned aerial vehicle, wherein the action space is used to represent action options taken by the unmanned aerial vehicle when performing a task, and the reward function is used to quantify an action effect of the unmanned aerial vehicle; an evaluation module, configured to evaluate the state information of the unmanned aerial vehicle according to the action space and the reward function; and an adjustment module, configured to adjust a flight path of the unmanned aerial vehicle according to an evaluation result.

[0150] Optionally, the apparatus further comprises: a classification module, configured to classify information of the target test scene to obtain a test scene type; an adding module, configured to add a semantic label and a risk level to the information of the target test scene corresponding to the test scene type; a fourth determination module, configured to determine a similarity between the information of the target test scene and information of a test scene in a test scene library according to the semantic label and the risk level; and a saving module, configured to save the information of the target test scene to the test scene library in response to the similarity being less than or equal to a similarity threshold.

[0151] In the embodiment of the present application, the initial image information of the road where the vehicle is located is acquired by the acquisition unit 402 using the plurality of unmanned aerial vehicles, wherein the initial image information contains road scene data of the road where the vehicle is located; the target image information is obtained by performing frame-level synchronization processing on the initial image information by the processing unit 404, wherein the target image information is used to represent the time stamp and the position information corresponding to each frame in the initial image information; the moving track of the vehicle is obtained by analyzing the target image information using the target detection model by the identification unit 406, wherein the target detection model is used to locate and classify the vehicle in the target image information; and the information of the target test scene where the vehicle is located is generated based on the moving track and the position information by the generation unit 408, thereby solving the technical problem of low accuracy of the information generation of the test scene of the vehicle and achieving the technical effect of improving the accuracy of the information generation of the test scene of the vehicle.

[0152] According to the embodiment of the present application, a computer readable storage medium is also provided, which comprises a stored program, wherein the program executes the method in the embodiment.

[0153] According to an embodiment of the present application, a processor is also provided, which is configured to run a program, wherein the program is configured to perform the method in the embodiments when running.

[0154] According to an embodiment of the present application, a vehicle is also provided, which is configured to perform the method in the embodiments of the present application.

[0155] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0156] In several embodiments provided in the present application, it should be understood that the disclosed technical contents can be implemented by other manners. Among them, the above-described device embodiments are only schematic, for example, the division of the units can be a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or modules shown or discussed can be indirect coupling or communication connection through some interfaces, units or modules, which can be electrical or other forms.

[0157] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment.

[0158] In addition, each functional unit in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0159] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0160] The above only describes the preferred embodiments of the present application. It should be noted that, for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should also be considered as the protection scope of the present application.

Claims

1. A method of generating information of a test scenario of a vehicle, characterized by, The method comprises the following steps: acquiring initial image information of a road where a vehicle is located by using a plurality of unmanned aerial vehicles, wherein the initial image information contains road scene data of the road where the vehicle is located; performing frame-level synchronization processing on the initial image information to obtain target image information, wherein the target image information is used to represent time stamps and position information corresponding to each frame in the initial image information; analyzing the target image information by using a target detection model to obtain a moving track of the vehicle, wherein the target detection model is used to locate and classify the vehicle in the target image information; generating information of a target test scene where the vehicle is located based on the moving track and the position information.

2. The method of claim 1, wherein, Generating information of a target test scene where the vehicle is located based on the moving track and the position information comprises: determining information of an initial test scene of the vehicle according to the moving track and the position information, wherein the information of the initial test scene is used to represent information of a test scene corresponding to the vehicle in the initial test scene; performing risk assessment on the vehicle in the initial test scene by using a driving risk scene model to obtain a risk score, wherein the risk score is used to represent a risk degree of the vehicle in the initial test scene; in response to the risk score being greater than a risk score threshold, determining that the information of the initial test scene is the information of the target test scene. The method further comprises: storing the target image information corresponding to the information of the target test scene.

3. The method of claim 2, wherein, Analyzing the target image information by using a target detection model to obtain a moving track of the vehicle comprises: performing identification processing on the target image information by using the target detection model, and performing frame selection marking on the vehicle in the target image information; tracking attribute information of the vehicle with frame selection marking by using a target tracking model, wherein the target tracking model is used to track the same vehicle in the target image information across multiple frames, and the attribute information is used to represent a speed and a direction of the vehicle; performing interpolation processing and smoothing processing on the attribute information to obtain the moving track of the vehicle in a three-dimensional space.

4. The method of claim 1, wherein, Acquiring initial image information of a road where a vehicle is located by using a plurality of unmanned aerial vehicles comprises: dividing the road where the vehicle is located into a plurality of sub-regions, wherein each sub-region can contain one or more unmanned aerial vehicles; deploying a plurality of unmanned aerial vehicles in a plurality of sub-regions and determining setting parameters of the unmanned aerial vehicles, wherein the setting parameters are used to represent a preset initial flight height, a waypoint coordinate, and a shooting angle of the unmanned aerial vehicles; in response to the setting parameters satisfying a setting parameter threshold, acquiring the initial image information of the road where the vehicle is located.

5. The method of claim 4, wherein, Performing frame-level synchronization processing on the initial image information to obtain target image information comprises: performing denoising processing and distortion correction processing on the initial image information to obtain processed initial image information, wherein an image quality of the processed initial image information is greater than an image quality of the initial image information before processing. The processed initial image information is subjected to frame-level synchronization processing to obtain the target image information.

6. The method of claim 1, wherein, The method further comprises: acquiring state information of the plurality of unmanned aerial vehicles, wherein the state information is used to represent at least the operating states of the plurality of unmanned aerial vehicles; based on the state information, determining the action space and the reward function of the unmanned aerial vehicle, wherein the action space is used to represent the action options taken by the unmanned aerial vehicle when performing a task, and the reward function is used to quantify the action effect of the unmanned aerial vehicle; according to the action space and the reward function, evaluating the state information of the unmanned aerial vehicle; adjusting the flight path of the unmanned aerial vehicle according to the evaluation result.

7. The method according to any one of claims 1 to 6, characterized in that, The method further comprises: classifying the information of the target test scene to obtain a test scene type; adding semantic labels and risk levels to the information of the target test scene corresponding to the test scene type; determining the similarity between the information of the target test scene and the information of the test scene in the test scene library according to the semantic labels and the risk levels; in response to the similarity being less than or equal to a similarity threshold, saving the information of the target test scene to the test scene library.

8. An information generation device of a test scenario of a vehicle, characterized by, comprises: an acquisition unit configured to acquire initial image information of a road where a vehicle is located by using a plurality of unmanned aerial vehicles, wherein the initial image information contains road scene data of the road where the vehicle is located; a processing unit configured to perform frame-level synchronization processing on the initial image information to obtain target image information, wherein the target image information is used to represent the time stamp and position information corresponding to each frame in the initial image information; an identification unit configured to analyze the target image information by using a target detection model to obtain a moving track of the vehicle, wherein the target detection model is used to locate and classify the vehicle in the target image information; a generation unit configured to generate information of a target test scene where the vehicle is located based on the moving track and the position information.

9. A processor, comprising: The processor is configured to run a program, wherein the program performs the method of any one of claims 1 to 7 when running.

10. A vehicle characterized by comprising: The vehicle is configured to perform the method of any one of claims 1 to 7.