A traffic event video slice setting method, device, equipment and product
By constructing a 3D traffic scene and setting the initial motion state, and using simulation to detect traffic events, the problem of low efficiency in setting traffic event video slices in existing technologies is solved, and efficient determination of traffic event video slices is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 苏州万集车联网技术有限公司
- Filing Date
- 2024-12-23
- Publication Date
- 2026-06-23
AI Technical Summary
In existing technologies, manually setting initial traffic scene images to obtain traffic event video slices is a cumbersome, inefficient, and time-consuming process.
By acquiring images of the target traffic scene, a 3D scene is constructed and the initial motion states of the foreground and virtual materials are set. The simulation is then used to detect traffic events and determine video slices of the traffic events.
It improves the efficiency of setting up traffic incident video slices, enabling efficient and convenient determination of initial traffic scene images, and accurate determination of traffic incident video slices based on simulation operation.
Smart Images

Figure CN122265900A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent transportation technology, and in particular to a method, apparatus, device, and product for setting up video slices of traffic incidents. Background Technology
[0002] With the rapid development of intelligent transportation technology, constructing simulation scenarios to simulate real-world traffic conditions in a virtual environment and detecting traffic events based on these simulation scenarios to efficiently detect traffic conditions and potential problems is one of the main development directions of current intelligent transportation technology.
[0003] Current technical solutions require technicians to manually set virtual objects and their corresponding virtual motion trajectories in a simulation map to obtain an initial traffic scene image, which is then used for simulation. During the simulation, corresponding traffic event video slices are determined based on detected traffic events, and a traffic event perception model can be trained using these video slices. However, this process of manually setting the initial traffic scene image to obtain traffic event video slices is cumbersome, inefficient, and time-consuming.
[0004] Therefore, how to improve the efficiency of setting up video slices for traffic incidents is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of this application is to provide a method, apparatus, training method, terminal device, computer-readable storage medium, and computer program product for setting up traffic incident video slices, aiming to improve the efficiency of setting up traffic incident video slices.
[0006] Firstly, this application provides a method for setting up video slices of traffic incidents. The method includes:
[0007] A target traffic scene image is acquired, and a three-dimensional scene is constructed in the target traffic scene image to obtain a three-dimensional traffic scene image;
[0008] In the three-dimensional traffic scene image, foreground material and virtual material are set, and initial motion states corresponding to the foreground material and the virtual material are set respectively to obtain an initial traffic scene image; the foreground material is material extracted from the original image;
[0009] The simulation is performed based on the initial traffic scene image, and traffic event video slices are determined based on the traffic events detected during the simulation.
[0010] In one specific embodiment, the step of performing a simulation based on the initial traffic scene image and determining traffic event video slices based on traffic events detected during the simulation includes:
[0011] The simulation operation is performed based on the initial traffic scene image;
[0012] If a traffic event is detected during the simulation, the target video frame corresponding to the traffic event is determined.
[0013] Traffic event video slices are determined based on the target video frame and adjacent video frames adjacent to the target video frame.
[0014] In one specific embodiment, the method further includes:
[0015] The conflict point, conflict target, and conflict trajectory corresponding to the traffic incident are marked in the video slice of the traffic incident.
[0016] In one specific embodiment, the method further includes:
[0017] The foreground material is scaled based on its initial size and target placement position.
[0018] In one specific embodiment, the method further includes:
[0019] The foreground material and / or the virtual material are subjected to texture processing, color processing, and lighting processing.
[0020] In one specific embodiment, the method further includes:
[0021] If no traffic event is detected during the simulation, the initial motion state is adjusted, and the simulation is returned to the steps of performing the simulation based on the initial traffic scene image and determining the traffic event video slice based on the traffic event detected during the simulation.
[0022] In one specific embodiment, the method further includes:
[0023] Obtain the 3D point cloud data corresponding to the foreground material, and determine the 3D model material corresponding to the foreground material based on the 3D point cloud data;
[0024] The process involves setting foreground and virtual materials in the 3D traffic scene image, and setting initial motion states corresponding to the foreground and virtual materials respectively, to obtain an initial traffic scene image; the foreground material is material extracted from the original image, including:
[0025] In the three-dimensional traffic scene image, three-dimensional model materials and virtual materials are set, and the initial motion states corresponding to the three-dimensional model materials and the virtual materials are set respectively to obtain the initial traffic scene image.
[0026] Secondly, this application also provides a method for training a traffic incident perception model, the method comprising:
[0027] Traffic incident video slices are obtained according to the above method for setting up traffic incident video slices, and a sample set of traffic incident video slices is determined.
[0028] The perception model is trained based on the traffic incident video slice sample set to obtain the traffic incident perception model.
[0029] Thirdly, this application also provides a device for setting up video slices of traffic incidents. The device includes:
[0030] The acquisition module is used to acquire a target traffic scene image and construct a three-dimensional scene in the target traffic scene image to obtain a three-dimensional traffic scene image.
[0031] The setting module is used to set foreground material and virtual material in the three-dimensional traffic scene image, and set the initial motion state corresponding to the foreground material and the virtual material respectively to obtain an initial traffic scene image; the foreground material is material extracted from the original image;
[0032] The simulation module is used to perform a simulation based on the initial traffic scene image and to determine traffic event video slices based on traffic events detected during the simulation.
[0033] Fourthly, this application also provides a terminal device. The terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described above.
[0034] Fifthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method described above.
[0035] Sixthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described above.
[0036] This application provides a method for setting up traffic event video slices. After acquiring a target traffic scene image and constructing a three-dimensional scene in the target traffic scene image, foreground material and virtual material are set in the target traffic scene image according to the three-dimensional scene, and initial motion states corresponding to the foreground material and the virtual material are set respectively to obtain an initial traffic scene image. The foreground material is material extracted from the original image. That is to say, the virtual objects in the initial traffic scene in this method can be extracted from the original image. Therefore, this method can efficiently and conveniently determine the initial traffic scene image, perform simulation operation based on the initial traffic scene image, and determine traffic event video slices based on traffic events detected during the simulation operation, which can improve the efficiency of setting up traffic event video slices.
[0037] It is understood that the traffic incident video slice setting device, traffic incident perception model training method, terminal device, computer-readable storage medium and computer program product provided in the embodiments of this application have the same beneficial effects as the traffic incident video slice setting method described above, and will not be repeated here. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0039] Figure 1 A flowchart illustrating a method for setting up video slices of a traffic incident, as provided in an embodiment of this application;
[0040] Figure 2 A schematic diagram illustrating how to extract foreground material from an original image, as provided in an embodiment of this application;
[0041] Figure 3 A schematic diagram of an initial traffic scene image provided for an embodiment of this application;
[0042] Figure 4 A flowchart illustrating a training method for a traffic incident perception model provided in an embodiment of this application;
[0043] Figure 5 The diagram shown is a schematic representation of a device for setting up video slices of a traffic incident according to an embodiment of this application.
[0044] Figure 6 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0045] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.
[0046] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0047] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0048] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0049] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0050] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. "A plurality" means "two or more."
[0051] The present application provides a method for setting up video slices of traffic incidents, which can be executed by the processor of a terminal device when running a corresponding computer program.
[0052] Figure 1 The flowchart illustrates a method for setting up traffic event video slices according to an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The method provided in this embodiment includes the following steps:
[0053] S100: Acquire the target traffic scene image and construct a 3D scene in the target traffic scene image to obtain a 3D traffic scene image.
[0054] Specifically, scene image frames identified from traffic scene monitoring videos can be used as target traffic scene images. Target traffic scene images include geographic information, road layout, and original materials, which are virtual objects such as vehicles and pedestrians included in the target traffic scene images.
[0055] Specifically, based on the geographic information and road layout in the target traffic scene image, a 3D model is constructed on the target traffic scene image to obtain a 3D traffic scene image. In other words, constructing a 3D scene from a target traffic scene image means transforming a 2D target traffic scene image into a 3D target traffic scene image, resulting in a 3D traffic scene image. For example, a certain location in the 3D traffic scene image represents its location in the corresponding virtual space.
[0056] More specifically, a 3D scene is constructed on top of the target traffic scene image, and the background of the 3D scene is set to transparent to obtain a 3D traffic scene image after the target traffic scene image and the 3D scene are fused.
[0057] S200: Set foreground material and virtual material in the 3D traffic scene image, and set the initial motion state corresponding to the foreground material and virtual material respectively to obtain the initial traffic scene image; the foreground material is the material extracted from the original image.
[0058] In this context, foreground elements refer to key elements in the actual traffic scene; virtual elements refer to key elements in the traffic scene manually set by technicians in the simulation map. Both foreground and virtual elements include vehicles and pedestrians.
[0059] In this embodiment, the foreground material is obtained by cutting out images from the original images. In practical applications, after cutting out multiple foreground materials from each original image, the foreground materials can be classified and stored according to their type or size to obtain a foreground material library.
[0060] In this embodiment, the virtual objects in the 3D traffic scene image include foreground material, virtual material, and preset original material in the target traffic scene image.
[0061] The initial motion state includes a first initial motion state corresponding to the foreground material, a second initial motion state corresponding to the virtual material, and a third initial motion state corresponding to the preset original material in the target traffic scene image.
[0062] Specifically, the initial motion state includes the initial position, preset operating parameters, and preset motion trajectory. For foreground and virtual materials, the initial position is the location where the foreground and virtual materials are set in the corresponding position in the 3D traffic scene image, i.e., the target placement position. Preset operating parameters include initial velocity, acceleration, and direction of travel.
[0063] Figure 2 A schematic diagram illustrating how to extract foreground material from an original image, as provided in an embodiment of this application; Figure 3 This is a schematic diagram of an initial traffic scene image provided in an embodiment of this application. (In conjunction with...) Figure 2 and Figure 3 As shown, the target placement positions (initial positions) of foreground and virtual materials can be set according to the traffic event to be simulated. When setting the target placement positions, the positions and occlusion relationships of the foreground and virtual materials relative to the original materials should be reasonably arranged to avoid occlusions in the target traffic scene image, such as trees, bushes and trash cans, so as to make the initial traffic scene image more natural, improve the detection algorithm's ability to handle occlusion, and improve the accuracy of determining traffic event video slices.
[0064] In a specific example, a target range is pre-defined within the target traffic scene image. This means that the target placement can only be set within the target range; that is, foreground and virtual materials can only be placed within the target range. The target range can be the boundary of a rectangle, such as setting vertical and horizontal boundaries in the target traffic scene image. The rectangle formed by the vertical and horizontal boundaries constitutes the target range of the target traffic scene image. Alternatively, the target range can also be circular, elliptical, or other irregular shapes; this embodiment does not limit this.
[0065] In practical applications, real vehicle trajectory data corresponding to the target traffic scene image can be obtained, and the initial motion state of the corresponding virtual object can be set according to the real vehicle trajectory data.
[0066] S300: Performs simulation based on initial traffic scene images and determines traffic event video slices based on traffic events detected during the simulation.
[0067] Among them, simulation operation refers to the process of simulating the operation of each virtual object based on the initial motion state by using the initial traffic scene image as the initial simulation state through simulation software. Specifically, it simulates the operation of the foreground material, virtual material and original material according to their respective initial motion states.
[0068] Specifically, based on the initial motion state, the trajectory of the virtual object is set by combining the kinematic model and traffic rules, and the virtual object is controlled to move according to the corresponding trajectory.
[0069] Traffic incidents refer to specific situations that occur in traffic flow, such as traffic accidents, traffic congestion, and violations.
[0070] Among them, the traffic incident video slice refers to a video segment extracted from the simulation process. The extracted video includes the traffic incident and the corresponding traffic scene for a period of time before and after the traffic incident.
[0071] Specifically, traffic events may occur during the process of each virtual object running along a preset trajectory; therefore, an event detection algorithm is used to detect events in the traffic flow corresponding to the simulation operation. When a traffic event is detected during the simulation operation, a traffic event video slice is determined based on the traffic event.
[0072] By generating various traffic events and identifying the corresponding traffic event video slices, the diversity of traffic event video slices can be significantly increased, improving the generalization ability of the traffic event perception model when training the model based on traffic event video slices. It can also simulate and generate extreme or rare traffic events, helping the algorithm learn and adapt to various complex scenarios. It can also be used to verify the performance of the traffic event perception model in these complex scenarios, ensuring its robustness.
[0073] This application provides a method for setting traffic event video slices. After acquiring a target traffic scene image and constructing a three-dimensional scene in the target traffic scene image, foreground material and virtual material are set in the target traffic scene image according to the three-dimensional scene, and initial motion states corresponding to the foreground material and virtual material are set respectively to obtain an initial traffic scene image. The foreground material is material extracted from the original image. That is to say, the virtual objects in the initial traffic scene in this method can be extracted from the original image. Therefore, this method can efficiently and conveniently determine the initial traffic scene image, perform simulation operation based on the initial traffic scene image, and determine traffic event video slices based on traffic events detected during the simulation operation, which can improve the efficiency of setting traffic event video slices.
[0074] Based on the above embodiments, this embodiment further explains and optimizes the technical solution. Specifically, in this embodiment, a simulation is performed based on an initial traffic scene image, and traffic event video slices are determined according to traffic events detected during the simulation, including:
[0075] Simulation operation is performed based on initial traffic scene images;
[0076] If a traffic event is detected during the simulation, the target video frame corresponding to the traffic event is determined.
[0077] Traffic event video slices are determined based on the target video frame and adjacent video frames adjacent to the target video frame.
[0078] In this context, the target video frame refers to a specific video frame corresponding to a detected traffic event during the simulation process; adjacent video frames refer to the video frames immediately before and after the target video frame, used to provide contextual information about the traffic event; this embodiment does not limit the number of adjacent video frames. For example, five video frames immediately before and after the target video frame can be defined as adjacent video frames.
[0079] In this embodiment, during the simulation operation based on the initial traffic scene image, the system continuously detects whether there are traffic events during the simulation operation. If a traffic event is detected, the system determines the target video frame based on the video frame corresponding to the traffic time. Then, based on the timing or frame identifier of the target video frame, a certain number of video frames before and after the target video frame are determined as adjacent video frames. The target video frame and the adjacent video frames are combined to form a complete traffic event video slice.
[0080] The method described in this embodiment can ensure that traffic incident video slices include complete traffic incidents while minimizing redundant information, thereby improving the quality of traffic incident video slices.
[0081] Based on the above embodiments, this embodiment further explains and optimizes the technical solution. Specifically, in this embodiment, the method further includes:
[0082] Mark the conflict points, conflict targets, and conflict trajectories corresponding to the traffic incidents in the video slices.
[0083] Among them, the conflict point refers to the specific location where a conflict occurs between traffic participants in a traffic incident; the conflict target refers to the traffic participants involved in the conflict in a traffic incident, such as vehicles involved in a collision or pedestrians who violate traffic rules; and the conflict trajectory refers to the movement path of the conflict target before and after the traffic incident.
[0084] Specifically, in the video frames corresponding to the traffic incident video slices, marking tools are used to add visual markings to the conflict points, conflict targets, and conflict trajectories, such as drawing a circle or square at the conflict point, marking the conflict target with a label box, and drawing the conflict trajectory.
[0085] In addition, the type of traffic incident can be marked in the text box, such as collision, illegal parking, motor vehicle occupying non-motor vehicle lane, etc. This embodiment does not limit this.
[0086] According to the method of this embodiment, users can more intuitively and accurately view traffic events during the simulation process. By marking the traffic event video slices with the corresponding conflict points, conflict targets and conflict trajectories, traffic event video slice samples are obtained, which can be used for training the traffic event perception model.
[0087] Based on the above embodiments, this embodiment further explains and optimizes the technical solution. Specifically, in this embodiment, the method further includes:
[0088] The foreground material is scaled based on its initial size and the target placement position.
[0089] The initial size refers to the size of the foreground material when cutting out the image from the original image.
[0090] The target placement location refers to the placement location of the foreground material in the 3D scene.
[0091] When setting foreground elements in a target traffic scene image, the foreground elements must conform to the rule of perspective (objects appear larger when closer and smaller when farther away). That is, the relative distance between the foreground elements and the virtual camera (virtual observer's viewpoint) must be considered. Figure 2 As shown, the closer the target is to the virtual camera, the smaller the relative distance, and the larger the size of the foreground material; the farther the target is from the virtual camera, the greater the relative distance, and the smaller the size of the foreground material.
[0092] In one specific embodiment, the relative distance between the target placement location and the virtual camera is determined; the virtual dimensions of other objects with the same relative distance in the target traffic scene are obtained; the scaling ratio of the foreground material is determined based on the ratio of the actual dimensions of other objects in the real environment to the virtual dimensions of other objects in the target traffic scene image; and the foreground material is scaled based on its initial size and scaling ratio.
[0093] After scaling the foreground material, the scaled foreground material is placed at the target location in the target traffic scene image.
[0094] According to the method of this embodiment, by scaling the foreground material, the accuracy of the initial traffic scene image can be improved, and the foreground material in the target traffic scene image can be made more visually harmonious and unified with other virtual objects.
[0095] Based on the above embodiments, this embodiment further explains and optimizes the technical solution. Specifically, in this embodiment, the method further includes:
[0096] Perform texture processing, color processing, and lighting and shadow processing on foreground materials and / or virtual materials.
[0097] In this embodiment, after setting the foreground material and / or virtual material in the target traffic scene image, the foreground material and / or virtual material are subjected to texture processing, color processing and lighting processing.
[0098] Texture processing refers to the operation of applying texture mapping to foreground materials and / or virtual materials to simulate the material and details of the surface of real objects.
[0099] Color processing refers to adjusting the color attributes of foreground and / or virtual materials. Color processing includes adjusting base color, saturation, and brightness. More specifically, color processing can be achieved by modifying the material properties of the foreground and / or virtual materials or by using color correction tools.
[0100] Lighting and shadow processing refers to the manipulation of simulating the effect of light sources on foreground and / or virtual materials. More specifically, it can be achieved by adjusting information such as light source type, light source intensity, light source color, and light source position; light source types include point lights and parallel lights, etc.
[0101] In practical applications, any one or more of the following operations—texture processing, color processing, and lighting processing—can be selected to process the foreground material and / or virtual material. This embodiment does not limit this.
[0102] The method described in this embodiment can further ensure that the foreground material and / or virtual material are as close as possible to the real world during the simulation process, thereby improving the realism and visual effect of the initial traffic scene image and traffic event video slices.
[0103] Based on the above embodiments, this embodiment further explains and optimizes the technical solution. Specifically, in this embodiment, the method further includes:
[0104] If no traffic event is detected during the simulation, the initial motion state is adjusted, and the simulation is returned to the initial traffic scene image. The steps for determining the traffic event video slice are then performed based on the traffic events detected during the simulation.
[0105] In this embodiment, during the process of using the event detection algorithm to detect events in the traffic flow corresponding to the simulation operation, if no traffic event is detected within the preset simulation time, it means that the operation will continue according to the current initial motion state and no traffic event may occur.
[0106] In this embodiment, in order to identify traffic event video slices that include traffic events, if no traffic event is detected during the simulation, the initial motion state of the foreground or virtual material needs to be adjusted. Specifically, the initial speed, acceleration, direction of travel, or preset trajectory of the foreground or virtual material needs to be adjusted.
[0107] It should be noted that after adjusting the initial motion state of the foreground or virtual material, the motion trajectories of other traffic participants related to the foreground or virtual material will be adjusted accordingly during the simulation model's operation; the event detection algorithm will continue to be used to detect events in the traffic flow corresponding to the adjusted simulation operation.
[0108] The method described in this embodiment can efficiently and accurately identify traffic event video slices, including traffic events.
[0109] Based on the above embodiments, this embodiment further explains and optimizes the technical solution. Specifically, in this embodiment, the method further includes:
[0110] Acquire the 3D point cloud data corresponding to the foreground material, and determine the 3D model material corresponding to the foreground material based on the 3D point cloud data;
[0111] In a 3D traffic scene image, foreground and virtual materials are set, and initial motion states corresponding to the foreground and virtual materials are set respectively to obtain an initial traffic scene image; the foreground material is material extracted from the original image, including:
[0112] In a 3D traffic scene image, 3D model materials and virtual materials are set, and the initial motion states corresponding to the 3D model materials and virtual materials are set respectively to obtain the initial traffic scene image.
[0113] In this embodiment, when acquiring foreground material, the corresponding 3D point cloud data is further acquired. Specifically, the 3D point cloud data corresponding to the foreground material can be acquired through laser scanning, photogrammetry, or other 3D scanning techniques.
[0114] Then, the 2D foreground material obtained from the image cutout is combined with the 3D point cloud data to obtain a 3D model material including both the foreground material and the 3D point cloud data. The 3D model material is then placed at the target location in the target traffic scene image, with preset running parameters and preset motion trajectories set. In other words, the 3D model material replaces the original 2D foreground material. The 3D model material can display more realistic dynamic effects in simulation. For example, when a 2D "paper doll" turns, it is no longer a planar rotation, but can display a 3D stereoscopic effect of a human turning.
[0115] The method described in this embodiment can make the foreground materials in the simulation scene more realistic, thereby improving the realism of the simulation and the accuracy of the analysis.
[0116] Figure 4 A flowchart illustrating a training method for a traffic incident perception model provided in this application embodiment. In this application embodiment, a training method for a traffic incident perception model includes:
[0117] S410: Obtain traffic event video slices according to the traffic event video slice setting method provided in any of the above embodiments, and determine the traffic event video slice sample set;
[0118] S420: Train the perception model based on the traffic incident video slice sample set to obtain the traffic incident perception model.
[0119] The traffic incident video slice sample set includes multiple traffic incident video slice samples. The traffic incident video slice sample refers to a traffic incident video slice with tagged information. That is, the corresponding traffic incident video slice sample is obtained by setting tag information in the traffic incident video slice. The tag information includes the conflict point, conflict target and conflict trajectory corresponding to the traffic incident.
[0120] After determining the sample set of traffic incident video slices, an initial neural network is used to train a perception model based on the sample set, thus obtaining a traffic incident perception model. The initial neural network can be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory network (LSTM), or a feed-forward neural network (FNN), etc. This embodiment does not limit the specific type of the initial neural network.
[0121] The traffic incident perception model can be applied to the vehicle end or the system control end; this embodiment does not limit this application.
[0122] In a specific example, assuming it is applied to the vehicle, after the vehicle collects the current traffic scene image corresponding to the current traffic flow, the current traffic scene image is input into the traffic event perception model, which can efficiently and accurately predict traffic events in the traffic flow. Based on the predicted traffic events, assisted driving can be performed, which can improve driving safety.
[0123] This application provides a training method for a traffic event perception model. Since the virtual objects in the initial traffic scene of a traffic event video slice setting method can be obtained by extracting images from the original image, the initial traffic scene image can be determined efficiently and conveniently. Simulation is performed based on the initial traffic scene image, and traffic event video slices are determined based on the traffic events detected during the simulation. Based on the traffic event video slices, a traffic event video slice sample set is determined, which can improve the efficiency of setting traffic event video slices. The perception model is trained based on the traffic event video slice sample set to obtain the traffic event perception model, which can efficiently train the traffic event perception model.
[0124] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0125] It should be noted that the information collection process (such as the facial image collection process, fingerprint information collection process, etc.) / feature extraction process involved in this application is carried out with the user's knowledge and permission. That is, the information collection process / feature extraction process complies with the requirements of laws and regulations and does not constitute an act that harms the public interest.
[0126] Figure 5 The diagram shown is a structural schematic of a device for setting up traffic incident video slices according to an embodiment of this application. Figure 5 As shown, the traffic incident video slice setting device in this embodiment includes an acquisition module 510, a setting module 520, and a simulation module 530; wherein,
[0127] The acquisition module 510 is used to acquire a target traffic scene image and construct a three-dimensional scene in the target traffic scene image to obtain a three-dimensional traffic scene image;
[0128] The setting module 520 is used to set foreground material and virtual material in a 3D traffic scene image, and set the initial motion state corresponding to the foreground material and virtual material respectively to obtain an initial traffic scene image; the foreground material is material extracted from the original image;
[0129] The simulation module 530 is used to perform simulation based on the initial traffic scene image and to determine traffic event video slices based on the traffic events detected during the simulation.
[0130] The device for setting up traffic event video slices provided in this application embodiment has the same beneficial effects as the traffic event video slice setting method described above.
[0131] In one embodiment, the simulation module 530 includes:
[0132] The simulation operation submodule is used to perform simulation operations based on the initial traffic scene images;
[0133] The target video frame determination submodule is used to determine the target video frame corresponding to the traffic event if a traffic event is detected during the simulation.
[0134] The traffic incident video slice determination submodule is used to determine traffic incident video slices based on the target video frame and adjacent video frames adjacent to the target video frame.
[0135] In one embodiment, a device for setting up traffic incident video slices further includes:
[0136] The tagging module is used to tag the conflict points, conflict targets, and conflict trajectories corresponding to traffic events in traffic event video slices.
[0137] In one embodiment, a device for setting up traffic incident video slices further includes:
[0138] The scaling module is used to scale the foreground material according to its initial size and the target placement position.
[0139] In one embodiment, a device for setting up traffic incident video slices further includes:
[0140] The image processing module is used to perform texture processing, color processing, and lighting and shadow processing on foreground materials and / or virtual materials.
[0141] In one embodiment, a device for setting up traffic incident video slices further includes:
[0142] If no traffic event is detected during the simulation, the initial motion state is adjusted and simulation module 530 is invoked.
[0143] In one embodiment, a device for setting up traffic incident video slices further includes:
[0144] The point cloud data acquisition module is used to acquire the 3D point cloud data corresponding to the foreground material, and determine the 3D model material corresponding to the foreground material based on the 3D point cloud data.
[0145] The settings module includes:
[0146] The settings submodule is used to set up 3D model materials and virtual materials in a 3D traffic scene image, and to set the initial motion state corresponding to the 3D model materials and virtual materials respectively, so as to obtain the initial traffic scene image.
[0147] This application embodiment also provides a training device for a traffic event perception model, which includes:
[0148] The sample set determination module is used to obtain traffic event video slices according to the traffic event video slice setting device in the above embodiment, and determine the traffic event video slice sample set.
[0149] The model training module is used to train the perception model based on a sample set of traffic event video slices to obtain a traffic event perception model.
[0150] The training device for a traffic event perception model provided in this application embodiment has the same beneficial effects as the training method for a traffic event perception model described above.
[0151] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0152] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0153] Figure 6 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 6 As shown, the terminal device 600 of this embodiment includes a memory 601, a processor 602, and a computer program 603 stored in the memory 601 and executable on the processor 602; when the processor 602 executes the computer program 603, it implements the steps in the above-described embodiments of the methods for setting up video slices of traffic events; or when the processor 602 executes the computer program 603, it implements the functions of each module / unit in the above-described embodiments of the devices.
[0154] For example, computer program 603 can be divided into one or more modules / units, one or more of which are stored in memory 601 and executed by processor 602 to implement the method of the embodiments of this application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of computer program 603 in terminal device 600. For example, computer program 603 can be divided into an acquisition module, a setting module, and a simulation module, with the specific functions of each module as follows:
[0155] The acquisition module is used to acquire the target traffic scene image and construct a 3D scene in the target traffic scene image to obtain a 3D traffic scene image;
[0156] The setup module is used to set foreground and virtual materials in a 3D traffic scene image, and to set the initial motion state corresponding to the foreground and virtual materials respectively, so as to obtain an initial traffic scene image; the foreground material is the material extracted from the original image;
[0157] The simulation module is used to perform simulation based on the initial traffic scene images and to determine traffic event video slices based on the traffic events detected during the simulation.
[0158] In applications, terminal device 600 can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. Terminal device 600 may include, but is not limited to, memory 601 and processor 602. Those skilled in the art will understand that... Figure 6 This is merely an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown, or combine certain components, or different components. For example, a terminal device may also include input / output devices, network access devices, buses, etc.; among which, input / output devices may include cameras, audio acquisition / playback devices, displays, etc.; network access devices may include communication modules for wireless communication with external devices.
[0159] In applications, the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0160] In applications, memory can be an internal storage unit of a terminal device, such as its hard drive or RAM; it can also be an external storage device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card; or it can include both internal and external storage units. Memory is used to store operating systems, applications, boot loaders, data, and other programs, such as computer program code. Memory can also be used to temporarily store data that has been output or will be output.
[0161] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps in the above-described method embodiments.
[0162] This application implements all or part of the processes in the methods of the above embodiments, which can be accomplished by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, such as a USB flash drive, a portable hard drive, a magnetic disk, or an optical disk.
[0163] The computer-readable storage medium provided in this application embodiment has the same beneficial effects as the above-described method for setting up traffic event video slices.
[0164] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps in the various method embodiments described above.
[0165] The computer program product provided in this application embodiment has the same beneficial effects as the above-described method for setting up traffic event video slices.
[0166] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0167] Those skilled in the art will recognize that the device and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0168] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interface, or the device may be indirectly coupled or communicated, and may be electrical, mechanical, or other forms.
[0169] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for setting up video slices of a traffic incident, characterized in that, The method includes: A target traffic scene image is acquired, and a three-dimensional scene is constructed in the target traffic scene image to obtain a three-dimensional traffic scene image; In the three-dimensional traffic scene image, foreground material and virtual material are set, and initial motion states corresponding to the foreground material and the virtual material are set respectively to obtain an initial traffic scene image; the foreground material is material extracted from the original image; The simulation is performed based on the initial traffic scene image, and traffic event video slices are determined based on the traffic events detected during the simulation.
2. The method according to claim 1, characterized in that, The simulation operation based on the initial traffic scene image, and the determination of traffic event video slices based on traffic events detected during the simulation operation, include: The simulation operation is performed based on the initial traffic scene image; If a traffic event is detected during the simulation, the target video frame corresponding to the traffic event is determined. Traffic event video slices are determined based on the target video frame and adjacent video frames adjacent to the target video frame.
3. The method according to claim 1, characterized in that, The method further includes: The conflict point, conflict target, and conflict trajectory corresponding to the traffic incident are marked in the video slice of the traffic incident.
4. The method according to claim 1, characterized in that, The method further includes: The foreground material is scaled based on its initial size and target placement position.
5. The method according to claim 1, characterized in that, The method further includes: The foreground material and / or the virtual material are subjected to texture processing, color processing, and lighting processing.
6. The method according to claim 3, characterized in that, The method further includes: If no traffic event is detected during the simulation, the initial motion state is adjusted, and the simulation is returned to the steps of performing the simulation based on the initial traffic scene image and determining the traffic event video slice based on the traffic event detected during the simulation.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain the 3D point cloud data corresponding to the foreground material, and determine the 3D model material corresponding to the foreground material based on the 3D point cloud data; The process involves setting foreground and virtual materials in the 3D traffic scene image, and setting initial motion states corresponding to the foreground and virtual materials respectively, to obtain an initial traffic scene image; the foreground material is material extracted from the original image, including: In the three-dimensional traffic scene image, three-dimensional model materials and virtual materials are set, and the initial motion states corresponding to the three-dimensional model materials and the virtual materials are set respectively to obtain the initial traffic scene image.
8. A training method for a traffic incident perception model, characterized in that, The method includes: Traffic incident video slices are obtained according to the method for setting traffic incident video slices according to any one of claims 1 to 7, and a traffic incident video slice sample set is determined. The perception model is trained based on the traffic incident video slice sample set to obtain the traffic incident perception model.
9. A device for setting up video slices of traffic incidents, characterized in that, The device includes: The acquisition module is used to acquire a target traffic scene image and construct a three-dimensional scene in the target traffic scene image to obtain a three-dimensional traffic scene image. The setting module is used to set foreground material and virtual material in the three-dimensional traffic scene image, and set the initial motion state corresponding to the foreground material and the virtual material respectively to obtain an initial traffic scene image; the foreground material is material extracted from the original image; The simulation module is used to perform a simulation based on the initial traffic scene image and to determine traffic event video slices based on traffic events detected during the simulation.
10. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7 or 8.
11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7 or 8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7 or claim 8.