Traffic accident detection method and apparatus, electronic device, and medium
Patent Information
- Application Number
- CN202310547670.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-16
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2043-05-16
Smart Images

Figure CN116563801B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, specifically to technologies such as deep learning, computer vision, and intelligent transportation, and in particular to methods, devices, electronic equipment, and media for traffic accident detection. Background Technology
[0002] With the rapid expansion of cities and the continuous increase in vehicle ownership, road traffic pressure has surged, frequently leading to urban road congestion and traffic accidents. Automated detection of traffic accidents, such as identifying the type of accident and the responsible party, can provide accident participants and traffic authorities with information for accident assessment, improving response and processing speeds and alleviating traffic congestion caused by accidents. Summary of the Invention
[0003] This disclosure provides a method, apparatus, electronic device, and medium for detecting traffic accidents.
[0004] According to one aspect of this disclosure, a traffic accident detection method is provided, applied on a server side, including:
[0005] Receive the video to be detected sent by the vehicle's in-vehicle terminal and the multimodal information collected from the cabin of the target vehicle;
[0006] The target video segment of the traffic accident is extracted from the video to be detected, and the collision type of the target video segment is classified to obtain the target collision type corresponding to the traffic accident.
[0007] Based on the multimodal information, semantic classification and accident liability detection are performed on the target video segments to obtain the target accident type and target liability party corresponding to the traffic accident;
[0008] Based on the target accident type, the target collision type, and the target responsible party, the detection result of the traffic accident is generated and sent to the vehicle-mounted terminal.
[0009] According to another aspect of this disclosure, an alternative method for detecting traffic accidents is provided, applied to the vehicle's infotainment system, including:
[0010] Send the video to be detected and the multimodal information collected from the cabin of the target vehicle to the server;
[0011] Receive the detection results sent by the server;
[0012] The detection result involves extracting a target video segment of a traffic accident from the video to be detected, classifying the target video segment by collision type to obtain the target collision type corresponding to the traffic accident, and performing semantic classification and accident liability detection on the target video segment based on the multimodal information to obtain the target accident type and target liability party corresponding to the traffic accident, and generating a result based on the target accident type, the target collision type, and the target liability party.
[0013] According to another aspect of this disclosure, a traffic accident detection device is provided for use on a server side, comprising:
[0014] The receiving module is used to receive the video to be detected sent by the vehicle's in-vehicle terminal and the multimodal information collected from the cabin of the target vehicle.
[0015] The interception module is used to intercept target video segments of traffic accidents from the video to be detected;
[0016] A classification module is used to classify the target video clips by collision type to obtain the target collision type corresponding to the traffic accident;
[0017] The processing module is used to perform semantic classification and accident liability detection on the target video segment based on the multimodal information, so as to obtain the target accident type and target liability party corresponding to the traffic accident;
[0018] The first generation module is used to generate the detection results of the traffic accident based on the target accident type, the target collision type, and the target responsible party;
[0019] The sending module is used to send the detection results to the vehicle-mounted terminal.
[0020] According to another aspect of this disclosure, another traffic accident detection device is provided, applied to the vehicle-mounted terminal of a target vehicle, comprising:
[0021] The sending module is used to send the video to be detected and the multimodal information collected from the cabin of the target vehicle to the server.
[0022] A receiving module is used to receive the detection results sent by the server;
[0023] The detection result involves extracting a target video segment of a traffic accident from the video to be detected, classifying the target video segment by collision type to obtain the target collision type corresponding to the traffic accident, and performing semantic classification and accident liability detection on the target video segment based on the multimodal information to obtain the target accident type and target liability party corresponding to the traffic accident, and generating a result based on the target accident type, the target collision type, and the target liability party.
[0024] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0025] At least one processor; and
[0026] A memory communicatively connected to the at least one processor; wherein,
[0027] The memory stores instructions that can be executed by the at least one processor, which, when executed, enable the at least one processor to perform either the traffic accident detection method proposed in one aspect of this disclosure or the traffic accident detection method proposed in another aspect of this disclosure.
[0028] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided, the computer instructions being used to cause the computer to perform the traffic accident detection method proposed in one aspect of this disclosure, or to perform the traffic accident detection method proposed in another aspect of this disclosure.
[0029] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the traffic accident detection method proposed in one aspect of this disclosure, or, when executed, implements the traffic accident detection method proposed in another aspect of this disclosure.
[0030] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0031] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0032] Figure 1 This is a flowchart illustrating the traffic accident detection method provided in Embodiment 1 of this disclosure;
[0033] Figure 2 This is a schematic flowchart of the traffic accident detection method provided in Embodiment 2 of this disclosure;
[0034] Figure 3 This is a flowchart illustrating the traffic accident detection method provided in Embodiment 3 of this disclosure;
[0035] Figure 4 This is a flowchart illustrating the traffic accident detection method provided in Embodiment 4 of this disclosure;
[0036] Figure 5 This is a schematic flowchart of the traffic accident detection method provided in Embodiment 5 of this disclosure;
[0037] Figure 6 This is a schematic flowchart of the traffic accident detection method provided in Embodiment Six of this disclosure;
[0038] Figure 7 This is a schematic flowchart of the traffic accident detection method provided in Embodiment 7 of this disclosure;
[0039] Figure 8 This is a schematic diagram illustrating the implementation process of the traffic accident detection method provided in this embodiment of the disclosure;
[0040] Figure 9 This is a schematic diagram of the display interface of the vehicle-mounted system provided in an embodiment of this disclosure. Figure 1 ;
[0041] Figure 10 This is a schematic diagram of the display interface of the vehicle-mounted system provided in an embodiment of this disclosure. Figure 2 ;
[0042] Figure 11 This is a schematic diagram of the display interface of the vehicle-mounted system provided in an embodiment of this disclosure. Figure 3 ;
[0043] Figure 12 This is a schematic diagram of the interaction process between the vehicle-mounted terminal and the server provided in an embodiment of this disclosure;
[0044] Figure 13 This is a schematic diagram of the structure of the traffic accident detection device provided in Embodiment 8 of this disclosure;
[0045] Figure 14 This is a schematic diagram of the structure of the traffic accident detection device provided in Embodiment 9 of this disclosure;
[0046] Figure 15 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0047] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0048] Currently, the usual way to handle traffic accidents is for the parties involved to report the accident to the traffic department, which then dispatches traffic police to reconstruct the accident scene or to determine the type of accident and handle it through vehicle dashcams and surveillance videos.
[0049] This approach has at least the following two drawbacks:
[0050] The first problem is that due to insufficient traffic police manpower, there are often multiple traffic accident scenes where no one responds in a timely manner. Moreover, while the traffic police are judging and handling the accident, the accident scene will still cause traffic congestion, making it impossible to guide traffic on surrounding roads in a timely manner, thus creating a vicious cycle.
[0051] The second type is that traffic accidents are usually caused by a variety of factors, and traffic police have difficulty judging the type of accident and the responsible party by visual inspection.
[0052] In addition, for a car dashcam to meet the above-mentioned user needs, the following three steps are required:
[0053] 1. Users can manually export accident-related videos through the vehicle's dashcam;
[0054] 2. Users select video clips before and after the accident;
[0055] 3. Users can determine the type of accident, the severity of the accident, and the responsible party by repeatedly watching the video.
[0056] However, the above three steps have at least the following drawbacks:
[0057] 1. High time cost: Users usually need to judge and search for the video clips where the accident occurred from long videos, and watch the videos repeatedly to make corresponding judgments, which is very time-consuming.
[0058] 2. High difficulty in judgment and low confidence: Car dashcams usually record video from only one angle, while traffic accidents are usually caused by multiple factors. It is difficult for users to judge the type of accident and the responsible party by watching the video alone. At the same time, the judgment is subjective and lacks persuasiveness and confidence.
[0059] In addition, the technical solutions for the detection and analysis of traffic accidents in related technologies mainly include the following two types:
[0060] The first method: based on digital image processing.
[0061] Using optical flow and target tracking models, motion information of each pixel in each video frame is extracted to identify static and / or dynamic targets as target objects. Based on the identified traffic accident scene features, the detection result of the traffic accident is determined.
[0062] The second approach is based on knowledge graphs.
[0063] Structured traffic accident data is extracted, and a knowledge graph is constructed based on the extracted data. This graph links various environmental features that affect traffic accidents, such as weather conditions, visibility, the number of lanes at intersections, and the time of the accident, to uncover the distribution patterns and triggering factors of accidents.
[0064] However, the first method mentioned above relies on a single video frame as the basis for judgment, which is insufficient in terms of accuracy. Furthermore, the model has weak generalization performance and cannot flexibly adapt to various different scenarios, making it difficult to implement in practice.
[0065] The second method mentioned above requires a large amount of data to construct a relatively complete knowledge graph, and the standards for knowledge extraction and representation in this application scenario are difficult to determine, making it difficult to implement.
[0066] In response to at least one of the aforementioned problems, this disclosure proposes a method, apparatus, electronic device, and storage medium for detecting traffic accidents.
[0067] The following description, with reference to the accompanying drawings, outlines a traffic accident detection method, apparatus, electronic device, and storage medium according to embodiments of the present disclosure.
[0068] Figure 1 This is a schematic flowchart of the traffic accident detection method provided in Embodiment 1 of this disclosure.
[0069] The traffic accident detection method provided in this embodiment can be applied to the server side.
[0070] like Figure 1 As shown, the traffic accident detection method may include the following steps:
[0071] Step S101: Receive the video to be detected sent by the vehicle's in-vehicle terminal and the multimodal information collected from the cabin of the target vehicle.
[0072] In this embodiment of the disclosure, the target vehicle can be any vehicle, such as the vehicle involved in a traffic accident or the vehicle of a participant in the accident.
[0073] In this embodiment of the disclosure, multimodal information, also known as multidimensional information, includes, but is not limited to, images, videos, audio, central control information, etc.
[0074] In this embodiment of the disclosure, the video to be detected can be acquired by an external camera of the target vehicle, wherein the target vehicle may be shown in the video to be detected. It should be understood that, in addition to showing the target vehicle, the video to be detected may also show other targets (such as vehicles, pedestrians, obstacles, etc.).
[0075] The external camera can be installed on the top of the target vehicle, the rearview mirror, the upper part of the windshield, the trunk, etc. The external camera can have a surround view function or a monitoring function. For example, the external camera can be a multi-angle camera.
[0076] The number of external cameras can be at least one. For example, in order to improve the comprehensiveness and accuracy of image or video information collection, the number of external cameras on the target vehicle can be multiple.
[0077] As an example, multiple cameras with different perspectives can be installed on the exterior of the target vehicle. The video frames captured by each camera can include partial information about the target vehicle (hereinafter, the video frames captured by each camera will be referred to as views). By stitching together the multiple views captured by the multiple cameras, the target vehicle can be displayed in the resulting stitched image. That is, in this disclosure, a video to be detected can be generated based on the video captured by cameras with multiple perspectives, wherein each frame of the video to be detected can include views captured by cameras with multiple perspectives at the same time.
[0078] In this embodiment of the disclosure, the server can receive the video to be detected and multimodal information sent by the vehicle's in-vehicle terminal, wherein the multimodal information is collected from inside the cabin of the target vehicle.
[0079] Step S102: Extract the target video segment of the traffic accident from the video to be detected, and classify the collision type of the target video segment to obtain the target collision type corresponding to the traffic accident.
[0080] In this embodiment of the disclosure, the target collision type may include, but is not limited to: scrape, minor collision, severe collision, major accident, etc.
[0081] It should be noted that the video to be detected may include video clips of traffic accidents, video clips of the target vehicle driving normally, video clips of the vehicle stopped, etc. In order to reduce the processing burden on the server and improve the accuracy and reliability of subsequent traffic accident detection results, in this embodiment of the disclosure, the target video clip of the traffic accident can be extracted from the video to be detected, and the target video clip can be classified by collision type to obtain the target collision type corresponding to the traffic accident.
[0082] Step S103: Based on multimodal information, semantic classification and accident liability detection are performed on the target video segments to obtain the target accident type and target liability party corresponding to the traffic accident.
[0083] In this embodiment of the disclosure, the target accident type may include, but is not limited to: vehicle-to-vehicle accidents, vehicle-to-pedestrian accidents, vehicle-to-bicycle accidents, vehicle-only accidents, collisions between vehicles and fixed objects, railway crossing accidents, etc.
[0084] In this embodiment of the disclosure, semantic classification of the target video segment can be performed based on multimodal information to obtain the target accident type corresponding to the traffic accident. That is, semantic understanding of the target video segment can be performed based on multimodal information to determine the target accident type corresponding to the traffic accident.
[0085] In this embodiment of the disclosure, the responsible party for the accident can also be detected based on multimodal information to obtain the target responsible party for the traffic accident.
[0086] For example, target tracking can be performed on each video frame in the target video segment to determine whether each target (including but not limited to the target vehicle, such as other vehicles, pedestrians, obstacles, etc.) is speeding or changing lanes. Based on audio information, video information, and central control information in the multimodal information, it can be determined whether the driver of the target vehicle is operating improperly or driving normally. By combining the above information, the target responsible party for the traffic accident can be determined from each target.
[0087] For example, based on the audio information in the multimodal information, if it is determined that the driver complained "Why is this car driving so fast? I want to overtake him" before the traffic accident occurred, and target tracking is performed on each video frame in the target video segment to determine that the target vehicle was speeding and caused the rear-end collision, then the target responsible party in the traffic accident can be determined to include the target vehicle.
[0088] Step S104: Generate the detection results of the traffic accident based on the target accident type, target collision type, and target responsible party, and send the detection results to the vehicle terminal.
[0089] In this embodiment of the disclosure, the detection results of a traffic accident can be generated based on the target accident type, target collision type, and target responsible party corresponding to the traffic accident.
[0090] Optionally, the detection results may also include other information, such as the time of the traffic accident, the motion information of the collided targets, etc., and this disclosure does not impose any limitations on this.
[0091] In this embodiment of the disclosure, the server can also send the detection results of traffic accidents to the vehicle terminal to provide accident participants and traffic departments with a basis for accident judgment, improve the response speed and processing speed of traffic accidents, and alleviate the traffic pressure caused by traffic accidents.
[0092] The traffic accident detection method of this disclosure involves a server extracting target video segments of a traffic accident from a video to be detected, classifying the target video segments by collision type to obtain the target collision type corresponding to the traffic accident, and then performing semantic classification and accident liability party detection on the target video segments based on multimodal information collected from inside the target vehicle's cabin to obtain the target accident type and target liability party. Based on the target accident type, target collision type, and target liability party, a traffic accident detection result is generated and sent to the vehicle's infotainment system. Therefore, by having the server perform traffic accident detection on video collected from outside the target vehicle based on multimodal information collected from inside the target vehicle's cabin, the accuracy of the detection results can be improved. This not only provides a basis for accident judgment for accident participants and traffic departments, but also improves the response and processing speed of traffic accidents, thereby alleviating traffic pressure caused by traffic accidents. Furthermore, server-side traffic accident detection reduces the processing burden on the vehicle's infotainment system and improves detection efficiency.
[0093] It should be noted that the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution disclosed herein are all carried out with the consent of the user, and all comply with the provisions of relevant laws and regulations, and do not violate public order and good morals.
[0094] To clearly illustrate how the target video segment is classified into collision types and the target collision type corresponding to the traffic accident is obtained in the above embodiments of this disclosure, this disclosure also proposes a traffic accident detection method.
[0095] Figure 2 This is a flowchart illustrating the traffic accident detection method provided in Embodiment 2 of this disclosure.
[0096] like Figure 2 As shown, the traffic accident detection method may include the following steps:
[0097] Step S201: Receive the video to be detected sent by the vehicle's in-vehicle terminal and the multimodal information collected from inside the target vehicle's cabin.
[0098] Step S202: Extract the target video segment of the traffic accident from the video to be detected.
[0099] The explanation of steps S201 to S202 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.
[0100] Step S203: Perform collision detection on each video frame in the target video segment to determine the first target video frame from each video frame, wherein the target in the first target video frame collides.
[0101] In this embodiment of the disclosure, collision detection of targets (such as vehicles, pedestrians, obstacles, etc.) can be performed on each video frame in the target video segment to determine a first target video frame from each video frame, wherein the target in the first target video frame collides.
[0102] It should be noted that the number of video frames in which the target collision occurs can be multiple frames. The first target video frame can be the video frame with the smallest acquisition time among the multiple video frames in which the target collision occurs, or it can be any one of the multiple video frames in which the target collision occurs, or it can be the video frame with the largest collision area among the multiple video frames in which the target collision occurs, etc. This disclosure does not impose any limitations in this regard. The collision area refers to the area in the video frame where the target collision occurs, that is, the overlapping area of the target or instance.
[0103] In one possible implementation of the present disclosure, the collision detection method for the target is, for example, to perform target detection on each video frame in the target video segment to determine candidate video frames containing multiple targets from each video frame. For any candidate video frame, target detection can be performed on the candidate video frame to obtain multiple detection boxes in the candidate video frame, and based on the position information of the multiple detection boxes in the candidate video frame, it is determined whether the target in the candidate video frame has collided.
[0104] As an example, based on the position information of multiple detection boxes in a candidate video frame, it can be determined whether the multiple detection boxes overlap. If the multiple detection boxes do not overlap, it can be determined that the target in the candidate video frame has not collided. However, if at least two detection boxes overlap, it indicates that the target in the candidate video frame may collide. In this case, semantic segmentation or instance (i.e., object or semantic object) segmentation can be further performed on the target within the at least two detection boxes to obtain instances (i.e., objects or semantic objects, referred to as second instances in this disclosure) within the at least two detection boxes. Based on the position information of the at least two second instances in any candidate video frame, it can be determined whether the at least two second instances overlap. If the at least two second instances do not overlap, it is determined that the target in the candidate video frame has not collided; if the at least two second instances overlap, it is determined that the target in the candidate video frame has collided.
[0105] Therefore, it is possible to determine whether a target in a video frame has collided based on whether multiple detection boxes overlap, and whether instances within overlapping detection boxes overlap when multiple detection boxes overlap, thereby improving the accuracy of collision detection results.
[0106] In this disclosure, when it is determined that a collision has occurred with a target in a candidate video frame, a first target video frame can be determined based on that candidate video frame. For example, the candidate video frame with the smallest acquisition time among the candidate video frames where a collision has occurred can be used as the first target video frame; or, any one of the candidate video frames where a collision has occurred can be used as the first target video frame; or, the candidate video frame with the largest collision area among the candidate video frames where a collision has occurred can be used as the first target video frame, and so on.
[0107] Step S204: Determine the collision area where the target collided from the first target video frame.
[0108] In this embodiment of the disclosure, the collision area where the target collides can be determined from the first video frame.
[0109] As one possible implementation, the collision region can be determined as follows: the collision target in the first target video frame is segmented into instances to obtain multiple first instances (i.e., objects or semantic objects, including but not limited to target vehicles, other vehicles, pedestrians, obstacles, etc.), and the overlapping area of multiple first instances in the first target video frame is determined as the collision region.
[0110] Therefore, it is possible to determine the collision area where a target collides based on the overlapping area between multiple instances in a video frame, thereby improving the accuracy of collision area detection.
[0111] Step S205: Determine the spatial position of the collision region in the world coordinate system based on the image position of the collision region in the first target video frame.
[0112] The image position of the collision region in the first target video frame can be either the position of the collision region in the image coordinate system or the position of the collision region in the pixel coordinate system.
[0113] In the image coordinate system, the origin is the center point of the first target video frame, the horizontal axis (X-axis) points horizontally to the right, and the vertical axis (Y-axis) points horizontally downwards, with the unit being pixels. In the pixel coordinate system, the origin is the top-left corner of the first target video frame, the horizontal axis (X-axis) points horizontally to the right, and the horizontal axis (Y-axis) points horizontally downwards, with the unit being pixels.
[0114] For ease of explanation, the image position of the collision region in the first target video frame will be used as an example to illustrate the position of the collision region in the image coordinate system.
[0115] In this embodiment of the disclosure, the server can calculate the spatial position of the collision region in the world coordinate system based on the image position of the collision region in the first target video frame.
[0116] In one possible implementation of this disclosure, the spatial location is calculated as follows: the server can map the image position of the collision region in the image coordinate system to the camera coordinate system according to the mapping relationship between the image coordinate system and the camera coordinate system to obtain the position of the collision region in the camera coordinate system, and then, according to the mapping relationship between the camera coordinate system and the world coordinate system, map the position of the collision region in the camera coordinate system to the world coordinate system to obtain the spatial location of the collision region in the world coordinate system.
[0117] In one possible implementation of this disclosure, the spatial location is calculated as follows: the server can obtain the installation location of the image sensor on the target vehicle that collects the video to be detected, wherein the installation location is used to indicate the position of the image sensor in the vehicle body coordinate system of the target vehicle, and obtain the intrinsic and extrinsic parameters of the image sensor. Based on the intrinsic and extrinsic parameters, the image location and the installation location, the position of the collision area in the vehicle body coordinate system is determined, and based on the position of the collision area in the vehicle body coordinate system, the spatial location of the collision area in the world coordinate system is determined.
[0118] The installation location, internal parameters, and external parameters can be actively sent from the vehicle terminal to the server, or they can be actively queried from the vehicle terminal by the server. This disclosure does not impose any restrictions on this.
[0119] For example, the coordinates of the collision area in the camera coordinate system can be calculated based on the image position, intrinsic parameters, and extrinsic parameters. The coordinates of the collision area in the vehicle body coordinate system can be calculated based on the installation position and the coordinates of the collision area in the camera coordinate system. The coordinates of the collision area in the world coordinate system can be calculated based on the motion trajectory of the target vehicle (which can be obtained by detecting the target video segment using optical flow) and the coordinates of the collision area in the vehicle body coordinate system.
[0120] Therefore, by combining the intrinsic and extrinsic parameters of the image sensor, its installation location, and the image position of the collision area in the image coordinate system, the spatial position of the collision area in the world coordinate system can be determined, which can improve the accuracy of spatial position calculation.
[0121] Step S206: Determine the collision area corresponding to the collision area based on the spatial location of the collision area, and determine the target collision type from multiple set collision types based on the collision area.
[0122] The collision type can be set to include, but is not limited to: scrape, minor collision, severe collision, major accident, etc.
[0123] In this embodiment of the disclosure, the collision area corresponding to the collision region can be determined based on the spatial location of the collision region in the world coordinate system.
[0124] As an example, spatial location can be used to indicate the position of the boundary of the collision area in the world coordinate system, and the collision area can be directly calculated based on the spatial location.
[0125] As another example, when the number of first target video frames is multiple, the coordinates of the collision area in the world coordinate system at different viewpoints can be calculated based on the motion trajectory of the target vehicle (which can be detected by optical flow method on the target video segment) and the coordinates of the collision area in the vehicle coordinate system in the multiple first target video frames. Based on the coordinates of the collision area in the world coordinate system at different viewpoints, the projected area of the collision area in the world coordinate system at different viewpoints can be calculated. Based on the projected area of the collision area in the world coordinate system at different viewpoints, 3D reconstruction is performed to obtain the actual collision area on the target vehicle corresponding to the collision area in the first target video frame. Then, based on the actual collision area obtained by 3D reconstruction, the collision area corresponding to the collision area can be calculated. That is, the area of the actual collision area can be used as the collision area corresponding to the collision area.
[0126] In this embodiment of the disclosure, the target collision type can be determined from a plurality of set collision types based on the collision area.
[0127] For example, each collision type can have a corresponding collision area range. When the collision area corresponding to the collision region is within the collision area range corresponding to a certain collision type, then that collision type can be used as the target collision type.
[0128] Step S207: Based on multimodal information, semantic classification and accident liability detection are performed on the target video segments to obtain the target accident type and target liability party corresponding to the traffic accident.
[0129] Step S208: Based on the target accident type, target collision type, and target responsible party, generate the detection results of the traffic accident and send the detection results to the vehicle terminal.
[0130] The explanation of steps S207 to S208 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.
[0131] The traffic accident detection method of this disclosure can determine the spatial position of the collision area in the world coordinate system based on the image location of the collision area, and determine the collision area of the collision area based on the spatial position of the collision area in the world coordinate system. Therefore, the target collision type corresponding to the traffic accident can be determined based on the collision area, which can improve the accuracy of the target collision type determination. For example, the smaller the collision area, the lighter the collision severity indicated by the target collision type; the larger the collision area, the heavier the collision severity indicated by the target collision type.
[0132] To clearly illustrate how the target collision type is determined from multiple preset collision types based on the collision area in the above embodiments, this disclosure also proposes a traffic accident detection method.
[0133] Figure 3 This is a schematic flowchart of the traffic accident detection method provided in Embodiment 3 of this disclosure.
[0134] like Figure 3 As shown, the traffic accident detection method may include the following steps:
[0135] Step S301: Receive the video to be detected sent by the vehicle terminal of the target vehicle and the multimodal information collected from the cabin of the target vehicle.
[0136] Step S302: Extract the target video segment of the traffic accident from the video to be detected.
[0137] Step S303: Perform collision detection on each video frame in the target video segment to determine the first target video frame from each video frame, wherein the target in the first target video frame collides.
[0138] Step S304: Determine the collision area where the target collided from the first target video frame.
[0139] Step S305: Determine the spatial position of the collision region in the world coordinate system based on the image position of the collision region in the first target video frame.
[0140] Step S306: Determine the collision area corresponding to the collision region based on the spatial location of the collision region.
[0141] The explanation of steps S301 to S306 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.
[0142] Step S307: Based on the first acquisition time of the first target video frame, query the historical driving data of the target vehicle to obtain the driving position of the target vehicle at the first acquisition time.
[0143] In this embodiment of the disclosure, the server can query the historical driving data of the target vehicle based on the acquisition time of the first target video frame (referred to as the first acquisition time in this disclosure) to obtain the driving position of the target vehicle at the first acquisition time.
[0144] Step S308: Determine the collision location on the target vehicle corresponding to the spatial location and driving location.
[0145] In this embodiment of the disclosure, the collision position on the target vehicle can be determined based on the spatial position of the collision area in the world coordinate system and the driving position of the target vehicle at the first acquisition time.
[0146] Step S309: Determine the target collision type from multiple preset collision types based on the collision location and collision area.
[0147] The collision type can be set to include, but is not limited to: scrape, minor collision, severe collision, major accident, etc.
[0148] In this embodiment of the disclosure, the target collision type can be determined from multiple preset collision types by comprehensively considering the collision location and collision area corresponding to the collision area on the target vehicle.
[0149] For example, the collision area range and collision position range corresponding to each collision type can be preset. When the collision area corresponding to the collision region is within the collision area range corresponding to a certain collision type, and the collision position is also within the collision position range corresponding to that collision type, then that collision type can be used as the target collision type.
[0150] For example, when the collision location is located in the edge area of the target vehicle and the collision area is small, the target collision type can be determined to be a scrape; when the collision location is located in the non-edge area of the target vehicle and the collision area is large, the target collision type can be determined to be a severe collision.
[0151] Step S310: Based on multimodal information, semantic classification and accident liability detection are performed on the target video segments to obtain the target accident type and target liability party corresponding to the traffic accident.
[0152] Step S311: Generate the detection results of the traffic accident based on the target accident type, target collision type, and target responsible party, and send the detection results to the vehicle terminal.
[0153] The explanation of steps S310 to S311 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.
[0154] The traffic accident detection method of this disclosure combines the actual collision area and collision position of the collision area on the target vehicle to determine the target collision type of the traffic accident, which can further improve the accuracy of the target collision type determination.
[0155] To clearly illustrate how multimodal information is used to detect the responsible party in a target video segment in any embodiment of this disclosure, this disclosure also proposes a traffic accident detection method.
[0156] Figure 4 This is a schematic flowchart of the traffic accident detection method provided in Embodiment 4 of this disclosure.
[0157] like Figure 4 As shown, the traffic accident detection method may include the following steps:
[0158] Step S401: Receive the video to be detected sent by the vehicle's in-vehicle terminal and the multimodal information collected from inside the target vehicle's cabin.
[0159] Step S402: Extract the target video segment of the traffic accident from the video to be detected, and classify the collision type of the target video segment to obtain the target collision type corresponding to the traffic accident.
[0160] Step S403: Based on multimodal information, perform semantic classification on the target video segment to obtain the target accident type corresponding to the traffic accident.
[0161] The explanation of steps S401 to S403 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.
[0162] Step S404: Perform optical flow detection on the target video segment to obtain optical flow change information of the target that has collided.
[0163] In this embodiment of the disclosure, optical flow detection can be performed on each video frame in the target video segment based on the optical flow method to obtain the optical flow change information of the target that has collided.
[0164] Step S405: Based on the optical flow change information of the collided target, determine the motion information of the collided target, wherein the motion information is used to indicate the driving position and driving speed of the collided target at multiple times.
[0165] In this embodiment of the disclosure, the motion information of the colliding target can be determined based on the optical flow change information of the colliding target. The motion information is used to indicate the driving position, driving speed and driving trajectory of the colliding target at multiple times.
[0166] For example, based on the driving position at multiple times, the driving trajectory of the target involved in the collision can be determined, such as whether it changed lanes. Based on the driving speed at multiple times, it can be determined whether the target involved in the collision was speeding.
[0167] Step S406: Based on the motion information of the collided targets, and the audio and multimodal information corresponding to the target video segments, determine the responsible party from the collided targets.
[0168] In this embodiment of the disclosure, the responsible party for the collision can be determined from the targets that have collided, based on the motion information of the targets and the audio and multimodal information corresponding to the target video segments.
[0169] For example, based on the motion information of the targets involved in the collision, it can be determined whether each target was speeding or changing lanes. Based on audio, video, and central control information in multimodal information, it can be determined whether the driver of the target vehicle was operating improperly or driving normally. By combining the above information, the responsible party for the traffic accident can be determined.
[0170] For example, suppose that based on the target vehicle's motion information, it is determined that the target vehicle was driving normally (e.g., not speeding, not changing lanes), but based on the video information in the multimodal information, it is determined that the driver of the target vehicle looked down to pick up his mobile phone before the time of the traffic accident, causing the target vehicle to rear-end other vehicles, then it can be determined that the target vehicle is among the parties responsible for the traffic accident.
[0171] Step S407: Based on the target accident type, target collision type, and target responsible party, generate the detection results of the traffic accident and send the detection results to the vehicle terminal.
[0172] The explanation of step S407 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.
[0173] The traffic accident detection method of this disclosure can comprehensively analyze the motion information of the collided target, the audio information corresponding to the target video segment, and the multimodal information to determine the responsible party from the collided targets, thereby improving the accuracy of the determination of the responsible party.
[0174] To clearly illustrate how a target video segment of a traffic accident is extracted from a video to be detected in any embodiment of this disclosure, this disclosure also proposes a traffic accident detection method.
[0175] Figure 5 This is a flowchart illustrating the traffic accident detection method provided in Embodiment 5 of this disclosure.
[0176] like Figure 5 As shown, the traffic accident detection method may include the following steps:
[0177] Step S501: Receive the video to be detected sent by the vehicle's in-vehicle terminal and the multimodal information collected from inside the target vehicle's cabin.
[0178] The explanation of step S501 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.
[0179] Step S502: Perform target collision detection on each video frame in the video to be detected to determine the second target video frame from each video frame; wherein, the target in the second target video frame collides.
[0180] In this embodiment of the disclosure, collision detection of the target can be performed on each video frame in the target video segment to determine a second target video frame from each video frame, wherein the target in the second target video frame collides.
[0181] It should be noted that the number of video frames in which the target collision occurs can be multiple frames. For example, the second target video frame can be the video frame with the smallest acquisition time among the multiple video frames in which the target collision occurs, or it can be the video frame with the second smallest acquisition time among the multiple video frames in which the target collision occurs, etc. This disclosure does not limit this.
[0182] Step S503: Extract the target video segment from the video to be detected based on the second acquisition time of the second target video frame.
[0183] In this embodiment of the disclosure, a target video segment can be extracted from the video to be detected based on the acquisition time of the second target video frame (referred to as the second acquisition time in this disclosure).
[0184] For example, a video segment within a set duration before and after the second acquisition time can be extracted from the video to be detected to obtain the target video segment. The set duration is a pre-set duration threshold, such as 14 seconds, 15 seconds, 16 seconds, etc., and this disclosure does not limit it.
[0185] Step S504: Classify the target video clips by collision type to obtain the target collision type corresponding to the traffic accident.
[0186] Step S505: Based on multimodal information, semantic classification and accident liability detection are performed on the target video segments to obtain the target accident type and target liability party corresponding to the traffic accident.
[0187] Step S506: Based on the target accident type, target collision type, and target responsible party, generate the detection results of the traffic accident and send the detection results to the vehicle terminal.
[0188] The explanation of steps S504 to S506 can be found in the relevant descriptions in any embodiment of this disclosure, and will not be repeated here.
[0189] The traffic accident detection method of this disclosure can extract the video segment of the target collision from the video to be detected and perform traffic accident detection only on the video segment, which can improve detection efficiency.
[0190] To clearly illustrate how collision detection of targets is performed on each video frame in the video to be detected in any embodiment of this disclosure in order to determine the second target video frame from each video frame, this disclosure also proposes a traffic accident detection method.
[0191] Figure 6 This is a flowchart illustrating the traffic accident detection method provided in Embodiment Six of this disclosure.
[0192] like Figure 6 As shown, the traffic accident detection method may include the following steps:
[0193] Step S601: Receive the video to be detected sent by the vehicle's in-vehicle terminal and the multimodal information collected from the cabin of the target vehicle.
[0194] The explanation of step S601 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.
[0195] Step S602: From each video frame in the video to be detected, determine the candidate video frames that contain multiple targets.
[0196] In this embodiment of the disclosure, target detection can be performed on each video frame in the video to be detected in order to determine candidate video frames containing multiple targets from each video frame.
[0197] Step S603: For any candidate video frame, perform target detection on the candidate video frame to obtain multiple detection boxes in the candidate video frame.
[0198] In this embodiment of the disclosure, for any candidate video frame, target detection can be performed on the candidate video frame to obtain multiple detection boxes in the candidate video frame, wherein each detection box contains a target (such as a vehicle, pedestrian, obstacle, etc.).
[0199] As an example, based on object detection technology, various semantic objects (i.e. targets) that may be related to traffic accidents in any candidate video frame can be detected, and detection boxes with minimum bounding rectangles can be marked to obtain each detection box.
[0200] Step S604: Based on the position information of multiple detection boxes in any candidate video frame, determine whether the target in any candidate video frame has collided.
[0201] In this embodiment of the disclosure, it can be determined whether the target in any candidate video frame has collided based on the position information of multiple detection boxes in any candidate video frame. If it is determined that the target in any candidate video frame has not collided, then any candidate video frame can not be used as the second target video frame. That is, the next candidate video frame can be obtained, and steps S603 to S605 can be executed on the next candidate video frame. If it is determined that the target in any candidate video frame has collided, then step S605 can be executed.
[0202] In one possible implementation of the present disclosure, the method for determining whether a target in any candidate video frame has collided is, for example, as follows: First, based on the position information of multiple detection boxes in any candidate video frame, it is determined whether the multiple detection boxes overlap. If the multiple detection boxes do not overlap, it can be determined that the target in any candidate video frame has not collided.
[0203] If at least two detection boxes overlap, it indicates that the target in any candidate video frame may collide. In this case, the target within the at least two detection boxes can be further segmented to obtain the second instance within the at least two detection boxes. Based on the position information of the at least two second instances in any candidate video frame, it is determined whether the at least two second instances overlap. If the at least two second instances do not overlap, it is determined that the target in any candidate video frame has not collided. If the at least two second instances overlap, it is determined that the target in any candidate video frame has collided.
[0204] Therefore, it is possible to determine whether a target in a video frame has collided based on whether multiple detection boxes overlap, and whether instances within overlapping detection boxes overlap when multiple detection boxes overlap. This can improve the accuracy of collision detection results and thus improve the accuracy of capturing target video segments of traffic accidents.
[0205] Step S605: If a collision occurs between the target and the target in any candidate video frame, determine the second target video frame based on any candidate video frame.
[0206] In this embodiment of the disclosure, if a collision occurs with a target in any candidate video frame, a second target video frame can be determined based on that candidate video frame. For example, the candidate video frame with the smallest acquisition time among the candidate video frames where a collision occurs can be used as the second target video frame, or the candidate video frame with the second smallest acquisition time among the candidate video frames where a collision occurs can be used as the second target video frame, and so on.
[0207] Step S606: Extract the target video segment from the video to be detected based on the second acquisition time of the second target video frame.
[0208] Step S607: Classify the target video clips by collision type to obtain the target collision type corresponding to the traffic accident.
[0209] Step S608: Based on multimodal information, semantic classification and accident liability detection are performed on the target video segments to obtain the target accident type and target liability party corresponding to the traffic accident.
[0210] Step S609: Based on the target accident type, target collision type, and target responsible party, generate the detection results of the traffic accident and send the detection results to the vehicle terminal.
[0211] The explanation of steps S606 to S609 can be found in the relevant description in any embodiment of this disclosure, and will not be repeated here.
[0212] In any embodiment of this disclosure, when the target responsible party is the target vehicle, the server can also obtain the time of occurrence of the traffic accident (such as the acquisition time of the second target video frame), and determine the driving behavior of the driver of the target vehicle at the time of occurrence based on multimodal information, audio information corresponding to the target video segment, and detection results. The server can also determine the cause of the traffic accident based on the driving behavior. That is, the server can analyze the driving behavior of the driver of the target vehicle based on the detection results, audio information, and multimodal information to obtain the cause of the traffic accident, and send the cause of the traffic accident to the vehicle terminal so that the driver of the target vehicle can regulate his driving behavior based on the cause of the accident to avoid similar accidents and improve the safety of vehicle driving.
[0213] For example, abnormal sounds (abnormal sounds produced by tires, brakes, engines, etc.) can be extracted from the audio information of the target video clip, and the cause of the traffic accident (accident caused by vehicle malfunction or improper operation by the driver) can be determined based on the extracted abnormal sounds.
[0214] For example, based on multimodal information collected or gathered in the cockpit (including the driver's voice or comments such as "Why is this car driving so fast? I'm going to overtake it," "Oh no, how did a bicycle suddenly appear?", "Scream," "Ah, we crashed," etc.), the time and cause of a traffic accident can be determined.
[0215] For example, assuming that based on multimodal information, it is determined that the driver of the target vehicle was looking at his mobile phone while driving, causing the target vehicle to rear-end other vehicles, then the cause of the traffic accident could be: Because you were driving at XX:XX:XX (the time of the traffic accident) on XX:XX:XX (the time of the traffic accident), a traffic accident occurred.
[0216] In any embodiment of this disclosure, when the target responsible party is not the target vehicle, the server can also generate driving suggestions based on multimodal information, audio information corresponding to the target video segment, and detection results. The driving suggestions are used to avoid or prevent traffic accidents corresponding to the target accident type and the target collision type, and send the driving suggestions to the vehicle terminal to enhance the safety awareness of the driver of the target vehicle, and enable the driver to optimize driving behavior according to the driving suggestions, reduce the probability of similar traffic accidents, and improve the safety of vehicle driving.
[0217] For example, suppose the motion information of vehicle A in the detection results indicates that vehicle A illegally changed lanes, resulting in a collision between the target vehicle and vehicle A, then the driving advice could be: While driving, pay attention to avoiding surrounding vehicles.
[0218] The traffic accident detection method of this disclosure can determine the second target video frame where a target collision occurs from each video frame of the video to be detected by performing target detection, instance segmentation, and instance overlap detection on the video to be detected, thereby improving the accuracy of determining the second target video frame and thus improving the accuracy of extracting the target video segment of the traffic accident.
[0219] The above are implementation examples of methods executed on the server side. This disclosure also proposes a traffic accident detection method executed by the vehicle terminal.
[0220] Figure 7 This is a flowchart illustrating the traffic accident detection method provided in Embodiment 7 of this disclosure.
[0221] The traffic accident detection method provided in this disclosure can be applied to the vehicle's infotainment system.
[0222] like Figure 7 As shown, the traffic accident detection method may include the following steps:
[0223] Step S701: Send the video to be detected and the multimodal information collected from the cabin of the target vehicle to the server.
[0224] Step S702: Receive the detection results sent by the server; wherein, the detection results are obtained by extracting the target video segment of the traffic accident from the video to be detected, classifying the target video segment by collision type to obtain the target collision type corresponding to the traffic accident, and performing semantic classification and accident liability detection on the target video segment based on multimodal information to obtain the target accident type and target liability type corresponding to the traffic accident, and generating the target accident type, target collision type and target liability type based on the target accident type, target collision type and target liability type.
[0225] It should be noted that the explanations of the various method embodiments executed on the server side in the foregoing embodiments also apply to the method embodiments executed on the vehicle terminal, as their implementation principles are similar and will not be repeated here.
[0226] In one possible implementation of this disclosure, in order to reduce the processing burden on the server, the video to be detected sent from the vehicle terminal to the server can be a relatively short video segment. For example, the user on the vehicle terminal can extract the video to be detected from the video stream collected by the image sensor based on the approximate time of the traffic accident.
[0227] As an example, a video stream captured by an image sensor outside the target vehicle can be acquired. In response to a user-triggered input operation, the start and end times of the capture to be performed can be determined, and the video to be detected can be extracted from the video stream based on the start and end times.
[0228] This allows for the sending of relatively short videos to the server, improving server processing efficiency and reducing the waiting time for vehicle-mounted systems to obtain traffic accident detection results.
[0229] The traffic accident detection method of this disclosure involves sending a video to be detected and multimodal information collected from the cabin of the target vehicle to a server via the vehicle-mounted terminal; receiving detection results from the server; wherein the detection results are obtained by extracting the target video segment of the traffic accident from the video to be detected, classifying the target video segment by collision type to obtain the target collision type corresponding to the traffic accident, and performing semantic classification and accident liability party detection on the target video segment based on the multimodal information to obtain the target accident type and target liability party corresponding to the traffic accident, and generating a system based on the target accident type, target collision type, and target liability party. Therefore, the server performs traffic accident detection on the video collected from the outside of the target vehicle based on the multimodal information collected from the cabin of the target vehicle, and obtains the detection results. This not only improves the accuracy of the detection results but also provides a basis for accident judgment for accident participants and traffic departments, improving the response and processing speed of traffic accidents and alleviating traffic pressure caused by traffic accidents. Furthermore, server-side traffic accident detection reduces the processing burden on the vehicle-mounted terminal and improves detection efficiency.
[0230] In any embodiment of this disclosure, traffic accident scene features can be extracted from videos captured by multi-angle cameras outside the vehicle. Multimodal capabilities and video semantic analysis capabilities can be used to determine the type of traffic accident, locate the collision position, calculate relevant motion information during the collision process, and identify the responsible party. Simultaneously, driver behavior analysis can be performed based on collected multimodal (or multi-dimensional) information from inside the cabin, providing safe driving suggestions. This can provide parties involved and traffic management departments with a basis for accident judgment, effectively improving the response and processing speed of traffic accidents, thereby effectively alleviating traffic pressure caused by traffic accidents and enhancing drivers' driving safety awareness.
[0231] The prerequisites for the traffic accident detection method provided in this disclosure are:
[0232] 1. Hardware environment.
[0233] Multi-angle cameras need to be installed on the exterior of the vehicle, and in-vehicle cameras need to be installed inside the vehicle. The exterior cameras can capture the vehicle's movement, while the interior cameras can capture the status of the occupants inside the vehicle.
[0234] 2. Software capabilities.
[0235] To improve the generalization performance of the model, knowledge-enhanced large models (such as Wenxin Yiyan, GPT (Generative Pre-trained Transformer)) can be used as the underlying model support to analyze traffic accidents and generate accident analysis reports and related driving safety suggestions.
[0236] Using a 15-second timeframe as an example, the implementation process of a traffic accident detection method can be as follows: Figure 8 As shown, the main steps include:
[0237] 1. Input the video captured by the multi-angle camera outside the vehicle (referred to as the video to be detected in this disclosure), and use a knowledge-enhanced large model to detect various targets (or semantic objects) that may be related to traffic accidents in each video frame, and mark the detection box with the minimum bounding rectangle.
[0238] 2. By checking whether the detection boxes overlap, it can be determined whether a target collision may occur in a certain video frame.
[0239] 3. If a target in a video frame may collide, then further segment the targets within the overlapping detection boxes, and then determine whether there is a target collision in the video frame by whether the segmented instances overlap.
[0240] 4. If a collision occurs with a target in the video frame, the knowledge-enhanced big model will intelligently extract video segments within 15 seconds before and after the video frame, and collect multimodal information (or multidimensional information) such as audio information, central control information and driver behavior information (images, audio) in the cockpit during the same period.
[0241] 5. Using a knowledge-enhanced large model, semantic understanding of video clips before and after a collision is performed based on multi-dimensional information to determine the type of accident.
[0242] Furthermore, based on the position of the collision region in the image coordinate system in the video frame and the installation position of the multi-angle camera, the position or coordinates of the collision region in the world coordinate system are calculated. Based on the position or coordinates of the collision region in the world coordinate system and the collision area corresponding to the collision region, the collision type is determined.
[0243] Furthermore, by calculating the optical flow changes of each party involved in the accident (i.e., each target involved in the collision), the motion information of each party involved in the accident is calculated during the collision. Then, a knowledge-enhanced large model is used to detect the party responsible for the accident based on multi-dimensional information and the motion information of each party involved in the accident, and the target responsible party is determined from the parties involved in the accident.
[0244] 6. Integrate the information from various dimensions in step 5 (such as accident type, collision type, motion information, and target responsible party) through a knowledge-enhanced large model, and output a graphic report of the entire event process through multimodal capabilities.
[0245] 7. Combining incident reports and multimodal information collected in the cockpit, a knowledge-enhanced large model is used for information integration and analysis to output safe driving suggestions to avoid or prevent similar traffic accidents.
[0246] Therefore, in traffic accident analysis, the multimodal capabilities and video semantic understanding of large knowledge-enhanced models enable multi-dimensional calculations, judgments, and analyses of videos, resulting in higher accuracy. Furthermore, in addition to analyzing video information collected by external cameras, it also collects and analyzes multimodal information from inside the cabin, providing safe driving suggestions beyond traffic accident analysis, thus helping drivers avoid or prevent similar accidents. In summary, this approach achieves higher precision, stronger generalization, and enhanced analytical capabilities.
[0247] As an application scenario, on the in-vehicle infotainment system, users can access the "Traffic Accident Analysis Assistant" application via the vehicle's central control screen (e.g., ...). Figure 9 (As shown).
[0248] Users can select or input the time range of the traffic accident in the
Traffic Accident Analysis Assistant
[0249] Once the test results are generated, users can view the accident analysis report, driving safety recommendations, and accident analysis videos (such as...). Figure 11 (As shown).
[0250] The interaction process between the vehicle-mounted system and the server can be as follows: Figure 12 As shown, specifically, after a traffic accident occurs, users can manually access the "Traffic Accident Analysis Assistant" application on the vehicle's infotainment system, select the time range of the traffic accident (accurate to the minute), and click confirm to generate an accident analysis report and safe driving suggestions.
[0251] After receiving the video from the vehicle's infotainment system, the server intelligently extracts 15-second video clips before and after the accident. It then uses the multimodal capabilities and video semantic understanding of a knowledge-enhanced model to analyze the data. The resulting accident analysis report includes four dimensions: "Accident Type," "Collision Type," "Collision-Related Motion Information," and "Target Liability Party." The output accident analysis video displays vehicle bounding boxes and the video analysis process, with a focus on predicting and analyzing driving behavior in the pre- and post-collision video clips. Simultaneously, combining collected multimodal information from within the cabin, the server utilizes the image processing capabilities of the knowledge-enhanced model to analyze driver behavior. If the driver is involved in an accident and is the target liability party, the server analyzes the driver's driving behavior to determine the cause of the accident; if the driver is involved in an accident but is not the target liability party, the server provides safe driving suggestions to avoid similar accidents.
[0252] Ultimately, users will receive a summary of traffic accident analysis results. Users can use these results to determine the outcome of an accident themselves, or provide information to traffic police for their assessment, thus improving the efficiency of traffic accident handling. Furthermore, users can improve their driving habits and raise their awareness of safe driving through safe driving suggestions, thereby reducing the occurrence of similar traffic accidents at their source.
[0253] In summary, the traffic accident detection method provided in this disclosure leverages the multimodal capabilities and image / video analysis capabilities of a knowledge-enhanced large-scale model to analyze traffic accidents and intelligently determine their causes. This significantly improves the response and processing speed of traffic accidents, thereby effectively alleviating road congestion. This method can be well applied to real-world scenarios. Through extensive training data and self-data augmentation, the knowledge-enhanced large-scale model achieves higher accuracy, stronger generalization, and enhanced analytical capabilities.
[0254] With the above Figures 1 to 6 Corresponding to the traffic accident detection method provided in the embodiments, this disclosure also provides a traffic accident detection device. Because the traffic accident detection device provided in the embodiments of this disclosure is similar to the one described above... Figures 1 to 6 The traffic accident detection method provided in the embodiments corresponds to the traffic accident detection method provided in the embodiments of this disclosure. Therefore, the implementation method of the traffic accident detection method is also applicable to the traffic accident detection device provided in the embodiments of this disclosure, and will not be described in detail in the embodiments of this disclosure.
[0255] Figure 13 This is a schematic diagram of the traffic accident detection device provided in Embodiment 8 of this disclosure.
[0256] like Figure 13 As shown, the traffic accident detection device 1300 can be applied to the server side and includes: a receiving module 1301, an interception module 1302, a classification module 1303, a processing module 1304, a first generation module 1305, and a sending module 1306.
[0257] The receiving module 1301 is used to receive the video to be detected sent by the vehicle's in-vehicle terminal and the multimodal information collected from inside the target vehicle's cabin.
[0258] The extraction module 1302 is used to extract target video segments of traffic accidents from the video to be detected.
[0259] The classification module 1303 is used to classify the collision type of the target video clip to obtain the target collision type corresponding to the traffic accident.
[0260] The processing module 1304 is used to perform semantic classification and accident liability detection on the target video segments based on multimodal information, so as to obtain the target accident type and target liability party corresponding to the traffic accident.
[0261] The first generation module 1305 is used to generate the detection results of traffic accidents based on the target accident type, target collision type, and target responsible party.
[0262] The sending module 1306 is used to send the detection results to the vehicle's infotainment system.
[0263] In one possible implementation of this disclosure, the classification module 1303 is configured to: perform collision detection on each video frame in the target video segment to determine a first target video frame from each video frame, wherein a target in the first target video frame collides; determine the collision region where the target collides from the first target video frame; determine the spatial position of the collision region in the world coordinate system based on the image position of the collision region in the first target video frame; determine the collision area corresponding to the collision region based on the spatial position of the collision region; and determine the target collision type from multiple preset collision types based on the collision area.
[0264] In one possible implementation of this disclosure, the classification module 1303 is configured to: obtain the installation position of the image sensor on the target vehicle that collects the video to be detected; wherein the installation position is used to indicate the position of the image sensor in the vehicle body coordinate system of the target vehicle; obtain the intrinsic and extrinsic parameters of the image sensor; determine the position of the collision area in the vehicle body coordinate system based on the intrinsic and extrinsic parameters, the image position, and the installation position; and determine the spatial position of the collision area in the world coordinate system based on the position of the collision area in the vehicle body coordinate system.
[0265] In one possible implementation of this disclosure, the classification module 1303 is configured to: perform instance segmentation on the targets that collide in the first target video frame to obtain multiple first instances; and determine the collision region based on the overlapping region of the multiple first instances in the first target video frame.
[0266] In one possible implementation of this disclosure, the classification module 1303 is configured to: query the historical driving data of the target vehicle based on the first acquisition time of the first target video frame to obtain the driving position of the target vehicle at the first acquisition time; determine the collision position of the collision area on the target vehicle based on the spatial position and the driving position; and determine the target collision type from multiple preset collision types based on the collision position and the collision area.
[0267] In one possible implementation of this disclosure, the processing module 1304 is configured to: perform optical flow detection on the target video segment to obtain optical flow change information of the colliding target; determine the motion information of the colliding target based on the optical flow change information of the colliding target, wherein the motion information is used to indicate the driving position and driving speed of the colliding target at multiple times; and determine the responsible party from the colliding targets based on the motion information of the colliding target, as well as the audio information and multimodal information corresponding to the target video segment.
[0268] In one possible implementation of this disclosure, the interception module 1302 is configured to: perform target collision detection on each video frame in the video to be detected, so as to determine a second target video frame from each video frame; wherein, the target in the second target video frame collides; and intercept a target video segment from the video to be detected according to the second acquisition time of the second target video frame.
[0269] In one possible implementation of this disclosure, the interception module 1302 or the classification module 1303 is configured to: determine candidate video frames containing multiple targets from each video frame; perform target detection on any candidate video frame to obtain multiple detection boxes in any candidate video frame; and determine whether a target in any candidate video frame has collided based on the position information of the multiple detection boxes in any candidate video frame.
[0270] In one possible implementation of this disclosure, the interception module 1302 or the classification module 1303 is configured to: determine whether multiple detection boxes overlap based on the position information of multiple detection boxes in any candidate video frame; if at least two detection boxes overlap, perform instance segmentation on the target within the at least two detection boxes to obtain a second instance within the at least two detection boxes; determine whether the at least two second instances overlap based on the position information of the at least two second instances in any candidate video frame; if at least two second instances overlap, determine that a target in any candidate video frame has collided; if none of the multiple detection boxes overlap, or if at least two second instances do not overlap, determine that a target in any candidate video frame has not collided.
[0271] In one possible implementation of this disclosure, the traffic accident detection device 1300 may further include:
[0272] The acquisition module is used to obtain the time of occurrence of a traffic accident when the target responsible party is the target vehicle.
[0273] The first determination module is used to determine the driving behavior of the driver of the target vehicle at the time of the incident based on multimodal information, audio information corresponding to the target video segment, and detection results.
[0274] The second determination module is used to determine the cause of a traffic accident based on driving behavior.
[0275] The sending module 1306 is also used to send the cause of the occurrence to the vehicle's infotainment system.
[0276] In one possible implementation of this disclosure, the traffic accident detection device 1300 may further include:
[0277] The second generation module is used to generate driving suggestions based on multimodal information, audio information corresponding to the target video segment, and detection results when the target responsible party is not the target vehicle.
[0278] The sending module 1306 is also used to send driving suggestions to the vehicle terminal; wherein, the driving suggestions are used to avoid traffic accidents corresponding to the target accident type and the target collision type.
[0279] The traffic accident detection device of this embodiment extracts target video segments of traffic accidents from videos to be detected via a server, classifies the target video segments by collision type to obtain the target collision type corresponding to the traffic accident, and performs semantic classification and accident liability party detection on the target video segments based on multimodal information collected from the cabin of the target vehicle to obtain the target accident type and target liability party corresponding to the traffic accident. Based on the target accident type, target collision type, and target liability party, a traffic accident detection result is generated and sent to the vehicle-mounted terminal. Therefore, by having the server perform traffic accident detection on videos collected from outside the target vehicle based on multimodal information collected from the cabin of the target vehicle, the detection result can not only improve the accuracy of the detection result but also provide a basis for accident judgment for accident participants and traffic departments, improving the response and processing speed of traffic accidents and alleviating traffic pressure caused by traffic accidents. Furthermore, having the traffic accident detection performed by the server can reduce the processing burden on the vehicle-mounted terminal and improve detection efficiency.
[0280] With the above Figure 7 Corresponding to the traffic accident detection method provided in the embodiments, this disclosure also provides a traffic accident detection device. Because the traffic accident detection device provided in the embodiments of this disclosure is similar to the one described above... Figure 7 The traffic accident detection method provided in the embodiments corresponds to the traffic accident detection method provided in the embodiments of this disclosure. Therefore, the implementation method of the traffic accident detection method is also applicable to the traffic accident detection device provided in the embodiments of this disclosure, and will not be described in detail in the embodiments of this disclosure.
[0281] Figure 14 This is a schematic diagram of the traffic accident detection device provided in Embodiment 9 of this disclosure.
[0282] like Figure 14As shown, the traffic accident detection device 1400 can be applied to the vehicle terminal of the target vehicle, including: a sending module 1401 and a receiving module 1402.
[0283] The sending module 1401 is used to send the video to be detected and the multimodal information collected from the cabin of the target vehicle to the server.
[0284] The receiving module 1402 is used to receive the detection results sent by the server. The detection results are obtained by extracting the target video segment of the traffic accident from the video to be detected, classifying the collision type of the target video segment to obtain the target collision type corresponding to the traffic accident, and performing semantic classification and accident liability detection on the target video segment based on multimodal information to obtain the target accident type and target liability type corresponding to the traffic accident. The results are generated based on the target accident type, target collision type and target liability type.
[0285] In one possible implementation of this disclosure, the traffic accident detection device 800 may further include:
[0286] The acquisition module is used to acquire video streams collected by image sensors on the exterior of the target vehicle.
[0287] The determination module is used to determine the start and end times of the interception in response to the input operation.
[0288] The extraction module is used to extract the video to be detected from the video stream based on the start and end extraction times.
[0289] The traffic accident detection device of this embodiment sends a video to be detected and multimodal information collected from the cabin of the target vehicle to a server via the vehicle-mounted terminal; it also receives detection results from the server. The detection results are derived by extracting target video segments of the traffic accident from the video to be detected, classifying the target video segments by collision type to obtain the target collision type corresponding to the traffic accident, and performing semantic classification and accident liability party detection on the target video segments based on the multimodal information to obtain the target accident type and target liability party corresponding to the traffic accident. This is then generated based on the target accident type, target collision type, and target liability party. Therefore, the server performs traffic accident detection on the video collected from the outside of the target vehicle based on the multimodal information collected from the cabin of the target vehicle, obtaining detection results. This not only improves the accuracy of the detection results but also provides a basis for accident judgment for accident participants and traffic departments, improving the response and processing speed of traffic accidents and alleviating traffic pressure caused by traffic accidents. Furthermore, server-side traffic accident detection reduces the processing burden on the vehicle-mounted terminal and improves detection efficiency.
[0290] To implement the above embodiments, this disclosure also provides an electronic device, which may include at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the traffic accident detection method proposed in any of the above embodiments of this disclosure.
[0291] To implement the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the traffic accident detection method proposed in any of the above embodiments of this disclosure.
[0292] To implement the above embodiments, this disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, implements the traffic accident detection method proposed in any of the above embodiments of this disclosure.
[0293] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0294] Figure 15 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. The electronic device may include the server and client described in the above embodiments. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0295] like Figure 15 As shown, the electronic device 1500 includes a computing unit 1501, which can perform various appropriate actions and processes according to a computer program stored in ROM (Read-Only Memory) 1502 or loaded from storage unit 1508 into RAM (Random Access Memory) 1503. The RAM 1503 may also store various programs and data required for the operation of the electronic device 1500. The computing unit 1501, ROM 1502, and RAM 1503 are interconnected via bus 1504. An I / O (Input / Output) interface 1505 is also connected to bus 1504.
[0296] Multiple components in electronic device 1500 are connected to I / O interface 1505, including: input unit 1506, such as keyboard, mouse, etc.; output unit 1507, such as various types of monitors, speakers, etc.; storage unit 1508, such as disk, optical disk, etc.; and communication unit 1509, such as network card, modem, wireless transceiver, etc. Communication unit 1509 allows electronic device 1500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0297] The computing unit 1501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1501 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 1501 performs the various methods and processes described above, such as the traffic accident detection method described above. For example, in some embodiments, the traffic accident detection method described above can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 1508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1500 via ROM 1502 and / or communication unit 1509. When the computer program is loaded into RAM 1503 and executed by the computing unit 1501, one or more steps of the traffic accident detection method described above can be performed. Alternatively, in other embodiments, the computing unit 1501 may be configured to perform the traffic accident detection method described above by any other suitable means (e.g., by means of firmware).
[0298] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0299] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0300] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0301] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0302] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.
[0303] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers integrated with blockchain technology.
[0304] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0305] Deep learning is a new research direction in the field of machine learning. It learns the inherent patterns and hierarchical representations of sample data. The information gained during this learning process is very helpful in interpreting data such as text, images, and sound. Its ultimate goal is to enable machines to have analytical and learning capabilities like humans, and to recognize data such as text, images, and sound.
[0306] According to the technical solution of this disclosure, the server extracts target video segments of traffic accidents from the video to be detected, classifies the target video segments by collision type, and obtains the target collision type corresponding to the traffic accident. Based on multimodal information collected from the cabin of the target vehicle, semantic classification and accident liability party detection are performed on the target video segments to obtain the target accident type and target liability party corresponding to the traffic accident. Based on the target accident type, target collision type, and target liability party, a traffic accident detection result is generated and sent to the vehicle terminal. Therefore, by having the server perform traffic accident detection on video collected from outside the target vehicle based on multimodal information collected from the cabin of the target vehicle, the detection result can not only improve the accuracy of the detection result, but also provide a basis for accident judgment for accident participants and traffic departments, improving the response and processing speed of traffic accidents and alleviating traffic pressure caused by traffic accidents. Furthermore, having the traffic accident detection performed by the server can reduce the processing burden on the vehicle terminal and improve detection efficiency.
[0307] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution proposed in this disclosure can be achieved, and this is not limited herein.
[0308] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A traffic accident detection method, applied on a server side, comprising: Receive the video to be detected sent by the vehicle's in-vehicle terminal and the multimodal information collected from the cabin of the target vehicle; The target video segment of the traffic accident is extracted from the video to be detected, and collision detection is performed on each video frame in the target video segment to classify the collision type, so as to obtain the target collision type corresponding to the traffic accident. Based on multimodal information collected at the same time as the target video segment, semantic classification and accident liability party detection are performed on the target video segment to obtain the target accident type and target liability party corresponding to the traffic accident; Based on the target accident type, the target collision type, and the target responsible party, the detection result of the traffic accident is generated and sent to the vehicle-mounted terminal. The step of performing collision detection on each video frame in the target video segment includes: From each video frame, candidate video frames containing multiple targets are identified; For any candidate video frame, target detection is performed on the candidate video frame to obtain multiple detection boxes in the candidate video frame, and based on the position information of the multiple detection boxes in the candidate video frame, it is determined whether the target in the candidate video frame has collided.
2. The method according to claim 1, wherein, The step of performing collision detection on each video frame in the target video segment to classify the collision type and obtain the target collision type corresponding to the traffic accident includes: Collision detection is performed on each video frame in the target video segment to determine a first target video frame from each video frame, wherein a collision occurs between the targets in the first target video frame; Determine the collision area where the target collided from the first target video frame; The spatial position of the collision region in the world coordinate system is determined based on the image position of the collision region in the first target video frame. Based on the spatial location of the collision region, the collision area corresponding to the collision region is determined, and based on the collision area, the target collision type is determined from multiple preset collision types.
3. The method according to claim 2, wherein, Determining the spatial position of the collision region in the world coordinate system based on the image position of the collision region in the first target video frame includes: The installation position of the image sensor on the target vehicle that collects the video to be detected is obtained; wherein, the installation position is used to indicate the position of the image sensor in the vehicle body coordinate system of the target vehicle; Obtain the intrinsic and extrinsic parameters of the image sensor; The position of the collision area in the vehicle body coordinate system is determined based on the intrinsic parameters, the extrinsic parameters, the image position, and the installation position. Based on the position of the collision area in the vehicle body coordinate system, the spatial position of the collision area in the world coordinate system is determined.
4. The method according to claim 2, wherein, Determining the collision region where the target collided from the first target video frame includes: The collision targets in the first target video frame are segmented into instances to obtain multiple first instances; The collision region is determined based on the overlapping areas of the plurality of first instances in the first target video frame.
5. The method according to claim 2, wherein, The step of determining the target collision type from multiple preset collision types based on the collision area includes: Based on the first acquisition time of the first target video frame, query the historical driving data of the target vehicle to obtain the driving position of the target vehicle at the first acquisition time; Based on the spatial location and the driving position, determine the collision location on the target vehicle corresponding to the collision area; The target collision type is determined from the plurality of preset collision types based on the collision location and the collision area.
6. The method according to claim 1, wherein, Based on multimodal information acquired simultaneously with the target video segment, the party responsible for the accident in the target video segment is detected, including: Optical flow detection is performed on the target video segment to obtain optical flow change information of the target that collided; Based on the optical flow change information of the target that collided, the motion information of the target that collided was determined, wherein the motion information was used to indicate the driving position and driving speed of the target at multiple times. Based on the motion information of the collided targets, the audio information corresponding to the target video segment, and the multimodal information, the responsible party for the collision is determined from the collided targets.
7. The method according to claim 1, wherein, The step of extracting the target video segment of the traffic accident from the video to be detected includes: Collision detection of targets is performed on each video frame in the video to be detected, so as to determine a second target video frame from each video frame; wherein, the targets in the second target video frame collide. Based on the second acquisition time of the second target video frame, a target video segment is extracted from the video to be detected.
8. The method according to claim 1, wherein, The step of determining whether a target in any candidate video frame has collided based on the position information of the plurality of detection boxes in any candidate video frame includes: Based on the position information of the multiple detection boxes in any candidate video frame, it is determined whether the multiple detection boxes overlap. If at least two of the multiple detection boxes overlap, the target within the at least two detection boxes is segmented to obtain a second instance within the at least two detection boxes; Based on the position information of at least two second instances in any candidate video frame, determine whether the at least two second instances overlap; In the case that at least two second instances overlap, it is determined that a collision occurs between the targets in any of the candidate video frames; If none of the multiple detection boxes overlap, or if at least two second instances do not overlap, it is determined that the target in any candidate video frame has not collided.
9. The method according to any one of claims 1-7, wherein, The method further includes: When the target responsible party is the target vehicle, obtain the time of occurrence of the traffic accident; Based on the multimodal information, the audio information corresponding to the target video segment, and the detection results, the driving behavior of the driver of the target vehicle at the time of occurrence is determined; Based on the driving behavior, determine the cause of the traffic accident; The cause of the occurrence is sent to the vehicle's infotainment system.
10. The method according to any one of claims 1-7, wherein, The method further includes: When the target responsible party is not the target vehicle, a driving suggestion is generated based on the multimodal information, the audio information corresponding to the target video segment, and the detection results; Send the driving suggestions to the vehicle's infotainment system; The driving advice is used to avoid traffic accidents corresponding to the target accident type and the target collision type.
11. A traffic accident detection method, applied to the vehicle-mounted terminal of a target vehicle, comprising: Send the video to be detected and the multimodal information collected from the cabin of the target vehicle to the server; Receive the detection results sent by the server; The detection result involves extracting a target video segment of a traffic accident from the video to be detected, performing collision detection on each video frame of the target video segment to classify the collision type, obtaining the target collision type corresponding to the traffic accident, and performing semantic classification and accident liability party detection on the target video segment based on multimodal information collected at the same time as the target video segment to obtain the target accident type and target liability party corresponding to the traffic accident, and generating a system based on the target accident type, the target collision type, and the target liability party. The step of performing collision detection on each video frame in the target video segment includes: From each video frame, candidate video frames containing multiple targets are identified; For any candidate video frame, target detection is performed on the candidate video frame to obtain multiple detection boxes in the candidate video frame, and based on the position information of the multiple detection boxes in the candidate video frame, it is determined whether the target in the candidate video frame has collided.
12. The method according to claim 11, wherein, Before sending the video to be tested to the server, the following steps are also included: Acquire the video stream captured by the image sensor outside the target vehicle; In response to an input operation, determine the start and end times for the truncation to be performed; The video to be detected is extracted from the video stream based on the start and end capture times.
13. A traffic accident detection device, applied to a server, comprising: The receiving module is used to receive the video to be detected sent by the vehicle's in-vehicle terminal and the multimodal information collected from the cabin of the target vehicle. The interception module is used to intercept target video segments of traffic accidents from the video to be detected; The classification module is used to perform collision detection on each video frame in the target video segment to classify the collision type and obtain the target collision type corresponding to the traffic accident. The collision detection on each video frame in the target video segment includes: determining candidate video frames containing multiple targets from each video frame; performing target detection on any candidate video frame to obtain multiple detection boxes in the candidate video frame; and determining whether a collision has occurred in the candidate video frame based on the position information of the multiple detection boxes in the candidate video frame. The processing module is used to perform semantic classification and accident liability detection on the target video segment based on multimodal information collected at the same time as the target video segment, so as to obtain the target accident type and target liability party corresponding to the traffic accident; The first generation module is used to generate the detection results of the traffic accident based on the target accident type, the target collision type, and the target responsible party; The sending module is used to send the detection results to the vehicle-mounted terminal.
14. The apparatus according to claim 13, wherein, The classification module is used for: Collision detection is performed on each video frame in the target video segment to determine a first target video frame from each video frame, wherein a collision occurs between the targets in the first target video frame; Determine the collision area where the target collided from the first target video frame; The spatial position of the collision region in the world coordinate system is determined based on the image position of the collision region in the first target video frame. Based on the spatial location of the collision region, the collision area corresponding to the collision region is determined, and based on the collision area, the target collision type is determined from multiple preset collision types.
15. The apparatus according to claim 14, wherein, The classification module is used for: The installation position of the image sensor on the target vehicle that collects the video to be detected is obtained; wherein, the installation position is used to indicate the position of the image sensor in the vehicle body coordinate system of the target vehicle; Obtain the intrinsic and extrinsic parameters of the image sensor; The position of the collision area in the vehicle body coordinate system is determined based on the intrinsic parameters, the extrinsic parameters, the image position, and the installation position. Based on the position of the collision area in the vehicle body coordinate system, the spatial position of the collision area in the world coordinate system is determined.
16. The apparatus according to claim 14, wherein, The classification module is used for: The collision targets in the first target video frame are segmented into instances to obtain multiple first instances; The collision region is determined based on the overlapping areas of the plurality of first instances in the first target video frame.
17. The apparatus according to claim 14, wherein, The classification module is used for: Based on the first acquisition time of the first target video frame, query the historical driving data of the target vehicle to obtain the driving position of the target vehicle at the first acquisition time; Based on the spatial location and the driving position, determine the collision location on the target vehicle corresponding to the collision area; The target collision type is determined from the plurality of preset collision types based on the collision location and the collision area.
18. The apparatus according to claim 13, wherein, The processing module is used for: Optical flow detection is performed on the target video segment to obtain optical flow change information of the target that collided; Based on the optical flow change information of the target that collided, the motion information of the target that collided was determined, wherein the motion information was used to indicate the driving position and driving speed of the target at multiple times. Based on the motion information of the collided targets, the audio information corresponding to the target video segment, and the multimodal information, the responsible party for the collision is determined from the collided targets.
19. The apparatus according to claim 13, wherein, The interception module is used for: Collision detection of targets is performed on each video frame in the video to be detected, so as to determine a second target video frame from each video frame; wherein, the targets in the second target video frame collide. Based on the second acquisition time of the second target video frame, a target video segment is extracted from the video to be detected.
20. The apparatus according to claim 13, wherein, The interception module or the classification module is used for: Based on the position information of the multiple detection boxes in any candidate video frame, it is determined whether the multiple detection boxes overlap. If at least two of the multiple detection boxes overlap, the target within the at least two detection boxes is segmented to obtain a second instance within the at least two detection boxes; Based on the position information of at least two second instances in any candidate video frame, determine whether the at least two second instances overlap; In the case that at least two second instances overlap, it is determined that a collision occurs between the targets in any of the candidate video frames; If none of the multiple detection boxes overlap, or if at least two second instances do not overlap, it is determined that the target in any candidate video frame has not collided.
21. The apparatus according to any one of claims 13-19, wherein, The device further includes: The acquisition module is used to acquire the time of occurrence of the traffic accident when the target responsible party is the target vehicle; The first determining module is used to determine the driving behavior of the driver of the target vehicle at the time of occurrence based on multimodal information collected at the same time as the target video segment, audio information corresponding to the target video segment, and the detection result. The second determining module is used to determine the cause of the traffic accident based on the driving behavior. The sending module is also used to send the cause of occurrence to the vehicle-mounted terminal.
22. The apparatus according to any one of claims 13-19, wherein, The device further includes: The second generation module is used to generate driving suggestions based on the multimodal information, the audio information corresponding to the target video segment, and the detection results when the target responsible party is not the target vehicle. The sending module is also used to send the driving suggestions to the vehicle's infotainment system. The driving advice is used to avoid traffic accidents corresponding to the target accident type and the target collision type.
23. A traffic accident detection device, applied to the vehicle-mounted terminal of a target vehicle, comprising: The sending module is used to send the video to be detected and the multimodal information collected from the cabin of the target vehicle to the server. A receiving module is used to receive the detection results sent by the server; The detection result involves extracting the target video segment of the traffic accident from the video to be detected, performing collision detection on each video frame of the target video segment to classify the collision type, obtaining the target collision type corresponding to the traffic accident, and performing semantic classification and accident liability party detection on the target video segment based on multimodal information collected simultaneously with the target video segment to obtain the target accident type and target liability party corresponding to the traffic accident, and generating a system based on the target accident type, the target collision type, and the target liability party. The step of performing collision detection on each video frame in the target video segment includes: From each video frame, candidate video frames containing multiple targets are identified; For any candidate video frame, target detection is performed on the candidate video frame to obtain multiple detection boxes in the candidate video frame, and based on the position information of the multiple detection boxes in the candidate video frame, it is determined whether the target in the candidate video frame has collided.
24. The apparatus according to claim 23, wherein, The device further includes: The acquisition module is used to acquire the video stream collected by the image sensor outside the target vehicle; The determination module is used to determine the start and end times of the truncation in response to the input operation. The extraction module is used to extract the video to be detected from the video stream according to the start extraction time and the end extraction time.
25. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the traffic accident detection method according to any one of claims 1-10, or to perform the traffic accident detection method according to claim 11 or 12.
26. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the traffic accident detection method according to any one of claims 1-10, or to execute the traffic accident detection method according to claim 11 or 12.
27. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the traffic accident detection method according to any one of claims 1-10, or, when executed, implements the steps of the traffic accident detection method according to claim 11 or 12.
Citation Information
Patent Citations
Vehicle collision handling method and device, HUD equipment and storage medium
CN110217187A