Expressway traffic incident whole-process fusion detection method and system and electronic equipment

Through real-time acquisition of video surveillance data and scene description generation of multimodal large models, combined with the correlation analysis of large language models, the full process fusion detection of highway traffic events is realized, solving the problem of multi-event elements fusion in the existing technology, and improving the accuracy and adaptability of detection.

CN120071273AActive Publication Date: 2025-05-30HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1

Patent Information

Application Number
CN202510550426.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing highway traffic event detection methods are difficult to effectively integrate and characterize multiple event elements, resulting in redundant alarms and maintenance of complex logical judgment rules, and are difficult to adapt when encountering new scenarios.

Method used

By obtaining the image data collected by the video surveillance camera on the highway in real time, using a pre-trained recognition model to detect event elements, combining multimodal large models to generate scene description information, and using a large language model to analyze the correlation between event elements and existing event tickets, and then generating event reports.

Benefits of technology

The full process fusion detection of highway traffic events has been realized, redundant alarms have been reduced, the need to maintain complex rules has been reduced, and the adaptability to complex scenarios and the integrity of event reporting has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071273A_ABST
    Figure CN120071273A_ABST
Patent Text Reader

Abstract

The invention provides an expressway traffic event whole-process fusion detection method and system and electronic equipment, and the method comprises the steps: obtaining expressway image data collected by a video monitoring camera in real time, and carrying out the event element detection of the expressway image data based on a recognition model. Outputting scene description information according to a preset cue word template on the basis of the recognized current event element by utilizing a multi-modal large model, analyzing the relevance between the current event element and an existing event ticket on the basis of the scene description information of the current event element by utilizing a preset large language model, and under the condition that the relevance exists, outputting the current event element according to the preset cue word template. And adding the current event element into an event to which the existing event ticket belongs until a preset requirement is met, and generating an event report based on the final event ticket. According to the scheme, the scene understanding ability of the multi-modal large model is utilized to generate detailed description about the event elements, then the large language model is combined to conduct event element incidence relation reasoning, and then fusion and representation of the multiple event elements can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent transportation systems. Specifically, it relates to a method, system, and electronic device for the whole-process integrated detection of highway traffic events. Background Technique

[0002] To monitor abnormal traffic events on highways, currently common methods include, for example, video-based detection techniques, indirect detection methods based on traffic operation states, multi-source information fusion methods, etc. Among them, the video-based detection method is the most direct way and conforms to the first principle. Such methods usually use a pre-trained YOLO model for object segmentation, detection, and tracking to identify event-related elements such as cones, engineering vehicles, etc., and then combine predefined rules to determine whether there is a road occupation and construction event. The YOLO model is widely used in real-time object detection services due to its fast detection speed and high accuracy. However, the YOLO model specializes in object detection, while for highway event detection, the detection of activities and behaviors is more important. For example, detecting cones and engineering vehicles does not necessarily mean there is a road occupation and construction event. It may also be that an engineering vehicle carrying cones is temporarily parked by the roadside. Similarly, when a camera detects a vehicle stopped, then detects pedestrians and warning triangles, and then the adjacent camera detects traffic congestion, the method based on the YOLO model and rule judgment often reports these as several separate events, but in fact, these events are the same event. To overcome redundant alarms, engineering technicians need to maintain a set of complex logical judgment rules. When encountering new scenarios, the general applicability and maintainability of this set of rules are very poor.

[0003] Therefore, there is an urgent need in the field of highway management and control for a method for the whole-process integrated detection of traffic events with strong scene generalization and low maintenance costs. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a method, system, and electronic device for the whole-process integrated detection of highway traffic events to solve the fusion and characterization of multiple event elements.

[0005] In the first aspect, the present invention provides a method for the whole-process integrated detection of highway traffic events, the method including: Obtain highway image data collected in real time by one or more video surveillance cameras set on the highway; Based on a pre-trained recognition model, perform event element detection on the highway image data to identify event elements related to preset events; Use a preset multi-modal large model to output scene description information according to a preset prompt word template based on the currently recognized event elements; Analyze the relevance between the current event elements and existing event tickets using a pre-set large language model based on the scenario description information of the current event elements; In the case where there is an association between the current event elements and existing event tickets, add the current event elements to the event to which the existing event tickets belong. When the preset requirements are met, generate an event report based on the final event tickets.

[0006] In an alternative embodiment, the event elements include at least one of vehicle stoppage, traffic congestion, and specific target objects; The step of detecting event elements related to preset events in the highway image data based on a pre-trained recognition model includes at least one of the following: Input the highway image data of continuously set number of frames into a pre-trained recognition model, and detect whether the position of any vehicle in the highway image data of the set number of frames has not changed in the highway image data of the set number of frames. If so, it is determined that vehicle stoppage related to the preset event is recognized; Input the highway image data into a pre-trained recognition model, and detect whether there is at least one lane with no less than a set number of vehicles having a position change amount less than a set length within a set time. If so, it is determined that traffic congestion related to the preset event is recognized; Input the highway image data of continuously set number of frames into a pre-trained recognition model, and detect whether specific target objects all appear in the highway image data of the set number of frames. If so, it is determined that specific target objects related to the preset event are recognized.

[0007] In an alternative embodiment, the step of using a pre-set multimodal large model to output scenario description information based on the recognized current event elements according to a pre-set prompt template includes: Pre-construct a traffic environment information prompt template and an event-related element information prompt template; Assemble the information of the recognized current event elements into the traffic environment information prompt template and the event-related element information prompt template to obtain the assembled prompt; Input the assembled prompt, the highway image data to which the current event elements belong, and the highway image data within the set time before and after it into a pre-set multimodal large model to output scenario description information.

[0008] In an alternative embodiment, the step of analyzing the relevance between the current event elements and existing event tickets using a pre-set large language model based on the scenario description information of the current event elements includes: Screen out target existing event tickets related to the current event elements from the existing event tickets included in the event ticket library that are not marked as in a completed state; Analyze the relevance between the current event elements and the target existing event tickets based on the scenario description information of the current event elements by means of a preset large language model.

[0009] In an alternative embodiment, the step of screening out target existing event tickets related to the current event elements from the existing event tickets included in the event ticket library that are not marked as in a completed state includes: For each existing event ticket included in the event ticket library that is not marked as in a completed state, screen out candidate existing event tickets having a spatio-temporal correlation relationship with the current event elements; Screen out target existing event tickets from the candidate existing event tickets whose similarity with the current event elements is greater than a preset threshold.

[0010] In an alternative embodiment, the step of screening out candidate existing event tickets having a spatio-temporal correlation relationship with the current event elements includes: Obtain the detection time of each existing event ticket and the location information of the video surveillance camera to which it belongs, as well as the detection time of the current event elements and the location information of the video surveillance camera to which it belongs; Screen out existing event tickets whose difference in detection time from the detection time of the current event elements is less than a preset duration and whose difference in location information from the location information of the current event elements is less than a preset distance as candidate existing event tickets.

[0011] In an alternative embodiment, the step of screening out target existing event tickets from the candidate existing event tickets whose similarity with the current event elements is greater than a preset threshold includes: Perform word vector embedding encoding on the scenario description information of each candidate existing event ticket and the current event elements respectively to obtain encoded vectors; Calculate the cosine similarity between the encoded vectors of each candidate existing event ticket and the current event elements; Screen out candidate existing event tickets whose cosine similarity is greater than a preset threshold as target existing event tickets.

[0012] In an alternative embodiment, the method further includes: In the case where the current event elements have no association relationship with the existing event tickets, create a new event for the current event elements and use the current event elements as the initial update item in this new event.

[0013] In a second aspect, the present invention provides a full-process fusion detection system for highway traffic events, and the system includes: An acquisition module, configured to acquire highway image data collected in real time by one or more video surveillance cameras set on a highway; An identification module, configured to perform event element detection on the highway image data based on a pre-trained identification model to identify event elements related to a preset event; An output module, configured to use a preset multimodal large model to output scene description information according to a preset prompt word template based on the identified current event elements; An analysis module, configured to analyze the relevance between the current event elements and existing event tickets based on the scene description information of the current event elements by using a preset large language model; A generation module, configured to, when there is an association relationship between the current event elements and existing event tickets, add the current event elements to the event to which the existing event tickets belong, and when the preset requirements are met, generate an event report based on the final event tickets.

[0014] In a third aspect, the present invention provides an electronic device, including a memory and a processor, where a computer program that can run on the processor is stored in the memory, and when the processor executes the computer program, the steps of the method according to any one of the foregoing embodiments are implemented.

[0015] The present application provides a method, a system and an electronic device for integrated detection of the whole process of highway traffic events. By acquiring highway image data collected in real time by video surveillance cameras, event element detection is performed on the highway image data based on an identification model. Using a multimodal large model to output scene description information according to a preset prompt word template based on the identified current event elements, using a preset large language model to analyze the relevance between the current event elements and existing event tickets based on the scene description information of the current event elements, and when there is an association relationship, adding the current event elements to the event to which the existing event tickets belong, and when the preset requirements are met, generating an event report based on the final event tickets. In this solution, the scene understanding ability of the multimodal large model is used to generate a detailed description of the event elements, and then the large language model is combined to perform event element association relationship reasoning, so as to solve the fusion and representation of multiple event elements.

[0016] When dealing with complex scenarios, this solution does not require manual maintenance of complex logic judgment rules, and when encountering new scenarios, there is no need to manually add new rules and handle conflicts with existing rules, improving the adaptability of highway event fusion scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0018] Figure 1 Flow chart of the whole-process integrated detection method for highway traffic incidents provided by the embodiments of the present application; Figure 2 Schematic diagram of camera layout in the embodiments of the present application; Figure 3 Schematic diagram of the monitoring screen of the camera in the embodiments of the present application; Figure 4 For Figure 1 Flow chart of the sub-steps included in S13 in Figure 5 Logical schematic diagram of the multi-modal large model generating scene description information in the embodiments of the present application; Figure 6 For Figure 1 Flow chart of the sub-steps included in S14 in Figure 7 For Figure 6 Flow chart of the sub-steps included in S141 in Figure 8 Logical schematic diagram of the large language model judging relevance in the embodiments of the present application; Figure 9 For Figure 7 Flow chart of the sub-steps included in S1411 in Figure 10 For Figure 7 Flow chart of the sub-steps included in S1412 in Figure 11 Schematic diagram of the event ticket mechanism in the embodiments of the present application; Figure 12 Another flow chart of the whole-process integrated detection method for highway traffic incidents provided by the embodiments of the present application; Figure 13 Functional module block diagram of the whole-process integrated detection system for highway traffic incidents provided by the embodiments of the present application; Figure 14 Structural block diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0019] The following will describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application.

[0020] Please refer toFigure 1 Figure 1 is a flowchart of a full - process integrated detection method for highway traffic incidents provided by an embodiment of the present invention. This full - process integrated detection method for highway traffic incidents can be executed by a full - process integrated detection system for highway traffic incidents. This full - process integrated detection system for highway traffic incidents can be implemented by software and / or hardware and can be configured in an electronic device. The electronic device can be a computer device, a server, etc. For example, it can be a server for traffic management monitoring in a traffic police platform. The detailed steps of this full - process integrated detection method for highway traffic incidents are introduced as follows.

[0021] S11, Obtain the highway image data collected in real - time by one or more video surveillance cameras set on the highway.

[0022] S12, Based on a pre - trained recognition model, perform event element detection on the highway image data to identify event elements related to preset events.

[0023] S13, Use a preset multi - modal large model to output scene description information based on the currently recognized event elements according to a preset prompt word template.

[0024] S14, Use a preset large language model to analyze the relevance between the currently recognized event elements and existing event tickets based on the scene description information of the currently recognized event elements.

[0025] S15, In the case where the currently recognized event elements have an associated relationship with existing event tickets, add the currently recognized event elements to the event to which the existing event tickets belong. When the preset requirements are met, generate an event report based on the final event tickets.

[0026] In this embodiment, multiple video surveillance cameras are set on the highway. An electronic device (such as a computer device, a server, etc. in a traffic police platform) can be communicatively connected to each video surveillance camera to receive the video images captured by each video surveillance camera.

[0027] The multiple video surveillance cameras are arranged in sequence on the highway, and each video surveillance camera faces the highway road surface to respectively capture real - time images of a section of the road surface, such as a section not less than 100 meters long, including but not limited to environmental information (such as weather), road surface, median strip, traffic signs and markings, and various vehicles.

[0028] For example, Figure 2 As shown in Figure 2, for cameras A and B, the two cameras can respectively capture real - time images of the road surface in the range from 1000 km + 460 m in the S direction to 1001 km + 520 m in the S direction on the highway. The road surface areas monitored by the two cameras are as shown in Figure 3. Figure 2 As shown.

[0029] The video images captured by the camera are as follows Figure 3 For example, real-time images of a section of road not less than 100 meters long can be captured, including environmental information such as weather, road surface, median strip, traffic signs and markings, etc., as well as various vehicles.

[0030] In this embodiment, a recognition model is pre-trained, and the recognition model can be a YOLO model. The recognition model can be obtained by continuing to train on the existing open-source models, such as the YOLOv11 model. Currently, the open-source YOLO model is pre-trained using public datasets such as COCO, Pascal VOC, Cityscapes, etc., and can recognize multiple target objects such as pedestrians, cars, trucks, motorcycles, traffic signs, animals, and daily items. In order to be able to recognize objects related to traffic events, such as event elements such as police cars, engineering vehicles, cones, warning triangles, traffic jams, spilled items, and flames, a pre-annotated private dataset is used to fine-tune the existing open-source YOLO model so that it can have the classification ability for the above-mentioned event element categories.

[0031] Specifically, in this embodiment, the configuration file of the existing YOLO model can be modified, and the nc (number of classes) in data.yml is updated to the number of new classes, and the names of new target categories are added to the names file. For example, the names of event elements such as police cars, engineering vehicles, cones, warning triangles, traffic jams, spilled items, and flames.

[0032] Start the training script, configure the image size (img), batch size (batch), number of training epochs (epochs), dataset (data), pre-trained weights (weights), etc., to continue training the YOLO model, so as to fine-tune the YOLO model to obtain the recognition model.

[0033] On this basis, the highway image data collected by the video surveillance cameras on the highway can be imported into the recognition model, and the recognition model is used to detect the event elements in the highway image data. Among them, the event elements include vehicle stoppage, traffic congestion, and specific target objects. The specific target objects include police cars, engineering vehicles, cones, warning triangles, spilled items, pedestrians, flames, etc.

[0034] After the event elements in the highway image data are detected by the recognition model, the information related to the event elements is recorded, including suspected event information, detection time, the stake number where the camera is located, etc.

[0035] In this embodiment, using the recognition model to detect event elements can be roughly divided into three major types of detections, including vehicle stoppage, traffic congestion, and detection of specific target objects. The detection methods for each major type of event element are different.

[0036] Specifically, as a possible implementation, when detecting event elements based on an identification model to identify event elements related to a preset event, it can be achieved through the following methods: Input the highway image data of a continuous set number of frames into the pre-trained identification model, and detect whether there is any vehicle whose position does not change in the highway image data of the set number of frames. If so, it is determined that a vehicle stop related to the preset event is identified.

[0037] Among them, the set number of frames can be, for example, multiple frames within 3 consecutive seconds. By using the identification model to compare the continuous multi-frame highway image data, if there is any vehicle whose position does not change in these multi-frames, it can be determined that there is an event element of vehicle stop.

[0038] In addition, in a possible implementation, when detecting event elements based on an identification model to identify event elements related to a preset event, it can be achieved through the following methods: Input the highway image data into the pre-trained identification model, and detect whether there is at least one lane with no less than a set number of vehicles whose position change amount is less than a set length within a set time. If so, it is determined that traffic congestion related to the preset event is identified; Among them, the set number can be, for example, 3 vehicles, the set time can be 3 seconds, and the set length can be a white solid line plus a set interval on the lane dividing line. For example, the set length can be 15m. That is, if there are no less than 3 vehicles on at least one lane whose position change amount is less than the set length within 3 seconds, at this time, it can be judged that the vehicle speed is less than 18 km / h. In this case, it can be determined that there is an event element of traffic congestion.

[0039] In addition, as a possible implementation, when detecting event elements based on an identification model to identify event elements related to a preset event, it can be achieved through the following methods: Input the highway image data of a continuous set number of frames into the pre-trained identification model, and detect whether there is a specific target object that appears in the highway image data of the set number of frames. If so, it is determined that the specific target object related to the preset event is identified.

[0040] As can be seen from the above, the specific target objects include police cars, engineering vehicles, cones, warning triangles, spilled items, pedestrians, flames, etc. Among them, the set number of frames can be 3 frames. That is, if any specific target object is detected in the continuous 3-frame highway image data, it can be determined that the event element of the specific target object is detected.

[0041] In this embodiment, through the above method, the recognition model obtained by continuous training can accurately detect event elements related to traffic events, such as vehicle stoppage, traffic congestion, police cars, engineering vehicles, cones, warning triangles, spilled items, pedestrians, flames, etc.

[0042] In the prior art, the D-S method of event information fusion is adopted to realize the whole-process fusion detection of traffic events. In this method, a method for event information fusion from multiple event detection information sources such as loop detectors, mobile phone reports, and mobile vehicle reports is given, which can solve the problem of the trust degree of multiple information sources for the same event element. However, this method cannot handle the internal correlation of multiple event elements that are temporally and spatially close, so as to judge whether they belong to the same event, and cannot solve the problem of redundant alarms. In addition, this method needs to calibrate the trust degree model for each scenario, and there is still the problem of poor maintainability.

[0043] In the actual scenario, there may be mutual correlations between traffic event elements detected by cameras on the highway. For example, the event elements detected by multiple cameras are the continuous distribution of the same event in time and space, and the multiple event elements detected by the same camera are different components of the same event.

[0044] Therefore, after the recognition model detects different event elements, although the multiple event elements seem to be independent, in fact, the multiple event elements may be jointly used to describe the same scenario.

[0045] Based on this, in this embodiment, a preset multimodal large model is used to output scenario description information according to a preset prompt word template based on the recognized current event elements.

[0046] Please refer to Figure 4 , in this embodiment, the multimodal large model can be Qwen QVQ-72B. The output of scenario description information based on the multimodal large model can be realized through the following methods: S131, pre-construct a traffic environment information prompt word template and an event-related element information prompt word template.

[0047] S132, assemble the information of the recognized current event elements into the traffic environment information prompt word template and the event-related element information prompt word template to obtain the assembled prompt words.

[0048] S133, input the assembled prompt words, the highway image data to which the current event elements belong, and the highway image data within a set time before and after it into a preset multimodal large model to output scenario description information.

[0049] In this embodiment, the traffic environment information prompt word template and the event-related element information prompt word template can be pre-configured.

[0050] Please refer to Figure 5 The prompt template may include information such as roles, instructions, prompts, context, questions, etc.

[0051] Among them, the traffic environment information prompt template can configure multiple description items of traffic environment, traffic facilities, and traffic order, which are used to guide the multi-modal large model to output information related to these multiple description items. For example, the traffic environment description item is used to instruct the multi-modal large model to describe the time, location, weather conditions, traffic flow density, traffic order, etc. in the picture based on highway image data. The traffic facilities description item is used to instruct the multi-modal large model to describe the highway direction, slope, lane composition, traffic signs, etc. The traffic order description item is used to instruct the multi-modal large model to describe the traffic flow density, traffic order, etc.

[0052] In addition, based on the detected event elements, the event-related element information prompt template can be configured to instruct the multi-modal large model to describe information related to the event elements. For example, describe the number of a certain event element in the picture, the spatial distribution characteristics of the event elements in the picture, the relationship with the lane lines, the dynamics of the vehicles around the event elements in the picture, the position and movement relationship between multiple event elements in the picture, etc.

[0053] Specifically, the configured traffic environment information prompt template and event-related element information prompt template can be as follows: You are an intelligent assistant good at analyzing highway traffic event scenarios, and always organize language in a concise markdown style.

[0054] First, you describe the overall traffic scenario from the following three aspects: 1. Traffic environment : Describe the time, location, weather conditions, etc., traffic flow density, traffic order, etc. in the picture; 2. Traffic facilities : Describe the highway direction, slope, lane composition, traffic signs, etc.; 3. Traffic order : Describe the traffic flow density, traffic order, etc.

[0055] Now it is suspected that a traffic event has occurred. The target detection model has discovered [detected event elements], the detection time is [detection time], and the occurrence location is [camera stake number]. Regarding this traffic event, you will select the key points related to the event from the following prompts for analysis. Note that these prompts may not all be relevant to the current event: 1. How many [detected event elements] are there in the picture; 2. Spatial distribution characteristics of [detection event elements] in the picture and their relationship with lane lines; 3. Dynamics of vehicles around [detection event elements] in the picture; 4. Positional and motion relationships among [detection event elements] (if there are multiple) in the picture.

[0056] Based on the configured traffic environment information prompt word template and event-related element information prompt word template, assemble the identified current event elements into the template, including the current event elements, suspected event information, detection time, the stake number where the camera is located, etc. into the template. In addition, extract the highway image data within the set time before and after the highway image data to which the current event elements belong, for example, the highway image data within 10 seconds before and after it. Then, input the assembled prompt words, the current highway image data, and the highway image data within the set time before and after it ( Figure 5 the event video in) into the multi-modal large model, and output the scene description information according to the prompt word template through the multi-modal large model ( Figure 5 the event scene description in).

[0057] In this embodiment, by pre-constructing the traffic environment information prompt word template and the event-related element information prompt word template, during actual implementation, only need to assemble the detected event elements into the template to quickly obtain the prompt words, which greatly improves the processing efficiency. In addition, use the multi-modal large model to combine the relatively independent event elements obtained to obtain detailed scene description information, avoid the isolation between event elements, and contribute to the full-process integration of traffic events.

[0058] Based on the obtained scene description information of the current event elements, it is necessary to determine whether the current event elements have an associated relationship with the existing event tickets. In this embodiment, use the preset large language model to implement the judgment of the associated relationship. Among them, the large language model can be Qwen2.5-72B. Specifically, please refer to Figure 6 and it can be implemented in the following ways: S141, Screen out the target existing event tickets related to the current event elements from the existing event tickets included in the event ticket library that are not marked as the completed state.

[0059] S142, Analyze the relevance between the current event elements and the target existing event tickets based on the scene description information of the current event elements by the preset large language model.

[0060] In this embodiment, the electronic device maintains an event ticket library, which includes multiple event tickets. An event ticket can be understood as an event element. Among them, if the event ticket has completed the closed-loop of the whole process fusion detection of the event or is manually marked as completed, the event ticket is marked as the completed state. The event tickets not marked as the completed state have not been added to the closed-loop whole process of the event yet.

[0061] Therefore, in this embodiment, the target existing event tickets related to the current event element are screened out from the existing event tickets not marked as the completed state. Specifically, please refer to Figure 7 and can be implemented in the following ways: S1411. For each existing event ticket included in the event ticket library that is not marked as the completed state, screen out the candidate existing event tickets that have a spatio-temporal correlation relationship with the current event element.

[0062] S1412. Screen out the target existing event tickets from the candidate existing event tickets whose similarity with the current event element is greater than a preset threshold.

[0063] As can be seen from the above, the events detected by multiple cameras may be the continuous distribution of the same event in time and space, and multiple event elements detected by the same camera may be different components of the same event. Please refer to Figure 8 together. There may be temporal and spatial associations between event elements. Existing event tickets that have a spatio-temporal correlation relationship with the current event element can be screened out as candidate existing event tickets. Specifically, please refer to Figure 9 and the candidate existing event tickets can be determined in the following ways: S14111. Obtain the detection time of each existing event ticket and the location information of the video surveillance camera to which it belongs, as well as the detection time of the current event element and the location information of the video surveillance camera to which it belongs.

[0064] S14112. Screen out the existing event tickets whose difference between the detection time and the detection time of the current event element is less than a preset duration and whose difference between the location information and the location information of the current event element is less than a preset distance as candidate existing event tickets.

[0065] After each event element is detected, the detection time related to the event element and the stake number of the video surveillance camera to which it belongs will be recorded. Based on the stake number of the video surveillance camera, the location information of the video surveillance camera can be determined.

[0066] If the distance between a video surveillance camera with an existing event ticket and the video surveillance camera of the current event element is less than a preset distance, it indicates that there is a spatial correlation between the two. If the difference between the detection time of an existing event ticket and the detection time of the current event element is less than a preset duration, it indicates that there is a temporal correlation between the two.

[0067] Based on this, candidate existing event tickets with spatio-temporal correlation with the current event element can be filtered out.

[0068] Considering that there should be similarities between event elements related to the same traffic event, therefore, based on the candidate existing event tickets, existing event tickets with a similarity greater than a preset threshold to the current event element are filtered out as target existing event tickets. Specifically, please refer to Figure 10 , this step can be implemented in the following way: S14121, perform word vector embedding encoding on the scene description information of each of the candidate existing event tickets and the current event element to obtain encoded vectors.

[0069] S14122, calculate the cosine similarity between the encoded vectors of each of the candidate existing event tickets and the current event element.

[0070] S14123, filter out candidate existing event tickets with a cosine similarity greater than the preset threshold as target existing event tickets.

[0071] In this embodiment, word vector encoding is respectively performed on the scene description information of the candidate existing event tickets and the scene description information of the current event element to obtain their encoded vectors. In this way, the cosine similarity between the two can be obtained by using the cosine similarity calculation method based on the encoded vectors. When the cosine similarity is greater than the preset threshold, it indicates that the two may point to the same traffic event. Therefore, the candidate existing event tickets in this case can be marked as target existing event tickets.

[0072] On the basis of filtering out the target existing event tickets, the large language model is used to analyze the correlation between the target existing event tickets and the current event element.

[0073] In this embodiment, similarly, a prompt template can be predefined, and the scene description information of the target existing event tickets and the current event element is assembled into the prompt template to obtain a prompt. The prompt is input into the large language model, so that the large language model analyzes the correlation between the two according to the instructions of the prompt template.

[0074] Among them, the prompt template can prompt the large language model to analyze from logical relationships such as sequence, causality, subordination, condition, similarity, and purpose between the two, and clearly judge whether the two belong to the same event based on the analysis structure.

[0075] Specifically, the pre-configured prompt template can be as follows: You are a highly logical AI assistant, proficient in analyzing the correlations behind multiple phenomena. The events we newly discovered are as follows: [Event Ticket 1]; [Event Ticket 2]; …; [Event Ticket n].

[0076] Now a new phenomenon has been discovered: [Detailed Scenario Description].

[0077] Please analyze whether there are any correlations between this phenomenon and the aforementioned events, such as logical relationships like sequence, causality, subordination, condition, similarity, purpose, etc., and based on the analysis results, clearly determine whether this phenomenon and the aforementioned events are the same event.

[0078] In this way, the large language model can be used to analyze the correlation between the current event elements and the existing event tickets based on the assembled prompt. If there is a correlation between the two, the current event elements will be added to the event to which the existing event tickets belong.

[0079] In addition, please refer to Figure 11 In this embodiment, in the case where there is no correlation between the current event elements and the existing event tickets, a new event is created for the current event elements, and the current event elements are used as the initial update item in this new event.

[0080] In addition, if the current event elements are the first event elements recognized, similarly, a new event ticket is distributed for the current event elements, and the current event elements are defined as the initial update item.

[0081] If there is a correlation between the current event elements and the existing event tickets, then the current event ticket serves as a sub-update in the event process to which the existing event tickets belong.

[0082] Among them, the initial update item information and sub-update item information are refined by the large language model, which can guide the large language model to extract materials according to the prompt template. Among them, the prompt template includes the initial update and each sub-update included in the event ticket process. In addition, it includes the newly detected current event elements, the scenario description information of the current event elements, etc. Specifically, the constructed prompt template can be as follows: You are an AI assistant good at refining key information of events. We use event tickets to organize event update information. The current event ticket information is as follows: [Initial Update]; [Sub-Update]; …; [Sub-Update].

[0083] The newly reported event elements are: [Detection event elements].

[0084] The detailed scene description information for it is: [Event scene description].

[0085] Please refine the newly reported event elements and scene description into event update information as an initial update or sub-update for the current event ticket. Please refer to the following event update information template: 1. Event type : Traffic congestion, construction, traffic accidents, fires, break-ins, etc.; 2. Urgency : Based on the event type and event impact; 3. Location of discovery : The mileage number in the detection event elements; 4. Detection time : The time in the detection event elements; 5. Event description : Describe the event scene in one sentence.

[0086] In this way, by using the large language model to perform processing in the above manner, the event elements in the same event process can be fused to complete the closed-loop of the whole-process fusion detection of the event until there is no need to continue the event fusion detection.

[0087] In summary, please refer to Figure 12 the overall logic schematic diagram of the whole-process fusion detection method provided in this embodiment shown in the figure. In this solution, the electronic device accesses each video surveillance camera to obtain the highway image data collected by the video surveillance camera. The recognition model is used to detect the event elements in the highway image data. When the corresponding event elements are detected, the multi-modal large model is used to generate the scene description information of the event elements. When there are unfinished event tickets in the event ticket library, the large language model is used to judge the relevance between the event elements and the event tickets. When there are no unfinished event tickets, or there is no association with the existing event tickets, a new event ticket is distributed for the event elements.

[0088] When there is an association between the event elements and the existing event tickets, the existing event tickets are updated, and the large language model is used to refine the event update information. Until the termination condition is triggered, the fusion process ends and the final event report is generated.

[0089] The whole-process fusion detection method provided in this embodiment uses the scene understanding ability of the multi-modal large model to generate scene description information about event elements, then uses the large language model to perform reasoning on the correlation relationships of event elements to determine whether it is a sub-update of an existing event, and finally uses the event ticket mechanism to organize and update event information. This method can solve the problems of fusion and representation of multiple event elements.

[0090] Compared with the existing event fusion method based on rule judgment, this solution uses the scene understanding of the multi-modal large model and the logical reasoning ability of the large language model based on the world model. When dealing with complex scenarios, it does not require manual maintenance of complex logical judgment rules, and when encountering new scenarios, it does not require manual addition of new rules and handling of conflicts with existing rules, improving the adaptability of highway event fusion scenarios.

[0091] In this solution, the logical reasoning ability of the large language model based on the world model is used, which can handle the mutual correlation between multiple event elements and improve the fusion ability of multiple different types of event elements.

[0092] Furthermore, this solution adopts the event ticket mechanism to organize and update a series of related event elements, improving the integrity of event reports through information integration.

[0093] Based on the same inventive concept, please refer to Figure 13 , the embodiment of the present invention also provides a schematic diagram of the functional modules of a whole-process fusion detection system for highway traffic events. This embodiment can divide the functional modules of the whole-process fusion detection system for highway traffic events according to the above method embodiment. For example, corresponding functional modules can be divided for each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present invention is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0094] For example, in the case of dividing corresponding functional modules for each function, Figure 13 The shown whole-process fusion detection system for highway traffic events is only a schematic diagram of a device. The whole-process fusion detection system for highway traffic events may include an acquisition module, an identification module, an output module, an analysis module, and a generation module. The functions of each functional module of the whole-process fusion detection system for highway traffic events will be elaborated in detail below.

[0095] The acquisition module is used to acquire the highway image data collected in real time by one or more video surveillance cameras set on the highway; An identification module, configured to perform event element detection on the highway image data based on a pre-trained identification model to identify event elements related to a preset event; An output module, configured to use a preset multimodal large model to output scene description information according to a preset prompt template based on the currently identified event elements; An analysis module, configured to use a preset large language model to analyze the relevance between the currently identified event elements and existing event tickets based on the scene description information of the currently identified event elements; A generation module, configured to, when there is an association between the currently identified event elements and existing event tickets, add the currently identified event elements to the event to which the existing event tickets belong, and when the preset requirements are met, generate an event report based on the final event tickets.

[0096] It can be understood that the above-mentioned acquisition module, identification module, output module, analysis module, and generation module can be used to execute the above S11 to S15. The detailed implementation methods of the acquisition module, identification module, output module, analysis module, and generation module can refer to the content related to the above S11 to S15.

[0097] In a possible implementation manner, the event elements include at least one of vehicle stoppage, traffic congestion, and specific target objects; the above-mentioned identification module can be used to: Input the highway image data of continuously set number of frames into a pre-trained identification model, and detect whether there is any vehicle whose position in the highway image data of the set number of frames has not changed. If so, it is determined that vehicle stoppage related to the preset event is identified; Input the highway image data into a pre-trained identification model, and detect whether there is at least one lane with no less than a set number of vehicles whose position change amount is less than a set length within a set time. If so, it is determined that traffic congestion related to the preset event is identified; Input the highway image data of continuously set number of frames into a pre-trained identification model, and detect whether specific target objects all appear in the highway image data of the set number of frames. If so, it is determined that specific target objects related to the preset event are identified.

[0098] In a possible implementation manner, the above-mentioned output module can be used to: Pre-construct a traffic environment information prompt template and an event-related element information prompt template; Assemble the information of the currently identified event elements into the traffic environment information prompt template and the event-related element information prompt template to obtain an assembled prompt; Input the assembled prompt words, the highway image data to which the current event element belongs, and the highway image data within the set time before and after it into a preset multi-modal large model to output scene description information.

[0099] In a possible implementation, the above analysis module can be used for: Screen out the target existing event tickets related to the current event element from the existing event tickets included in the event ticket library that are not marked as in the completed state; Analyze the relevance between the current event element and the target existing event tickets based on the scene description information of the current event element by a preset large language model.

[0100] In a possible implementation, the above analysis module can be used for: For each existing event ticket included in the event ticket library that is not marked as in the completed state, screen out the candidate existing event tickets that have a spatio-temporal correlation relationship with the current event element; Screen out the target existing event tickets with a similarity greater than a preset threshold between the candidate existing event tickets and the current event element.

[0101] In a possible implementation, the above analysis module can be used for: Obtain the detection time of each existing event ticket and the location information of the video surveillance camera to which it belongs, as well as the detection time of the current event element and the location information of the video surveillance camera to which it belongs; Screen out the existing event tickets with a difference in detection time less than a preset duration between the detection time of the existing event ticket and the detection time of the current event element, and a difference in location information less than a preset distance between the location information of the existing event ticket and the location information of the current event element as candidate existing event tickets.

[0102] In a possible implementation, the above analysis module can be used for: Perform word vector embedding encoding on the scene description information of each candidate existing event ticket and the current event element respectively to obtain encoding vectors; Calculate the cosine similarity between the encoding vectors of each candidate existing event ticket and the current event element; Screen out the candidate existing event tickets with a cosine similarity greater than a preset threshold as the target existing event tickets.

[0103] In a possible implementation, the highway traffic event whole-process fusion detection system further includes a creation module, and this creation module is used for: In the case that the current event element has no association relationship with the existing event ticket, create a new event for the current event element and use the current event element as the initial update item in this new event.

[0104] Please refer to Figure 14 , which is a structural block diagram of the electronic device provided by the embodiment of the present invention. The electronic device can be a computer device, a server, etc. in a traffic police platform. The electronic device includes a memory, a processor, and a communication module. Each element of the memory, the processor, and the communication module is electrically connected directly or indirectly to each other to realize data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines.

[0105] Among them, the memory is used to store computer programs or data. The memory can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0106] The processor is used to read / write the data or programs stored in the memory and execute the whole-process fusion detection method for highway traffic events provided by any embodiment of the present invention.

[0107] The communication module is used to establish a communication connection between the electronic device and other communication terminals through the network and is used to send and receive data through the network.

[0108] It should be understood that Figure 14 the structure shown is only a schematic diagram of the structure of the electronic device, and the electronic device may further include more or fewer components than those shown in Figure 14 , or have a different configuration from that shown in Figure 14 .

[0109] Furthermore, the embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed, the whole-process fusion detection method for highway traffic events provided by the above embodiment is realized.

[0110] Specifically, the computer-readable storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the computer-readable storage medium runs, it can execute the above-mentioned whole-process fusion detection method for highway traffic events. Regarding the process involved when the computer-readable storage medium and its executable instructions run, reference can be made to the relevant descriptions in the above method embodiments, which will not be elaborated here.

[0111] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.

[0112] In addition, the units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0113] Furthermore, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0114] It should be noted that if the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs and other various media that can store program codes.

[0115] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0116] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A highway traffic incident full-process fusion detection method, characterized in that: The method comprises: Acquire highway image data collected in real time by one or more video surveillance cameras installed on the highway; Performing event element detection on the highway image data based on the pre-trained recognition model to identify event elements related to a preset event; Outputting scene description information according to a preset prompt word template based on the identified current event elements using a preset multimodal large model; Analyzing the correlation between the current event element and the existing event ticket based on the scene description information of the current event element by using a preset large language model; In the case where the current event element is associated with an existing event ticket, the current event element is added to the event to which the existing event ticket belongs until the preset requirements are met, and then an event report is generated based on the final event ticket.

2. The whole-process fusion detection method for highway traffic incidents according to claim 1 is characterized in that: The event element includes at least one of vehicle stoppage, traffic congestion and a specific target object; The step of performing event element detection on the highway image data based on the pre-trained recognition model to identify event elements related to a preset event includes at least one of the following: Inputting a set number of consecutive frames of highway image data into a pre-trained recognition model, detecting whether there is any vehicle whose position has not changed in the set number of frames of highway image data, and if so, determining that a vehicle stop related to a preset event is recognized; Inputting the highway image data into a pre-trained recognition model to detect whether there is at least one lane with a set number of vehicles having a position change amount less than a set length within a set time, and if so, determining that a traffic congestion related to a preset event is recognized; A set number of consecutive frames of highway image data are input into a pre-trained recognition model to detect whether a specific target object appears in the set number of frames of highway image data. If so, the specific target object related to the preset event is determined and identified.

3. The whole-process fusion detection method for highway traffic incidents according to claim 1 is characterized in that: The step of using the preset multimodal large model to output scene description information based on the identified current event elements according to the preset prompt word template includes: Pre-build traffic environment information prompt word templates and event-related element information prompt word templates; Assembling the identified current event element information into the traffic environment information prompt word template and the event-related element information prompt word template to obtain assembled prompt words; The assembled prompt words, the highway image data to which the current event element belongs, and the highway image data within a set time before and after the current event element are input into a preset multimodal large model, and the scene description information is output.

4. The whole-process fusion detection method for highway traffic incidents according to claim 1 is characterized in that: The step of analyzing the correlation between the current event element and the existing event ticket based on the scene description information of the current event element by using the preset large language model includes: Filter out target existing event tickets related to the current event element from existing event tickets that are not marked as completed and are included in the event ticket library; Based on a preset large language model and based on the scene description information of the current event element, the correlation between the current event element and the target existing event ticket is analyzed.

5. The whole-process fusion detection method for highway traffic incidents according to claim 4 is characterized in that: The step of selecting target existing event tickets related to the current event element from existing event tickets not marked as completed and included in the event ticket library comprises: For each existing event ticket in the event ticket database that is not marked as completed, select candidate existing event tickets that have a temporal and spatial association relationship with the current event element; Filter out target existing event tickets from the candidate existing event tickets, the target existing event tickets having a similarity with the current event element greater than a preset threshold.

6. The whole-process fusion detection method for highway traffic incidents according to claim 5 is characterized in that: The step of screening out candidate existing event tickets having a spatiotemporal association relationship with the current event element comprises: Obtaining the detection time of each existing event ticket and the location information of the video surveillance camera to which it belongs, as well as the detection time of the current event element and the location information of the video surveillance camera to which it belongs; Existing event tickets whose difference between the detection time and the detection time of the current event element is less than a preset duration, and whose difference between the location information and the location information of the current event element is less than a preset distance are screened out as candidate existing event tickets.

7. The whole-process fusion detection method for highway traffic incidents according to claim 5 is characterized in that: The step of selecting target existing event tickets whose similarity with the current event element is greater than a preset threshold from the candidate existing event tickets comprises: Performing word vector embedding encoding on each of the candidate existing event tickets and the scene description information of the current event element to obtain an encoding vector; Calculating the cosine similarity between the encoding vectors of each candidate existing event ticket and the current event element; The candidate existing event tickets whose cosine similarity is greater than a preset threshold are selected as the target existing event tickets.

8. The highway traffic incident full-process fusion detection method according to claim 1 is characterized in that: The method further comprises: In the case that the current event element has no association relationship with the existing event ticket, a new event is created for the current event element, and the current event element is used as an initial update item in the new event.

9. A highway traffic incident full-process fusion detection system, characterized in that: The system comprises: An acquisition module is used to acquire highway image data collected in real time by one or more video surveillance cameras installed on the highway; A recognition module, used to perform event element detection on the highway image data based on a pre-trained recognition model to identify event elements related to a preset event; An output module, used to output scene description information according to a preset prompt word template based on the identified current event elements using a preset multimodal large model; An analysis module, configured to analyze the correlation between the current event element and the existing event ticket based on the scene description information of the current event element by using a preset large language model; A generation module is used to add the current event element to the event to which the existing event ticket belongs when the current event element has an association relationship with the existing event ticket, until the preset requirements are met, and then generate an event report based on the final event ticket.

10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Information situation mining method based on event extraction

    CN116757221A

  • Fine adjustment method of large language model, resource recommendation method, device and equipment

    CN118626717A

  • Fire fighting access occupation early warning grade analysis system and method based on large model

    CN119723469A

  • Natural driving accident scene key element extraction method based on visual large model

    CN119832478A

  • Summarizing Events Over a Time Period

    US20250111674A1

Cited By

  • Method and device for generating traffic vertical domain model and autonomous intelligent traffic system

    CN120851148A