Artificial Intelligence-Based Intelligent Identification Method and System for Tunnel Traffic Incidents

By constructing a multi-source spatiotemporal response tensor and utilizing a cross-source temporal attention structure, the problem of truncated cross-sensor temporal response chains in tunnel traffic event discrimination models is solved, achieving high accuracy discrimination and improved stability of tunnel traffic events.

CN122416745APending Publication Date: 2026-07-17ZHONGNAN TRANSPORT

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHONGNAN TRANSPORT
Filing Date
2026-06-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing tunnel traffic event discrimination models suffer from reduced accuracy because the branch-independent extraction structure leads to the truncation of the cross-sensor temporal response chain, making it impossible to effectively distinguish between fixed visual interference in tunnels and real events.

Method used

A cross-source temporal attention structure is adopted. By constructing a multi-source spatiotemporal response tensor, attention weights between different sensor channels are learned, and the cross-source response propagation aggregation degree and propagation path continuity integrity are calculated. Event discrimination is performed by combining the traffic flow response characteristics of the tunnel entrance detection section.

Benefits of technology

It significantly improved the accuracy of identifying tunnel traffic events, reduced the false alarm rate at tunnel entrance and exit sections, and enhanced the model's stability in different tunnel sections and event types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416745A_ABST
    Figure CN122416745A_ABST
Patent Text Reader

Abstract

This invention relates to the field of tunnel traffic monitoring technology, specifically to an intelligent method and system for identifying tunnel traffic events based on artificial intelligence. The method includes: acquiring multi-source sensor data; determining a response collection window in response to a video detection model outputting candidate traffic events in any detection segment; extracting response intensity within the response collection window and constructing a multi-source spatiotemporal response tensor; inputting the multi-source spatiotemporal response tensor into a discrimination model, and outputting an event discrimination conclusion through the discrimination model to complete the discrimination of the traffic event. This invention can reduce the false alarm rate at tunnel entrance and exit sections and improve the overall accuracy of tunnel traffic event discrimination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tunnel traffic monitoring technology, specifically to an intelligent method and system for identifying tunnel traffic incidents based on artificial intelligence. Background Technology

[0002] Current methods for identifying traffic incidents in tunnels primarily employ deep learning models to detect events in a single video stream. While these models can identify anomalies such as parked vehicles, vehicles driving in the wrong direction, smoke, and pedestrians on open roads, when deployed in tunnel scenarios, they encounter three types of fixed visual interferences: drastic changes in lighting at entrances and exits, reflections from vehicle headlights on curved walls, and obstruction by large vehicles. These interferences closely resemble the visual characteristics of real events in a single-channel video stream, making it difficult for the model to distinguish them based solely on single-channel visual information, resulting in a significant drop in accuracy.

[0003] There are two main existing improvement methods: one is to re-label the data and fine-tune the model in the tunnel scenario. However, the number of interference variants under different tunnel orientations, different lighting conditions at different times, and different traffic flow densities is enormous, making it difficult to cover all interference patterns with fine-tuning. The second is to introduce auxiliary sensors such as radar and traffic flow detection for multi-source fusion. Current multi-source fusion models process the data from each sensor through independent feature extraction branches and then stitch them together at a deep feature layer. This branch-independent extraction structure causes the core feature of tunnel traffic events, namely the cross-sensor temporal response chain formed by the same physical event between video, radar, and traffic flow detection, to be truncated when each branch is convolved independently. Taking a collision event as an example, there are definite time differences and causal directions between the four stages: the sudden change in video motion state followed by a synchronous change in radar trajectory, then an increase in traffic flow cross-section occupancy, and finally a decrease in speed in the upstream section forming a queue. However, branch-independent extraction compresses the temporal changes of each sensor into local features within the branch, and the time differences and causal directions between different branches are irreversibly lost in their respective feature vectors.

[0004] Some models concatenate multi-source data into multi-channel tensors at the input end according to time steps. However, the concatenated multi-channels are treated as equivalent dimensions and convolved together by the model. The model does not endow the temporal response relationship between cross-sensor channels with an independent learning structure that is different from the temporal changes within the channel, making it difficult for the model to stably capture the cross-sensor temporal response chain of tunnel traffic events.

[0005] Therefore, due to the structure of independent branch extraction, the existing tunnel multi-source fusion model truncates the cross-sensor temporal response chain of tunnel traffic events in the feature extraction stage. This makes it impossible for the model to learn the cross-sensor temporal transmission pattern unique to real events from multi-source data, and it is difficult to effectively separate tunnel fixed visual interference from real events in the model's feature space. Summary of the Invention

[0006] This invention provides an intelligent identification method and system for tunnel traffic incidents based on artificial intelligence, in order to solve existing problems.

[0007] The intelligent identification method for tunnel traffic incidents based on artificial intelligence of the present invention adopts the following technical solution: One embodiment of the present invention provides an intelligent method for identifying tunnel traffic incidents based on artificial intelligence, the method comprising the following steps: Acquire multi-source sensor data from multiple detection sections within the tunnel; In response to the video detection model outputting candidate traffic events in any detection segment, the response collection window is determined based on the physical transmission speed corresponding to the type of the candidate traffic event, with the detection segment as the central segment and the triggering time of the candidate traffic event as the starting point. Based on multi-source sensor data, the response intensity of each detection segment, each sensor type, and each sampling time is extracted within the response collection window. A multi-source spatiotemporal response tensor is constructed with the detection segment as the first dimension, the sensor type as the second dimension, and the sampling time as the third dimension. The response intensity is used to characterize the degree of anomaly perceived by each type of sensor in each detection segment at each sampling time for the candidate traffic event. The multi-source spatiotemporal response tensor is input into the discrimination model, and the discrimination model outputs the event discrimination conclusion to complete the discrimination of traffic events; The discrimination model includes a cross-source temporal attention structure, which is used to perform attention calculation on the time response sequences of different channels with sensor type as the grouping dimension, obtain attention weights, calculate the cross-source response transmission aggregation degree based on the attention weights, and then calculate the transmission path continuity integrity. Combined with the traffic flow response characteristics of the tunnel entrance detection section, the event discrimination confidence is obtained, and the event discrimination conclusion is output accordingly.

[0008] Furthermore, taking the detection section as the central section and the triggering time of the candidate traffic event as the starting point, the response collection window is determined based on the physical transmission velocity corresponding to the type of the candidate traffic event, specifically including: When the candidate traffic event is a collision event or a parking event, the response collection window is determined according to the speed at which the traffic flow shock wave propagates upstream. The range of the response collection window covers the central section and its upstream adjacent detection section. When the candidate traffic event is a fire event or a smoke event, the response collection window is determined according to the speed at which the smoke spreads longitudinally along the tunnel. The scope of the response collection window covers the central section and its downstream adjacent detection section.

[0009] Furthermore, the sensor types include video detectors, radar target detectors, traffic flow detection sections, visibility meters, carbon monoxide sensors, and anemometers; Based on multi-source sensor data, the response intensity of each detection segment, each sensor type, and each sampling time is extracted within the response collection window, specifically including: For video detectors, the confidence level of candidate events at each sampling time point output by the video detection model is used as the response strength. For radar target detectors, the probability of target presence and the degree of abnormality of target speed or direction of motion at each sampling time detected by the radar are obtained, and the product of the probability of target presence and the degree of abnormality is used as the response intensity. For traffic flow detection sections, the time occupancy rate at each sampling time is obtained, the change in time occupancy rate relative to the historical normal value is calculated, and the normalized change is used as the response intensity. For the visibility meter, the visibility value at each sampling time is acquired, the decrease in visibility value relative to the normal reference value is calculated, and the decrease is normalized and used as the response intensity. For a carbon monoxide sensor, the carbon monoxide concentration value at each sampling time is acquired, the increase of the carbon monoxide concentration value relative to the normal reference value is calculated, and the increase is normalized and used as the response intensity. For the anemometer, the longitudinal wind speed and wind direction are obtained at each sampling time. When the wind direction is downstream along the tunnel longitudinal direction, the change in the longitudinal wind speed value relative to the normal reference value is normalized and used as the response intensity; otherwise, the response intensity is zero. For any sensor, if it does not respond to a candidate traffic event at a sampling time, the response intensity at that sampling time is set to zero.

[0010] Furthermore, a multi-source spatiotemporal response tensor is constructed using the detection segment as the first dimension, the sensor type as the second dimension, and the sampling time as the third dimension, specifically including: The detection segments of the response collection window are arranged from near to far from the central segment, and used as the first dimension of the multi-source spatiotemporal response tensor. The sensor types are arranged in the order of video detector, radar target detector, traffic flow detection section, visibility meter, carbon monoxide sensor and anemometer, which is used as the second dimension of the multi-source spatiotemporal response tensor. Arrange the sampling moments within the response collection window in chronological order as the third dimension of the multi-source spatiotemporal response tensor. The response intensity value corresponding to each detection segment, each sensor type, and each sampling time is filled into the corresponding position in the multi-source spatiotemporal response tensor to obtain the multi-source spatiotemporal response tensor. The shape of the multi-source spatiotemporal response tensor is the product of the number of detection segments, the number of sensor types, and the number of sampling times.

[0011] Furthermore, attention calculations are performed on the time response sequences of different channels, grouped by sensor type, to obtain attention weights, specifically including: For each sensor type, the response intensity of that sensor type in each detection segment and at each sampling time is arranged in chronological order to form the time response sequence of that sensor type; Using sensor type as the grouping dimension, calculate the attention weight between the time response sequence of the first sensor type and the time response sequence of the second sensor type at each sampling time, for any two different sensor types. The attention weight is used to represent the probability that when the first sensor type generates a response at the first sampling time, the response generated by the second sensor type at the second sampling time will have a transmission relationship with the former.

[0012] Furthermore, the cross-source response propagation clustering degree is calculated based on the attention weights, specifically including: Based on the type of candidate traffic event, determine the sensor response chain that matches the type, determine each pair of adjacent sensor types in the sensor response chain as a connection, obtain at least one connection position, and obtain the preset time difference window corresponding to the connection position. For each connection point, within the preset time difference window corresponding to the connection point, the attention weights between a pair of sensor types corresponding to the connection point are iterated, and the maximum value of the attention weight is determined as the response conduction attention of the connection point. Arrange the response conduction attention of all successive positions according to the conduction direction to obtain the successive attention sequence; Calculate the arithmetic mean of the successive attention sequences as the average response transmission level at each successive position; Calculate the absolute deviation of each value in the successive attention sequence from the average response conduction level, and calculate the average of all absolute deviations as the degree of dispersion of the successive attention sequence. The average response conduction level is used as the basic quantity of conduction. The ratio of the degree of dispersion to the average response conduction level is used as the dispersion suppression ratio. The product of the average response conduction level and one minus the dispersion suppression ratio is determined as the cross-source response conduction aggregation degree.

[0013] Furthermore, the integrity of the conduction path is calculated, specifically including: Obtain the maximum and minimum values ​​in the successive attention sequence, and calculate the difference between the maximum and minimum values ​​as the extreme value span; Calculate the ratio of extreme span to average response conduction level as the relative fracture degree; The cross-source response conduction concentration is used as the basis for conduction quality, the relative fracture degree is used as the denominator correction term, and the cross-source response conduction concentration is divided by one and the sum of the relative fracture degree is used to determine the conduction path continuity integrity. Specifically, when the integrity of the transmission path connection is higher than the preset integrity threshold, it is determined that there is a cross-sensor time-series response chain; otherwise, it is determined that there is no cross-sensor time-series response chain.

[0014] Furthermore, the event discrimination confidence is obtained by combining the traffic flow response characteristics of the detection section at the tunnel entrance, specifically including: The detection segment of the output candidate traffic event is determined as the trigger detection segment; Obtain the location marker of the trigger detection section. If the trigger detection section is located at the tunnel entrance or tunnel exit, the location marker is set to the first value; otherwise, it is set to the second value. Obtain the maximum and minimum values ​​of the time occupancy rate of the traffic flow detection section corresponding to the triggered detection section within the response collection window, and use the difference between the maximum and minimum values ​​as the change range of the time occupancy rate; Obtain the standard deviation of the time occupancy fluctuation of the triggered detection section under historical normal traffic conditions; The threshold for the magnitude of change in occupancy is determined based on the type of candidate traffic event. Specifically, when the type of candidate traffic event is a collision event or a parking event, the threshold for the magnitude of change in occupancy is equal to three times the standard deviation. When the type of candidate traffic event is a fire event or a smoke event, the threshold for the magnitude of change in occupancy is set to zero. When the location is marked as the first value and the change in time occupancy is less than the threshold of the change in occupancy, the product of the integrity of the transmission path and the attenuation factor is determined as the event discrimination confidence. The attenuation factor is negatively correlated with the degree of insufficiency of the change in time occupancy. The degree of insufficiency of the change in time occupancy is the ratio of the difference between the threshold of the change in occupancy and the change in time occupancy to the threshold of the change in occupancy. When the location is marked as the second value, or when the change in time occupancy reaches or exceeds the threshold of the change in occupancy, the integrity of the transmission path is used as the confidence level for event discrimination.

[0015] Furthermore, the discriminant model outputs event discrimination conclusions, specifically including: When the event discrimination confidence level reaches or exceeds the preset high confidence threshold, the candidate traffic event is determined to be a real traffic event, and the event type, trigger detection section, and event discrimination confidence level are output. When the confidence level of an event is lower than the preset low confidence threshold, the candidate traffic event is determined to be a false alarm caused by tunnel interference, and the output of the candidate traffic event is suppressed. When the event discrimination confidence level is between the low confidence threshold and the high confidence threshold, the candidate traffic event is marked as an observation state. After extending the response collection window, the construction of the multi-source spatiotemporal response tensor and the discrimination model are re-executed.

[0016] This invention proposes an intelligent identification system for tunnel traffic incidents based on artificial intelligence, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of an intelligent identification method for tunnel traffic incidents based on artificial intelligence.

[0017] The beneficial effects of the technical solution of the present invention are: This invention organizes the response intensity of each detection section, sensor type, and sampling time into a multi-source spatiotemporal response tensor that retains the sensor type dimension. It then utilizes a cross-source temporal attention structure grouped by sensor type to directly learn the attention weights of different sensor channels at different sampling times. This transforms the inherent cross-sensor temporal response chain of tunnel traffic events (such as the definite time difference and causal direction between sudden changes in video motion, radar trajectory, traffic flow occupancy increase, and upstream speed decrease in collision events) from implicit information truncated by independent branch extraction into structured features that the model can directly learn. Based on this, the completeness of the response chain is quantified by calculating the cross-source response transmission aggregation degree and the transmission path continuity completeness. Combined with the traffic flow response characteristics of the tunnel entrance section for scene constraints, the model can effectively distinguish between the coordinated cross-sensor response of real events and the isolated response of fixed visual interference in the tunnel. Simultaneously, by continuously accumulating response chain identification records and periodically recalibrating attention weights and section thresholds, the model's discrimination stability under different tunnel sections and event types is continuously improved, thereby significantly reducing the false alarm rate at the tunnel entrance and exit sections and improving the overall discrimination accuracy of tunnel traffic events. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating an artificial intelligence-based intelligent identification method for tunnel traffic incidents provided in an embodiment of the present invention; Figure 2 This is a structural diagram of an artificial intelligence-based intelligent judgment system for tunnel traffic incidents provided in an embodiment of the present invention. Detailed Implementation

[0020] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the intelligent identification method for tunnel traffic incidents based on artificial intelligence proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0022] The specific solution of the intelligent tunnel traffic incident identification method based on artificial intelligence provided by the present invention will be described in detail below with reference to the accompanying drawings.

[0023] This invention provides an intelligent method and system for identifying tunnel traffic incidents based on artificial intelligence. Please refer to [link / reference]. Figure 1 The diagram illustrates a flowchart of an artificial intelligence-based intelligent identification method for tunnel traffic incidents according to an embodiment of the present invention. The method includes the following steps: S101. Acquire multi-source sensor data from multiple detection sections within the tunnel.

[0024] In this embodiment, the tunnel is divided into multiple continuous detection sections at fixed intervals. Each detection section is equipped with a video detector, a radar target detector, a visibility meter, a carbon monoxide sensor, an anemometer, and a traffic flow detection section.

[0025] The video detector is equipped with an artificial intelligence video detection model, also known as a video detection model. This AI video detection model continuously performs target detection and motion analysis on traffic scenes within the detection area, and outputs candidate traffic events such as parking, wrong-way driving, smoke, and pedestrians, as well as the triggering section and triggering time of the candidate traffic event.

[0026] The radar target detector continuously outputs trajectory information of multiple targets within the detection zone, including the target's position, velocity, and direction of motion.

[0027] The visibility meter and carbon monoxide sensor output visibility and carbon monoxide concentration values ​​within the detection section at fixed intervals.

[0028] The anemometer outputs the longitudinal wind speed and direction within the detection section.

[0029] Traffic flow detection sections output lane speed, time occupancy, and flow rate of the corresponding sections of the detection zone at fixed intervals.

[0030] The data from each sensor are continuously synchronized and aligned according to the detection section number and timestamp to form multi-source sensor data.

[0031] S102. In response to the video detection model outputting candidate traffic events in any detection segment, the response collection window is determined based on the physical transmission speed corresponding to the type of the candidate traffic event, with the detection segment as the central segment and the triggering time of the candidate traffic event as the starting point.

[0032] In this embodiment, the detection section is taken as the central section, and the triggering time of the candidate traffic event is taken as the starting point. The response collection window is determined according to the physical transmission velocity corresponding to the type of the candidate traffic event, specifically including: When the candidate traffic event is a collision event or a parking event, the response collection window is determined according to the speed at which the traffic flow shock wave propagates upstream. The range of the response collection window covers the central section and its upstream adjacent detection section. When the candidate traffic event is a fire event or a smoke event, the response collection window is determined according to the speed at which the smoke spreads longitudinally along the tunnel. The scope of the response collection window covers the central section and its downstream adjacent detection section.

[0033] For example, the trigger condition for the response collection window is set to the output of candidate traffic events by the AI ​​video detection model in any detection segment, rather than by radar target detectors, traffic flow detection sections, or other sensors, for the following reasons: The core challenge in identifying tunnel traffic incidents lies in three types of fixed visual interference: drastic changes in lighting at the entrance and exit sections, reflections from vehicle headlights on curved walls, and obstruction by large vehicles. These interferences closely resemble the visual characteristics of real events in single-channel video, leading to numerous false alarms in video detection models. However, the fundamental difference between real events and tunnel interference is that real events form a cross-sensor temporal response chain among sensors such as video, radar, and traffic flow, while tunnel interference only manifests as an isolated response in the video channel and does not produce subsequent responses in other sensors such as radar and traffic flow that conform to the laws of physical transmission.

[0034] Therefore, using a video detection model as a trigger condition essentially means treating the model's output as a "candidate event to be verified," and then performing cross-source temporal verification of this candidate event using other sensors such as radar and traffic flow. However, using radar or traffic flow detection as the trigger condition presents the following problems: First, although radar target detectors can detect the target's position and speed, they cannot directly identify the event type (such as parking, driving in the wrong direction, smoke, pedestrians, etc.). Radar triggers make it difficult to determine the specific type of candidate event, and thus it is impossible to determine the sensor type sequence and preset time difference window expected to participate in the response chain.

[0035] Second, traffic flow detection sections reflect macroscopic changes in traffic flow status, and the response time lags behind the moment the event occurs. Traffic flow triggering will cause a serious delay in the response collection window, making it impossible to capture video and radar responses in the early stages of the event.

[0036] Third, tunnel interference exists only in the video channel. By triggering the video and introducing other sensors for verification, the fundamental difference that "real events have cross-source response chains while interference does not" can be used to suppress false alarms. If triggered by other sensors, the problem of video false alarms cannot be specifically solved.

[0037] When the AI ​​video detection model outputs a candidate traffic event in any detection segment, the system takes the detection segment triggered by the candidate traffic event as the central segment, the triggering time of the candidate traffic event as the starting point, and determines the response collection window based on the physical transmission speed corresponding to the type of candidate traffic event.

[0038] If the candidate traffic event is a collision event or a parking event, the response collection window is determined according to the speed at which the traffic flow shock wave propagates upstream, and the spatial range of the response collection window covers the central section and its upstream adjacent detection section.

[0039] It should be noted that vehicles inside the tunnel travel in a one-way direction from the entrance to the exit. The direction in which vehicles come is defined as upstream, and the direction in which vehicles go is defined as downstream.

[0040] When a collision or parking incident occurs, vehicles in front of the incident point gradually move away, while vehicles behind the incident point slow down, stop, and form a queue due to obstruction. This deceleration and queuing does not instantaneously cover the entire tunnel; instead, it propagates continuously upstream from the incident point as a traffic flow shockwave, affecting vehicles in the upstream section sequentially. Therefore, in the core sensor response chain for collision or parking incidents (video abrupt change—radar trajectory abrupt change—increased traffic flow occupancy—decreased upstream speed), the increase in traffic flow occupancy and the decrease in speed first appear in the incident section, then sequentially in adjacent upstream sections and even further upstream. Based on this physical law, this embodiment defines the spatial range of the response collection window for collision or parking incidents as covering the central section and its adjacent upstream detection sections to ensure the capture of cross-sensor responses generated during the upstream propagation of the traffic flow shockwave.

[0041] If the candidate traffic event is a fire event or a smoke event, the response collection window is determined according to the speed at which the smoke spreads longitudinally along the tunnel, and the spatial range of the response collection window covers the central section and its downstream adjacent detection section.

[0042] It should be noted that in the event of a fire or smoke-related incident, the smoke generated by the fire spreads downstream along the longitudinal airflow within the tunnel. During the smoke diffusion process, visibility in the incident section decreases first, and carbon monoxide concentration increases first. Subsequently, the smoke spreads to adjacent downstream sections, causing visibility in those downstream sections to decrease sequentially. Anemometers are used to detect longitudinal wind speed and direction to confirm the airflow conditions for the downstream smoke diffusion. Therefore, in the core sensor response chain for fire or smoke-related incidents (video smoke detection—visibility decrease—carbon monoxide concentration increase—wind speed and direction change—downstream visibility decrease), each link occurs sequentially downstream. Based on this physical law, in this embodiment, the spatial range of the response collection window for fire or smoke-related incidents is determined to cover the central section and its adjacent downstream detection sections to ensure that cross-sensor responses generated during the downstream smoke diffusion process can be captured.

[0043] S103. Based on multi-source sensor data, extract the response intensity of each detection segment, each sensor type, and each sampling time within the response collection window, and construct a multi-source spatiotemporal response tensor with the detection segment as the first dimension, the sensor type as the second dimension, and the sampling time as the third dimension. The response intensity is used to characterize the degree of anomaly perceived by each type of sensor in each detection segment at each sampling time for the candidate traffic event.

[0044] In this embodiment, the sensor types include video detectors, radar target detectors, traffic flow detection sections, visibility meters, carbon monoxide sensors, and anemometers. Based on multi-source sensor data, the response intensity of each detection segment, each sensor type, and each sampling time is extracted within the response collection window, specifically including: For video detectors, the confidence level of candidate events at each sampling time point output by the video detection model is used as the response strength. For radar target detectors, the probability of target presence and the degree of abnormality of target speed or direction of motion at each sampling time detected by the radar are obtained, and the product of the probability of target presence and the degree of abnormality is used as the response intensity. For traffic flow detection sections, the time occupancy rate at each sampling time is obtained, the change in time occupancy rate relative to the historical normal value is calculated, and the normalized change is used as the response intensity. For the visibility meter, the visibility value at each sampling time is acquired, the decrease in visibility value relative to the normal reference value is calculated, and the decrease is normalized and used as the response intensity. For a carbon monoxide sensor, the carbon monoxide concentration value at each sampling time is acquired, the increase of the carbon monoxide concentration value relative to the normal reference value is calculated, and the increase is normalized and used as the response intensity. For the anemometer, the longitudinal wind speed and wind direction are obtained at each sampling time. When the wind direction is downstream along the tunnel longitudinal direction, the change in the longitudinal wind speed value relative to the normal reference value is normalized and used as the response intensity; otherwise, the response intensity is zero. For any sensor, if it does not respond to a candidate traffic event at a sampling time, the response intensity at that sampling time is set to zero.

[0045] A multi-source spatiotemporal response tensor is constructed using the detection segment as the first dimension, the sensor type as the second dimension, and the sampling time as the third dimension. Specifically, it includes: The detection segments of the response collection window are arranged from near to far from the central segment, and used as the first dimension of the multi-source spatiotemporal response tensor. The sensor types are arranged in the order of video detector, radar target detector, traffic flow detection section, visibility meter, carbon monoxide sensor and anemometer, which is used as the second dimension of the multi-source spatiotemporal response tensor. Arrange the sampling moments within the response collection window in chronological order as the third dimension of the multi-source spatiotemporal response tensor. The response intensity value corresponding to each detection segment, each sensor type, and each sampling time is filled into the corresponding position in the multi-source spatiotemporal response tensor to obtain the multi-source spatiotemporal response tensor. The shape of the multi-source spatiotemporal response tensor is the product of the number of detection segments, the number of sensor types, and the number of sampling times.

[0046] For example, in this embodiment, the response intensity is used to characterize the degree of anomaly perceived by each type of sensor in each detection segment and at each sampling time for the candidate traffic event. The response intensity of each sensor is determined as follows: The video detector incorporates an AI-powered video detection model. This model analyzes the video image at each sampling time and outputs the confidence scores for candidate traffic events such as parked vehicles, vehicles driving against traffic, smoke, and pedestrians. This confidence score is directly used as the response strength of the video detector at the current sampling time. The confidence score ranges from 0 to 1; a higher value indicates greater confidence in the existence of a candidate traffic event. If the AI-powered video detection model does not output any candidate traffic events, the response strength is set to 0.

[0047] The radar target detector outputs trajectory information of multiple targets within the detection zone at each sampling time, including the target's position, velocity, and direction of motion. Based on the radar output data, it calculates the probability of target presence and the degree of anomaly in the target's velocity or direction of motion. The probability of target presence represents the confidence level of the radar in detecting a valid target, ranging from 0 to 1. The degree of anomaly is determined by whether the target's velocity suddenly drops or its direction of motion abruptly changes, also ranging from 0 to 1. The product of the probability of target presence and the degree of anomaly is used as the response intensity of the radar target detector at the current sampling time. If the radar does not detect any target, or if the detected target's velocity and direction of motion are both normal, the response intensity is 0.

[0048] Traffic flow detection sections output the time occupancy rate of the corresponding sections at fixed intervals. First, the normal baseline value and standard deviation of the time occupancy rate for the detection section under historical normal traffic conditions are obtained. For each sampling time, the change in the current time occupancy rate relative to the normal baseline value is calculated. This change is divided by three times the standard deviation and truncated to the interval of 0 to 1, which is used as the response intensity of the traffic flow detection section. If the time occupancy rate does not increase or the increase is less than one standard deviation, the response intensity is set to 0.

[0049] The visibility meter outputs visibility values ​​for the detection section at fixed intervals. A baseline visibility value for this detection section under normal weather conditions is pre-obtained. For each sampling time, the difference between the baseline visibility value and the current visibility value is calculated. This difference is divided by the baseline value and truncated to the range of 0 to 1, serving as the visibility meter's response intensity. This value indicates the degree of visibility decrease relative to normal conditions; a higher value indicates a more severe decrease in visibility. If visibility has not decreased, the response intensity is set to 0.

[0050] The carbon monoxide sensor outputs the carbon monoxide concentration value within the detection zone at fixed intervals. A baseline value for the carbon monoxide concentration in this detection zone under normal traffic conditions is pre-observed. For each sampling time, the difference between the current carbon monoxide concentration value and the baseline value is calculated. This difference is divided by a preset maximum allowable concentration change value, and the result is truncated to the range of 0 to 1, which is used as the response intensity of the carbon monoxide sensor. This value indicates the degree of increase in carbon monoxide concentration relative to the normal state; a higher value indicates a more severe exceedance of the limit. If the carbon monoxide concentration does not increase, the response intensity is set to 0.

[0051] The anemometer outputs longitudinal wind speed and direction within the detection section at fixed intervals. A baseline wind speed value for this detection section under normal ventilation conditions is pre-acquired. For each sampling moment, it is first determined whether the wind direction is downstream along the tunnel's longitudinal direction. If the wind direction is downstream, the change in the current longitudinal wind speed value relative to the baseline value is calculated, and this change is normalized and used as the anemometer's response intensity. If the wind direction is not downstream, the response intensity is directly set to 0. The reason for this design is that smoke from a fire will diffuse downstream with the airflow; only when the wind direction is downstream is the wind speed change relevant to the fire event.

[0052] For any sensor, if it does not respond to a candidate traffic event at a certain sampling time, the response intensity at that sampling time is set to zero.

[0053] Within the response collection window, the response intensity of each detection segment, each sensor type, and each sampling time is extracted from the multi-source sensor data, and it is constructed into a multi-source spatiotemporal response tensor with the detection segment as the first dimension, the sensor type as the second dimension, and the sampling time as the third dimension.

[0054] The first dimension of the multi-source spatiotemporal response tensor is the detection segment axis, which is arranged from near to far according to the distance between each detection segment and the central segment.

[0055] The second dimension of the multi-source spatiotemporal response tensor is the sensor type axis, which is arranged in a fixed order: video detector, radar target detector, visibility meter, carbon monoxide sensor, anemometer and traffic flow detection section.

[0056] The third dimension of the multi-source spatiotemporal response tensor is the time axis, which is arranged in chronological order of the sampling times within the response collection window.

[0057] Each position in the multi-source spatiotemporal response tensor stores the response intensity value of the corresponding sensor type in the detection segment and at the sampling time, and positions where no response is generated are set to zero.

[0058] The multi-source spatiotemporal response tensor constructed in this way retains the sensor type dimension as an independent dimension, so that the time difference between the three cross-sensor time-series events in the collision event—the video detector's motion state change at sampling time 1, the radar target detector's trajectory change at sampling time 2, and the traffic flow detection section's occupancy increase at sampling time 3—is completely preserved in the form of the position difference of the three sensor channels on the time axis.

[0059] S104. Input the multi-source spatiotemporal response tensor into the discrimination model, and output the event discrimination conclusion through the discrimination model to complete the discrimination of traffic events. The discrimination model includes a cross-source temporal attention structure, which is used to perform attention calculation on the time response sequences of different channels with sensor type as the grouping dimension to obtain attention weights. Based on the attention weights, the cross-source response transmission aggregation degree is calculated, and then the transmission path continuity integrity is calculated. Combined with the traffic flow response characteristics of the tunnel entrance detection section, the event discrimination confidence is obtained, and the event discrimination conclusion is output accordingly.

[0060] In this embodiment, attention calculation is performed on the time response sequences of different channels using sensor type as the grouping dimension to obtain attention weights, specifically including: For each sensor type, the response intensity of that sensor type in each detection segment and at each sampling time is arranged in chronological order to form the time response sequence of that sensor type; Using sensor type as the grouping dimension, calculate the attention weight between the time response sequence of the first sensor type and the time response sequence of the second sensor type at each sampling time, for any two different sensor types. The attention weight is used to represent the probability that when the first sensor type generates a response at the first sampling time, the response generated by the second sensor type at the second sampling time will have a transmission relationship with the former.

[0061] The cross-source response propagation clustering degree is calculated based on attention weights, specifically including: Based on the type of candidate traffic event, determine the sensor response chain that matches the type, determine each pair of adjacent sensor types in the sensor response chain as a connection, obtain at least one connection position, and obtain the preset time difference window corresponding to the connection position. For each connection point, within the preset time difference window corresponding to the connection point, the attention weights between a pair of sensor types corresponding to the connection point are iterated, and the maximum value of the attention weight is determined as the response conduction attention of the connection point. Arrange the response conduction attention of all successive positions according to the conduction direction to obtain the successive attention sequence; Calculate the arithmetic mean of the successive attention sequences as the average response transmission level at each successive position; Calculate the absolute deviation of each value in the successive attention sequence from the average response conduction level, and calculate the average of all absolute deviations as the degree of dispersion of the successive attention sequence. The average response conduction level is used as the basic quantity of conduction. The ratio of the degree of dispersion to the average response conduction level is used as the dispersion suppression ratio. The product of the average response conduction level and one minus the dispersion suppression ratio is determined as the cross-source response conduction aggregation degree.

[0062] Calculating the integrity of the conduction path connection includes: Obtain the maximum and minimum values ​​in the successive attention sequence, and calculate the difference between the maximum and minimum values ​​as the extreme value span; Calculate the ratio of extreme span to average response conduction level as the relative fracture degree; The cross-source response conduction concentration is used as the basis for conduction quality, the relative fracture degree is used as the denominator correction term, and the cross-source response conduction concentration is divided by one and the sum of the relative fracture degree is used to determine the conduction path continuity integrity. Specifically, when the integrity of the transmission path connection is higher than the preset integrity threshold, it is determined that there is a cross-sensor timing response chain; otherwise, it is determined that there is no cross-sensor timing response chain.

[0063] The event discrimination confidence level is obtained by combining the traffic flow response characteristics of the tunnel entrance detection section, specifically including: The detection segment of the output candidate traffic event is determined as the trigger detection segment; Obtain the location marker of the trigger detection section. If the trigger detection section is located at the tunnel entrance or tunnel exit, the location marker is set to the first value; otherwise, it is set to the second value. Obtain the maximum and minimum values ​​of the time occupancy rate of the traffic flow detection section corresponding to the triggered detection section within the response collection window, and use the difference between the maximum and minimum values ​​as the change range of the time occupancy rate; Obtain the standard deviation of the time occupancy fluctuation of the triggered detection section under historical normal traffic conditions; The threshold for the magnitude of change in occupancy is determined based on the type of candidate traffic event. Specifically, when the type of candidate traffic event is a collision event or a parking event, the threshold for the magnitude of change in occupancy is equal to three times the standard deviation. When the type of candidate traffic event is a fire event or a smoke event, the threshold for the magnitude of change in occupancy is set to zero. When the location is marked as the first value and the change in time occupancy is less than the threshold of the change in occupancy, the product of the integrity of the transmission path and the attenuation factor is determined as the event discrimination confidence. The attenuation factor is negatively correlated with the degree of insufficiency of the change in time occupancy. The degree of insufficiency of the change in time occupancy is the ratio of the difference between the threshold of the change in occupancy and the change in time occupancy to the threshold of the change in occupancy. When the location is marked as the second value, or when the change in time occupancy reaches or exceeds the threshold of the change in occupancy, the integrity of the transmission path is used as the confidence level for event discrimination.

[0064] The discriminant model outputs event discrimination conclusions, specifically including: When the event discrimination confidence level reaches or exceeds the preset high confidence threshold, the candidate traffic event is determined to be a real traffic event, and the event type, trigger detection section, and event discrimination confidence level are output. When the confidence level of an event is lower than the preset low confidence threshold, the candidate traffic event is determined to be a false alarm caused by tunnel interference, and the output of the candidate traffic event is suppressed. When the event discrimination confidence level is between the low confidence threshold and the high confidence threshold, the candidate traffic event is marked as an observation state. After extending the response collection window, the construction of the multi-source spatiotemporal response tensor and the discrimination model are re-executed.

[0065] For example, the essential difference between tunnel traffic event discrimination and open road event discrimination is that: fixed visual interference in a tunnel is highly similar to real events in a single video channel, but real events will form a cross-sensor temporal response chain between the video detector, radar target detector and traffic flow detection section, while tunnel interference will not.

[0066] The response chain of the collision event is as follows: the video motion state changes abruptly, the radar trajectory then changes abruptly in sync, which triggers an increase in the traffic flow section occupancy rate of the detection segment, and finally the speed of the upstream segment decreases and queues form.

[0067] The response chain for the fire incident is as follows: the appearance of smoke in the video triggers a decrease in visibility in the detection section, followed by an increase in carbon monoxide concentration at the sensor, a change in longitudinal wind speed detected by the anemometer, and a subsequent decrease in visibility in the downstream detection section.

[0068] The sensor type, conduction direction, and time difference range between each link in the above response chain are determined by the physical propagation characteristics of the event, and the response chain forms are different for different event types.

[0069] In existing technologies, the fundamental reason for failing to utilize this feature is not insufficient training data, but rather that when multi-source sensor data enters the model, each sensor is sent to an independent branch to extract features separately, and the temporal relationship across sensors is truncated during independent convolution of the branches.

[0070] Therefore, the core processing of this embodiment is not to extract features from each sensor branch and then design fusion rules, but to first organize multi-source spatiotemporal response tensors with sensor type as an independent dimension, and then directly learn the response transmission relationship of different sensor channels on the time axis through cross-source temporal attention structure, so that the cross-sensor temporal response chain of tunnel traffic events becomes a structured feature that the model can learn and discriminate.

[0071] In a multi-source spatiotemporal response tensor, the response relationship between different sensor channels at the same or close sampling times is fundamental to determining the existence of a cross-sensor temporal response chain. However, the response intensity dimensions of each sensor channel in the tensor are different: the video detector outputs the detection confidence level, the radar target detector outputs the target presence probability, and the traffic flow detection section outputs the speed and occupancy change magnitude. Directly comparing the response intensity of different sensor channels cannot reflect whether a transmission relationship exists between them.

[0072] In this embodiment, the role of the cross-source temporal attention structure is not to normalize the response intensity of each sensor channel before comparison, but to learn: when a sensor channel responds at a certain sampling time, at which sampling times do the responses of other sensor channels form a physical transmission relationship with it, and to assign higher attention weights to the sensor-sampling time positions that form a transmission relationship.

[0073] In the specific processing, the cross-source temporal attention structure first obtains the time response sequence of each sensor channel along the sensor type dimension. Then, using sensor type as the grouping dimension, temporal attention weights are calculated between any two different sensor channels. For the time response sequence value of the first sensor channel at the first sampling time and the time response sequence value of the second sensor channel at the second sampling time, the attention weight represents the probability that the response of the second sensor channel at the second sampling time forms a transmission relationship with the first sensor channel when the first sensor channel generates a response at the first sampling time. This attention weight is optimized during training through supervised learning with response chain integrity labels: samples carrying response chain integrity labels drive the attention weight to concentrate on the sensor-sampling time position pairs that constitute the real physical transmission relationship, while interfering samples drive the attention weight to be dispersed or concentrated only on the video detector channel.

[0074] For example, a time response sequence refers to a data sequence formed by arranging the response intensities at each sampling moment within a response collection window for a single sensor type in chronological order. If the response intensities of the video detector at the five sampling moments within the response collection window are 0.1, 0.2, 0.9, 0.3, and 0.1, then the time response sequence of the video detector is... .

[0075] The time response sequence fully preserves the response change process of this sensor type over time. The magnitude of the values ​​in the sequence represents the intensity of the response, and the order of the values ​​indicates the time sequence of the response. Through the time response sequence, the entire process of the response intensity of a certain sensor type changing over time can be observed, such as when the response intensity begins to rise, when it reaches its peak, and when it decreases.

[0076] In the multi-source spatiotemporal response tensor, each sensor type corresponds to a time response sequence. The time response sequences of different sensor types are correlated through a cross-source temporal attention structure to discover cross-sensor temporal response relationships between different sensors.

[0077] Cross-source response transmission clustering is used to represent the clustering level and balance of cross-source temporal attention weights at each successive position along the expected transmission path corresponding to the current candidate traffic event type. Its calculation is not a simple average of the response transmission attention at each successive position, but rather first extracts the response transmission attention sequence for each successive position along the expected transmission path, then uses the average of this sequence as the transmission baseline, and suppresses this baseline based on the dispersion of the sequence. The more unbalanced the succession, the stronger the dispersion suppression, and the lower the clustering. The reason for this approach is that in real collision events, the three successive positions from the video detector to the radar target detector and then to the traffic flow detection section should all have high attention and be close to each other. However, in scenarios where large vehicles obstruct the view, the attention between the video detector and the radar target detector may be high, but the attention from the radar target detector to the traffic flow detection section drops sharply to zero. Simply using the average cannot reflect the transmission interruption caused by this sudden drop between successive positions.

[0078] The specific steps for calculating the cross-source response conduction aggregation degree are as follows: The first step is to determine the sequence of sensor types expected to participate in the response chain based on the type of candidate traffic event. For collision or parking events, the expected sensor type sequence is, in order: video detector, radar target detector, and traffic flow detection section. For fire or smoke events, the expected sensor type sequence is, in order: video detector, visibility meter, carbon monoxide sensor, anemometer, and visibility meter in the downstream detection section. Each pair of adjacent sensor types in this sequence is recorded as one connection, forming multiple connection points.

[0079] The second step involves performing cross-source temporal attention computation on the multi-source spatiotemporal response tensor to obtain the attention weights. The elements in the attention weights represent the cross-source temporal attention weights between the first sensor type at the first sampling time and the second sensor type at the second sampling time.

[0080] The third step involves, for each successive location in the expected sensor type sequence, connecting two adjacent sensor types, iterating through the attention weights between the pair of sensor types corresponding to that successive location within a preset time difference window, and extracting the maximum value as the response transmission attention for that successive location. The preset time difference window is determined by the physical propagation time range of the candidate traffic event type at that successive location. For collision-type events, the preset time difference window from the video detector to the radar target detector is 0 to 3 seconds, and the preset time difference window from the radar target detector to the traffic flow detection section is 2 to 15 seconds.

[0081] The fourth step is to arrange the response conduction attention at all successive positions according to the conduction direction to obtain a successive attention sequence. The arithmetic mean of this sequence is calculated as the average response conduction level at each successive position.

[0082] The fifth step is to calculate the absolute deviation of each value in the successive attention sequence from the average response conduction level, and to calculate the average of all absolute deviations as the degree of dispersion of the successive attention sequence.

[0083] The sixth step is to use the average response conduction level as the basic quantity of conduction, the ratio of the degree of dispersion to the average response conduction level as the dispersion suppression ratio, and the product of the average response conduction level and one minus the dispersion suppression ratio is determined as the cross-source response conduction aggregation degree.

[0084] Cross-source response propagation aggregation The calculation formula can be specifically as follows: ; in, Indicates the first The cross-source response transmission clustering degree of candidate traffic events. This value is used to measure the level and balance of cross-source temporal attention weights at each successive position on the expected transmission path corresponding to the event type. The value ranges from 0 to 1. The higher the value, the stronger and more balanced the attention weights are on the expected transmission path, and the more likely it is to be a real traffic event.

[0085] Indicates the first The arithmetic mean of the successive attention sequences corresponding to each candidate traffic event. The successive attention sequence is formed by arranging the response transmission attention at each successive position in the direction of transmission, and the sequence length is equal to the number of successive positions.

[0086] Indicates the first The mean absolute deviation of the successive attention sequence corresponding to each candidate traffic event is calculated as follows: first, calculate the absolute deviation of each value in the sequence from the mean of the sequence, then sum all the absolute deviations and take the average value. This value is used to represent the degree of dispersion between the response transmission attention at each successive position and the mean. The value ranges from 0 to 1. The smaller the value, the closer the attention at each successive position is to the mean. The larger the value, the more unbalanced the attention at each successive position is.

[0087] It represents a very small positive number and is used to prevent the denominator from being 0.

[0088] Cross-source response conduction clustering reflects the overall level and balance of attention at each successive position along the conduction path. However, there is a situation where cross-source response conduction clustering cannot fully distinguish: when only one successive position has extremely low response conduction attention while the others have extremely high response conduction attention, due to the structure of the mean absolute deviation in cross-source response conduction clustering, the contribution of a single extremely low value to the mean absolute deviation is distributed among the number of successive positions, and the cross-source response conduction clustering may still remain at a moderate level.

[0089] However, for cross-sensor temporal response chains, the absence of any connecting point signifies a break in the transmission chain at that location, preventing subsequent connections from forming a complete cross-sensor transmission even with high attention levels. For example, in a collision event, the response transmission attention from the video detector to the radar target detector is 0.9, and the response transmission attention from the radar target detector to the traffic flow detection section is 0.05, with a cross-source response transmission aggregation of approximately 0.50. However, the response transmission attention of 0.05 from the radar target detector to the traffic flow detection section means that almost no transmission has occurred between the radar target detector and the traffic flow detection section, and the response chain is essentially broken.

[0090] Therefore, the integrity of the conduction path continuity is constrained by introducing the extreme span of the continuity attention sequence as a constraint, based on the cross-source response conduction aggregation. The larger the span between the strongest and weakest continuity in the sequence, the more significant the strong-weak break in the conduction path; even if the cross-source response conduction aggregation is at a moderate level, the integrity should be significantly reduced. This constraint is not multiplied in parallel with the cross-source response conduction aggregation, but rather serves as a correction to the denominator of the cross-source response conduction aggregation: the larger the span, the larger the denominator, and the greater the discount on integrity relative to the cross-source response conduction aggregation.

[0091] The transmission path continuity integrity is used to represent whether the cross-sensor temporal response chain of the current candidate traffic event is completely closed at the continuity level. Its calculation uses cross-source response transmission clustering as the basis for transmission quality and the extreme span of the continuity attention sequence as a completeness constraint, thus correcting the denominator of the cross-source response transmission clustering. This approach is based on the fact that while the cross-source response transmission clustering already incorporates the average response transmission level and dispersion, it cannot distinguish transmission breaks caused by a single extremely low value; the extreme span directly exposes the difference between the weakest and strongest continuity in the denominator. A larger span results in a larger denominator, and the completeness is compressed further, consistent with the actual physical consequences of such a continuity break.

[0092] The specific calculation steps for the integrity of the conduction path continuation are as follows: The first step is to obtain the cross-source response propagation aggregation degree of the current candidate traffic events.

[0093] The second step is to read the continuation attention sequence, extract the maximum and minimum values ​​in the sequence, and calculate the difference between the maximum and minimum values ​​as the extreme value span of the sequence.

[0094] The third step is to read the arithmetic mean of the successive attention sequences, calculate the ratio of the extreme value span to the arithmetic mean, and use it as the relative break of the sequence.

[0095] The fourth step is to use the cross-source response conduction concentration as the basis of conduction quality, the relative fracture degree as the denominator correction term, and the sum of the cross-source response conduction concentration degree and the relative fracture degree to determine the conduction path continuity integrity.

[0096] Conduction path continuity integrity The calculation formula can be: ; in, Indicates the first The transmission path continuity integrity of candidate traffic events. This value measures whether the cross-sensor temporal response chain is completely closed at the continuity level. The value ranges from 0 to 1. A higher value indicates that each continuity position on the transmission path has a high level of response transmission attention, and the response chain is completely closed; a lower value indicates that there is a lack of response transmission attention at a certain continuity position, and the response chain is broken at that point.

[0097] Indicates the first The cross-source response convergence degree of candidate traffic events. This value serves as the basis for transmission quality, reflecting the overall level and balance of attention at each successive location along the transmission path. The value ranges from 0 to 1, with higher values ​​indicating a higher and more balanced average level of attention at each successive location.

[0098] This represents the relative discontinuity of the attention sequence. It is calculated as follows: first, extract the maximum and minimum values ​​from the attention sequence; then, calculate the difference between the maximum and minimum values ​​as the extreme value span; finally, divide this extreme value span by the arithmetic mean of the attention sequence. It is used to quantify the difference between the strongest and weakest attention sequences along the transmission path relative to the average level. The value ranges from 0 to positive infinity; a smaller value indicates that the attention at each attention position is closer, while a larger value indicates a more significant difference between the strongest and weakest attention sequences.

[0099] In the formula First, it provides the overall quality foundation of the transmission path, reflecting the average level and balance of attention at each successive point. The denominator in the formula... Based on the extreme value span of the sequence Integrity discounting is applied. When the response conduction attention at each successive position is close, the maximum and minimum values ​​are close, the extreme value span approaches zero, and the relative discontinuity... As the denominator approaches zero, the fraction approaches 1. Approaching This indicates that the conduction path is both strong and complete. When an extreme break occurs in the sequence, with one high and one low value, the extreme value span is large, indicating a relatively high degree of breakage. Larger, denominator increases. The signal was significantly suppressed, accurately reflecting the fact that the conduction path was essentially incomplete due to the breakage of a single connection.

[0100] The integrity of the transmission path continuity measures whether the responses from multiple sensors have formed a complete cross-sensor transmission at the attention level. However, for tunnel entrance detection sections, there exists a boundary situation that the integrity of the transmission path continuity cannot cover: an abnormal video response triggered by a sudden change in entrance illumination may coincide in time with the natural increase in the time occupancy of the traffic flow detection section caused by a vehicle normally entering the tunnel. In this case, the transmission attention of the response from the video detector to the radar target detector is at a moderate level because the radar target detector also captures the vehicle, while the transmission attention of the response from the radar target detector to the traffic flow detection section also shows a non-zero value due to the natural increase in time occupancy. The integrity of the transmission path continuity may be pushed up to near the integrity threshold.

[0101] However, the key difference between this coincidental co-occurrence and real collision events lies in the fact that the time occupancy rate of traffic flow detection sections changes far beyond the normal traffic fluctuation range in real collision events, while the time occupancy rate changes only at the normal fluctuation level in scenarios where sudden changes in illumination coincide with normal vehicle traffic. Therefore, in the tunnel entrance detection section, it is necessary to introduce the response magnitude of traffic flow detection sections as a scenario constraint, based on the integrity of the transmission path continuity. That is, the traffic flow detection sections within the tunnel entrance detection section must exhibit strong changes exceeding normal fluctuations to support the establishment of a response chain.

[0102] This constraint does not simply multiply by a suppression coefficient between 0 and 1, but rather acts on the integrity of the transmission path continuity in an exponentially decaying manner: when the response magnitude of the traffic flow detection section is sufficient, the decay approaches zero, and the integrity of the transmission path continuity is basically preserved; when the response magnitude of the traffic flow detection section is insufficient, the decay increases exponentially with the degree of insufficiency, and the integrity of the transmission path continuity is significantly compressed. This exponential decay structure makes the constraint's sensitivity to insufficient traffic flow increase with the degree of insufficiency; the decay is mild for slight insufficiency and rapid for severe insufficiency, consistent with the practical judgment logic of the tunnel entrance detection section: "the more insufficient the traffic flow response, the greater the possibility of coincidental co-occurrence, and the stronger the exclusion should be."

[0103] Event discrimination confidence is used to represent the probability that a current candidate traffic event is a real traffic event after comprehensively considering the cross-sensor transmission integrity and tunnel scene constraints. Its calculation uses the transmission path continuity integrity as the discrimination basis, and uses the response magnitudes of the tunnel entrance detection section markers and traffic flow detection sections to form an exponential decay constraint, thus performing scene-adaptive corrections to the transmission path continuity integrity. The reason for this approach is that the tunnel entrance detection section is a high-incidence location where illumination interference and normal traffic overlap, and the transmission path continuity integrity itself cannot distinguish between complete response chains and coincidental co-occurrence; whether the time occupancy change of the traffic flow detection section exceeds the normal fluctuation range is the only quantifiable boundary that distinguishes the two.

[0104] The specific steps for calculating the confidence level of event discrimination are as follows: The first step is to obtain the continuity of the transmission path of the current candidate traffic event.

[0105] The second step is to obtain the location markers of the trigger detection sections of the candidate traffic events. If the trigger detection section is located at the tunnel entrance or exit, the tunnel entrance marker is set to the first value (e.g., 1); otherwise, the tunnel entrance marker is set to the second value (e.g., 0).

[0106] The third step is to obtain the maximum and minimum time occupancy rates of the traffic flow detection section corresponding to the triggered detection segment within the response collection window, and use the difference between the maximum and minimum values ​​as the variation range of the time occupancy rate. Simultaneously, the standard deviation of the time occupancy rate fluctuation of the triggered detection segment under historical normal traffic conditions is obtained. This standard deviation is statistically derived from the time occupancy rate data of the same period over several consecutive days in the past, and is used to represent the natural fluctuation range of the time occupancy rate under normal traffic conditions.

[0107] The fourth step is to determine the threshold for the magnitude of time occupancy change based on the type of candidate traffic event. When the candidate traffic event is a collision or parking event, the threshold for the magnitude of time occupancy change is equal to three standard deviations. This means that the time occupancy change at the traffic flow detection section must reach three times the standard deviation of normal fluctuations to exclude random fluctuations in normal traffic flow. When the candidate traffic event is a fire or smoke event, the threshold for the magnitude of time occupancy change is set to zero. In this case, the judgment is not based on the time occupancy change, and the scenario constraint step is skipped.

[0108] Fifth, when the opening is marked as the first value and the change in time occupancy is less than the threshold value for the change in time occupancy, the integrity of the transmission path continuity is multiplied by an attenuation factor to obtain the event discrimination confidence level. This attenuation factor decreases exponentially as the degree of inadequacy of the change in time occupancy increases. The degree of inadequacy is determined by the ratio of the difference between the threshold value for the change in time occupancy and the change in time occupancy to the threshold value for the change in time occupancy.

[0109] When the opening is marked as the second value, or when the change in time occupancy reaches or exceeds the threshold of the change in time occupancy, the confidence level of the event judgment is directly equal to the integrity of the transmission path continuity, and no scene constraints are imposed.

[0110] Optionally, the attenuation intensity coefficient is set according to the statistical intensity of light interference in the detection section at the tunnel entrance, and the value ranges from 0.5 to 2.0.

[0111] Event discrimination confidence level The calculation formula can be: ; in, Indicates the first The event discrimination confidence score for each candidate traffic event. This score represents the probability that a candidate traffic event is a real traffic event after comprehensively considering cross-sensor transmission integrity and tunnel scenario constraints. The value ranges from 0 to 1; a higher value indicates a greater probability that the candidate traffic event is a real event, while a lower value indicates a greater probability that the candidate traffic event is tunnel interference.

[0112] Indicates the first The completeness of the transmission path continuity for each candidate traffic event. This value serves as the basis for determining the event's confidence level and reflects the degree of complete closure of the cross-sensor temporal response chain at the continuity level. The value ranges from 0 to 1, with higher values ​​indicating a more complete response chain.

[0113] Indicates the first The marker for the tunnel entrance detection section of each candidate traffic event. This marker indicates whether the trigger detection section of the candidate traffic event is located in the tunnel entrance area. The value is 1 when the trigger detection section is located at the tunnel entrance or exit; otherwise, the value is 0. The purpose of this parameter is to apply scenario constraints only to candidate traffic events in the tunnel entrance detection section, and not to non-entrance detection sections.

[0114] This represents the attenuation intensity coefficient. This coefficient is used to control the attenuation intensity of scene constraints and is set based on the statistical intensity of light interference in the detection section at the tunnel entrance. The value ranges from 0.5 to 2.0; the larger the value, the stronger the attenuation and the more severe the penalty for insufficient traffic flow response. The specific value can be adjusted according to the actual tunnel operating environment (such as the amplitude of light variation at the tunnel entrance, traffic volume, etc.).

[0115] Indicates the first The change in time occupancy of traffic flow detection sections corresponding to the trigger detection segments of candidate traffic events within the response collection window. The calculation method is as follows: obtain the maximum and minimum values ​​of the time occupancy of traffic flow detection sections corresponding to the trigger detection segments within the response collection window, and use the difference between the maximum and minimum values ​​as the change in time occupancy. This is used to quantify the degree to which traffic flow is affected by the event; a larger value indicates a more drastic change in traffic density.

[0116] Indicates the first The threshold for the magnitude of time occupancy change corresponding to the event type of each candidate traffic event. This threshold is used to determine whether the time occupancy change at the traffic flow detection section exceeds the normal fluctuation range. When the candidate traffic event type is a collision event or a parking event, it is equal to three times the standard deviation of time occupancy fluctuation under historical normal traffic conditions. That is, the time occupancy change must reach three times the standard deviation of normal fluctuation to exclude random fluctuations in normal traffic. When the candidate traffic event type is a fire event or a smoke event, it is set to zero, indicating that the judgment is not based on the time occupancy change.

[0117] This represents an exponential function with the natural constant e as the base. This function is used to convert the decay term into a decay factor between 0 and 1. When the exponent is 0, it equals 1, and no decay occurs; when the exponent is negative, it is less than 1, and the larger the absolute value of the exponent, the smaller the decay factor.

[0118] In the formula First, a basis for judgment is provided, indicating the completeness of the cross-sensor response chain.

[0119] In the formula It is a scenario constraint used to correct candidate traffic events in the tunnel entrance detection section.

[0120] When the candidate traffic incident is not in the tunnel entrance detection section The scenario constraint term is equal to 1. No scene constraints are imposed.

[0121] When a candidate traffic event is detected in the tunnel entrance detection zone and the change in its time occupancy rate reaches or exceeds a certain threshold, , The scenario constraint term is equal to 1. No constraints are imposed.

[0122] When a candidate traffic event is detected in the tunnel entrance area and the change in its time occupancy is less than a threshold value. , >0, scenario constraint term is less than 1. .at this time The smaller the value, the greater the insufficiency, and the smaller the scenario constraint. The lower it is pressed down.

[0123] The system presets a high confidence threshold and a low confidence threshold, for example, a high confidence threshold of 0.7 and a low confidence threshold of 0.3.

[0124] when When a candidate traffic event is determined to be a real traffic event, the event type, the trigger detection section, and the event discrimination confidence level are output.

[0125] when If a candidate traffic event is determined to be a false alarm caused by tunnel interference, the output of that candidate traffic event is suppressed.

[0126] when At that time, candidate traffic events are marked as unobservable states, and the construction of multi-source spatiotemporal response tensors and the discrimination of the discrimination model are re-executed after the response collection window is extended.

[0127] It should be noted that, in this embodiment, the discrimination model may include the following functional modules: Multi-source Spatiotemporal Response Tensor Construction Module: This module extracts the response intensity of each detection segment, each sensor type, and each sampling time from multi-source sensor data within the response collection window. It then constructs a multi-source spatiotemporal response tensor with the detection segment as the first dimension, the sensor type as the second dimension, and the sampling time as the third dimension. This tensor retains the sensor type as an independent dimension, ensuring that the temporal response relationship between different sensors is fully preserved as a positional difference within the tensor.

[0128] Temporal Feature Extraction Module: This module extracts temporal features from the time response sequence of each sensor channel along the sensor type dimension. Specifically, for each sensor type, the response intensities of that sensor type in each detection segment and at each sampling time are arranged in chronological order to form the time response sequence for that sensor type. The temporal feature extraction module performs feature encoding on each time response sequence and outputs the time response sequence values ​​of each sensor channel at each sampling time for subsequent attention calculation.

[0129] Cross-source temporal attention module: This module is the core component of the discriminative model. It calculates attention weights between any two different sensor types, grouping them by sensor type. Specifically, for the time response sequence value of the first sensor type at the first sampling time and the time response sequence value of the second sensor type at the second sampling time, the cross-source temporal attention module calculates an attention weight. This weight represents the probability that the response of the second sensor type at the second sampling time will have a transmission relationship with the first sensor type when the first sensor type produces a response at the first sampling time. This attention weight is optimized during training through supervised learning using response chain integrity labels.

[0130] Response Chain Parsing Module: This module determines sensor response chains that match the type of candidate traffic event. Each pair of adjacent sensor types in the response chain is identified as a sequence, resulting in multiple sequence locations. Then, for each sequence location, within a corresponding preset time difference window, the attention weights between the pair of sensor types corresponding to that sequence location are traversed. The maximum value is extracted as the response transmission attention for that sequence location. All response transmission attentions for each sequence location are then arranged according to the transmission direction to obtain a sequence attention sequence.

[0131] Conduction Aggregation Calculation Module: This module is used to calculate the arithmetic mean of the successive attention sequence as the average response conduction level, calculate the average absolute deviation of each value in the successive attention sequence from the average response conduction level as the dispersion, then use the average response conduction level as the conduction basis, use the ratio of the dispersion to the average response conduction level as the discrete suppression ratio, and finally determine the cross-source response conduction aggregation degree by multiplying the average response conduction level by one minus the discrete suppression ratio.

[0132] The conduction integrity calculation module is used to obtain the maximum and minimum values ​​of the successive attention sequence, calculate the difference between the maximum and minimum values ​​as the extreme value span, calculate the ratio of the extreme value span to the arithmetic mean of the successive attention sequence as the relative fracture degree, then use the cross-source response conduction aggregation degree as the basis of conduction quality, use the relative fracture degree as the denominator correction term, divide the cross-source response conduction aggregation degree by one and add the sum of the relative fracture degree to determine the conduction path continuity integrity degree.

[0133] Scene Constraint Module: This module is used to impose scene constraints on candidate traffic events in the tunnel entrance detection section. Specifically, it obtains the location marker of the trigger detection section to determine whether it is located at the tunnel entrance or exit; it obtains the change in time occupancy of the traffic flow detection section corresponding to the trigger detection section within the response collection window; it obtains the standard deviation of the time occupancy fluctuation of the detection section under historical normal traffic conditions; and it determines the threshold for the magnitude of time occupancy change based on the type of candidate traffic event. When a candidate traffic event is located in the tunnel entrance detection section and the change in time occupancy is less than the magnitude threshold, the integrity of the transmission path is multiplied by a decay factor to obtain the event discrimination confidence; otherwise, the event discrimination confidence is directly equal to the integrity of the transmission path.

[0134] Discriminant Output Module: This module outputs a discrimination conclusion based on the event discrimination confidence level. When the event discrimination confidence level reaches or exceeds the high confidence threshold, the candidate traffic event is determined to be a real traffic event, and the event type, trigger detection section, and event discrimination confidence level are output. When the event discrimination confidence level is below the low confidence threshold, the candidate traffic event is determined to be a false alarm caused by tunnel interference, and the output of the candidate traffic event is suppressed. When the event discrimination confidence level is between the low confidence threshold and the high confidence threshold, the candidate traffic event is marked as an unobservable state, and the construction of the multi-source spatiotemporal response tensor and the discrimination model are re-executed after triggering an extended response collection window.

[0135] Recalibration Module: This module continuously accumulates response chain identification records from each discrimination result after model deployment and periodically performs statistical analysis on the accumulated records. When real collision events repeatedly occur in a certain detection segment but the response transmission attention at the traffic flow detection section connection position remains low, the weight coefficient of that connection link in the loss function is increased, and the sampling ratio of collision event samples in that detection segment is increased. When a sudden change in entrance illumination in a certain detection segment is repeatedly and correctly identified as interference, the attention baseline of the video channel in that detection segment is lowered, and the transmission path connection integrity threshold of that detection segment is increased accordingly. After multiple rounds of recalibration, the system solidifies the calibrated model parameters and detection segment thresholds into the deployment configuration of the current tunnel instance.

[0136] The connection relationships are as follows: The output of the multi-source spatiotemporal response tensor construction module serves as the input to the temporal feature extraction module. The output of the temporal feature extraction module serves as the input to the cross-source temporal attention module. The output of the cross-source temporal attention module serves as the input to the response chain parsing module. The output of the response chain parsing module is input to the conduction aggregation calculation module and the conduction integrity calculation module, respectively. The output of the conduction integrity calculation module is input to the scene constraint module. The output of the scene constraint module is input to the discrimination output module. The recalibration module runs independently, analyzes the recognition records accumulated by the discrimination output module, and provides feedback to adjust the parameters and thresholds of the cross-source temporal attention module and the conduction integrity calculation module.

[0137] After the discrimination model is deployed and running, it continuously accumulates response chain identification records from each discrimination result. Each record includes at least the trigger detection segment number, candidate traffic event type, transmission path continuity completeness, event discrimination confidence, response transmission attention at each connection position, and the final discrimination conclusion.

[0138] The system periodically performs statistical analysis on the accumulated records.

[0139] When real collision events repeatedly occur in a certain detection segment, but the response transmission attention at the traffic flow detection section junction in the model remains consistently lower than that at the junction from the video detector to the radar target detector, it indicates that the response chain annotation weights of the traffic flow detection link in that detection segment are insufficient in the training samples. To address this, in the next round of training, the weight coefficient of the traffic flow detection junction in the loss function is increased, and the sampling ratio of collision event samples in that detection segment in the training batch is also increased.

[0140] When a sudden change in entrance illumination in a detection segment is repeatedly and correctly identified as interference, the attention baseline of the video detector channel in that segment is lowered within the cross-source temporal attention structure, and the subsequent conduction path continuity integrity threshold for that segment is correspondingly increased. The same conduction path continuity integrity value might be sufficient to determine a response chain in non-hole detection segments, but in hole detection segments, a higher ratio of the time occupancy change magnitude to the magnitude threshold is needed to achieve a high confidence threshold for event discrimination through exponential decay constraints.

[0141] After multiple rounds of recalibration, the decision boundaries of the response chains under different detection sections and event types gradually stabilized. The false alarm rate of the model continued to decrease in the tunnel entrance and exit sections, while the detection rate remained stable in the middle section of the tunnel.

[0142] The system will solidify the calibrated model parameters and detection segment thresholds into the deployment configuration of the current tunnel instance for continuous identification of subsequent real-time traffic events.

[0143] In summary, in this embodiment of the invention, by organizing the response intensities of each detection section, each sensor type, and each sampling time within the tunnel into a multi-source spatiotemporal response tensor that retains the sensor type dimension, the temporal response relationship between different sensor channels is directly learned using a cross-source temporal attention structure with sensor type as the grouping dimension. This makes the inherent cross-sensor temporal response chains of real events such as collisions and fires into structured features that the model can directly discriminate. The completeness of the response chain is quantified by the cross-source response transmission aggregation degree and the continuity of the transmission path. Combined with the traffic flow response characteristics of the tunnel entrance section, exponential decay constraints are applied to effectively suppress false alarms caused by fixed visual interference such as sudden changes in illumination at the tunnel entrance. At the same time, by continuously accumulating response chain identification records and periodically recalibrating model parameters, the model's discrimination accuracy and stability under different tunnel sections and different event types are significantly improved.

[0144] This invention also proposes an intelligent identification system for tunnel traffic incidents based on artificial intelligence. Please refer to [link / reference]. Figure 2 The diagram shows a structural diagram of an artificial intelligence-based intelligent identification system for tunnel traffic incidents provided in an embodiment of the present invention. The system includes a data acquisition module 101 and an anomaly identification module 102.

[0145] Data acquisition module 101 is used to acquire multi-source sensor data from multiple detection sections within the tunnel; The anomaly detection module 102 is used to respond to the video detection model outputting candidate traffic events in any detection segment, taking the detection segment as the central segment and the triggering time of the candidate traffic event as the time starting point, and determining the response collection window according to the physical transmission speed corresponding to the type of the candidate traffic event; Based on multi-source sensor data, the response intensity of each detection segment, each sensor type, and each sampling time is extracted within the response collection window. A multi-source spatiotemporal response tensor is constructed with the detection segment as the first dimension, the sensor type as the second dimension, and the sampling time as the third dimension. The response intensity is used to characterize the degree of anomaly perceived by each type of sensor in each detection segment at each sampling time for the candidate traffic event. The multi-source spatiotemporal response tensor is input into the discrimination model, and the discrimination model outputs the event discrimination conclusion to complete the discrimination of traffic events; The discrimination model includes a cross-source temporal attention structure, which is used to perform attention calculation on the time response sequences of different channels with sensor type as the grouping dimension, obtain attention weights, calculate the cross-source response transmission aggregation degree based on the attention weights, and then calculate the transmission path continuity integrity. Combined with the traffic flow response characteristics of the tunnel entrance detection section, the event discrimination confidence is obtained, and the event discrimination conclusion is output accordingly.

[0146] It should be noted that the system provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the intelligent tunnel traffic incident identification system based on artificial intelligence and the intelligent tunnel traffic incident identification method based on artificial intelligence provided in the above embodiments belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0147] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0148] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0149] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An intelligent method for identifying tunnel traffic incidents based on artificial intelligence, characterized in that: include: Acquire multi-source sensor data from multiple detection sections within the tunnel; In response to the video detection model outputting candidate traffic events in any detection segment, the response collection window is determined based on the physical transmission speed corresponding to the type of the candidate traffic event, with the detection segment as the central segment and the triggering time of the candidate traffic event as the starting point. Based on multi-source sensor data, the response intensity of each detection segment, each sensor type, and each sampling time is extracted within the response collection window. A multi-source spatiotemporal response tensor is constructed with the detection segment as the first dimension, the sensor type as the second dimension, and the sampling time as the third dimension. The response intensity is used to characterize the degree of anomaly perceived by each type of sensor in each detection segment at each sampling time for the candidate traffic event. The multi-source spatiotemporal response tensor is input into the discrimination model, and the discrimination model outputs the event discrimination conclusion to complete the discrimination of traffic events; The discrimination model includes a cross-source temporal attention structure, which is used to perform attention calculation on the time response sequences of different channels with sensor type as the grouping dimension, obtain attention weights, calculate the cross-source response transmission aggregation degree based on the attention weights, and then calculate the transmission path continuity integrity. Combined with the traffic flow response characteristics of the tunnel entrance detection section, the event discrimination confidence is obtained, and the event discrimination conclusion is output accordingly.

2. The intelligent identification method for tunnel traffic incidents based on artificial intelligence according to claim 1, characterized in that, The process of determining the response collection window based on the physical transmission velocity corresponding to the type of the candidate traffic event, with the detection section as the central section and the trigger time of the candidate traffic event as the starting point, specifically includes: When the candidate traffic event is a collision event or a parking event, the response collection window is determined according to the speed at which the traffic flow shock wave propagates upstream. The range of the response collection window covers the central section and its upstream adjacent detection section. When the candidate traffic event is a fire event or a smoke event, the response collection window is determined according to the speed at which the smoke spreads longitudinally along the tunnel. The scope of the response collection window covers the central section and its downstream adjacent detection section.

3. The intelligent identification method for tunnel traffic incidents based on artificial intelligence according to claim 1, characterized in that, The sensor types include video detectors, radar target detectors, traffic flow detection sections, visibility meters, carbon monoxide sensors, and anemometers. The step of extracting the response intensity of each detection segment, each sensor type, and each sampling time within the response collection window based on multi-source sensor data specifically includes: For video detectors, the confidence level of candidate events at each sampling time point output by the video detection model is used as the response strength. For radar target detectors, the probability of target presence and the degree of abnormality of target speed or direction of motion at each sampling time detected by the radar are obtained, and the product of the probability of target presence and the degree of abnormality is used as the response intensity. For traffic flow detection sections, the time occupancy rate at each sampling time is obtained, the change in time occupancy rate relative to the historical normal value is calculated, and the normalized change is used as the response intensity. For the visibility meter, the visibility value at each sampling time is acquired, the decrease in visibility value relative to the normal reference value is calculated, and the decrease is normalized and used as the response intensity. For a carbon monoxide sensor, the carbon monoxide concentration value at each sampling time is acquired, the increase of the carbon monoxide concentration value relative to the normal reference value is calculated, and the increase is normalized and used as the response intensity. For the anemometer, the longitudinal wind speed and wind direction are obtained at each sampling time. When the wind direction is downstream along the tunnel longitudinal direction, the change in the longitudinal wind speed value relative to the normal reference value is normalized and used as the response intensity; otherwise, the response intensity is zero. For any sensor, if it does not respond to a candidate traffic event at a sampling time, the response intensity at that sampling time is set to zero.

4. The intelligent identification method for tunnel traffic incidents based on artificial intelligence according to claim 1, characterized in that, The construction of the multi-source spatiotemporal response tensor, using the detection segment as the first dimension, sensor type as the second dimension, and sampling time as the third dimension, specifically includes: The detection segments of the response collection window are arranged from near to far from the central segment, and used as the first dimension of the multi-source spatiotemporal response tensor. The sensor types are arranged in the order of video detector, radar target detector, traffic flow detection section, visibility meter, carbon monoxide sensor and anemometer, which is used as the second dimension of the multi-source spatiotemporal response tensor. Arrange the sampling moments within the response collection window in chronological order as the third dimension of the multi-source spatiotemporal response tensor. The response intensity value corresponding to each detection segment, each sensor type, and each sampling time is filled into the corresponding position in the multi-source spatiotemporal response tensor to obtain the multi-source spatiotemporal response tensor. The shape of the multi-source spatiotemporal response tensor is the product of the number of detection segments, the number of sensor types, and the number of sampling times.

5. The intelligent identification method for tunnel traffic incidents based on artificial intelligence according to claim 1, characterized in that, The attention calculation, performed on the time response sequences of different channels grouped by sensor type to obtain attention weights, specifically includes: For each sensor type, the response intensity of that sensor type in each detection segment and at each sampling time is arranged in chronological order to form the time response sequence of that sensor type; Using sensor type as the grouping dimension, calculate the attention weight between the time response sequence of the first sensor type and the time response sequence of the second sensor type at each sampling time, for any two different sensor types. The attention weight is used to represent the probability that when the first sensor type generates a response at the first sampling time, the response generated by the second sensor type at the second sampling time will have a transmission relationship with the former.

6. The intelligent identification method for tunnel traffic incidents based on artificial intelligence according to claim 1, characterized in that, The calculation of cross-source response propagation clustering based on attention weights specifically includes: Based on the type of candidate traffic event, determine the sensor response chain that matches the type, determine each pair of adjacent sensor types in the sensor response chain as a connection, obtain at least one connection position, and obtain the preset time difference window corresponding to the connection position. For each connection point, within the preset time difference window corresponding to the connection point, the attention weights between a pair of sensor types corresponding to the connection point are iterated, and the maximum value of the attention weight is determined as the response conduction attention of the connection point. Arrange the response conduction attention of all successive positions according to the conduction direction to obtain the successive attention sequence; Calculate the arithmetic mean of the successive attention sequences as the average response transmission level at each successive position; Calculate the absolute deviation of each value in the successive attention sequence from the average response conduction level, and calculate the average of all absolute deviations as the degree of dispersion of the successive attention sequence. The average response conduction level is used as the basic quantity of conduction. The ratio of the degree of dispersion to the average response conduction level is used as the dispersion suppression ratio. The product of the average response conduction level and one minus the dispersion suppression ratio is determined as the cross-source response conduction aggregation degree.

7. The intelligent identification method for tunnel traffic incidents based on artificial intelligence according to claim 6, characterized in that, The calculation of the integrity of the conduction path connection specifically includes: Obtain the maximum and minimum values ​​in the successive attention sequence, and calculate the difference between the maximum and minimum values ​​as the extreme value span; Calculate the ratio of extreme span to average response conduction level as the relative fracture degree; The cross-source response conduction concentration is used as the basis for conduction quality, the relative fracture degree is used as the denominator correction term, and the cross-source response conduction concentration is divided by one and the sum of the relative fracture degree is used to determine the conduction path continuity integrity. Specifically, when the integrity of the transmission path connection is higher than the preset integrity threshold, it is determined that there is a cross-sensor timing response chain; otherwise, it is determined that there is no cross-sensor timing response chain.

8. The intelligent identification method for tunnel traffic incidents based on artificial intelligence according to claim 6, characterized in that, The event discrimination confidence level is obtained by combining the traffic flow response characteristics of the tunnel entrance detection section, specifically including: The detection segment of the output candidate traffic event is determined as the trigger detection segment; Obtain the location marker of the trigger detection section. If the trigger detection section is located at the tunnel entrance or tunnel exit, the location marker is set to the first value; otherwise, it is set to the second value. Obtain the maximum and minimum values ​​of the time occupancy rate of the traffic flow detection section corresponding to the triggered detection section within the response collection window, and use the difference between the maximum and minimum values ​​as the change range of the time occupancy rate; Obtain the standard deviation of the time occupancy fluctuation of the triggered detection section under historical normal traffic conditions; The threshold for the magnitude of change in occupancy is determined based on the type of candidate traffic event. Specifically, when the type of candidate traffic event is a collision event or a parking event, the threshold for the magnitude of change in occupancy is equal to three times the standard deviation. When the type of candidate traffic event is a fire event or a smoke event, the threshold for the magnitude of change in occupancy is set to zero. When the location is marked as the first value and the change in time occupancy is less than the threshold of the change in occupancy, the product of the integrity of the transmission path and the attenuation factor is determined as the event discrimination confidence. The attenuation factor is negatively correlated with the degree of insufficiency of the change in time occupancy. The degree of insufficiency of the change in time occupancy is the ratio of the difference between the threshold of the change in occupancy and the change in time occupancy to the threshold of the change in occupancy. When the location is marked as the second value, or when the change in time occupancy reaches or exceeds the threshold of the change in occupancy, the integrity of the transmission path is used as the confidence level for event discrimination.

9. The intelligent identification method for tunnel traffic incidents based on artificial intelligence according to claim 1, characterized in that, The event discrimination conclusions output by the discrimination model specifically include: When the event discrimination confidence level reaches or exceeds the preset high confidence threshold, the candidate traffic event is determined to be a real traffic event, and the event type, trigger detection section, and event discrimination confidence level are output. When the confidence level of an event is lower than the preset low confidence threshold, the candidate traffic event is determined to be a false alarm caused by tunnel interference, and the output of the candidate traffic event is suppressed. When the event discrimination confidence level is between the low confidence threshold and the high confidence threshold, the candidate traffic event is marked as an observation state. After extending the response collection window, the construction of the multi-source spatiotemporal response tensor and the discrimination model are re-executed.

10. An intelligent identification system for tunnel traffic incidents based on artificial intelligence, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the steps of the intelligent identification method for tunnel traffic incidents based on artificial intelligence as described in any one of claims 1-9.