An intelligent traffic guidance system and method based on smart signs

By deploying a lightweight dual-path network and millimeter-wave radar on intelligent signs, and combining multimodal sensor fusion verification and collaborative guidance priority calculation, the problems of event detection delay and guidance incoordination in intelligent transportation systems are solved, thereby improving the response speed and guidance effect of sudden traffic events.

CN121963491BActive Publication Date: 2026-07-17CHONGQING JIHENG LOGO CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING JIHENG LOGO CO LTD
Filing Date
2026-04-03
Publication Date
2026-07-17

Smart Images

  • Figure CN121963491B_ABST
    Figure CN121963491B_ABST
Patent Text Reader

Abstract

This application provides an intelligent traffic guidance system and method based on intelligent signage, relating to the field of intelligent transportation technology. The system includes: a first intelligent signage deployed on the roadside identifies suspected sudden traffic events from real-time road video streams via a lightweight dual-path network and generates a first confidence level; it calls millimeter-wave radar data to detect abrupt changes in the motion state of target objects in the suspected sudden traffic event and generates a second confidence level; it performs dual verification based on the first and second confidence levels to confirm the occurrence of a sudden traffic event and generates sudden event information; it broadcasts the sudden event information to other intelligent signages; each intelligent signage calculates a collaborative guidance priority factor; and each intelligent signage matches a corresponding guidance strategy according to the collaborative guidance priority factor and executes traffic guidance actions. This solves the technical problems of high detection latency and insufficient guidance in existing technologies for sudden traffic events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent transportation technology, and in particular to an intelligent traffic guidance system and method based on intelligent signage. Background Technology

[0002] With the acceleration of urbanization and the continuous growth of motor vehicle ownership, road traffic flow is becoming increasingly dense, and the frequency of traffic emergencies is also rising. If traffic emergencies are not detected and effectively guided in a timely manner, they can easily lead to secondary accidents and widespread traffic congestion, seriously threatening driving safety and road traffic efficiency.

[0003] Currently, intelligent transportation systems commonly employ video surveillance and radar detection to monitor road traffic conditions. However, existing technologies have the following shortcomings: First, video surveillance relies on centralized processing platforms, transmitting large amounts of video data to the cloud or central servers for analysis, resulting in high bandwidth pressure and processing latency, making it difficult to meet the needs of real-time event response. Second, single sensors are easily affected by environmental interference, leading to insufficient accuracy and reliability in event detection. Third, after an event occurs, there is a lack of effective collaborative guidance mechanisms; traffic guidance devices often operate independently, failing to dynamically adjust guidance strategies based on the event location and their own location, resulting in delayed guidance information or poor guidance effectiveness. Summary of the Invention

[0004] This invention addresses the technical problems in existing technologies, such as high event detection latency and difficulty in meeting real-time response requirements due to reliance on centralized video surveillance platforms, susceptibility to environmental interference and insufficient reliability due to reliance on single sensors, and independent responses and poor guidance effects of individual guidance devices due to the lack of a collaborative guidance mechanism, resulting in delayed identification of sudden traffic events and a high misjudgment rate. The invention provides an intelligent traffic guidance system and method based on intelligent signage.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0006] In a first aspect, the present invention provides an intelligent traffic guidance system based on intelligent signage, comprising:

[0007] The visual initial perception module is used to identify suspected sudden traffic events and generate a first confidence level from the collected real-time video stream of the road by the first smart sign in the smart sign array deployed on the roadside through a pre-trained lightweight dual-path network;

[0008] The radar motion detection module is used to synchronously call millimeter-wave radar data, detect the abrupt change characteristics of the motion state of the target object in the suspected sudden traffic incident, and generate a second confidence level based on the abrupt change characteristics of the motion state.

[0009] The fusion verification and confirmation module is used to perform dual verification and judgment based on the first confidence level and the second confidence level. After the dual verification is passed, it confirms that a sudden traffic incident has occurred and generates sudden incident information including the incident type, incident location and lane occupation.

[0010] An event information broadcasting module is used to broadcast the emergency event information to other smart signs in the smart sign array;

[0011] The collaborative priority calculation module is used for each smart sign in the smart sign array to calculate the collaborative guidance priority factor based on its own location information and the emergency information through a collaborative guidance priority calculation model deployed on the edge.

[0012] The differentiated guidance execution module is used for each smart sign to match the corresponding guidance strategy based on the collaborative guidance priority factor calculated by itself, and to execute differentiated traffic guidance actions.

[0013] Secondly, the present invention provides an intelligent traffic guidance method based on intelligent signage, comprising:

[0014] The first smart sign in the smart sign array deployed on the roadside identifies suspected traffic emergencies from the collected real-time video stream of the road through a pre-trained lightweight dual-path network and generates a first confidence level.

[0015] Simultaneously call millimeter-wave radar data to detect the abrupt change characteristics of the motion state of the target object in the suspected sudden traffic incident, and generate a second confidence level based on the abrupt change characteristics of the motion state;

[0016] Based on the first confidence level and the second confidence level, a dual verification judgment is performed. After the dual verification is passed, it is confirmed that a sudden traffic incident has occurred, and sudden incident information including the incident type, incident location and lane occupation is generated.

[0017] The emergency information is broadcast to all other smart signs in the smart sign array;

[0018] Each smart sign in the smart sign array calculates a collaborative guidance priority factor based on its own location information and the emergency event information through a collaborative guidance priority calculation model deployed on the edge.

[0019] Each smart sign matches a corresponding guidance strategy based on the collaborative guidance priority factor it calculates, and executes differentiated traffic guidance actions.

[0020] The beneficial effects of this invention are:

[0021] Compared to existing technologies, this application first utilizes a lightweight dual-path network to identify suspected sudden traffic events from real-time road video streams using a first smart sign deployed in a roadside smart sign array, generating a first confidence level. This achieves preliminary edge-side perception of the event, effectively reducing the latency and bandwidth pressure of centralized processing. Simultaneously, millimeter-wave radar data is used to detect abrupt changes in the motion state of the target object and generate a second confidence level. Multimodal sensor fusion compensates for the blind spots of a single visual sensor in harsh environments, improving the accuracy and anti-interference capability of sudden traffic event detection. Furthermore, a dual verification judgment is performed based on the first and second confidence levels. After successful dual verification, the event is confirmed, and sudden event information including event type, event location, and occupied lane is generated, further eliminating potential misjudgments from a single sensor and ensuring the reliability of the sudden event information. Finally, this sudden event information is broadcast to other smart signs in the smart sign array, enabling rapid sharing of event information and laying a data foundation for collaborative guidance. Finally, each smart sign in the smart sign array calculates a collaborative guidance priority factor based on its own location information and emergency information through a collaborative guidance priority calculation model deployed on the edge. Based on this factor, it matches the corresponding guidance strategy and executes differentiated traffic guidance actions. This enables each sign to form a hierarchical and coordinated guidance network according to its relative relationship with the event, avoiding redundancy or conflict in guidance information and optimizing the allocation efficiency of guidance resources.

[0022] Through the above technical solutions, this application constructs a complete closed loop from real-time edge perception, model lightweighting and edge computing combined with model edge deployment, multi-sensor fusion verification to distributed collaborative guidance. It effectively solves the problems of high detection latency, insufficient accuracy and lack of coordination in guidance of sudden traffic events in the existing technology, improves the response speed and guidance effect of sudden traffic events, and ultimately improves the overall safety and traffic efficiency of road traffic. Attached Figure Description

[0023] Figure 1 A schematic diagram of the structure of an intelligent traffic guidance system based on intelligent signage provided by the present invention;

[0024] Figure 2 This is a flowchart illustrating an intelligent traffic guidance method based on intelligent signage provided by the present invention.

[0025] In the attached diagram, the components represented by each number are as follows:

[0026] Visual preliminary perception module 11, radar motion detection module 12, fusion verification and confirmation module 13, event information broadcasting module 14, collaborative priority calculation module 15, and differentiated guidance execution module 16. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0029] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0030] Example 1, as Figure 1 As shown, this embodiment of the invention provides an intelligent traffic guidance system based on intelligent signage, comprising:

[0031] The visual preliminary perception module 11 is used by the first smart sign in the smart sign array deployed on the roadside to identify suspected sudden traffic events from the collected real-time video stream of the road and generate a first confidence level through a pre-trained lightweight dual-path network.

[0032] In existing technologies, deep learning models used for traffic incident recognition are mostly heavy networks with a large number of parameters and high computational complexity. They cannot be deployed on the edge of smart signs and rely on transmitting massive amounts of video data back to the cloud or central server for analysis. This method not only consumes a lot of communication bandwidth but also has high transmission and processing latency, making it difficult to meet the real-time response requirements for emergencies and resulting in a problem of delayed recognition of sudden traffic incidents.

[0033] To address the aforementioned issues, this application utilizes a first intelligent sign in a roadside intelligent sign array to identify suspected traffic emergencies from real-time road video streams via a pre-trained lightweight dual-path network and generate a first confidence score. The first confidence score quantifies the credibility of a suspected traffic emergency, ranging from 0 to 1; a higher value indicates a higher probability that the suspected emergency is indeed a real traffic emergency.

[0034] The roadside smart sign array consists of multiple smart signs evenly distributed at preset intervals on both sides of the road or in the central median, forming a comprehensive traffic sensing and guidance network. Each smart sign integrates a high-definition camera, millimeter-wave radar, edge computing unit, and communication module, enabling it to independently collect and process surrounding traffic data. The preset interval can be dynamically set according to road grade, number of lanes, and traffic flow. For example, it can be set at 300-500 meters per sign on urban arterial roads and 500-1000 meters per sign on highways.

[0035] The first intelligent sign in the array is the first sign to identify a suspected traffic emergency using real-time road video streams it collects. This first intelligent sign is not fixed but rather a dynamic role triggered by a suspected traffic emergency. When any intelligent sign in the array detects a suspected traffic emergency through its initial visual perception module, it automatically becomes the first intelligent sign in the current suspected traffic emergency handling process, responsible for event identification, fusion verification, and event information broadcasting. At this time, other intelligent signs in the array, upon receiving the emergency information broadcast by the first intelligent sign, perform collaborative priority calculations and differentiated guidance actions based on their relative positions to the emergency information.

[0036] Specifically, the visual preliminary perception module 11 is used for:

[0037] The collected real-time video streams of the road are input into the slow path network branch and the fast path network branch of the pre-trained lightweight dual-path network, respectively.

[0038] Static scene features in video frames are extracted using a slow path network branch at the first sampling frame rate. These static scene features include at least lane line positions, vehicle appearance, and obstacle positions.

[0039] Motion temporal features in a video frame sequence are extracted using a fast path network branch at a second sampling frame rate, wherein the second sampling frame rate is higher than the first sampling frame rate, and the motion temporal features include at least vehicle speed changes, vehicle trajectory deviations, and changes in relative distances between vehicles.

[0040] The motion time sequence features and static scene features are fused to obtain fused spatiotemporal features. Suspected sudden traffic events are then identified based on the fused spatiotemporal features, and the corresponding first confidence score is generated.

[0041] In this embodiment, the collected real-time road video stream is first input into the slow path network branch and the fast path network branch of a pre-trained lightweight dual-path network. Specifically, after the high-definition camera on the first smart sign captures the real-time road video stream at a fixed frame rate, the real-time road video stream is simultaneously input into the slow path network branch and the fast path network branch of the lightweight dual-path network. The slow path network branch focuses on extracting static scene features, while the fast path network branch focuses on extracting motion temporal features.

[0042] Secondly, static scene features are extracted from video frames using a slow path network branch at a first sampling frame rate. Specifically, the slow path network branch extracts video frames from the real-time road video stream for processing at a preset first sampling frame rate. The first sampling frame rate refers to the frequency at which the slow path network branch samples frames from the real-time road video stream; its value is lower than the second sampling frame rate of the fast path network branch, allowing the slow path branch to focus on extracting static information in the spatial dimension. Static scene features refer to visual information in the video frames that changes slowly or relatively steadily over time, used to characterize the fixed structures and long-term objects in the road environment. These features include at least lane line positions, vehicle appearance, and obstacle positions: lane line positions refer to the pixel coordinate range of the lane lines in the image coordinate system or their position in the world coordinate system after coordinate transformation; vehicle appearance includes at least visual attributes such as the vehicle's outline, color, and model, usually represented in the form of feature vectors; obstacle positions refer to the positions of fixed or temporary obstacles on the road surface other than normally traveling vehicles, such as broken-down vehicles, debris, and construction facilities.

[0043] For example, the first sampling frame rate can be dynamically set according to the actual application scenario. For instance, in urban expressway scenarios with high traffic volume and frequent events, the first sampling frame rate can be set to 5fps (i.e., processing 5 frames per second, one frame every 200ms) to reduce the computational load while ensuring sufficient extraction of static scene features. The static scene features are encoded into feature vectors for subsequent fusion analysis with motion temporal features.

[0044] Next, motion temporal features in the video frame sequence are extracted by the fast path network branch at a second sampling frame rate. The second sampling frame rate is higher than the first sampling frame rate, and the motion temporal features include at least vehicle speed changes, vehicle trajectory deviations, and changes in relative distances between vehicles. Specifically, the fast path network branch continuously extracts video frames from the real-time road video stream at a preset second sampling frame rate, forming a dense frame sequence, which is then input into the network for processing. The second sampling frame rate refers to the frequency at which the fast path network branch samples frames from the real-time road video stream; its value is higher than the first sampling frame rate to ensure that the continuous motion trajectory and instantaneous changes of objects can be captured. Among them, motion temporal features refer to information extracted from a continuous video frame sequence that reflects the dynamic changes of the target object over time. It is used to characterize the motion state and behavior patterns of traffic participants, including at least vehicle speed changes, vehicle trajectory deviations, and changes in relative distances between vehicles. Vehicle speed changes refer to the rate of change of the displacement of the same vehicle in consecutive frames over time, which can be used to estimate the instantaneous speed, acceleration, and the degree of acceleration and deceleration of the vehicle. Vehicle trajectory deviations refer to the degree to which the vehicle's trajectory deviates from the expected lane centerline or normal driving path, which can be obtained by calculating the change in the lateral distance between the vehicle's position and the lane line. Changes in relative distances between vehicles refer to the trend of changes in the spatial distance between adjacent vehicles over time, which can be used to determine whether the distance between vehicles suddenly decreases or increases.

[0045] For example, the second sampling frame rate can be dynamically set according to the actual application scenario and the requirements for motion capture accuracy. For instance, in a highway scenario, in order to accurately capture the dynamic changes of high-speed vehicles, the second sampling frame rate can be set to 20fps (i.e., processing 20 frames per second, one frame every 50ms). The fast path network branch extracts a continuous frame sequence from the real-time video stream of the road at a second sampling frame rate of 20fps and extracts motion temporal features from the frame sequence.

[0046] For example, suppose that within a 2-second time window, the fast path network branch extracts the following temporal motion features from 40 consecutive frames of images: the target vehicle's speed drops sharply from 80 km / h to 30 km / h, manifested as a rapid decrease in displacement between consecutive frames; the vehicle's trajectory shifts 0.5 meters to the right, manifested as a gradual increase in the lateral distance between the vehicle's center point and the lane centerline; the relative distance between the vehicle and the following vehicle decreases from 30 meters to 5 meters, manifested as a rapid decrease in the distance between the vehicles in the image. These temporal motion features are encoded into feature vectors for subsequent fusion analysis with static scene features.

[0047] Finally, the motion temporal features and static scene features are fused to obtain fused spatiotemporal features. These fused spatiotemporal features are then used to identify suspected sudden traffic events and generate a corresponding first confidence score. Specifically, because the slow path network branch and the fast path network branch have different sampling frame rates, their output features have inconsistent resolution in the temporal dimension. Therefore, the motion temporal features output by the fast path network branch are first upsampled in the temporal dimension, for example, through linear interpolation or repeated sampling, to align their temporal resolution with the static scene features. Then, the aligned motion temporal features and static scene features are concatenated along the channel dimension to form fused spatiotemporal features. Subsequently, the existence of suspected sudden traffic events is identified based on these fused spatiotemporal features, and a corresponding first confidence score is generated based on the identification results, completing the preliminary identification of suspected events.

[0048] Feature fusion refers to integrating the static scene features extracted by the slow path network branch and the motion temporal features extracted by the fast path network branch into a unified feature representation, so as to make comprehensive use of spatial static information and temporal dynamic information; the first confidence level is between 0 and 1, which is used to quantify the credibility of the lightweight dual-path network's preliminary judgment on the occurrence of suspected sudden traffic events through visual modality, and to complete the preliminary identification of suspected sudden traffic events.

[0049] Specifically, the construction process of the lightweight dual-path network includes:

[0050] A dual-path network architecture is constructed, which includes a slow path network branch, a fast path network branch, and a lateral connection layer. The lateral connection layer is used to connect the slow path network branch and the fast path network branch. The slow path network branch and the fast path network branch each contain multiple convolutional layers. Each convolutional layer contains multiple feature channels. Each feature channel is used to extract different dimensional features of the input data.

[0051] Video stream samples containing historical traffic accidents were collected from a historical traffic monitoring video database, and the video stream samples were labeled with the type of emergency to obtain a training dataset and a corresponding supervision label set.

[0052] The dual-path network architecture is trained in a supervised manner using the training dataset and the supervised label set until it is verified to converge, thus obtaining the pre-trained dual-path network.

[0053] A lightweight dual-path network is obtained by performing a lightweight modification on the pre-trained dual-path network.

[0054] A lightweight dual-path network is deployed in the edge computing unit of the first smart sign, wherein the slow path network branch is deployed in the main processor of the edge computing unit, and the fast path network branch is deployed in the neural network acceleration unit of the edge computing unit.

[0055] In this embodiment, a dual-path network architecture is first constructed, comprising a slow-path network branch, a fast-path network branch, and lateral connection layers. The lateral connection layers connect the slow-path and fast-path network branches, each containing multiple convolutional layers. Each convolutional layer contains multiple feature channels, and each feature channel is used to extract different dimensional features from the input data. Specifically, the dual-path network architecture is the basic framework of a lightweight dual-path network. Its design aims to extract static scene features and motion temporal features from real-time road video streams, respectively, and fuse them through lateral connection layers to enhance the network's ability to represent spatiotemporal information. The convolutional layers are the basic feature extraction units of the network. Through the stacking of multiple convolutions, high-level semantic features can be learned progressively from low-level edge features. The multiple feature channels in each convolutional layer can be considered as parallel learning feature extractors, with each channel focusing on different patterns in the input data, such as edges in different directions, color textures, and object parts. The lateral connection layer achieves information flow by upsampling the features of the fast path branch in the time dimension and concatenating or adding them element by element with the features of the slow path branch. This allows the slow path branch to obtain high-frequency motion information, while the fast path branch can obtain fine spatial structure information, thereby compensating for the lack of features of a single branch and improving the expressive ability of fused spatiotemporal features.

[0056] For example, the construction process and parameter configuration of the dual-path network architecture can be referenced as follows: The slow path network branch sets up 4 convolutional layers, with the number of convolutional kernels (i.e., the number of feature channels) set to 64, 128, 256, and 256 respectively, the kernel size is 3×3, the stride is 2, and the activation function is ReLU; the fast path network branch sets up 4 convolutional layers, with the number of convolutional kernels set to 32, 64, 128, and 128 respectively, and the kernel size and stride settings are consistent with the corresponding layers of the slow path branch; the lateral connection layer is set between the 2nd, 3rd, and 4th layers of the slow path branch and the 2nd, 3rd, and 4th layers of the fast path branch. For each layer to be connected, the feature map of the fast path branch is first upsampled in the time dimension through linear interpolation to align its time resolution with the feature map of the slow path branch. Then, the upsampled fast path features and slow path features are concatenated along the channel dimension to obtain the fused feature map, which is used as the input of the subsequent layers to construct the complete dual-path network architecture.

[0057] Secondly, video stream samples containing historical traffic accidents are collected from historical traffic monitoring video databases, and these video stream samples are labeled with the type of emergency to obtain a training dataset and a corresponding supervision label set. Specifically, firstly, a large number of video clips containing real emergency traffic events are extracted from the historical traffic monitoring video database of traffic management departments or publicly available traffic accident datasets. These video clips can cover different scene types, different accident types, and different shooting angles and distances to ensure the diversity and representativeness of the training data. All extracted video stream samples together constitute the training dataset. Then, professional annotators perform detailed annotations on each video stream sample. The annotation content includes at least: the type of emergency (e.g., rear-end collision, rollover, foreign object intrusion, etc.), the timestamps of the start and end frames of the event, and the position of key targets in the event (e.g., accident vehicles, obstacles, etc.) in each frame. All annotation information that corresponds one-to-one with the video samples in the training dataset together constitutes the supervision label set.

[0058] For example, 10,000 video stream samples containing sudden traffic events can be collected from a historical traffic monitoring video database. Each sample is between 10 and 30 seconds long, with a resolution of 1920×1080 and a frame rate of 30fps. The collected samples cover different weather and lighting conditions and include various accident types. For each video stream sample, its event type is labeled, along with the start and end frames of the event. Rectangular bounding boxes are used to mark the positions of each key vehicle or obstacle involved in the event. All video samples form the training dataset, and all corresponding annotation information forms the supervised label set for subsequent network training.

[0059] Next, supervised training of the dual-path network architecture is performed using the training dataset and the supervised label set until validation convergence, resulting in a pre-trained dual-path network. For example, the training process of the dual-path network can refer to the following steps: 1. Data preparation: The training dataset and the supervised label set are randomly divided into training, validation, and test sets in a ratio of 7:1.5:1.5. The training set is used for updating model parameters, the validation set is used for adjusting hyperparameters and determining convergence, and the test set is used for final evaluation of model performance. 2. Model training: Using video stream samples from the training set as input features and corresponding event type labels as supervision signals, the cross-entropy loss function is used to calculate the error between the model output and the true labels. The optimizer is Adam, the initial learning rate is set to 0.001, and a cosine annealing strategy is used to dynamically adjust the learning rate. The batch size is set to 8, with each batch containing 8 video segments. The model is trained for a total of 100 epochs. After each epoch, the loss value and classification accuracy are calculated on the validation set. 3. Convergence determination: When the loss value on the validation set no longer decreases for 10 consecutive rounds (for example, a decrease of less than 0.001 is considered to be no longer decreasing) and the classification accuracy no longer improves, the training is considered to have converged, training is stopped, and the model parameters that perform best on the validation set at this time are saved to obtain the pre-trained dual-path network.

[0060] Furthermore, the pre-trained dual-path network is subjected to lightweight modification to obtain a lightweight dual-path network. Lightweight modification refers to reducing the number of network parameters and computational complexity while maintaining network recognition accuracy as much as possible, making it adaptable to the limited computing resources and storage space of the edge computing unit on the smart signage device. The lightweight dual-path network, after the above modification, has significantly reduced parameter and computational costs while still retaining its core spatiotemporal feature extraction capabilities. It can be deployed on the main processor and neural network acceleration unit of the edge computing unit for real-time operation. Specifically, lightweight modification can be achieved using weight quantization technology: converting the floating-point weight parameters of each layer in the pre-trained dual-path network into a low-bit-width integer format, thereby significantly reducing model storage space and memory usage, and accelerating inference using integer operations.

[0061] Finally, a lightweight dual-path network is deployed in the edge computing unit of the first smart sign. The slow path network branch is deployed in the main processor of the edge computing unit, while the fast path network branch is deployed in the neural network acceleration unit of the edge computing unit. The edge computing unit refers to an embedded computing module integrated within the smart sign for local data processing and model inference. It typically consists of a general-purpose main processor and a dedicated neural network acceleration unit: the main processor (e.g., an ARM architecture CPU) is responsible for running control logic, lightweight computing tasks, and computations that cannot be hardware-accelerated; the neural network acceleration unit (e.g., an NPU, GPU, or DSP) provides hardware acceleration for intensive computations such as convolution and pooling in deep learning models, significantly improving inference speed and energy efficiency.

[0062] Specifically, in deployment, heterogeneous computing power is allocated based on the computational characteristics of the two branches of the dual-path network: the slow path branch processes static scenes at a low frame rate, with relatively small computational load but high requirements for feature accuracy, so it is deployed on the main processor, leveraging the CPU's flexibility and high-precision floating-point arithmetic capabilities to ensure accurate feature extraction; the fast path branch processes dynamic temporal information at a high frame rate, is computationally intensive and highly parallel, suitable for deployment on the neural network acceleration unit, utilizing its parallel computing capabilities to achieve high throughput processing. The two branches execute in parallel and interact with features through shared memory or lightweight communication mechanisms, ultimately achieving efficient real-time inference for the entire lightweight dual-path network.

[0063] It should be noted that the hardware architecture of the edge computing unit and its driver adaptation are well-known technologies in the field. Those skilled in the art can perform corresponding model conversion and deployment optimization according to the specific chip platform (e.g., Rockchip RK3588, Horizon Robotics Journey series, etc.) and deployment framework (e.g., TensorRT, Tengine, OpenVINO, etc.), which will not be elaborated here.

[0064] Furthermore, the pre-trained dual-path network is modified to be lightweight, resulting in a lightweight dual-path network, including:

[0065] In the training dataset, the image regions containing lane lines and vehicles in each video frame are labeled as regions of interest, and the image regions other than the regions of interest are labeled as background regions.

[0066] The weight parameters of each convolutional layer in the pre-trained dual-path network are statistically analyzed, and the first floating-point number distribution range of the weight parameters corresponding to the region of interest and the second floating-point number distribution range of the weight parameters corresponding to the background region are obtained respectively.

[0067] Based on the first floating-point number distribution range, the first quantization parameter corresponding to the region of interest is calculated, and based on the second floating-point number distribution range, the second quantization parameter corresponding to the background region is calculated, wherein the quantization bit width corresponding to the first quantization parameter is higher than the quantization bit width corresponding to the second quantization parameter.

[0068] The first quantization parameter is used to quantize the weight parameters corresponding to the region of interest, converting the weight parameters corresponding to the region of interest from floating-point format to low-bit-width integer format.

[0069] The second quantization parameter is used to quantize the weight parameter corresponding to the background region, converting the weight parameter corresponding to the background region from floating-point format to low-bit-width integer format.

[0070] The weight parameters corresponding to the region of interest after quantization transformation and the weight parameters corresponding to the background region after quantization transformation are merged and stored according to their respective region identifiers to obtain a lightweight dual-path network.

[0071] In this embodiment, firstly, in the training dataset, the image region containing lane lines and vehicles in each video frame image is labeled as the region of interest, and the image region other than the region of interest is labeled as the background region. The region of interest refers to the region in the video frame image that is highly relevant to the identification of sudden traffic events, and at least contains key targets such as lane lines and vehicles; the background region refers to redundant regions in the video frame image that are unrelated to the identification of sudden traffic events, such as roadside vegetation, sky, and buildings.

[0072] Specifically, specialized image annotation tools can be used to perform fine-grained pixel-level or region-level annotations on each frame of video images in the training dataset: regions containing key targets such as lane lines and vehicles are selected and marked as regions of interest (ROIs); all other regions in the image not selected are considered background regions by default. During annotation, it is crucial to ensure that the ROI completely covers all targets that might affect event recognition, while avoiding mislabeling irrelevant background elements into the ROI, to guarantee that subsequent region-aware quantization operations can accurately distinguish the importance of weights.

[0073] For example, open-source annotation tools such as LabelImg or CVAT can be used for annotation. Taking LabelImg as an example, the operation steps are as follows: First, load a video frame image from the training dataset; then, use the bounding box tool to select each car and each clearly visible lane line in the image, and select a predefined "Region of Interest" category for each selected area in the pop-up label options; after the selection is completed, save the annotation results and generate an XML file with the same name as the image, which records the coordinates and category information of each region of interest bounding box. The area in the video frame image not covered by any bounding box is the background area and does not require additional annotation. In this way, by traversing all video frame images in the training dataset, an annotated dataset with region of interest / background region division can be obtained for subsequent lightweight modification.

[0074] Secondly, the weight parameters of each convolutional layer in the pre-trained dual-path network are statistically analyzed. The first floating-point number distribution range of the weight parameters corresponding to the region of interest (ROI) and the second floating-point number distribution range of the weight parameters corresponding to the background region are obtained. Specifically, for each convolutional layer, its weight parameters correspond to different spatial regions of the input image within the receptive field. By recording the type of image region mainly activated by each weight during forward propagation (i.e., whether the weight responds more strongly to the ROI or the background region), the weight parameters can be divided into a part related to the ROI and a part related to the background region. Then, the floating-point number distributions of these two parts of the weights are statistically analyzed to obtain the corresponding first and second floating-point number distribution ranges.

[0075] Here, the weight parameters refer to the core learnable parameters of the convolutional layers in the pre-trained dual-path network, which are used to extract features from the input data; the first floating-point distribution range refers to the value range of all weight parameters related to the region of interest; and the second floating-point distribution range refers to the value range of all weight parameters related to the background region.

[0076] For example, a network parameter analysis tool can be used to statistically analyze a convolutional layer. This tool can be implemented based on a deep learning framework (such as PyTorch or TensorFlow) and a numerical computation library (such as NumPy). First, a pre-trained dual-path network is loaded, and the weight parameters of the target convolutional layer are read. Then, by combining the mapping relationship between the pre-annotated region of interest mask and the receptive field of the network, the image region type (i.e., region of interest or background region) corresponding to each weight parameter is determined, thereby dividing the weights into a part related to the region of interest and a part related to the background region. Finally, the floating-point values ​​of these two parts of the weights are counted separately, and their minimum and maximum values ​​are obtained to obtain the corresponding distribution range.

[0077] For example, statistical results from network parameter analysis tools show that the weights responsible for extracting features such as vehicles and lane lines (i.e., related to the region of interest) are mainly concentrated in the range of [-2.5, 3.2], while the weights responsible for extracting features such as the sky and roadside vegetation (i.e., related to the background region) are mainly concentrated in the range of [-0.8, 0.6]. Thus, the first floating-point number distribution range is [-2.5, 3.2], and the second floating-point number distribution range is [-0.8, 0.6].

[0078] Next, based on the first floating-point number distribution range, the first quantization parameter corresponding to the region of interest is calculated, and based on the second floating-point number distribution range, the second quantization parameter corresponding to the background region is calculated. The quantization bit width corresponding to the first quantization parameter is higher than the quantization bit width corresponding to the second quantization parameter. Specifically, for the region of interest, because its weight distribution range is wider and it is sensitive to recognition accuracy, a higher quantization bit width is used for quantization. The scaling factor and zeros are calculated based on the first floating-point number distribution range to obtain the first quantization parameter. For the background region, because its weight distribution range is narrower and its accuracy requirements are lower, a lower quantization bit width is used for quantization. The corresponding scaling factor and zeros are calculated based on the second floating-point number distribution range to obtain the second quantization parameter.

[0079] Among them, quantization parameters refer to the parameters required to map floating-point weights to a low-bit-width integer format, including at least scaling factors and zeros, which are used to determine the linear mapping relationship between floating-point values ​​and integer values; quantization bit width refers to the number of bits of the integer weights after quantization. The higher the bit width, the higher the quantization accuracy, but the model storage and computational costs also increase accordingly.

[0080] For example, the specific value of the quantization bit width can be dynamically set by those skilled in the art based on actual hardware support, tolerance for accuracy loss, and model compression requirements. For instance, the quantization bit width corresponding to the first quantization parameter can be set to 8 bits (i.e., quantization range [-128, 127]), and the quantization bit width corresponding to the second quantization parameter can be set to 4 bits (i.e., quantization range [-8, 7]), thereby effectively reducing the storage space and computational overhead occupied by the background region weights while ensuring the accuracy of region of interest recognition.

[0081] Furthermore, the weight parameters corresponding to the region of interest are quantized using the first quantization parameter, converting them from floating-point format to low-bit-width integer format. Low-bit-width integer format refers to an integer data type that uses fewer bits (e.g., 8 bits, 16 bits) to represent values. Compared to the original floating-point format, this reduces model storage space and memory usage, while also accelerating the inference process using integer operations.

[0082] Specifically, for each weight parameter belonging to the region of interest, a linear mapping is performed according to the scaling factor and zero point in the first quantization parameter to obtain the low-bit width integer representation of the target bit width. The mapping formula is: low-bit width integer value = round(floating-point value / scaling factor + zero point), and the result is restricted to the integer range corresponding to the target bit width. For example, if 8-bit quantization is used, the range is [-128, 127] or [0, 255].

[0083] For example, the quantization bit width can be dynamically set by those skilled in the art based on hardware support capabilities, tolerance for accuracy loss, and model compression requirements. For instance, 8 bits, 16 bits, or other bit widths can be selected, and it is not fixed to a certain value. Through quantization conversion, the core features of the region of interest are preserved, while the number of parameters and computational cost are significantly reduced, laying the foundation for subsequent real-time inference at the edge.

[0084] Furthermore, a second quantization parameter is used to quantize and convert the weight parameters corresponding to the background region from floating-point format to low-bit-width integer format. Specifically, since the background region has a relatively small impact on recognition accuracy, a lower quantization bit width can be used to further compress the model, such as 4 bits, with a range of [-8, 7] or [0, 15]. The specific bit width can be flexibly selected according to actual needs and is not limited to 4 bits. After quantization conversion, redundant information of the background region weights is effectively removed, and the overall model size is further reduced. The core advantages of using the low-bit-width integer format are: on the one hand, it significantly reduces the model storage requirements, enabling the lightweight dual-path network to adapt to the limited storage resources of the edge computing unit of the smart sign; on the other hand, it utilizes the hardware acceleration capability of integer operations to improve inference speed and meet the requirements of real-time detection of sudden traffic events.

[0085] Finally, the weight parameters corresponding to the quantized region of interest (ROI) and the weight parameters corresponding to the quantized background region are merged and stored according to their respective region identifiers to obtain a lightweight dual-path network. The region identifier is a label used to distinguish between the ROI and the background region; for example, the ROI is identified as "ROI" and the background region as "BG". Merging and storing refers to organizing the two types of weight parameters according to the original network structure and integrating them into the same network through additional identifiers or implicit storage layout to form a complete lightweight model.

[0086] Specifically, the quantized weights retain the same hierarchical structure as the pre-trained network, but each weight parameter is associated with a region identifier. During inference, the edge computing unit dynamically selects to use either the first or second quantization parameter for dequantization based on the region type corresponding to the input feature, and then calls the corresponding weights to complete the calculation. This storage method significantly reduces storage space and computational overhead while maintaining the functionality of the dual-path network.

[0087] Furthermore, the motion time-series features and static scene features are fused to obtain fused spatiotemporal features. These fused spatiotemporal features are then used to identify suspected sudden traffic incidents and generate corresponding first confidence scores, including:

[0088] The motion temporal features are upsampled in the time dimension to obtain time-aligned motion features that are aligned with the static scene features in the time dimension.

[0089] The time-aligned motion features and static scene features are concatenated to construct multiple spatiotemporal feature pairs. Among them, the spatiotemporal feature pairs include at least the lane-speed feature pair consisting of the lane line position and the corresponding vehicle speed change, the obstacle-trajectory feature pair consisting of the obstacle position and the corresponding vehicle trajectory offset, and the appearance-trajectory feature pair consisting of the vehicle appearance and the corresponding vehicle trajectory offset.

[0090] For each spatiotemporal feature pair, calculate the matching degree between static scene features and time-aligned motion features in the spatiotemporal feature pair to obtain multiple feature matching degrees;

[0091] The minimum value of multiple feature matching degrees is selected as the association consistency index. When the association consistency index is lower than the preset consistency threshold, it is determined that there is an abnormal association pattern, and the scene corresponding to the abnormal association pattern is identified as a suspected sudden traffic event.

[0092] Based on the difference between the correlation consistency index and the preset consistency threshold, the first confidence level corresponding to the suspected sudden traffic incident is generated.

[0093] Specifically, for each spatiotemporal feature pair, the matching degree between the static scene features and the time-aligned motion features in the spatiotemporal feature pair is calculated to obtain multiple feature matching degrees, including:

[0094] For lane-speed feature pairs, the normalized lane-speed matching degree is calculated based on the ratio of the vehicle's current lateral offset to the lane's allowed lateral offset range, and the ratio of the vehicle's current speed to the lane's historical average speed.

[0095] For obstacle-trajectory feature pairs, the normalized obstacle-trajectory matching degree is calculated based on the ratio of the shortest distance between the vehicle's current trajectory and the obstacle's position to the preset safe distance, and the ratio of the angle between the vehicle's current speed direction and the obstacle's direction to the maximum allowable angle.

[0096] For appearance-trajectory feature pairs, the normalized appearance-trajectory matching degree is calculated based on the similarity score between the current appearance features and historical appearance features of the vehicle, and the ratio of the deviation between the current trajectory and historical trajectory to the maximum allowable deviation.

[0097] In this embodiment, the motion temporal features are first upsampled in the temporal dimension to obtain time-aligned motion features that are aligned with the static scene features in the temporal dimension. Specifically, because the second sampling frame rate of the fast path network branch is higher than the first sampling frame rate of the slow path network branch, the motion temporal features output by the fast path have a denser time step, while the time step of the static scene features output by the slow path is relatively sparse. To enable subsequent fusion of the two types of features, the motion temporal features need to be upsampled in the temporal dimension to make their temporal resolution consistent with that of the static scene features.

[0098] Upsampling refers to adding time dimension sampling points to the motion temporal features through interpolation (e.g., linear interpolation) or duplication, thereby obtaining features with the same time step as the static scene features. The features obtained after upsampling are time-aligned motion features.

[0099] For example, if the time step of the motion time series feature is 20 (i.e., 20 feature vectors per second) and the time step of the static scene feature is 5 (i.e., 5 feature vectors per second), then the time step of the motion time series feature is adjusted to 5 by linear interpolation to ensure that each static feature has a corresponding motion feature at any given time.

[0100] Secondly, time-aligned motion features and static scene features are concatenated to construct multiple spatiotemporal feature pairs. These pairs include at least three types: lane-speed feature pairs (comprising lane line positions and corresponding vehicle speed changes), obstacle-trajectory feature pairs (comprising obstacle positions and corresponding vehicle trajectory offsets), and appearance-trajectory feature pairs (comprising vehicle appearance and corresponding vehicle trajectory offsets). Specifically, time-aligned motion features and static scene features at the same time step are concatenated along the channel dimension to form joint features containing both static and dynamic information. Based on this, spatiotemporal feature pairs with specific semantic relationships are further extracted, including at least three categories: lane-speed feature pairs (paired with lane line position static features and corresponding vehicle speed change motion features to characterize the relationship between vehicle speed and lane position); obstacle-trajectory feature pairs (paired with obstacle position static features and corresponding vehicle trajectory offset motion features to characterize the interaction between vehicle trajectory and obstacles); and appearance-trajectory feature pairs (paired with vehicle appearance static features and corresponding vehicle trajectory offset motion features to characterize the consistency between the vehicle's appearance and its trajectory).

[0101] Feature concatenation refers to merging feature vectors along the channel dimension to form a higher-dimensional feature representation; spatiotemporal feature pairs refer to feature units formed by combining time-aligned motion features and static scene features at the same point in time. Each feature pair reflects the relationship between specific static elements and dynamic behaviors.

[0102] Next, for each spatiotemporal feature pair, the matching degree between the static scene features and the time-aligned motion features in the spatiotemporal feature pair is calculated to obtain multiple feature matching degrees. Specifically, the matching degree is used to quantify the degree of consistency or anomalousness between static scene features and time-aligned motion features (i.e., static elements and dynamic behaviors). The value is normalized to between 0 and 1. The closer the value is to 1, the more consistent the two are, and the more normal the motion state of the target object; conversely, the more anomalous the value is. Appropriate calculation methods are used depending on the type of spatiotemporal feature pair.

[0103] 1. For lane-speed feature pairs, the normalized lane-speed matching degree is calculated based on the ratio of the vehicle's current lateral offset to the lane's allowed lateral offset range, and the ratio of the vehicle's current speed to the lane's historical average speed. The vehicle's current lateral offset refers to the vertical distance between the vehicle's center point and the lane centerline, which can be calculated from the vehicle's position and lane line detection results in the image. The lane's allowed lateral offset range refers to the maximum lateral offset allowed for normal vehicle movement within the lane, which can be determined based on the lane width and vehicle width. For example, if the lane width is 3.5 meters and the vehicle width is 1.8 meters, the allowed lateral offset range is ±(3.5 / 2 - 1.8 / 2) = ±0.85 meters. Alternatively, it can be set to a fixed value, such as ±0.5 meters, according to safe driving regulations. The lane's historical average speed refers to the average speed of all vehicles passing through the lane within a past time window, which can be obtained in real-time from the traffic monitoring system.

[0104] For example, the formula for calculating lane-speed matching degree can be set as: Lane-speed matching degree = 0.5 × (1 - |current lateral offset of vehicle / allowable lateral offset range|) + 0.5 × (1 - |current speed of vehicle / historical average speed of the lane - 1|).

[0105] For example, assuming the current lateral offset is 0.2 meters, the allowable lateral offset range is ±0.5 meters, the current speed is 55 km / h, and the historical average speed is 60 km / h, substituting these values ​​into the formula, we get the lane-speed matching degree: 0.5 × (1 - 0.2 / 0.5) + 0.5 × (1 - |55 / 60 - 1|) ≈ 0.76. The closer the lane-speed matching degree is to 1, the more consistent the vehicle's speed and position are with normal driving patterns; conversely, a lower degree indicates a greater probability of abnormal events.

[0106] 2. For obstacle-trajectory feature pairs, the normalized obstacle-trajectory matching degree is calculated based on the ratio of the shortest distance between the vehicle's current trajectory and the obstacle's position to the preset safe distance, and the ratio of the angle between the vehicle's current speed direction and the obstacle's direction to the maximum permissible angle. The shortest distance between the vehicle's current trajectory and the obstacle's position refers to the minimum distance between the vehicle's current travel trajectory (which can be approximated as an extension of the vehicle's current position along the speed direction) and the obstacle's boundary, and can be obtained through coordinate calculation. The preset safe distance refers to the minimum safe distance that should be maintained between the vehicle and the obstacle, which can be set according to traffic regulations or industry standards, such as 5 meters on highways and 3 meters on urban roads. The angle between the vehicle's current speed direction and the obstacle's direction refers to the angular difference between the vehicle's travel direction and the direction the vehicle points towards the obstacle, which can be calculated using the vehicle's speed vector and the obstacle's relative position vector. The maximum permissible angle refers to the maximum steering angle allowed when the vehicle is normally avoiding obstacles, and can be set according to vehicle dynamics and road curvature; for example, 90 degrees represents the maximum range of steering that the vehicle can take. If the angle exceeds 90 degrees, it is considered a serious mismatch.

[0107] For example, the formula for calculating the obstacle-trajectory matching degree can be set as: obstacle-trajectory matching degree = 0.5 × (shortest distance between the current vehicle trajectory and the obstacle position / preset safe distance) + 0.5 × (1 - angle between the current vehicle speed direction and the obstacle direction / maximum allowable angle).

[0108] For example, assuming the closest distance between the vehicle and the obstacle is 4 meters, the preset safe distance is 5 meters, the current angle is 30°, and the maximum allowable angle is 90°, substituting into the formula, we can calculate the obstacle-trajectory matching degree as 0.5×(4 / 5)+0.5×(1-30 / 90)≈0.73.

[0109] 3. For appearance-trajectory feature pairs, a normalized appearance-trajectory matching degree is calculated based on the similarity score between the current and historical appearance features of the vehicle, and the ratio of the deviation between the current and historical trajectories to the maximum permissible deviation. The similarity score between the current and historical appearance features refers to the degree of similarity between the current vehicle appearance feature vector and the historical appearance feature vector calculated using a feature matching algorithm (e.g., cosine similarity, Euclidean distance), with a value ranging from 0 to 1, where 1 indicates complete similarity. The appearance feature vector and the historical appearance feature vector can be obtained by uniformly encoding the vehicle appearance in the static scene features and the historical vehicle appearance in the same dimension.

[0110] For example, cosine similarity or Euclidean distance can be used for calculation. For instance, when using cosine similarity, the current appearance feature vector and the historical appearance feature vector are denoted as A and B respectively, then the similarity score = (A·B) / (|A|×|B|). If A and B have been normalized, the cosine similarity can be directly used as the score. When using Euclidean distance, the distance d can be calculated first, and then converted into a score between 0 and 1 using the similarity score = 1 / (1+d).

[0111] Among them, the deviation between the current trajectory and the historical trajectory refers to the lateral or longitudinal offset distance of the current vehicle's driving trajectory relative to the vehicle's historical normal driving trajectory; the maximum allowable deviation refers to the range of trajectory fluctuation allowed when the vehicle is driving normally, which can be set according to lane width, driving habits, etc., for example, 1 meter.

[0112] For example, the formula for calculating the appearance-trajectory matching degree can be set as: appearance-trajectory matching degree = 0.5 × appearance similarity score + 0.5 × (1 - deviation between the current trajectory and the historical trajectory of the vehicle / maximum allowable deviation).

[0113] For example, assuming the current appearance feature vector is [0.2, 0.8, 0.3] and the historical appearance feature vector is [0.3, 0.7, 0.4], the similarity score calculated using cosine similarity is approximately 0.98. If the current trajectory deviation is 0.3 meters and the maximum allowable deviation is 1 meter, substituting these values ​​into the appearance-trajectory matching degree calculation formula, we get the appearance-trajectory matching degree = 0.5 × 0.98 + 0.5 × (1 - 0.3 / 1) = 0.84.

[0114] Thus, using the above calculation method, the corresponding feature matching degree is obtained for each type of spatiotemporal feature pair. These matching degrees collectively reflect the degree of coordination between the static environment and dynamic behavior in the current scenario, providing a basis for subsequent calculation of correlation consistency indicators and judgment of suspected events.

[0115] Furthermore, the minimum value among multiple feature matching degrees is selected as the association consistency index. When the association consistency index is lower than a preset consistency threshold, an abnormal association pattern is determined to exist, and the scene corresponding to the abnormal association pattern is identified as a suspected sudden traffic incident. Specifically, the minimum value among all feature matching degrees is taken as the association consistency index to characterize the overall coordination degree between static scene features and motion temporal features in the current scene. The smaller the association consistency index, the more inconsistent the association between at least one static element and dynamic behavior, and the more likely the motion state of the target object is to be abnormal. The association consistency index is then compared with the preset consistency threshold: if the association consistency index is lower than the preset consistency threshold, an abnormal association pattern is determined to exist in the current scene, and it is initially identified as a suspected sudden traffic incident; conversely, if the association consistency index is not lower than the preset consistency threshold, the current scene is determined to be normal.

[0116] The preset consistency threshold is an empirical parameter that can be dynamically set according to the actual application scenario, used to define the abnormal threshold value of static and dynamic correlation. The setting principle of this threshold can be determined based on the following factors: for example, selecting the quantile that maximizes the distinction between the two types of events based on the statistical distribution of matching degrees between normal and abnormal events in historical traffic accident data; or adjusting it according to the traffic management department's tolerance for false alarm and missed alarm rates; or it can be set differently based on the safety requirements of different road grades (e.g., highways, urban expressways, and ordinary roads). Those skilled in the art can determine the appropriate preset consistency threshold value for a specific scenario through experience, experiments, or simulations. For example, setting it to 0.6 in highway scenarios to improve sensitivity, and setting it to 0.5 in urban roads to balance false alarms and missed alarms.

[0117] For example, if the preset consistency threshold is 0.6, and the calculated matching degrees of the three features are 0.875, 0.73, and 1.0 respectively, the minimum value of 0.73 is selected as the correlation consistency index. Since the correlation consistency index is higher than the preset consistency threshold of 0.6, it is determined that there is no suspected sudden traffic incident. Conversely, if the calculated matching degrees of the three features are 0.8, 0.5, and 0.7 respectively, the minimum value of 0.5 is selected as the correlation consistency index. Since the correlation consistency index is lower than the preset consistency threshold of 0.6, it is determined that there is a suspected sudden traffic incident.

[0118] Finally, based on the difference between the correlation consistency index and the preset consistency threshold, a first confidence level is generated for the suspected traffic emergency. Specifically, when a suspected traffic emergency is determined to exist, the first confidence level is used to quantify the credibility of the suspected traffic emergency. It is calculated as follows: First Confidence Level = (Preset Consistency Threshold - Correlation Consistency Index) / Preset Consistency Threshold, ensuring the confidence level ranges from 0 to 1. The lower the correlation consistency index and the larger the difference between the correlation consistency index and the preset consistency threshold, the higher the calculated first confidence level, indicating a greater likelihood of the event occurring.

[0119] For example, if the preset consistency threshold is 0.6 and the association consistency index is 0.5, then the first confidence level = (0.6-0.5) / 0.6≈0.167.

[0120] Thus, through the above steps, suspected traffic emergencies can be identified based on the fusion of spatiotemporal features and a corresponding first confidence level can be generated. The preliminary determination of suspected traffic emergencies in the visual modality can be completed through a pre-trained lightweight dual-path network.

[0121] In summary, compared to existing technologies, this application utilizes a first intelligent sign in a roadside intelligent sign array to identify suspected sudden traffic events from real-time road video streams and generate a first confidence score through a pre-trained lightweight dual-path network. Thus, by deploying a lightweight dual-path network at the edge of the intelligent sign to analyze the video stream in real time, rapid initial perception of sudden traffic events at the edge is achieved, generating a quantified first confidence score. This provides a reliable initial screening basis for subsequent fusion verification, effectively reducing the latency and bandwidth pressure of traditional centralized processing.

[0122] The radar motion detection module 12 is used to synchronously call millimeter-wave radar data, detect the abrupt change characteristics of the motion state of the target object in the suspected sudden traffic incident, and generate a second confidence level based on the abrupt change characteristics of the motion state.

[0123] The aforementioned visual preliminary perception module 11 detects suspected sudden traffic incidents from the collected real-time road video stream through a pre-trained lightweight dual-path network. However, video detection may produce false alarms due to factors such as lighting and occlusion.

[0124] To address the aforementioned issues, this application utilizes a radar motion detection module 12 to synchronously access millimeter-wave radar data, detect abrupt changes in the motion state of a target object in a suspected sudden traffic incident, and generate a second confidence level based on these abrupt changes. The millimeter-wave radar can accurately measure the target's speed and distance, and is unaffected by weather or lighting conditions, precisely capturing the target object's motion state, such as speed and acceleration.

[0125] Specifically, the radar motion detection module 12 is used for:

[0126] Based on the spatial location of the target object in a suspected traffic emergency, millimeter-wave radar data of the corresponding area is retrieved;

[0127] Extract the velocity sequence of the target object within a continuous time window from the millimeter-wave radar data, and obtain the historical velocity data sequence of the target object from the historical trajectory data;

[0128] Based on the velocity sequence and historical velocity data sequence, the velocity decrease of the target object within a preset time window is calculated, and the velocity decrease is used as a feature of abrupt change in motion state.

[0129] Based on historical speed data series, the mean and standard deviation of the historical speed data series are statistically analyzed, and the threshold for the speed decrease is calculated based on the linear combination of the mean and standard deviation.

[0130] Calculate the ratio of the speed decrease to the speed decrease threshold, and use this ratio as the second confidence level.

[0131] In this embodiment, the system first calls up millimeter-wave radar data for the corresponding area based on the spatial location of the target object in a suspected traffic emergency. The spatial location of the target object refers to the precise location of key targets involved in the suspected traffic emergency, such as accident vehicles or obstacles, in the road coordinate system, obtained by the first intelligent signage through visual recognition. The millimeter-wave radar data for the corresponding area refers to real-time point data collected by millimeter-wave radar deployed in conjunction with the first intelligent signage, corresponding to that spatial location, including information such as the target's speed, distance, and azimuth.

[0132] Specifically, after the first intelligent signboard visually identifies a suspected traffic emergency and determines the spatial location of the target object, it indexes the millimeter-wave radar in the corresponding coverage area based on the location and calls up the millimeter-wave radar data collected by that radar within the same time window to ensure that subsequent analysis is based on synchronous multimodal data of the same target object.

[0133] Secondly, the velocity sequence of the target object within a continuous time window is extracted from the millimeter-wave radar data, and the historical velocity data sequence of the target object is obtained from the historical trajectory data. The continuous time window refers to the length of time used to analyze the target's motion state, such as 0.5 seconds before and after a suspected traffic incident, totaling 1 second; the velocity sequence refers to the set of instantaneous velocity values ​​of the target object arranged in chronological order within this continuous time window. The historical trajectory data refers to the driving trajectory and corresponding speed records of the target object over a past period, continuously tracked and stored by the first intelligent signage; the historical velocity data sequence is the velocity value sequence of the target object under normal driving conditions extracted from this historical trajectory data.

[0134] Next, based on the velocity sequence and historical velocity data sequence, the velocity decrease of the target object within a preset time window is calculated, and the velocity decrease is used as a feature of abrupt change in motion state. The preset time window is a sub-interval within a continuous time window, used to focus on the moment of abrupt change and calculate the time range of the velocity decrease (e.g., 0.5 seconds). The velocity decrease refers to the difference between the initial velocity and the final velocity of the target object within the preset time window, used to characterize the degree of abrupt change in the target object's motion state. Abrupt change in motion state features refers to characteristics that reflect sudden changes in the target object's motion state, such as the velocity decrease (because in sudden traffic incidents, the target object often experiences a sudden drop in speed).

[0135] For example, the preset time window is set to 0.5 seconds. From the historical speed data sequence [30km / h, 25km / h, 10km / h, 5km / h, 0km / h], the initial speed of 25km / h and the final speed of 0km / h within this time window are extracted. The speed decrease is calculated as 25km / h - 0km / h = 25km / h, and this speed decrease is used as the feature of the sudden change in motion state.

[0136] Furthermore, based on historical speed data sequences, the mean and standard deviation of the historical speed data sequences are statistically analyzed, and a speed reduction threshold is calculated based on a linear combination of the mean and standard deviation. The mean of the historical speed data sequence reflects the normal driving speed level of the target object, and the standard deviation reflects its normal speed fluctuation range. The linear combination refers to combining the mean and standard deviation according to a preset ratio to calculate the speed reduction threshold. For example, the linear combination can be in the form of mean - k × standard deviation, where k is a preset coefficient; preferably, k can be 2 or 3. The speed reduction threshold is used to determine whether there is a sudden change in the target object's motion state. If the speed reduction exceeds this threshold, a significant change in motion state is considered to exist.

[0137] For example, if the mean of the statistical historical speed data sequence is 28 km / h, the standard deviation is 5 km / h, and the linear combination coefficient k is set to 2, then the speed decrease threshold = mean - 2 × standard deviation = 28 - 2 × 5 = 18 km / h. That is, when the speed decrease exceeds 18 km / h, the target object is considered to have a significant change in motion state.

[0138] Finally, the ratio of the speed decrease magnitude to the speed decrease magnitude threshold is calculated, and this ratio is used as the second confidence level. The second confidence level is a numerical value that quantifies the credibility of a suspected sudden traffic incident based on the characteristics of abrupt changes in motion state. Its value range is constrained to 0-1. The larger the ratio of the speed decrease magnitude to the speed decrease magnitude threshold, the more obvious the change in motion state, and the higher the probability that the suspected sudden traffic incident is a real sudden traffic incident. Specifically, if the ratio of the speed decrease magnitude to the speed decrease magnitude threshold is greater than or equal to 1, the second confidence level is 1.0, indicating an extremely obvious change in motion state; if the ratio is less than or equal to 0, the second confidence level is 0, indicating no change in motion state; if the ratio is between 0 and 1, the second confidence level is equal to that ratio.

[0139] For example, if the speed reduction threshold is 18 km / h, when the speed reduction is 25 km / h, the ratio = 25 / 18 ≈ 1.38, so the second confidence level is 1.0; when the speed reduction is 15 km / h, the ratio = 15 / 18 ≈ 0.83, so the second confidence level is 0.83.

[0140] In summary, compared to existing technologies, this application simultaneously utilizes millimeter-wave radar data to detect abrupt changes in the motion state of target objects in suspected sudden traffic incidents, and generates a second confidence level based on these abrupt changes. This achieves independent cross-validation of the preliminary visual recognition results, effectively compensating for potential blind spots in visual sensors due to lighting conditions and weather, and improving the overall reliability and anti-interference capability of sudden traffic incident detection.

[0141] The fusion verification and confirmation module 13 is used to perform dual verification and judgment based on the first confidence level and the second confidence level. After the dual verification is passed, it confirms that a sudden traffic incident has occurred and generates sudden incident information including the incident type, incident location and lane occupation.

[0142] The aforementioned visual preliminary perception module 11 and radar motion detection module 12 determine whether a sudden traffic incident has occurred from a visual or radar perspective, respectively. However, either visual or radar judgment alone may result in false alarms.

[0143] To address the aforementioned issues, this application employs a fusion verification and confirmation module 13 to perform dual verification based on a first confidence level and a second confidence level. After the dual verification is passed, it confirms that a sudden traffic incident has occurred and generates incident information including the incident type, incident location, and lane occupancy.

[0144] Specifically, the fusion verification and confirmation module 13 is used for:

[0145] The first confidence level and the second confidence level are input into the fusion judgment model, and the fusion judgment model outputs the judgment result of whether the double verification is passed or failed. The fusion judgment model is a binary classification model constructed based on the joint distribution characteristics of the first confidence level and the second confidence level in historical accident data.

[0146] When the fusion judgment model outputs a judgment result that passes dual verification, a sudden traffic incident is confirmed to have occurred.

[0147] After confirming the occurrence of a traffic emergency, the event type is identified based on static scene features and motion time sequence features, the event location is determined based on the spatial location of the target object, and the occupied lane is determined based on the relative positional relationship between the target object and the lane line, generating emergency event information that includes event type, event location, and occupied lane.

[0148] In this embodiment, the first confidence level and the second confidence level are first input into the fusion judgment model. The fusion judgment model outputs a judgment result indicating whether the dual verification is passed or failed. The fusion judgment model is a binary classification model constructed based on the joint distribution characteristics of the first and second confidence levels in historical accident data. Specifically, the fusion judgment model is used to comprehensively verify the preliminary judgment results of the visual modality and the radar modality to eliminate false alarms or missed alarms that may be caused by single-modality judgment. The fusion judgment model uses the first and second confidence levels as input features and outputs a binary classification result (pass or fail), indicating whether a sudden traffic incident is confirmed.

[0149] The fusion decision model can be constructed using binary classification algorithms such as logistic regression, support vector machine, or shallow neural network. Taking logistic regression as an example, the fusion decision model is in the form of: P(event|C1,C2)=1 / (1+exp(-(w0+w1·C1+w2·C2))), where C1 is the first confidence level, C2 is the second confidence level, and w0, w1, and w2 are the parameters to be trained. When the output probability P is greater than the preset threshold, the double verification is considered successful; otherwise, it fails. The preset threshold is an empirical parameter that can be dynamically set according to the actual application scenario. It is used to balance the accuracy and recall of event detection. The preset threshold can be determined based on the performance indicators on the validation set: for example, iterating through different thresholds on the validation set, such as 0.1 to 0.9, with a step size of 0.05, calculating the accuracy and recall at each threshold, and selecting the threshold that maximizes the F1 score or meets the business requirements as the final preset threshold. Those skilled in the art can flexibly adjust it according to different tolerances for false positive and false negative rates.

[0150] For example, the training process of the fusion judgment model includes, taking logistic regression as an example: 1. Data preparation: Collect a large number of historical accident samples from the historical traffic event database. Each sample contains the first confidence level, the second confidence level and the true label of the corresponding event (for example, 1 represents a sudden traffic event and 0 represents a false alarm or no sudden traffic event). The collected sample set is randomly divided into training set, validation set and test set in a ratio of 7:1.5:1.5.

[0151] 2. Model Construction: An initial binary classification model is constructed based on the logistic regression algorithm. The input layer is 2-dimensional (corresponding to two confidence levels), and the output layer is 1-dimensional. It is mapped to 0~1 through the sigmoid function. The fusion judgment model is in the form of: P(event|C1,C2)=1 / (1+exp(-(w0+w1·C1+w2·C2))), where C1 is the first confidence level, C2 is the second confidence level, and w0, w1, and w2 are the parameters to be trained.

[0152] 3. Model Training: Using the first and second confidence scores from the training set as input features and the corresponding true labels as supervision signals, a binary classification cross-entropy loss function is employed. The optimizer is Adam, with an initial learning rate of 0.01, a batch size of 32, and 100 training epochs. After each training epoch, the loss and accuracy are calculated on the validation set. When the validation set loss no longer decreases for 10 consecutive epochs (e.g., the decrease is less than 0.001) and the accuracy no longer improves, the model is considered converged. The model parameters w0, w1, and w2 at this point are saved, training is stopped, and the pre-trained fusion decision model is obtained.

[0153] For example, assuming the first confidence level of the input is 0.9 and the second confidence level is 1.0, the output probability P = 0.98 is calculated by the trained logistic regression model, which is greater than the preset threshold of 0.5, so it is determined that the double verification has passed.

[0154] Secondly, when the fusion judgment model outputs a judgment result indicating that both verifications have passed, a sudden traffic incident is confirmed. Specifically, passing dual verification means that both the first confidence level (preliminary confidence level of the visual modality) and the second confidence level (confidence level of motion abrupt change in the radar modality) meet the collaborative verification requirements, effectively eliminating possible misjudgments from a single sensor, and thus determining that a sudden traffic incident has occurred. Conversely, if dual verification fails, it is determined as a suspected misjudgment, and the occurrence of a sudden traffic incident is not confirmed.

[0155] Finally, after confirming the occurrence of a sudden traffic incident, the event type is identified based on static scene features and motion temporal features, the event location is determined based on the spatial position of the target object, and the occupied lane is determined based on the relative positional relationship between the target object and the lane lines, generating sudden incident information containing the event type, event location, and occupied lane. Specifically, event type identification can combine static scene features extracted from the slow path network branch and motion temporal features extracted from the fast path network branch. For example, based on the contour deformation after a vehicle collision, the shape of obstacles, the distribution of debris, and the sudden drop in speed, abrupt change in trajectory, and sharp reduction in distance before the collision, the specific type of the event is output, such as vehicle collision, lane closure due to a breakdown, or fallen obstacle.

[0156] Specifically, the location of an event can be determined based on the spatial position of the target object in the image, combined with the distance and azimuth measured by millimeter-wave radar, as well as the GPS coordinates and road orientation of the first intelligent signboard itself. The precise latitude and longitude or road mileage of the event can be calculated through coordinate transformation, such as 500 meters east to west on XX Avenue.

[0157] Specifically, lane occupancy can be determined by the relative position of the target object and the lane line in the image: if the center of the target object is located within a lane line and its boundary does not exceed the lane line, then the target is determined to occupy the corresponding lane, and the lane number can be output, such as occupying lane 3, or the lane description.

[0158] Specifically, the event type, event location, and occupied lane identified above are integrated into structured emergency event information, such as: {Event type: vehicle collision, Event location: 500 meters east to west on XX Avenue, Occupied lane: Lane 3}, which is then broadcast to other smart signs.

[0159] In summary, compared to existing technologies, this application employs dual verification based on a first confidence level and a second confidence level. Upon successful dual verification, a sudden traffic incident is confirmed, and incident information including the incident type, location, and occupied lane is generated. This achieves multimodal collaborative verification, improving the accuracy and reliability of sudden traffic incident detection. Simultaneously, it provides precise incident information support for subsequent distributed collaborative guidance, effectively avoiding invalid warnings or guidance errors caused by single-modal misjudgments.

[0160] The event information broadcasting module 14 is used to broadcast emergency information to other smart signs in the smart sign array.

[0161] In this embodiment, the first smart signboard, based on the event information broadcasting module 14, broadcasts emergency event information to all smart signs within a certain range via vehicle-to-everything (V2X), 5G, Internet of Things (IoT), and Dedicated Short Range Communication (DSRC).

[0162] The broadcast range can be set according to the road topology, such as smart signs within 5 kilometers before and after the event point. The broadcast messages can use a standard data format to ensure that each smart sign can be quickly parsed and processed.

[0163] For example, the intelligent sign array consists of 10 intelligent signs evenly deployed at 500-meter intervals along the K3+000m to K3+900m section of XX Avenue, forming a continuously covered traffic sensing network. When the visual preliminary perception module of any intelligent sign in the array identifies a suspected sudden traffic incident from the collected real-time video stream of the road, that intelligent sign automatically becomes the "first intelligent sign" in the incident handling process. For instance, if the visual preliminary perception module of the intelligent sign located at K3+500m is the first to identify a rear-end collision 200 meters ahead, then that intelligent sign is identified as the first intelligent sign. After the first intelligent sign completes the identification of the suspected sudden traffic incident, its radar motion detection module synchronously calls millimeter-wave radar data to detect the abrupt change characteristics of the target object's motion state and generates a second confidence level. Then, the fusion verification and confirmation module performs a double verification judgment. After confirming that a sudden traffic incident has occurred, it generates incident information including the incident type, incident location, and occupied lane. Finally, the incident information is broadcast to the remaining 9 intelligent signs in the intelligent sign array through the incident information broadcasting module. During this process, the role of the first smart sign is dynamically determined by the trigger of a suspected traffic emergency. Any smart sign in the smart sign array can become the first smart sign, and its identity is unique in one event handling process. After the event is handled, all smart signs return to their initial state, waiting for the next suspected traffic emergency to trigger to re-determine the first smart sign.

[0164] The collaborative priority calculation module 15 is used for each smart sign in the smart sign array to calculate the collaborative guidance priority factor based on its own location information and emergency information through the collaborative guidance priority calculation model deployed on the edge.

[0165] The guidance priority of smart signs in different locations varies depending on their distance from the sudden traffic incident and their relationship with the occupied lane: the closer the sign is to the incident and the higher the correlation between the sign and the occupied lane, the higher the guidance priority, and guidance actions should be performed first to quickly evacuate surrounding vehicles; the farther the sign is from the incident and the lower the correlation, the lower the guidance priority, and auxiliary guidance actions can be performed.

[0166] To address the aforementioned issues, this application utilizes a collaborative priority calculation module 15 to combine the location information of each smart sign in the smart sign array with information on unforeseen events. Through a collaborative guidance priority calculation model deployed on the edge, a collaborative guidance priority factor is calculated, providing a basis for matching differentiated guidance strategies.

[0167] Specifically, the collaborative priority calculation module 15 is used for:

[0168] Calculate the distance along the road between the location of each smart sign and the location of the event in the emergency information to obtain the distance factor;

[0169] Determine the lane association relationship between the lane where each smart sign is located and the lane occupied in the emergency information, wherein the lane association relationship includes at least the same lane, adjacent lane or opposite lane;

[0170] The distance factor and lane association relationship are input into a pre-trained collaborative guidance priority calculation model deployed on the edge computing unit of the end side, and the collaborative guidance priority factor is output.

[0171] The pre-training process of the collaborative guidance priority calculation model includes:

[0172] Multiple historical emergency samples are collected from the historical traffic incident database. Each historical emergency sample includes the location of the historical incident, the lanes occupied in the past, and the historical response data of multiple smart signs in the historical emergency.

[0173] For each historical emergency sample, for each smart sign participating in the response, calculate the historical distance factor between the smart sign and the location of the historical event, determine the historical lane association between the lane where the smart sign is located and the historically occupied lane, and construct a feature sample set;

[0174] Based on the actual response level of the smart sign in historical emergencies, and by labeling the corresponding historical priority factors according to the actual response level, a tag sample set is constructed.

[0175] An initial regression model is constructed based on the regression algorithm. Then, the initial regression model is trained in a supervised manner using the feature sample set and the label sample set until it is verified to converge, thus obtaining a pre-trained collaborative guidance priority calculation model.

[0176] In this embodiment, the road distance between the location of each smart sign and the location of the event in the emergency information is first calculated to obtain a distance factor. The road distance refers to the actual distance along the road between the location of the smart sign and the location of the event in the emergency information (not the straight-line distance), to more accurately reflect the distance a vehicle needs to travel to reach the emergency traffic incident point, and the road distance is used as the distance factor.

[0177] For example, if a smart sign is located at K3+400 meters on XX Avenue and the event location is located at K3+500 meters on XX Avenue, then the distance along the road is 100 meters.

[0178] Secondly, the lane association relationship between each smart sign's lane and the lanes occupied in the emergency incident information is determined. This lane association relationship includes at least same-lane, adjacent-lane, or opposite-lane relationships. Specifically, the lane association relationship refers to the relationship between the smart sign's lane and the lane occupied by the incident, characterizing the targeted guidance of the smart sign: same-lane means the smart sign's lane and the incident-occupied lane are in the same lane; adjacent-lane means the smart sign's lane and the incident-occupied lane are adjacent (one on each side); opposite-lane means the smart sign's lane and the incident-occupied lane travel in opposite directions. In detail, the lane association relationship is determined by combining the occupied lanes in the emergency incident information with the lane where each smart sign is located.

[0179] For example, if the event occupies lane 3 from east to west, and a smart sign is located in lane 3 from east to west, the relationship is the same lane; if it is located in lane 2 or 4 from east to west, the relationship is adjacent lane; if it is located in a lane from west to east, the relationship is opposite lane.

[0180] Finally, the distance factor and lane association relationship are input into a pre-trained collaborative guidance priority calculation model deployed on the edge computing unit, and the collaborative guidance priority factor is output. Specifically, a lightweight collaborative guidance priority calculation model is deployed on each smart sign. This model takes the distance factor and lane association relationship as input and outputs a collaborative guidance priority factor between 0 and 1. The collaborative guidance priority factor is a numerical value that quantifies the guidance priority of the smart sign, ranging from 0 to 1. The larger the value, the higher the guidance priority.

[0181] For example, a smart sign in the same lane with a distance factor of 100 meters outputs a collaborative guidance priority factor of 0.95 through a pre-trained collaborative guidance priority calculation model deployed on the edge computing unit; a smart sign in an adjacent lane with a distance factor of 600 meters outputs a collaborative guidance priority factor of 0.85 through the pre-trained collaborative guidance priority calculation model deployed on the edge computing unit.

[0182] This involves collecting multiple historical emergency samples from a historical traffic incident database. Each historical emergency sample includes the location of the historical incident, the lanes occupied in the past, and historical response data from multiple smart signs during the historical emergency. Specifically, the historical traffic incident database refers to a database storing a large amount of historical emergency traffic data; the historical emergency sample refers to a historical emergency traffic incident selected from the historical traffic incident database that contains complete information; and the historical response data refers to the actual response actions and response times performed by each smart sign during the historical emergency.

[0183] For example, 5,000 historical emergency samples are collected from the historical traffic incident database. Each historical emergency sample includes the location of the historical incident, the lane occupied in the past, and the historical response data of the smart signs involved in the response, such as response time and guidance action type.

[0184] For each historical emergency sample, for each smart sign participating in the response, a historical distance factor between the smart sign and the historical event location is calculated. This determines the historical lane association between the lane where the smart sign is located and the previously occupied lane, thus constructing a feature sample set. Specifically, the historical distance factor is the distance factor calculated based on the roadside distance between the smart sign and the historical event location, using the same calculation method as the current distance factor. The historical lane association refers to the association between the lane where the smart sign is located and the previously occupied lane, with the classification consistent with the current lane association. The feature sample set is a set of samples composed of the historical distance factor and the historical lane association, used for model training.

[0185] This involves constructing a tag sample set based on the actual response level of the smart signage during historical emergencies, and labeling each sample with a corresponding historical priority factor. Specifically, based on historical response data, each sample is labeled with a historical priority factor as a supervisory tag. For example, if a smart signage is actually activated with high priority during an event, such as displaying an accident warning and flashing, its historical priority factor can be marked as 0.9; if it only displays a general prompt, its historical priority factor can be marked as 0.5; and if it does not respond, its historical priority factor can be marked as 0.1.

[0186] The process involves constructing an initial regression model based on a regression algorithm, followed by supervised training using a feature sample set and a label sample set until convergence is verified, resulting in a pre-trained collaborative guidance priority calculation model. For example, the training process of the collaborative guidance priority calculation model includes the following steps: 1. Data preparation: The feature sample set and label sample set are randomly divided into a training set, a validation set, and a test set in a ratio of 7:1.5:1.5. The training set is used for model parameter updates, the validation set is used for hyperparameter tuning and convergence judgment, and the test set is used for final model performance evaluation.

[0187] 2. Model Construction: Gradient Boosting Decision Tree (GBDT) is used as the regression algorithm to construct the initial regression model. The main model parameters are configured as follows: number of trees set to 100, maximum depth set to 5, learning rate set to 0.1, subsampling ratio set to 0.8, and mean squared error (MSE) used as the loss function. The input layer receives two features: historical distance factor and historical lane correlation. The output layer is a continuous value, namely the predicted collaborative guidance priority factor.

[0188] 3. Model Training: Using historical distance factors and historical lane associations from the training set as input features, and corresponding historical priority factors as supervision labels, mean squared error is used as the loss function. The model is trained iteratively using a gradient boosting algorithm. In each iteration, a new decision tree is added to fit the residuals of the current model. During training, after every 5 iterations, the mean squared error is calculated on the validation set. When the validation set loss no longer decreases for 10 consecutive iterations (e.g., the decrease is less than 0.001), the model is considered to have converged, training is stopped, and the model parameters at this point are saved, resulting in a pre-trained collaborative guidance priority calculation model.

[0189] In summary, compared to existing technologies, each smart sign in the smart sign array of this application calculates a collaborative guidance priority factor based on its own location information and emergency event information through a collaborative guidance priority calculation model deployed on the edge. This achieves a distributed, adaptive collaborative response for the smart sign array, providing reliable and accurate technical support for precise and efficient guidance after sudden traffic incidents.

[0190] The differentiated guidance execution module 16 is used for each smart sign to match the corresponding guidance strategy based on the collaborative guidance priority factor calculated by itself, and to execute differentiated traffic guidance actions.

[0191] The aforementioned collaborative priority calculation module 15 calculates the collaborative guidance priority factor for each smart sign in the smart sign array. This application further utilizes a differentiated guidance execution module 16 to enable each smart sign to match its calculated collaborative guidance priority factor with a corresponding guidance strategy and execute differentiated traffic guidance actions. This avoids driver confusion caused by all smart signs issuing the same instructions simultaneously, while concentrating resources on effective guidance at key locations.

[0192] Specifically, the differentiated boot execution module 16 is used for:

[0193] A mapping table between collaborative guidance priority factor ranges and guidance strategies is pre-constructed. The mapping table contains a one-to-one correspondence between multiple collaborative guidance priority factor ranges and multiple guidance strategies.

[0194] Each smart sign determines the target interval of the collaborative guidance priority factor in the mapping table based on its own calculated collaborative guidance priority factor, and matches the guidance strategy corresponding to the target interval from the mapping table;

[0195] Based on the matched guidance strategy, execute the traffic guidance actions corresponding to the guidance strategy.

[0196] In this embodiment, a mapping table between collaborative guidance priority factor intervals and guidance strategies is first pre-constructed. This mapping table contains a one-to-one correspondence between multiple collaborative guidance priority factor intervals and various guidance strategies. Specifically, a collaborative guidance priority factor interval refers to dividing the value range (0~1) of the collaborative guidance priority factor into multiple consecutive intervals, such as 0~0.5, 0.5~0.8, 0.8~1.0, etc.; a guidance strategy refers to a specific plan for guiding traffic based on the guidance priority, which may include guidance display content, display frequency, voice prompts, and whether to link with traffic lights; the mapping table is a table used to store the correspondence between collaborative guidance priority factor intervals and guidance strategies, used for the intelligent signage to quickly match guidance strategies. The mapping table is pre-constructed and can be stored in the local memory of the intelligent signage.

[0197] For example, the mapping table can be constructed as follows: The collaborative guidance priority factor range [0.8, 1.0] represents high priority: the guidance strategy is to display "Accident ahead, please slow down," flash warning lights, display frequency 1 time / second, and loop the voice prompt. The collaborative guidance priority factor range [0.5, 0.8) represents medium priority: the guidance strategy can be to display "Caution: Accident ahead, detour recommended," display frequency 1 time / 2 seconds, and play a voice prompt every 10 seconds. The collaborative guidance priority factor range [0.0, 0.5) represents low priority: the guidance strategy can be to not display accident information or display "Sudden traffic incident ahead, please drive carefully," display frequency 1 time / 5 seconds, and not play a voice prompt.

[0198] Secondly, each smart sign determines the target interval of its calculated collaborative guidance priority factor in the mapping table, and matches the guidance strategy corresponding to the target interval from the mapping table. For example, if a smart sign has a collaborative guidance priority factor of 0.95, its target interval in the mapping table is the high-priority interval [0.8, 1.0], and it matches the guidance strategy corresponding to the high priority interval in the mapping table.

[0199] Finally, based on the matched guidance strategy, the intelligent signage executes the corresponding traffic guidance actions. Specifically, the intelligent signage executes specific actions according to the guidance strategy: for example, if the guidance strategy requires displaying specific text and flashing lights, the signage controls the LED display to show "Accident ahead, please slow down," while simultaneously illuminating surrounding warning lights to enhance the warning effect.

[0200] Traffic guidance actions refer to the specific operations performed by intelligent signs according to guidance strategies to guide traffic, including displaying guidance information, playing voice prompts, and linking with traffic lights; differentiated traffic guidance actions refer to intelligent signs of different priorities performing different traffic guidance actions to achieve coordinated guidance and avoid confusion.

[0201] In summary, compared to existing technologies, each intelligent sign in this application matches a corresponding guidance strategy based on its own calculated collaborative guidance priority factor and executes differentiated traffic guidance actions. This achieves distributed, precise, and collaborative response of the intelligent sign array in sudden traffic incidents, effectively avoiding redundancy or conflict in guidance information, optimizing the allocation efficiency of guidance resources, and thus improving the overall effectiveness and safety of road traffic guidance.

[0202] In summary, the embodiments of this application have at least the following technical effects:

[0203] Compared to existing technologies, this application first uses a first smart sign in an array of smart signs deployed on the roadside to identify suspected traffic emergencies from real-time road video streams through a pre-trained lightweight dual-path network and generate a first confidence score. This enables rapid initial perception of traffic emergencies at the edge, providing a reliable preliminary screening basis for subsequent fusion verification and effectively reducing the latency and bandwidth pressure of centralized processing.

[0204] Secondly, this application simultaneously calls millimeter-wave radar data to detect the abrupt change characteristics of the motion state of the target object in a suspected sudden traffic incident and generates a second confidence level, realizing independent cross-validation of the preliminary visual recognition results. This effectively compensates for the blind spots of visual sensors in environments such as lighting and weather, and improves the overall reliability and anti-interference capability of event detection.

[0205] Furthermore, this application uses a dual verification method based on a first confidence level and a second confidence level. After the dual verification is passed, it confirms that a sudden traffic incident has occurred and generates incident information including the incident type, incident location, and lane occupancy. This achieves multimodal collaborative verification, significantly improves the accuracy and reliability of incident detection, and provides accurate incident information support for subsequent distributed collaborative guidance.

[0206] Furthermore, this application broadcasts emergency information to other smart signs in the smart sign array, enabling rapid sharing of event information within the array and laying a data foundation for collaborative response among the signs.

[0207] Finally, each smart sign in the smart sign array calculates a collaborative guidance priority factor based on its own location information and emergency information through a collaborative guidance priority calculation model deployed on the edge, and executes differentiated traffic guidance actions according to the corresponding guidance strategy matched by the factor. This realizes the distributed and precise collaborative response of the smart sign array, effectively avoids redundancy or conflict of guidance information, and optimizes the configuration efficiency of guidance resources.

[0208] Through the above technical solutions, this application constructs a complete closed loop from real-time edge perception, multimodal fusion verification to distributed collaborative guidance, effectively solving the problems of high detection latency, insufficient accuracy and lack of coordination in guidance in the existing technology, improving the response speed and guidance effect of sudden traffic events, and ultimately enhancing the overall safety and traffic efficiency of road traffic.

[0209] Example 2, as Figure 2 As shown, this embodiment of the invention also provides an intelligent traffic guidance method based on intelligent signage, including:

[0210] The first smart sign in the smart sign array deployed on the roadside identifies suspected traffic emergencies from the collected real-time video stream of the road through a pre-trained lightweight dual-path network and generates a first confidence level.

[0211] Simultaneously call millimeter-wave radar data to detect abrupt changes in the motion state of target objects in suspected sudden traffic incidents, and generate a second confidence level based on the abrupt changes in motion state.

[0212] A dual verification decision is made based on the first confidence level and the second confidence level. After the dual verification is passed, it is confirmed that a traffic emergency has occurred, and emergency information including the event type, event location and lane occupancy is generated.

[0213] Broadcast emergency information to all other smart signs in the smart sign array;

[0214] Each smart sign in the smart sign array calculates a collaborative guidance priority factor based on its own location information and emergency event information through a collaborative guidance priority calculation model deployed on the edge.

[0215] Each smart sign matches a corresponding guidance strategy based on its own calculated collaborative guidance priority factor and executes differentiated traffic guidance actions.

[0216] Specifically, the first smart sign in the roadside smart sign array identifies suspected traffic emergencies from the collected real-time road video stream through a pre-trained lightweight dual-path network and generates a first confidence score, including:

[0217] The collected real-time video streams of the road are input into the slow path network branch and the fast path network branch of the pre-trained lightweight dual-path network, respectively.

[0218] Static scene features in video frames are extracted using a slow path network branch at the first sampling frame rate. These static scene features include at least lane line positions, vehicle appearance, and obstacle positions.

[0219] Motion temporal features in a video frame sequence are extracted using a fast path network branch at a second sampling frame rate, wherein the second sampling frame rate is higher than the first sampling frame rate, and the motion temporal features include at least vehicle speed changes, vehicle trajectory deviations, and changes in relative distances between vehicles.

[0220] The motion time sequence features and static scene features are fused to obtain fused spatiotemporal features. Suspected sudden traffic events are then identified based on the fused spatiotemporal features, and the corresponding first confidence score is generated.

[0221] Furthermore, the construction process of the lightweight dual-path network includes:

[0222] A dual-path network architecture is constructed, which includes a slow path network branch, a fast path network branch, and a lateral connection layer. The lateral connection layer is used to connect the slow path network branch and the fast path network branch. The slow path network branch and the fast path network branch each contain multiple convolutional layers. Each convolutional layer contains multiple feature channels. Each feature channel is used to extract different dimensional features of the input data.

[0223] Video stream samples containing historical traffic accidents were collected from a historical traffic monitoring video database, and the video stream samples were labeled with the type of emergency to obtain a training dataset and a corresponding supervision label set.

[0224] The dual-path network architecture is trained in a supervised manner using the training dataset and the supervised label set until it is verified to converge, thus obtaining the pre-trained dual-path network.

[0225] A lightweight dual-path network is obtained by performing a lightweight modification on the pre-trained dual-path network.

[0226] A lightweight dual-path network is deployed in the edge computing unit of the first smart sign, wherein the slow path network branch is deployed in the main processor of the edge computing unit, and the fast path network branch is deployed in the neural network acceleration unit of the edge computing unit.

[0227] Furthermore, the pre-trained dual-path network is modified to be lightweight, resulting in a lightweight dual-path network, including:

[0228] In the training dataset, the image regions containing lane lines and vehicles in each video frame are labeled as regions of interest, and the image regions other than the regions of interest are labeled as background regions.

[0229] The weight parameters of each convolutional layer in the pre-trained dual-path network are statistically analyzed, and the first floating-point number distribution range of the weight parameters corresponding to the region of interest and the second floating-point number distribution range of the weight parameters corresponding to the background region are obtained respectively.

[0230] Based on the first floating-point number distribution range, the first quantization parameter corresponding to the region of interest is calculated, and based on the second floating-point number distribution range, the second quantization parameter corresponding to the background region is calculated, wherein the quantization bit width corresponding to the first quantization parameter is higher than the quantization bit width corresponding to the second quantization parameter.

[0231] The first quantization parameter is used to quantize the weight parameters corresponding to the region of interest, converting the weight parameters corresponding to the region of interest from floating-point format to low-bit-width integer format.

[0232] The second quantization parameter is used to quantize the weight parameter corresponding to the background region, converting the weight parameter corresponding to the background region from floating-point format to low-bit-width integer format.

[0233] The weight parameters corresponding to the region of interest after quantization transformation and the weight parameters corresponding to the background region after quantization transformation are merged and stored according to their respective region identifiers to obtain a lightweight dual-path network.

[0234] Furthermore, the motion time-series features and static scene features are fused to obtain fused spatiotemporal features. These fused spatiotemporal features are then used to identify suspected sudden traffic incidents and generate corresponding first confidence scores, including:

[0235] The motion temporal features are upsampled in the time dimension to obtain time-aligned motion features that are aligned with the static scene features in the time dimension.

[0236] The time-aligned motion features and static scene features are concatenated to construct multiple spatiotemporal feature pairs. Among them, the spatiotemporal feature pairs include at least the lane-speed feature pair consisting of the lane line position and the corresponding vehicle speed change, the obstacle-trajectory feature pair consisting of the obstacle position and the corresponding vehicle trajectory offset, and the appearance-trajectory feature pair consisting of the vehicle appearance and the corresponding vehicle trajectory offset.

[0237] For each spatiotemporal feature pair, calculate the matching degree between static scene features and time-aligned motion features in the spatiotemporal feature pair to obtain multiple feature matching degrees;

[0238] The minimum value of multiple feature matching degrees is selected as the association consistency index. When the association consistency index is lower than the preset consistency threshold, it is determined that there is an abnormal association pattern, and the scene corresponding to the abnormal association pattern is identified as a suspected sudden traffic event.

[0239] Based on the difference between the correlation consistency index and the preset consistency threshold, the first confidence level corresponding to the suspected sudden traffic incident is generated.

[0240] Specifically, for each spatiotemporal feature pair, the matching degree between the static scene features and the time-aligned motion features in the spatiotemporal feature pair is calculated to obtain multiple feature matching degrees, including:

[0241] For lane-speed feature pairs, the normalized lane-speed matching degree is calculated based on the ratio of the vehicle's current lateral offset to the lane's allowed lateral offset range, and the ratio of the vehicle's current speed to the lane's historical average speed.

[0242] For obstacle-trajectory feature pairs, the normalized obstacle-trajectory matching degree is calculated based on the ratio of the shortest distance between the vehicle's current trajectory and the obstacle's position to the preset safe distance, and the ratio of the angle between the vehicle's current speed direction and the obstacle's direction to the maximum allowable angle.

[0243] For appearance-trajectory feature pairs, the normalized appearance-trajectory matching degree is calculated based on the similarity score between the current appearance features and historical appearance features of the vehicle, and the ratio of the deviation between the current trajectory and historical trajectory to the maximum allowable deviation.

[0244] Specifically, millimeter-wave radar data is simultaneously invoked to detect abrupt changes in the motion state of target objects in suspected sudden traffic incidents, and a second confidence level is generated based on these abrupt changes, including:

[0245] Based on the spatial location of the target object in a suspected traffic emergency, millimeter-wave radar data of the corresponding area is retrieved;

[0246] Extract the velocity sequence of the target object within a continuous time window from the millimeter-wave radar data, and obtain the historical velocity data sequence of the target object from the historical trajectory data;

[0247] Based on the velocity sequence and historical velocity data sequence, the velocity decrease of the target object within a preset time window is calculated, and the velocity decrease is used as a feature of abrupt change in motion state.

[0248] Based on historical speed data series, the mean and standard deviation of the historical speed data series are statistically analyzed, and the threshold for the speed decrease is calculated based on the linear combination of the mean and standard deviation.

[0249] Calculate the ratio of the speed decrease to the speed decrease threshold, and use this ratio as the second confidence level.

[0250] Specifically, a dual verification process is performed based on a first confidence level and a second confidence level. Upon successful dual verification, a traffic emergency is confirmed, and emergency information including the event type, location, and lane occupancy is generated, including:

[0251] The first confidence level and the second confidence level are input into the fusion judgment model, and the fusion judgment model outputs the judgment result of whether the double verification is passed or failed. The fusion judgment model is a binary classification model constructed based on the joint distribution characteristics of the first confidence level and the second confidence level in historical accident data.

[0252] When the fusion judgment model outputs a judgment result that passes dual verification, a sudden traffic incident is confirmed to have occurred.

[0253] After confirming the occurrence of a traffic emergency, the event type is identified based on static scene features and motion time sequence features, the event location is determined based on the spatial location of the target object, and the occupied lane is determined based on the relative positional relationship between the target object and the lane line, generating emergency event information that includes event type, event location, and occupied lane.

[0254] Specifically, each smart sign in the smart sign array, based on its own location information and emergency event information, calculates a collaborative guidance priority factor through a collaborative guidance priority calculation model deployed on the edge, including:

[0255] Calculate the distance along the road between the location of each smart sign and the location of the event in the emergency information to obtain the distance factor;

[0256] Determine the lane association relationship between the lane where each smart sign is located and the lane occupied in the emergency information, wherein the lane association relationship includes at least the same lane, adjacent lane or opposite lane;

[0257] The distance factor and lane association relationship are input into a pre-trained collaborative guidance priority calculation model deployed on the edge computing unit of the end side, and the collaborative guidance priority factor is output.

[0258] The pre-training process of the collaborative guidance priority calculation model includes:

[0259] Multiple historical emergency samples are collected from the historical traffic incident database. Each historical emergency sample includes the location of the historical incident, the lanes occupied in the past, and the historical response data of multiple smart signs in the historical emergency.

[0260] For each historical emergency sample, for each smart sign participating in the response, calculate the historical distance factor between the smart sign and the location of the historical event, determine the historical lane association between the lane where the smart sign is located and the historically occupied lane, and construct a feature sample set;

[0261] Based on the actual response level of the smart sign in historical emergencies, and by labeling the corresponding historical priority factors according to the actual response level, a tag sample set is constructed.

[0262] An initial regression model is constructed based on the regression algorithm. Then, the initial regression model is trained in a supervised manner using the feature sample set and the label sample set until it is verified to converge, thus obtaining a pre-trained collaborative guidance priority calculation model.

[0263] Specifically, each smart sign matches a corresponding guidance strategy based on its calculated collaborative guidance priority factor and executes differentiated traffic guidance actions, including:

[0264] A mapping table between collaborative guidance priority factor ranges and guidance strategies is pre-constructed. The mapping table contains a one-to-one correspondence between multiple collaborative guidance priority factor ranges and multiple guidance strategies.

[0265] Each smart sign determines the target interval of the collaborative guidance priority factor in the mapping table based on its own calculated collaborative guidance priority factor, and matches the guidance strategy corresponding to the target interval from the mapping table;

[0266] Based on the matched guidance strategy, execute the traffic guidance actions corresponding to the guidance strategy.

[0267] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0268] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0269] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0270] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0271] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0272] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.

[0273] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. An intelligent traffic guidance system based on intelligent signage, characterized in that, The system includes: The visual preliminary perception module is used to identify suspected sudden traffic events from the collected real-time video stream of the road by the first smart sign in the smart sign array deployed on the side of the road through a pre-trained lightweight dual-path network and generate a first confidence score. The smart sign array is a traffic perception network formed by multiple smart signs evenly deployed at a preset interval. The first smart sign is the smart sign in the smart sign array that first identifies the suspected sudden traffic event. The radar motion detection module is used to synchronously call millimeter-wave radar data, detect the abrupt change characteristics of the motion state of the target object in the suspected sudden traffic incident, and generate a second confidence level based on the abrupt change characteristics of the motion state. The fusion verification and confirmation module is used to perform dual verification and judgment based on the first confidence level and the second confidence level. After the dual verification is passed, it confirms that a sudden traffic incident has occurred and generates sudden incident information including the incident type, incident location and lane occupation. An event information broadcasting module is used to broadcast the emergency event information to other smart signs in the smart sign array; The collaborative priority calculation module is used for each smart sign in the smart sign array to calculate the collaborative guidance priority factor based on its own location information and the emergency information through a collaborative guidance priority calculation model deployed on the edge. The differentiated guidance execution module is used for each smart sign to match the corresponding guidance strategy based on the collaborative guidance priority factor calculated by itself, and to execute differentiated traffic guidance actions. The construction process of the lightweight dual-path network includes lightweighting the pre-trained dual-path network to obtain the lightweight dual-path network, including: In the training dataset, the image region containing lane lines and vehicles in each video frame image is labeled as the region of interest, and the image region other than the region of interest is labeled as the background region. The weight parameters of each convolutional layer in the pre-trained dual-path network are statistically analyzed, and the first floating-point number distribution range of the weight parameters corresponding to the region of interest and the second floating-point number distribution range of the weight parameters corresponding to the background region are obtained respectively. Based on the first floating-point number distribution range, a first quantization parameter corresponding to the region of interest is calculated, and based on the second floating-point number distribution range, a second quantization parameter corresponding to the background region is calculated, wherein the quantization bit width corresponding to the first quantization parameter is higher than the quantization bit width corresponding to the second quantization parameter. The first quantization parameter is used to quantize and convert the weight parameter corresponding to the region of interest, converting the weight parameter corresponding to the region of interest from floating-point format to low-bit-width integer format; The second quantization parameter is used to quantize and convert the weight parameter corresponding to the background region from floating-point format to low-bit-width integer format. The weight parameters corresponding to the region of interest after quantization and the weight parameters corresponding to the background region after quantization are merged and stored according to their respective region identifiers to obtain a lightweight dual-path network.

2. The intelligent traffic guidance system based on intelligent signage according to claim 1, characterized in that, The preliminary visual perception module is specifically used for: The collected real-time video streams of the road are input into the slow path network branch and the fast path network branch of the pre-trained lightweight dual-path network, respectively. Static scene features in video frames are extracted through the slow path network branch at a first sampling frame rate, wherein the static scene features include at least lane line positions, vehicle appearance, and obstacle positions. Motion temporal features in the video frame sequence are extracted through the fast path network branch at a second sampling frame rate, wherein the second sampling frame rate is higher than the first sampling frame rate, and the motion temporal features include at least vehicle speed changes, vehicle trajectory deviations, and changes in relative distances between vehicles. The motion time sequence features and the static scene features are fused to obtain fused spatiotemporal features. Suspected sudden traffic events are then identified based on the fused spatiotemporal features, and a corresponding first confidence level is generated.

3. The intelligent traffic guidance system based on intelligent signage according to claim 2, characterized in that, The construction process of a lightweight dual-path network includes: A dual-path network architecture is constructed, which includes a slow path network branch, a fast path network branch, and a lateral connection layer. The lateral connection layer is used to connect the slow path network branch and the fast path network branch. The slow path network branch and the fast path network branch each contain multiple convolutional layers. Each convolutional layer contains multiple feature channels. Each feature channel is used to extract different dimensional features of the input data. Video stream samples containing historical traffic accidents are collected from a historical traffic monitoring video database, and the video stream samples are labeled with the type of sudden event to obtain a training dataset and a corresponding supervision label set. The dual-path network architecture is trained in a supervised manner using the training dataset and the supervised label set until convergence is verified, thus obtaining a pre-trained dual-path network. A lightweight dual-path network is obtained by performing a lightweight modification on the pre-trained dual-path network. The lightweight dual-path network is deployed in the edge computing unit of the first smart sign, wherein the slow path network branch is deployed in the main processor of the edge computing unit, and the fast path network branch is deployed in the neural network acceleration unit of the edge computing unit.

4. The intelligent traffic guidance system based on intelligent signage according to claim 2, characterized in that, The motion time-series features and the static scene features are fused to obtain fused spatiotemporal features. Based on these fused spatiotemporal features, suspected sudden traffic events are identified, and a corresponding first confidence level is generated, including: The motion temporal features are upsampled in the time dimension to obtain time-aligned motion features that are aligned with the static scene features in the time dimension. The time-aligned motion features and the static scene features are concatenated to construct multiple spatiotemporal feature pairs. The spatiotemporal feature pairs include at least a lane-speed feature pair consisting of lane line position and corresponding vehicle speed change, an obstacle-trajectory feature pair consisting of obstacle position and corresponding vehicle trajectory offset, and an appearance-trajectory feature pair consisting of vehicle appearance and corresponding vehicle trajectory offset. For each spatiotemporal feature pair, calculate the matching degree between the static scene features and the time-aligned motion features in the spatiotemporal feature pair to obtain multiple feature matching degrees; The minimum value is selected from multiple feature matching degrees as the association consistency index. When the association consistency index is lower than the preset consistency threshold, it is determined that there is an abnormal association pattern, and the scene corresponding to the abnormal association pattern is identified as a suspected sudden traffic event. Based on the difference between the correlation consistency index and the preset consistency threshold, a first confidence level is generated corresponding to the suspected sudden traffic incident. Specifically, for each spatiotemporal feature pair, the matching degree between the static scene features and the time-aligned motion features in the spatiotemporal feature pair is calculated to obtain multiple feature matching degrees, including: For lane-speed feature pairs, the normalized lane-speed matching degree is calculated based on the ratio of the vehicle's current lateral offset to the lane's allowed lateral offset range, and the ratio of the vehicle's current speed to the lane's historical average speed. For obstacle-trajectory feature pairs, the normalized obstacle-trajectory matching degree is calculated based on the ratio of the shortest distance between the vehicle's current trajectory and the obstacle's position to the preset safe distance, and the ratio of the angle between the vehicle's current speed direction and the obstacle's direction to the maximum allowable angle. For appearance-trajectory feature pairs, the normalized appearance-trajectory matching degree is calculated based on the similarity score between the current appearance features and historical appearance features of the vehicle, and the ratio of the deviation between the current trajectory and historical trajectory to the maximum allowable deviation.

5. The intelligent traffic guidance system based on intelligent signage according to claim 1, characterized in that, The radar motion detection module is specifically used for: Based on the spatial location of the target object in the suspected traffic emergency, millimeter-wave radar data of the corresponding area is retrieved; Extract the velocity sequence of the target object within a continuous time window from the millimeter-wave radar data, and obtain the historical velocity data sequence of the target object from the historical trajectory data; Based on the velocity sequence and the historical velocity data sequence, the velocity decrease of the target object within a preset time window is calculated, and the velocity decrease is used as a feature of a sudden change in motion state. Based on the historical speed data sequence, the mean and standard deviation of the historical speed data sequence are statistically analyzed, and the speed decrease threshold is calculated based on the linear combination of the mean and the standard deviation. Calculate the ratio of the speed decrease magnitude to the speed decrease magnitude threshold, and use the ratio as the second confidence level.

6. The intelligent traffic guidance system based on intelligent signage according to claim 1, characterized in that, The fusion verification and confirmation module is specifically used for: The first confidence level and the second confidence level are input into the fusion judgment model, and the fusion judgment model outputs the judgment result of whether the double verification passes or fails. The fusion judgment model is a binary classification model constructed based on the joint distribution characteristics of the first confidence level and the second confidence level in historical accident data. When the fusion judgment model outputs a judgment result that passes dual verification, a sudden traffic incident is confirmed to have occurred. After confirming the occurrence of a traffic emergency, the event type is identified based on static scene features and motion time sequence features, the event location is determined based on the spatial location of the target object, and the occupied lane is determined based on the relative positional relationship between the target object and the lane line, generating emergency event information that includes event type, event location, and occupied lane.

7. The intelligent traffic guidance system based on intelligent signage according to claim 1, characterized in that, The collaborative priority calculation module is specifically used for: Calculate the road distance between the location of each smart sign and the location of the event in the emergency information to obtain the distance factor; Determine the lane association relationship between the lane where each smart sign is located and the lane occupied in the emergency information, wherein the lane association relationship includes at least the same lane, adjacent lane or opposite lane; The distance factor and the lane association relationship are input into a pre-trained collaborative guidance priority calculation model deployed on the edge computing unit, and the collaborative guidance priority factor is output. The pre-training process of the collaborative guidance priority calculation model includes: Multiple historical emergency samples are collected from the historical traffic incident database. Each historical emergency sample includes the location of the historical incident, the lanes occupied in the past, and the historical response data of multiple smart signs in the historical emergency. For each historical emergency sample, for each smart sign participating in the response, calculate the historical distance factor between the smart sign and the location of the historical event, determine the historical lane association between the lane where the smart sign is located and the historically occupied lane, and construct a feature sample set; Based on the actual response level of the smart sign in the historical emergencies, and based on the corresponding historical priority factor labeled according to the actual response level, a tag sample set is constructed; An initial regression model is constructed based on the regression algorithm. Then, the initial regression model is trained in a supervised manner using the feature sample set and the label sample set until it is verified to converge, thus obtaining a pre-trained collaborative guidance priority calculation model.

8. The intelligent traffic guidance system based on intelligent signage according to claim 1, characterized in that, The differentiated boot execution module is specifically used for: A mapping table between collaborative guidance priority factor intervals and guidance strategies is pre-constructed, wherein the mapping table contains a one-to-one correspondence between multiple collaborative guidance priority factor intervals and multiple guidance strategies; Each smart sign determines the target interval to which the collaborative guidance priority factor belongs in the mapping table based on the collaborative guidance priority factor calculated by itself, and matches the guidance strategy corresponding to the target interval from the mapping table; Based on the matched guidance strategy, execute the traffic guidance action corresponding to the guidance strategy.

9. A smart traffic guidance method based on intelligent signage, characterized in that, The intelligent traffic guidance system based on intelligent signage as described in any one of claims 1 to 8 includes: The first smart sign in the smart sign array deployed on the roadside identifies a suspected sudden traffic incident from the collected real-time video stream of the road through a pre-trained lightweight dual-path network and generates a first confidence score. The smart sign array is a traffic perception network formed by multiple smart signs evenly deployed at a preset interval. The first smart sign is the smart sign in the smart sign array that first identifies the suspected sudden traffic incident. Simultaneously call millimeter-wave radar data to detect the abrupt change characteristics of the motion state of the target object in the suspected sudden traffic incident, and generate a second confidence level based on the abrupt change characteristics of the motion state; Based on the first confidence level and the second confidence level, a dual verification judgment is performed. After the dual verification is passed, it is confirmed that a sudden traffic incident has occurred, and sudden incident information including the incident type, incident location and lane occupation is generated. The emergency information is broadcast to all other smart signs in the smart sign array; Each smart sign in the smart sign array calculates a collaborative guidance priority factor based on its own location information and the emergency event information through a collaborative guidance priority calculation model deployed on the edge. Each smart sign matches a corresponding guidance strategy based on the collaborative guidance priority factor it calculates, and executes differentiated traffic guidance actions.