A multi-agent traffic incident real-time response system and method

The multi-agent traffic incident real-time response system utilizes a lightweight language model and multi-source data fusion to achieve second-level identification and dynamic response to traffic incidents. This solves the problems of data silos and insufficient real-time performance in traditional technologies, improves identification accuracy and response efficiency, and enhances system adaptability and security.

CN120612843BActive Publication Date: 2025-12-12ZHEJIANG SUPCON INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511102196.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-12-12
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

Traditional traffic incident identification and response technologies suffer from problems such as data silos, insufficient real-time performance, poor adaptability, and limited response strategies. In particular, they struggle to achieve second-level responses to multi-vehicle rear-end collisions, especially in severe weather conditions.

Method used

A multi-agent traffic event real-time response system is adopted. By deploying lightweight language models on vehicle-mounted and roadside devices, and combining multi-source data fusion, edge decision-making and central collaborative decision-making, a local decision-making closed loop is realized, and response strategies are dynamically generated. The system uses a lightweight recognition model to generate unstructured traffic behavior sets in real time, and optimizes the strategy library through federated learning and reinforcement learning.

Benefits of technology

It improves the accuracy and real-time performance of traffic incident identification, dynamically generates multi-measure collaborative response schemes, reduces edge computing load, enhances system adaptability, reduces false alarm and false negative rates and secondary accident rates, and improves the efficiency and safety of handling complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612843B_ABST
    Figure CN120612843B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-agent traffic event real-time response system and method, system includes: event perception fusion module, is deployed in vehicle and roadside equipment, fuses vehicle motion state, environmental perception data and meteorological data, and non-structured traffic behavior set is generated in real time by lightweight identification model;Edge decision module adopts dynamic role binding mechanism to define decision responsibility, and response instruction is generated by task decomposition based on hierarchical memory storage architecture;Center collaborative decision module, strategy decoupling is realized by layered role adapter, and three-level conflict resolution mechanism is used to coordinate cross-domain response measures;Strategy optimization module, update center parameters by federated learning, and optimize response strategy by combining reinforcement learning reward model, can realize the second-level identification and dynamic response of traffic event, reduce congestion dissipation time, reduce false alarm rate, reduce secondary accident rate, comprehensively improve the disposal efficiency and safety of complex traffic scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent transportation systems, in particular to a multi-agent traffic incident real-time response system and method. BACKGROUND

[0002] Traditional traffic incident identification relies on a single data source, and the response measures are mostly static strategies, which have problems of data island, insufficient real-time performance, poor adaptability and single response strategy. Current traffic incident identification and response technology mainly focuses on vehicle-mounted single sensor, road-side multi-device fusion, cloud centralized processing and fixed rule library, but all have significant limitations. For example, the patent with publication number CN117557570B, which takes a vehicle-mounted camera as the core, can achieve basic target identification, but is easily disturbed by light, rain and fog due to a single data source, resulting in an increased missed detection rate, and a static threshold setting that is difficult to distinguish between school area speed limits and real congestion scenarios, leading to frequent false positives. In addition, large parameter models such as YOLO have insufficient real-time inference capabilities on vehicle-mounted terminals, and there is a significant contradiction between local computing power and model complexity, and the response measures are limited to vehicle-mounted warning prompts, without linkage with road-side traffic lights or cloud platforms, resulting in low disposal efficiency.

[0003] To further improve the reliability of perception, some technologies have shifted to road-side radar fusion solutions. For example, the patent with publication number CN115685185A achieves vehicle trajectory tracking through target-level fusion of millimeter-wave radar and cameras, and predicts abnormal behavior based on Kalman filtering. Although such systems have been applied to intersection signal control, their data sources are still limited to road-side devices, lacking the fine input of vehicle-mounted sensors, making it difficult to capture driver intent or vehicle dynamics. At the same time, trajectory prediction relies on complex models such as cloud LSTM, resulting in a response delay of more than 5 seconds, and when the radar false positive rate increases in bad weather, the system cannot correct the misjudgment in real time through driver voice feedback and other interaction mechanisms, lacking dynamic adaptability. The response level is also limited to single-point regulation of traffic lights, without forming a coordinated strategy with variable lanes and information board information dissemination.

[0004] To meet the needs of multi-source data integration, cloud centralized processing platforms attempt to gather vehicle, road-side and weather data, and perform event analysis and signal timing optimization through multi-modal spatio-temporal feature extraction models. However, the core bottleneck is that the data upload, cloud computing and instruction delivery link delay exceeds 10 seconds, which cannot meet the needs of high real-time scenarios such as emergency braking, and the transmission of raw video streams exacerbates bandwidth pressure and privacy leakage risks. In addition, centrally trained models are difficult to adapt to differentiated scenarios such as mountain roads and urban expressways, lacking the local decision-making capabilities of edge-side lightweight models, resulting in limited generalization.

[0005] Another type of traditional solution completely relies on manual pre-set rule library. Although this type of system can achieve a response within seconds, the rule library needs to be frequently maintained manually, cannot dynamically respond to multi-event superimposed scenarios such as "abnormal parking + large vehicle passing", and lacks the ability to autonomously optimize strategies from historical data. The coarse-grained regulation of fixed parameters is also difficult to match real-time traffic flow changes, and is prone to cause conflicts between local optimization and global congestion in complex road networks. SUMMARY

[0006] In order to solve the problem of chain response failure of multiple vehicle rear-end events in heavy rain scenarios, the present application proposes a multi-agent traffic event real-time response system and method, which loads agents with lightweight language models on vehicle-mounted and roadside devices to realize local decision-making loop, bypass cloud delay, and improve the identification and response speed of the system.

[0007] A further object of the present application is to achieve a second-level response to complex traffic events through hierarchical coordinated decision-making of the device layer, roadside edge layer and center coordination layer.

[0008] In order to achieve the above object, the present application adopts the following technical solution: a multi-agent traffic event real-time response system, characterized in that it comprises:

[0009] An event perception fusion module is deployed on vehicle-mounted and roadside devices, which fuses vehicle motion state, environmental perception data and meteorological data, and generates an unstructured traffic behavior set in real time through a lightweight identification model;

[0010] An edge decision-making module defines decision-making responsibilities using a dynamic role binding mechanism, and generates response instructions based on a hierarchical memory storage architecture;

[0011] A center coordination decision-making module realizes strategy decoupling through a role adapter hierarchy, and coordinates cross-domain response measures using a three-level conflict resolution mechanism;

[0012] A strategy optimization module updates center parameters through federated learning, and optimizes response strategies in combination with a reinforcement learning reward model.

[0013] In this technical solution, by constructing a multi-agent architecture of vehicle-road cloud cooperation, combining a lightweight edge perception module, a role customization decision-making mechanism, a dynamic conflict resolution engine and a federated reinforcement learning optimization strategy, the second-level identification and dynamic response of traffic events are realized. Specifically, multi-source data fusion is used to improve the robustness of perception, role adapter hierarchy is used to inject domain knowledge to realize strategy decoupling, a three-level conflict resolution strategy mechanism is used to coordinate cross-domain response measures, and based on federated learning and reinforcement learning, the strategy library is continuously optimized, and finally a "perception-decision-making-coordination-optimization" closed-loop system is formed, which significantly improves the disposal efficiency and safety of complex traffic scenarios.

[0014] Preferably, the event perception fusion module comprises a collision risk warning unit, which calculates a collision time estimate in real time through an adaptive Kalman filter model, calculates a collision risk probability value when the collision time estimate is less than a critical safety time corresponding to the current vehicle speed and road conditions, and generates an alarm text for collision risk warning.

[0015] Preferably, calculating the collision risk probability value comprises: obtaining a maximum deceleration stopping time based on the road friction coefficient, dynamically adjusting a mean correction term and a variance correction term of a lognormal distribution to obtain a driver reaction time based on the real-time attention score and the line-of-sight focus degree, wherein the mean correction term is negatively correlated with the real-time attention score, and the variance correction term is negatively correlated with the line-of-sight focus degree; and obtaining the collision risk probability value according to the driver reaction time, the maximum deceleration stopping time, and the system response delay, in combination with the collision time estimate.

[0016] Preferably, the event perception fusion module comprises a voice feedback unit, which comprises:

[0017] Lightweight large model edge deployment, pre-training large model loaded on roadside and vehicle terminal;

[0018] Unstructured text fusion interface, multiple source alarm texts are uniformly encoded into natural language sequence input;

[0019] Dynamic prompt engine, when the event confidence is lower than a set threshold, an inquiry text for ambiguous semantics is generated;

[0020] Voice feedback closed-loop correction, real-time update of event judgment result according to vehicle side user voice input.

[0021] Preferably, the event perception fusion module comprises a vehicle slip detection unit, which establishes a slip rate dynamic calculation model through cross verification of four-wheel rotational speed sensors and visual speed measurement, obtains an actual slip rate based on the Pacejka magic formula, and generates an alarm text for a vehicle slip event when any wheel slip rate is greater than a first threshold and lasts for more than a second threshold time.

[0022] Preferably, the vehicle slip detection unit comprises:

[0023] Multi-source verification engine, through cross verification of wheel speed sensors and visual speed measurement, detects data conflict, and triggers IMU inertial unit arbitration when the data conflict reaches a fourth threshold;

[0024] Dynamic weight distributor, dynamically adjusts the multi-sensor data fusion weight according to the real-time road friction coefficient;

[0025] Slip event decision maker, when the arbitrated slip rate lasts for more than 500 ms, a slip alarm text is generated; the dynamic threshold is negatively correlated with the road friction coefficient.

[0026] Preferably, the edge decision module comprises:

[0027] a role definition layer that dynamically binds the responsibilities of the agent through a prompt word template;

[0028] a hierarchical memory storage layer that includes a short-term memory cache unit that stores real-time event context using a sliding window and a long-term memory retrieval unit that associates historical event handling schemes based on feature similarity;

[0029] a task sub-solution strategy layer that extracts spatiotemporal features using Bi-LSTM and ranks task priorities by improving the TOPSIS algorithm.

[0030] Preferably, the central collaborative decision-making module has a role adapter, which comprises:

[0031] a feature extraction unit that extracts global context features of the input event;

[0032] a role transformation unit that compresses the features to a bottleneck dimension using a dimension reduction matrix, recovers them after nonlinear activation by a dimension increasing matrix; and a residual fusion unit that weights and fuses the output of the role transformation unit and the original input features to generate role decision features.

[0033] Preferably, the reinforcement learning reward model of the strategy optimization module uses a negative expected logarithmic Sigmoid difference as the training loss function; and the loss function for parameter learning of the agent model according to artificial evaluation comprises: a reward driving item that maximizes the reward score of the human-recognized response; a strategy anchoring item that constrains the deviation of the current strategy from the basic safety strategy by weighting the logarithm of the strategy probability ratio; and a dynamic balance mechanism that adjusts the constraint strength coefficient according to the emergency level of the event.

[0034] The application also adopts the following technical solution: a multi-agent traffic event real-time response method based on the above-mentioned multi-agent traffic event real-time response system, comprising the following steps:

[0035] S1, real-time identification of traffic events and generation of an unstructured event feature set;

[0036] S2, dynamic adjustment of role responsibilities according to the event type and construction of preliminary response instructions according to the event features;

[0037] S3, multi-agent strategy collaborative decoupling, dynamic generation of globally optimized instructions through three-level conflict resolution;

[0038] S4, updating of parameters through federated learning and continuous iteration and evolution of the strategy library in combination with the reinforcement learning reward model.

[0039] The application has the following beneficial effects:

[0040] 1) Improve recognition accuracy and real-time performance: through vehicle-road cloud multi-source data fusion and Kalman filter denoising, enhance the robustness of event detection;

[0041] 2) Dynamic response strategy optimization: based on lightweight large language model to build agent, combined with federated learning and reinforcement learning, dynamically generate multi-measure collaborative response scheme;

[0042] 3) Reduce edge computing load: deploy lightweight models on vehicle or roadside devices, reduce computing power demand through unstructured text interaction;

[0043] 4) Enhance system adaptability: introduce user voice feedback and dynamic threshold, such as slip rate magic formula, TTC collision time compensation coefficient, improve adaptability in complex scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 is a whole system block diagram of a multi-agent traffic event real-time response system of embodiment 1 of the present application.

[0045] Figure 2 is a system block diagram of the event perception fusion module of embodiment 1 of the present application.

[0046] Figure 3 is a system block diagram of the edge decision module of embodiment 1 of the present application.

[0047] Figure 4 is a system block diagram of the central collaborative decision module of embodiment 1 of the present application.

[0048] Figure 5 is a logic diagram of the three-level conflict resolution mechanism of embodiment 1 of the present application.

[0049] Figure 6 is a system block diagram of the strategy optimization module of embodiment 1 of the present application.

[0050] Figure 7 is a training flowchart of the strategy optimization module of embodiment 1 of the present application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical scheme and advantages of the present application clearer and more apparent, the present application will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only the best mode of the present application, which are used to explain the present application and do not limit the protection scope of the present application. All other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0052] Embodiment 1

[0053] This embodiment provides a multi-agent traffic event real-time response system, such as... Figure 1 As shown, it includes an event perception fusion module, an edge decision-making module, a central collaborative decision-making module, and a strategy optimization module. It can achieve second-level identification and dynamic response to traffic events, reduce congestion dissipation time, reduce false alarm and missed alarm rates and secondary accident rates, and comprehensively improve the efficiency and safety of handling complex traffic scenarios.

[0054] The event perception fusion module is deployed in vehicle-mounted and roadside equipment, integrating multi-source traffic data and generating structured event features in real time through a lightweight recognition model, referencing... Figure 2 .

[0055] The vehicle-side sensors mainly include accelerator, brake, steering wheel direction, camera, radar, inertial unit, GPS, voice, etc., and take data from 9-axis IMU and GPS as input. When both types of data are available, Kalman filtering is used to fuse and denoise them.

[0056] The surrounding environment is dynamically identified by fusion of visual or radar-visual data from vehicles or roadsides. The target recognition model uses, but is not limited to, YOLO 11, to determine the type, location, speed, road channelization, speed limit, traffic lights, signs, etc. of surrounding vehicles.

[0057] By deploying various lightweight recognition models on vehicle terminals or roadside edge devices, unstructured traffic behavior sets can be identified and generated. The lightweight recognition models include, but are not limited to, the following sub-modules.

[0058] Abnormal parking monitoring unit: When the vehicle speed is less than 3km / h, it generates an alarm text for an abnormal parking event, and at the same time supplements the driving speed of surrounding vehicles.

[0059] Low-speed slow-moving detection unit: When the vehicle speed is below 20km / h, it generates an alarm text for a low-speed slow-moving event and adds information on the speed of surrounding vehicles.

[0060] Emergency braking detection unit: Based on an improved sliding window variance algorithm, it analyzes triaxial acceleration data in real time, calculates acceleration abrupt change characteristics within the window, and generates an alarm text for an emergency braking event when the variance value in the vehicle's forward direction exceeds a third threshold. Taking N=18 and Δt=300ms as an example, in this embodiment, the third threshold is set to 3.5m / s. 2 .

[0061] It should be noted that the acceleration variance of emergency braking under heavy rain weather will increase in both collision and non-collision scenarios, so a fixed threshold will not cause missed detection, but false positives may occur. In the present embodiment, false positives can be effectively prevented by multi-event cooperative detection. In some other embodiments, a dynamic adjustment mechanism can also be set to associate the threshold with the weather, such as increasing the third threshold according to an empirical coefficient in heavy rain weather to improve the accuracy of the emergency braking detection unit single detection result.

[0062] Abnormal lane changing detection unit: taking the original trajectory P as input, predicting the future 2-second trajectory P' through a recurrent neural network model (such as GRU), adopting Hausdorff distance d H Quantifying trajectory deviation, when Hausdorff distance d H is greater than 1.2 m, generating an abnormal lane changing event alarm text, and supplementing event labels according to turn signals, target steering wheel angle, etc.

[0063] Vehicle skid detection unit: through cross verification of four-wheel rotational speed sensors and visual speed measurement, a slip rate dynamic calculation model is established, and based on the Pacejka magic formula, the actual slip rate λ is derived. When any wheel λ is greater than 15% and lasts more than 500 ms, an alarm text of a vehicle skid event is generated.

[0064] Large vehicle monitoring unit: through radar perception data or video data to identify the size of the target, if the target speed is greater than a threshold (such as 40 km / h on urban roads and 80 km / h on high-speed roads), an alarm text of a large vehicle driving event is generated.

[0065] Special vehicle monitoring unit: through traffic management publishing, surrounding vehicle type identification, special vehicle voiceprint matching, etc., to discover special vehicle (such as police cars, ambulances, fire trucks, etc.) passing needs, and generate an alarm text of a special vehicle smooth event.

[0066] Collision risk warning unit: through an adaptive Kalman filter model to calculate the predicted collision time TTC in real time, specifically, taking the ratio of relative distance and speed as the base, and taking the function of acceleration as the product correction term to get the predicted collision time. The product correction term is negatively related to acceleration, for example, the product correction term can be set to 2 / [1+γsign(a rel )], where a rel represents acceleration, and γ is a dynamic compensation coefficient.

[0067] When the predicted collision time is less than 2.5 seconds, the collision risk is estimated according to the predicted collision time and the driver reaction time, when the critical safety time corresponding to the current vehicle speed and road conditions is greater than the predicted collision time, the collision risk probability value is the complement percentage of the ratio of the predicted collision time to the critical safety time corresponding to the current vehicle speed and road conditions, and when the critical safety time corresponding to the current vehicle speed and road conditions is not greater than the predicted collision time, the collision risk probability value is 0.

[0068] Specifically, the critical safety time corresponding to the current vehicle speed and road conditions is affected by the vehicle stopping time, system response delay and human reaction delay time. In the embodiment, the three are added to obtain the critical safety time corresponding to the current vehicle speed and road conditions.

[0069] Wherein, the vehicle stopping time varies according to the current vehicle speed v and road conditions, and can be specifically expressed as v / μg, μ is the road friction coefficient, generally between 0.3-0.9, the actual value is dynamically adjusted according to the real-time precipitation intensity data obtained from the real-time meteorological platform, and g is the current position gravity acceleration; the system response delay is obtained according to the empirical value, which can be taken as 0.3s in the embodiment.

[0070] The human reaction delay time, i.e. the driver reaction time, is subject to a lognormal distribution, and the mean correction term and the variance correction term are formed by the real-time attention score of the driver and the line of sight focus degree respectively. The mean correction term is negatively correlated with the real-time attention score, and the variance correction term is negatively correlated with the line of sight focus degree.

[0071] Specifically, the mean correction term can be expressed as α(1-S attention / 100), and the variance correction term can be expressed as e -βSfocus The mean correction term is added to the mean to obtain the corrected mean, and the variance correction term is multiplied by the variance to obtain the corrected variance.

[0072] Wherein, S attention is the real-time attention score, which can be obtained according to the driver's facial features captured by the vehicle-mounted camera in real time; Sfocus is the line of sight focus degree, which can be obtained based on the matching degree of the eye gaze point and the key area of the road; α and β are learning parameters, both of which are taken as 0.15 in the embodiment.

[0073] According to the collision risk probability value, the warning text of the collision risk warning is finally generated.

[0074] Pedestrian intrusion warning unit: through thermal imaging or video data, an improved YOLOv5-Human detection network is used to generate the warning text of the pedestrian intrusion warning.

[0075] Abnormal weather warning unit: access to public meteorological platform, and output warning text for strong wind, heavy rain, ice and snow, high temperature and other severe weather.

[0076] User voice feedback unit: edge deployment of pre-trained lightweight large language model, acting as edge data fusion analysis agent, unified acceptance of multi-source and multi-type alarm text, through unstructured text input form to construct the target of each topological section, comprehensive analysis of the spatio-temporal correlation of section information, extraction of the location, type and adverse impact of abnormal events. Through prompt words to remind the model to find out the key problems of fuzzy judgment, ask the user side, and correct the judgment result according to the user voice feedback.

[0077] The knowledge base is constructed in seven dimensions of measure type, execution mode, configuration parameter, response speed, event correlation, combination rule and traffic impact, and complete prompt words for each response measure are written.

[0078] The edge decision module connects the event perception fusion module, adopts a dynamic role binding mechanism to define decision responsibilities, and generates response instructions based on a hierarchical memory storage architecture, referring to Figure 3 .

[0079] The edge decision module can build a basic single agent, and automatically build and guide single agents to respond to traffic events using a four-layer progressive architecture to achieve autonomous decision-making capabilities of agents in edge computing environments.

[0080] The four-layer progressive architecture includes a role definition layer, a hierarchical memory storage layer, a three-level task decomposition decision layer, and an execution feedback layer.

[0081] The role definition layer generates an agent identity feature matrix through pre-trained prompt word templates. Specifically, the prompt word template of the identity feature matrix is "As {agent type}, your core responsibilities include: 1. Real-time monitoring of {data type} in {jurisdiction area}; 2. Based on {decision model} to evaluate event risk level (1-5); 3. Trigger the top 3 priority in {response measure library}; 4. Constraints: response delay <{max_latency} ms, decision confidence >{min_confidence}%; 5. Dynamic correspondence between device ID and responsibility area".

[0082] Dynamic identity binding of single agents is achieved through the identity feature matrix, which can solve the problem of fixed roles of traditional edge devices that cannot dynamically adjust responsibilities according to event types.

[0083] The hierarchical memory storage layer includes a short-term memory cache unit and a long-term memory retrieval unit. The short-term memory cache unit uses a sliding window to store real-time event context; the long-term memory retrieval unit correlates historical event handling schemes based on feature similarity.

[0084] Specifically, in the present embodiment, the short-term memory adopts the context caching mechanism of the Transformer-XL architecture, the window length is adjustable (default 128 tokens), the long-term memory is based on the vector database of FAISS, the feature similarity retrieval is implemented, and the associative memory is triggered when the cosine similarity is > 0.85.

[0085] The three-level task solution strategy layer assists the single-agent decision by calling the subdivision model, for example, the Bi-LSTM extracts the spatio-temporal feature vector of the event, the improved TOPSIS algorithm (entropy weight method optimizes the weight) is applied to prioritize the tasks, and the complex instructions are decomposed into API call sequences according to the measure implementation knowledge base.

[0086] The execution feedback layer presets 12 standard operation interfaces (signal control, unmanned aerial vehicle scheduling, etc.).

[0087] In the present embodiment, the edge decision module is also provided with an exception handling mechanism, which starts the backup scheme selector when the API returns an error code.

[0088] The central collaborative decision module connects multiple edge decision modules, realizes strategy decoupling through role adapter layering, and coordinates cross-domain response measures through a three-level conflict resolution mechanism, referring to Figure 4 .

[0089] The role adapter layering architecture will be described in detail below.

[0090] The agent roles are divided into edge event recognition agents, central decision agents, edge traffic control agents, edge information interaction agents, edge emergency rescue agents, and edge rapid disposal agents. On the basis of a pre-trained large model (such as DeepSeek 671B), a pre-trained understanding model LLM pre (·) is established.

[0091] To ensure the independence of central decision and different single-agent thinking, a lightweight role adapter is trained for central agents and edge response agents.

[0092] The specific implementation is that, on the premise of freezing the parameters of the pre-trained large model, the role-specific knowledge injection is realized through a lightweight residual structure. Taking the original event feature vector x (containing event type, location, impact range, etc.) as input, the adaptive feature A(x) injected with role knowledge is obtained through a role-customized feature enhancement model, for example, the emergency rescue role focuses on "injury classification", and the signal control role focuses on "time coordination".

[0093] The process of obtaining the adaptive feature A(x) injected with role knowledge through the role-customized feature enhancement model is as follows.

[0094] First, the self-attention layer of the pre-trained large model (such as DeepSeek 671B) is called to extract the global context features of the input x, and the relevance of the key elements of the event is identified, such as the spatio-temporal binding of "location K12+300" and "east-west direction".

[0095] Then, the attention features are transformed nonlinearly and embedded with domain knowledge through the feedforward network layer of the pre-trained large model, activating traffic event-specific features such as the decision correlation of "congestion slowdown" and "signal cycle extension".

[0096] Next, down-sampling compression is performed to reduce the feature dimension of the pre-trained model from d to the bottleneck dimension k, for example, the bottleneck dimension is 8 in this embodiment, reducing the computational load; nonlinear activation is performed by introducing the nonlinear characteristics required for role decision through the activation function σ(·); up-sampling recovery is performed to restore the features to the original dimension, preserving the role knowledge. Through the above steps, role-sensitive features are generated, such as the "ambulance dispatch priority" feature enhanced by the edge emergency rescue agent.

[0097] Finally, in the residual feature fusion layer, the weighted sum of the original input x and the role-adapted features is calculated to preserve the integrity of the original event information, prevent catastrophic forgetting, and only add role difference features, such as the "road network global impact" weight added by the central decision-making agent.

[0098] It should be noted that after down-sampling compression, a bias calibration term b can be added first to learn the bias vector and compensate for the feature distribution shift of role adaptation to adapt to the decision preferences of different agents, such as the signal control agent's preference for short latency strategies.

[0099] In this technical solution, the role adapter performs a three-stage transformation of feature compression-role activation-feature reconstruction to inject lightweight role-specific decision-making features while preserving the general capabilities of the pre-trained model. The feature compression layer reduces the high-dimensional features to an 8-dimensional bottleneck space; the nonlinear activation layer gives the role different decision-making characteristics; and the residual fusion mechanism prevents the loss of original event information.

[0100] Then, the normalization layer of the role adapter eliminates the semantic deviation between the role features and the original model, standardizes the distribution of the adapter output, and eliminates the feature distribution shift caused by role adaptation. Then, the pre-trained language model generates response instructions based on the normalized features to perform the final decision-making.

[0101] In this technical solution, the parameters of the pre-trained large model are completely frozen, and only the front role adapter is used to realize decision customization, which can significantly reduce the computational load.

[0102] The role fine-tuner continues three-step hierarchical training on the traffic event corpus.

[0103] First, pre-train the base parameters such as attention layers and feedforward layers on 100,000 traffic event texts, and inject domain knowledge.

[0104] Second, freeze the base parameters, and independently train each adapter with 50,000 role-labeled data.

[0105] Third, implement knowledge distillation through KL divergence constraint and feature alignment (L KD ), and constrain the sub-agents to ensure policy coordination among agents.

[0106] The third step is implemented through a multi-agent collaborative decision-making unit, which includes a probability distribution constraint, a feature space aligner, and a dynamic weight distributor.

[0107] The probability distribution constraint enforces the output decision probability distribution of the edge agent to approximate the center agent, eliminating decision conflicts such as signal control and variable lane instruction conflicts. The feature space aligner minimizes the Euclidean distance between the intermediate layer features of the center and edge agents, unifying event semantic understanding, such as consistent global 'congestion slow down' judgment criteria. The dynamic weight distributor adjusts the feature alignment strength according to the event urgency.

[0108] In this technical solution, the probability distribution constraint and feature space alignment mechanism achieve multi-agent policy coordination.

[0109] The center-coordinated edge multi-agent architecture is described in detail below.

[0110] Through historical event identification and response data recorded by manual input or system, according to the occurrence time, location, event type, event status, and response measures library supported by real conditions, feature encoding is formed, and time-space feature vectors are extracted through Bi-LSTM. The edge event recognition agent encodes real-time events and sends them to the center agent, which searches for similar features in the event memory library.

[0111] The top ten most relevant event memories are obtained, and the center agent is guided by prompt words to analyze the logical relationship between events and response measures, current events and historical events, and response measures and effects, to determine the type of response measures that can be selected. Further, the establishment of edge agents of this type is initiated to analyze executable response measures within the event impact range, and the edge agents are sent to the edge agents through prompt words, including executable measure types, execution terminals, mutual influence between existing measures, and related historical events. Each agent generates corresponding response actions and parameters based on the prompt dialogue.

[0112] The conflict detection and resolution method is described in detail below, as shown in Figure 5 .

[0113] A three-level conflict processing mechanism is established, respectively L1, L2 and L3. L1 is resource mutual exclusion, for example, different agents need to use an ambulance, but there is only one in the nearby range; L2 is strategy conflict, for example, the ambulance needs to ensure the green light to pass through, and the induced vehicle needs to shorten the green light time, and the variable lane; L3 is the overlap of space-time influence domain, for example, the nearby ground intersection and ramp entrance are congested.

[0114] More than 100 conflict rules have been preloaded on the central agent, for example, "IF measure A contains "signal control", AND measure B contains "variable lane" AND influence radius overlap > 60% THEN conflict level = L2", the central agent collects the decision results of each agent in real time, and judges the conflict level according to the knowledge base and event logic.

[0115] For different conflict situations, the central agent processes different conflicts according to the following process according to the conflict level.

[0116] For L1 level conflict, trigger "man-in-the-loop" feedback response mechanism, the central agent generates visual event and impact analysis report, and pushes it to the management and control personnel for decision-making.

[0117] For L2 level conflict, the central agent is guided by prompt words to evaluate the risks that different choices may bring, and the risk types and priorities are sorted, with the priority of solving life safety risks, followed by traffic safety.

[0118] For L3 level conflict, according to the response measures, confidence and decision-making benefits of independent decision-making of different agents, a hybrid parameter method is used to generate a compromise measure.

[0119] Specifically, first, fuse the confidence and expected return to generate the measure reliability weight, then normalize the weight distribution, convert the absolute weight to probability distribution form, and finally generate a compromise measure according to the original response measure.

[0120] By dynamically weighting the independent decisions of each agent, a globally optimal compromise measure is generated, and each measure is linearly combined according to the weight, rather than simply voting or priority coverage, which can solve the conflict of overlapping space-time influence domain in multi-agent collaborative decision-making.

[0121] At the same time, the execution side sets all-round safety constraints for the specific instructions issued by the edge agent, such as the step of the signal light, the switchable flow direction of the variable lane, etc., and if the issued instructions are abnormal, the specific abnormal information is fed back to prompt the edge agent to correct.

[0122] The strategy optimization module updates the central collaborative decision-making module parameters through federated learning, optimizes the response strategy combined with the reinforcement learning reward model, and refers to Figure 6and Figure 7 .

[0123] The supervised fine-tuning of the single-agent adapter is described in detail below.

[0124] The adapter of different agents is trained with the historical record as the anticipation library, and the cross-entropy of the next word prediction of each position in the sequence is calculated as the loss function. The parameter update of the central agent uses federated parameter learning.

[0125] The interactive reinforcement learning of multiple agents is described in detail below.

[0126] The reinforcement learning reward model is established to calculate the scalar reward of text quality for different input and output. The reward model takes the negative expected logarithmic Sigmoid difference as the training loss function. Specifically, input the traffic event description x, the high-quality response recognized by humans, and the set of responses generated by the agent to be evaluated, calculate the difference between the reward score of the positive sample response and the reward score of the negative sample response set, and convert the score difference into a probability value through the Sigmoid function. The expected value of the negative logarithmic probability is taken as the optimization target to drive the reward model to improve the response quality differentiation capability.

[0127] Based on human preference, the agent decision strategy is optimized, and the model parameters are dynamically adjusted through artificial evaluation data to improve the rationality of event response. Specifically, define the loss function for parameter learning of the agent model according to artificial evaluation, input the traffic event description x and the response y generated by the agent, calculate the scalar reward value given by the reward model, minimize the expected negative reward of the agent response strategy, and drive the strategy to generate high-yield responses that meet human preferences; Calculate the logarithm of the probability ratio of the reinforcement learning strategy and the supervised fine-tuning base strategy, and take the negative coefficient -β weighted expected value as the strategy anchoring term to constrain the strategy update amplitude and prevent the decision logic from mutating and deviating from the safety benchmark; Calculate the logarithmic probability of the response generated by the reinforcement learning strategy, and take the positive coefficient γ weighted expected value as the exploration incentive term to maintain the diversity of the strategy to avoid local optimum.

[0128] The three-part weighted sum drives the strategy to dynamically generate efficient response schemes that meet human evaluation under the premise of ensuring safety benchmarks, reducing traffic event response delay and strategy conflict rate.

[0129] Embodiment 2

[0130] This embodiment takes the rapid identification and response of a multi-vehicle rear-end event on an urban expressway in rainy weather as an example to further illustrate the multi-agent traffic event real-time response system provided in Embodiment 1.

[0131] During the morning rush hour in a certain city, a sudden heavy rain caused the road surface to be slippery at the curve of the expressway. A car skidded due to emergency braking, and the vehicles behind it failed to slow down in time due to obstructed vision, resulting in a 5-car rear-end collision accident.

[0132] Traditional traffic management systems rely on single vehicle-mounted sensors (such as cameras) and fixed threshold rules (such as sudden speed drop triggering alarms), making it difficult to quickly identify the location and scope of the accident under the interference of rain, leading to delayed signal light adjustment, rescue vehicles unable to arrive quickly, and congestion spreading.

[0133] The application process of a multi-agent traffic incident real-time response system of the present application is described in detail below.

[0134] First, real-time collection and fusion of multi-source data are performed.

[0135] Access real-time precipitation intensity data (15mm / h) from the meteorological platform, and dynamically correct the road friction coefficient μ to 0.35; the accident truck triggers emergency braking (acceleration variance exceeds threshold 3.5m / s 2 ), the vehicle-mounted IMU detects skidding (slip ratio s = 0.25), and the GPS locates the accident at the K12+300 curve of the expressway; the radar and vision fusion device detects a sudden speed drop to 8km / h in the accident area, the trajectories of vehicles behind are offset (Hausdorff distance d = 4.2m), and the thermal imaging camera captures the behavior of pedestrians getting off the vehicle to avoid danger.

[0136] Second, edge agents quickly analyze and generate events.

[0137] The roadside edge computing node completes multi-source data fusion within 2 seconds and generates an unstructured alarm text: "[Event Type] Multi-vehicle Rear-End Collision Accident, [Location] Expressway K12 (East to West), [Impact] 3-lane congestion, 800-meter queue length behind, average speed 8km / h, [Associated Risk] Pedestrian Intrusion, Secondary Collision Risk".

[0138] Finally, a lightweight large language model analyzes the text, extracts key information, and generates preliminary response suggestions: "1) Upstream intersection variable lane switching (left turn to straight lane, open emergency lane); 2) Trigger rescue vehicle priority green wave; 3) Linkage navigation platform pushes rerouting information; 4) Roadside variable information board displays 'Accident ahead, reduce speed to avoid secondary collision'".

[0139] The implementation effect of the present embodiment is as follows.

[0140] Response efficiency is improved: event identification delay is shortened from 30 seconds in the traditional scheme to within 5 seconds, response strategy generation time ≤3 seconds, and congestion dissipation time is reduced by 38% (traditional scheme 120 minutes → present scheme 75 minutes).

[0141] False positive and false negative optimization: Through multi-source data cross-validation (such as IMU + radar + weather), the false positive rate is reduced from 15% to 2%; due to the assistance of thermal imaging, the false negative rate of pedestrian intrusion detection is reduced from 10% to 1.5%.

[0142] Resource utilization improvement: Emergency lane opening decision accuracy rate increased to 95%, avoiding resource waste caused by invalid lane switching; signal timing dynamic optimization increases auxiliary road traffic volume by 22%.

[0143] User interaction enhancement: Driver voice feedback corrects misjudgment events (such as correcting "vehicle slow down" to "courtesy to pedestrians"); attention score model linked to warning system, secondary accident rate down 45%.

Claims

1. A multi-agent traffic incident real-time response system, characterized in that, The application relates to an event perception fusion module, which is arranged on a vehicle and a roadside device, fuses vehicle motion state, environment perception data and meteorological data, and generates an unstructured traffic behavior set in real time through a lightweight identification model; an edge decision module, which adopts a dynamic role binding mechanism to define decision responsibilities, including: generating an intelligent agent identity feature matrix through a pre-trained prompt word template according to the traffic behavior set; storing real-time event context based on a hierarchical memory storage architecture and a historical event disposal scheme based on feature similarity; decomposing a task to generate a response instruction, including: extracting a time-space feature vector of the event; performing task priority sorting; decomposing a complex instruction into an API calling sequence according to a measure implementation knowledge base; a central collaborative decision module, which realizes strategy decoupling through a role adapter hierarchical implementation, including: extracting global context features of an original event feature vector and identifying the relevance of event key elements; performing nonlinear transformation and domain knowledge embedding on attention features to activate traffic event exclusive features; weighting and summing the original event feature vector and the role adaptation features; and adopting a three-level conflict resolution mechanism to coordinate cross-domain response measures; and a strategy optimization module, which updates parameters of a central intelligent agent through federated learning and optimizes a response strategy in combination with a reinforcement learning reward model. The event perception fusion module comprises a collision risk early warning unit, which calculates a collision time estimate value in real time through an adaptive Kalman filter model, calculates a collision risk probability value when the collision time estimate value is less than a critical safety time corresponding to a current vehicle speed and road surface condition, and generates an alarm text of the collision risk early warning. The collision risk probability value is calculated by: obtaining a maximum deceleration stopping time based on a road surface friction coefficient, dynamically adjusting a mean correction term and a variance correction term of a logarithmic normal distribution to obtain a driver reaction time based on a real-time attention score and a line of sight focus degree, wherein the mean correction term is negatively correlated with the real-time attention score, and the variance correction term is negatively correlated with the line of sight focus degree; and obtaining the collision risk probability value based on the driver reaction time, the maximum deceleration stopping time and a system response delay in combination with the collision time estimate value. The event perception fusion module comprises a voice feedback unit, including: lightweight large model edge deployment, pre-trained large model loading on a roadside terminal and a vehicle terminal; an unstructured text fusion interface, which uniformly encodes multiple source alarm texts into natural language sequence inputs; a dynamic prompt word engine, which generates inquiry texts for ambiguous semantics when an event confidence is lower than a set threshold; and a voice feedback closed-loop correction, which updates an event judgment result in real time according to vehicle-side user voice input. The event perception fusion module comprises a vehicle slip detection unit, which establishes a slip rate dynamic calculation model through cross verification of four-wheel rotation speed sensors and visual speed measurement, obtains an actual slip rate based on a Pacejka magic formula, and generates an alarm text of a vehicle slip event when any wheel slip rate is greater than a first threshold value and lasts for more than a second threshold time.

2. The multi-agent traffic incident real-time response system of claim 1, wherein, The vehicle slip detection unit comprises: a multi-source verification engine, which detects data conflicts through cross verification of four-wheel rotation speed sensors and visual speed measurement and triggers arbitration when the data conflicts reach a fourth threshold value; 3. The multi-agent traffic incident real-time response system of claim 2, wherein, ​ 4. The multi-agent traffic incident real-time response system of claim 1, wherein, ​ ​ ​ ​ ​ 5. The multi-agent traffic incident real-time response system of claim 1 or 2 or 4, wherein, ​ 6. The multi-agent traffic incident real-time response system of claim 5, wherein, ​ ​ A dynamic weight allocator dynamically adjusts the weight of multi-sensor data fusion according to the real-time road friction coefficient; A slip event decision maker generates a slip warning text when the arbitrated slip rate continues to exceed the dynamic threshold for 500 ms; the dynamic threshold is negatively related to the road friction coefficient.

7. The multi-agent traffic incident real-time response system of claim 1, wherein, The edge decision module comprises: A role definition layer dynamically binds the responsibilities of intelligent agents through prompt templates; A hierarchical memory storage layer includes a short-term memory cache unit that stores real-time event context using a sliding window and a long-term memory retrieval unit that associates historical event handling schemes based on feature similarity; A task resolution decision layer uses Bi-LSTM to extract spatio-temporal features and improves the TOPSIS algorithm to rank task priorities.

8. The multi-agent traffic incident real-time response system of claim 1, wherein, The central collaborative decision module has a role adapter, which includes: A feature extraction unit extracts global context features of the original event feature vector and identifies the relevance of event key elements; A role transformation unit performs nonlinear transformation and domain knowledge embedding on attention features to activate traffic event-specific features; it uses a dimension reduction matrix to compress features to a bottleneck dimension, and after nonlinear activation, it recovers them using an upsampling matrix; A residual fusion unit weights and fuses the output of the role transformation unit with the original event feature vector to generate role decision features and achieve policy decoupling.

9. The multi-agent traffic incident real-time response system of claim 1, wherein, The reinforcement learning reward model of the policy optimization module uses the negative expected logarithmic Sigmoid difference as the training loss function; The loss function for parameter learning of the agent model based on human evaluation includes: A reward-driven item that maximizes the reward score of human-recognized responses; 10. A multi-agent traffic incident real-time response method based on the multi-agent traffic incident real-time response system of any one of claims 1-9, characterized in that, A policy anchoring item that constrains the deviation of the current policy from the basic safety policy by weighting the logarithm of the policy probability ratio; a dynamic balancing mechanism that adjusts the constraint intensity coefficient according to the event urgency. The method comprises the following steps: S1, real-time identification of traffic events and generation of an unstructured event feature set; S2, dynamically adjusting the role responsibilities according to the event type and constructing preliminary response instructions based on event features; S3, multi-agent policy collaborative decoupling, dynamically generating globally optimized instructions through three-level conflict resolution; S4, updating parameters through federated learning and continuously iterating and evolving the policy library in combination with the reinforcement learning reward model.

Citation Information

Patent Citations

  • 4D millimeter wave radar and vision fusion perception method

    CN115685185A

  • A rail vehicle abnormality detection method and system

    CN117557570B

  • Three-dimensional road network traffic guidance and emergency rescue collaborative management method for underground road

    CN117373243A

  • Intelligent agent role switching method and system based on multi-modal perception and related components

    CN120197139A