Traffic subject collaborative cognition method based on semantic representation
By collecting and transforming traffic data into structured semantic representations in intelligent transportation systems and establishing a multi-modal collaborative interaction architecture, the problems of perception limitations and insufficient real-time response are solved, enabling efficient collaborative decision-making and information sharing among traffic entities, and improving the system's safety and traffic efficiency.
Patent Information
- Application Number
- CN202511637523.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-02-13
AI Technical Summary
Existing intelligent transportation systems suffer from limitations in perception, insufficient real-time response capabilities, incomplete risk assessment, and low communication efficiency in complex intersection scenarios, leading to high misjudgment rates, decision delays, and systemic failures.
By collecting traffic scene and subject status information, and using data fusion technology to transform it into structured semantic representation, a multi-mode collaborative interaction architecture is established to realize autonomous decision-making by peripheral traffic subjects and overall optimization of central traffic subjects, thus forming consistent decisions both locally and globally.
It enhances the perception capabilities among traffic entities, enables efficient fusion and collaborative decision-making of multi-source data, breaks down information barriers, improves system security and traffic efficiency, and has good scalability and adaptability.
Smart Images

Figure CN121527992A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of transportation technology, and specifically relates to a collaborative cognition method for transportation subjects based on semantic representation. Background Technology
[0002] Current intelligent transportation systems primarily rely on vehicle-mounted intelligent sensing technology, using onboard sensors to perceive the environment and make decisions. However, vehicle-mounted sensing has inherent limitations: its field of view is severely restricted by physical obstructions and weather conditions. For example, the recognition rate of visual sensors drops significantly in rainy or foggy weather, and blind spots on the sides of the vehicle cause delays in detecting suddenly appearing pedestrians or non-motorized vehicles. Each vehicle needs to independently process massive amounts of sensing data, making it difficult to meet millisecond-level response requirements in sudden emergency scenarios. Existing risk assessment models are mostly based on single-dimensional parameters such as relative distance, lacking comprehensive analysis of vehicle speed, road topology, and participant behavioral intentions, resulting in a persistently high misjudgment rate in complex intersection scenarios.
[0003] Vehicle-road cooperative technology expands the perception range through "vehicle-road-cloud" interaction, but existing solutions face severe challenges. Vehicle-to-everything (V2X) networks require high-frequency transmission of heterogeneous data from multiple sources, such as roadside cameras and radar point clouds. Direct transmission of raw perception data causes severe bandwidth overload. In typical urban intersection scenarios, data transmission per second can reach several gigabytes, far exceeding the capacity of existing communication protocols, leading to decision-making delays. Roadside devices and vehicle-mounted sensors experience spatiotemporal discrepancies due to clock asynchrony and coordinate system differences. Traditional data fusion methods struggle to eliminate systematic errors from asynchronous heterogeneous data, significantly reducing the robustness and consistency of environmental perception.
[0004] Existing solutions suffer from fundamental contradictions in their architectural design. Centralized cloud-based decision-making models, due to excessively long data transmission distances, struggle to meet the low-latency requirements of dynamic traffic scenarios; fully distributed vehicle-side computing, limited by the computing power of a single vehicle, cannot support the complex computational needs of multi-entity collaborative scenarios. At the risk perception level, some solutions employ simplified models applicable only to specific scenarios, lacking sufficient ability to model complex interactions such as multi-vehicle game theory and mixed pedestrian-vehicle traffic; while some deep learning-based solutions can improve prediction accuracy, their computational complexity increases exponentially when targets are dense, making it difficult to meet real-time requirements.
[0005] Inefficient communication protocols further constrain system performance. Current mainstream vehicle communication standards primarily transmit raw data or low-level perception results, failing to achieve efficient compression of semantic-level information. For example, converting high-level cognitive information such as "a pedestrian is crossing 30 meters ahead" into a structured representation can significantly reduce communication load, but existing technologies lack lightweight semantic encoding mechanisms. Simultaneously, the lack of a unified cognitive framework for group decision-making leads to differing understandings of the same scenario among traffic participants, resulting in conflicts in collaborative behavior. A typical example is that vehicle A, based on local perception, determines it is safe to proceed, while vehicle B, based on different data sources, concludes there is a collision risk; such cognitive discrepancies can trigger systemic failures.
[0006] In summary, traditional solutions suffer from systemic deficiencies in terms of comprehensive perception, real-time response capabilities, depth of risk assessment, and communication efficiency. A semantic representation-based collaborative cognition method for traffic subjects is needed to achieve semantic-level fusion and efficient interaction of multi-source perception data, and to establish a collaborative decision-making mechanism that balances distributed computing and global optimization, thereby improving system safety, traffic efficiency, and reliability in complex and dynamic traffic environments. Summary of the Invention
[0007] The purpose of this invention is to provide a traffic subject collaborative cognition method based on semantic representation, comprising the following steps:
[0008] Step S1: Collect raw data including the state information of traffic scenarios and traffic entities, integrate it into perception information through data fusion technology, and transform the perception information into structured semantic representation information;
[0009] Step S2: Establish a multi-mode collaborative interaction architecture for different traffic scenarios, dynamically select the collaborative interaction architecture based on real-time traffic conditions, and conduct orderly information interaction according to the task.
[0010] Step S3: The peripheral traffic entities autonomously generate decisions, which are then optimized as a whole by the central traffic entity to achieve collaborative cognition and behavioral consistency among multiple traffic entities.
[0011] Furthermore, in step S1, data fusion technology is used to integrate the information into perceptual information, and the perceptual information is transformed into structured semantic representation information, specifically including:
[0012] The fusion algorithm model integrates state information from different sources into a unified form, and the recognition algorithm model identifies and marks key elements of traffic entities, organizing them into structured semantic representation information.
[0013] Furthermore, multi-modal collaborative interaction architectures include centralized collaborative interaction architectures, distributed collaborative interaction architectures, and layered collaborative interaction architectures.
[0014] Furthermore, the centralized collaborative interaction architecture has a central node. All edge traffic entities upload their perception information to the central node, which then performs global data processing, fusion, and cognitive computing before distributing it to each edge traffic entity.
[0015] Furthermore, the distributed collaborative interaction architecture does not have a central node. All traffic entities are equal and communicate directly with other neighboring entities to exchange perception and cognitive information, and negotiate with each other to achieve local collaboration.
[0016] Furthermore, the hierarchical collaborative interaction architecture divides the system into different layers, with nodes at each layer processing perceptual and cognitive information.
[0017] Further, step S3 includes:
[0018] Step S31: Each edge traffic entity predicts and models the vehicle-obstacle interaction game based on locally perceived semantic information and received collaborative cognitive information, and generates a safe, smooth trajectory plan that complies with traffic rules.
[0019] Step S32: The central traffic entity coordinates and optimizes the decisions of the peripheral traffic entities from a global perspective, resolves decision conflicts, and improves overall traffic efficiency;
[0020] Step S33: Establish an evaluation and feedback mechanism for decision-making effectiveness. By continuously collecting data on the actual implementation effects of traffic decisions, optimize decision model parameters, and achieve continuous improvement of the system.
[0021] Furthermore, the main components of edge transportation are intelligent vehicles.
[0022] Furthermore, the central traffic entity is either a roadside edge computing node or a regional traffic control center.
[0023] The beneficial effects of this invention are as follows:
[0024] 1. The present invention provides a traffic subject collaborative cognition method based on semantic representation, which breaks down information barriers and enhances perception capabilities: by establishing a unified semantic representation framework, multi-source heterogeneous traffic data is transformed into structured semantic representation information, realizing semantic mutual understanding between different traffic subjects and breaking down information barriers.
[0025] 2. This method employs lightweight semantic encoding and a dynamic collaborative structure (centralized, distributed, or hierarchical), which can promote the formation of collaborative and comprehensive cognition among multiple traffic entities and drive effective cooperation among them. By combining the autonomous decision-making of peripheral traffic entities with the overall optimization of the central traffic entity, a balance between local decision-making and global optimization is achieved, while simultaneously improving decision-making efficiency.
[0026] 3. This invention supports continuous system evolution: Through decision feedback and learning mechanisms, the system can continuously learn and optimize from real-world traffic scenarios, adapting to diverse traffic environments. The system developed based on this method exhibits excellent performance in both open-loop and closed-loop tests, demonstrating good scalability and adaptability. Attached Figure Description
[0027] Figure 1 This is a flowchart illustrating a traffic subject collaborative cognition method based on semantic representation according to the present invention.
[0028] Figure 2 This is a schematic diagram of the structure of a traffic subject collaborative cognition method based on semantic representation according to an embodiment of the present invention;
[0029] Figure 3 This is a schematic diagram of a highway ramp merging collaborative guidance scenario according to an embodiment of the present invention. Detailed Implementation
[0030] The present invention will be further described in detail below with reference to the accompanying drawings and using a scenario of coordinated guidance for merging of highway ramps as an example.
[0031] like Figure 1 The diagram shows a flowchart of a traffic subject collaborative cognition method based on semantic representation according to the present invention, which includes the following steps:
[0032] Step S1: Collect raw data including the state information of traffic scenarios and traffic entities, integrate it into perception information through data fusion technology, and transform the perception information into structured semantic representation information.
[0033] In this embodiment, multi-source state information of traffic scenes and traffic entities is collected through vehicle-mounted sensors and roadside equipment. This information is then integrated into standardized perception information using a fusion algorithm model. A recognition and labeling algorithm model is used to identify and label key elements of traffic entities, constructing structured semantic representation information to achieve semantic mutual understanding among traffic entities. A multi-mode collaborative cognitive architecture, including centralized, distributed, and hierarchical structures, is created. The optimal collaborative structure is dynamically selected based on real-time traffic conditions, constructing a collaborative sharing process from edge perception to global cognition. A collaborative cognitive mechanism combining local decision-making and global optimization is established. Edge traffic entities autonomously generate decisions based on semantic information, while central traffic entities coordinate and optimize from a global perspective. Finally, an evaluation and feedback mechanism enables continuous system evolution, achieving collaborative cognition and decision support for multiple traffic entities in complex environments.
[0034] Taking the scenario of coordinated guidance for highway ramp merging as an example, a collaborative cognition experiment was conducted. Vehicles interacted with each other, and the status information obtained from interacting with roadside units was verified, ultimately completing the scenario of coordinated guidance for ramp merging. Figure 3The diagram illustrates a reverse overtaking scenario, explained as follows: In the merging zone of a highway ramp and mainline, a connected vehicle (HV) from the ramp aims to merge into the mainline. Vehicles (M1, M2) on the mainline are traveling in the rightmost lane. Roadside edge computing nodes (RSUs) are deployed above the merging zone, equipped with cameras and millimeter-wave radar, to monitor traffic flow in the area. This is a typical high-risk scenario involving vehicles and roadside equipment, requiring precise coordination to avoid collisions and achieve a smooth merge.
[0035] By applying the semantic representation-based collaborative cognition method for traffic subjects disclosed in this invention, traffic semantic information is expressed in a structured manner, and a collaborative interaction architecture is constructed to complete multiple rounds of information interaction and decision-making, thereby realizing information sharing and collaborative cognition among traffic subjects.
[0036] In this embodiment, vehicle-mounted sensors and roadside equipment capture real-time vehicle and lane status information. A fusion algorithm model integrates status information from different sources into a unified format. Simultaneously, an identification and labeling algorithm model identifies and labels key elements of the traffic subject, transforming perceived information into structured semantic representations, such as <HV, intent, merge_into_mainline> and <M1, behavior, maintain_speed>. Those skilled in the art should understand that the selection of sensors and fusion, identification, and labeling algorithms needs to be flexibly adjusted according to specific traffic scenarios; this embodiment does not impose specific limitations.
[0037] Step S2: Design various collaborative structures (such as centralized, distributed, and hierarchical) for different traffic scenarios, and dynamically select the most suitable collaborative structure based on real-time traffic conditions, while conducting orderly information interaction according to the task.
[0038] Figure 2 This is a schematic diagram of a traffic subject collaborative cognition method based on semantic representation according to an embodiment of the present invention; step S2, which designs multiple collaborative structures and dynamically selects the most suitable collaborative structure based on real-time traffic conditions, mainly includes:
[0039] A centralized collaborative interaction architecture is established, with a central node. All edge traffic entities upload their perception information to the central node, which then performs global data processing, fusion, and cognitive computing before distributing it to each edge entity.
[0040] A distributed collaborative interaction architecture is established, with no central node. All traffic entities are of equal status and exchange perception and cognitive information with other neighboring entities through direct communication between traffic entities, achieving local collaboration through mutual consultation.
[0041] A hierarchical collaborative interaction architecture is established, which is a hybrid model that integrates the advantages of centralized and distributed systems. The system is divided into different layers, and nodes at different layers assume different responsibilities.
[0042] In the highway ramp merging collaborative guidance system of this embodiment, the implementation of step S2 relies on a highly intelligent dynamic collaborative architecture. This architecture uses roadside edge computing nodes (RSUs) as the collaborative core, and dynamically adjusts the collaborative structure by continuously monitoring multi-dimensional real-time parameters. The system collects four categories of key indicators in real time: traffic flow parameters including vehicle density, average speed, and lane occupancy in the merging zone; communication status parameters covering communication latency, packet loss rate, and available bandwidth; computing resource parameters involving the computing capabilities of RSUs and edge vehicles; and scenario complexity parameters including the number of conflict points, the diversity of traffic participant types, and environmental visibility. Based on these parameters, the system uses a logical decision-making model to select the collaborative structure, assessing the scenario as "local area, multi-subject, high real-time requirement," and therefore dynamically selects a hierarchical collaborative structure. The RSU, as the "center" of the area, undertakes the core coordination responsibility. Vehicles (V2V) also conduct local negotiation based on semantic information to confirm each other's intentions.
[0043] The collaborative process achieves closed-loop optimization through multiple rounds of semantic interaction. In the initial phase, on-ramp vehicles (HVs) encapsulate their merging intentions into semantic messages and upload them to the Traffic Management Unit (RSU). In the first round of processing, the RSU integrates multi-source information, performs conflict detection, and generates a preliminary coordination strategy. Subsequently, multiple rounds of interaction begin: the RSU issues coordination instructions to relevant vehicles (e.g., suggesting a slight acceleration on M2), and the vehicles immediately feed back the results to the RSU after execution. Based on the feedback data, the RSU reassesses the traffic conditions, optimizes and adjusts the strategy, and issues instructions in the second round. This iterative interaction continues until a safe and efficient merging scheme is formed, typically achieving optimal coordination after 2-3 rounds of adjustments. The entire multi-round interaction process forms a dynamic optimization closed loop, ensuring the system can adapt to real-time changes in the merging zone's traffic conditions.
[0044] Step S3: The peripheral traffic entities autonomously generate decisions, which are then optimized as a whole by the central traffic entity to achieve collaborative cognition and behavioral consistency among multiple traffic entities.
[0045] Specifically, it includes:
[0046] Step S31: Each edge traffic entity (such as intelligent vehicle) predicts and models the vehicle-obstacle interaction game based on locally perceived semantic information and received collaborative cognitive information, and generates a safe, smooth trajectory plan that complies with traffic rules.
[0047] Step S32: The central traffic entity (such as a roadside edge computing node or a regional traffic control center) coordinates and optimizes the decisions of multiple edge entities from a global perspective, resolves possible decision conflicts, and improves overall traffic efficiency.
[0048] Step S33: Establish an evaluation and feedback mechanism for decision-making effectiveness. By continuously collecting data on the actual implementation effects of traffic decisions, optimize decision model parameters, and achieve continuous improvement of the system.
[0049] In the highway ramp merging collaborative guidance system of this embodiment, the core of step three lies in realizing a collaborative mechanism between edge autonomous decision-making and central global optimization. This process begins with the edge traffic subject (i.e., intelligent connected vehicles) generating autonomous decisions based on locally perceived semantic information and received collaborative cognitive information. Taking the ramp vehicle (HV) as an example, its onboard computing unit first parses the semantic instructions issued by the RSU (e.g., <RSU, HV, merge_timing, {wait: 2.3s, target_gap: front}>), while simultaneously integrating real-time data from local sensors (e.g., distance to the vehicle in front, its own speed, etc.). By integrating game theory and prediction models, the HV constructs a vehicle-obstacle interaction game scenario, simulating the possible behavioral responses of surrounding vehicles (e.g., M2). Based on this, the vehicle generates a safe, smooth trajectory planning scheme that complies with traffic rules, specifically including longitudinal speed control (acceleration / deceleration curves) and lateral path planning (lane keeping or gradual merging), ensuring that the merging operation is completed within a predetermined time window. The entire decision-making process must meet strict real-time requirements.
[0050] The central traffic entity, namely the roadside edge computing node RSU in this example, coordinates and optimizes the decisions of multiple edge entities from a global perspective. The RSU continuously monitors the planned trajectories and real-time status of all relevant vehicles, identifying potential decision conflicts through conflict detection algorithms. When a risk of intersection is detected between the merging trajectory of the HV and the acceleration trajectory of M2, the RSU immediately initiates a coordination mechanism: calculating the optimal adjustment scheme using predictive control algorithms, such as fine-tuning the acceleration parameters of M2 or providing alternative merging gap suggestions for the HV; subsequently, optimization instructions are sent to the corresponding vehicles in the form of semantic messages. Central optimization not only resolves immediate conflicts but also improves overall efficiency through macro-level traffic flow optimization.
[0051] To ensure continuous system evolution, this step establishes an evaluation and feedback mechanism for decision-making effectiveness. The RSU collects data on the actual effectiveness of decision execution, including key indicators such as trajectory tracking deviation, minimum safe distance, merging time, and acceleration change rate. This data is uploaded to the regional cloud control center via an encrypted channel for evaluating decision-making effectiveness. This entire step, through efficient collaboration between the edge and center, achieves a balance between local decision-making autonomy and global optimization consistency.
[0052] In this embodiment, based on a semantic representation-based traffic subject collaborative cognition method disclosed in the present invention, a traffic subject collaborative cognition simulation scenario interaction verification is designed. Based on the Python environment, taking the highway ramp merging collaborative guidance scenario as an example, traffic information is semantically expressed, a collaborative interaction architecture is constructed, and a collaborative cognition simulation experiment is conducted to verify the effectiveness of the proposed method, thereby improving traffic efficiency and ensuring traffic safety.
[0053] The specific simulation experiment process in this embodiment is as follows:
[0054] (1) Experimental environment
[0055] Experiments were conducted using Python as the programming language in a Windows system environment.
[0056] (2) Experimental procedure
[0057] First, a simulation environment for a highway ramp merging collaborative guidance scenario is established. The highway ramp merging collaborative guidance scenario includes merging lane Lane 1 and main lane Lane 2. On main lane Lane 1, vehicles HV that will be merging on the ramp are set up, and vehicles M1 and M2 are set up on main lane Lane 2. The specific details are as follows... Figure 3 As shown. Each vehicle is 4m long, with an initial speed of 10m / s, and the acceleration during gear changes is 2.5m / s². 2 The time for the main vehicle HV to change lanes and merge is 1 second, ensuring that dangerous situations such as loss of control and lateral movement do not occur during lane changes or merges. The initial distance between the main vehicle HV and the merging point is set to 15m. The initial distance between M1 and M2 is randomly generated in a uniform distribution between 40m and 60m. When M2 is to the left of the merging point, the distance between it and the merging point is set to a negative number, and to a positive number when it is to the right. The distance is randomly generated in a uniform distribution between -5m and 20m, generating a total of 100 cases of two different types of distances.
[0058] The above scenario verification shows that when vehicles merge on ramps, they can achieve mutual avoidance and effective coordination through information exchange, thus improving road traffic safety. The final simulation results are shown in Table 1. In 100 case verifications, without semantic information exchange, the lead vehicle successfully completed the merge 90 times, and risk alarms occurred 59 times; after semantic information collaborative cognition, in the set simulation scenarios, the lead vehicle could complete overtaking in all cases, and the number of risk alarms decreased to 23.
[0059] Table 1 Summary of Simulation Experiment Results
[0060]
[0061] Based on the data results, this embodiment achieves collaborative cognition among traffic subjects in the scenario of coordinated guidance for highway ramp merging by transmitting semantic interaction messages, which ultimately increases the number of successful ramp merging processes and significantly reduces the number of safety risk alarms.
[0062] This invention enables semantic mutual understanding between different traffic entities, breaking down information barriers; it achieves a balance between local decision-making and global optimization, while improving decision-making efficiency, and has good scalability and adaptability.
Claims
1. A traffic subject collaborative cognition method based on semantic representation, characterized in that, Includes the following steps: Step S1: Collect raw data including the state information of traffic scenarios and traffic entities, integrate it into perception information through data fusion technology, and transform the perception information into structured semantic representation information; Step S2: Establish a multi-mode collaborative interaction architecture for different traffic scenarios, dynamically select the collaborative interaction architecture based on real-time traffic conditions, and conduct orderly information interaction according to the task. Step S3: The peripheral traffic entities autonomously generate decisions, which are then optimized as a whole by the central traffic entity to achieve collaborative cognition and behavioral consistency among multiple traffic entities.
2. The traffic subject collaborative cognition method based on semantic representation according to claim 1, characterized in that, In step S1, the data fusion technology is used to integrate the information into perceptual information, and the perceptual information is then transformed into structured semantic representation information. This specifically includes: The fusion algorithm model integrates state information from different sources into a unified form, and the recognition algorithm model identifies and marks key elements of traffic entities, organizing them into structured semantic representation information.
3. The traffic subject collaborative cognition method based on semantic representation according to claim 1, characterized in that, Multi-mode collaborative interaction architectures include centralized collaborative interaction architecture, distributed collaborative interaction architecture, and layered collaborative interaction architecture.
4. The traffic subject collaborative cognition method based on semantic representation according to claim 3, characterized in that, The centralized collaborative interaction architecture has a central node. All edge traffic entities upload their perception information to the central node, which then performs global data processing, fusion, and cognitive computing before distributing it to each edge traffic entity.
5. The traffic subject collaborative cognition method based on semantic representation according to claim 3, characterized in that, The distributed collaborative interaction architecture has no central node, and all traffic entities are equal. They communicate directly with other neighboring entities to exchange perception and cognitive information and negotiate with each other to achieve local collaboration.
6. The traffic subject collaborative cognition method based on semantic representation according to claim 3, characterized in that, The hierarchical collaborative interaction architecture divides the system into different layers, with nodes at each layer processing perceptual and cognitive information.
7. The traffic subject collaborative cognition method based on semantic representation according to claim 1, characterized in that, Step S3 includes: Step S31: Each edge traffic entity predicts and models the vehicle-obstacle interaction game based on locally perceived semantic information and received collaborative cognitive information, and generates a safe, smooth trajectory plan that complies with traffic rules. Step S32: The central traffic entity coordinates and optimizes the decisions of the peripheral traffic entities from a global perspective, resolves decision conflicts, and improves overall traffic efficiency; Step S33: Establish an evaluation and feedback mechanism for decision-making effectiveness. By continuously collecting data on the actual implementation effects of traffic decisions, optimize decision model parameters, and achieve continuous improvement of the system.
8. The traffic subject collaborative cognition method based on semantic representation according to claim 7, characterized in that, The main body of the edge transportation system is intelligent vehicles.
9. The traffic subject collaborative cognition method based on semantic representation according to claim 7, characterized in that, The central traffic entity is either a roadside edge computing node or a regional traffic control center.
Citation Information
Cited By
Zero-contact traffic monitoring method, system and equipment and computer medium
CN121904993A