Traffic congestion prediction method, system, electronic device and computer program product

By constructing a time-series heterogeneous graph and combining specific prior knowledge and multimodal data to dynamically update edge weights, the problem of difficulty in representing the relationships between heterogeneous entities in traffic networks is solved, resulting in more accurate traffic prediction and improving the accuracy and reliability of traffic congestion prediction.

CN121191327BActive Publication Date: 2026-03-27CETC NEW SMART CITY RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing traffic congestion prediction methods are unable to effectively characterize the relationships between various heterogeneous entities in the traffic network, resulting in insufficient prediction accuracy and reliability.

Method used

By introducing specific prior knowledge and multimodal traffic data, the feature representation and association strength of traffic entities are determined, a temporal heterogeneous graph is constructed, and the edge weights are dynamically updated to reflect the association between entities, thereby achieving accurate traffic congestion prediction.

Benefits of technology

It improves the accuracy and reliability of traffic congestion prediction, enabling more refined quantification of traffic entities and the dynamic relationships between them, and providing more accurate traffic condition analysis and prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121191327B_ABST
    Figure CN121191327B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of traffic control, and provides a traffic congestion prediction method, system, electronic device and computer program product, wherein the method comprises: determining a plurality of types of traffic entities in a target area and a specific type of traffic correlation relationship between each pair of traffic entities based on obtained specific prior knowledge; determining a feature representation of each traffic entity and a traffic correlation strength between each pair of traffic entities based on obtained multi-modal traffic data and the traffic correlation relationship; taking the feature representation as a node feature representation of a traffic entity node corresponding to the plurality of types of traffic entities, and taking the traffic correlation strength as an edge weight of a traffic correlation edge between the traffic entity nodes representing the traffic correlation relationship, to obtain a time sequence heterogeneous graph containing the traffic entity nodes and the traffic correlation edges, so as to predict traffic congestion information of the target area. The scheme quantitatively characterizes heterogeneous entities and the correlation relationship between the entities, and improves the accuracy and reliability of traffic congestion prediction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of traffic control, and particularly relates to a traffic congestion prediction method and system, an electronic device, and a computer program product. BACKGROUND

[0002] With the rapid growth of urbanization and the number of motor vehicles, traffic congestion has become a common phenomenon. By accurately predicting traffic congestion, decision support can be provided for traffic management, path planning, and other aspects. However, existing traffic congestion prediction methods are difficult to effectively represent the correlation between various heterogeneous entities in the traffic network, which restricts the accuracy and reliability of traffic congestion prediction. SUMMARY

[0003] The embodiments of the present application provide a traffic congestion prediction method, system, electronic device, and computer program product to solve the problem that the existing technology is difficult to effectively represent the correlation between various heterogeneous entities in the traffic network, which restricts the accuracy and reliability of traffic congestion prediction.

[0004] The first aspect of the embodiments of the present application provides a traffic congestion prediction method, comprising:

[0005] acquiring multi-modal traffic data corresponding to specific prior knowledge and a target area;

[0006] determining a specific type of traffic correlation between each pair of traffic entities in the target area based on the specific prior knowledge;

[0007] determining a feature representation of each traffic entity and a traffic correlation strength between each pair of traffic entities based on the multi-modal traffic data and the traffic correlation;

[0008] taking the feature representation as a node feature representation of a traffic entity node corresponding to the traffic entity of the plurality of types, and taking the traffic correlation strength as an edge weight of a traffic correlation edge between the traffic entity nodes representing the traffic correlation, to obtain a time-series heterogeneous graph containing the traffic entity nodes of the plurality of types and the traffic correlation edges of the plurality of types;

[0009] predicting traffic congestion information of the target area based on the time-series heterogeneous graph.

[0010] The second aspect of the embodiments of the present application provides a traffic congestion prediction system, comprising:

[0011] an acquisition module configured to acquire multi-modal traffic data corresponding to specific prior knowledge and a target area;

[0012] The first determining module is configured to determine a plurality of types of traffic entities in the target area and a specific type of traffic correlation between each pair of the traffic entities based on the specific prior knowledge.

[0013] The second determining module is configured to determine a feature representation of each of the traffic entities and a traffic correlation strength between each pair of the traffic entities based on the traffic data in multiple modalities and the traffic correlation.

[0014] The constructing module is configured to take the feature representation as a node feature representation of a traffic entity node corresponding to the plurality of types of traffic entities, take the traffic correlation strength as an edge weight of a traffic correlation edge between the traffic entity nodes indicating the traffic correlation, and obtain a time-series heterogeneous graph including the plurality of types of traffic entity nodes and the plurality of types of traffic correlation edges.

[0015] The predicting module is configured to predict traffic congestion information of the target area based on the time-series heterogeneous graph.

[0016] The third aspect of the embodiments of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the method according to the first aspect when executing the computer program.

[0017] The fourth aspect of the embodiments of the present application provides a computer program product, which includes a computer program, and the computer program implements the steps of the method according to the first aspect when executed by a processor.

[0018] The fifth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program implements the steps of the method according to the first aspect when executed by a processor.

[0019] As can be seen from the above, the present application introduces specific prior knowledge, accurately defines a plurality of types of traffic entities in a target area and a traffic correlation between traffic entities, determines a feature representation of each traffic entity and a traffic correlation strength between each pair of traffic entities in combination with traffic data in multiple modalities, and constructs a time-series heterogeneous graph including a plurality of types of traffic entity nodes and a plurality of types of traffic correlation edges. The time-series heterogeneous graph combines the real-time features of the multi-modal traffic data to finely quantify the representation of heterogeneous traffic entities and the dynamic correlation between the heterogeneous traffic entities. On this basis, the traffic congestion information of the target area is predicted based on the constructed time-series heterogeneous graph, that is, the time-series heterogeneous graph with rich information is used for prediction, which effectively improves the accuracy and reliability of traffic congestion prediction. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only some of the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0021] Figure 1 is a flow chart of a traffic congestion prediction method provided by an embodiment of the present application;

[0022] Figure 2 is a topological structure schematic diagram containing traffic entity nodes and traffic associated edges provided by an embodiment of the present application;

[0023] Figure 3 is a structural diagram of a traffic congestion prediction system provided by an embodiment of the present application;

[0024] Figure 4 is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0025] In the following description, for the purpose of explanation and not limitation, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary detail.

[0026] It should be understood that the term "comprising" as used in the specification and the appended claims indicates the presence of the recited features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0027] It should also be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0028] It should be further understood that the term "and / or" as used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0029] As used in the specification and the appended claims, the term "if' can be interpreted as meaning "when," or "upon," or "in response to determining," or "in response to detecting" depending on the context. Similarly, the phrase "if it is determined" or "if [the recited condition or event] is detected" can be interpreted as meaning "upon determining" or "in response to determining" or "upon detecting [the recited condition or event]" or "in response to detecting [the recited condition or event]" depending on the context.

[0030] In particular implementations, the terminals described in the embodiments of the present application include, but are not limited to, other portable devices such as mobile telephones, laptop computers, or tablet computers with touch-sensitive surfaces (e.g., touch screen displays and / or touch pads). It should also be understood that, in some embodiments, the device is not a portable communication device, but is a desktop computer with a touch-sensitive surface (e.g., a touch screen display and / or a touch pad).

[0031] In the following discussion, a terminal that includes a display and a touch-sensitive surface is described. It should be understood, however, that a terminal can include one or more other physical user-interface devices, such as a physical keyboard, a mouse and / or a joystick.

[0032] The terminal supports a variety of applications, such as one or more of the following: a drawing application, a presentation application, a word processing application, a website creation application, a disk authoring application, a spreadsheet application, a game application, a telephone application, a video conferencing application, an e-mail application, an instant messaging application, a workout support application, a photo management application, a digital camera application, a digital camcorder application, a web browsing application, a digital music player application, and / or a digital video player application.

[0033] The various applications that can be executed on the terminal can use at least one common physical user-interface device, such as the touch-sensitive surface. One or more functions of the touch-sensitive surface, as well as the display of information on the terminal, can be adjusted and / or changed in accordance with the application that is being executed by the terminal. In this way, the user experiences a consistent user interface when switching between various applications that are executed on the terminal.

[0034] It should be understood that the size of the serial number of each step in the embodiments does not mean the order of execution, the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0035] In order to illustrate the technical solutions described in the present application, the following will be described by specific embodiments.

[0036] Referring to Figure 1 , Figure 1 is a flowchart of a traffic congestion prediction method provided by an embodiment of the present application. As shown in Figure 1 , a traffic congestion prediction method includes the following steps:

[0037] Step 101, acquiring specific prior knowledge and multi-modal traffic data corresponding to a target area.

[0038] The target area is a geographical space range with a clear boundary to be predicted.

[0039] The specific prior knowledge refers to regularized information in the field of traffic for defining entity types and association relationships, including entity definition information and association relationship definition information. The specific prior knowledge is obtained based on traffic engineering theory, industry standards and actual scene experience summary.

[0040] Among them, the entity definition information defines various traffic entities, and the association relationship definition information clearly defines the association definition information of the possible traffic association relationship between traffic entities.

[0041] In some embodiments, a traffic entity refers to an individual in a traffic scene that has independent attributes and participates in traffic operation, including but not limited to the following types: vehicle, road segment, traffic event, sensor, and traffic light.

[0042] In some embodiments, the entity definition information also defines the core attributes corresponding to each type of traffic entity. For example, the core attributes of a vehicle include speed, license plate number, etc., the core attributes of a road segment include length, number of lanes, real-time density, etc., the core attributes of a traffic event include type, occurrence time, location, etc., the core attributes of a sensor include traffic flow, rainfall, visibility, etc., and the core attributes of a traffic light include phase, cycle, and phase duration, etc. Subsequently, before generating node feature representation, target features can be filtered based on these core attributes.

[0043] In some embodiments, the association relationship definition information includes but is not limited to the following types: following relationship definition information, containing relationship definition information, causal relationship definition information, subsidiary relationship definition information, and control relationship definition information. These definition information respectively defines the following relationship, the containing relationship, the causal relationship, the subsidiary relationship and the control relationship.

[0044] Correspondingly, the traffic association relationship refers to the association attributes between entities based on spatial location or functional logic, including but not limited to the following types: following relationship between vehicles, containing relationship between road segments and vehicles, causal relationship between traffic events and road segments, subsidiary relationship between sensors and road segments, and control relationship between traffic lights and road segments.

[0045] The multi-modal traffic data refers to traffic-related data of different collection methods and formats, including visual and time-series data, text data, sensor numerical data, and signal light state data, covering real-time running state, control logic, and dynamic events in the traffic scene.

[0046] The visual and time-series data refers to image pictures collected by image collection devices such as traffic monitoring cameras, which are used to extract vehicle trajectories and identify traffic events.

[0047] The text data refers to document information such as traffic event reports and traffic control notices published by traffic management departments. For example, traffic control notices formulated for road collapse, construction maintenance, or large-scale activities.

[0048] The sensor numerical data refers to physical measurement values collected by various sensors deployed along the road. For example, vehicle flow monitored by ground coils, speed recorded by radar speedometers, rainfall and visibility collected by weather sensors.

[0049] The signal light state data refers to real-time control information obtained from traffic signal control systems, mainly including the current phase (red / yellow / green) of the signal light, the remaining time of the phase, the signal cycle, and the like.

[0050] The multi-modal traffic data is obtained to define specific prior knowledge of the basic components (traffic entities and traffic correlation relationships) in the traffic scene, and to reflect the traffic state of the target area.

[0051] The acquisition operation provides a solid rule basis and data basis for subsequent graph construction and traffic congestion prediction, ensuring that the definition of traffic entities and traffic correlation relationships conforms to the traffic scene logic and avoids ruleless entity identification. The multi-modal data covers multi-dimensional influencing factors in the traffic scene, and can capture congestion correlation information more comprehensively than single data.

[0052] In some embodiments, the obtained multi-modal traffic data is data that has been preprocessed.

[0053] In step 102, based on the specific prior knowledge, a plurality of types of traffic entities in the target area and a specific type of traffic correlation relationship between each pair of traffic entities are determined.

[0054] According to the specific prior knowledge, a plurality of types of traffic entities contained in the target area and the traffic correlation relationship between the associated traffic entities are determined, that is, the target area is represented by structured information with clear semantics, and the construction of the time-series heterogeneous graph is promoted.

[0055] In some embodiments, the determining, based on the specific prior knowledge, of the multiple types of traffic entities in the target region and the specific types of traffic correlation between each pair of the traffic entities comprises: identifying the multiple types of the traffic entities in the target region according to entity definition information in the specific prior knowledge; and determining the traffic correlation between each pair of the traffic entities according to correlation definition information in the specific prior knowledge; the correlation definition information comprises follow-up relationship definition information, containment relationship definition information, cause-effect relationship definition information, subsidiary relationship definition information, and control relationship definition information.

[0056] In some embodiments, when determining the traffic entities contained in the target region, the traffic entities can be determined in combination with the multi-modal traffic data of the target region and deployment planning information. The deployment planning information comprises a regional map, a device deployment list, and the like.

[0057] In some embodiments, according to the entity definition information in the specific prior knowledge, vehicles are identified from the visual and time-series data, traffic events are identified from the visual and time-series data and text data, road segments are identified from electronic maps, and signal lights and sensors are confirmed from a device registration list.

[0058] According to the correlation definition information in the specific prior knowledge, the traffic correlation between each pair of the traffic entities is further determined. The follow-up relationship definition information indicates that there is a follow-up relationship between vehicles that follow each other in space, the containment relationship definition information indicates that there is a containment relationship between a vehicle and a road segment on which the vehicle is located, the cause-effect relationship definition information indicates that there is a cause-effect relationship between a traffic event and an affected road segment, the subsidiary relationship definition information indicates that there is a subsidiary relationship between a sensor and a monitored road segment, and the control relationship definition information indicates that there is a control relationship between a signal light and a controlled road segment.

[0059] By using the predefined entity rules and relationship rules, the traffic entities and the traffic correlation related to traffic congestion in the target region are accurately identified, the confusion in entity identification is avoided, and the rationality of the correlation is ensured, so that the feature representation determination and the traffic correlation strength calculation in the subsequent steps can be adaptively fitted.

[0060] In step 103, based on the multi-modal traffic data and the traffic correlation, the feature representation of each traffic entity and the traffic correlation strength between each pair of the traffic entities are determined.

[0061] The feature representation refers to a vector used to quantitatively describe the core attributes of a traffic entity.

[0062] The traffic correlation strength refers to a numerical value used to quantitatively describe the correlation closeness between two traffic entities having a traffic correlation.

[0063] According to the multi-modal traffic data, a numerical feature representation is generated for each traffic entity, and a dynamic quantitative traffic correlation strength is given to each pair of traffic entities with a traffic correlation relationship, and the multi-modal traffic data and the identified traffic correlation relationship are converted into machine-readable and structured features rich in semantics.

[0064] In some embodiments, the multi-modal traffic data and the traffic correlation relationship are used to determine the feature representation of each traffic entity and the traffic correlation strength between each pair of traffic entities, including: performing feature extraction on the multi-modal traffic data to obtain initial features of the traffic data of each modality; mapping the initial features of each modality to a feature space with a set dimension to obtain multi-modal set dimension features; selecting target features for each traffic entity from the multi-modal set dimension features that match the feature representation requirements of the entity type of the traffic entity, and performing feature fusion based on the target features to obtain the feature representation of the traffic entity; and selecting target data for each pair of traffic entities from the multi-modal traffic data that matches the quantitative requirements of the traffic correlation relationship between the traffic entities, and calculating the traffic correlation strength between the traffic entities based on the target data.

[0065] After performing exclusive feature extraction on the traffic data of each modality according to the modality type, the original feature vector corresponding to the traffic data of each modality, i.e., the initial feature, is obtained.

[0066] In some embodiments, visual and temporal initial features are extracted from visual and temporal data .

[0067] In some embodiments, a target detection model such as YOLOv8 is used to identify each vehicle from visual and temporal data, and a multi-target tracking algorithm such as the Deep Simple Online and Realtime Tracking (DeepSORT) algorithm is used to realize continuous tracking of each vehicle to obtain a spatiotemporal trajectory sequence of each vehicle. The spatiotemporal trajectory sequence is input into a convolutional neural network (CNN) model, such as a temporal convolutional neural network (TCNN) model with a convolution kernel of 3x3 and a step size of 1, to obtain vehicle trajectory initial features output by the model , such as 256-dimensional vehicle trajectory initial features.

[0068] In some embodiments, the road area is segmented from the visual and time series data, and the density value is calculated by the proportion of vehicle pixels. After smoothing processing by a sliding window (such as a 5-minute window), the density value is input into a TCNN model to obtain the initial feature of the road segment density output by the model , such as the initial feature of the road segment density of 64 dimensions.

[0069] The initial feature of the vehicle trajectory and the initial feature of the road segment density are dimensionally aligned, and the dimensionally aligned initial feature of the vehicle trajectory and the initial feature of the road segment density are spliced to obtain the initial visual time series feature .

[0070] In some embodiments, a pre-trained language model is used to encode text data such as traffic control notices and accident reports to obtain initial text semantic features. For example, a Bidirectional Encoder Representations from Transformers (BERT) model is used to encode the text data. When processing text, the BERT model adds a special token [CLS] at the initial position of the text. This token is used to summarize the global semantic information of the entire text through the bidirectional attention mechanism of the model. Accordingly, the output vector corresponding to the [CLS] token is used as the initial text semantic feature.

[0071] In some embodiments, . Wherein, is text data, such as an event description text that lane 1 of road segment A is closed due to road construction, is the initial text semantic feature.

[0072] In some embodiments, the sensor numerical data such as traffic volume, rainfall, and visibility are standardized by Z-score to eliminate dimensional differences. The time series data (such as continuous 30 5-minute traffic volumes) obtained by the standardization are input into a TCNN model to obtain the initial sensor time series feature output by the TCNN model .

[0073] In some embodiments, the phase (red / green / yellow) of the signal light is encoded into a 3-dimensional one-hot vector, and is spliced with the remaining time of the phase and the green ratio (green light duration / period) to obtain the initial signal light state feature of 16 dimensions .

[0074] In some embodiments, for each modality of traffic data, a modality-specific feature extractor, such as a feature encoder, is designed to extract initial features that preserve the core information of the modality.

[0075] In some embodiments, each individual is assigned a unique identifier at the time of feature extraction, such as a vehicle's license plate number, to facilitate the selection of matching features before feature fusion.

[0076] In some embodiments, due to the large difference in the dimensions of initial features from different modalities (e.g., 768-dimensional text semantic initial features and 128-dimensional sensor time series initial features), direct feature fusion is not possible. Therefore, the initial features from multiple modalities are mapped to a feature space with a set dimension, resulting in set-dimensional features for each modality, such as set-dimensional visual time series features , set-dimensional text semantic features , set-dimensional sensor time series features , and set-dimensional signal light state features .

[0077] In some embodiments, all initial features are uniformly mapped to a feature space with a set dimension (e.g., 256 dimensions) through a modality alignment layer. wherein, represents different modalities, is the initial feature of different modalities, is the set-dimensional feature of different modalities, is the learnable mapping matrix of different modalities, is the bias of different modalities.

[0078] In some embodiments, all initial features are uniformly mapped to a feature space with a set dimension (e.g., 256 dimensions) through the learnable mapping matrix contained in the modality alignment layer.

[0079] Uniformly mapping initial features of different dimensions to a set dimension breaks down the dimensional barriers of visual, text, sensor, and other modalities, providing feasibility for cross-modal feature fusion and avoiding the problem of features not being integrated due to inconsistent dimensions.

[0080] After mapping, for each traffic entity, the target feature that matches the feature representation requirement of the entity type of the traffic entity is selected from the set-dimensional features of multiple modalities. The set-dimensional features of multiple modalities may contain multiple traffic entity nodes of the same type, so when selecting the target feature, the entity-specific identifier of the traffic entity should also be considered to avoid selecting the wrong target feature, such as selecting the target feature for vehicle A1 and avoiding selecting the target feature for vehicle A2.

[0081] In some embodiments, the feature representation requirement can be a core attribute defined in the entity definition information.

[0082] In some embodiments, the target feature matching the feature representation requirement of the vehicle includes a set-dimension visual timing feature and a set-dimension sensor timing feature; the target feature matching the feature representation requirement of the road section includes a set-dimension visual timing feature, a set-dimension sensor timing feature, and a set-dimension signal light state feature; the target feature matching the feature representation requirement of the traffic event includes a set-dimension text semantic feature and a set-dimension sensor timing feature; the target feature matching the feature representation requirement of the sensor includes a set-dimension sensor timing feature; and the target feature matching the feature representation requirement of the signal light includes a set-dimension signal light state feature.

[0083] The feature fusion is performed based on the target feature of each traffic entity to obtain a feature representation of the traffic entity.

[0084] Through the above feature extraction, uniform dimension mapping, and targeted feature fusion operations, the modality information can be comprehensively retained, and the feature can be accurately matched with the representation requirement of the traffic entity, thereby improving the integrity and adaptability of the feature representation of the traffic entity.

[0085] In some embodiments, the selecting, for each traffic entity, a target feature matching the feature representation requirement of the entity type from the set-dimension features of multiple modalities and performing feature fusion based on the target feature to obtain the feature representation of the traffic entity includes: selecting, for each traffic entity, the target feature matching the feature representation requirement of the entity type from the set-dimension features of multiple modalities; determining, according to a target modality corresponding to the target feature, a non-target modality not matching the feature representation requirement; obtaining a mask feature corresponding to the non-target modality; and performing feature fusion on the target feature and the mask feature to obtain the feature representation corresponding to the traffic entity.

[0086] After selecting the target feature for each traffic entity, according to the target modality corresponding to the target feature, a non-target modality not matching the feature representation requirement of the traffic entity can be determined in multiple modalities. The set-dimension feature corresponding to the non-target modality does not contribute to the feature representation of the traffic entity, but if it is directly excluded, it will destroy the structural consistency of the multi-modal feature (such as dimension alignment and modality position correspondence). Therefore, the feature of the non-target modality is masked, for example, a zero vector consistent with the set-dimension or a specific mask value vector is generated, and the zero vector consistent with the set-dimension or the specific mask value vector is used as a mask feature.

[0087] The target feature and the mask feature of each traffic entity are subjected to feature fusion, and the feature representation of the traffic entity is obtained.

[0088] In some embodiments, the target feature and the mask feature of each traffic entity are spliced and nonlinearly fused to obtain the feature representation of the traffic entity. Specifically, . 、 、 、 is the feature to be fused in the set dimension for each modality, which includes the target feature and may also include the mask feature, is an activation function, is a modality fusion mapping matrix, is a bias term in the feature fusion process, is the feature representation of each traffic entity, refers to the type of traffic entity, including vehicles , road segments , traffic events , sensors , traffic lights , refers to different traffic entity individuals under the same entity type, such as refers to the feature representation of the 3rd vehicle, refers to the feature representation of the 5th sensor.

[0089] The feature representation of the traffic entity is generated by feature fusion, realizing the collaborative modeling of multi-modal semantics.

[0090] The target feature is accurately selected for the traffic entity to match the feature representation requirement, which can reduce the interference of redundant features. Mask processing is performed on the non-target modality and it participates in the fusion, which can ensure that the generated feature representation is matched with the entity attribute, and the structural consistency of the feature representation is maintained, thereby improving the accuracy and effectiveness of the feature representation of the traffic entity.

[0091] The traffic correlation strength is introduced to quantify the correlation tightness between each pair of traffic entities (such as vehicles and road segments, road segments and traffic lights).

[0092] From the multi-modal traffic data, the corresponding target data (such as the traffic record in the visual and time series data, the signal duration in the signal light state data) is selected according to the quantitative requirement of the traffic correlation relationship between each pair of traffic entities (such as the number of passes and the length of stay of vehicles and road segments, the signal timing matching degree of road segments and traffic lights). Based on the selected target data, the traffic correlation strength between each pair of traffic entities is obtained by statistical analysis or correlation calculation.

[0093] In some embodiments, the traffic correlation strength between vehicles and vehicles is calculated by . Wherein, is the vehicle With vehicles distance between vehicles and vehicles and vehicles instantaneous speed, For distance attenuation parameters, For velocity decay parameters, This refers to the traffic association strength reflected in the car-following relationship between vehicles. The traffic association strength of the car-following relationship is negatively correlated with the distance between vehicles and the speed difference. The smaller the distance between vehicles and the closer their speeds, the higher the traffic association strength and the stronger the car-following relationship.

[0094] In some embodiments, the traffic association strength of the inclusion relationship between road segments and vehicles is determined by... Calculation. Among them, For vehicles Entering the section Time, This is the maximum dwell time threshold, for example, 300 seconds. This represents the traffic association strength reflected by the inclusion relationship between road segments and vehicles. This traffic association strength is positively correlated with the dwell time of vehicles on the road segment; the longer the dwell time, the higher the traffic association strength and the stronger the inclusion relationship.

[0095] In some embodiments, the traffic association strength between traffic events and road segments is determined by... Calculation. Among them, The semantic matching degree between traffic events and road segments (e.g., the semantic matching degree between construction events and main roads is 0.8, and the semantic matching degree with branch roads is 0.3). The straight-line distance between the traffic incident and the road segment. For spatial attenuation scale, This represents the strength of traffic association, reflecting the causal relationship between a traffic event and a road segment. This traffic association strength is positively correlated with the semantic matching degree of the traffic event and negatively correlated with distance. Higher semantic matching degree and closer distance result in a stronger traffic association and a stronger causal relationship.

[0096] In some embodiments, the traffic association strength between the sensor and the road segment is determined by... Calculation. Among them, This represents the shortest spatial distance between the sensor and the center point of the road segment. For sensor coverage ratio, and For adjustable weight parameters, This refers to the traffic association strength reflected in the affiliation between the sensor and the road segment. This traffic association strength reflects the coverage and data correlation of the sensor's observations of that road segment. The closer the distance, the higher the coverage ratio, the higher the traffic association strength, and the stronger the affiliation.

[0097] In some embodiments, the traffic correlation strength of the signal light and the control relationship between the road segments is calculated by . The phase weight of the signal light controls the road segment, with a value range of [0, 1], The distance from the signal light to the entrance of the road segment is used to reflect the control delay, And is a learnable adjustment parameter, The traffic correlation strength reflected by the control relationship between the signal light and the road segment. The control relationship represents the control effect of the signal light on the traffic state of the corresponding road segment, and reflects the signal control strength and the traffic flow response sensitivity. The closer the distance, the higher the phase weight, the higher the traffic correlation strength, and the stronger the control relationship.

[0098] Design dedicated calculation logic for different correlation relationship types, such as following distance for following relationship and phase for control relationship, so that the traffic correlation strength can truly reflect the correlation closeness between traffic entities, and provide reliable basis for the edge weight required for subsequent construction of the time-varying heterogeneous graph.

[0099] Step 104, taking the feature representation as the node feature representation of the traffic entity nodes corresponding to the multiple types of traffic entities, and taking the traffic correlation strength as the edge weight of the traffic correlation edges between the traffic entity nodes, to obtain a time-varying heterogeneous graph containing multiple types of traffic entity nodes and multiple types of traffic correlation edges.

[0100] The time-varying heterogeneous graph is a dynamic graph structure containing multiple types of nodes and multiple types of edges, and the state of the nodes and edges or the relationship between them will change over time.

[0101] In this application, each entity individual in various types of traffic entities (such as vehicles, road segments, traffic events, sensors, signal lights, etc.) is mapped to a traffic entity node in the time-varying heterogeneous graph, and the feature representation of each entity individual obtained by the multi-modal feature extraction, screening and fusion in the foregoing is directly taken as the node feature representation of each specific traffic entity node of the corresponding type, to embody the exclusive attribute information of each node.

[0102] At the same time, each type of traffic correlation (such as following relationship, containing relationship, causal relationship, subsidiary relationship, control relationship) between each pair of traffic entities (such as vehicle and vehicle, vehicle and road segment, traffic event and road segment, sensor and road segment, signal light and road segment, etc.) is mapped to a traffic correlation edge in the graph, and the traffic correlation strength calculated by the target data is taken as the edge weight of the corresponding traffic correlation edge, to quantitatively reflect the closeness of the correlation relationship between the nodes.

[0103] Accordingly, we obtain various types of traffic entity nodes (vehicle nodes, road segment nodes, traffic event nodes, sensor nodes, traffic light nodes) and various types of traffic-related edges (carriage-following edges). , including edges Causal edge , Subsidiary side Control edge The graph is a temporally heterogeneous graph where traffic entity nodes carry node feature representations and traffic association edges carry traffic association strengths. Traffic association edges are directed edges.

[0104] like Figure 2 As shown, Figure 2 This is a schematic diagram of a topology structure including traffic entity nodes and traffic-related edges provided in an embodiment of this application. Figure 2 It includes various types of traffic entity nodes, namely vehicle node 1, vehicle node 2, road segment node, traffic event node, sensor node, and traffic light node. Figure 2 It also includes various types of traffic-related edges, namely, car-following edges between vehicle node 1 and vehicle node 2, containment edges between vehicle node 1 and road segment nodes, containment edges between vehicle node 2 and road segment nodes, causal edges between traffic event nodes and road segment nodes, dependent edges between sensor nodes and road segment nodes, and control edges between traffic light nodes and road segment nodes. These traffic entity nodes and interactive edges constitute a network such as... Figure 2 The topology shown.

[0105] The state of traffic entities and the traffic relationships between them in a traffic scenario will change dynamically (e.g., changes in vehicle position lead to updates in the car-following correlation strength, and changes in road segment density change the causal correlation strength between them and traffic events). Therefore, the node feature representation of traffic entity nodes and the edge weights of traffic-related edges in the temporal heterogeneous graph will also change dynamically.

[0106] In the temporal heterogeneous graph of this application, edge weights are not only affected by the spatial structure and semantic relationships between traffic entities (such as the spatial distance of vehicle following and the causal semantics of traffic events and road segments), but also need to be dynamically adjusted as the interaction time of entities changes. For example, if two vehicles maintain a short distance and similar speed for a short period of time, the edge weight between their vehicle nodes should be enhanced; if a sensor does not report data for a long time, the edge weight between the sensor node and the corresponding road segment node should be weakened; after the traffic light phase changes, the edge weight between the traffic light node and the corresponding road segment node needs to be updated in real time.

[0107] This application defines a time-decay-based adaptive edge weight update function, enabling edge weights to automatically adjust according to the time difference of traffic entity interactions, thus achieving continuous evolution of temporal heterogeneous graphs from static weighted graphs to dynamic temporal heterogeneous graphs. The time-decay-based adaptive edge weight update function is as follows: . Wherein, representing a node with a node a time interval of the last interaction, i.e. a local time difference; is a time decay function, which can be set as an exponential decay, to quantify the decay effect of time on edge weights; is a time fusion factor, which controls the fusion ratio of edge weights before and after updating, ; representing a global time step or sampling interval, i.e. the time interval from the last update to the current, used for historical weight indexing; is the edge weight before updating, is the edge weight after updating.

[0108] The adaptive update of edge weights is realized through time decay, ensuring that the time-heterogeneous graph "moves" in time and reflects the dynamic association between traffic entities, providing more accurate dynamic graph data support for subsequent traffic state analysis and congestion prediction.

[0109] The time-heterogeneous graph of the present application is a dynamic graph structure containing multiple types of nodes, multiple types of edges, and attributes of nodes and edges changing over time, providing dynamic and comprehensive graph data basis for subsequent traffic state analysis or event prediction based on graph structure.

[0110] Step 105, predicting traffic congestion information of the target area based on the time-heterogeneous graph.

[0111] After obtaining the time-heterogeneous graph that can reflect the spatio-temporal association relationship between traffic entities in real time, traffic congestion prediction can be performed based on the time-heterogeneous graph, and the traffic congestion information of the target area is obtained accordingly.

[0112] Traditional static homogeneous graphs can only depict single-type entities and fixed edge weights, and cannot realize the characteristics of traffic state changing over time. The time-heterogeneous graph of the present application can integrate dynamic information of traffic scenes (such as vehicle position update, event influence decay, signal light phase switching) in real time, so that the model can more accurately restore the dynamic evolution process of traffic state, and finally output more realistic traffic congestion prediction information, providing support for traffic control and diversion decisions.

[0113] In some embodiments, the predicting the traffic congestion information of the target area based on the time-evolving heterogeneous graph comprises: performing weighted aggregation on the node feature representation of each traffic entity node and the node feature representation of neighbor nodes connected by the traffic-related edges based on the topology of the time-evolving heterogeneous graph and the edge weights of the traffic-related edges, to generate a global traffic feature representation representing the states and mutual influences of all traffic entities in the target area; and predicting the traffic congestion information of the target area in a target time period based on the global traffic feature representation.

[0114] The feature weighted aggregation is performed based on the topology of the time-evolving heterogeneous graph and the edge weights of the traffic-related edges. The topology defines the connection relationships between the traffic entity nodes (e.g., the association between vehicle nodes and road segment nodes, and the connection between road segment nodes and signal nodes), and the edge weights are used to quantify the closeness of the connection relationships between different traffic entity nodes (e.g., the higher the edge weight of the following edge between two vehicles, the closer the following of the two vehicles). The edge weights are dynamically updated over time.

[0115] In some embodiments, the weighted aggregation is performed on the node feature representation of each traffic entity node and the node feature representation of neighbor nodes connected by the traffic-related edges based on the edge weights.

[0116] The neighbor nodes with larger edge weights have higher weights in the aggregation process, so as to highlight the influence of the strongly connected nodes on the current node.

[0117] The weighted aggregation generates a global traffic feature representation that can comprehensively represent the real-time states of all traffic entities in the target area and the mutual influence relationships between the traffic entities.

[0118] The global traffic feature representation integrates the spatiotemporal dynamic information of the traffic entities in the target area (e.g., road traffic volume, vehicle speed, and event influence range) and the connection logic between the traffic entities (e.g., the conduction relationship of congestion from an upstream road segment to a downstream road segment). Based on the global traffic feature representation, the traffic congestion information of the target area in a target time period (e.g., 15 to 30 minutes in the future, from 5 p.m. to 6 p.m.) is obtained, which can include the location of the congested road segment, the congestion degree (light, moderate, or severe), and the congestion duration, etc.

[0119] In some embodiments, the step of weighted aggregation of the node feature representations of each traffic entity node and the node feature representations of neighboring nodes connected to the traffic-related edges, based on the topology of the temporal heterogeneous graph and the edge weights of the traffic-related edges, to generate a global traffic feature representation characterizing the state and mutual influence of all traffic entities in the target area, includes: calculating the attention weight of a traffic-related edge of a corresponding type based on the node feature representation of each traffic entity node, the edge weight of each type of traffic-related edge connected to the traffic entity node, and the node feature representations of neighboring nodes connected to the corresponding type of traffic-related edge; weighted aggregation of the node feature representations of neighboring nodes connected to all traffic-related edges according to the attention weights of all traffic-related edges of each type, to generate a single-relationship feature representation corresponding to the traffic-related edge of the corresponding type; feature fusion of the single-relationship feature representations corresponding to all types of traffic-related edges of each traffic entity node to generate a multi-relationship feature representation of the traffic entity node; and feature fusion of the multi-relationship feature representations of all traffic entity nodes to generate the global traffic feature representation.

[0120] Edge weights are adopted express, This refers to the weights of each edge. Node feature representation uses... express, Refers to various traffic entity nodes.

[0121] The edge weights are normalized using methods such as min-max normalization to obtain the normalized standard edge weights. The normalization formula is as follows: ,in, These are the normalized standard edge weights. Minimum edge weight The maximum edge weight is determined by normalization. This eliminates the dimensional differences between different edge weights, making them comparable in subsequent attention calculations and other operations, thereby improving the stability and accuracy of feature aggregation and inference prediction.

[0122] Node feature representation of each traffic entity node Perform linear projection to obtain the node projection feature representation. ,in, This represents the projection features of the nodes. This is a trainable matrix. Dimensional adaptation and semantic enhancement of node feature representations are achieved through projection operations.

[0123] The node attention score of each neighbor node connected by each type of traffic-related edge of each traffic entity node is calculated. For a traffic entity node and a traffic entity node , wherein, denotes the edge type of the traffic-related edge, is the node attention score, and are the node projection feature representations of the traffic entity node and the traffic entity node respectively, is a tunable parameter, is the semantic similarity between nodes, is the difference in the latest interaction time between the traffic entity node and the traffic entity node , is a time decay function, , is the normalized standard edge weight.

[0124] The node attention scores of each neighbor node connected by each type of traffic-related edge of each traffic entity node are normalized to obtain the attention weight of each traffic-related edge in the corresponding type of traffic-related edge.

[0125] In some embodiments, the node attention scores of each neighbor node connected by the same type of traffic-related edge are normalized using a softmax function. wherein, is the attention weight of the traffic-related edge.

[0126] The edge weight is used to represent the strength and reliability of the relationship between traffic entity nodes, which can be used as a direct scaling factor for attention weight calculation here. In addition, the edge weight is also used to control the sparsification and sampling strategy of message passing, which can be used as an explanatory index and basis for anomaly detection.

[0127] For each traffic entity node, the attention weights of all traffic-related edges in each type of traffic-related edge and the node feature representations of the neighbor nodes connected by all traffic-related edges are weighted and aggregated to generate a single relationship feature representation corresponding to the neighbor nodes connected by the corresponding type of traffic-related edge, realizing the aggregation of the same type of edge. Each type of traffic-related edge represents a type of traffic-related relationship, is the single relationship feature representation.

[0128] In some embodiments, the edge type weight of each type of traffic-related edge of each traffic entity node is calculated. wherein, To learn the edge type weight, through the parameter, the importance of the corresponding type edge can be learned, The normalized edge type weight. The edge type weight is a weight index representing the importance of different types of traffic-related edges, and is used to distinguish the contribution difference of various traffic-related edges in traffic scene analysis.

[0129] According to the edge type weight of all types of interactive associated edges of each traffic entity node and the single relationship feature representation corresponding to all types of interactive associated edges, the multi-relationship feature representation of the traffic entity node is calculated. wherein, The optional residual term, The activation function is, for example, a rectified linear unit (ReLU), The multi-relationship feature representation of the current time step.

[0130] In some embodiments, the above fusion process can be implemented by a time-aware heterogeneous attention network (T-HAN) to obtain the multi-relationship feature representation of each traffic entity node.

[0131] In the above process, the information of different types of neighbor nodes in the time sequence heterogeneous graph is weighted aggregated according to semantic similarity, edge weight, etc., and then merged according to edge type importance to obtain the fusion representation of each traffic entity node, which realizes the effective fusion of multi-dimensional and differentiated associated information of traffic entity nodes, and significantly improves the accuracy of node feature representation in describing traffic situation.

[0132] In addition to semantic similarity and edge weight, time freshness can also be combined to realize weighted aggregation.

[0133] In some embodiments, the multi-relationship feature representations of all traffic entity nodes in the time sequence heterogeneous graph are fused to generate a global traffic feature representation representing the traffic situation of the time sequence heterogeneous graph. Then, traffic congestion prediction can be performed based on the global traffic feature representation.

[0134] In some embodiments, the predicting the traffic congestion information of the target region in the target time period based on the global traffic feature representation comprises: predicting a target traffic congestion index of the target region in the target time period, a first traffic congestion index of the target region in a time period before the target time period, and a second traffic congestion index of the target region in a time period after the target time period based on the global traffic feature representation; determining a target traffic congestion evolution stage of the target region in the target time period from a set of traffic congestion evolution stages according to a time sequence variation trend among the target traffic congestion index, the first traffic congestion index, and the second traffic congestion index; and the set of traffic congestion evolution stages comprises a forming stage, a sustaining stage, and a dissipating stage.

[0135] The global traffic feature representation has integrated state information and mutual influence relationships of all traffic entities in the target region.

[0136] The target traffic congestion index is a congestion quantification value of the target time period (e.g., 17:20-17:40), which is used to directly represent the congestion degree of the time period, and the higher the value, the more serious the congestion.

[0137] The first traffic congestion index is a congestion quantification value of a time period before the target time period (e.g., 17:00-17:20), which reflects a historical basic state of the congestion.

[0138] The second traffic congestion index is a congestion quantification value of a time period after the target time period (e.g., 17:40-18:00), which reflects a subsequent development trend of the congestion.

[0139] Here, a single statistical value (i.e., one traffic congestion index) is used to represent the traffic state of a time period, instead of outputting multiple values per second / minute. In this way, the phase judgment can be completed by comparing the statistical values of the three time periods (i.e., the front, middle, and rear time periods), and data redundancy caused by a time period corresponding to hundreds of scattered indexes is avoided.

[0140] In some embodiments, the traffic congestion index is a summary data of the traffic state of a time period, for example, an average congestion index of the entire time period, a peak congestion index in the time period, or a weighted congestion index corresponding to the time period.

[0141] In some embodiments, the global traffic feature representation of the target time period, the previous time period of the target time period, and the next time period of the target time period are respectively input into a time series prediction model, such as a long short-term memory (LSTM) network with attention mechanism fusion, a time series Transformer model, etc., for prediction to obtain the target traffic congestion index, the first traffic congestion index, and the second traffic congestion index output by the time series prediction model. The global traffic feature representation of the target time period, the previous time period of the target time period, and the next time period of the target time period can be obtained by adjusting the time decay scale to optimize the feature fusion process, and the accuracy of the feature representation is high.

[0142] By analyzing the time series trend (increasing, stable, decreasing) of the target traffic congestion index, the first traffic congestion index, and the second traffic congestion index, the stage attribute of the target time period is matched from the set traffic congestion evolution stage, i.e., the target traffic congestion evolution stage of the target time period is obtained.

[0143] In some embodiments, in the formation stage, if the first traffic congestion index is low (slight congestion or no congestion), the target traffic congestion index increases significantly (congestion intensifies), and the second traffic congestion index continues to increase (congestion will further develop), it is determined that the target time period is in the congestion formation stage (e.g., a road section gradually changes from smooth to congested during the morning peak); in the sustained stage, if the first traffic congestion index, the target traffic congestion index, and the second traffic congestion index are all at a high level and have a small change range (fluctuation within a preset threshold), it is determined that the target time period is in the congestion sustained stage (e.g., a congested road section maintains a serious congestion state during the peak period); in the dissipation stage, if the first traffic congestion index is high (serious congestion), the target traffic congestion index decreases significantly (congestion is alleviated), and the second traffic congestion index continues to decrease (congestion will further dissipate), it is determined that the target time period is in the congestion dissipation stage (e.g., a road section gradually recovers from congestion to smooth after the evening peak).

[0144] In some embodiments, the target traffic congestion evolution stage of the target time period can be a composite stage, i.e., the congestion state in the same target time period successively experiences a continuous combination of two or more basic evolution stages (formation stage, sustained stage, dissipation stage).

[0145] Through the above process, not only the congestion degree of the target time period in the target area can be predicted, but also the evolution degree in the entire congestion evolution cycle can be determined, which provides more targeted decision basis for traffic control (e.g., early diversion in the formation stage, intensified shunting in the sustained stage, and adjustment of signal timing in the dissipation stage).

[0146] In some embodiments, the traffic congestion of a road section in the target area can be predicted.

[0147] For each road segment node , take the multi-relation feature representation of the last time steps as the input parameters of the time series prediction model, The multi-relation feature representation of the last time steps can be represented as Perform time step recursion on the input parameters , take the hidden layer state of the last time as the time series aggregation representation, and use the time series aggregation representation to predict the traffic congestion index of the road segment

[0148] In some embodiments, a prediction target function of the traffic congestion index is constructed based on the time series aggregation representation, , is the traffic congestion index, is the weight, is the bias.

[0149] In some embodiments, to quantify the deviation of the predicted traffic congestion index from the actual traffic congestion index, the mean square error is used as the loss function, . Wherein, is the real traffic congestion index in the time step, is the predicted traffic congestion index in the time step, is the prediction time span. By minimizing the loss function and combining the back propagation algorithm, the joint learning of the LSTM network parameters, the output layer weight and the bias can ensure continuous optimization of prediction accuracy.

[0150] In some embodiments, the latest multi-modal traffic data can be used to construct a time series heterogeneous graph to generate the latest global traffic feature representation, ensuring that the time series heterogeneous graph can capture the dynamic changes of the traffic scene in real time and provide accurate dynamic graph data support for subsequent continuous traffic congestion prediction.

[0151] In this application, multi-modal heterogeneous data is fused to construct a time series heterogeneous graph, which accurately models the dynamic changes of the traffic network and the complex relationships of heterogeneous entities. Not only can it comprehensively capture congestion influencing factors from multiple dimensions and discover potential associations (such as the synergistic effect of business districts and bus stops on congestion), but it can also greatly improve prediction accuracy (such as accurately determining waterlogged road congestion by combining real-time traffic and weather warnings), and dynamically update the prediction results in real time to adapt to sudden traffic events (such as accidents and temporary restrictions).

[0152] At the application level, it can provide accurate prediction for traffic management department to advance dredge, optimize resource allocation (such as signal timing, police deployment), assist travelers in planning routes and saving time cost, provide data support for intelligent transportation system modules such as autonomous driving and intelligent bus scheduling to promote system collaborative development, reduce idling emission by guiding vehicles to avoid congestion, help energy saving and emission reduction, and ultimately realize multi-dimensional value from traffic management efficiency improvement, travel experience optimization to urban intelligent transportation ecological perfection and environmental protection benefit.

[0153] In the embodiments of the present application, specific prior knowledge is introduced, and the traffic entities of multiple types in the target area and the traffic correlation relationship between the traffic entities are accurately defined. In combination with the multi-modal traffic data, the feature representation of each traffic entity and the traffic correlation strength between each pair of traffic entities are determined, so as to construct a time-sequential heterogeneous graph containing traffic entity nodes of multiple types and traffic correlation edges of multiple types. The time-sequential heterogeneous graph realizes fine quantification of the heterogeneous traffic entities and the dynamic correlation relationship between the heterogeneous traffic entities in combination with the real-time characteristics of the multi-modal traffic data. On this basis, the traffic congestion information of the target area is predicted based on the constructed time-sequential heterogeneous graph, that is, the time-sequential heterogeneous graph with rich information is used for prediction, which effectively improves the accuracy and reliability of traffic congestion prediction.

[0154] Referring to Figure 3 , Figure 3 is a structure diagram of a traffic congestion prediction system provided by an embodiment of the present application. Only parts related to the embodiments of the present application are shown for ease of description.

[0155] The traffic congestion prediction system 300 comprises an acquisition module 301, a first determination module 302, a second determination module 303, a construction module 304, and a prediction module 305.

[0156] The acquisition module 301 is configured to acquire specific prior knowledge and multi-modal traffic data corresponding to a target area.

[0157] The first determination module 302 is configured to determine traffic entities of multiple types in the target area and specific types of traffic correlation relationships between each pair of the traffic entities based on the specific prior knowledge.

[0158] The second determination module 303 is configured to determine feature representations of each of the traffic entities and traffic correlation strengths between each pair of the traffic entities based on the multi-modal traffic data and the traffic correlation relationships.

[0159] The constructing module 304 is configured to construct a time-sequential heterogeneous graph including traffic entity nodes corresponding to the traffic entities of the multiple types and traffic association edges corresponding to the traffic association relationships between the traffic entities.

[0160] The predicting module 305 is configured to predict traffic congestion information of the target area based on the time-sequential heterogeneous graph.

[0161] In some embodiments, the first determining module is specifically configured to:

[0162] According to entity definition information in the specific prior knowledge, identify the traffic entities of the multiple types in the target area;

[0163] According to association relationship definition information in the specific prior knowledge, determine the traffic association relationship between each pair of the traffic entities; the association relationship definition information includes car-following relationship definition information, containment relationship definition information, cause-effect relationship definition information, subsidiary relationship definition information, and control relationship definition information.

[0164] In some embodiments, the second determining module is specifically configured to:

[0165] Perform feature extraction on the traffic data of the multiple modalities to obtain initial features of the traffic data of each modality;

[0166] Map the initial features of each modality to a feature space of a set dimension to obtain set-dimension features of the multiple modalities;

[0167] Select, from the set-dimension features of the multiple modalities, target features of each traffic entity that match a feature representation requirement of the entity type of the traffic entity, and perform feature fusion based on the target features to obtain the feature representation of the traffic entity; and

[0168] Select, from the traffic data of the multiple modalities, target data of each pair of the traffic entities that match a quantization requirement of the traffic association relationship of the traffic entities, and calculate the traffic association strength between the traffic entities based on the target data.

[0169] In some embodiments, the second determining module is further configured to:

[0170] Select, from the set-dimension features of the multiple modalities, the target features of each traffic entity that match the feature representation requirement of the entity type of the traffic entity;

[0171] According to a target modality corresponding to the target features, determine a non-target modality that does not match the feature representation requirement;

[0172] obtain mask features corresponding to the non-target modality;

[0173] perform feature fusion on the target features and the mask features to obtain the feature representation corresponding to the traffic entity.

[0174] In some embodiments, the prediction module is specifically configured to:

[0175] based on the topology of the time-series heterogeneous graph and the edge weights of the traffic-related edges, perform weighted aggregation on the node feature representation of each traffic entity node and the node feature representation of the neighbor nodes connected by the traffic-related edges of the traffic entity node, to generate a global traffic feature representation representing the states and mutual influences of all traffic entities in the target region;

[0176] based on the global traffic feature representation, predict the traffic congestion information of the target region in a target time period.

[0177] In some embodiments, the prediction module is further configured to:

[0178] based on the node feature representation of each traffic entity node, the edge weight of each traffic-related edge of each type connected by the traffic entity node, and the node feature representation of the neighbor nodes connected by the traffic-related edge of the corresponding type, calculate the attention weight of the traffic-related edge of the corresponding type;

[0179] perform weighted aggregation on the node feature representation of the neighbor nodes connected by all traffic-related edges of each type based on the attention weights of all traffic-related edges of each type, to generate a single relationship feature representation corresponding to the traffic-related edge of the corresponding type;

[0180] perform feature fusion on the single relationship feature representations corresponding to all types of traffic-related edges of each traffic entity node, to generate a multi-relationship feature representation of the traffic entity node;

[0181] perform feature fusion on the multi-relationship feature representations of all traffic entity nodes, to generate the global traffic feature representation.

[0182] In some embodiments, the prediction module is further configured to:

[0183] based on the global traffic feature representation, predict a target traffic congestion index of the target region in the target time period, a first traffic congestion index in a previous time period of the target time period, and a second traffic congestion index in a subsequent time period of the target time period;

[0184] According to a time sequence change trend among the target traffic congestion index, the first traffic congestion index and the second traffic congestion index, a target traffic congestion evolution stage of the target area in the target time period is determined from set traffic congestion evolution stages; the set traffic congestion evolution stages include a forming stage, a sustaining stage and a dissipating stage.

[0185] The traffic congestion prediction system provided by the embodiments of the present application can implement each process of the embodiments of the traffic congestion prediction method described above and achieve the same technical effects. To avoid repetition, the traffic congestion prediction system will not be described here.

[0186] Figure 4 is a structural diagram of an electronic device provided by an embodiment of the present application. As shown in the diagram, the electronic device 4 of the embodiment includes at least one processor 40 (only one is shown in the diagram), a memory 41, and a computer program 42 stored in the memory 41 and executable on the at least one processor 40, wherein the processor 40 implements the steps in any of the method embodiments described above when executing the computer program 42. Figure 4

[0187] The electronic device 4 can be a desktop computer, a notebook computer, a palm computer, a cloud server and the like. The electronic device 4 can include, but is not limited to, the processor 40 and the memory 41. Those skilled in the art can understand that the electronic device 4 is only an example of the electronic device 4 and does not constitute a limitation on the electronic device 4, and can include more or fewer components than those shown in the diagram, or combine certain components, or different components, for example, the electronic device can also include an input / output device, a network access device, a bus and the like. Figure 4

[0188] The processor 40 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0189] ​​The memory 41 can be an internal storage unit of the electronic device 4, such as a hard disk or a memory of the electronic device 4. The memory 41 can also be an external storage device of the electronic device 4, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 4. Further, the memory 41 can also include both the internal storage unit and the external storage device of the electronic device 4. The memory 41 is used to store the computer program and other programs and data required by the electronic device. The memory 41 can also be used to temporarily store data that has been output or will be output.

[0190] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or software. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0191] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0192] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the application.

[0193] In the embodiments of the present application, it should be understood that the disclosed system / electronic device and method can be implemented in other manners. For example, the embodiments of the system / electronic device described above are merely schematic. For example, the division of the modules or units is only a logical function division. There can be another division manner for the actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, can be indirect couplings or communication connections through some interfaces, and electrical, mechanical or other forms.

[0194] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0195] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0196] The integrated module / unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the computer readable medium can include appropriate contents according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0197] The application can realize all or part of the processes in the above-mentioned embodiment methods, and can also be realized by a computer program product. When the computer program product runs on an electronic device, the electronic device is caused to perform the steps in the above-mentioned various method embodiments.

[0198] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, rather than limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents. The modifications or replacements do not change the essence of the corresponding technical solutions, and should be included in the protection scope of the present application.

Claims

1. A traffic congestion prediction method, characterized in that, include: Acquire multimodal traffic data corresponding to specific prior knowledge and target areas; Based on the specific prior knowledge, determine multiple types of traffic entities in the target area and specific types of traffic associations between each pair of traffic entities; Based on the multimodal traffic data and the traffic association relationships, the feature representation of each traffic entity and the traffic association strength between each pair of traffic entities are determined. The feature representation is used as the node feature representation of traffic entity nodes corresponding to multiple types of traffic entities, and the traffic association strength is used as the edge weight of the traffic association edge between the traffic entity nodes that refers to the traffic association relationship, to obtain a temporal heterogeneous graph containing multiple types of traffic entity nodes and multiple types of traffic association edges. Based on the temporal heterogeneous graph, predicting traffic congestion information in the target area specifically includes: based on the topology of the temporal heterogeneous graph and the edge weights of the traffic-related edges, weighted aggregation of the node feature representations of each traffic entity node and the node feature representations of the neighboring nodes connected by the traffic-related edges of the traffic entity nodes, generating a global traffic feature representation characterizing the state and mutual influence of all traffic entities in the target area; and based on the global traffic feature representation, predicting the traffic congestion information in the target area during a target time period. The step of weighted aggregation of the node feature representations of each traffic entity node and the node feature representations of the neighboring nodes connected to the traffic association edges, based on the topology of the temporal heterogeneous graph and the edge weights of the traffic association edges, to generate a global traffic feature representation characterizing the state and mutual influence of all traffic entities in the target area, includes: calculating the attention weight of the traffic association edges of the corresponding type based on the node feature representation of each traffic entity node, the edge weight of each type of traffic association edge connected to the traffic entity node, and the node feature representations of the neighboring nodes connected to the traffic association edges of the corresponding type; weighted aggregation of the node feature representations of the neighboring nodes connected to all traffic association edges of each type according to the attention weights of all traffic association edges of each type, to generate a single-relationship feature representation corresponding to the traffic association edges of the corresponding type; feature fusion of the single-relationship feature representations corresponding to all types of traffic association edges of each traffic entity node to generate a multi-relationship feature representation of the traffic entity node; and feature fusion of the multi-relationship feature representations of all traffic entity nodes to generate the global traffic feature representation.

2. The method according to claim 1, characterized in that, The process of determining multiple types of traffic entities in the target area and specific types of traffic associations between each pair of traffic entities based on the specific prior knowledge includes: Based on the entity definition information in the specific prior knowledge, identify multiple types of traffic entities in the target area; Based on the association relationship definition information in the specific prior knowledge, the traffic association relationship between each pair of traffic entities is determined; the association relationship definition information includes car-following relationship definition information, inclusion relationship definition information, causal relationship definition information, subordinate relationship definition information, and control relationship definition information.

3. The method according to claim 1, characterized in that, The traffic data and traffic associations based on multimodal data are used to determine the feature representation of each traffic entity and the traffic association strength between each pair of traffic entities, including: Feature extraction is performed on the multimodal traffic data to obtain the initial features of the traffic data for each modality; The initial features of each modality are mapped to a feature space of a set dimension to obtain the set dimension features of the multimodality; From the defined dimensional features of the multimodal traffic entity, target features matching the feature representation requirements of its entity type are selected, and feature fusion is performed based on the target features to obtain the feature representation of the traffic entity; and, From the multimodal traffic data, target data matching the quantitative requirements of the traffic association relationship between each pair of traffic entities is selected, and the traffic association strength between the traffic entities is calculated based on the target data.

4. The method according to claim 3, characterized in that, The step of selecting target features matching the feature representation requirements of the traffic entity type from the defined dimensional features of the multimodal traffic entity, and performing feature fusion based on the target features to obtain the feature representation of the traffic entity, includes: From the defined dimensional features of the multimodal traffic entity, select the target feature that matches the feature representation requirements of its entity type for each traffic entity; Based on the target modality corresponding to the target feature, determine the non-target modality that does not match the feature representation requirements; Obtain the mask features corresponding to the non-target mode; The target feature and the mask feature are fused to obtain the feature representation corresponding to the traffic entity.

5. The method according to claim 1, characterized in that, The prediction of traffic congestion information in the target area during the target time period based on the global traffic feature representation includes: Based on the global traffic feature representation, predict the target traffic congestion index of the target area in the target time period, the first traffic congestion index in the time period before the target time period, and the second traffic congestion index in the time period after the target time period. Based on the temporal variation trend among the target traffic congestion index, the first traffic congestion index, and the second traffic congestion index, the target traffic congestion evolution stage of the target area in the target time period is determined from the set traffic congestion evolution stages; the set traffic congestion evolution stages include the formation stage, the duration stage, and the dissipation stage.

6. A traffic congestion prediction system, said system being used to perform the method as described in any one of claims 1-5, characterized in that, include: The acquisition module is used to acquire multimodal traffic data corresponding to specific prior knowledge and target areas; The first determining module is used to determine, based on the specific prior knowledge, multiple types of traffic entities in the target area and specific types of traffic associations between each pair of traffic entities; The second determining module is used to determine the feature representation of each traffic entity and the traffic association strength between each pair of traffic entities based on the multimodal traffic data and the traffic association relationship. The construction module is used to use the feature representation as the node feature representation of traffic entity nodes corresponding to multiple types of traffic entities, and to use the traffic association strength as the edge weight of the traffic association edge between the traffic entity nodes that refers to the traffic association relationship, so as to obtain a temporal heterogeneous graph containing multiple types of traffic entity nodes and multiple types of traffic association edges. The prediction module is used to predict traffic congestion information in the target area based on the time-series heterogeneous graph.

7. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device performs the method as described in any one of claims 1 to 5.

8. A computer program product, characterized in that, Includes a computer program, which, when run, causes the method as described in any one of claims 1 to 5 to be performed.

Citation Information

Patent Citations

  • Multi-intersection traffic signal control method fusing traffic flow prediction

    CN120071647A

  • Urban traffic flow pre-judgment system

    CN120636157A