Highway area risk intelligent early warning method and system based on space-time coupling model
Through the space-time coupling model, the road area monitoring data characteristics and risk warning rule characteristics are obtained and integrated, and the problem of insufficient adaptability caused by single-dimensional analysis in the existing technology is solved, and intelligent and accurate warning of highway road area risks is achieved.
Patent Information
- Application Number
- CN202510978643.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-16
AI Technical Summary
In the prior art, highway road risk warning methods are mostly based on single-dimensional analysis, and it is difficult to capture the spatial and temporal correlation of multi-source data, resulting in insufficient feature utilization and weak adaptability of early warning models to dynamic road domain scenarios.
The method based on the space-time coupling model is adopted to obtain the time dimension and spatial dimension characteristics of the road domain monitoring data, fuse it into the space-time coupling characteristics of the first road domain, and obtain the first rule embedded characteristics of the risk warning rule. Through heterogeneous feature coupling processing, combined with the learning parameter optimization, the risk warning model is called for risk prediction.
It improves the multi-dimensional feature expression ability of road domain risks, enhances the model's adaptability to dynamic scenarios, improves the accuracy of risk warnings, and realizes the efficient integration of intelligent warnings.
Smart Images

Figure CN120472694A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of traffic management, and in particular to a method and system for intelligent early warning of highway road area risks based on a spatiotemporal coupling model. Background Art
[0002] As critical transportation infrastructure, accurate early warning of highway risks is crucial for ensuring safe and efficient travel. Existing approaches to early warning of highway risks often suffer from the following shortcomings: Traditional methods often analyze monitoring data from a single dimension, failing to capture the spatiotemporal correlations among multi-source data and resulting in inadequate feature utilization; risk warning rules are insufficiently integrated with the heterogeneous features of monitoring data, resulting in limited adaptability of warning models to dynamic road scenarios. Therefore, an intelligent early warning method is urgently needed that can efficiently integrate spatiotemporal and rule-based features to enhance model adaptability. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for intelligent early warning of highway road area risks based on a spatiotemporal coupling model.
[0004] In a first aspect, an embodiment of the present invention provides an intelligent early warning method for highway road area risks based on a spatiotemporal coupling model, comprising: Acquiring road domain monitoring data, extracting time dimension features and space dimension features of the road domain monitoring data, and fusing the time dimension features and the space dimension features to obtain a first road domain spatiotemporal coupling feature of the road domain monitoring data;
[0005] Acquire a risk warning rule, extract a first rule embedding feature of the risk warning rule, and perform heterogeneous feature coupling processing on the first rule embedding feature and the first road domain spatiotemporal coupling feature, wherein the risk warning rule is used in a risk warning model to perform intelligent risk warning on the road domain monitoring data;
[0006] Optimizing the heterogeneous feature coupling processing result based on a first learnable parameter to obtain a target rule embedding feature, wherein the first learnable parameter is determined by training after fixing the parameters of the risk warning model;
[0007] The risk warning model is called to perform risk prediction based on the target rule embedding feature to obtain a risk warning result of the road domain monitoring data.
[0008] In a second aspect, an embodiment of the present invention provides a server system, including a server, wherein the server is configured to execute the method described in the first aspect.
[0009] Compared to existing technologies, the present invention provides the following beneficial effects: Using the disclosed method and system for intelligent early warning of highway road-domain risks based on a spatiotemporal coupling model, the present invention acquires road-domain monitoring data, extracts and fuses its temporal and spatial dimension features to obtain a first road-domain spatiotemporal coupling feature; secondly, acquires risk warning rules and extracts their first rule embedding features, performing heterogeneous feature coupling on the two; then, based on the first learnable parameters determined by training after fixing the risk warning model parameters, optimizes the coupling result to obtain the target rule embedding feature; finally, calls the risk warning model to predict risks based on the target rule embedding feature and outputs a warning result. Through spatiotemporal feature fusion and heterogeneous rule feature coupling, this method enhances the multi-dimensional feature expression capability of road-domain risks. Combined with learnable parameter optimization, it enhances the model's adaptability to dynamic scenarios and effectively improves the accuracy of risk warnings. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly describes the drawings required for use in the embodiments. It should be understood that the following drawings illustrate only certain embodiments of the present invention and should not be construed as limiting the scope of the present invention. Those skilled in the art can, without inventive effort, derive other relevant drawings from these drawings.
[0011] Figure 1 A schematic block diagram of the steps of an intelligent early warning method for highway road area risks based on a spatiotemporal coupling model provided by an embodiment of the present invention;
[0012] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more apparent, the technical solutions of the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the embodiments of the present invention. It should be understood that the described embodiments are only a portion of the embodiments of the present invention, not all of them. Generally, the components of the embodiments of the present invention described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.
[0014] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0015] In order to solve the technical problems in the above background technology, Figure 1 This is a flow chart of the intelligent early warning method for highway road domain risks based on the time-space coupling model provided in an embodiment of the present disclosure. The intelligent early warning method for highway road domain risks based on the time-space coupling model is introduced in detail below.
[0016] Step S201: Acquire road monitoring data, extract time dimension features and space dimension features of the road monitoring data, and fuse the time dimension features and the space dimension features to obtain a first road spatiotemporal coupling feature of the road monitoring data;
[0017] Step S202: Acquire a risk warning rule, extract a first rule embedding feature of the risk warning rule, and perform heterogeneous feature coupling processing on the first rule embedding feature and the first road domain spatiotemporal coupling feature, wherein the risk warning rule is used in a risk warning model to perform intelligent risk warning on the road domain monitoring data;
[0018] Step S203, optimizing the heterogeneous feature coupling processing result based on a first learnable parameter to obtain a target rule embedding feature, wherein the first learnable parameter is determined by training after fixing the parameters of the risk warning model;
[0019] Step S204 : calling the risk warning model to perform risk prediction based on the target rule embedding feature to obtain a risk warning result of the road area monitoring data.
[0020] In an embodiment of the present invention, for example, the server first receives a continuous road monitoring image sequence (7,500 frames, 5 minutes) of the road section from 2:00 PM to 2:05 PM on a weekday. The images contain clear spatial information: lane markings, bounding boxes of approximately 30 vehicles per frame, shallow water accumulation in some lanes after a light rain (approximately 1 / 5 of the area of a single lane), and the status of road guardrails. They also carry dynamic temporal information: the real-time speed of vehicles calculated from the displacement of adjacent frames (most vehicles maintain a speed of 80-100 km / h, but vehicles on flooded sections slow to below 60 km / h), the area change of the flooded area extracted by the semantic segmentation model (increases by approximately 0.5 square meters every 10 seconds), and the slow decrease in light intensity caused by cloud cover (brightness decreases by 10% over 15 minutes).
[0021] To extract spatiotemporal features, the server first uses a pre-trained ResNet-50 model (without the fully connected layer) to extract a 2048-dimensional spatial feature vector for each image frame. Then, sliding window mean pooling (50-frame window, 25-frame step) is performed to preserve the temporal continuity of spatial information, resulting in a 2048×300 spatial monitoring image feature sequence (corresponding to 300 time windows). Next, temporal dynamic data (vehicle speed, rate of change of flooded area, and changes in light intensity) is extracted from the image and fed into a bidirectional LSTM model (with a 1024-dimensional hidden layer) for time series modeling, resulting in 1024×300 temporal features aligned with the spatial feature sequence time window. To capture the spatiotemporal synergy between water expansion and vehicle deceleration, the server uses a spatiotemporal attention mechanism to fuse the two: it calculates the dot product association matrix (300×300) of the spatial and temporal feature vectors, obtains the spatiotemporal attention weights through Softmax normalization, and then sums the spatial feature sequence with the weights to obtain the first road-domain spatiotemporal coupling feature (2048×300), which fully preserves the association information in the spatiotemporal dimension.
[0022] Subsequently, the server needs to associate historical scenes to generate risk warning rules. It accesses the road risk feature library (which stores 100,000 benchmark data of the road section in the past three years, each containing a reference image sequence, a second road domain spatiotemporal coupling feature that integrates spatiotemporal features, and manually annotated benchmark risk status), and uses the KD tree algorithm to quickly retrieve candidate data similar to the current first road domain spatiotemporal coupling feature index (spatial mean + time mean, 3072 dimensions), and then calculates the cosine similarity (threshold 0.7) to filter out the most similar historical scene (a waterlogging scene in a summer afternoon, with a similarity of 0.85), whose benchmark risk status is "moderate road waterlogging risk". The server generates according to the rule format: add a start symbol ( <start>) and the terminator ( <end>), get the rule element text" <start>Moderate road flooding risk <end>", and then add the start symbol at the end to form a complete risk warning rule" <start>Moderate road flooding risk <end> <start>” (used to associate the current scenario with historical rules). Then, the pre-trained BERT-base model is used to embed the rules: the rules are split into " <start>Moderate road surface flooding and wind risk <end> ”" <start>There are 11 tokens in total. The first rule embedding feature (768 dimensions) of each token is obtained through word embedding and position embedding, forming an 11×768 rule embedding matrix, which retains the semantic and sequential information of the rule.
[0023] To fuse rule embedding (semantic) with spatiotemporal coupling features (data), the server performs heterogeneous feature coupling. First, the current first-domain spatiotemporal coupling features are superimposed with the retrieved historical second-domain spatiotemporal coupling features (2048×300) to generate the integrated spatiotemporal coupling features (which associate the spatiotemporal information of the current and historical scenarios). Then, the text part of the rule elements (the first 10 tokens) is processed token by token: taking the "product" token as an example, the learnable linear layer (W_r∈R^768×512, W_s∈R^2048×512) is used to map its rule embedding features (768 dimensions) and the historical spatiotemporal features (2048 dimensions) in the integrated spatiotemporal features into low-dimensional vectors (512 dimensions), and the dot product is calculated to obtain the association strength vector (indicating the degree of association between the rule semantics and the historical spatiotemporal features); irrelevant features are suppressed through the feature masking vector (the element is 1 when the association strength exceeds 0.5, otherwise it is 0), and then the feature coupling weight vector (512 dimensions) is obtained by Softmax normalization; finally, the historical spatiotemporal features are mapped to a representation vector (512 dimensions), which is weighted and fused with the weight vector to obtain the heterogeneous coupling result (512 dimensions) of the token. For the start symbol at the end of the rule (the 11th token, corresponding to the current scene), its rule embedding features and the current scene spatiotemporal features in the integrated spatiotemporal features are processed in the same way, and finally an 11×512 heterogeneous coupling result matrix is obtained.
[0024] To adapt the coupling results to the risk warning model, the server uses pre-trained learnable parameter optimization: First, the first learnable parameter (W_1∈R^512×768) is used to transform the coupling result dimension to match the rule embedding (768 dimensions), resulting in an optimized coupling result (11×768). This is then superimposed with the first rule embedding feature (preserving the original semantics of the rule while integrating spatiotemporal information), resulting in a second rule embedding feature (11×768). Finally, a second learnable parameter (W_2∈R^768×768, with ReLU activation) is used to perform a nonlinear transformation to enhance the feature representation, resulting in the target rule embedding feature (11×768). These parameters were obtained by training on 10,000 annotated historical data points for this road section while keeping the parameters of the risk warning model (Transformer encoder) fixed, ensuring that the features can be effectively adjusted to adapt to the model input.
[0025] The server invokes a risk warning model (3 cascaded Transformer encoder layers) for prediction. The target rule embedding features are used as the input to the first encoder layer. Through the self-attention mechanism, long-range dependencies between tokens are captured (such as the semantic association between "accumulation" and "water"), and the first-layer feature extraction result (11×768) is obtained. For the second encoder (not the last one), the server performs heterogeneous coupling again between the first-layer result and the first road domain spatio-temporal coupling feature (the spatio-temporal feature is the current scenario), optimizes the coupling result with the learnable parameters of the corresponding layer, superimposes and transforms it with the first-layer result, and then inputs it into the second encoder to obtain the second-layer feature extraction result (11×768). The third encoder (the last one) directly processes the second-layer result and outputs the final features (11×768). Through a linear classification layer (R^768×4, where 4 is the risk level: none, mild, moderate, severe), the risk probability distribution of each token is obtained. The server takes the probability distribution of the starting symbol at the end of the rule (corresponding to the current scenario) (risk-free 0.12, mild 0.21, moderate 0.63, severe 0.04), and selects the category with the highest probability as the final risk warning result - moderate road surface water accumulation risk.
[0026] Finally, the server stores the results in a structured manner (road section identification, time range 14:00 - 14:05, risk type "road surface water accumulation", risk level "moderate", warning basis "historical similar water accumulation scenarios"), and pushes them to the highway management monitoring center: on the large screen, it shows "There is a moderate road surface water accumulation risk on a certain six-lane plain road section from 14:00 to 14:05", marks the spatial location of the water accumulation area (through semantic segmentation positioning) and the time change trend (the water accumulation area increases from 12㎡ to 28㎡); at the same time, it sends instructions to the variable message sign 2 kilometers ahead of the road section, showing "There is water accumulation on the road ahead, slow down (recommended ≤60km / h)"; it also notifies the maintenance team through the SMS platform, providing the water accumulation location, risk level and treatment suggestions ("Immediately go to the site to lay anti-slip mats and turn on warning lights"). The maintenance personnel arrived at the site at 14:10, confirmed that the water depth was about 3cm, which was consistent with the warning result, and took timely measures to avoid potential risks.
[0027] Through spatio-temporal coupling feature extraction, historical rule association, heterogeneous feature fusion and learnable parameter optimization, the entire process realizes the intelligent and accurate warning of road domain risks, effectively improving the operation safety and management efficiency of expressways.
[0028] Furthermore, the embodiments of the present invention also relate to a dynamic collaborative warning mechanism and hierarchical control measures to achieve efficient linkage disposal of risks.
[0029] In terms of dynamic collaborative warning, after generating the risk warning result, the server does not only push through a single channel, but starts a three-level collaborative response process of "perception-edge-cloud":
[0030] Perception-layer collaboration: The server sends warning information (risk type, level, location, and trend) in real time to all intelligent sensing devices (such as millimeter-wave radar, lidar, weather stations, and road condition sensors) within a 5-kilometer radius upstream and downstream of the warning section. Upon receiving these instructions, these devices automatically adjust their sampling frequency (for example, from the standard 1Hz to 10Hz) and monitoring range, focusing on tracking changes in the risk area and associated road sections, forming a closed-loop monitoring loop for risk development. For example, for a "moderate flooding risk," road moisture sensors upstream and downstream of the flooded area will encrypt data collection and provide real-time feedback on the spread of water.
[0031] Edge-layer collaboration: Edge computing nodes deployed in road monitoring subcenters receive warning results pushed by the server and real-time data streams uploaded by the perception layer, initiating local rapid response plans. For example, in addition to controlling local variable information boards to issue warning information, they can also directly interact with traffic signal control systems within the road section (if any) to implement appropriate traffic flow control upstream of the risky road section, such as temporarily restricting the entry of large vehicles or reducing lane speed limits, while ensuring safety.
[0032] Cloud-based Collaboration: The cloud-based management platform integrates risk warning information from the entire road network to conduct a global situation assessment. When the risk level of a road section reaches a preset threshold (such as "severe") or when multiple adjacent sections experience correlated risks (such as flooding on multiple roads due to continuous rainfall), the cloud-based platform coordinates cross-regional and cross-departmental responses across multiple departments (such as traffic police, maintenance, and emergency rescue), centrally deploying resources and developing an optimal emergency plan.
[0033] In terms of hierarchical control measures, the embodiment of the present invention triggers different levels of control measures according to the risk level of the risk warning result (such as no risk, mild, moderate, severe):
[0034] No risk (probability ≤ 0.2): The system only performs routine monitoring data recording and trend analysis and does not trigger active warnings.
[0035] Mild risk (0.2 < probability ≤ 0.4): Information prompt: Send risk alert information to the highway management and monitoring center and mark it in the internal management system. Strengthen monitoring: Instruct relevant sensing equipment to increase monitoring frequency and closely monitor changes in risk factors. Prepare emergency plans: Notice relevant maintenance teams to monitor the situation on this section of road and prepare emergency supplies.
[0036] Moderate risk (0.4 < probability ≤ 0.7): Multi-level warning push: The warning information is highlighted on the large screen of the monitoring center, and specific risk warnings and recommended speeds are issued to the variable information board 2-5 kilometers ahead of the road section (such as "There is water accumulation on the road ahead, slow down (recommended ≤ 60km / h)"), and a text message / APP push containing the precise location, risk level and preliminary treatment suggestions is sent to the maintenance team leader (such as "Go to the scene immediately to lay anti-skid mats and turn on the warning lights"). Traffic guidance: Edge computing nodes can be linked to LED induction screens along the route to guide vehicles to drive cautiously or choose alternative lanes in advance. Personnel dispatch: Based on the warning information, the on-duty personnel of the monitoring center can dispatch maintenance personnel and equipment to the scene in advance for standby or preliminary treatment.
[0037] Severe risk (probability > 0.7): Emergency warning response: In addition to all measures for moderate risk, variable information boards will display more prominent warning messages (such as "Severe flooding ahead, no passage"), and detour instructions will be issued to vehicles through traffic broadcasts, navigation apps, and other channels. Traffic control: Immediately report to the traffic police department, and temporary traffic control (such as closing some lanes or the entire road section) will be implemented based on the actual situation. Emergency coordination: The cloud platform will initiate a high-level emergency response, coordinating the efforts of multiple departments, including traffic police, firefighters, and medical rescue teams, to rush to the scene for emergency response and rescue preparations. Resource prioritization: Ensure that emergency supplies (such as large drainage equipment and rescue vehicles) are prioritized for high-risk sections.
[0038] In the embodiment of the present invention, the optimization of the heterogeneous feature coupling processing result based on the first learnable parameter to obtain the target rule embedding feature can be implemented through the following example.
[0039] Optimizing a heterogeneous feature coupling processing result based on the first learnable parameter, and superimposing the optimized heterogeneous feature coupling processing result and the first rule embedding feature to obtain a second rule embedding feature;
[0040] The second rule embedding feature is subjected to feature conversion, and the feature conversion result is optimized based on a second learnable parameter to obtain a target rule embedding feature, wherein the second learnable parameter is determined by training after fixing the parameters of the risk warning model.
[0041] In this embodiment of the present invention, the server processes the risk warning process of a highway section from 14:00 to 14:05 as an example, and describes in detail the execution process of "optimizing the heterogeneous feature coupling processing results based on the first learnable parameter to obtain the target rule embedding feature". In the previous process, the server has completed two key tasks: heterogeneous feature coupling processing: integrating the risk warning rules (such as " <start>Moderate road flooding risk <end> <start>The first rule embedding feature (11 tokens, each token containing 768-dimensional semantic information, retaining the core semantics of "moderate road surface waterlogging risk") is fused with the integrated road domain spatio-temporal coupling feature (associating the spatio-temporal data of the current scene with historical similar waterlogging scenes, such as the spatial location and time change trend of historical waterlogging areas), resulting in a heterogeneous feature coupling processing result (11 tokens, each token containing 512-dimensional features, integrating rule semantics and spatio-temporal information). The first rule embedding feature: The original rule semantic feature (11×768 matrix) extracted by the BERT model is the "semantic benchmark" for risk warning, directly reflecting the core meaning of the rule. The first task of the server is to adjust the dimension of the heterogeneous coupling result to be consistent with the first rule embedding feature (from 512 dimensions to 768 dimensions) for subsequent fusion. The "first learnable parameter" here is a pre-trained linear layer (composed of weights and biases), specifically used to convert the data-type spatio-temporal-rule fusion feature (512 dimensions) into the semantic-type rule embedding feature dimension (768 dimensions). For example, the heterogeneous coupling result of the "accumulation" token in the rule (512 dimensions, containing the spatio-temporal features of the historical waterlogging scene corresponding to "accumulation", such as the expansion speed of the waterlogging area and the vehicle deceleration situation), after being processed by this linear layer, becomes a 768-dimensional feature, consistent with the original rule embedding feature of "accumulation" (768 dimensions, containing the semantic information of "accumulation", such as "pile up, gather"). After processing, the server obtains an optimized heterogeneous coupling result (11×768 matrix), and the features of each token not only retain the spatio-temporal coupling information but also adapt to the dimension requirements of the rule embedding, preparing for the next fusion. Next, the server performs an element-wise addition of the optimized heterogeneous coupling result and the first rule embedding feature (that is, directly adding the 768-dimensional vectors corresponding to each token). The core purpose of this step is to retain the original semantics of the rule while integrating spatio-temporal coupling information. For example, the original rule embedding feature of the "accumulation" token (768 dimensions) contains the semantics of "accumulation" (such as "pile up, gather"), while the optimized heterogeneous coupling result (768 dimensions) contains the spatio-temporal features corresponding to "accumulation" (such as the expansion speed of the "accumulation" waterlogging area and the vehicle deceleration situation in the historical waterlogging scene) in the historical waterlogging scene. After adding the two, the feature of the "accumulation" token (768 dimensions) not only retains the semantics of "accumulation" but also integrates spatio-temporal information, achieving the synergy of "semantics + data" - for example, "accumulation" not only represents "pile up" but also associates with spatio-temporal scenes such as "expansion of the waterlogging area" and "vehicle deceleration". After superposition, the server obtains the second rule embedding feature (11×768 matrix), which is the preliminary fusion result of semantics and spatio-temporal information, retaining the core meaning of the rule while adding the support of spatio-temporal data. To make the second rule embedding feature more suitable for the input of the risk warning model (such as the Transformer encoder), the server needs to enhance its non-linear expression ability.The "second learnable parameter"—a pre-trained linear layer with a ReLU activation function—is used here. This linear layer performs a nonlinear transformation on the second rule embedding, suppressing irrelevant features and enhancing key information. For example, after processing the second rule embedding feature (768 dimensions) of the "accumulation" token through this layer, it emphasizes the semantic associations between "accumulation" and tokens such as "water" and "risk" (e.g., the logical chain from "accumulation" to "risk"). It also strengthens spatiotemporal synergies with the current scenario (e.g., the association between the current flooding area's expansion rate and vehicle deceleration and "accumulation"). For example, the association between "accumulation" and "water" is strengthened, while the association between "accumulation" and irrelevant words (e.g., "sky") is suppressed. After processing, the server obtains the target rule embedding feature (an 11×768 matrix), which is the final input feature for the risk warning model. This feature not only contains the core semantics of the rule (e.g., "moderate road flooding risk") but also incorporates spatiotemporal information from current and historical scenarios (e.g., the current flooding area's location and expansion rate, and historical experience with similar scenarios), providing enhanced risk prediction capabilities. The first and second learnable parameters are trained by fixing the parameters of the risk warning model. Training data: Historical road monitoring data from the previous year for the road section (including 10,000 samples labeled with true risk status, such as "moderate road flooding risk" and "mild vehicle congestion risk") is used. Training process: Historical data is fed into the risk warning model (e.g., a Transformer encoder), and the difference between the predicted results and the true labels (cross-entropy loss) is calculated. Backpropagation is then used to update only these two learnable parameters until the loss converges (e.g., after 100 epochs of training, the loss drops from 2.1 to 0.4). The goal is to ensure that these two parameters can specifically adjust the fused features, adapt them to the existing risk warning model, and improve the accuracy of risk prediction for the current scenario. For example, by making the features of the "accumulation" token more prominently associated with "water" and "risk," more accurately reflecting the risk status of the current flooding scenario. Through these steps, the server converts the heterogeneously coupled spatiotemporal-rule features (512 dimensions) into a dimensionality consistent with the original rule embedding (768 dimensions), overlaying them to preserve semantic information. Feature representation is then enhanced through nonlinear transformations, ultimately yielding the target rule embedding. This feature not only embodies the core semantics of the risk rule but also integrates the spatiotemporal information of current and historical scenarios, laying the foundation for accurate predictions in subsequent risk warning models. For example, when the server embeds the target rule into the feature and inputs it into the risk warning model, the model can more accurately identify the "moderate road flooding risk" in the current scenario and issue a timely warning.
[0042] In an embodiment of the present invention, the risk warning model is provided with a plurality of feature extraction components operating in series, and the risk warning model is called to perform risk prediction based on the target rule embedded features to obtain the risk warning results of the road area monitoring data, which can be implemented through the following examples.
[0043] Extracting features from the target rule embedding features based on the plurality of feature extraction components, performing risk prediction based on the extracted feature representation, and obtaining a risk warning result of the road area monitoring data;
[0044] Wherein, the target rule embedding feature is the input of the first feature extraction component;
[0045] For any remaining feature extraction component except the feature extraction component at the last position, heterogeneous feature coupling processing is performed on the feature extraction result of the current feature extraction component and the first road domain spatiotemporal coupling feature, the heterogeneous feature coupling processing result is optimized based on the first learnable parameter, feature superposition is performed on the optimized heterogeneous feature coupling processing result and the feature extraction result of the current feature extraction component, feature conversion is performed on the features obtained after feature superposition, and after optimizing the feature conversion result based on the second learnable parameter, it is loaded into the subsequent feature extraction component for feature extraction.
[0046] In this embodiment of the present invention, for example, the server processes a risk warning process for a highway section road monitoring image sequence from 14:00 to 14:05. Assuming that the risk warning model consists of three serially connected Transformer encoder layers (feature extraction components 1, 2, and 3, with component 3 being the last), the execution process of "calling the risk warning model to perform risk prediction based on the target rule embedding feature" is described in detail. In the previous process, the server has completed: target rule embedding feature (11×768 matrix): by fusing the risk rule semantics (such as " <start>Moderate road flooding risk <end> <start>") and the spatiotemporal information of the current / historical scenes (such as the location of the flooded area and the expansion speed), and is obtained through learnable parameter optimization. Each token (a total of 11) contains 768-dimensional features, which not only retains the core semantics of the rule but also integrates spatiotemporal data. The first road domain spatiotemporal coupling feature (2048×300 matrix): It integrates the spatiotemporal information of the current scene (such as the spatial distribution of the flooded area and the time variation of the vehicle speed) and serves as the "data benchmark" for the current scene. The server embeds the target rule feature as component 1 (the first Transformer encoding The core function of component 1 is to capture long-range associations within the semantics of rules (such as the logical relationship between "accumulation" and "water" and "risk"). For example, the "accumulation" token (768 dimensions) in the target rule embedding feature contains the semantics of "accumulation" (such as "accumulation") and historical spatiotemporal information (such as the expansion rate of historical waterlogging scenes). Component 1 uses the self-attention mechanism to calculate the association strength between "accumulation" and other tokens (such as "water", "wind", and "risk"), strengthening the semantic chain of "accumulation→water→risk" and weakening the association between "accumulation" and irrelevant tokens (such as " <start>Association with "). After processing, Component 1 outputs the feature extraction result (11×768 matrix), and the features of each token are more prominent in the internal association of the rule semantics - for example, the association between the feature of "accumulation" and the feature of "water" has increased by 30% (calculated based on attention weights). Component 2 is not the last component and needs to further fuse the result of Component 1 with the spatio-temporal information of the current scene (the first road domain spatio-temporal coupling feature) to enhance the scene adaptability of the features. The processing steps are as follows: Heterogeneous feature coupling: The server performs heterogeneous fusion of the feature extraction result of Component 1 (11×768, semantic type) and the first road domain spatio-temporal coupling feature (2048×300, data type) - the process is the same as the logic of fusing the rule embedding feature and the spatio-temporal feature before: fuse each token feature of Component 1 (such as the 768-dimensional semantic feature of "accumulation") with the corresponding spatio-temporal data in the first road domain spatio-temporal coupling feature (such as the expansion speed of the current water accumulation area and the 2048-dimensional feature of vehicle deceleration situation), and obtain the heterogeneous coupling processing result (11×512 matrix). For example, the heterogeneous coupling result (512-dimensional) of the "accumulation" token contains both the semantic association between "accumulation" and "water" and "risk" (from Component 1), and spatio-temporal information such as the real-time expansion speed of the current water accumulation area (such as increasing by 0.5 square meters every 10 seconds) and the deceleration ratio of vehicles passing through the water accumulation section (such as 30% of the vehicles decelerating to below 60 km / h). First learnable parameter optimization: Use the first learnable parameter (pre-trained linear layer) to convert the heterogeneous coupling result (512-dimensional) into the same dimension as the result of Component 1 (768-dimensional), and obtain the optimized heterogeneous coupling result (11×768 matrix). For example, after being processed by this layer, the heterogeneous coupling result (512-dimensional) of "accumulation" becomes a 768-dimensional feature, adapting to the dimension of the result of Component 1. Feature stacking: Element-wise add the optimized heterogeneous coupling result (11×768) and the feature extraction result of Component 1 (11×768) to obtain the stacked feature (11×768 matrix). This step retains the semantic association information of Component 1 (such as "accumulation→water→risk"), and at the same time fuses the spatio-temporal data of the current scene (such as the water accumulation expansion speed) - for example, the stacked feature (768-dimensional) of "accumulation" contains both the semantic association between "accumulation" and "water" and the real-time expansion situation of the current water accumulation area. Second learnable parameter optimization: Use the second learnable parameter (pre-trained linear layer with ReLU activation) to perform a non-linear transformation on the stacked feature, suppress irrelevant features (such as the weak association between "accumulation" and "sky"), and strengthen key information (such as the strong association between "accumulation" and "water accumulation expansion" and "vehicle deceleration"), and obtain the output feature of Component 2 (11×768 matrix). For example, after being processed by this layer, the association strength between "accumulation" and the "water accumulation expansion speed" of the stacked feature of "accumulation" increases from 0.4 to 0.7 (calculated based on the cosine similarity of the feature vectors), which can better reflect the actual situation of the current scene.Component 3 is the final component and eliminates the need for further heterogeneous coupling (the spatiotemporal information of the current scene has already been integrated by Component 2). Its core function is to enhance the expressive power of the final features. The server inputs the output features of Component 2 (11×768) into Component 3, where a self-attention mechanism further strengthens long-range dependencies between tokens (e.g., the relationship between "moderate" and "waterlogging" and "risk"). For example, after processing by Component 3, the correlation between the features of the "moderate" token and "waterlogging" and "risk" increases by 25%, strengthening the semantic link of "moderate road flooding risk." After processing, Component 3 outputs the final feature representation (an 11×768 matrix). The server inputs the final feature representation from Component 3 into the linear classification layer (which outputs four risk levels: no risk, mild, moderate, and severe), resulting in a risk probability distribution for each token (an 11×4 matrix). Based on the rule generation logic, the last start symbol (the 11th token) is used. <start>) corresponds to the risk status of the current scenario (because the <start>To correlate the current scenario, the server extracts the token's probability distribution: no risk: 0.12; mild risk: 0.21; moderate risk: 0.63; severe risk: 0.04. The server selects the category with the highest probability (moderate risk) as the risk warning result for the road monitoring data—moderate road flooding risk. The risk warning model's multiple serial components achieve accurate predictions by progressively integrating semantic and spatiotemporal information: Component 1 (the first component): captures associations within the rule semantics, laying the semantic foundation; Component 2 (the middle component): integrates the spatiotemporal information of the current scenario to enhance the feature's scenario adaptability; Component 3 (the last component): strengthens the final feature expression and improves risk prediction accuracy. Through this process, the server deeply integrates the target rule's embedded features (semantics + historical spatiotemporal information) with the current scenario's spatiotemporal information (the first road-domain spatiotemporal coupling feature), ultimately outputting a risk warning result (moderate road flooding risk) that reflects the current scenario's actual conditions, providing precise decision-making for highway management.
[0047] In an embodiment of the present invention, each of the feature extraction components is respectively configured with the corresponding first learnable parameter and the corresponding second learnable parameter. Before optimizing the heterogeneous feature coupling processing result based on the first learnable parameter, the embodiment of the present invention also provides the following implementation method.
[0048] Obtaining a road monitoring data instance, fixing the parameters of the risk warning model, and initially assigning zero values to each of the first learnable parameters and each of the second learnable parameters;
[0049] Each of the first learnable parameters and each of the second learnable parameters are trained based on the road domain monitoring data instance.
[0050] In the embodiment of the present invention, the server training of learnable parameters of a highway risk warning model (containing three serially connected Transformer encoder layers, i.e., feature extraction components 1, 2, and 3) is taken as an example to describe in detail the execution process of "configuring each feature extraction component corresponding to the first / second learnable parameter" and "parameter initialization and training". The three feature extraction components of the risk warning model (components 1, 2, and 3) are independently configured with corresponding first learnable parameters (denoted as (W1_i), linear layer weights) and second learnable parameters (denoted as (W2_i), linear layer weights with ReLU activation). Among them: Component 1 (the first one, processing the target rule embedding feature): corresponding to (W1_1) (used to optimize the heterogeneous coupling results. Although component 1 does not require heterogeneous coupling, the parameters still need to be initialized), (W2_1) (used to optimize the feature conversion results); Component 2 (middle, non-last, requires integration of spatiotemporal information): corresponding to (W1_2) (optimize heterogeneous coupling results), (W2_2) (optimize feature conversion results); Component 3 (last, strengthen the final feature): corresponding to (W1_3), (W2_3) (although the last position does not require heterogeneous coupling, the parameters still need to be initialized). The server retrieves 10,000 instances of historical road monitoring data for the road section over the past year. Each instance includes: 1. A 5-minute sequence of road monitoring images (e.g., images from 4:00 PM to 4:05 PM on a weekday in the summer of 2023, including flooded areas and vehicle trajectories); 2. The first spatiotemporal coupling feature of the road section (a 2048×300 matrix, integrating spatiotemporal information from the image sequence, such as the spatial distribution of flooded areas and temporal variations in vehicle speed); and 3. Manually annotated real-world risk status (e.g., "moderate road flooding risk" and "mild vehicle congestion risk," with four categories: none, mild, moderate, and severe). The server fixes the core parameters of the Transformer encoder layer in the model (which do not participate in this training), including the query (Q), key (K), and value (V) weight matrices of the self-attention mechanism (used to calculate the strength of associations between tokens); and the weights and biases of the feedforward neural network (FFN) (used to enhance feature representation). These parameters are obtained by pre-training the model on large-scale text / image data and have general feature extraction capabilities. This training only adjusts the learnable parameters corresponding to the components ((W1_i), (W2_i)).The server initializes the first and second learnable parameters of each component to zero: Component 1's (W1_1) (shape 512×768, used to convert the 512-dimensional heterogeneous coupling result to 768-dimensional), (W2_1) (shape 768×768, used for nonlinear transformation), and the corresponding bias terms (b1_1, b2_1)) are all set to all-zero matrices / vectors. The same applies to component 2's (W1_2) (512×768), (W2_2) (768×768) and bias terms. Component 3's (W1_3) (512×768), (W2_3) (768×768) and bias terms are also initialized to zero. The purpose of zero initialization is to start with no prior knowledge of the parameters and adapt them to the current task through training.The server uses historical data instances and only updates (W1_i) and (W2_i) for each component. The training process is as follows (taking Component 2 as an example): 1. Input data: Embed the target rule features (11×768, integrating rule semantics and historical spatio-temporal information) of a certain data instance (such as the water accumulation scenario from 16:00 to 16:05 on August 15, 2023) into the model; 2. Processing by Component 1: Component 1 captures rule semantic associations (such as the relationship between "accumulation" and "water") through the self-attention mechanism and outputs features (11×768); 3. Processing by Component 2 (not the last one): Heterogeneous coupling: Integrate the output of Component 1 with the first-way domain spatio-temporal coupling features (2048×300, spatio-temporal data of the current scenario) to obtain the heterogeneous coupling result (11×512, including the current water accumulation expansion speed corresponding to "accumulation" and the vehicle deceleration situation); Optimization of (W1_2): Use (W1_2) of Component 2 (512×768) to convert the heterogeneous coupling result (512 dimensions) into 768 dimensions consistent with the output of Component 1, obtaining the optimized heterogeneous coupling result (11×768); Feature superposition: Add the optimized heterogeneous coupling result to the output of Component 1 (11×768), integrating the current spatio-temporal information while retaining semantic associations (such as the features of "accumulation" contain both the semantics of "accumulation → water" and the current water accumulation expansion speed); Optimization of (W2_2): Use (W2_2) of Component 2 (768×768) to perform a non-linear transformation (ReLU activation) on the superimposed features, suppressing irrelevant features (such as the weak association between "accumulation" and "sky") and strengthening key information (such as the strong association between "accumulation" and "water accumulation expansion"), obtaining the output of Component 2 (11×768); 4. Processing by Component 3 (the last one): Component 3 strengthens the final features (such as the relationship between "medium", "water accumulation", and "risk") through the self-attention mechanism and outputs the final features (11×768); 5. Calculate the loss: Input the final features into the linear classification layer to obtain the risk probability distribution of each token (11×4), take the probability distribution of the last start symbol (corresponding to the current scenario) (such as the medium risk probability of 0.7), and calculate the cross-entropy loss with the true label ("medium road surface water accumulation risk") ((-ln(0.7)≈0.36)); 6. Backpropagation and parameter update: Calculate the gradients of the loss with respect to (W1_2) and (W2_2) of Component 2 through backpropagation (such as the gradient of (W1_2) represents its influence on the loss), and use the Adam optimizer to update these parameters (similarly for (W1_1), (W2_1) of Component 1, and (W1_3), (W2_3) of Component 3), while the core parameters of the Transformer (such as self-attention weights) remain unchanged; 7. Iterative training: Repeat the above process, perform 100 rounds of training on 10,000 data instances until the loss converges (such as the loss before training is 2.1 and drops to 0.4 after training).After training, each component's (W1_i) and (W2_i) are adapted to the corresponding component's processing tasks: Component 2's (W1_2) accurately converts the heterogeneous coupling results (512 dimensions) into 768 dimensions, preserving the spatiotemporal relationship between "accumulation" and "water expansion." Component 2's (W2_2) strengthens the semantic chain of "accumulation → water → risk" and suppresses irrelevant information (such as the association between "accumulation" and "cloud layer"). Component 1's (W1_1) and (W2_1) (although Component 1 does not require heterogeneous coupling) are also adjusted through training to make Component 1's output more suitable for subsequent processing. The first and second learnable parameters of each feature extraction component are obtained by fixing the core model parameters and training with historical data examples. This allows each component to better process the corresponding features (such as Component 2's integration of spatiotemporal information), thereby improving the accuracy of risk prediction. For example, after training, components 2 (W1_2) and (W2_2) can more accurately integrate the spatiotemporal information of the current flooding scene with the rule semantics, increasing the predicted probability of "moderate road flooding risk" output by the model from 0.5 before training to 0.7, which is closer to the true label.
[0051] In an embodiment of the present invention, the acquisition of risk warning rules, extraction of the first rule embedding features of the risk warning rules, and heterogeneous feature coupling processing of the first rule embedding features and the first road domain spatiotemporal coupling features can be implemented through the following examples.
[0052] Extracting a second road-domain spatiotemporal coupling feature similar to the first road-domain spatiotemporal coupling feature, and a baseline road-domain risk state associated with the second road-domain spatiotemporal coupling feature, wherein the baseline road-domain risk state is a risk state of baseline road-domain monitoring data, the second road-domain spatiotemporal coupling feature is obtained by fusing time dimension features and spatial dimension features of the baseline road-domain monitoring data, and the spatiotemporal dimensions of the second road-domain spatiotemporal coupling feature are the same as the spatiotemporal dimensions of the first road-domain spatiotemporal coupling feature;
[0053] Constructing a risk warning rule based on the benchmark road risk state, and extracting a first rule embedding feature of the risk warning rule;
[0054] The second road domain spatiotemporal coupling feature is superimposed on the first road domain spatiotemporal coupling feature to obtain an integrated road domain spatiotemporal coupling feature, and heterogeneous feature coupling processing is performed on the first rule embedding feature and the integrated road domain spatiotemporal coupling feature.
[0055] In this embodiment of the present invention, the risk warning process for a server processing a sequence of road monitoring images from 2:00 PM to 2:05 PM on a particular highway section (the current scene, including spatiotemporal information such as shallow flooding and vehicle deceleration) is described in detail. The server fuses the temporal features of the current scene (vehicle speed changes and flooded area expansion rate) with the spatial features (spatial distribution of flooded areas and lane markings) to produce a first spatiotemporal coupling feature (a 2048×300 matrix, with 2048 representing the feature dimension and 300 representing the number of time windows). This fully preserves the spatiotemporal coordination information of the current scene (e.g., the correlation between "flooded area expansion" and "vehicle deceleration"). The server then retrieves historical data from the road risk feature library (which stores baseline road monitoring data for the same section over the past three years) that matches the current scene's spatiotemporal features to correlate the historical risk status. 1. Risk signature database content: Each baseline data item in the database includes: baseline road monitoring data (e.g., an image sequence from 2:00 PM to 2:05 PM on a weekday in the summer of 2023, consistent with the current time period and weather (after light rain)); secondary road domain spatiotemporal coupling features (a 2048 × 300 matrix that integrates the temporal and spatial dimensions of the baseline data and is identical to the primary road domain spatiotemporal coupling features); and baseline road domain risk status (manually annotated actual risk type and level, such as "moderate road flooding risk"). 2. Similarity Search: The server first calculates the index features of the first road-domain spatiotemporal coupling feature (mean spatial dimension features + mean temporal dimension features, 3072 dimensions) for rapid candidate screening. Using the KD tree algorithm, the server retrieves the 10 candidate data items from the database whose index features are most similar to the current index features (e.g., flooding scenarios from August 10, 2023, September 5, 2023, etc.). The server then calculates the cosine similarity (a measure of spatiotemporal similarity) between the second road-domain spatiotemporal coupling feature of the candidate data and the current first road-domain spatiotemporal coupling feature, ultimately selecting the most similar benchmark data item (the scenario from 2:00 PM to 2:05 PM on August 10, 2023, with a similarity of 0.85). The server then extracts the second road-domain spatiotemporal coupling feature of this benchmark data item (2048 × 300, containing spatiotemporal information such as the spatial distribution of historical flooded areas and vehicle deceleration) and the benchmark road-domain risk status ("moderate road flooding risk"). The server then converts the historical benchmark risk status into risk warning rules to link the current scenario with the historical risk logic. 1. Construct risk warning rules: The rule format follows the "start symbol + baseline risk status + end symbol + start symbol" format, which is used to clarify the semantic boundaries of the rule and the connection points of the current scenario. For example, if the baseline risk status is "moderate road flooding risk", the constructed rule is: [ <start>Moderate road flooding risk <end> <start>];in, <start>Indicates the start of the rule. <end>Indicates the end of the rule. <start>Used to associate the current scenario (i.e., the "continuation" of the rule, indicating that the current scenario needs to refer to historical rules to determine risk). 2. Extract the first rule embedding feature: The server uses a pre-trained BERT-base model (vocabulary contains 30,000 tokens, hidden layer dimension 768) to perform semantic embedding on the rule: split the rule into a token sequence: <start>"、"Medium、"Degree、"Road、"Surface、"Accumulation、"Water、"Wind、"Danger、" <end> ”、" <start>” (a total of 11 tokens); through BERT's word embedding layer (converting tokens into 768-dimensional vectors) and position embedding layer (retaining the order information of tokens), the first rule embedding feature (768 dimensions) of each token is generated, and finally an 11×768 rule embedding matrix is formed. This matrix retains the core semantics of the rule (such as "moderate" represents the risk level, "water accumulation" represents the risk type) and sequential logic (such as the semantic chain of "water accumulation → risk"). The server needs to fuse the historical spatiotemporal features (the second road domain spatiotemporal coupling features) with the current spatiotemporal features (the first road domain spatiotemporal coupling features), and then perform heterogeneous fusion (cooperation of data and semantic features) with the rule embedding features (semantic type). 1. Generate integrated road domain spatiotemporal coupling features: Combine the second road domain spatiotemporal coupling features (2048×300, spatiotemporal information of historical similar scenes) with the first road domain spatiotemporal coupling features The feature vectors of the corresponding time windows (2048×300, the spatiotemporal information of the current scene) are superimposed (i.e., the 2048-dimensional feature vectors of each time window are added together) to obtain the integrated road-domain spatiotemporal coupling feature (2048×300). For example, the flooded area feature (one of the 2048 dimensions) in the "14:00-14:01" time window in the historical scenario is 0.6 (after normalization), while the same feature in the same time window in the current scenario is 0.5. After superposition, the value is 1.1 (normalization will be performed later). This step integrates the spatiotemporal information of the historical and current scenarios (for example, the synergy between the historical and current flood expansion rates). 2. Heterogeneous Feature Coupling Processing: The server must fuse the first rule embedding feature (11×768, semantic) with the integrated road-domain spatiotemporal coupling feature (2048×300, data) to achieve synergy between "rule semantics" and "spatiotemporal data."The processing logic is as follows (taking the "product" token in the rule as an example): Feature mapping: Use two learnable linear layers to respectively map the rule embedding features of the "product" token (768 dimensions, containing the semantic information of "product") and the spatiotemporal features corresponding to "product" in the integrated road domain spatiotemporal coupling features (2048 dimensions, containing the spatial distribution of historical and current waterlogged areas, expansion speed and other information) to a low-dimensional space (such as 512 dimensions), and obtain the rule feature mapping vector (512 dimensions) and the spatiotemporal feature mapping vector (512 dimensions); Association strength calculation: Calculate the dot product (512 dimensions) of the two mapping vectors to obtain the association strength vector (indicating the degree of association between the semantics of "product" and the spatiotemporal features, such as the association strength between "product" and "expansion of waterlogged areas"). The degree of correlation is 0.7, and the correlation strength with "vehicle speed" is 0.3. Feature masking and weight normalization: A feature masking vector (elements are 1 when the correlation strength exceeds the threshold of 0.5, and 0 otherwise) is used to suppress irrelevant features (such as the weak correlation between "product" and "vehicle speed"). Softmax normalization is then performed on the masked correlation strength vector to obtain a feature coupling weight vector (512 dimensions, representing the contribution of spatiotemporal features to the semantics of "product"). Weighted fusion: A linear layer is used to map the spatiotemporal features of "product" (2048 dimensions) from the integrated spatiotemporal coupling features of the road domain into a spatiotemporal feature representation vector (512 dimensions). This is then dot-producted with the feature coupling weight vector to obtain a heterogeneous coupling result (512 dimensions) for the "product" token. This result contains both the semantics of "product" (such as accumulation, aggregation) and historical and current spatiotemporal information (such as the expansion rate of the flooded area and vehicle deceleration). This process is repeated to process all 11 tokens in the rule (including the last one). <start>, corresponding to the current scenario), ultimately resulting in an 11×512 matrix of heterogeneous feature coupling results. Through this process, the server implements the complete chain of "historical scenario association → rule construction → semantic embedding → spatiotemporal fusion → heterogeneous coupling": 1. Find historical data with similar spatiotemporal characteristics to the current scenario from the risk feature library to obtain a baseline risk state; 2. Convert the baseline risk state into a risk warning rule and extract semantic embedding features; 3. Fuse the spatiotemporal features of the historical and current scenarios to obtain an integrated spatiotemporal coupling feature; 4. Heterogeneously fuse the rule semantic features with the integrated spatiotemporal features to obtain a coupling result with both semantic and data support. This result lays the foundation for subsequent learnable parameter optimization and risk warning model prediction, ensuring that the model can combine historical experience with current data to accurately determine road area risks (such as the "moderate road flooding risk" in the current scenario).
[0056] In an embodiment of the present invention, the risk warning rule constructed based on the reference road risk status may be implemented through the following examples.
[0057] Adding a start symbol and a stop symbol at the beginning and end of the reference road risk state to obtain a rule element text corresponding to the reference road risk state;
[0058] The start symbol is added after the rule element text at the end to obtain the risk warning rule.
[0059] In an embodiment of the present invention, exemplarily, taking the risk warning process of the server processing a road domain monitoring image sequence from 14:00 to 14:05 on a certain highway section (the current scene, including spatiotemporal information such as shallow water accumulation areas and vehicle deceleration), the execution process of "constructing risk warning rules based on the baseline road domain risk state" is described in detail. The server has retrieved historical benchmark data (the water accumulation scene from 14:00 to 14:05 on August 10, 2023, with a cosine similarity of 0.85) similar to the current scene (the first road domain spatiotemporal coupling feature) from the road domain risk feature library, and extracted the baseline road domain risk state of the benchmark data - "moderate road surface waterlogging risk" (the real risk level manually labeled, reflecting the risk type and severity of the historical scene). The server's primary task is to clarify the semantic boundaries of the baseline risk state so that subsequent models can accurately identify the start and end of the rule. To this end, the server adds start symbols (denoted as ) at the beginning and end of the baseline road domain risk state according to the preset format. <start>) and the terminator (denoted as <end>). Action: The server takes the baseline risk status "moderate road flooding risk" as the core content and adds <start>, add <end>; Result: Get the rule element text—— <start>Moderate road flooding risk <end>; Purpose: Semantic boundary division: <start>Indicates the "start" of the rule, <end>Indicates the "end" of the rule, allowing the model to clearly distinguish the rule content from other information (such as subsequent current scene data); Sequential information preservation: The fixed position of the symbol (first → core content → end) conforms to the sequence logic of natural language, making it easier to retain the sequential semantics of the rule (such as the logical chain of "moderate" → "road flooding" → "risk") when extracting embedded features using models such as BERT. In order to associate the current scene with historical rules, the server needs to place the last position of the rule element text (i.e. <end>After that) add the start symbol again ( <start>). Action: The server will rule element text <start>Moderate road flooding risk <end>As a base, add <start>; Result: Get complete risk warning rules—— <start>Moderate road flooding risk <end> <start>; Purpose: Current scene association: End <start>Indicates the "continuation" of the rule, that is, the historical rule ( <start>Moderate road flooding risk <end>) is completed, it is necessary to connect the information of the current scene (such as the current first-path spatiotemporal coupling feature); the model processing logic guides: in the subsequent rule embedding and heterogeneous feature coupling process, the model can recognize: the end <start>Corresponding to the risk judgment of the current scenario (that is, the semantics of historical rules need to be combined with the spatiotemporal data of the current scenario to predict the current risk). The risk warning rules built by the server strictly follow the preset format specifications ( <start>+Baseline Risk Status+ <end> + <start>), ensure: the accuracy of rule embedding: when the BERT model is used to extract the first rule embedding features, the symbol ( <start> 、 <end>) will be treated as a special token, preserving the semantic structure of the rule (such as <start>The "moderate road flooding risk" at the end is the core rule content. <end>Behind <start>is the connection point of the current scene); Targeted heterogeneous coupling: In the subsequent heterogeneous feature coupling processing, the model can accurately distinguish the historical semantic part of the rule ( <start>Moderate road flooding risk <end>) associated with the current scene (the <start>), and fuse the corresponding spatiotemporal features (historical spatiotemporal coupling features of the second road domain and current spatiotemporal coupling features of the first road domain). The process of building risk warning rules based on the benchmark road domain risk status is essentially to transform historical risk experience into a semantic sequence that can be understood by the model: the first step is to add <start>and <end>, clarifying the semantic boundaries of historical rules; the second step is to add the final <start>, establish the association between historical rules and current scenarios; the final generated rules ( <start>Moderate road flooding risk <end> <start>) not only retains the core semantics of historical risks ("moderate road flooding risk"), but also reserves an interface for the subsequent integration of spatiotemporal data of the current scene, ensuring that the model can combine historical experience with current data to accurately judge road risks. For example, in the subsequent rule embedding process, the BERT model will <start>It is identified as the "beginning of the current scene" and, during heterogeneous coupling, the embedded features of the token are fused with the first-path spatiotemporal coupling features of the current scene, thereby achieving collaborative risk prediction of "historical rules + current data".
[0060] In the embodiment of the present invention, the heterogeneous feature coupling processing of the first rule embedding feature and the integrated road domain spatiotemporal coupling feature can be implemented through the following examples.
[0061] performing heterogeneous feature coupling processing on a first feature component of the rule element text in the first rule embedding feature and a second road domain spatiotemporal coupling feature corresponding to the first feature component in the integrated road domain spatiotemporal coupling feature;
[0062] Heterogeneous feature coupling processing is performed on the second feature component of the start symbol located at the last position in the first rule embedding feature and the first road-domain spatiotemporal coupling feature in the integrated road-domain spatiotemporal coupling feature.
[0063] In this embodiment of the present invention, the server processes the risk warning process of a highway section from 14:00 to 14:05 on a road monitoring image sequence (the current scene) as an example. The execution process of "heterogeneous feature coupling processing of the first rule embedded features and the integrated road domain spatiotemporal coupling features" is described in detail. The first rule embedded features: The server constructs a risk warning rule based on the historical baseline risk status ("moderate road flooding risk") - <start>Moderate road flooding risk <end> <start>, and extracted an 11×768 rule embedding matrix (11 tokens, each token contains 768-dimensional semantic features) through the BERT model. Among them: Rule element text: the first 10 tokens ( <start>,"medium","degree","road","surface","area","water","wind","risk", <end>), corresponding to the first feature component (semantic information of historical rules); the last starting symbol: the 11th token ( <start>), corresponding to the second feature component (associated points in the current scene). Integrate the road domain spatio-temporal coupling features: The server superimposes the first road domain spatio-temporal coupling feature of the current scene (2048×300, including spatio-temporal information such as the spatial distribution of the current water accumulation area and the vehicle deceleration situation) with the second road domain spatio-temporal coupling feature of the historical similar scene (2048×300, including the spatio-temporal information of the water accumulation scene on August 10, 2023), and obtains an integrated spatio-temporal feature of 2048×300 (fusing historical and current spatio-temporal data). The first feature component of the rule element text (768-dimensional embedding of the first 10 tokens) corresponds to the second road domain spatio-temporal coupling feature (spatio-temporal data of the historical scene). The server needs to fuse the two to achieve the coordination of "historical rule semantics" and "historical spatio-temporal data". Taking the "accumulate" token in the rule (the 6th token, with the semantic meaning of "pile up, gather") as an example, the coupling process is described in detail as follows: Feature mapping: The server processes the embedding feature of "accumulate" and the corresponding historical spatio-temporal feature with two learnable linear layers (preset to output 512 dimensions) respectively: maps the 768-dimensional semantic embedding of "accumulate" (including the semantic information of "accumulate", such as "pile up") to a rule feature mapping vector (512 dimensions); maps the second road domain spatio-temporal coupling feature related to "accumulate" in the integrated spatio-temporal feature (2048 dimensions, including data such as the spatial distribution and expansion speed of the water accumulation area on August 10, 2023) to a spatio-temporal feature mapping vector (512 dimensions). Association strength calculation: The server calculates the dot product of the two mapping vectors (512 dimensions) to obtain an association strength vector (indicating the degree of association between the semantics of "accumulate" and the historical spatio-temporal feature). For example, the association strength between the semantics of "accumulate" and the "expansion speed of the water accumulation area" is 0.7 (high association), and the association strength with the "vehicle brand" is 0.2 (low association). Feature masking and weight normalization: The server uses a feature masking vector (the element is 1 when the association strength exceeds the threshold of 0.5, otherwise 0) to suppress irrelevant features (such as the weak association between "accumulate" and "vehicle brand"), and then performs Softmax normalization on the masked association strength vector to obtain a feature coupling weight vector (512 dimensions, indicating the contribution weight of the historical spatio-temporal feature to the semantics of "accumulate"). For example, the weight of the "expansion speed of the water accumulation area" is 0.6 (major contribution), and the weight of the "road surface temperature" is 0.1 (minor contribution). Weighted fusion: The server uses a linear layer to map the historical spatio-temporal feature related to "accumulate" in the integrated spatio-temporal feature (2048 dimensions) to a spatio-temporal feature representation vector (512 dimensions), and then performs a dot product with the feature coupling weight vector to obtain a heterogeneous coupling result of the "accumulate" token (512 dimensions). This result contains both the semantics of "accumulate" (such as "pile up") and the spatio-temporal information of the historical scene (such as the expansion speed of the water accumulation area on August 10, 2023).The server repeats the above process for the first 10 tokens of the rule element text (such as "中", "度", "水", etc.), and finally obtains a 10×512 coupling result matrix (corresponding to the fusion of historical rule semantics and historical spatio-temporal data). The second feature component of the end start symbol (the 768-dimensional embedding of the 11th token). <start>) corresponds to the first domain spatiotemporal coupling feature (the spatiotemporal data of the current scene). The server needs to fuse the two to achieve the coordination of "current scene association" and "current spatiotemporal data". <start>Token as an example, the coupling process is described in detail: Feature mapping: The server uses two learnable linear layers (512-dimensional output) to process the last bit respectively. <start>The embedding features and the corresponding current spatiotemporal features: <start>The 768-dimensional embedding (including the semantic information of "the beginning of the current scene") is mapped into a regular feature mapping vector (512 dimensions); the integration of spatiotemporal features and the last <start>The relevant first road domain spatiotemporal coupling features (2048 dimensions, including the spatial distribution of flooded areas and vehicle deceleration data from 14:00 to 14:05) are mapped into a spatiotemporal feature mapping vector (512 dimensions). Calculation of association strength: The server calculates the dot product (512 dimensions) of the two mapping vectors to obtain an association strength vector (indicating the degree of association between the semantics of "the start of the current scene" and the current spatiotemporal features). For example, the last <start>The correlation strength with "the current area of flooded areas" is 0.8 (high correlation), and the correlation strength with "light intensity" is 0.3 (low correlation). Feature masking and weight normalization: The server uses a feature masking vector (threshold 0.5) to suppress irrelevant features (such as "light intensity"), and then performs Softmax normalization on the masked correlation strength vector to obtain a feature coupling weight vector (512 dimensions, representing the contribution weight of the current spatiotemporal features to the semantics of "the beginning of the current scene"). For example, the weight of "the current area of flooded areas" is 0.7 (main contribution), and the weight of "the number of vehicles" is 0.2 (secondary contribution). Weighted fusion: The server uses a linear layer to map the current spatiotemporal features (2048 dimensions) in the integrated spatiotemporal features into a spatiotemporal feature representation vector (512 dimensions), and then performs a dot product with the feature coupling weight vector to obtain the final value. <start>The heterogeneous coupling result of the token (512 dimensions) contains both the semantics of "the beginning of the current scenario" (e.g., "need to determine the current risk") and the spatiotemporal information of the current scenario (e.g., the current expansion rate of the flooded area). After the two-step processing described above, the server generates an 11×512 heterogeneous feature coupling matrix: the first 10 rows represent the fusion of the first feature component corresponding to the rule element text and the second road domain spatiotemporal coupling feature (historical rule semantics + historical spatiotemporal data); the 11th row represents the fusion of the second feature component corresponding to the last start symbol and the first road domain spatiotemporal coupling feature (current scenario association + current spatiotemporal data). This matrix combines both semantics and data support: the historical component (the first 10 rows) provides historical experience with "moderate road flooding risk" (e.g., the semantic association between "accumulation" and "water" and the spatiotemporal characteristics of historical flooding scenarios); the current component (the 11th row) provides real-time data for the current scenario (e.g., the current expansion rate of the flooded area and vehicle deceleration). The server then feeds this coupling result matrix into a learnable parameter optimization module (e.g., the linear layer of the first / second learnable parameters) to further adjust the feature dimensionality and expressiveness. This matrix is then fed into a risk warning model (e.g., a Transformer encoder) to achieve accurate risk predictions based on a combination of historical experience and current data (e.g., "moderate road flooding risk" for the current scenario). The core of heterogeneous feature coupling processing is to distinguish between historical and current feature associations: the first feature component (historical semantics) of the rule element text is fused with the second spatiotemporal feature (historical data) to preserve the empirical logic of historical risks; the second feature component (current association) of the final start symbol is fused with the first spatiotemporal feature (current data) to incorporate real-time information from the current scenario. This processing approach ensures that the model can both "learn from history" and "adapt to the current situation," improving the accuracy and adaptability of road risk warnings.
[0064] In an embodiment of the present invention, the performing heterogeneous feature coupling processing on the first feature component of the rule element text in the first rule embedding feature and the second road domain spatiotemporal coupling feature corresponding to the first feature component in the integrated road domain spatiotemporal coupling feature includes:
[0065] Determine a rule feature mapping vector based on the first rule embedding feature, determine a spatiotemporal feature mapping vector based on the integrated road domain spatiotemporal coupling feature, and obtain an association strength vector based on a multiplication result of the rule feature mapping vector and the spatiotemporal feature mapping vector after feature space transformation;
[0066] Performing feature masking on the association strength vector based on a feature masking vector, and normalizing the masked association strength vector to obtain a feature coupling weight vector, wherein the feature masking vector is used to dynamically suppress remaining vector feature values in the association strength vector except for a target feature value, and the target feature value is determined based on a product operation result of a first feature component of the rule element text in the first rule embedding feature and a second road-domain spatiotemporal coupling feature corresponding to the first feature component in the integrated road-domain spatiotemporal coupling feature;
[0067] A spatiotemporal feature representation vector is determined according to the second road domain spatiotemporal coupling feature, and a heterogeneous feature coupling processing result is obtained according to a multiplication result of the feature coupling weight vector and the spatiotemporal feature representation vector.
[0068] In the embodiment of the present invention, for example, the server processes the "product" token (the first feature component, corresponding to the historical rule semantics "accumulation, aggregation") of the rule element text in the road monitoring image sequence (current scene) of a certain highway section from 14:00 to 14:05 as an example, and describes the execution process of "heterogeneous feature coupling processing" in detail. The first feature component: the rule element text is <start>Moderate road flooding risk <end>The first 10 tokens, where the first rule embedding feature of the "accumulation" token (the 6th token) is a 768-dimensional vector (extracted by BERT, containing the semantic information of "accumulation", such as "pile up, gather"). Corresponding to the second spatio-temporal feature: The integrated road domain spatio-temporal coupling feature (2048×300) is the superposition of the current first spatio-temporal feature (2048×300, containing the spatio-temporal information of the current waterlogging area) and the historical second spatio-temporal feature (2048×300, containing the spatio-temporal information of the waterlogging scene on August 10, 2023). Among them, the second spatio-temporal feature corresponding to the "accumulation" token is dimensions such as "spatial distribution of waterlogging area" and "expansion speed" in the historical second spatio-temporal feature (a specific subset of the 2048 dimensions, located by the feature alignment algorithm). The server needs to map the high-dimensional rule semantic features and spatio-temporal data features to a low-dimensional space to calculate the correlation strength between the two. Rule feature mapping vector: The server processes the 768-dimensional rule embedding feature of the "accumulation" token with a pre-trained linear layer (input 768 dimensions, output 512 dimensions) to obtain a rule feature mapping vector (512 dimensions). This vector retains the core semantics of "accumulation" (such as "pile up"), while reducing the dimension for subsequent calculations. Spatio-temporal feature mapping vector and space transformation: The server processes the 2048-dimensional second spatio-temporal feature corresponding to "accumulation" (spatio-temporal data of the historical waterlogging scene) with another pre-trained linear layer (input 2048 dimensions, output 512 dimensions) to obtain an initial spatio-temporal feature mapping vector (512 dimensions). To make the spatio-temporal features more adaptable to the feature space of the rule semantics, the server performs a feature space transformation on the initial spatio-temporal feature mapping vector (through a third pre-trained linear layer, 512-dimensional input and output) to obtain a transformed spatio-temporal feature mapping vector (512 dimensions). Calculate the correlation strength vector: The server performs a dot product operation (multiplying corresponding elements and then summing) on the rule feature mapping vector (512 dimensions) and the transformed spatio-temporal feature mapping vector (512 dimensions) to obtain a correlation strength vector (512 dimensions). Each element of this vector represents the degree of correlation between the semantics of "accumulation" and a certain dimension in the historical spatio-temporal features - for example, the correlation strength between "accumulation" and "historical expansion speed of waterlogging area" is 0.7 (high correlation), and the correlation strength with "historical vehicle brand distribution" is 0.2 (low correlation). To suppress irrelevant features (such as "historical vehicle brand distribution") and retain spatio-temporal features strongly correlated with the semantics of "accumulation" (such as "historical expansion speed of waterlogging area"), the server performs the following processing: Generate a feature masking vector: The server calculates the product operation result of the first rule embedding feature (768 dimensions) of the "accumulation" token and the corresponding second spatio-temporal feature (2048 dimensions) (i.e., the dot product matrix of 768×2048, taking the maximum value of each row as the correlation degree of this dimension) to obtain a target feature value (768 dimensions, representing the original correlation degree between the semantics of "accumulation" and each dimension of the second spatio-temporal feature).The server then sets a threshold (e.g., 0.5) and generates a 512-dimensional feature masking vector. When an element in the association strength vector is greater than or equal to the threshold, the corresponding position in the masking vector is set to 1 (retaining the feature); otherwise, it is set to 0 (suppressing the feature). For example, if the association strength corresponding to "historical waterlogged area expansion rate" is 0.7 (greater than 0.5), the corresponding position in the masking vector is 1; if the association strength corresponding to "historical vehicle brand distribution" is 0.2 (less than 0.5), the corresponding position in the masking vector is 0. Feature masking: The server performs element-wise multiplication of the association strength vector (512-dimensional) with the feature masking vector (512-dimensional) to generate a masked association strength vector (512-dimensional). This vector retains only spatiotemporal feature dimensions that are strongly associated with the semantic meaning of "accumulation" (e.g., "historical waterlogged area expansion rate") and suppresses irrelevant dimensions (e.g., "historical vehicle brand distribution"). Normalization Obtains a Feature Coupling Weight Vector: The server performs Softmax normalization on the masked correlation strength vector (converting vector elements to probability values between 0 and 1, summing to 1), resulting in a 512-dimensional feature coupling weight vector. This vector represents the contribution of each dimension of the historical spatiotemporal features to the semantic meaning of "accumulation." For example, the weight of "historical expansion rate of flooded areas" is 0.6 (primary contribution), the weight of "historical spatial distribution of flooded areas" is 0.3 (secondary contribution), and the weight of "historical vehicle brand distribution" is 0 (no contribution). To integrate the semantic information of historical spatiotemporal features with the semantics of "accumulation," the server performs the following processing: Generating a Spatiotemporal Feature Representation Vector: The server processes the 2048-dimensional second spatiotemporal feature corresponding to "accumulation" (the spatiotemporal data of historical flooding scenarios) using a pre-trained linear layer (2048-dimensional input, 512-dimensional output), resulting in a 512-dimensional spatiotemporal feature representation vector. This vector converts the historical spatiotemporal data into a semantically meaningful feature representation (for example, "historical rapid expansion rate of flooded areas" corresponds to a high value in the vector), facilitating integration with regular semantic features.
[0069] Weighted fusion to obtain the coupling result: The server performs an element-wise multiplication of the feature coupling weight vector (512-dimensional) and the spatio-temporal feature representation vector (512-dimensional) to obtain the heterogeneous feature coupling result (512-dimensional) of the "product" token. This result combines the semantic information of the "product" and the contribution of historical spatio-temporal data - for example, the semantics of the "product" ("accumulation") is fused with the spatio-temporal feature of "rapid expansion of the historical waterlogging area" (weight 0.6) to form a semantic feature representation of "accumulated water is expanding rapidly". The server repeats the above process for all the first feature components of the rule element text (such as tokens like "medium", "degree", "water", etc.), and finally obtains a 10×512 coupling result matrix (corresponding to the fusion of historical rule semantics and historical spatio-temporal data). Each element of this matrix represents the associated contribution of a certain semantic token in the historical rule and a certain dimension in the historical spatio-temporal feature. Subsequently, the server merges this coupling result matrix with the coupling result of the end-start symbol (corresponding to the fusion of the current scenario association and the current spatio-temporal data) to obtain an 11×512 complete heterogeneous coupling result matrix, and inputs it into a learnable parameter optimization module (such as a linear layer of the first / second learnable parameters) to further adjust the feature dimension and expression ability, and finally inputs it into a risk warning model (such as a Transformer encoder) to achieve accurate risk prediction of "historical experience + current data" (such as the "medium road surface waterlogging risk" of the current scenario). The heterogeneous coupling process of the first feature component of the rule element text and the corresponding second spatio-temporal feature is essentially to accurately associate and fuse the semantic information of the historical rule and the spatio-temporal data of the historical scenario: by feature mapping and spatial transformation, reducing the high-dimensional features to low-dimensional for easy calculation of associations; by feature masking and weight normalization, retaining the spatio-temporal features strongly associated with the semantics and suppressing irrelevant features; by weighted fusion, integrating the contributions of the spatio-temporal features into the rule semantics according to the weights to form a coupling result with both semantic and data support. This processing method ensures that the model can "draw on historical experience" and lays a foundation for subsequent risk prediction by combining current data.
[0070] In the embodiment of the present invention, the extraction of the second road domain spatio-temporal coupling feature similar to the first road domain spatio-temporal coupling feature and the reference road domain risk state associated with the second road domain spatio-temporal coupling feature can be implemented through the following examples.
[0071] Obtain multiple pieces of the reference road domain monitoring data and the reference road domain risk state of each piece of the reference road domain monitoring data, and fuse the time dimension feature and the space dimension feature of each piece of the reference road domain monitoring data to obtain the second road domain spatio-temporal coupling feature of each piece of the reference road domain monitoring data;
[0072] Each second road domain spatiotemporal coupling feature is associated with the corresponding baseline road domain risk state and stored in a risk feature library, and the second road domain spatiotemporal coupling feature similar to the first road domain spatiotemporal coupling feature and the baseline road domain risk state associated with the second road domain spatiotemporal coupling feature are extracted from the risk feature library.
[0073] In this embodiment of the present invention, the risk warning process for extracting similar second-domain spatiotemporal coupling features and their associated baseline risk states is described in detail, using the server's processing of a sequence of road-area monitoring images from 2:00 PM to 2:05 PM on a particular highway section (for the current scene, a first-domain spatiotemporal coupling feature has been generated, a 2048×300 matrix containing spatiotemporal information about the current flooded area). The server fuses the current scene's spatial features (2048-dimensional, representing the spatial distribution of flooded areas and vehicle positions in the monitoring images, extracted using a ResNet-50) with temporal features (1024-dimensional, representing vehicle speed changes and flooded area expansion rates, extracted using a bidirectional LSTM). Using a spatiotemporal attention mechanism, the server generates the first-domain spatiotemporal coupling feature (2048×300, preserving spatiotemporal synergy information, such as the association between "flooded area expansion" and "vehicle deceleration"). The server needs to retrieve multiple baseline road monitoring data sets from a historical database. These data sets represent real-world monitoring records for the road section over the past three years. Each data set includes: baseline road monitoring data: 5 minutes of continuous high-definition monitoring footage (1920×1080 resolution), vehicle speed sensor data (collected every 1 second), and road surface moisture sensor data (collected every 10 seconds); baseline road risk status: manually annotated risk types and levels (such as "moderate flooding risk," "mild congestion risk," and "no risk"). This annotation is performed by inspectors from the highway management department to ensure accuracy. For example, the server retrieved baseline data from 2:00 PM to 2:05 PM on August 10, 2023: Monitoring footage showed shallow flooding (occupying one-fifth of a single lane) on the road section, and vehicle speeds dropped from 80 km / h to below 60 km / h. Sensor data indicated road surface moisture was 85% (after a light rain). The manually annotated baseline risk status was "moderate flooding risk."The server fuses the spatiotemporal features of each benchmark data to generate a second road domain spatiotemporal coupling feature (the dimension is consistent with the first road domain spatiotemporal coupling feature, 2048×300). The specific steps are as follows: Extract spatial dimension features: Use the pre-trained ResNet-50 model (remove the fully connected layer) to process each frame of the benchmark monitoring image and extract the spatial feature vector (2048 dimensions) to represent the road condition (such as the spatial distribution of the flooded area and the spatial position of the vehicle); Extract temporal dimension features: Use the bidirectional LSTM model (hidden layer 1024 dimensions) to process the temporal sensor of the benchmark data Data (vehicle speed, road surface humidity) is extracted, and a temporal feature vector (1024 dimensions) is extracted to represent dynamic changes (such as the temporal decrease in vehicle speed or the expansion rate of the flooded area). Spatiotemporal fusion: The spatial feature vector (2048 dimensions) is fused with the temporal feature vector (1024 dimensions) through a spatiotemporal attention mechanism. The dot product correlation matrix (300×300, where 300 represents the number of time windows) is calculated. The spatiotemporal attention weights are then calculated using Softmax normalization. The spatial feature vector and the weights are then weighted and summed to produce a second road-domain spatiotemporal coupling feature (2048×300). For example, for the baseline data from August 10, 2023, the spatial feature vector extracted by the server contains the spatial information that the flooded area is located on the right side of the lane, while the temporal feature vector contains the temporal information that the flooded area increases by 0.5 square meters every 10 seconds. The fused second road-domain spatiotemporal coupling feature (2048×300) retains the spatiotemporal synergy that the flooded area is on the right side and is expanding. The server associates the second-domain spatiotemporal coupling feature of each benchmark data entry with the corresponding benchmark road risk status and stores it, along with metadata (such as time, road section, and weather) in a risk feature database (using the FAISS vector database for fast similarity retrieval). For example, the benchmark data entry for August 10, 2023, would be: Second-domain spatiotemporal coupling feature: 2048×300 matrix; Benchmark road risk status: "Moderate road flooding risk"; Metadata: Time (2023-08-10 14:00-14:05), Road section (K120+300-K121+500), Weather (Light rain turning sunny). The server needs to find the second road domain spatiotemporal coupling feature similar to the first road domain spatiotemporal coupling feature (2048×300) of the current scenario from the risk feature library and extract the corresponding baseline risk status. The specific steps are as follows: Calculate similarity: The server uses the cosine similarity algorithm to calculate the similarity between the current first road domain spatiotemporal coupling feature and all second road domain spatiotemporal coupling features in the risk feature library (range 0-1, the larger the value, the more similar); Filter similar results: The server filters out the benchmark data with the highest similarity (the similarity threshold is set to 0.7 to ensure that the similarity is high enough); Extract related information: The server extracts the second road domain spatiotemporal coupling feature and the corresponding baseline road domain risk status of the benchmark data.For example, the cosine similarity between the first road domain spatiotemporal coupling feature (2048×300) of the current scene and the second road domain spatiotemporal coupling feature of the benchmark data on August 10, 2023 is 0.85 (much higher than the threshold of 0.7). The server extracts the following from the benchmark data: the second road domain spatiotemporal coupling feature: a 2048×300 matrix (containing the spatiotemporal information that "the waterlogged area is on the right and is continuously expanding"); the benchmark road domain risk status: "moderate road flooding risk". The server extracts similar second-domain spatiotemporal coupling features (August 10, 2023) and the baseline risk state (moderate flooding), providing historical empirical support for risk warnings in the current scenario. The high similarity (0.85) between the second-domain spatiotemporal coupling features (historical) and the first-domain spatiotemporal coupling features (current) indicates that the spatiotemporal patterns of the current and historical scenarios are highly similar (e.g., both are after light rain, the flooded area is on the right, and vehicles slow down). The baseline risk state (moderate flooding) indicates the risk level of the historical scenario and can provide a reference for risk prediction in the current scenario (e.g., the current scenario is likely to also have a moderate flooding risk). Subsequently, the server superimposes the second-domain spatiotemporal coupling features with the first-domain spatiotemporal coupling features to obtain an integrated spatiotemporal coupling feature (integrating historical and current spatiotemporal information). This feature is then combined with the risk warning rules (built based on the baseline risk state) for heterogeneous feature processing, ultimately inputting it into the risk warning model for accurate prediction. The process of extracting similar second-domain spatiotemporal coupling features and associating them with baseline risk states essentially involves searching historical data for cases with similar spatiotemporal patterns to the current scenario and drawing on their risk experience. By integrating the spatiotemporal features of the baseline data, a second coupling feature consistent with the current feature dimension is generated. This second coupling feature is then associated and stored with the baseline risk state to form a risk feature library. A similarity search is then performed to identify the historical case most similar to the current scenario, extracting its spatiotemporal features and risk state. This approach ensures that the model can "learn from history," improving the accuracy and reliability of risk warnings for the current scenario.
[0074] In an embodiment of the present invention, the benchmark road area monitoring data is a reference road area monitoring image sequence, and the associating each of the second road area spatiotemporal coupling features with the corresponding benchmark road area risk state and storing them in a risk feature library can be implemented through the following examples.
[0075] For each of the reference road monitoring image sequences, performing mean aggregation on the monitoring image features of the reference road monitoring image sequence in the spatial dimension to obtain a risk value sequence feature of the same dimension as the time series data feature of the reference road monitoring image sequence in the temporal dimension;
[0076] Performing mean aggregation on the risk value sequence features and the time series data features of the reference road monitoring image sequence in the time dimension to obtain an index feature of each reference road monitoring image sequence;
[0077] Each of the index features, the second road domain spatiotemporal coupling features and the corresponding reference road domain risk state are associated and stored in a risk feature library.
[0078] In an embodiment of the present invention, for example, the server processes a reference road domain monitoring image sequence (benchmark data, including spatiotemporal information such as shallow water accumulation areas and vehicle deceleration) of a certain highway section from 14:00 to 14:05 on August 10, 2023, and the execution process of "associating the second road domain spatiotemporal coupling features with the benchmark risk state and storing them in the risk feature library" is described in detail. The reference image sequence is a 5-minute high-definition image (7500 frames, 25fps), and the server has completed the following processing: Spatial dimension monitoring image features: Use the pre-trained ResNet-50 model to extract 2048-dimensional spatial features per frame, and then divide the 7500 frames into 300 time windows (25 frames per window), take the average of the spatial features in each window, and obtain a 300×2048 spatial feature sequence (retain the mean of the spatial information in each time window, such as the average distribution position of the water accumulation area). Temporal dimension time series data features: A bidirectional LSTM model is used to extract time series data such as vehicle speed (calculated by displacement between adjacent frames) and waterlogged area changes (semantic segmentation results), resulting in a 300×1024 temporal feature sequence (retaining the mean of dynamic changes within each time window, such as the average rate of decrease in vehicle speed and the average rate of expansion of waterlogged area). Secondary road domain spatiotemporal coupling features: A spatiotemporal attention mechanism is used to fuse these spatial and temporal feature sequences to generate a 2048×300 second spatiotemporal coupling feature sequence (retaining the spatiotemporal synergy between "waterlogged expansion" and "vehicle deceleration"). To associate spatial and temporal features, the server performs mean aggregation on the spatial dimension monitoring image features to generate a risk value sequence feature (300×1024) of the same dimension as the temporal dimension time series data features. The specific implementation is as follows: For each time window (300 in total) of the spatial feature sequence (300×2048), the mean of the 2048-dimensional spatial feature is taken (a 1-dimensional feature representing the average strength of the spatial feature within that window, such as the average coverage of waterlogged areas). This 1-dimensional mean feature for each time window is repeated 1024 times (expanded to 1024 dimensions) to match the dimensionality of the time series data features (1024 dimensions). The resulting 300×1024 risk value sequence feature is obtained (the 1024-dimensional feature for each time window incorporates the mean information of the spatial features). For example, if the mean of the spatial feature for a time window is 0.6 (after normalization, indicating high coverage of waterlogged areas), after expansion to 1024 dimensions, each element is 0.6, forming the risk value sequence feature for that window. To facilitate subsequent retrieval of similar scenarios from the risk feature database, the server performs mean aggregation on the risk value sequence feature and the time series data features to generate a fixed-dimensional index feature (1×1024).Specific implementation: For each time window of the risk value sequence feature (300×1024) and the time series data feature (300×1024), take the element-wise mean of the corresponding positions of the two (for example, the j-th dimension of the i-th window of the risk value sequence feature is 0.6, and the j-th dimension of the i-th window of the time series data feature is 0.5, with a mean of 0.55), to obtain a 300×1024 fused feature sequence. The time dimension (300 windows) of the fused feature sequence is averaged to obtain a 1×1024 index feature (a fixed-dimensional vector that integrates the mean of the spatiotemporal features of the reference image sequence). For example, the mean of the 300 windows of the fused feature sequence is 0.58 (1024 dimensions), forming the index feature of the reference image sequence, which is used to calculate similarity with the index feature of the current scene. The server associates the index features of the reference image sequence, the second road domain spatiotemporal coupling features, and the baseline road domain risk status (manually labeled "moderate flooding risk"), along with metadata (time, road section, and weather) in a risk feature database (using the FAISS vector database to support fast similarity retrieval). The stored entries are as follows: Index features: a 1×1024 vector (used for fast cosine similarity calculation with the current scene's index features); Second road domain spatiotemporal coupling features: a 2048×300 matrix (used for subsequent fusion with the current scene's spatiotemporal features); Baseline road domain risk status: "moderate flooding risk" (used for developing risk warning rules); Metadata: Time (2023-08-10 14:00-14:05), Road section (K120+300-K121+500), Weather (light rain turning sunny). The stored entries in the risk signature library provide historical references for subsequent risk warnings for the current scenario: Index signature: Serving as a "quick retrieval key," when the server processes the current scenario (e.g., the image sequence from 2:00 PM to 2:05 PM on May 10, 2024), it generates an index signature (1 × 1024) for the current scenario. The server then uses the FAISS database to quickly retrieve the benchmark data most similar to the current index signature (e.g., the benchmark data from August 10, 2023, with a similarity of 0.85). Second road domain spatiotemporal coupling signature: The second spatiotemporal coupling signature of the retrieved benchmark data (August 10, 2023) is superimposed with the first spatiotemporal coupling signature of the current scenario to form an integrated spatiotemporal signature (fusing historical and current spatiotemporal information). Baseline road domain risk status: The retrieved baseline risk status ("moderate road flooding risk") is used to construct risk warning rules (e.g., <start>Moderate road flooding risk <end> <start>), providing historical experience support for risk prediction in the current scenario. The process of associating and storing reference image sequences in the risk signature database essentially structures the spatiotemporal characteristics of historical scenarios and risk experience, enabling rapid retrieval of historical cases similar to the current scenario. Risk value sequence features are generated through mean aggregation, matching the feature dimensions of the time series data to link spatial and temporal features. Index features are generated through mean aggregation, providing a fixed-dimensional "search key" for rapid retrieval. The index features, the second spatiotemporal coupling feature, and the baseline risk status are associated and stored to form a complete historical case entry. This approach ensures that the server can efficiently "draw on historical experience," improving the accuracy and efficiency of risk warnings for the current scenario.
[0079] In an embodiment of the present invention, the road monitoring data is a sequence of road monitoring images to be evaluated, and the extraction of temporal and spatial dimension features of the road monitoring data can be implemented by the following example: performing feature extraction on each road monitoring image in the sequence of road monitoring images to be evaluated to obtain monitoring image features of the sequence of road monitoring images to be evaluated in the spatial dimension; extracting temporal feature information from the sequence of road monitoring images to be evaluated, and performing feature extraction on the temporal feature information to obtain time series data features of the sequence of road monitoring images to be evaluated in the temporal dimension.
[0080] In this embodiment of the present invention, the process of extracting temporal and spatial features is described in detail, using a server processing a road monitoring image sequence (5 minutes, 7500 frames, 25 fps, containing dynamic information such as shallow flooding and vehicle deceleration) from 2:00 PM to 2:05 PM on a weekday on a certain highway section. This image sequence covers a six-lane, two-way, plain road. For example, the following spatial information is visible: shallow flooding in the right lane (occupying one-fifth of the lane area, with clear boundaries), 20-30 vehicles moving (mostly concentrated in the left lane, with fewer vehicles on the right lane), intact lane markings, and intact guardrails. Temporal information also shows that vehicles decelerate when passing through the flooded area (from 80 km / h to below 60 km / h), the flooded area slowly expands over time (increasing by approximately 0.5 square meters every 10 seconds), and light intensity decreases slightly due to cloud cover (the average brightness per frame decreases from 180 to 160, normalized). The server's goal is to extract spatial features (such as the location of flooded areas and the spatial distribution of vehicles) from each image frame while preserving their temporal continuity (i.e., the change in spatial information across different time windows). Specific implementation: Single-frame image feature extraction: The server invokes a pretrained ResNet-50 model (removing the last fully connected layer, retaining the convolutional layer features) to perform feature extraction on each frame of the image sequence (7,500 frames total), generating a 2048-dimensional spatial feature vector for each frame. This vector encodes the spatial information in the image—for example, the spatial distribution of flooded areas corresponds to high-value areas in the vector, and the spatial location of vehicles corresponds to a specific dimension in the vector. Time window mean aggregation: To reduce data volume and preserve the temporal trend of spatial information, the server divides the 7,500 image frames into 300 time windows (each window is 25 frames, corresponding to 1 second). The element-wise mean (i.e., the mean value of each dimension) of the 25 spatial feature vectors within each time window is taken, resulting in a 300×2048 spatial dimension monitoring image feature sequence. For example, in the spatial feature mean vector for a certain time window (2:00:00 PM - 2:00:01 PM), the value of the "Waterlogged Area Location" dimension is 0.7 (after normalization, indicating that the waterlogged area is located in the right lane), and the value of the "Vehicle Spatial Distribution" dimension is 0.6 (indicating that vehicles are mostly concentrated in the left lane). The server's goal is to extract dynamic temporal features (such as vehicle speed changes and waterlogged area expansion rate) from image sequences and capture their temporal dependencies (i.e., how features change over time) using a time series model.Specific implementation: Extracting temporal features: The server extracts the following temporal features from the image sequence (collected every 1 second, corresponding to 300 time windows): Vehicle speed: Calculated by the displacement of the vehicle bounding box between adjacent frames (unit: km / h). For example, the vehicle speed in the right lane is 75 km / h at 2:00:00 PM and drops to 70 km / h at 2:00:01 PM. Waterlogged area: The number of pixels in each frame of the waterlogged area is extracted using a pre-trained U-Net semantic segmentation model and converted to actual area (unit: square meters). For example, the waterlogged area is 12 square meters at 2:00:00 PM and increases to 12.5 square meters at 2:00:10 PM. Light intensity: Calculate the mean brightness of each frame (normalized to the range 0-1). For example, it is 0.85 at 2:00:00 PM and drops to 0.82 at 2:00:05 PM. These temporal features form a 300×3 raw time series data matrix (300 time windows, each with 3D features). Time Series Feature Extraction: The server invokes a bidirectional LSTM model (hidden layer dimension 1024, two stacked layers) to process the raw time series data matrix. The bidirectional LSTM model can capture temporal dependencies from both the past to the present (forward) and the present to the past (reverse). For example, a decrease in vehicle speed is not only related to the current flooded area, but also to the expansion of the flooded area in the previous few seconds. After processing, a 300×1024 time-dimensional time series data feature sequence is generated. This sequence encodes the dynamic trends of temporal features—for example, the feature value of the "vehicle speed" dimension gradually decreases over time, the feature value of the "flooded area" dimension gradually increases over time, and the feature value of the "light intensity" dimension slightly decreases over time. The spatial feature sequence (300×2048) of monitoring images and the temporal feature sequence (300×1024) of time series data extracted by the server serve as the foundation for subsequent spatiotemporal coupling feature fusion. The spatial feature sequence preserves the second-by-second mean of spatial information (such as the location of flooded areas and the spatial distribution of vehicles); the temporal feature sequence captures the dynamic trends of each second (such as vehicle speed decreases and the expansion of flooded areas). Subsequently, the server fuses these two feature sequences using a spatiotemporal attention mechanism to generate the first road-domain spatiotemporal coupling feature (2048×300). This feature captures the spatiotemporal synergy between "expansion of flooded areas" and "vehicle deceleration" (e.g., increasing flooded areas → decreasing vehicle speeds), providing critical spatiotemporal information support for subsequent risk warnings. The spatiotemporal feature extraction of the road-domain monitoring image sequence to be evaluated essentially separates and structures the spatial and dynamic temporal information within the image. Spatial feature extraction preserves the temporal continuity of spatial information through single-frame image processing and time window aggregation. Temporal feature extraction captures the dynamic trends of temporal features through dynamic feature parsing and temporal model processing. This processing method ensures that the server can fully obtain the spatiotemporal information of the scenario to be evaluated, laying the foundation for accurate prediction of subsequent risk warning models.
[0081] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned intelligent early warning method for highway road area risks based on the spatiotemporal coupling model. Figure 2 As shown, Figure 2 This is a block diagram of the structure of a computer device 100 provided in an embodiment of the present invention. Computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or exchange, memory 111, processor 112, and communication unit 113 are electrically connected to each other, directly or indirectly. For example, these components can be electrically connected via one or more communication buses or signal lines.
[0082] In terms of engineering deployment, the embodiments of the present invention also involve the following key technologies:
[0083] Smart Sensing Pole Hardware Interface Standard: To ensure the efficient collection and transmission of heterogeneous monitoring data from multiple sources (such as video, radar, weather, and road conditions), this embodiment of the present invention defines a unified smart sensing pole hardware interface standard. This standard includes: Sensor Power Supply Standard: It provides DC12V / 24V and PoE (Power over Ethernet) power supply interfaces to meet the power supply requirements of different sensor types and includes overload protection.
[0084] Data Interface Design: Integrated Gigabit Ethernet, RS485 / RS232 serial ports, and LoRa / NB-IoT wireless communication modules support multiple wired and wireless data transmission methods. The data interface complies with the OpenAPI specification, facilitating access and data analysis for sensors from different manufacturers.
[0085] Mechanical mounting interface: Standardizes sensor mounting bracket dimensions, screw hole locations, and load-bearing standards to facilitate rapid installation and maintenance of various sensors (such as cameras, radars, and weather stations).
[0086] The task allocation logic of the "perception-edge-cloud" three-level computing architecture: The perception layer is primarily responsible for collecting raw data and performing preliminary preprocessing (such as filtering, format conversion, and simple feature extraction). The processed data is then uploaded to the edge layer via standard interfaces. This layer focuses on low power consumption, real-time performance, and data compression.
[0087] The edge layer (e.g., the server at the road section monitoring center): This layer handles tasks such as localized data fusion, real-time risk detection, rapid warning issuance, and coordinated control with local equipment. For example, it integrates video images and radar data from the same area to identify abnormal vehicle behavior; makes local decisions based on warning information, and rapidly controls variable information boards. This layer prioritizes low latency, localized response, and bandwidth conservation.
[0088] The cloud layer (e.g., the highway management center cloud platform) is responsible for global data storage and management, complex model training and updating (such as the risk warning model and regular optimization of learnable parameters in the embodiments of this invention), network-wide risk assessment, cross-regional coordinated scheduling, and long-term trend analysis and decision support. This layer focuses on large-scale data processing, complex computation, and global optimization.
[0089] Communication Protocol for Multi-Source Data Fusion (V2X Message Format Definition): Based on the V2X (Vehicle to Everything) technical framework, this embodiment of the present invention defines a dedicated message format and communication protocol for communication between roadside units (RSUs), onboard units (OBUs), intelligent sensing devices, and cloud platforms. This protocol includes messages such as Risk Warning, Road Condition Report, and Cooperative Control Command, ensuring efficient, secure, and reliable exchange of risk-related information between different entities. For example, a Risk Warning message includes key fields such as Warning ID, road section code, risk type, risk level, location coordinates, impact range, and estimated duration.
[0090] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the present disclosure and to utilize various embodiments with various modifications as appropriate for the specific application contemplated.< / start> < / end> < / start> < / end> < / start> < / start> < / start> < / start> < / start> < / start> < / start> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / end> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / start> < / end> < / end> < / start> < / end> < / start> < / end> < / start> < / end> < / start> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / start> < / start> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / start> < / end> < / start> < / end> < / start> < / end> < / start>
Claims
1. An intelligent early warning method for highway road risk based on a spatiotemporal coupling model, characterized by: include: Acquiring road domain monitoring data, extracting time dimension features and space dimension features of the road domain monitoring data, and fusing the time dimension features and the space dimension features to obtain a first road domain spatiotemporal coupling feature of the road domain monitoring data; Acquire a risk warning rule, extract a first rule embedding feature of the risk warning rule, and perform heterogeneous feature coupling processing on the first rule embedding feature and the first road domain spatiotemporal coupling feature, wherein the risk warning rule is used in a risk warning model to perform intelligent risk warning on the road domain monitoring data; Optimizing the heterogeneous feature coupling processing result based on a first learnable parameter to obtain a target rule embedding feature, wherein the first learnable parameter is determined by training after fixing the parameters of the risk warning model; The risk warning model is called to perform risk prediction based on the target rule embedding feature to obtain a risk warning result of the road domain monitoring data.
2. The method according to claim 1, characterized in that The step of optimizing the heterogeneous feature coupling processing result based on the first learnable parameter to obtain the target rule embedding feature includes: Optimizing a heterogeneous feature coupling processing result based on the first learnable parameter, and performing feature superposition on the optimized heterogeneous feature coupling processing result and the first rule embedding feature to obtain a second rule embedding feature; The second rule embedding feature is subjected to feature conversion, and the feature conversion result is optimized based on a second learnable parameter to obtain a target rule embedding feature, wherein the second learnable parameter is determined by training after fixing the parameters of the risk warning model.
3. The method according to claim 2, characterized in that The risk warning model is provided with a plurality of feature extraction components operating in series. The risk warning model is called to perform risk prediction based on the target rule embedded features to obtain the risk warning result of the road monitoring data, including: Extracting features from the target rule embedding features based on the plurality of feature extraction components, performing risk prediction based on the extracted feature representation, and obtaining a risk warning result of the road area monitoring data; Wherein, the target rule embedding feature is the input of the first feature extraction component; For any remaining feature extraction component except the feature extraction component at the last position, heterogeneous feature coupling processing is performed on the feature extraction result of the current feature extraction component and the first road domain spatiotemporal coupling feature, the heterogeneous feature coupling processing result is optimized based on the first learnable parameter, feature superposition is performed on the optimized heterogeneous feature coupling processing result and the feature extraction result of the current feature extraction component, feature conversion is performed on the features obtained after feature superposition, and after optimizing the feature conversion result based on the second learnable parameter, it is loaded into the subsequent feature extraction component for feature extraction.
4. The method according to claim 3, characterized in that Each of the feature extraction components is respectively configured with a corresponding first learnable parameter and a corresponding second learnable parameter. Before optimizing the heterogeneous feature coupling processing result based on the first learnable parameter, the method further includes: Obtaining a road monitoring data instance, fixing the parameters of the risk warning model, and initially assigning zero values to each of the first learnable parameters and each of the second learnable parameters; Each of the first learnable parameters and each of the second learnable parameters are trained based on the road domain monitoring data instance.
5. The method according to claim 1, wherein The acquiring of the risk warning rule, extracting the first rule embedding feature of the risk warning rule, and performing heterogeneous feature coupling processing on the first rule embedding feature and the first road domain spatiotemporal coupling feature include: Extracting a second road-domain spatiotemporal coupling feature similar to the first road-domain spatiotemporal coupling feature, and a baseline road-domain risk state associated with the second road-domain spatiotemporal coupling feature, wherein the baseline road-domain risk state is a risk state of baseline road-domain monitoring data, the second road-domain spatiotemporal coupling feature is obtained by fusing time dimension features and spatial dimension features of the baseline road-domain monitoring data, and the spatiotemporal dimensions of the second road-domain spatiotemporal coupling feature are the same as the spatiotemporal dimensions of the first road-domain spatiotemporal coupling feature; Adding a start symbol and a stop symbol at the beginning and end of the reference road risk state to obtain a rule element text corresponding to the reference road risk state; Adding the start symbol after the last rule element text to obtain a risk warning rule, and extracting the first rule embedding feature of the risk warning rule; The second road-domain spatiotemporal coupling feature is superimposed on the first road-domain spatiotemporal coupling feature to obtain an integrated road-domain spatiotemporal coupling feature, and heterogeneous feature coupling processing is performed on the first feature component of the rule element text in the first rule embedding feature and the second road-domain spatiotemporal coupling feature corresponding to the first feature component in the integrated road-domain spatiotemporal coupling feature; Heterogeneous feature coupling processing is performed on the second feature component of the start symbol located at the last position in the first rule embedding feature and the first road-domain spatiotemporal coupling feature in the integrated road-domain spatiotemporal coupling feature.
6. The method according to claim 5, characterized in that The performing heterogeneous feature coupling processing on the first feature component of the rule element text in the first rule embedding feature and the second road domain spatiotemporal coupling feature corresponding to the first feature component in the integrated road domain spatiotemporal coupling feature includes: Determine a rule feature mapping vector based on the first rule embedding feature, determine a spatiotemporal feature mapping vector based on the integrated road domain spatiotemporal coupling feature, and obtain an association strength vector based on a multiplication result of the rule feature mapping vector and the spatiotemporal feature mapping vector after feature space transformation; Performing feature masking on the association strength vector based on a feature masking vector, and normalizing the masked association strength vector to obtain a feature coupling weight vector, wherein the feature masking vector is used to dynamically suppress remaining vector feature values in the association strength vector except for a target feature value, and the target feature value is determined based on a product operation result of a first feature component of the rule element text in the first rule embedding feature and a second road-domain spatiotemporal coupling feature corresponding to the first feature component in the integrated road-domain spatiotemporal coupling feature; A spatiotemporal feature representation vector is determined according to the second road domain spatiotemporal coupling feature, and a heterogeneous feature coupling processing result is obtained according to a multiplication result of the feature coupling weight vector and the spatiotemporal feature representation vector.
7. The method according to claim 5, characterized in that The extracting of a second road domain spatiotemporal coupling feature similar to the first road domain spatiotemporal coupling feature, and a reference road domain risk state associated with the second road domain spatiotemporal coupling feature, includes: Acquire a plurality of the reference road domain monitoring data and a reference road domain risk status of each of the reference road domain monitoring data, fuse the time dimension feature and the space dimension feature of each of the reference road domain monitoring data, and obtain the second road domain spatiotemporal coupling feature of each of the reference road domain monitoring data; Each second road domain spatiotemporal coupling feature is associated with the corresponding baseline road domain risk state and stored in a risk feature library, and the second road domain spatiotemporal coupling feature similar to the first road domain spatiotemporal coupling feature and the baseline road domain risk state associated with the second road domain spatiotemporal coupling feature are extracted from the risk feature library.
8. The method according to claim 7, characterized in that The reference road monitoring data is a reference road monitoring image sequence, and the associating each of the second road spatiotemporal coupling features with the corresponding reference road risk state and storing them in a risk feature library includes: For each of the reference road monitoring image sequences, performing mean aggregation on the monitoring image features of the reference road monitoring image sequence in the spatial dimension to obtain a risk value sequence feature of the same dimension as the time series data feature of the reference road monitoring image sequence in the temporal dimension; Performing mean aggregation on the risk value sequence features and the time series data features of the reference road monitoring image sequence in the time dimension to obtain an index feature of each reference road monitoring image sequence; Each of the index features, the second road domain spatiotemporal coupling features and the corresponding reference road domain risk state are associated and stored in a risk feature library.
9. The method according to claim 1, characterized in that The road monitoring data is a road monitoring image sequence to be evaluated, and extracting the time dimension features and space dimension features of the road monitoring data includes: Performing feature extraction on each road area monitoring image in the road area monitoring image sequence to be evaluated to obtain monitoring image features of the road area monitoring image sequence to be evaluated in a spatial dimension; Extracting time feature information from the road area monitoring image sequence to be evaluated, performing feature extraction on the time feature information, and obtaining time series data features of the road area monitoring image sequence to be evaluated in a time dimension.
10. A server system, characterized in that: The method comprises a server, wherein the server is configured to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Urban road traffic risk prediction method and device, electronic equipment and storage medium
CN119359012A
Dangerous behavior identification and early warning method based on multi-modal analysis
CN119360278A
Intelligent terminal environment monitoring method based on combination of multi-source data fusion and deep learning
CN120180046A
Smart city multi-modal data acquisition and fusion method and system
CN120197130A
Sensing intelligent driving complex traffic scene dynamic risk prediction method
CN120220390A