Edge calculation signal lamp control method and system based on traffic participant behavior analysis

Through the edge computing signal light control method, the spatial and temporal feature encoding network is used to analyze the behavior of traffic participants and adjust the signal light phase switching timing, solving the response lag and cloud computing delay problems of traditional signal light control methods, and achieving efficient and real-time traffic flow optimization and safety control.

CN120526475AActive Publication Date: 2025-08-22HEBEI JOY SMART TECH CO LTD

Patent Information

Application Number
CN202510608338.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-22
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

Traditional signal light control methods are difficult to cope with dynamic changes in the behavior of complex traffic participants, resulting in response lag and misjudgment, and the adaptive optimization of dynamic traffic flow cannot be achieved. In addition, there are delays and resource bottlenecks in the cloud computing architecture, which cannot meet the real-time signal control needs.

Method used

Through the edge computing signal light control method, the video stream data is analyzed using the spatiotemporal feature encoding network, the behavioral state characteristics of traffic participants are extracted, the behavioral trigger marks are generated, and the phase switching timing is adjusted based on the signal light phase state, so as to realize millisecond localization decisions, combining reverse phase conflict pre-check and dynamic clearing time calculation, the traffic efficiency of intersections is optimized.

Benefits of technology

It significantly improves the control accuracy and real-time response in complex traffic scenarios, reduces the risk of traffic accidents, optimizes the overall traffic efficiency of intersections, avoids phase switching lag or redundancy problems, and realizes highly robust intelligent signal control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526475A_ABST
    Figure CN120526475A_ABST
Patent Text Reader

Abstract

The invention provides a traffic participant behavior analysis-based edge calculation signal lamp control method and system, and the method comprises the steps: obtaining video stream data of a target intersection, carrying out the frame-by-frame behavior state analysis of the video stream data through a spatial-temporal feature coding network, extracting the behavior state features of each traffic participant, and carrying out the analysis of the behavior state features of each traffic participant; and inputting the behavior state characteristics into a preset behavior matching model, generating a behavior triggering identifier of the traffic participant in the current time window, generating a signal lamp control instruction of an edge computing node according to the time sequence relevance between the behavior triggering identifier and the phase state of the current signal lamp, and sending the signal lamp control instruction to the edge computing node. And finally, a signal lamp phase switching time sequence of the target intersection is adjusted based on the signal lamp control instruction, so that the traffic participants meeting the traffic rule conflict condition obtain the traffic priority under the signal lamp phase switching time sequence. According to the method, the traffic accident risk is reduced, and meanwhile, the overall traffic efficiency of the intersection is optimized by dynamically adjusting the minimum response period and the phase duration parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to an edge computing traffic light control method and system based on traffic participant behavior analysis. Background Art

[0002] With the rapid growth of urban traffic volume, intelligent traffic signal control technology has become a key means to improve road traffic efficiency and safety. Traditional traffic light control methods mainly rely on a phase switching mechanism triggered by a fixed timing scheme, and adjust the signal timing through a preset cycle or a simple vehicle queue length detection. However, such methods are difficult to cope with complex scenarios with dynamic changes in the behavior of traffic participants. The fixed timing mechanism cannot perceive sudden behaviors such as pedestrians crossing illegally and vehicles changing lanes abnormally, resulting in delayed risk response, difficulty in distinguishing between normal traffic and potential conflict behaviors, and prone to misjudgment. Although the multi-source data fusion solution can improve the judgment accuracy, it is limited by the centralized cloud computing architecture, with data transmission delays and computing resource bottlenecks, and cannot meet the millisecond-level response requirements of real-time signal control. The above defects make it difficult for existing technologies to achieve adaptive optimization of dynamic traffic flow while ensuring traffic safety. Summary of the Invention

[0003] The present invention provides an edge computing signal light control method and system based on traffic participant behavior analysis.

[0004] In the first aspect, an embodiment of the present invention provides an edge computing traffic light control method based on traffic participant behavior analysis, the method comprising: obtaining video stream data of a target intersection, the video stream data comprising the motion trajectory of at least one traffic participant and the phase state of the current traffic light; performing frame-by-frame behavior state analysis on the video stream data through a spatiotemporal feature coding network, and extracting the behavior state features of each traffic participant, wherein the behavior state features comprise a motion direction vector, a speed change sequence, and a spatial position offset; inputting the behavior state features into a preset behavior matching model, and generating a behavior trigger identifier of the traffic participant within the current time window, the behavior trigger identifier being used to indicate whether the traffic participant meets a preset traffic rule conflict condition; generating a traffic light control instruction for an edge computing node based on the temporal correlation between the behavior trigger identifier and the phase state of the current traffic light; and adjusting the traffic light phase switching timing of the target intersection based on the traffic light control instruction, so that traffic participants who meet the traffic rule conflict condition obtain passage priority under the traffic light phase switching timing.

[0005] In a second aspect, an embodiment of the present invention provides a traffic light control system, comprising: a memory storing a computer program; and a processor for loading the computer program to implement the above-mentioned edge computing traffic light control method based on traffic participant behavior analysis.

[0006] The edge computing traffic light control method based on traffic participant behavior analysis provided by the present invention significantly improves the control accuracy and real-time response in complex traffic scenarios by integrating multi-dimensional dynamic behavior characteristics with the temporal correlation modeling of the signal light phase state. Compared with the traditional fixed timing control method, this method utilizes the collaborative analysis of the motion direction vector, speed change sequence and spatial position offset to dynamically capture the abnormal behavior trend of traffic participants, and conducts conflict priority assessment in combination with the remaining time of the signal light, effectively avoiding the phase switching lag or redundancy problem caused by feature misjudgment. Through the end-to-end closed-loop processing mechanism of the edge computing node, millisecond-level localized decision-making from behavior feature extraction to control instruction generation is achieved, breaking through the delay bottleneck caused by cloud computing dependence. At the same time, the risk-driven phase switching strategy, under the premise of ensuring priority passage in the conflicting direction, takes into account the traffic safety of oncoming vehicles through reverse phase conflict pre-inspection and dynamic clearing time calculation, forming an adaptive control effect that unifies local optimization and global coordination. While reducing the risk of traffic accidents, this method optimizes the overall traffic efficiency of the intersection by dynamically adjusting the minimum response period and phase duration parameters, and achieves highly robust intelligent signal control in scenarios without relying on multi-source sensor equipment fusion and manual rule presets. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 This is a flowchart of an edge computing traffic light control method based on traffic participant behavior analysis provided by an embodiment of the present invention.

[0008] Figure 2 It is a schematic diagram of the composition of a traffic light control system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0009] See also Figure 1 , Figure 1 A flowchart of an edge computing signal light control method based on traffic participant behavior analysis provided in an embodiment of the present invention. The edge computing signal light control method based on traffic participant behavior analysis can be executed by a signal light control system. The edge computing signal light control method based on traffic participant behavior analysis may include the following steps:

[0010] Step S100: Acquire video stream data of a target intersection, where the video stream data includes a motion trajectory of at least one traffic participant and a phase state of a current traffic light.

[0011] In an embodiment of the present invention, the target intersection is a specific traffic intersection that requires signal light control, such as crossroads, T-junctions, and other traffic nodes with signal light control in cities. Video stream data is continuous image sequence data continuously collected by camera equipment installed around the target intersection. Traffic participants include pedestrians and vehicles (such as cars, motorcycles, bicycles, etc.). The motion trajectory refers to their movement path in the intersection space over a period of time, which can be obtained by tracking and recording the positions of traffic participants in the video stream data frame by frame. The phase state of the current traffic light indicates the display state of the traffic light at a certain moment, such as red light, green light, yellow light, and the remaining time for each light to be on.

[0012] Step S200: Perform frame-by-frame behavior state analysis on the video stream data through a spatiotemporal feature coding network to extract the behavior state features of each traffic participant, where the behavior state features include motion direction vector, speed change sequence and spatial position offset.

[0013] The spatiotemporal feature coding network is a deep learning network used to process data with temporal and spatial dimensions. It can extract meaningful spatiotemporal features from video stream data. Frame-by-frame behavioral state analysis refers to analyzing each frame of video stream data to determine the behavioral state of traffic participants in that frame. Behavioral state features are key information for describing the behavior of traffic participants. The motion direction vector represents the direction of motion of a traffic participant at a certain moment and can be obtained by analyzing its motion trajectory. The speed change sequence records the changes in the speed of a traffic participant over a period of time, reflecting the dynamic characteristics of its motion. The spatial position offset represents the change in the position of a traffic participant between adjacent frames.

[0014] In practical applications, a spatiotemporal feature encoding network can employ an architecture that combines a convolutional neural network (CNN) and a recurrent neural network (RNN). First, the CNN extracts spatial features from video frames, while the RNN processes time series information, thereby encoding the spatiotemporal features of the video stream data. For example, for a car traveling at an intersection, frame-by-frame analysis of the video stream using the spatiotemporal feature encoding network can reveal the car's motion direction vector, velocity change sequence, and spatial position offset for each frame. Specifically, the motion direction vector can be derived by calculating the car's position change between adjacent frames; the velocity change sequence can be calculated by differentially calculating the car's position change in consecutive frames; and the spatial position offset can be determined by comparing the car's specific coordinate positions in adjacent frames. These behavioral state features can accurately describe the behavior patterns of traffic participants, providing a basis for subsequent traffic rule violation detection and signal control.

[0015] As an embodiment, step S200, performing frame-by-frame behavior state analysis on the video stream data through a spatiotemporal feature coding network to extract the behavior state features of each traffic participant, may include the following steps S210-S250:

[0016] Step S210: constructing a three-dimensional space-time tensor of the video stream data, where the three-dimensional space-time tensor includes a time dimension, a spatial position dimension, and a pixel feature dimension.

[0017] A three-dimensional spatiotemporal tensor is a multidimensional data structure used to represent video stream data. It integrates temporal, spatial, and pixel feature information. The temporal dimension represents the sequence of video stream frames, reflecting the changes in traffic scenes over time. The spatial position dimension represents the spatial coordinates of each pixel in the video frame, identifying the position of traffic participants in the image. The pixel feature dimension represents the characteristic information of each pixel, such as color and brightness.

[0018] A specific example of constructing a three-dimensional spatiotemporal tensor can be implemented as follows: first, the video stream data is divided into frames to obtain a series of video frames. Then, for each frame of video, its pixel information is extracted and encoded to form a two-dimensional pixel feature matrix. Next, these two-dimensional pixel feature matrices are stacked in chronological order to form a three-dimensional spatiotemporal tensor. For example, for a video stream data containing 100 frames, the resolution of each frame of video is 640×480 pixels, and each pixel has 3 color channels (RGB), then the dimension of the constructed three-dimensional spatiotemporal tensor is 100 (time dimension) × 640 (spatial position dimension-width) × 480 (spatial position dimension-height) × 3 (pixel feature dimension). By constructing a three-dimensional spatiotemporal tensor, the video stream data can be converted into a format suitable for spatiotemporal feature coding network processing, which facilitates the extraction of spatiotemporal features therein.

[0019] Step S220: extracting spatiotemporal features from the three-dimensional spatiotemporal tensor through a convolutional long short-term memory network to obtain a motion feature map of each traffic participant.

[0020] The convolutional long short-term memory network (ConvLSTM) is a deep learning model that combines the characteristics of convolutional neural networks and long short-term memory networks (LSTM). It can effectively process data with spatiotemporal structures. In an embodiment of the present invention, ConvLSTM is used to extract spatiotemporal features from a three-dimensional spatiotemporal tensor to capture the motion information of traffic participants in time and space. The motion feature map is a two-dimensional feature map that reflects the motion patterns and features of traffic participants in a video frame. For example, for a car turning at an intersection, the corresponding three-dimensional spatiotemporal tensor is processed by ConvLSTM to obtain a motion feature map of the car. Different channels in the motion feature map may represent the activation intensity of the car's motion modes such as straight-line driving, turning, acceleration, and deceleration. By analyzing the values ​​of these channels, the motion state of the car can be accurately identified.

[0021] Specifically, step S220 may include the following steps S221 to S229:

[0022] Step S221: Divide the three-dimensional spatiotemporal tensor into continuous time step slices according to the time dimension, each time step slice contains the spatial position dimension and pixel feature dimension of the current frame.

[0023] A time-step slice is a two-dimensional data block obtained by segmenting a three-dimensional spatiotemporal tensor along the time dimension. Each time-step slice corresponds to a frame in the video stream. By segmenting the three-dimensional spatiotemporal tensor into time-step slices, processing of spatiotemporal data can be transformed into processing of a series of two-dimensional data, facilitating subsequent convolution operations and feature extraction.

[0024] In practice, the 3D spatiotemporal tensor can be segmented at fixed time intervals. For example, a 3D spatiotemporal tensor containing 100 frames can be segmented frame by frame, resulting in 100 time-step slices. Each time-step slice has a spatial position dimension (e.g., 640×480) and a pixel feature dimension (e.g., 3 color channels). In this way, the 3D spatiotemporal tensor can be converted into a series of 2D data blocks suitable for convolution operations, laying the foundation for subsequent spatiotemporal feature extraction.

[0025] Step S222: Perform a three-dimensional convolution operation on each time step slice to generate a spatiotemporal convolution feature map of the current time step. The convolution kernel of the three-dimensional convolution operation spans adjacent frames in the time dimension to capture motion continuity.

[0026] A 3D convolution operation is performed in three-dimensional space. Its convolution kernel not only convolves the input data in the spatial dimension but also spans adjacent frames in the temporal dimension, thereby capturing temporal and spatial variations in the data. The spatiotemporal convolution feature map is the output of the 3D convolution operation and reflects the spatiotemporal characteristics of the current time step.

[0027] When performing a three-dimensional convolution operation, the convolution kernel slides in the time and space dimensions, performing convolution calculations on each time step slice and part of its adjacent frames. For example, assuming the size of the convolution kernel is 3×3×3 (time dimension × space width × space height), it will span 3 adjacent frames in the time dimension and perform convolution calculations on a 3x3 area in the space dimension. In this way, the motion continuity of traffic participants between adjacent frames can be captured. For each time step slice, after the three-dimensional convolution operation, a spatiotemporal convolution feature map is generated. The dimension of this feature map depends on the number of convolution kernels and the dimension of the input data. By generating a spatiotemporal convolution feature map, the spatiotemporal features in the video stream data can be extracted, providing a basis for subsequent feature processing and analysis.

[0028] Step S223: Input the spatiotemporal convolution feature map into the forget gate structure of the convolutional long short-term memory network, and generate the forget gate output through element-by-element multiplication and activation function processing to control the retention ratio of the previous hidden state.

[0029] The forget gate structure is a key component of the Convolutional Long Short-Term Memory (ConvLSTM) network. It controls the retention ratio of the previous hidden state, thereby achieving selective forgetting of historical information. The spatiotemporal convolutional feature map is the output of the previous three-dimensional convolution operation and contains the spatiotemporal feature information of the current time step. In the forget gate structure, the spatiotemporal convolutional feature map is first concatenated with the previous hidden state and then processed through a convolutional layer and a sigmoid activation function. The output value of the sigmoid activation function ranges from 0 to 1, representing the retention ratio of each element. By element-by-element multiplication of the previous hidden state with the forget gate output, selective retention of the previous hidden state is achieved. For example, if the forget gate output value of a certain element is 0.8, it means that the corresponding previous hidden state information will be retained at a ratio of 80%. In this way, the forget gate structure can dynamically adjust the retention ratio of the previous hidden state based on the current input information, preventing the network from overfitting or forgetting important historical information, thereby improving the network's ability to process spatiotemporal sequence data.

[0030] Step S224: The spatiotemporal convolution feature map is concatenated with the forget gate output and then input into the input gate structure. The candidate cell state of the current time step is generated through convolution operation and activation function processing.

[0031] The input gate structure is also a key part of the ConvLSTM. It is used to determine how much of the current input information should be added to the cell state. In this embodiment of the present invention, the spatiotemporal convolution feature map is concatenated with the forget gate output to obtain a feature vector that contains the current input information and historical information.

[0032] The specific operation is as follows: First, the spatiotemporal convolution feature map and the forget gate output are spliced ​​in the channel dimension to form a new feature matrix. Then, this feature matrix is ​​input into the convolution layer of the input gate structure to extract features through the convolution operation. Next, the convolution output is processed by the Sigmoid activation function and the hyperbolic tangent activation function respectively. The output of the Sigmoid activation function is used to control whether the input information is allowed to enter the cell state, and the output of the hyperbolic tangent activation function is a preliminary estimate of the candidate cell state. Finally, these two outputs are multiplied element by element to obtain the candidate cell state for the current time step. For example, for a vehicle traveling at an intersection, through the input gate structure, it can dynamically decide whether to add the vehicle's new motion information to the cell state based on the current spatiotemporal features and historical information, thereby more accurately tracking the vehicle's trajectory.

[0033] Step S225: Perform weighted superposition of the candidate cell state and the previous cell state adjusted by the forget gate to generate an updated current cell state.

[0034] The cell state is a data structure used to store and transmit historical information in ConvLSTM. In this embodiment of the present invention, the candidate cell state generated in the previous step is weightedly superimposed with the previous cell state adjusted by the forget gate to update the current cell state.

[0035] The specific process is as follows: First, the previous cell state is processed by the forget gate, retaining some important historical information. Then, the candidate cell state is added element-by-element to the previous cell state adjusted by the forget gate to obtain the updated current cell state. This weighted superposition method combines historical information with current input information, allowing the cell state to continuously update and evolve, thereby better adapting to the changes in the movement of traffic participants. For example, for a vehicle that makes multiple turns at an intersection, this method can combine historical information about previous turns with current motion information to accurately reflect the vehicle's overall motion state.

[0036] Step S226: The current cell state is input into the output gate structure, processed by convolution operation and hyperbolic tangent activation function to generate the hidden state of the current time step.

[0037] The output gate structure is used to determine how much information in the current cell state should be output as the hidden state. The hidden state is the output of ConvLSTM, which contains important information of the current time step and can be used for subsequent feature fusion and analysis.

[0038] The specific operation is as follows: the current cell state is input into the convolution layer of the output gate structure, and features are extracted through the convolution operation. Then, the convolution output is processed through the Sigmoid activation function to obtain a control vector, each element of which represents the output ratio of the corresponding element in the cell state. At the same time, the current cell state is processed through the hyperbolic tangent activation function, mapping its value to the range of -1 to 1. Finally, these two outputs are multiplied element by element to obtain the hidden state of the current time step. For example, if a vehicle suddenly accelerates at an intersection, the output gate structure can accurately output the vehicle's current motion state information based on the information in the cell state, providing a basis for subsequent motion feature analysis.

[0039] Step S227: Traverse all time step slices, perform cross-layer feature fusion on the hidden state of each time step along the channel dimension, and generate a multi-scale motion feature map.

[0040] Cross-layer feature fusion is a method that integrates feature information from different layers. It can fully utilize the hidden state information at different time steps to extract richer and more comprehensive motion features. The multi-scale motion feature map is the result of cross-layer feature fusion, which contains motion feature information at different scales and layers.

[0041] In its implementation, all time-step slices are traversed, and the hidden states of each time step are concatenated along the channel dimension. These concatenated features are then fused through a convolutional layer to generate a multi-scale motion feature map. For example, for a video stream consisting of 100 time-step slices, the hidden states of each time step are concatenated along the channel dimension to produce a larger-dimensional feature matrix. This feature matrix is ​​then reduced in dimensionality and fused through convolution operations to generate a feature map containing motion features at different scales. Cross-layer feature fusion effectively integrates motion information from different time steps, improving the expressiveness of motion features and providing richer feature information for subsequent local motion pattern feature extraction.

[0042] Step S228: performing spatial pyramid pooling on the multi-scale motion feature map to extract local motion pattern features in different field of view ranges.

[0043] Spatial pyramid pooling (SPP) is a method for extracting features of different scales. It can extract local features of different visual ranges without changing the size of the input feature map. In an embodiment of the present invention, spatial pyramid pooling is performed on the multi-scale motion feature map to extract local motion pattern features of different visual ranges. The specific operation is as follows: the multi-scale motion feature map is input into the spatial pyramid pooling layer, which divides the feature map into regions of different sizes and then performs a pooling operation on each region. Possible pooling operations include maximum pooling and average pooling. For example, the feature map can be divided into regions of different sizes such as 1×1, 2×2, 4×4, and then a maximum pooling operation is performed on each region to obtain local feature vectors of different scales. Finally, these local feature vectors are spliced ​​to obtain a feature vector containing local motion pattern features of different visual ranges. Through spatial pyramid pooling, local motion pattern features of different visual ranges in the multi-scale motion feature map can be extracted. These features can more comprehensively describe the motion patterns of traffic participants and provide richer information for subsequent motion feature analysis and behavior recognition.

[0044] Step S229: Perform channel attention weighted fusion on the local motion pattern features and the hidden state of the last time step to generate a motion feature map of each traffic participant, where each channel of the motion feature map corresponds to the activation intensity of the preset motion pattern.

[0045] Channel-attention weighted fusion is a method for weighted combination of different features. It dynamically adjusts feature weights based on the importance of different channels. In this embodiment of the present invention, channel-attention weighted fusion is performed on the local motion pattern features and the hidden state of the last time step to generate motion feature maps for each traffic participant.

[0046] Specifically, for example, it can be implemented as follows: First, calculate the channel attention weights of the local motion pattern features and the last time step hidden state. The channel attention weights can be calculated through an attention mechanism, such as using a fully connected layer and a Sigmoid activation function to map the input features to a weight vector between 0 and 1. Then, the local motion pattern features and the last time step hidden state are respectively multiplied element-by-element by the corresponding channel attention weights to obtain weighted features. Finally, the weighted local motion pattern features and the last time step hidden state are added to generate a motion feature map for each traffic participant. Each channel of the motion feature map corresponds to the activation intensity of a preset motion pattern. By analyzing the values ​​of these channels, the motion pattern of the traffic participant can be accurately identified. For example, for a vehicle traveling at an intersection, a channel of the motion feature map may correspond to the vehicle's accelerating motion pattern. The higher the activation intensity of the channel, the greater the possibility of vehicle acceleration.

[0047] Step S230: Track key point trajectories on the motion feature map to generate motion direction vectors and spatial position offsets.

[0048] Keypoint trajectory tracking is a method used to track the positional changes of key points of traffic participants within video frames. By performing keypoint trajectory tracking on a motion feature map, the motion trajectory information of traffic participants can be obtained. The motion direction vector indicates the direction of movement of a traffic participant at a given moment, and the spatial position offset indicates the change in the position of the traffic participant between adjacent frames. For example, for a pedestrian walking at an intersection, keypoint trajectory tracking can be used to obtain the pedestrian's motion direction vector and spatial position offset. The motion direction vector indicates the pedestrian's walking direction, and the spatial position offset reflects the distance and direction of movement between adjacent frames.

[0049] Specifically, step S230 may include the following steps S231 to S236:

[0050] Step S231: extracting a set of candidate key points of each traffic participant from the motion feature map. The set of candidate key points is obtained by screening the response value threshold, wherein each candidate key point includes spatial coordinates and motion intensity features.

[0051] The candidate keypoint set is a set of keypoints with high response values ​​selected from the motion feature map. These keypoints can represent important characteristics and motion information of traffic participants. The response value threshold is a pre-set value used to select keypoints with response values ​​above the threshold. The spatial coordinates represent the position of the keypoint in the motion feature map, and the motion intensity feature represents the motion intensity corresponding to the keypoint.

[0052] The specific extraction can be implemented as follows: first, the motion feature map is traversed and the response value of each point is calculated. The response value can be obtained by calculating the numerical value of a certain channel or multiple channels of the motion feature map, for example, by using the maximum value, average value and other methods. Then, the points with response values ​​higher than the response value threshold are regarded as candidate key points, and their spatial coordinates and motion intensity characteristics are recorded. For example, for a motion feature map representing vehicle motion, by setting an appropriate response value threshold, key points such as the front and rear of the vehicle can be screened out. The spatial coordinates and motion intensity characteristics of these key points can be used for subsequent trajectory tracking and behavior analysis.

[0053] Step S232: perform bidirectional matching on the candidate key point sets of two adjacent frames, calculate the matching cost matrix based on the motion intensity feature similarity and the spatial coordinate offset distance, and retain the first K pairs of key points with the smallest matching cost in the cost matrix as the initial trajectory segments, K ≥ 1.

[0054] Bidirectional matching is a method used to match keypoints between two adjacent frames, improving matching accuracy and reliability. The matching cost matrix, calculated based on the similarity of motion intensity features and the spatial coordinate offset distance, measures the degree of matching between keypoints in two adjacent frames. The initial trajectory segments are obtained by retaining the first K keypoint pairs with the lowest matching cost in the matching cost matrix. These keypoint pairs form the preliminary motion trajectory of the traffic participant.

[0055] The calculation of the matching cost matrix can be implemented as follows: for the candidate key point sets of two adjacent frames, calculate the motion intensity feature similarity and spatial coordinate offset distance between each pair of key points. The motion intensity feature similarity can be obtained by calculating the cosine similarity between the motion intensity feature vectors of the two key points, and the spatial coordinate offset distance can be obtained by calculating the Euclidean distance between the spatial coordinates of the two key points. Then, the motion intensity feature similarity and the spatial coordinate offset distance are weightedly combined to obtain the matching cost between each pair of key points. The matching costs of all key point pairs are combined into a matrix, namely the matching cost matrix. Finally, the top K pairs of key points with the smallest matching cost are selected from the matching cost matrix and used as the initial trajectory segments. For example, for the key point set representing the vehicle in two adjacent frames, the corresponding key points of the vehicle in the adjacent frames can be accurately found through bidirectional matching and matching cost calculation, thereby constructing the initial motion trajectory of the vehicle.

[0056] Step S233: Input the initial trajectory segment into the trajectory continuity verification unit to detect whether there is a trajectory point sequence with continuously changing spatial positions within three consecutive frames, eliminate abnormal trajectory segments with jump amplitudes exceeding a preset threshold, and generate an optimized trajectory set.

[0057] The trajectory continuity verification unit is a module used to verify trajectory continuity. It detects whether a trajectory point sequence exhibits continuous spatial position changes within three consecutive frames. Abnormal trajectory segments are those whose jump amplitude exceeds a preset threshold. These segments may be mismatched due to noise, occlusion, and other factors and need to be eliminated. The optimized trajectory set is the trajectory set obtained after trajectory continuity verification and contains more accurate and reliable trajectory information.

[0058] The specific verification can be implemented as follows: input the initial trajectory segment into the trajectory continuity verification unit, and for each trajectory segment, check its spatial position change within three consecutive frames. Calculate the spatial coordinate offset distance of the trajectory point between two adjacent frames. If the offset distance change within three consecutive frames exceeds a preset threshold, it is considered that there is a jump in the trajectory segment, and it is eliminated as an abnormal trajectory segment. For example, for a trajectory segment representing pedestrian movement, if the position of the pedestrian suddenly jumps greatly within three consecutive frames, and the jump amplitude exceeds the preset threshold, the trajectory segment will be considered an abnormal trajectory segment and will be eliminated. Through trajectory continuity verification, the accuracy and reliability of trajectory tracking can be improved, and more accurate trajectory information can be provided for subsequent calculations of motion direction vectors and spatial position offsets.

[0059] Step S234: Calculate the direction angle of each trajectory in the optimized trajectory set, generate an initial direction vector based on the spatial coordinate difference between the starting point and the end point of the trajectory, and fit the direction change trend of N consecutive trajectory points through a sliding window to generate a smoothed motion direction vector.

[0060] Direction angle calculation is used to determine the direction angle of a trajectory. The initial direction vector is generated based on the spatial coordinate difference between the trajectory's starting and ending points, representing the approximate direction of the trajectory. Sliding window fitting is used to smooth changes in trajectory direction. By fitting the directional trends of N consecutive trajectory points, a smoother and more accurate motion direction vector can be generated.

[0061] The specific calculation can be implemented as follows: for each trajectory in the optimized trajectory set, the spatial coordinate difference between its starting point and end point is calculated, and the direction angle is calculated based on the difference to generate an initial direction vector. Then, a sliding window of fixed length is used to slide on the trajectory, and for each consecutive N trajectory points in the window, the direction change trend is fitted by the least squares method and other methods to obtain a smoothed direction vector. These smoothed direction vectors are combined in sequence to obtain a smoothed motion direction vector. For example, for a trajectory representing vehicle travel, the initial direction vector is obtained by calculating the coordinate difference between the starting point and the end point of the trajectory, and then the sliding window is used to fit the direction change trend of 10 consecutive trajectory points to generate a smoothed motion direction vector, which can more accurately reflect the vehicle's travel direction.

[0062] Step S235: Calculate the cumulative displacement in the X-axis and Y-axis directions based on the spatial coordinate difference between the starting point and the end point of the trajectory corresponding to the smoothed motion direction vector, and generate the instantaneous spatial position offset by combining the position difference of the trajectory point between the current frame and the previous frame.

[0063] Cumulative displacement refers to the change in position between the starting and ending points of a trajectory in the X and Y directions, reflecting the total distance a traffic participant has moved over a period of time. Instantaneous spatial offset refers to the difference in trajectory point position between the current frame and the previous frame, reflecting the instantaneous distance a traffic participant has moved between adjacent frames.

[0064] The specific calculation can be implemented as follows: for the trajectory corresponding to the smoothed motion direction vector, calculate the coordinate difference between its starting point and end point in the X-axis and Y-axis directions to obtain the cumulative displacement. At the same time, for the trajectory points of the current frame and the previous frame, calculate their coordinate difference in the X-axis and Y-axis directions to obtain the instantaneous spatial position offset. For example, for a vehicle traveling at an intersection, by calculating the coordinate difference between the starting point and end point of its trajectory, the cumulative displacement of the vehicle in the X-axis and Y-axis directions is obtained; by calculating the coordinate difference between the vehicle trajectory points in the current frame and the previous frame, the instantaneous spatial position offset of the vehicle is obtained. This displacement information can be used to analyze the vehicle's movement speed and movement trend.

[0065] Step S236: Perform weighted fusion on the accumulated displacement and the instantaneous spatial position offset to generate the final spatial position offset. The weight coefficient of the weighted fusion is dynamically adjusted according to the trajectory length and the stability of the motion direction vector. The motion direction vector and the final spatial position offset are grouped and stored according to the traffic participant identification to form a trajectory tracking result sequence aligned with the video stream data timestamp.

[0066] By weightedly fusing the cumulative displacement with the instantaneous spatial position offset, a more accurate and reliable final spatial position offset can be obtained. The weight coefficient can be pre-set and dynamically adjusted based on the trajectory length and the stability of the motion direction vector. The longer the trajectory length and the more stable the motion direction vector, the greater the weight of the cumulative displacement; conversely, the greater the weight of the instantaneous spatial position offset. The trajectory tracking result sequence is a sequence obtained by grouping the motion direction vector and the final spatial position offset by traffic participant identification and aligning it with the video stream data timestamp. It contains the motion trajectory information of each traffic participant.

[0067] For example, the weight coefficients of the cumulative displacement and the instantaneous spatial position offset are calculated based on the trajectory length and the stability of the motion direction vector. Then, the cumulative displacement and the instantaneous spatial position offset are multiplied by the corresponding weight coefficients respectively, and the results are added to obtain the final spatial position offset. Finally, the motion direction vector and the final spatial position offset are grouped and stored according to the traffic participant identification, and aligned with the timestamp of the video stream data to form a trajectory tracking result sequence. For example, for a traffic scene containing multiple vehicles and pedestrians, the motion trajectory information of each vehicle and pedestrian can be obtained through weighted fusion and group storage. This information can be used for subsequent traffic rule conflict detection and signal light control.

[0068] Step S240: performing differential calculation on the motion feature maps of consecutive frames through a sliding window to generate a speed change sequence.

[0069] Sliding windows are a method used to process time series data, analyzing and calculating data within continuous time windows. Differential calculation involves calculating the difference between adjacent data points. By performing differential calculations on motion feature maps of consecutive frames, we can obtain information about the speed changes of traffic participants within different time windows. A speed change sequence, consisting of speed change values ​​within each time window, reflects the speed trend of traffic participants.

[0070] For example, for a vehicle traveling at an intersection, by performing differential calculations on the motion feature maps of consecutive frames using a sliding window, we can obtain a speed variation sequence for the vehicle within different time windows. Larger values ​​in the speed variation sequence indicate more dramatic speed changes within that time window. By analyzing the speed variation sequence, we can accurately understand the vehicle's acceleration, deceleration, and other motion states.

[0071] Specifically, step S240 may include the following steps S241 to S247:

[0072] Step S241: a sliding window of fixed length is set in the time dimension of the motion feature map, and the sliding window covers the motion feature maps of consecutive T frames.

[0073] The sliding window is a fixed-length time window that slides along the time dimension of the motion feature map and is used to process T consecutive frames of motion feature maps. The purpose of setting a sliding window is to analyze the motion changes of traffic participants within a preset time range.

[0074] The specific settings can be implemented as follows: Determine the length T of the sliding window based on actual needs and analysis accuracy. For example, T can be set to 5 frames, meaning that the sliding window covers 5 consecutive frames of motion feature maps. Then, the sliding window starts from the first frame of the motion feature map and slides sequentially in chronological order until the motion feature maps of all frames are covered. By setting the sliding window, the continuous motion feature map can be divided into multiple fixed-length time windows, facilitating subsequent differential calculations and velocity change analysis.

[0075] Step S242: performing a channel-level difference operation on the motion feature maps of two adjacent frames in the sliding window to generate a channel difference feature map for each pair of adjacent frames in the window.

[0076] Channel-level differencing is the process of calculating the difference between two adjacent frames on each channel of the motion feature map. This process yields a channel-level difference feature map for each pair of adjacent frames within the window. The channel-level difference feature map reflects the motion changes in each channel between two adjacent frames.

[0077] The specific calculation can be implemented as follows: for each pair of adjacent frames in the sliding window, element-by-element subtraction is performed on each channel to obtain a channel differential feature map. For example, for a motion feature map containing three channels, a subtraction operation is performed on the two adjacent frames of each channel to obtain three channel differential feature maps. Through channel-level difference operations, the motion change information between adjacent frames can be extracted, providing a basis for subsequent cumulative motion change feature map generation and speed change analysis.

[0078] Step S243: The channel differential feature maps are accumulated and summed along the time dimension to generate a cumulative motion change feature map within the window. The cumulative motion change feature map reflects the sum of the motion changes of each channel in the sliding window.

[0079] Cumulative summation involves adding the channel differential feature maps over time. This cumulative summation yields a cumulative motion change feature map within the window. This cumulative motion change feature map reflects the sum of the motion changes across each channel within the sliding window, providing a more comprehensive description of the motion changes of traffic participants within that time window.

[0080] The specific calculation can be implemented as follows: all channel differential feature maps within the sliding window are element-by-element added along the time dimension to obtain a cumulative motion change feature map. For example, for a sliding window containing five frames of motion feature maps, the channel differential feature maps of each pair of adjacent frames are added along the time dimension to obtain a cumulative motion change feature map. By generating a cumulative motion change feature map, the motion change information between adjacent frames can be integrated, providing more accurate information for subsequent speed change index extraction.

[0081] Step S244: Divide the cumulative motion change feature map into spatial regions, and extract the cumulative change amount of the corresponding region of each traffic participant as the initial speed change index.

[0082] Spatial region segmentation involves dividing the cumulative motion change feature map into distinct regions, with each region corresponding to a traffic participant. The initial speed change index (ISI) is speed change information extracted from the cumulative change in each region corresponding to each traffic participant. It reflects the speed change of the traffic participant within that time window. For example, the segmentation and extraction can be implemented by dividing the cumulative motion change feature map into distinct regions based on the position of the traffic participant in the motion feature map, with each region corresponding to a traffic participant. The cumulative change in each region is then summed or averaged to obtain the cumulative change for that region as the initial speed change index. For example, for a traffic scene containing multiple vehicles and pedestrians, the cumulative motion change feature map is divided into multiple regions based on their positions in the motion feature map. The cumulative change in each region is calculated separately to obtain the initial speed change index for each vehicle and pedestrian. Through spatial region segmentation and initial speed change index extraction, accurate speed change information for each traffic participant can be obtained, providing a foundation for subsequent speed change trend analysis and normalization processing.

[0083] Step S245: performing a trend comparison between the initial speed change index of the current sliding window and the initial speed change index of the previous sliding window to generate an inter-window speed change trend vector.

[0084] Trend comparison involves comparing the changing trends of the initial speed change index of the current sliding window with the initial speed change index of the previous sliding window. This trend comparison generates an inter-window speed change trend vector. The inter-window speed change trend vector reflects the speed change trend of a traffic participant within two adjacent time windows. Specifically, for example, the comparison and generation can be implemented as follows: for each traffic participant, the initial speed change index of the current sliding window is subtracted or divided element-by-element from the initial speed change index of the previous sliding window to generate the inter-window speed change trend vector. For example, if the initial speed change index of the current sliding window is [10, 20, 30] and the initial speed change index of the previous sliding window is [5, 15, 25], then the inter-window speed change trend vector is [5, 5, 5]. By generating the inter-window speed change trend vector, the speed change trends of traffic participants can be analyzed, providing a basis for subsequent speed change index adjustments.

[0085] Step S246: adjusting the initial speed change index of the current sliding window according to the speed change trend vector between windows to generate a normalized speed change vector.

[0086] Adjusting the initial speed change index of the current sliding window is to make it more accurately reflect the actual speed changes of traffic participants. The normalized speed change vector is the adjusted speed change vector, and its value range is usually between 0 and 1, which facilitates subsequent analysis and comparison.

[0087] The specific adjustment can be implemented as follows: according to the speed change trend vector between windows, the initial speed change index of the current sliding window is adjusted. For example, if the speed change trend vector between windows shows that the speed is increasing, the initial speed change index of the current sliding window can be appropriately increased; if the speed change trend vector between windows shows that the speed is decreasing, the initial speed change index of the current sliding window can be appropriately reduced. Then, the adjusted speed change index is normalized so that its value range is between 0 and 1 to obtain a normalized speed change vector. For example, the maximum-minimum normalization method can be used to subtract the minimum value from the adjusted speed change index, and then divide it by the difference between the maximum and minimum values ​​to obtain a normalized speed change vector. Through adjustment and normalization, the speed change index can be made more comparable and accurate, providing reliable data for subsequent speed change sequence generation.

[0088] Step S247: Slide the window along the time axis until the motion feature map of all frames is covered, and splice the normalized speed change vectors of each window in chronological order to form a speed change sequence. Each element of the speed change sequence corresponds to the speed change intensity of the traffic participant within the preset time window.

[0089] Sliding the window along the time axis involves sliding the window from the first frame of the motion feature map in chronological order until all frames of the motion feature map are covered. Splicing involves concatenating the normalized speed change vectors of each window in chronological order to form a continuous sequence. The speed change sequence reflects the speed changes of traffic participants throughout the entire video stream, with each element corresponding to the intensity of the speed change within a preset time window.

[0090] For example, the splicing process can be implemented as follows: starting from the first frame of the motion feature map, a sliding window is set up, and differential calculation, cumulative summation, region division, trend comparison, indicator adjustment, and normalization are performed to obtain the normalized speed change vector for the first window. Then, the sliding window is moved back one frame, and the above process is repeated to obtain the normalized speed change vector for the second window. This process is repeated in this way until the sliding window covers all frames of the motion feature map. Finally, the normalized speed change vectors for each window are concatenated in chronological order to form a speed change sequence. For example, for a video stream containing 100 frames of motion feature maps, the sliding window length is set to 5 frames. After the above processing, the normalized speed change vectors for 20 windows can be obtained. These vectors are concatenated in chronological order to form a speed change sequence containing 20 elements. By generating a speed change sequence, a comprehensive understanding of the speed change trends of traffic participants can be achieved, providing important information for subsequent traffic rule violation detection and signal light control.

[0091] Step S250: normalize and concatenate the motion direction vector, the speed change sequence, and the spatial position offset to form a behavior state feature matrix.

[0092] Normalization converts data of varying scopes and scales into data of the same scope and scale, facilitating subsequent processing and analysis. Concatenation sequentially connects the motion direction vector, velocity change sequence, and spatial position offset to form a matrix. The behavioral state feature matrix, composed of the normalized motion direction vector, velocity change sequence, and spatial position offset, comprehensively describes the behavioral state of traffic participants.

[0093] Specific normalization and splicing can be implemented as follows: first, the motion direction vector, speed change sequence and spatial position offset are normalized separately. For the motion direction vector, each element can be divided by the modulus of the vector to make its length 1; for the speed change sequence and spatial position offset, the maximum-minimum normalization method can be used to convert its value range to between 0 and 1. Then, the normalized motion direction vector, speed change sequence and spatial position offset are spliced ​​in columns or rows to form a behavior state feature matrix. For example, assuming that the dimension of the motion direction vector is 3, the length of the speed change sequence is 10, and the dimension of the spatial position offset is 2, the dimension of the spliced ​​behavior state feature matrix can be (3+10+2) rows and 1 column or 1 row (3+10+2) columns. Through normalized splicing, the formed behavior state feature matrix can be more effectively used for subsequent behavior matching and traffic rule conflict detection.

[0094] As an implementation method, the training process of the spatiotemporal feature coding network includes the following steps S201 to S206:

[0095] Step S201: Collect sample video stream data of historical traffic scenes, and annotate the behaviors of traffic participants in the sample video stream data to generate a behavior label sequence.

[0096] Sample video stream data for historical traffic scenes refers to video stream data collected at the target intersection or other similar intersections over a period of time. This data contains rich information about traffic participant behavior. Behavior labeling involves classifying and labeling the behaviors of traffic participants in the sample video stream data. For example, pedestrians can be labeled as walking normally or running a red light, while vehicles can be labeled as driving normally or making an illegal turn. The behavior label sequence consists of the behavior labels for each traffic participant and serves as supervisory information for training the spatiotemporal feature encoding network.

[0097] The specific collection and labeling can be implemented as follows: by installing a camera device at the intersection, sample video stream data of historical traffic scenes is collected. Then, the behavior of traffic participants in the sample video stream data is annotated manually or using an automated labeling tool. For each traffic participant, a corresponding behavior label is assigned according to his or her behavior performance in the video stream data. The behavior labels of all traffic participants are arranged in chronological order to form a behavior label sequence. For example, for a sample video stream data containing 100 traffic participants, after labeling, a behavior label sequence containing 100 behavior labels can be obtained. By collecting sample video stream data and generating a behavior label sequence, supervised learning data can be provided for the training of the spatiotemporal feature coding network, thereby improving the training effect and accuracy of the network.

[0098] Step S202: constructing a hybrid architecture of a three-dimensional convolutional neural network and a recurrent neural network, wherein the hybrid architecture includes a spatiotemporal feature extraction branch and a behavior classification branch.

[0099] A three-dimensional convolutional neural network (3D CNN) is a convolutional neural network used to process data with three-dimensional structures. It can effectively extract spatiotemporal features from video streams. A recurrent neural network (RNN) is a neural network used to process sequential data, capturing temporal dependencies within the data. A hybrid architecture combines 3D CNN and RNN, leveraging the strengths of both to better process video streams. The spatiotemporal feature extraction branch extracts spatiotemporal features from video streams, and the behavior classification branch classifies traffic participants' behaviors based on these extracted spatiotemporal features.

[0100] For example, the specific construction can be implemented as follows: First, construct the 3D CNN component, which uses multiple 3D convolutional layers and pooling layers to process the input video stream data and extract spatiotemporal features. Then, the output of the 3D CNN is fed into the RNN component, which uses recurrent units such as LSTM or GRU to process the spatiotemporal features and capture temporal dependencies. The spatiotemporal feature extraction branch consists of a 3D CNN and RNN, and its output is a spatiotemporal feature vector. The behavior classification branch is a fully connected layer or multi-layer perceptron, which takes the spatiotemporal feature vector as input and outputs a classification of the traffic participant's behavior. For example, for an input video stream, after processing it through the 3D CNN and RNN, a spatiotemporal feature vector is generated. This vector is then fed into the behavior classification branch, which outputs a traffic participant behavior label, such as normal driving or illegal turning. By constructing a hybrid architecture, spatiotemporal features can be effectively extracted from video stream data and accurately classified.

[0101] Step S203: a multi-scale feature fusion module is set in the spatiotemporal feature extraction branch. The multi-scale feature fusion module is used to integrate motion features under different receptive fields.

[0102] The multi-scale feature fusion module integrates feature information at different scales and levels. It fully utilizes motion features within different receptive fields to improve feature expressiveness. The receptive field refers to the size of the input region that each neuron in a convolutional neural network can perceive. Different convolutional layers have different receptive fields, allowing them to extract features at different scales.

[0103] For example, a specific configuration can be implemented as follows: In the spatiotemporal feature extraction branch, the output features of different convolutional layers are input into a multi-scale feature fusion module. The multi-scale feature fusion module can fuse features of different scales using methods such as convolution, pooling, or an attention mechanism. For example, a 1x1 convolutional layer can be used to reduce the dimensionality of features of different scales, and then the reduced features are concatenated and fused through a convolutional layer. By setting up a multi-scale feature fusion module, motion features under different receptive fields can be integrated to extract richer and more comprehensive spatiotemporal features, providing more accurate feature information for subsequent behavior classification.

[0104] Step S204: Optimizing the feature discrimination of the spatiotemporal feature extraction branch by using a contrastive learning loss function. The contrastive learning loss function causes the feature vectors of similar behavior samples to cluster in the embedding space.

[0105] The contrastive learning loss function is used to optimize feature discrimination. By comparing the feature similarities between different samples, it encourages feature vectors of similar behavior samples to cluster in the embedding space, while separating feature vectors of different behavior samples. The embedding space is the low-dimensional space into which feature vectors are mapped. This space allows for more intuitive observation and analysis of the relationships between feature vectors.

[0106] For example, the optimization can be implemented as follows: First, randomly select positive and negative pairs from the sample video stream data. The positive pairs contain different instances of the same behavior category, while the negative pairs contain instances of different behavior categories. Then, data augmentation is used to generate transformed views of the positive pairs, including spatial cropping, temporal slicing, and color jittering to increase sample diversity. The cosine similarity between the feature vectors of the positive pairs is calculated, and a positive similarity distribution matrix is ​​constructed. The cosine similarity between the feature vectors of the negative pairs is calculated, and a negative similarity distribution matrix is ​​constructed. A temperature-scaled cross-entropy loss function is used to maximize the separation between the positive and negative similarity distribution matrices. Finally, backpropagation is used to adjust the convolution kernel parameters of the spatiotemporal feature extraction branch and the gating weights of the recurrent neural network to optimize feature discrimination. For example, for normal driving and illegal turning, by optimizing the contrastive learning loss function, the feature vectors of normal driving samples are clustered together in the embedding space, while the feature vectors of illegal turning samples are clustered together in the embedding space, with the distance between the two feature vectors being as large as possible. By optimizing the feature discrimination of the spatiotemporal feature extraction branch, the accuracy and reliability of behavior classification can be improved.

[0107] As an implementation method, the optimization process of the contrastive learning loss function includes the following steps S2041 to S2046:

[0108] Step S2041: randomly selecting positive sample pairs and negative sample pairs from the sample video stream data, where the positive sample pairs contain different instances of the same behavior category, and the negative sample pairs contain instances of different behavior categories.

[0109] A positive pair is composed of different instances of the same behavior category, for example, a pair of two normally moving vehicles. A negative pair is composed of instances of different behavior categories, for example, a pair of one normally moving vehicle and one making an illegal turn. Positive and negative pairs are randomly selected to ensure sample diversity and randomness, improving the effectiveness of contrastive learning.

[0110] The specific selection can be implemented as follows: first, classify the behaviors of traffic participants in the sample video stream data into different behavior categories. Then, randomly select different instances from each behavior category to form a positive sample pair. At the same time, randomly select instances from different behavior categories to form a negative sample pair. For example, for vehicle behaviors in the sample video stream data, they are divided into behavior categories such as normal driving, illegal turning, and running a red light. Two different vehicle instances are randomly selected from the normal driving category to form a positive sample pair, and one vehicle instance is selected from the normal driving category and the illegal turning category to form a negative sample pair. By randomly selecting positive and negative sample pairs, diverse samples can be provided for comparative learning, which helps to optimize the feature discrimination of the spatiotemporal feature extraction branch.

[0111] Step S2042: Generate a transformed view of the positive sample pair through data enhancement, where the transformed view includes spatial cropping, temporal slicing, and color jittering.

[0112] Data augmentation is a method used to increase the number and diversity of samples. By augmenting positive sample pairs, more samples can be generated, improving the model's generalization ability. Transformed views refer to sample views obtained after data augmentation, including operations such as spatial cropping, temporal slicing, and color jittering.

[0113] The specific generation can be implemented as follows: for each sample in the positive sample pair, a spatial cropping operation is performed to randomly crop a portion of the sample area to change the spatial characteristics of the sample. A temporal slicing operation is performed to randomly select a portion of the temporal segment of the sample to change the temporal characteristics of the sample. A color jittering operation is performed to randomly adjust the color of the sample, such as changing the brightness, contrast, saturation, etc., to change the color characteristics of the sample. Through these operations, a transformed view of the positive sample pair is generated. For example, for a sample video containing a driving vehicle, spatial cropping can be performed to crop out part of the background around the vehicle, temporal slicing can be performed to select a certain time segment during the vehicle's driving process, and color jittering can be performed to change the color hue of the video. By generating transformed views, the diversity of the positive sample pairs can be increased and the effect of contrastive learning can be improved.

[0114] Step S2043: Calculate the cosine similarity between the feature vectors of the positive sample pairs and construct a positive sample similarity distribution matrix.

[0115] Cosine similarity is a metric used to measure the similarity between two vectors. It expresses similarity by calculating the cosine of the angle between the two vectors. The positive sample similarity distribution matrix is ​​a matrix composed of the cosine similarities between the feature vectors of positive sample pairs. It can reflect the similarity distribution between positive sample pairs.

[0116] The specific calculation and construction can be implemented as follows: for each positive sample pair, input it into the spatiotemporal feature extraction branch to obtain a feature vector. Then, calculate the cosine similarity between the feature vectors of each positive sample pair. Arrange the cosine similarities of all positive sample pairs in order to construct a positive sample similarity distribution matrix. For example, for a sample set containing 100 positive sample pairs, calculate the cosine similarity between the feature vectors of each positive sample pair to obtain 100 cosine similarity values, and arrange these values ​​in order to form a 100x1 positive sample similarity distribution matrix. By constructing the positive sample similarity distribution matrix, the similarity distribution between positive sample pairs can be intuitively observed, providing a basis for the subsequent calculation of the contrastive learning loss function.

[0117] Step S2044: Calculate the cosine similarity between the feature vectors of the negative sample pairs and construct a negative sample similarity distribution matrix.

[0118] Similar to step S2043, the cosine similarity between the feature vectors of the negative sample pairs is calculated, and a negative sample similarity distribution matrix is ​​constructed. The negative sample similarity distribution matrix reflects the similarity distribution between the negative sample pairs.

[0119] The specific calculation and construction can be implemented as follows: for each negative sample pair, input it into the spatiotemporal feature extraction branch to obtain a feature vector. Then, calculate the cosine similarity between the feature vectors of each negative sample pair. Arrange the cosine similarities of all negative sample pairs in order to construct a negative sample similarity distribution matrix. For example, for a sample set containing 100 negative sample pairs, calculate the cosine similarity between the feature vectors of each negative sample pair to obtain 100 cosine similarity values, and arrange these values ​​in order to form a 100×1 negative sample similarity distribution matrix. By constructing the negative sample similarity distribution matrix, the similarity distribution between negative sample pairs can be intuitively observed, providing a basis for the subsequent calculation of the contrastive learning loss function.

[0120] Step S2045: maximizing the separation between the positive sample similarity distribution matrix and the negative sample similarity distribution matrix by temperature scaling the cross entropy loss function.

[0121] The temperature-scaled cross entropy loss function is a loss function used to optimize classification problems. It adjusts the separation between the similarity distribution matrix of positive samples and the similarity distribution matrix of negative samples, so that the feature vectors of similar behavior samples are aggregated in the embedding space and the feature vectors of different behavior samples are separated in the embedding space.

[0122] The specific optimization can be implemented as follows: the positive sample similarity distribution matrix and the negative sample similarity distribution matrix are input as input to the temperature scaling cross entropy loss function. The temperature scaling cross entropy loss function calculates the loss value based on the similarity distribution of positive samples and negative samples. Through the back propagation algorithm, the convolution kernel parameters of the spatiotemporal feature extraction branch and the gating weights of the recurrent neural network are adjusted to minimize the loss value, thereby maximizing the separation between the positive sample similarity distribution matrix and the negative sample similarity distribution matrix. For example, by adjusting the parameters, the values ​​in the positive sample similarity distribution matrix are made as large as possible, and the values ​​in the negative sample similarity distribution matrix are made as small as possible, thereby achieving the aggregation of feature vectors of similar behavior samples and the separation of feature vectors of different behavior samples. By using the temperature scaling cross entropy loss function, the feature discrimination of the spatiotemporal feature extraction branch can be effectively optimized, thereby improving the accuracy of behavior classification.

[0123] Step S2046: Back propagation adjusts the convolution kernel parameters of the spatiotemporal feature extraction branch and the gating weights of the recurrent neural network.

[0124] In an embodiment of the present invention, the convolution kernel parameters of the spatiotemporal feature extraction branch and the gating weights of the recurrent neural network are adjusted by the back propagation algorithm to optimize the performance of the network.

[0125] The specific adjustment can be implemented as follows: based on the calculation result of the temperature-scaled cross entropy loss function, the gradient of the loss function on the convolution kernel parameters of the spatiotemporal feature extraction branch and the gating weights of the recurrent neural network is calculated. Then, the convolution kernel parameters and gating weights are updated according to the gradient descent algorithm. For example, for the convolution kernel parameters, optimization algorithms such as stochastic gradient descent (SGD) and Adam can be used to update them. By continuously performing backpropagation and parameter updates, the spatiotemporal feature extraction branch can better extract discriminative features and improve the accuracy of behavior classification.

[0126] Step S205: A dynamic focus loss function is used in the behavior classification branch. The dynamic focus loss function automatically adjusts the classification weight according to the category distribution of the samples.

[0127] The dynamic focus loss function is used to address class imbalance. It automatically adjusts classification weights based on the class distribution of samples, allowing the model to focus more on samples from minority classes. In the behavior classification branch, since the number of samples for different behavior classes can vary significantly, using the dynamic focus loss function can improve the model's classification performance for minority class behaviors.

[0128] Specifically, for example, it can be implemented as follows: in the output layer of the behavior classification branch, the prediction results and the true labels are input into the dynamic focus loss function. The dynamic focus loss function automatically adjusts the classification weight of each category according to the category distribution of the samples. For samples of the minority category, its classification weight is increased so that the model pays more attention to these samples; for samples of the majority category, its classification weight is reduced to prevent the model from being too biased towards the majority category. For example, in the classification of traffic participant behavior, the number of samples of illegal behavior may be small, and the number of samples of normal behavior may be large. Using the dynamic focus loss function can increase the classification weight of illegal behavior samples and improve the model's classification accuracy of illegal behavior. By adopting the dynamic focus loss function, the problem of category imbalance can be effectively solved and the overall performance of behavior classification can be improved.

[0129] Step S206: using the output of the optimized spatiotemporal feature extraction branch as the input feature of the preset behavior matching model.

[0130] The preset behavior matching model is a model used to determine whether the behavior of traffic participants meets the preset traffic rules conflict conditions. It uses the output of the optimized spatiotemporal feature extraction branch as input features to perform behavior matching and conflict detection. Specifically, for example, it can be implemented as follows: after training and optimization in the above steps, the spatiotemporal feature extraction branch can extract spatiotemporal features with high discrimination. These features are used as input to the preset behavior matching model, and the preset behavior matching model will determine whether the behavior of traffic participants meets the preset traffic rules conflict conditions based on these features. For example, the preset behavior matching model can determine whether a vehicle runs a red light, whether a pedestrian crosses the road illegally, etc. based on the input spatiotemporal features. By using the output of the optimized spatiotemporal feature extraction branch as the input feature of the preset behavior matching model, the accuracy and reliability of behavior matching and conflict detection can be improved.

[0131] Step S300: Input the behavior state characteristics into a preset behavior matching model to generate a behavior trigger identifier of the traffic participant in the current time window. The behavior trigger identifier is used to indicate whether the traffic participant meets the preset traffic rule conflict condition.

[0132] The behavior state features are extracted through steps S200-S250 and contain information such as the traffic participant's motion direction vector, speed change sequence, and spatial position offset. The preset behavior matching model is a trained and optimized model that can determine whether the traffic participant's behavior meets the preset traffic rule conflict conditions based on the input behavior state features. The behavior trigger flag is an identifier used to indicate whether the traffic participant's behavior triggers a traffic rule conflict. It can be a binary value (such as 0 for not triggered and 1 for triggered) or an identifier containing the conflict type and trigger timestamp.

[0133] Specifically, the preset behavior matching model can be constructed based on machine learning or deep learning algorithms, such as a decision tree model, a neural network model, etc. In this embodiment, a decision tree model is combined with rule judgment to achieve behavior matching.

[0134] As an embodiment, step S300, inputting the behavior state characteristics into a preset behavior matching model to generate a behavior trigger identifier of the traffic participant in the current time window, may include the following steps S310-S390:

[0135] Step S310: Calculate the direction angle between the motion direction vector in the behavior state feature and the set of allowed passing directions corresponding to the phase state of the current traffic light to generate a direction deviation angle sequence.

[0136] The direction of motion vector represents the direction of motion of a traffic participant at a given moment. The set of permitted directions corresponding to the current signal light phase state is the legal direction of travel determined by the signal light display state. Direction angle calculation calculates the angle between two vectors using the dot product and the modulus of the vectors. The direction deviation angle sequence is a sequence of angles between the direction of motion vector and the permitted direction of travel within each time window.

[0137] The specific calculation can be implemented as follows: for the motion direction vector in the behavior state feature, perform a dot product operation on it with each allowed direction vector in the allowed direction set corresponding to the current signal light phase state, and then divide it by the product of the module lengths of the two vectors to obtain the cosine value of the angle. The angle value is calculated by the inverse cosine function, and the angle value is used as the direction deviation angle within the time window. Each time window is calculated in turn to obtain a direction deviation angle sequence. For example, at an intersection, the current signal light shows a green light for straight ahead, and the allowed direction of travel is due east. The motion direction vector of a vehicle points to the northeast. By calculating the angle between the two vectors, the direction deviation angle of the vehicle within the time window is obtained. By generating a direction deviation angle sequence, the deviation between the motion direction of traffic participants and the allowed direction can be intuitively understood, providing a basis for subsequent direction conflict detection.

[0138] Step S320: comparing the value of each time window in the speed change sequence with the speed threshold interval associated with the phase state, and generating a speed out-of-bounds mark sequence.

[0139] The speed change sequence, generated in step S240, reflects the speed changes of traffic participants within different time windows. The phase-state-associated speed threshold interval is the legal speed range determined based on the current signal phase state and traffic regulations. For example, during a red light, the vehicle's speed should be 0, while during a green light, the vehicle's speed has upper and lower limits. The speed violation flag sequence consists of flags indicating whether the speed violates the speed limit within each time window. The flags can be binary values ​​(e.g., 0 for no violation, 1 for violation).

[0140] The specific comparison can be implemented as follows: for each time window in the speed change sequence, the value is compared with the upper and lower limits of the speed threshold interval associated with the phase state. If the value is less than the lower limit or greater than the upper limit, it is marked as out of bounds, and the mark value is 1; otherwise, it is marked as not out of bounds, and the mark value is 0. Each time window is compared in turn to obtain a speed out-of-bounds mark sequence. For example, in the green light straight-through phase state, the speed threshold interval is [20,60] km / h, and the speed of a vehicle in a certain time window is 70 km / h, then the mark value of the time window is 1. By generating a speed out-of-bounds mark sequence, it is possible to quickly determine whether the speed of traffic participants meets the requirements of the current phase state, providing a basis for subsequent speed anomaly detection.

[0141] Step S330: performing sliding window accumulation of the direction deviation angle sequence in the time dimension, and generating a direction conflict flag when the accumulated angles of three consecutive windows exceed a preset deviation threshold.

[0142] Sliding window accumulation in the time dimension involves setting a fixed-length sliding window on the direction deviation angle sequence and summing the direction deviation angles within the window. A preset deviation threshold is a pre-set angle value. When the cumulative angle of three consecutive windows exceeds this threshold, a conflict in the direction of movement of the traffic participant is considered, and a direction conflict indicator is generated.

[0143] The specific accumulation and judgment can be implemented as follows: a sliding window of length 3 is set on the direction deviation angle sequence. Starting from the first element of the sequence, the cumulative value of the direction deviation angle within each window is calculated in sequence. When the cumulative angle of a window exceeds the preset deviation threshold, a direction conflict flag is generated. For example, if the preset deviation threshold is 60°, and the direction deviation angles in three consecutive time windows are 20°, 30°, and 25° respectively, the cumulative angle is 75°, which exceeds the preset deviation threshold, and a direction conflict flag is generated. By performing a sliding window accumulation and judgment on the direction deviation angle sequence, directional conflict behavior of traffic participants can be effectively detected.

[0144] Step S340: Perform continuous positive event detection on the speed crossing mark sequence, and generate a speed abnormality mark when it is detected that the speed exceeds the same direction threshold twice in a row.

[0145] Continuous positive event detection involves detecting consecutive positive signs (i.e., a speed violation value of 1) within a speed violation sequence. If the speed exceeds the same direction threshold twice in a row, the traffic participant's speed is considered abnormal and a speed anomaly flag is generated. The same direction threshold refers to the speed exceeding the upper limit or falling below the lower limit.

[0146] The specific detection can be implemented as follows: traverse the speed out-of-bounds mark sequence, and when a positive mark is encountered, check whether the next mark is also a positive mark, and whether the two out-of-bounds directions are the same (i.e., both exceed the upper limit or both fall below the lower limit). If the conditions are met, a speed anomaly flag is generated. For example, in the speed out-of-bounds mark sequence, the 3rd and 4th mark values ​​are both 1, and both are due to the speed exceeding the upper limit, then a speed anomaly flag is generated. By performing continuous positive event detection on the speed out-of-bounds mark sequence, abnormal speed behavior of traffic participants can be accurately detected.

[0147] Step S350: inputting the direction conflict flag and the speed abnormality flag into a decision tree model associated with the phase state, and calculating the priority weight of the conflict behavior according to the remaining duration of the phase state.

[0148] The phase state-associated decision tree model makes decisions based on the current signal's phase state and conflict behavior information. It calculates the priority weights of conflicting behaviors based on different phase states and conflict types. The remaining duration of a phase state refers to the remaining time in the current signal's phase state.

[0149] The specific calculation can be implemented as follows: the direction conflict flag and the speed abnormality flag are input into the decision tree model associated with the phase state. The decision tree model performs node judgment and branch selection based on the current phase state of the traffic light and the conflict behavior information. For each conflict behavior, its priority weight is calculated based on the remaining duration of the phase state. For example, in the red light phase state, if the remaining time is short and there is a direction conflict and speed abnormality behavior, the decision tree model may consider the direction conflict to have a higher priority because the direction conflict may lead to more serious traffic conflicts. By calculating the priority weights of the conflicting behaviors, different types of conflicting behaviors can be sorted, providing a basis for subsequent dynamic fusion.

[0150] Step S360: Dynamically merge the direction conflict flag and the speed abnormality flag based on the priority weight to generate an initial behavior trigger probability.

[0151] Dynamic fusion involves combining direction conflict indicators and speed anomaly indicators based on their priority to generate an initial maneuver trigger probability. This initial maneuver trigger probability represents the likelihood that a traffic participant will trigger a traffic rule conflict within the current time window.

[0152] Specifically, fusion can be implemented by taking a weighted sum of the direction conflict flag and the speed anomaly flag based on the priority weights calculated in step S350. For example, if the direction conflict flag has a priority weight of 0.6, the speed anomaly flag has a priority weight of 0.4, the direction conflict flag has a value of 1, and the speed anomaly flag has a value of 0, then the initial behavior trigger probability is 0.6 × 1 + 0.4 × 0 = 0.6. Through dynamic fusion, both direction conflict and speed anomaly conflict behaviors can be comprehensively considered to generate a more accurate initial behavior trigger probability.

[0153] Step S370: aligning the initial behavior trigger probability with the trigger probability sequence of the historical time window, and generating a corrected behavior trigger probability through an exponential smoothing algorithm.

[0154] Time series alignment involves aligning the initial behavior trigger probability with the trigger probability sequence of the historical time window in the temporal dimension to ensure data consistency. Exponential smoothing is an algorithm used to smooth time series data. It can generate smoother and more accurate forecasts based on historical and current data.

[0155] The specific processing can be implemented as follows: the initial behavior trigger probability and the trigger probability sequence of the historical time window are arranged in chronological order to align the time series. Then, the exponential smoothing algorithm is used to process the initial behavior trigger probability and the historical trigger probability sequence. The formula of the exponential smoothing algorithm is: S t =αy t +(1-α)S t-1 , where S t is the corrected behavior trigger probability, y t is the initial behavior triggering probability, S t-1 is the corrected trigger probability for the previous time window, and α is the smoothing coefficient, ranging from [0 to 1]. For example, if the smoothing coefficient α is 0.3, the corrected trigger probability for the previous time window is 0.5, and the initial behavior trigger probability is 0.6, then the corrected behavior trigger probability is 0.3 × 0.6 + (1 - 0.3) × 0.5 = 0.53. Time series alignment and exponential smoothing can reduce the impact of noise and fluctuations, generating a more stable and accurate corrected behavior trigger probability.

[0156] Step S380: When the corrected behavior trigger probability exceeds a dynamic trigger threshold that is negatively correlated with the current intersection traffic density, a valid behavior trigger flag is generated.

[0157] The dynamic trigger threshold is a threshold that is dynamically adjusted according to the current traffic density at the intersection. It is negatively correlated with the traffic density, that is, the greater the traffic density, the smaller the dynamic trigger threshold. When the corrected behavior trigger probability exceeds the dynamic trigger threshold, it is considered that the behavior of the traffic participant has triggered a traffic rule conflict, and a valid behavior trigger identifier is generated. The specific judgment can be implemented, for example, as follows: real-time monitoring of the traffic density at the current intersection, and calculating the dynamic trigger threshold based on the traffic density. For example, traffic density information can be obtained through vehicle detectors installed at the intersection or video analysis technology, and then the dynamic trigger threshold is calculated based on a pre-set functional relationship. The corrected behavior trigger probability is compared with the dynamic trigger threshold. If the corrected behavior trigger probability exceeds the dynamic trigger threshold, a valid behavior trigger identifier is generated. By using the dynamic trigger threshold, it is possible to flexibly judge whether the behavior of the traffic participant triggers a conflict based on the actual traffic conditions at the intersection, thereby improving the accuracy of conflict detection.

[0158] Step S390: Perform logical constraint verification on the valid behavior trigger identifier and the remaining time of the phase state, eliminate trigger events whose remaining time is less than the safety response threshold, and generate a final behavior trigger identifier set. Each element in the final behavior trigger identifier set contains the conflict type and trigger timestamp.

[0159] Logical constraint verification involves logically evaluating valid behavior trigger identifiers and the remaining time of phase states to ensure that the triggering event is meaningful and manageable. The safety response threshold is a pre-set time value. When the remaining time of a phase state is less than this threshold, the triggering event is considered unresolved even if a conflict event is triggered, and therefore eliminated. The final behavior trigger identifier set is the set of triggering events obtained after logical constraint verification. It includes the conflict type and trigger timestamp for each triggering event.

[0160] The specific verification and generation can be implemented as follows: for each valid behavior trigger identifier, compare it with the remaining time of the phase state. If the remaining time of the phase state is less than the safety response threshold, the trigger event is eliminated; otherwise, the conflict type and trigger timestamp of the trigger event are added to the final behavior trigger identifier set. For example, if the safety response threshold is 5 seconds and the remaining time of the phase state corresponding to a valid behavior trigger identifier is 3 seconds, the trigger event is eliminated; if the remaining time is 8 seconds, the conflict type (such as direction conflict, speed abnormality) and trigger timestamp of the trigger event are added to the final behavior trigger identifier set. By verifying the logical constraints and generating the final behavior trigger identifier set, it can be ensured that the trigger event has actual processing value and provides accurate information for subsequent traffic light control.

[0161] Step S400: Generate a signal light control instruction for the edge computing node according to the timing correlation between the behavior trigger identifier and the phase state of the current signal light.

[0162] The behavior trigger flag contains information about conflicting behaviors of traffic participants. The current signal light phase state indicates the signal light's display status and remaining time at a specific moment. Timing correlation refers to the temporal relationship between the trigger time of the behavior trigger flag and the current signal light phase state. Edge computing nodes are computing devices located at the edge of the network that can process and make decisions based on local data in real time. Signal light control instructions are used to control signal light phase switching and contain information such as the phase switching direction and duration.

[0163] As an embodiment, step S400, generating a signal light control instruction for an edge computing node based on a temporal correlation between the behavior trigger identifier and the current phase state of the signal light, may include the following steps S410-S4100:

[0164] Step S410: Map and match each behavior trigger identifier in the final behavior trigger identifier set with the remaining time of the phase state to generate a priority parameter set including the conflict type, trigger timestamp and remaining time.

[0165] Mapping matches associate each action trigger identifier in the final action trigger identifier set with the remaining time of the current signal light phase state, generating a priority parameter set containing the conflict type, trigger timestamp, and remaining time. The priority parameter set is used for subsequent phase adjustment rule activation and priority calculation.

[0166] The specific mapping matching can be implemented as follows: for each behavior trigger identifier in the final behavior trigger identifier set, obtain its conflict type and trigger timestamp. At the same time, obtain the remaining time of the current signal light phase state. Combine the conflict type, trigger timestamp and remaining time into a priority parameter, arrange all priority parameters in order, and generate a priority parameter set. For example, there is a direction conflict trigger identifier in the final behavior trigger identifier set, the trigger timestamp is 10:00:00, and the remaining time of the current signal light phase state is 30 seconds, then the generated priority parameter is (direction conflict, 10:00:00, 30 seconds). By mapping matching and generating a priority parameter set, the behavior trigger identifier and the time information of the phase state can be integrated to provide more comprehensive information for subsequent signal light control decisions.

[0167] Step S420: activating corresponding phase adjustment rules according to the conflict type in the priority parameter set, where the phase adjustment rules include a direction conflict priority strategy and a speed anomaly degradation strategy.

[0168] Phase adjustment rules are traffic light phase adjustment strategies tailored to different conflict types. They include a directional conflict priority strategy and a speed anomaly degradation strategy. The directional conflict priority strategy prioritizes adjusting the traffic light phase to resolve directional conflicts. The speed anomaly degradation strategy appropriately reduces the priority of resolving speed anomalies.

[0169] Specific activation can be implemented, for example, by traversing the priority parameter set and activating the corresponding phase adjustment rules based on the conflict type. If the conflict type is a directional conflict, the directional conflict priority strategy is activated; if the conflict type is a speed anomaly, the speed anomaly degradation strategy is activated. For example, the priority parameter set contains a directional conflict priority parameter and a speed anomaly priority parameter. For the directional conflict priority parameter, the directional conflict priority strategy is activated; for the speed anomaly priority parameter, the speed anomaly degradation strategy is activated. By activating the corresponding phase adjustment rules based on the conflict type, different processing strategies can be adopted for different conflict situations, thereby improving the effectiveness of traffic light control.

[0170] Step S430: spatially aggregate multiple direction conflict identifiers within the same time window using a direction conflict priority strategy to generate a phase switching direction candidate set.

[0171] The directional conflict priority strategy prioritizes directional conflicts. Spatial orientation aggregation integrates the spatial orientation information of multiple directional conflict markers within the same time window. The phase switching direction candidate set is a set of possible phase switching directions obtained after spatial orientation aggregation.

[0172] The specific aggregation and generation can be implemented as follows: for multiple direction conflict identifiers within the same time window, extract their spatial orientation information, such as the vehicle's driving direction, lane, etc. These spatial orientation information are classified and merged to find direction conflict identifiers with the same or similar spatial orientations. Based on these aggregated spatial orientation information, the possible phase switching directions are determined, and a phase switching direction candidate set is generated. For example, at an intersection, there are direction conflict identifiers of multiple vehicles in the same time window. After spatial orientation aggregation, it is found that most of the conflicting vehicles are from the east-west direction. Then the phase switching direction candidate set may include the green light phase that switches from the current phase to the east-west direction. By spatial orientation aggregation and generating the phase switching direction candidate set, the direction in which the phase switching needs to be performed can be determined more accurately, thereby improving the targetedness of traffic light control.

[0173] Step S440: Calculate the duration of the direction conflict based on the remaining time of the phase state and the trigger timestamp, and generate a phase adjustment parameter in combination with the occurrence frequency of the speed anomaly indicator in the speed anomaly degradation strategy.

[0174] The directional conflict duration refers to the length of time from the directional conflict trigger time to the current time or the end time of the phase state. It can be calculated using the remaining time of the phase state and the trigger timestamp. The speed anomaly indicator frequency in the speed anomaly degradation policy refers to the number of times the speed anomaly indicator appears within the specified time range. The phase adjustment parameter is used to adjust the traffic light phase, taking into account both the directional conflict duration and the speed anomaly indicator frequency.

[0175] The specific calculation and generation can be implemented as follows: For each directional conflict indicator, the duration of the directional conflict is calculated based on the remaining time of the phase state and the trigger timestamp. For example, if the trigger timestamp is 10:00:00, the current time is 10:00:10, and the remaining time of the phase state is 20 seconds, then the duration of the directional conflict is 30 seconds. Simultaneously, the frequency of speed anomaly indicators is counted. A phase adjustment parameter is generated based on the duration of the directional conflict and the frequency of speed anomaly indicators. For example, a weighted summation method can be used, with a weight of 0.7 for the duration of the directional conflict and a weight of 0.3 for the frequency of speed anomaly indicators. If the duration of the directional conflict is 30 seconds and the frequency of speed anomaly indicators is 2, then the phase adjustment parameter is 0.7 × 30 + 0.3 × 2 = 21.6. By calculating the duration of the directional conflict and combining it with the frequency of speed anomaly indicators to generate the phase adjustment parameter, the impact of different types of conflicts can be more comprehensively considered, providing a more reasonable basis for traffic light phase adjustment.

[0176] Step S450: Input the phase switching direction candidate set and the phase adjustment parameters into the time series optimizer to generate a phase switching timing plan that meets the minimum phase switching interval and maximum traffic efficiency constraints.

[0177] The Time Series Optimizer is a model used to optimize traffic light phase switching timing. Based on a set of candidate phase switching directions and phase adjustment parameters, it generates a phase switching timing plan that satisfies the constraints of minimum phase switching interval and maximum traffic efficiency. Minimum phase switching interval refers to the minimum time interval between two consecutive phase switches, which ensures traffic stability and safety. Maximum traffic efficiency refers to the phase switching plan that maximizes traffic flow while maintaining the minimum phase switching interval.

[0178] The specific optimization can be implemented as follows: inputting the phase switching direction candidate set and phase adjustment parameters into the time series optimizer. The time series optimizer can use optimization algorithms such as genetic algorithms and simulated annealing algorithms to search and optimize the phase switching timing. During the optimization process, the constraints of minimum phase switching interval and maximum traffic efficiency are considered. For example, through a genetic algorithm, a set of phase switching timing schemes are randomly generated, and the traffic efficiency of each scheme and whether it meets the minimum phase switching interval requirements are calculated. Then, through operations such as selection, crossover and mutation, better schemes are continuously evolved until the optimal phase switching timing scheme that meets the constraints is found. By generating a phase switching timing scheme through a time series optimizer, traffic efficiency can be improved while ensuring traffic stability and safety.

[0179] Step S460: Generate a phase switching direction code according to the difference between the target phase direction in the phase switching timing scheme and the current phase state.

[0180] Phase switch direction encoding is used to represent the phase switch direction. It is generated based on the difference between the target phase direction in the phase switch timing scheme and the current phase state. The target phase direction is the phase direction to be switched to in the phase switch timing scheme, and the current phase state is the current display status of the signal light.

[0181] The specific generation can be implemented as follows: compare the target phase direction in the phase switching timing scheme with the current phase state, and calculate the difference between them. Depending on the difference, different encoding methods are used to generate the phase switching direction code. For example, at an intersection, the current phase state is a green light in the north-south direction, and the target phase direction is a green light in the east-west direction. The difference is switching from the north-south direction to the east-west direction. Binary encoding can be used to represent the north-south direction as 00 and the east-west direction as 01, and the phase switching direction code is 01. By generating the phase switching direction code, the phase switching direction information can be converted into a digital code, which is convenient for the subsequent generation and transmission of signal light control instructions.

[0182] Step S470: Binary-concatenate the phase switching direction code and the duration in the phase adjustment parameter to form a traffic light control instruction including a direction identification bit and a duration identification bit.

[0183] Binary bit concatenation converts the phase switch direction code and the duration of the phase adjustment parameter into binary bits and concatenates them to form a complete signal light control instruction. The direction bit indicates the direction of the phase switch, and the duration bit indicates the duration after the phase switch.

[0184] The specific splicing can be implemented as follows: convert the phase switching direction code into binary bits, for example, the phase switching direction code is 01, which is converted into binary bits 00000001. The duration in the phase adjustment parameter is also converted into binary bits, for example, the duration is 30 seconds, which is converted into binary bits 00011110. Then, the direction identification bit and the duration identification bit are spliced ​​together to form a traffic light control instruction. For example, the direction identification bit 00000001 and the duration identification bit 00011110 are spliced ​​together to obtain the traffic light control instruction 0000000100011110. Through binary bit splicing, the traffic light control instruction formed can accurately convey the phase switching direction and duration information, which is convenient for the edge computing node to control the traffic light.

[0185] Step S480: before generating the signal light control instruction, perform a conflict pre-check to verify whether the phase switching direction code has a spatial overlap risk with the trajectory of the unfinished vehicle in the opposite lane.

[0186] Conflict pre-check refers to checking the phase switching direction code and the trajectory of the unfinished vehicle in the opposite lane before generating the signal light control instruction to verify whether there is a risk of spatial overlap. The risk of spatial overlap means that after the phase switch, the unfinished vehicle in the opposite lane may collide or conflict with the vehicle in the new phase direction. The specific pre-check can be implemented as follows: obtaining the trajectory information of the unfinished vehicle in the opposite lane through video stream data and trajectory tracking results. The phase switching direction corresponding to the phase switching direction code and the trajectory of the unfinished vehicle are spatially analyzed to determine whether there is a possibility of spatial overlap. For example, by calculating the intersection of the vehicle trajectory and the lane line of the new phase direction, it is determined whether there is a vehicle that will enter the lane of the new phase direction after the phase switch. If there is a risk of spatial overlap, the phase switching plan needs to be adjusted or other measures need to be taken to avoid conflict. By performing conflict pre-check, potential conflict risks can be discovered in advance, improving the safety of signal light control.

[0187] Step S490: When a spatial overlap risk is detected, a forced waiting period flag is embedded in the signal light control instruction, and the forced waiting period flag extends the duration of the current phase state until the risk is resolved.

[0188] The mandatory wait period flag is used to extend the duration of the current phase state. When a spatial overlap risk is detected, it is embedded in the signal light control instruction. By extending the duration of the current phase state, vehicles in the oncoming lane that have not yet completed their journey have sufficient time to pass through the intersection, avoiding collisions with vehicles traveling in the new phase direction. Specifically, when a spatial overlap risk is detected, the mandatory wait period flag is added to the signal light control instruction. The mandatory wait period flag can be a binary code, such as 1111. Simultaneously, the required extension time is calculated based on the trajectory and speed of the vehicle in the oncoming lane that has not yet completed its journey. This extended time is then added to the duration flag of the signal light control instruction. For example, if the original signal light control instruction is 0000000100011110, and a spatial overlap risk is detected, the mandatory wait period flag 1111 is embedded, and the extended time (e.g., 10 seconds, converted to binary 00001010) is added to the duration flag, resulting in a new signal light control instruction of 0000000111110001111000001010. By embedding mandatory waiting period signs, the risk of spatial overlap can be effectively avoided and traffic safety and smoothness can be ensured.

[0189] Step S4100: Encapsulate the verified traffic light control instruction into an instruction frame structure executable by the edge computing node. The instruction frame structure includes a timestamp synchronization field and an instruction effectiveness countdown field.

[0190] The instruction frame structure is used to encapsulate traffic light control instructions. It contains a timestamp synchronization field and an instruction validity countdown field. The timestamp synchronization field is used to ensure time synchronization between the edge computing node and the traffic light device, and the instruction validity countdown field is used to indicate the time when the traffic light control instruction takes effect.

[0191] The specific encapsulation can be implemented as follows: adding the verified traffic light control instruction to the instruction frame structure. Adding a timestamp synchronization field to the instruction frame structure, the timestamp synchronization field can be the current timestamp, such as 10:00:00. At the same time, adding an instruction effective countdown field, the instruction effective countdown field can be calculated based on the duration in the traffic light control instruction and the current time. For example, if the duration in the traffic light control instruction is 30 seconds and the current time is 10:00:00, the instruction effective countdown field is 10:00:30. These fields are combined into a complete instruction frame structure, such as [timestamp synchronization field: 10:00:00, traffic light control instruction: 0000000111110001111000001010, instruction effective countdown field: 10:00:30]. By encapsulating the traffic light control instruction into an instruction frame structure, the communication between the edge computing node and the traffic light device can be ensured to be accurate and effective, and the correct execution of the traffic light control instruction can be guaranteed.

[0192] Step S500: adjusting the signal light phase switching sequence of the target intersection based on the signal light control instruction so that traffic participants who meet the traffic rule conflict conditions obtain passage priority under the signal light phase switching sequence.

[0193] Traffic light control instructions contain information such as phase switching direction and duration. By executing these instructions, the signal light phase switching sequence at the target intersection can be adjusted. Traffic participants meeting traffic rule conflict conditions are those marked as triggering traffic rule conflicts in the behavior trigger indicator. By adjusting the signal light phase switching sequence, these traffic participants receive priority at the appropriate time, thereby reducing traffic conflicts and congestion.

[0194] As an embodiment, step S500, adjusting the signal light phase switching timing of the target intersection based on the signal light control instruction, may include the following steps S510-S560:

[0195] Step S510: parsing the phase switching direction identifier and phase duration parameter in the signal light control instruction.

[0196] Parsing is the process of converting the binary information encoded in the signal light control command into specific phase switching direction and phase duration parameters. The phase switching direction identifier indicates the phase direction the signal light needs to switch to, and the phase duration parameter indicates the duration of the phase state.

[0197] Specific analysis can be implemented as follows: according to the coding rules of the traffic light control instruction, the binary bits in the instruction are segmented to extract the phase switching direction identifier and phase duration parameter. For example, the traffic light control instruction is 0000000111110001111000001010. According to the coding rules, the first 8 bits 00000001 represent the phase switching direction identifier, which is converted to the specific phase direction of the east-west direction; the following binary bits 11110001111000001010 represent the phase duration parameter, which is converted to a decimal number of 40 seconds. By parsing the traffic light control instruction, accurate phase switching direction and duration information can be obtained, providing a basis for subsequent phase switching operations.

[0198] Step S520: determining whether the next phase state is a reverse phase sequence according to the phase switching direction identifier.

[0199] A reverse phase sequence refers to a phase sequence that is opposite to or conflicts with the current phase state. For example, if the current phase state is a green light in the north-south direction and the next phase state is a green light in the east-west direction, then the east-west green light phase sequence is a reverse phase sequence relative to the north-south green light phase sequence.

[0200] Specifically, this determination can be implemented by comparing the phase switching direction indicator with the current signal light phase state to determine whether the next phase state is in the reverse phase sequence. For example, if the current phase state is a north-south green light and the phase switching direction indicator is a east-west green light, the next phase state is determined to be in the reverse phase sequence. By determining whether the next phase state is in the reverse phase sequence, appropriate measures can be taken in advance to avoid phase conflicts and traffic congestion.

[0201] Step S530: When a reverse phase sequence is detected, a phase conflict detection mechanism is started to verify whether there is an unfinished vehicle that conflicts with the reverse phase sequence.

[0202] The phase conflict detection mechanism is used to detect whether phase switching will cause conflicts. When a reverse phase sequence is detected, the mechanism is activated to verify whether there are any unfinished vehicles that conflict with the reverse phase sequence. Unfinished vehicles are vehicles that have not yet passed the intersection in the current phase state.

[0203] Specific detection methods, for example, can be implemented by obtaining vehicle trajectory information at the current intersection through video stream data and trajectory tracking results. The direction of travel in the reverse phase sequence is compared with the trajectory of vehicles that have not yet completed the process to determine whether there is a conflict. For example, if the reverse phase sequence is an east-west green light, the system detects whether there are vehicles in the north-south direction that have not yet completed the process and will conflict with vehicles traveling in the east-west direction when the east-west green light is on. If a conflict exists, the system flags the vehicle as having a phase conflict risk. By activating the phase conflict detection mechanism, potential phase conflict risks can be identified in advance, providing a basis for subsequent processing.

[0204] Step S540: Calculate the estimated clearing time of the conflicting vehicle and compare the estimated clearing time with the phase duration parameter.

[0205] Conflicting vehicles are vehicles that have not yet completed their journey and are in conflict with the reverse phase sequence. The estimated clearance time is the time required for all conflicting vehicles to pass through the intersection. Comparing the estimated clearance time with the phase duration parameter can determine whether there is sufficient time for conflicting vehicles to pass through the intersection, thus avoiding traffic conflicts caused by phase switching.

[0206] As an embodiment, step S540, calculating the estimated clearing time of the conflicting vehicles and comparing the estimated clearing time with the phase duration parameter, may include the following steps S541-S5411:

[0207] Step S541: extracting a position coordinate sequence from the historical trajectory segment of the conflicting vehicle, where the position coordinate sequence includes the coordinates of the vehicle center point in the past N consecutive frames, where N>1.

[0208] A historical trajectory segment refers to the movement trajectory of the conflicting vehicle over a period of time. A position coordinate sequence is a sequence of vehicle center point coordinates extracted from this historical trajectory segment. It contains the coordinates of the vehicle center point for the past N consecutive frames. By extracting the position coordinate sequence, we can obtain historical movement information of the conflicting vehicle, providing a basis for subsequent movement trend prediction.

[0209] Specifically, extraction can be implemented by, for example, obtaining the historical trajectory segments of the conflicting vehicles from the video stream data and trajectory tracking results. Then, the vehicle center coordinates for the past N consecutive frames are extracted from the historical trajectory segments to form a position coordinate sequence. For example, if N is 10, the vehicle center coordinates for the past 10 frames are extracted from the historical trajectory segments of the conflicting vehicles, resulting in the position coordinate sequence [(x1, y1), (x2, y2), ..., (x10, y10)]. By extracting the position coordinate sequence, the historical position information of the conflicting vehicles can be accurately recorded, providing data support for subsequent motion analysis.

[0210] Step S542: Input the position coordinate sequence into the spatiotemporal coupled encoder, extract the vehicle motion trend features through cyclic convolution in the time dimension, and capture the relative position relationship between the vehicle and the adjacent lanes through the attention mechanism in the spatial dimension.

[0211] The spatiotemporal coupled encoder is an encoder used to process spatiotemporal data, extracting features from both the temporal and spatial dimensions simultaneously. The temporal recurrent convolution is a convolution operation used to process time series data, capturing temporal dependencies and trends in vehicle motion. The spatial attention mechanism focuses on important parts of the data, capturing the relative position of the vehicle and adjacent lanes.

[0212] The specific processing can be implemented as follows: input the position coordinate sequence into the spatiotemporal coupling encoder. In the time dimension, the position coordinate sequence is processed using a recurrent convolution layer to extract the vehicle's motion trend features. For example, a recurrent neural network such as a long short-term memory network (LSTM) or a gated recurrent unit (GRU) is used to process the position coordinate sequence to obtain the vehicle's motion trend feature vector. In the spatial dimension, an attention mechanism is used to capture the relative position relationship between the vehicle and the adjacent lanes. The attention mechanism can calculate the attention weight of each position based on the vehicle's position coordinates and the position information of the adjacent lanes, and then perform a weighted summation of the vehicle's position information based on the attention weight to obtain a feature vector of the relative position relationship between the vehicle and the adjacent lanes. Through the processing of the spatiotemporal coupling encoder, the vehicle's motion trend features and the relative position relationship features with the adjacent lanes can be extracted simultaneously, providing more comprehensive information for subsequent trajectory prediction.

[0213] Step S543: Perform feature cross-fusion on the vehicle motion trend feature and the relative position relationship to generate a spatiotemporal coupling feature vector.

[0214] Feature cross-fusion combines and fuses vehicle motion trend features and relative position relationship features to generate a more expressive spatiotemporal coupling feature vector. This cross-fusion leverages the vehicle's motion trend and relative position relationship information with adjacent lanes, improving trajectory prediction accuracy.

[0215] Specifically, fusion can be implemented by concatenating or weighted summing the vehicle motion trend feature vector and the relative position relationship feature vector to generate a spatiotemporal coupling feature vector. For example, the vehicle motion trend feature vector and the relative position relationship feature vector are concatenated along the channel dimension to produce a feature vector with a larger dimension. This concatenated feature vector is then processed through a fully connected layer to further fuse the feature information and generate a spatiotemporal coupling feature vector. This cross-feature fusion allows different types of feature information to be integrated, providing a richer feature representation for subsequent trajectory prediction.

[0216] Step S544: Input the spatiotemporal coupling feature vector into the trajectory prediction decoder, and generate a predicted position coordinate sequence of the next M frames frame by frame through a deconvolution operation, where M≥1.

[0217] The trajectory prediction decoder is used to predict the vehicle's future trajectory based on the input feature vector. The deconvolution operation is used to recover spatial information from the feature vector. Through the deconvolution operation, a sequence of predicted position coordinates for the next M frames can be generated frame by frame.

[0218] The specific prediction can be implemented as follows: inputting the spatiotemporal coupling feature vector into the trajectory prediction decoder. The trajectory prediction decoder can adopt models such as deconvolution neural network (DeconvNet) or generative adversarial network (GAN). In the decoder, the spatiotemporal coupling feature vector is processed by deconvolution operation, and a sequence of predicted position coordinates for the next M frames is generated frame by frame. For example, M is 5, and through deconvolution operation, the predicted position coordinates of the next 5 frames [(x1', y1'), (x2', y2'), ..., (x5', y5')] are generated in sequence. Through the processing of the trajectory prediction decoder, the future motion trajectory of the conflicting vehicle can be predicted, providing a basis for the subsequent calculation of the estimated clearance time.

[0219] Step S545: Calculate the longitudinal distance between each frame position in the predicted position coordinate sequence and the stop line of the target lane, and generate a distance change curve.

[0220] The target lane stop line is the stop line position of the lane that the conflicting vehicle needs to reach. The longitudinal distance is the longitudinal distance between each frame position in the predicted position coordinate sequence and the target lane stop line. The distance change curve is a curve composed of the longitudinal distance between each frame position and the target lane stop line. It can reflect the change in the distance between the conflicting vehicle and the target lane stop line over time.

[0221] For example, the calculation can be implemented as follows: for each predicted position coordinate in the predicted position coordinate sequence, calculate its longitudinal distance from the stop line of the target lane. For example, if the longitudinal coordinate of the stop line of the target lane is y0 and the predicted position coordinate is (xi,yi), then the longitudinal distance is |yi-y0|. The longitudinal distances of each frame are arranged in chronological order to generate a distance change curve. By calculating the distance change curve, the changing trend of the distance between the conflicting vehicle and the stop line of the target lane can be visually observed, providing a basis for subsequent instantaneous speed feature extraction and estimated clearance time calculation.

[0222] Step S546: Perform a first-order derivative operation on the distance change curve to extract the instantaneous speed characteristics of the vehicle in each prediction frame.

[0223] The first-order derivative operation is to take the derivative of the distance change curve to obtain the slope of the curve. Through the first-order derivative operation, the instantaneous speed characteristics of the vehicle in each prediction frame can be extracted. The instantaneous speed characteristics represent the instantaneous movement speed of the vehicle in each prediction frame.

[0224] The specific operation can be implemented as follows: discretize the distance change curve and convert it into a set of discrete data points. Then, use numerical differentiation methods, such as forward difference, backward difference or central difference, to perform first-order derivative operations on the discrete data points. For example, using the central difference method, for the i-th data point, its derivative is approximately (yi+1-yi-1) / (2Δt), where yi+1 and yi-1 are adjacent data points and Δt is the time interval. The derivative of each data point is used as the instantaneous speed feature of the vehicle in the prediction frame. Through the first-order derivative operation, the instantaneous speed feature of the vehicle in each prediction frame can be accurately extracted, providing key information for the subsequent estimated clearance time calculation.

[0225] Step S547: Calculate the remaining time sequence required for the vehicle to reach the stop line based on the instantaneous speed characteristics and the current frame position coordinates.

[0226] The remaining time series refers to the time series required for the vehicle to reach the stop line of the target lane from the current frame position, which can be calculated based on the instantaneous speed characteristics and the current frame position coordinates.

[0227] The specific calculation can be implemented as follows: for each predicted frame's instantaneous speed feature and the current frame's position coordinates, calculate the remaining distance for the vehicle to reach the target lane's stop line. Then, divide the remaining distance by the instantaneous speed feature to obtain the remaining time required for the vehicle to reach the stop line. Arrange the remaining time of each predicted frame in order to generate a remaining time series. For example, if the current frame's position coordinates are (xi, yi), the vertical coordinate of the target lane's stop line is y0, and the instantaneous speed feature is vi, then the remaining distance is |yi-y0|, and the remaining time is |yi-y0| / vi. By calculating the remaining time series, the time it takes for the vehicle to reach the stop line can be predicted, providing a basis for the subsequent calculation of the estimated clearing time.

[0228] Step S548: Compare the minimum value in the remaining time series with the phase duration parameter, and mark it as a traffic conflict risk when the minimum value is less than the phase duration parameter.

[0229] The minimum value in the remaining time series represents the fastest time a vehicle would need to reach the stop line. Comparing this value with the phase duration parameter can determine whether there is sufficient time for the vehicle to pass through the intersection. If the minimum value is less than the phase duration parameter, it indicates that the vehicle may not be able to pass through the intersection within the current phase duration, posing a risk of traffic conflict.

[0230] For example, the comparison and marking process can be implemented by finding the minimum value in the remaining time series. This minimum value is then compared to the phase duration parameter. If the minimum value is less than the phase duration parameter, a traffic conflict risk is flagged as present; otherwise, no traffic conflict risk is flagged as present. For example, if the remaining time series is [5, 6, 7, 8, 9], the minimum value is 5 seconds, and the phase duration parameter is 4 seconds, then a traffic conflict risk is flagged as present. This comparison and marking process allows for the timely identification of potential traffic conflict risks, providing a basis for the subsequent generation of risk level indicator parameters and adjustment of estimated clearance times.

[0231] Step S549: associate and map the traffic conflict risk flag with the switching direction code of the reverse phase sequence to generate a risk level indication parameter.

[0232] Correlation mapping involves mapping the traffic conflict risk flag to the switching direction code of the reverse phase sequence to generate a risk level indicator parameter. The risk level indicator parameter is used to indicate the traffic conflict risk level under different switching directions.

[0233] The specific mapping and generation can be implemented as follows: establishing an associated mapping relationship based on the traffic conflict risk mark and the switching direction code of the reverse phase sequence. For example, a traffic conflict risk mark of 1 indicates that there is a traffic conflict risk, and 0 indicates that there is no traffic conflict risk; the switching direction code of the reverse phase sequence is 01, which indicates switching from the north-south direction to the east-west direction. The traffic conflict risk mark and the switching direction code are combined into a risk level indication parameter, such as (01,1) indicating that there is a traffic conflict risk when switching from the north-south direction to the east-west direction. By associating mapping and generating risk level indication parameters, the traffic conflict risk situation under different switching directions can be clearly indicated, providing a basis for the subsequent adjustment of the estimated clearing time calculation benchmark.

[0234] Step S5410: Adjust the calculation basis of the estimated clearing time based on the risk level indication parameter, and recalculate the remaining time series using the middle value of the predicted position coordinate series when the risk level exceeds the preset threshold.

[0235] The Estimated Clearance Time calculation basis refers to the underlying data and method used to calculate the Estimated Clearance Time. The Risk Level indicator parameter represents the level of conflict risk. When the risk level exceeds the preset threshold, indicating a high conflict risk, the Estimated Clearance Time calculation basis needs to be adjusted to more accurately calculate the Estimated Clearance Time.

[0236] Specific adjustments can be implemented, for example, as follows: determine whether the risk level exceeds a preset threshold based on the risk level indicator parameter. If the preset threshold is exceeded, the remaining time series is recalculated using the middle value of the predicted position coordinate sequence. For example, the predicted position coordinate sequence is [(x1, y1), (x2, y2), ..., (x5, y5)], and the middle value is (x3, y3). Based on the middle value, the remaining distance and remaining time series for the vehicle to reach the stop line are recalculated. By adjusting the calculation basis of the estimated clearing time, the estimated clearing time can be calculated more accurately when the risk of traffic conflict is high, avoiding traffic conflicts caused by phase switching.

[0237] Step S5411: Update the minimum value in the recalculated remaining time series as the final estimated clearing time, and write the final estimated clearing time into the phase conflict detection result.

[0238] The final estimated clearing time refers to the time required for all conflicting vehicles to pass through the intersection after adjustment and calculation. It is written into the phase conflict detection result to provide a basis for subsequent phase switching decisions.

[0239] The specific update and write can be implemented as follows: find the minimum value in the recalculated remaining time series and use it as the final estimated clearing time. The final estimated clearing time is added to the phase conflict detection result. The phase conflict detection result can be a record containing conflicting vehicle information, risk level indication parameters, and the final estimated clearing time. For example, the phase conflict detection result is [(vehicle ID: 123, risk level indication parameters: (01, 1), final estimated clearing time: 8 seconds)]. By updating the final estimated clearing time and writing the phase conflict detection result, accurate time information can be provided for subsequent phase switching operations to ensure safe and smooth traffic.

[0240] Step S550: When the estimated clearing time is less than the phase duration parameter, a switching operation of the reverse phase sequence is performed.

[0241] When the estimated clearing time is less than the phase duration parameter, it means that the conflicting vehicles have enough time to pass through the intersection within the current phase duration. At this time, the reverse phase sequence switching operation can be performed to adjust the phase state of the traffic light to meet the traffic needs of traffic participants.

[0242] The specific switching operation can be implemented as follows: controlling the signal light to switch phases based on the phase switching direction indicator and phase duration parameter in the signal light control instruction. For example, if the phase switching direction indicator is switching from north-south to east-west and the phase duration parameter is 30 seconds, the signal light will switch from a north-south green light to an east-west green light, and set the green light duration to 30 seconds. By performing the switching operation in the reverse phase sequence, the phase state of the signal light can be properly adjusted, improving traffic efficiency.

[0243] Step S560: When the estimated clearing time is greater than or equal to the phase duration parameter, the current phase state is maintained until a preset safety switching condition is met.

[0244] If the estimated clearing time is greater than or equal to the phase duration parameter, it means that not all conflicting vehicles can pass through the intersection within the current phase duration. In this case, the current phase state must be maintained until the preset safety switching conditions are met. The preset safety switching conditions can be that all conflicting vehicles have passed the intersection, the estimated clearing time is less than the phase duration parameter, etc.

[0245] Specifically, maintaining and waiting can be implemented as follows: continuing to maintain the current phase state of the traffic light while continuously monitoring the movement of conflicting vehicles and the estimated clearance time. When the preset safe switching conditions are met, the phase switching operation is performed. For example, if the current phase state is a green light in the north-south direction, the estimated clearance time is 40 seconds, and the phase duration parameter is 30 seconds, the north-south green light state will be maintained until all conflicting vehicles pass through the intersection or the estimated clearance time is less than the new phase duration parameter. By maintaining the current phase state until the safe switching conditions are met, traffic conflicts caused by phase switching can be avoided, ensuring safe and orderly traffic.

[0246] In summary, this edge computing traffic light control method based on traffic participant behavior analysis comprehensively analyzes and accurately judges the behavioral status of traffic participants, combines the phase status of traffic lights, generates reasonable traffic light control instructions, and realizes intelligent adjustment of traffic light phase switching timing, which can effectively reduce traffic conflicts and congestion, and improve traffic efficiency and safety.

[0247] It should be noted that, when reading the above embodiments of the invention, those skilled in the art can implement the unrefined technical details without obstacles based on their own technical knowledge in the field. For example, when it comes to the calculation of variables of different dimensions, those skilled in the art can use a general normalization or standardization method to eliminate the dimensional differences and then perform subsequent operations. For another example, for scenarios not involved, the technical means disclosed in the present invention can be used to continue adaptive extension. For example, when detecting the reverse phase sequence in step S530, if there is a delay between the estimated clearing time of the conflicting vehicle (step S540) and the phase duration parameter, a dynamic priority adjustment mechanism for real-time trajectory prediction can be introduced, combined with the low-latency characteristics of edge computing, to shorten the feedback cycle of prediction and decision-making. In step S350, the priority weights of the direction conflict indicator and the speed anomaly indicator depend only on the remaining time. A multi-factor weight distribution model can be introduced to dynamically adjust the priority based on the conflict type, traffic flow, historical accident data, etc. The models used can also be adaptively selected based on the specific usage scenario. For example, in step S231, in order to overcome trajectory jumps (such as the potential loss of valid data by removing abnormal segments in step S232), the bidirectional key point matching in occluded or dense scenes can be combined with object detection (such as YOLO) and optical flow methods to enhance key point correlation, and graph neural networks can be introduced to optimize trajectory continuity.

[0248] See also Figure 2 , Figure 2This is a schematic diagram of the structure of a signal light control system provided in an embodiment of the present invention. The signal light control system is, for example, an edge computer system, which includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, the communication interface 102, and the memory 103 may be connected via a bus or other means. The processor 101 (or central processing unit (CPU)) is the computing and control core of the signal light control system, which can parse various instructions within the signal light control system and process various data within the signal light control system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as Wi-Fi, a mobile communication interface, etc.), which can be used to send and receive data under the control of the processor 101. The communication interface 102 can also be used for data transmission and interaction within the signal light control system. The memory 103 (Memory) is a memory device in the signal light control system used to store programs and data. It is understood that the memory 103 here can include both the built-in memory of the signal light control system and the extended memory supported by the signal light control system. The memory 103 provides a storage space, which stores an operating system of the traffic light control system, and the present invention does not limit this.

[0249] In one embodiment, the processor 101 executes the edge computing traffic light control method based on traffic participant behavior analysis provided in the above embodiment of the present invention by running the computer program in the memory 103.

Claims

1. An edge computing traffic light control method based on traffic participant behavior analysis, characterized in that: The method comprises: Acquire video stream data of a target intersection, wherein the video stream data includes a motion trajectory of at least one traffic participant and a phase state of a current traffic light; Performing frame-by-frame behavior state analysis on the video stream data through a spatiotemporal feature coding network to extract behavior state features of each traffic participant, wherein the behavior state features include a motion direction vector, a speed change sequence, and a spatial position offset; Inputting the behavior state feature into a preset behavior matching model to generate a behavior triggering identifier of the traffic participant in the current time window, wherein the behavior triggering identifier is used to indicate whether the traffic participant meets a preset traffic rule conflict condition; Generate a signal light control instruction for an edge computing node according to a temporal correlation between the behavior trigger identifier and the phase state of the current signal light; The signal light phase switching sequence of the target intersection is adjusted based on the signal light control instruction, so that traffic participants who meet the traffic rule conflict condition obtain passage priority under the signal light phase switching sequence.

2. The edge computing traffic light control method based on traffic participant behavior analysis according to claim 1 is characterized in that: The process of performing frame-by-frame behavior state analysis on the video stream data through a spatiotemporal feature coding network to extract the behavior state features of each traffic participant includes: Constructing a three-dimensional space-time tensor of the video stream data, wherein the three-dimensional space-time tensor includes a time dimension, a spatial position dimension, and a pixel feature dimension; Extracting spatiotemporal features from the three-dimensional spatiotemporal tensor using a convolutional long short-term memory network to obtain a motion feature map of each traffic participant; Performing key point trajectory tracking on the motion feature map to generate the motion direction vector and the spatial position offset; Performing differential calculation on the motion feature maps of consecutive frames through a sliding window to generate the speed change sequence; The motion direction vector, the speed change sequence and the spatial position offset are normalized and spliced ​​to form the behavior state feature matrix.

3. The edge computing traffic light control method based on traffic participant behavior analysis according to claim 2 is characterized in that: The three-dimensional spatiotemporal tensor is subjected to spatiotemporal feature extraction by a convolutional long short-term memory network to obtain a motion feature map of each traffic participant, including: Splitting the three-dimensional spatiotemporal tensor into continuous time step slices according to the time dimension, each time step slice includes the spatial position dimension and pixel feature dimension of the current frame; Performing a three-dimensional convolution operation on each time step slice to generate a spatiotemporal convolution feature map of the current time step, wherein the convolution kernel of the three-dimensional convolution operation spans adjacent frames in the time dimension to capture motion continuity; Inputting the spatiotemporal convolution feature map into the forget gate structure of the convolutional long short-term memory network, generating a forget gate output by element-by-element multiplication and activation function processing to control the retention ratio of the previous hidden state; The spatiotemporal convolution feature map and the forget gate output are spliced ​​and input into the input gate structure, and the candidate cell state of the current time step is generated through convolution operation and activation function processing; Performing weighted superposition of the candidate cell state and the previous cell state adjusted by the forget gate to generate an updated current cell state; The current cell state is input into the output gate structure and processed by convolution operation and hyperbolic tangent activation function to generate the hidden state of the current time step; Traverse all time step slices, perform cross-layer feature fusion on the hidden state of each time step along the channel dimension, and generate a multi-scale motion feature map; Performing spatial pyramid pooling on the multi-scale motion feature map to extract local motion pattern features in different field of view ranges; The local motion pattern features are fused with the hidden state of the last time step using a channel-attention weighted fusion method to generate a motion feature map for each traffic participant, where each channel of the motion feature map corresponds to the activation intensity of a preset motion pattern.

4. The edge computing traffic light control method based on traffic participant behavior analysis according to claim 2 is characterized in that: The step of inputting the behavior state characteristics into a preset behavior matching model to generate a behavior triggering identifier of the traffic participant within the current time window includes: Calculating the direction angle between the motion direction vector in the behavior state feature and the set of allowed passing directions corresponding to the phase state of the current traffic light to generate a direction deviation angle sequence; Comparing the value of each time window in the speed change sequence with the speed threshold interval associated with the phase state, and generating a speed out-of-bounds mark sequence; Performing a sliding window accumulation of the direction deviation angle sequence in a time dimension, and generating a direction conflict flag when the accumulated angles of three consecutive windows exceed a preset deviation threshold; Performing continuous positive event detection on the speed violation mark sequence, and generating a speed abnormality mark when detecting that the speed exceeds the same direction threshold twice in a row; Inputting the direction conflict flag and the speed abnormality flag into a decision tree model associated with a phase state, and calculating the priority weight of the conflict behavior according to the remaining duration of the phase state; Dynamically integrating the direction conflict indicator and the speed abnormality indicator based on the priority weight to generate an initial behavior trigger probability; Aligning the initial behavior trigger probability with the trigger probability sequence of the historical time window, and generating a corrected behavior trigger probability through an exponential smoothing algorithm; When the corrected behavior trigger probability exceeds a dynamic trigger threshold that is negatively correlated with the current intersection traffic density, a valid behavior trigger flag is generated; The valid behavior trigger identifier and the remaining time of the phase state are subjected to logical constraint verification, and trigger events whose remaining time is less than the safety response threshold are eliminated to generate a final behavior trigger identifier set, in which each element contains a conflict type and a trigger timestamp.

5. The edge computing traffic light control method based on traffic participant behavior analysis according to claim 4 is characterized in that: The generating, according to the temporal correlation between the behavior trigger identifier and the phase state of the current signal light, a signal light control instruction of the edge computing node includes: Mapping and matching each behavior trigger identifier in the final behavior trigger identifier set with the remaining time of the phase state to generate a priority parameter set including a conflict type, a trigger timestamp, and a remaining time; activating a corresponding phase adjustment rule according to the conflict type in the priority parameter set, wherein the phase adjustment rule includes a direction conflict priority strategy and a speed abnormality degradation strategy; The direction conflict priority strategy is used to perform spatial orientation aggregation on multiple direction conflict identifiers in the same time window to generate a phase switching direction candidate set; Calculating the duration of the direction conflict based on the remaining time of the phase state and the trigger timestamp, and generating a phase adjustment parameter in combination with the frequency of occurrence of the speed anomaly indicator in the speed anomaly degradation strategy; Inputting the phase switching direction candidate set and the phase adjustment parameters into a time series optimizer to generate a phase switching timing plan that satisfies the minimum phase switching interval and maximum traffic efficiency constraints; Generate a phase switching direction code according to the difference between the target phase direction in the phase switching timing scheme and the current phase state; Binary concatenation of the phase switching direction code and the duration in the phase adjustment parameter to form a traffic light control instruction including a direction identification bit and a duration identification bit; Before generating the signal light control instruction, a conflict pre-check is performed to verify whether the phase switching direction code has a spatial overlap risk with the trajectory of an unfinished vehicle in the opposite lane; When a spatial overlap risk is detected, a mandatory waiting period flag is embedded in the signal light control instruction, wherein the mandatory waiting period flag extends the duration of the current phase state until the risk is resolved; The verified traffic light control instruction is encapsulated into an instruction frame structure executable by the edge computing node, and the instruction frame structure includes a timestamp synchronization field and an instruction effectiveness countdown field.

6. The edge computing traffic light control method based on traffic participant behavior analysis according to claim 1 is characterized in that: The adjusting the signal light phase switching timing of the target intersection based on the signal light control instruction includes: parsing the phase switching direction identifier and phase duration parameters in the signal light control instruction; determining whether the next phase state is a reverse phase sequence according to the phase switching direction identifier; When the reverse phase sequence is detected, a phase conflict detection mechanism is activated to verify whether there is an unfinished vehicle that conflicts with the reverse phase sequence; calculating an estimated clearing time for the conflicting vehicle and comparing the estimated clearing time with the phase duration parameter; When the estimated clearing time is less than the phase duration parameter, performing a switching operation of the reverse phase sequence; When the estimated clearing time is greater than or equal to the phase duration parameter, the current phase state is maintained until a preset safety switching condition is met.

7. The edge computing traffic light control method based on traffic participant behavior analysis according to claim 1 is characterized in that: The training method of the spatiotemporal feature coding network includes: Collecting sample video stream data of historical traffic scenes, and annotating the behaviors of traffic participants in the sample video stream data to generate a behavior label sequence; Constructing a hybrid architecture of a three-dimensional convolutional neural network and a recurrent neural network, wherein the hybrid architecture includes a spatiotemporal feature extraction branch and a behavior classification branch; A multi-scale feature fusion module is provided in the spatiotemporal feature extraction branch, wherein the multi-scale feature fusion module is used to integrate motion features under different receptive fields; Optimizing the feature discrimination of the spatiotemporal feature extraction branch by using a contrastive learning loss function, wherein the contrastive learning loss function causes feature vectors of similar behavior samples to cluster in the embedding space; A dynamic focus loss function is used in the behavior classification branch, and the dynamic focus loss function automatically adjusts the classification weight according to the category distribution of the samples; The output of the optimized spatiotemporal feature extraction branch is used as the input feature of the preset behavior matching model.

8. The edge computing traffic light control method based on traffic participant behavior analysis according to claim 7 is characterized in that: The optimization process of the contrastive learning loss function includes: Randomly selecting positive sample pairs and negative sample pairs from the sample video stream data, wherein the positive sample pairs contain different instances of the same behavior category, and the negative sample pairs contain instances of different behavior categories; generating a transformed view of the positive sample pair by data augmentation, wherein the transformed view includes spatial cropping, temporal slicing, and color jittering; Calculate the cosine similarity between the feature vectors of the positive sample pairs and construct the positive sample similarity distribution matrix; Calculate the cosine similarity between the feature vectors of the negative sample pairs and construct the negative sample similarity distribution matrix; Maximizing the separation between the positive sample similarity distribution matrix and the negative sample similarity distribution matrix by temperature scaling the cross entropy loss function; Back propagation adjusts the convolution kernel parameters of the spatiotemporal feature extraction branch and the gating weights of the recurrent neural network.

9. The edge computing traffic light control method based on traffic participant behavior analysis according to claim 5 is characterized in that: The calculating the estimated clearing time of the conflicting vehicle and comparing the estimated clearing time with the phase duration parameter comprises: Extracting a position coordinate sequence from a historical trajectory segment of the conflicting vehicle, wherein the position coordinate sequence includes the coordinates of the center point of the vehicle in the past N consecutive frames, where N>1; The position coordinate sequence is input into a spatiotemporal coupled encoder, which extracts the vehicle motion trend features through cyclic convolution in the time dimension, while capturing the relative position relationship between the vehicle and adjacent lanes through an attention mechanism in the spatial dimension; Cross-fusing the vehicle motion trend feature with the relative position relationship to generate a spatiotemporal coupling feature vector; Input the spatiotemporal coupling feature vector into the trajectory prediction decoder, and generate a predicted position coordinate sequence of the next M frames frame by frame through deconvolution operation, where M≥1; Calculating the longitudinal distance between each frame position in the predicted position coordinate sequence and the stop line of the target lane to generate a distance change curve; Performing a first-order derivative operation on the distance change curve to extract the instantaneous speed characteristics of the vehicle in each prediction frame; Calculating the remaining time sequence required for the vehicle to reach the stop line based on the instantaneous speed characteristics and the current frame position coordinates; Comparing the minimum value in the remaining time series with the phase duration parameter, and marking it as a traffic conflict risk when the minimum value is less than the phase duration parameter; Associating and mapping the passage conflict risk mark with the switching direction code of the reverse phase sequence to generate a risk level indication parameter; Adjusting the calculation basis of the estimated clearing time based on the risk level indicator parameter, and recalculating the remaining time series using the middle value of the predicted position coordinate series when the risk level exceeds a preset threshold; The minimum value in the recalculated remaining time series is updated as the final estimated clearing time, and the final estimated clearing time is written into the phase conflict detection result.

10. A signal light control system, characterized in that: include: a memory storing a computer program; A processor for loading the computer program to implement the edge computing traffic light control method based on traffic participant behavior analysis as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method, program product and system for determining state of traffic light

    CN114613129A

  • Test evaluation method for bus signal priority control

    CN115099599A

  • Traffic signal optimization method and system based on artificial intelligence

    CN119811077A

  • Traffic signal lamp control method and system based on edge calculation

    CN119889063A

  • Method and apparatus for controlling traffic light, method and apparatus for navigating unmanned vehicle, and method and apparatus for training model

    US20240071222A1

Cited By

  • Method and system for predicting red light running behavior of pedestrian in sparse scene

    CN121053612A

  • A method and system for predicting the behavior of pedestrians running red lights in sparse scenarios

    CN121053612B

  • Road traffic AI adaptive edge computing server

    CN121583119A

  • Traffic signal management and control method and device based on AI edge calculation

    CN121905005A

  • Traffic signal management and control method and device based on AI edge computing

    CN121905005B