Signal light duration estimation method based on physical information guided deep image prior

By dividing the signalized intersection into a spatiotemporal grid and generating a sparse queue contour image, and combining a depth image prior model and a physical loss function, the adaptability and accuracy of traffic light duration estimation under different scenarios are solved, and efficient traffic light duration estimation is achieved in the absence of historical annotation information.

CN121600496BActive Publication Date: 2026-05-01NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV
Filing Date
2026-01-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies are difficult to adapt to dynamic control environments where the period and phase can change in real time in fixed timing scenarios, and rely on a large amount of historical labeled data in dynamic timing scenarios, resulting in insufficient accuracy and generalization ability in traffic light duration estimation.

Method used

By dividing the signalized intersection into a spatiotemporal grid, a sparse queue contour image is generated. Then, by using a depth image prior model combined with reconstruction loss, traffic wave physical loss, and traffic light control loss, the complete queue contour is adaptively repaired to achieve traffic light status recognition.

Benefits of technology

Even in the absence of a large amount of historical annotation information, it can accurately estimate traffic light duration in both fixed and dynamic timing scenarios, improving adaptability and reliability, and enhancing the practicality and generalization ability of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600496B_ABST
    Figure CN121600496B_ABST
Patent Text Reader

Abstract

The application discloses a signal lamp duration estimation method based on physical information guided depth image prior, which comprises the following steps: constructing a space-time grid at a target entrance, mapping vehicle position and speed information into a sparse queue profile image; constructing a depth image prior model, taking the sparse image as input, iteratively optimizing through a composite loss of fusion reconstruction, traffic wave physical constraints and signal control information, and repairing to obtain a complete queue profile; identifying the signal lamp state based on the queue time sequence at the stop line, merging the continuous time period and calculating the red lamp, green lamp and cycle duration; the signal lamp duration estimation is reconstructed as a sparse queue profile image repair problem, combined with the depth image prior model guided by physical information, without relying on a large amount of historical labeled information, the cycle level signal lamp duration estimation can be realized in the fixed timing and dynamic timing scenes.
Need to check novelty before this filing date? Find Prior Art

Description

A method for estimating traffic light duration based on physical information-guided depth image priors Technical Field

[0001] This invention belongs to the field of traffic light duration estimation technology, and particularly relates to a traffic light duration estimation method based on physical information-guided depth image prior. Background Technology

[0002] Traffic light duration estimation is a core foundation of intelligent transportation systems, and its accuracy directly impacts the performance of key applications such as signal control optimization, traffic condition perception, eco-driving guidance, and dynamic route planning. Currently, with the increasing demands for intelligent transportation, higher requirements are being placed on high-precision, real-time, and universal traffic light duration estimation capabilities. However, in practical applications, due to inconsistent signal control system standards, limited data openness, and the complexity of dynamic control strategies, external systems struggle to directly obtain accurate and real-time signal timing information, thus limiting the effectiveness and widespread adoption of upper-level applications.

[0003] In existing technologies, for fixed timing scenarios, researchers typically assume that the timing scheme remains unchanged within a specific time period, and use historical observation data to fit or extract parameters such as period and phase. This approach is further extended to the identification and estimation of fixed schemes across multiple time periods. For dynamic timing scenarios, probabilistic statistical models are mainly used for offline estimation, or machine learning methods are employed to train supervised models using labeled historical signal states and timing data to predict period duration and phase structure.

[0004] However, the aforementioned existing technologies still have significant limitations. In fixed timing scenarios, the methods rely on the strong assumption that the timing remains unchanged over a certain period, making it difficult to adapt to dynamic control environments where the period and phase can change in real time, and lacking the ability to estimate single-period level changes. In dynamic timing scenarios, probabilistic statistical methods often rely on long-term stable statistical laws, making it difficult to cope with real-time changing signal strategies; while supervised learning methods, although possessing certain predictive capabilities, heavily rely on a large amount of accurate historical labeled data for model training, often facing problems such as a lack of labeled data, high acquisition costs, and insufficient scenario generalization ability in practical applications, making it difficult to achieve accurate and reliable periodic-level traffic light duration estimation under conditions of no labeling or scarce labeling. Summary of the Invention

[0005] Purpose of the invention: The purpose of this invention is to provide a traffic light duration estimation method based on physical information-guided depth image priors, which can achieve periodic traffic light duration estimation in both fixed and dynamic timing scenarios without relying on a large amount of historical annotation information.

[0006] Technical solution: The signal light duration estimation method based on physical information-guided depth image prior as described in this invention includes the following steps:

[0007] S1. Divide the target approach lane corresponding to the given signal phase in the signal-controlled intersection into multiple spatial units at equal intervals from upstream to the stop line along the vehicle travel direction, and discretize the study time interval into multiple moments to construct a spatiotemporal grid.

[0008] S2. Collect the position and speed information of vehicles on the target entrance lane, map the position and speed information to the spatiotemporal grid, binarize the vehicle status according to the preset speed threshold, generate queue contours, and obtain sparse queue contour images by combining data observability information.

[0009] S3. Construct a depth image prior model. Using the sparse queue contour image as input, iteratively optimize the depth image prior model through a composite loss function that includes reconstruction loss, traffic wave physical loss, and traffic light control loss to repair and obtain a complete queue contour image.

[0010] S4. Based on the time series changes of the spatial units corresponding to the stop line in the complete queue contour image, identify the traffic light status at each moment, merge the time periods of consecutive traffic light statuses with the same status, and calculate the corresponding red light duration, green light duration, and signal cycle duration.

[0011] This invention effectively captures vehicle state changes using limited observation data by dividing the spatiotemporal grid and generating sparse queue contour images. By leveraging a depth image prior model and combining reconstruction, traffic wave physics, and traffic light control loss, it can adaptively repair the complete queue contour even in the absence of a large amount of historical annotation information. Finally, based on the queue dynamics, it accurately identifies the traffic light state and achieves stable estimation of periodic red light duration, green light duration, and signal cycle in both fixed and dynamic timing scenarios, significantly improving the adaptability and estimation reliability for different signal control scenarios.

[0012] Preferably, step S2 includes:

[0013] In the spatiotemporal grid, the vehicle's operating state at spatial cell i and sampling time t is collected, and the corresponding average vehicle speed is recorded. ,in Indicates the spatial unit number, Indicates the sampling time number; based on a preset speed threshold. The average speed of the vehicle Binarization is performed to characterize the vehicle's driving or stationary state at the corresponding spatial cell and sampling time, thus constructing the vehicle state matrix. The vehicle state matrix Vehicle state variables corresponding to each spatial unit and each sampling time. Composition, the vehicle state variables Defined as:

[0014]

[0015] in, This indicates that the vehicles are stationary or in a queue. This indicates that the vehicle is in motion; based on the vehicle state matrix. Construct a binary queue profile to characterize the vehicle queuing evolution process at the target entrance lane.

[0016] By binarizing vehicle operating states based on preset speed thresholds, vehicle states are clearly divided into two categories: driving and queuing. This effectively constructs a binary queue outline that can intuitively reflect the evolution of vehicle queues. This process transforms complex speed information into a stable and easily parsed state matrix, significantly improving data interpretability and subsequent processing efficiency, and providing a reliable structured input for accurately depicting the dynamic changes of vehicle queues under different traffic conditions.

[0017] Preferably, in constructing the vehicle state matrix In the process, an observability matrix M corresponding one-to-one with the spatiotemporal grid is further constructed, wherein the observability matrix M consists of observability markers corresponding to each spatial unit and each sampling time. The components are used to indicate whether the vehicle state at the corresponding spatial unit and sampling time was actually collected; the observability markers The definition of is:

[0018]

[0019] in, This indicates that the vehicle state at the corresponding spatial unit and sampling time was actually collected. This indicates that the corresponding vehicle state has not been collected; based on the observability matrix M and the vehicle state matrix Construct the sparse queue contour matrix obtained from actual data collection. , is represented as:

[0020]

[0021] Among them, symbols This represents element-wise multiplication. This represents the observational disturbance term used to characterize vehicle detection errors, communication delays, or sensor noise.

[0022] By introducing an observability matrix to accurately identify the actual acquisition of vehicle state data in the spatiotemporal grid, and combining it with a sparse queue contour matrix generation mechanism that takes into account observational disturbances, the influence of factors such as real vehicle queue information, data loss, detection errors, and communication noise is effectively separated. This significantly enhances the robustness of contour data to complex real-world observation conditions and provides high-quality sparse input that truly reflects the reliability of the data for subsequent depth image prior restoration.

[0023] Preferably, the construction of the depth image prior model in step S3 includes:

[0024] Constructing a deep image prior generator network ,in The trainable parameters of the network are represented by the generator network, which employs a symmetric encoder-decoder structure and is composed of... Each downsampling stage and It consists of an upsampling stage that is symmetrical to it in terms of spatial scale; in the encoder, the first... The input features of each downsampling stage are represented as follows: Its number of channels Defined as:

[0025]

[0026] in, Basic number of channels, This is the upper limit of the channel;

[0027] Each downsampling stage includes a convolutional block with a stride of 1 that maintains the spatial resolution. And convolutional blocks with a stride of 2 that halve the spatial resolution and increase the number of channels :

[0028]

[0029]

[0030] in, This represents a convolution with a stride of 1 that maintains the spatial resolution. This represents a convolution with a stride of 2 that halves the spatial resolution, where k represents the kernel size, s represents the stride, BN represents batch normalization, and LeakyReLU represents the linear rectified unit activation function with leakage correction; it also increases the number of channels to... ;No. Each downsampling stage can be written as: In the decoder, each upsampling stage first performs a 2x nearest neighbor interpolation upsampling on the input features, followed by two convolutional blocks with a stride of 1. , Feature refinement and channel recovery:

[0031]

[0032] in, For the features corresponding to the decoding stage, Upsample represents the nearest neighbor interpolation upsampling operation; this continues until the highest resolution features are obtained. ;pass The convolutional layer maps the highest resolution features to a single-channel reconstruction result, and then obtains a normalized output through the Sigmoid function. :

[0033]

[0034] in, This represents a convolution operation with a kernel size of 1×1 and a stride of 1. This represents the Sigmoid activation function;

[0035] For the normalized output Threshold binarization is performed to obtain the repaired queue contour matrix. The input P of the generator network in It is composed of the sparse queue contour matrix P observed in practice and random noise δ with added regularization perturbation:

[0036]

[0037] in, and These represent the mean and standard deviation of the disturbance noise, respectively.

[0038] This generator network achieves multi-level effective extraction and fusion of high-dimensional spatiotemporal features through an encoder-decoder symmetric architecture and progressive downsampling and upsampling design. Combined with controllable noise perturbation input and a flexible channel capacity mechanism, it significantly enhances the model's ability to capture complex queue evolution patterns and its robustness. Finally, through refined feature reconstruction and threshold binarization, it can repair and output a clear and complete queue contour matrix with high quality, laying a reliable image foundation for accurate recognition of traffic light status.

[0039] Preferably, the composite loss function described in step S3 is composed of a weighted sum of three parts: reconstruction loss, traffic wave physical loss, and traffic light control loss, and its expression is:

[0040] For intersections using fixed timing, the total model loss is... Defined as:

[0041]

[0042] For intersections using dynamic timing, the total model loss Defined as:

[0043] .

[0044] in, To reconstruct the loss, Traffic wave physical losses and traffic light control losses include Duration range of loss For time consistency loss in fixed timing scenarios and Addressing the loss of time smoothness in dynamic timing scenarios; , , , , These are the non-negative weighting coefficients corresponding to each loss term.

[0045] This composite loss function integrates three types of constraints: reconstruction, traffic wave physics, and traffic light control. It achieves multi-dimensional physical guidance and rule embedding for the queue contour repair process. The reconstruction loss ensures the fidelity of the repair results to the observed data, the traffic wave physics loss ensures that the queue evolution conforms to the laws of traffic flow dynamics, and the traffic light control loss designed for different timing scenarios strengthens the periodic consistency or dynamic smoothness constraints, thereby significantly improving the model's adaptability and the physical rationality and logical coherence of the estimation results in both fixed and dynamic timing scenarios.

[0046] Preferably, the reconstruction loss The expression is:

[0047]

[0048] in, Indicates the total number of spatial units. Indicates the total number of time units. Represents the mask matrix In spacetime unit<j, t> The element at that location, with a value of 1 or 0, indicates whether the cell has been observed. The repaired queue contours, representing the output of the depth image prior network, are represented in the cells.<j, t> The state value at that location; This represents the original binary queue contour obtained based on actual observation data in the cell.<j, t> The state value at the location; ||·|| represents calculating the square of the difference between the two.

[0049] The reconstruction loss uses a mask matrix to precisely weight the observed units, ensuring that the repair contour strictly matches the actual data in the observed area, while providing flexible repair space for the unobserved area. This design effectively balances data fidelity and model generalization requirements, guiding the deep network to reasonably infer the complete and reliable queue evolution process while fully respecting the original observations.

[0050] Preferably, the traffic wave physical loss It consists of four parts: traffic wave propagation loss, waiting time loss, formation boundary loss, and dissipation boundary loss.

[0051]

[0052] in, , For traffic wave propagation loss, To compensate for the loss of waiting time, To form boundary loss, To dissipate boundary losses, the expression for the traffic wave propagation loss is:

[0053]

[0054]

[0055] in, Indicates the total number of spatial units; The total number of time units is represented by ; j represents the index of the spatial unit; t represents the index of the time unit. This indicates the repaired queue outline in the spacetime unit.<j, t> The state value at this location can be either 0 or 1. Indicates spatially adjacent downstream units<j+1, t> The state value at that location; Indicates the unit of time that is adjacent to the previous time.<j, t-1> The state value at the specified point; the max(0, ·) function represents the value when the calculated value within the parentheses is positive, otherwise it is 0; the expression for the waiting time loss is:

[0056]

[0057] The expression for the boundary loss is:

[0058]

[0059] in, It represents the propagation step size in the spatial dimension, determined based on the wave formation velocity; This represents the propagation step size in the time dimension, determined based on the wave formation velocity. This indicates that the same spatial location j is earlier in time. The state value at each moment; This indicates that it is spatially closer to the downstream than position j. Each unit, earlier in time The state value at each time step; the expression for the dissipation boundary loss is:

[0060]

[0061] in, This represents the propagation step size in the spatial dimension, determined based on the dissipation wave velocity. This represents the propagation step size in the time dimension, determined based on the dissipation wave velocity. This represents the state value at the same spatial location j, which is one time later in time. This indicates that it is spatially closer to the downstream than position j. Each unit, earlier in time The state value at each moment.

[0062] The traffic wave physical loss integrates multiple refined constraints such as propagation, waiting time, and formation and dissipation boundaries to comprehensively model the spatiotemporal evolution rules of vehicle states during queue formation and dissipation. This loss function effectively guides the repair profile to conform to the propagation characteristics and continuity requirements in traffic wave theory, significantly enhancing the physical consistency and rationality of queue dynamics, and ensuring that the generated complete profile can truly reflect the fluctuation patterns and boundary behaviors of traffic flow.

[0063] Preferably, the calculation of the traffic light control loss includes duration range loss, time consistency loss for fixed control scenarios, and time smoothness loss for dynamic control scenarios;

[0064] Among them, the time range loss Including the range of red light duration loss And the loss of green light duration :

[0065]

[0066]

[0067]

[0068] in, , These represent the estimated red light duration and green light duration for the nth cycle, respectively. , These represent the total number of red light cycles and the total number of green light cycles detected, respectively. , These are the preset lower and upper limits for the red light duration; , These are the preset reasonable lower and upper limits for green light duration; max(0, ·) represents the positive value function; the time consistency loss for fixed control scenarios... The calculation includes: for the nth period, constructing an eigenvector based on the estimated timing parameters. where norm(·) denotes normalization. The nth period represents the period length; a connectivity constraint matrix W is constructed to constrain clustering to not span time-discontinuous periods, where the elements... exist The value is 1 if the condition is met, and 0 otherwise; hierarchical clustering is performed based on the Ward connectivity criterion and connectivity constraints, and the inter-cluster distance is defined as:

[0069]

[0070] in and There are two clusters. and For each, the number of cycles it contains. and For each of their respective feature mean vectors, The L2 norm is represented; the optimal number of clusters is determined by the silhouette coefficient SI. ,in and The preset range for the number of clusters; The silhouette coefficient is used to calculate the consistency measure within each cluster. Clusters:

[0071]

[0072]

[0073] in, and Let the red light duration and green light duration be the values ​​for the nth cycle within the kth cluster. and The true average of the red and green light durations for all cycles within this cluster; the final time consistency loss is:

[0074]

[0075] in, Mean absolute deviation;

[0076] The time smoothness loss for dynamic control scenarios The calculations include: calculating the difference in duration between adjacent periods:

[0077]

[0078]

[0079] Where n = 1, 2, ..., N-1, and N is the total number of detected cycles. , Represent the difference in duration between adjacent cycles of red and green lights, respectively; calculate the time smoothness loss. :

[0080]

[0081]

[0082]

[0083] in, and These represent the smoothness loss for red light duration and the smoothness loss for green light duration, respectively. and These are the preset upper limits for the red light duration and green light duration, respectively.

[0084] The traffic light control loss achieves structured physical constraints on signal timing parameters through a multi-layered design of duration range loss, fixed-scene time consistency loss, and dynamic-scene time smoothness loss. Duration range loss ensures that the estimated traffic light duration is within a reasonable range, time consistency loss strengthens the periodic stability and regularity in fixed-time-matching scenarios through constrained clustering, and time smoothness loss effectively suppresses the drastic fluctuations between adjacent cycles in dynamic-time-matching scenarios. This improves the overall realism of the traffic light duration estimation results and the scenario adaptability under different control strategies.

[0085] Preferably, step S4 includes:

[0086] Based on the repaired complete queue contour image, extract the spatial units corresponding to the stop line. The state values ​​at each point constitute a time series data;

[0087] Map this time series data to the traffic light status at each time point, where, when When, it is mapped to a red light state; when When, it is mapped to a green light state, where For each spatial unit and each sampling time, there are vehicle state variables;

[0088] Merge traffic light states that are adjacent in time and have the same state to form multiple consecutive traffic light segments with the same state;

[0089] In the merged traffic light segments, complete red and green light segments are identified, and the corresponding red and green light durations are calculated based on the number of time units that each segment lasts.

[0090] By extracting the stop line state sequence from the complete queue profile and mapping it to the traffic light state, the correspondence between vehicle state and traffic light state is accurately identified. By merging adjacent identical states, continuous red and green light segments can be reliably reconstructed, thereby accurately calculating the actual duration of red and green lights in each signal cycle, ensuring the clarity, completeness and high reliability of the final traffic light duration estimation result.

[0091] Preferably, the present invention further includes an evaluation step:

[0092] The mean absolute error (MAE) and mean absolute percentage error (MAPE) are used as evaluation indicators.

[0093] For any given duration parameter type, including red light duration, green light duration, or total cycle duration, the MAE and MAPE are calculated as follows:

[0094] Let the actual duration of the nth evaluation sample be... The corresponding model estimate is The total number of samples evaluated is ,but:

[0095]

[0096]

[0097] The method for constructing the evaluation samples varies depending on the timing scenario:

[0098] In a fixed timing scenario, the signal period is first clustered to obtain K timing scheme categories; then, the actual average duration of each cluster category and the model-estimated average duration are used to form sample pairs. ,at this time Equal to the number of timing schemes obtained from clustering ;

[0099] In dynamic timing scenarios, the actual duration of each signal cycle and the model-estimated duration are directly used to form sample pairs. ,at this time It equals the total number of signal cycles.

[0100] This evaluation step, by designing differentiated sample construction strategies applicable to both fixed and dynamic timing scenarios, and combining mean absolute error and mean absolute percentage error as dual metrics, achieves a comprehensive quantitative evaluation of the accuracy and consistency of traffic light duration estimation results. This approach not only accurately measures the model's fine-grained estimation capability for each cycle in dynamic scenarios but also effectively evaluates its overall accuracy in capturing the overall timing scheme pattern in fixed scenarios. Thus, it provides a reliable and practically significant evaluation standard for the objective verification and comparison of method performance.

[0101] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: 1. This invention reconstructs traffic light duration estimation into a sparse queue contour image restoration problem. Combined with a depth image prior model guided by physical information, it can achieve periodic-level traffic light duration estimation in both fixed and dynamic timing scenarios without relying on a large amount of historical annotation information; 2. The model introduces physical regularization terms such as traffic wave propagation and waiting time into the loss function, and combines them with constraints on the duration range, consistency, or smoothness of traffic light control, making the estimation results more consistent with the actual traffic evolution and control logic, thus enhancing the interpretability of the method; 3. Based on sparse trajectory data such as floating cars, a binary queue contour is constructed, and the complete queuing process is restored through a physical-guided depth image prior, effectively overcoming data loss and noise interference. It exhibits high accuracy in estimating red light, green light, and cycle duration in both simulation and real-world scenarios; 4. By designing differentiated traffic light control loss terms, constraints are imposed on the time consistency of fixed timing and the time smoothness of dynamic timing, enabling the model to flexibly adapt to intersections with different control strategies, thus improving the practicality and generalization ability of the method. Attached Figure Description

[0102] Figure 1 is a schematic diagram of the method flow of the present invention;

[0103] Figure 2 is a queue outline diagram before repair under the fixed timing control scenario of the present invention (penetration rate 10%, sampling interval 10s).

[0104] Figure 3 shows the queue outline after repair under the fixed timing control scenario of the present invention (penetration rate 10%, sampling interval 10s).

[0105] Figure 4 shows the queue outline before repair under the dynamic timing control scenario of the present invention (penetration rate 10%, sampling interval 10s).

[0106] Figure 5 shows the queue outline after repair under the dynamic timing control scenario of the present invention (penetration rate 10%, sampling interval 10s). Detailed Implementation

[0107] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0108] This embodiment provides a traffic light duration estimation method based on physical information-guided depth image priors, as shown in Figure 1, including the following steps:

[0109] Step 1: Divide the target entrance lane into spatial units at equal intervals from upstream to the stop line, and discretize the study time period at equal intervals.

[0110] Consider a signalized intersection. For an approach lane with a given signal phase, discretize it along the travel direction as follows: Each spatial unit is numbered sequentially from upstream to downstream (stop line). The study time interval is discretized as follows: Each equally spaced sampling time is numbered as follows: .

[0111] Step 2: Collect vehicle speed and position data and map them to a spatiotemporal grid. Use binarization to mark whether the vehicle is stationary or in motion, and generate a sparse queue contour image.

[0112] For each spatiotemporal unit Record the average speed of the vehicle Based on the speed threshold The velocity can be binarized to characterize the vehicle's moving and stationary states:

[0113]

[0114] The binary matrix constructed from this This is denoted as the queue outline, used to characterize the queue evolution process at intersections. However, in practical applications, due to limitations such as floating car penetration rate, sampling frequency, and communication reliability, a large number of spatiotemporal units cannot be observed. Therefore, a mask matrix M is introduced to represent the observability of the data:

[0115]

[0116] The actual sparse queue outline collected It can be represented as:

[0117]

[0118] in Represents element-wise product. This is for potential observation noise.

[0119] Step 3: Construct a depth image prior (DIP) model, design a composite loss function that includes reconstruction loss, traffic wave physical loss and traffic light control loss, and iteratively optimize and repair the queue contour image.

[0120] Step 3.1: Construction of the Depth Image Prior Network

[0121] Constructing a Deep Image Prior (DIP) Generator Network It has a symmetrical encoder-decoder structure. The overall network consists of... Each downsampling stage and It consists of a symmetrical upsampling stage.

[0122] Let the encoder be the first Each downsampling stage The input features are The number of channels is defined as

[0123]

[0124] in Based on the number of channels, This represents the upper limit of the channel. Each downsampling stage consists of two concatenated convolutional blocks, defined with strides of 1 and 2 as follows:

[0125]

[0126]

[0127] in This represents a convolution with a stride of 1 that maintains the spatial resolution. This represents a convolution with a stride of 2 that halves the spatial resolution while increasing the number of channels. Therefore, the first Each downsampling stage can be written as:

[0128]

[0129] The decoder architecture is spatially symmetrical to the encoder. Two convolutional blocks are defined for the decoder. and Its structure is similar to that of the encoder. Both are the same, using convolutional blocks with a stride of 1, but the number of output channels of the convolutional layer is set to... This is done to gradually reduce the number of channels. Let the characteristic corresponding to the decoding stage be... Each upsampling stage first performs a 2x upsampling using nearest neighbor interpolation, followed by sequentially using... and Feature refinement and channel recovery:

[0130]

[0131] The final output layer adopts Convolution will extract the highest resolution features Project onto a single-channel reconstructed image and obtain a normalized output using the Sigmoid function:

[0132]

[0133] Subsequently Threshold binarization is performed to obtain the repaired queue outline.

[0134] Network input P in It consists of a fixed portion of observed data P and a portion of random noise δ with regularized perturbation:

[0135]

[0136] in, and These represent the mean and standard deviation of the disturbance noise, respectively.

[0137] Step 3.2, Reconstructing Loss

[0138] The DIP network learns its weight parameters by minimizing data consistency, which is defined as the reconstruction loss:

[0139]

[0140] This loss function calculates the mean square error only in the observed region, ensuring that the network output matches the true observations in these locations. Maintain consistency.

[0141] Step 3.3, Traffic Wave Physical Loss

[0142] Step 3.3.1, Traffic wave propagation loss: For any cell in the queue outline If the position is determined to be in a queuing state (i.e. If so, the queue must originate from its downstream position. Or the same position at the previous moment Furthermore, if queues exist downstream and in the previous moment, then queues will inevitably exist at the current location due to traffic flow propagation. Based on this principle of causality, the traffic wave propagation loss is defined as:

[0143]

[0144]

[0145] Step 3.3.2, Waiting Time Loss: Under signal control conditions, the waiting time of vehicles decreases along the direction away from the stop line. The waiting time loss is defined as:

[0146]

[0147] Step 3.3.3, Boundary Loss: During queue formation, the queue wave travels forward at a certain velocity. Propagation upstream from the stopping line. In a discrete-time grid, if the element... At any moment First time entering the queue Therefore, the queue should have expanded from the queue at an earlier time point, which was located further downstream. If this condition is violated, it means that the queue expansion rate exceeds the theoretically allowed maximum wave formation velocity. Therefore, the formation boundary loss is defined as...

[0148]

[0149] in and These represent the permissible propagation step lengths of the formed wave in the spatial and temporal dimensions, respectively, which can be determined based on the actual formed wave velocity. Perform calibration.

[0150] The queue dissipation phase corresponds to the release of queued vehicles during the green light period. The dissipation wave has a velocity... Propagation upstream from the stopping line. In a discrete-time grid, if the position... At any moment arrive Dissipation transformation occurs ( and If this dissipation occurs at an earlier time downstream, then the dissipation must have originated from the dissipation at a downstream location. Otherwise, it means the dissipation velocity exceeds the theoretically allowed maximum dissipation wave velocity. Accordingly, the dissipation boundary loss is defined as...

[0151]

[0152] in and These represent the allowable propagation steps of the dissipating wave in the spatial and temporal dimensions, respectively, which can be determined based on the actual dissipating wave velocity. set up.

[0153] Step 3.3.4, Traffic Wave Physical Loss: Combining the above-mentioned traffic wave propagation loss, waiting time loss, formation boundary loss, and dissipation boundary loss, the complete traffic wave physical loss function is expressed as:

[0154]

[0155] Step 3.4, Traffic Light Control Loss

[0156] Step 3.4.1, Duration Range Loss: Based on the repaired queue profile In the stop line unit The values ​​on the signal are smoothed and binarized to obtain the signal state sequence S:

[0157]

[0158] Based on this, the traffic light status can be merged into A complete cycle, and define the timing scheme set D:

[0159]

[0160] in, , , They represent the first The duration of red lights, green lights, and total cycle length for each cycle.

[0161] Define the reasonable duration range of traffic lights as follows: and The time-range loss can be expressed as:

[0162]

[0163]

[0164]

[0165] in and This represents the total number of traffic light cycles detected.

[0166] Step 3.4.2: For fixed control scenarios, design time consistency loss:

[0167] For the For each period, based on the estimated timing scheme parameters, a cluster feature vector is constructed and normalized using norm(·).

[0168]

[0169] Considering that the periods within each cluster are continuous in time, a clustering periodicity connectivity matrix is ​​constructed to ensure that clustering does not skip time steps.

[0170]

[0171] Based on the Ward connectivity criterion and connectivity constraints, the hierarchical clustering algorithm selects the periodic cluster pairs that satisfy the connectivity condition and have the smallest merging distance for merging in each iteration. The inter-cluster distance is defined as:

[0172]

[0173] in and There are two clusters. and For each, the number of cycles it contains. and Each element represents its feature mean vector. The optimal number of clusters can be determined by calculating the silhouette coefficient (SI).

[0174]

[0175] in and For the preset range of cluster numbers, This refers to the profile coefficient.

[0176] For the clustering of the th For each timing scheme, the mean absolute deviation (MAD) is used to measure the consistency of the duration of n cycles within the scheme:

[0177]

[0178]

[0179] in and N are respectively within this scheme k The true average of the red light duration and green light duration for each cycle. Ultimately, the time consistency loss is defined as:

[0180]

[0181] in, Mean absolute deviation;

[0182] Step 3.4.3: For dynamic control scenarios, design a time smoothness loss:

[0183] Assuming the detected The signal cycles are arranged in chronological order, with two adjacent cycles... The differences in red and green light durations between them are defined as follows:

[0184]

[0185]

[0186] Define reasonable upper limits for the duration of red and green lights as follows: and The time smoothness loss can be expressed as:

[0187]

[0188]

[0189]

[0190] Step 3.5, Total Model Loss

[0191] For intersections using fixed timing, the total model loss is... Defined as:

[0192]

[0193] For intersections using dynamic timing, the total model loss Defined as:

[0194]

[0195] Step 4: Based on the repair results, map the stop line position time sequence data to the traffic light status at each time, identify the red and green light segments by merging the status and calculate the duration.

[0196] To evaluate the model's performance, mean absolute error (MAE) and mean absolute percentage error (MAPE) are used as evaluation metrics. Let the th... The actual duration of each evaluation sample is The model estimate is ,total For each sample, then:

[0197]

[0198]

[0199] Let the red light duration, green light duration, and total cycle time be respectively set as follows: By taking the corresponding true and estimated values, we can obtain their respective MAE and MAPE. In a fixed timing scenario, the periods are first clustered, and then the true duration and estimated average duration of each cluster are used as sample pairs. ,at this time Equal to the number of timing schemes obtained from clustering In dynamic timing scenarios, the evaluation sample is taken as each signal cycle, and its actual duration and estimated duration directly constitute... , It equals the total number of cycles.

[0200] The numerical experiments in this embodiment used the VISSIM and SUMO simulation platforms to construct two types of single-intersection simulation scenarios: fixed-time and dynamic-time. In the fixed-time scenario, a three-lane approach road was set up, with a total phase flow of 1350 veh / h. Two representative timing schemes were considered: Scheme I, with a red light of 60 s and a green light of 40 s (cycle 100 s), and Scheme II, with a red light of 100 s and a green light of 60 s (cycle 160 s). Each scheme was simulated for 15 consecutive cycles, with a total simulation time of 3900 s. In the dynamic-time scenario, a two-lane approach road was set up, with the traffic flow exhibiting a process of first rising, then peaking, and then falling back: approximately 1440 veh / h at the beginning, approximately 1964 veh / h at the peak, and approximately 1448 veh / h at the end, for a total of 20 cycles, with a total simulation time of 2792 s, as shown in Table 1. The signal timing is adaptively adjusted according to the traffic flow, and a smoothing constraint is applied to adjacent cycles (the duration of red and green lights does not vary by more than 10 seconds).

[0201] Table 1. Traffic light duration configuration for each cycle in dynamic timing scenarios.

[0202]

[0203] Taking a penetration rate of 10% and a sampling interval of 10 seconds as an example, Figures 2 and 3 show the queue contour results before and after repair under a fixed timing control scenario, while Figures 4 and 5 show the queue contour results before and after repair under a dynamic timing control scenario. Comparing the results before and after repair shows that even under sparse data conditions, the model still has good repair capabilities and can generate queue contour shapes that conform to the physical laws of traffic waves.

[0204] Tables 2 and 3 present the model estimation performance evaluation results for fixed and dynamic timing scenarios, respectively, under the conditions of 10% penetration rate and a sampling interval of 10 seconds. In the fixed timing scenario, after three independent experiments, the mean absolute errors of cycle length, red light duration, and green light duration were all controlled within 1 second. In the dynamic timing scenario, the mean absolute errors of the above three indicators were 2.68 seconds, 2.80 seconds, and 3.18 seconds, respectively. The results show that the model exhibits good estimation accuracy in both timing scenarios.

[0205] Table 2. Performance Evaluation Results of Model Estimation in Fixed Timing Control Scenarios

[0206]

[0207] Table 3. Performance Evaluation Results of Model Estimation in Dynamic Timing Control Scenarios

[0208]

[0209] This invention proposes a depth image prior method based on physical regularization, reconstructing signal phase duration estimation from a traditional time-series problem into a sparse queue contour image inpainting problem, achieving self-supervised estimation. By incorporating traffic wave physical information and signal control rules into the loss function as regularization terms, the interpretability of the estimation results is improved. Simulation experiments show that, relying only on a small amount of sparse floating car data, the model can achieve accurate phase duration estimation in both fixed and dynamic timing scenarios. The above description is only a preferred embodiment of this invention. It should be noted that for those skilled in the art, several improvements and substitutions can be made without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this invention.

Claims

1. A method for estimating traffic light duration based on physical information-guided depth image priors, characterized in that, Includes the following steps: S1. Divide the target approach lane corresponding to the given signal phase in the signal-controlled intersection into multiple spatial units at equal intervals from upstream to the stop line along the vehicle travel direction, and discretize the study time interval into multiple moments to construct a spatiotemporal grid. S2. Collect the position and speed information of vehicles on the target entrance lane, map the position and speed information to the spatiotemporal grid, binarize the vehicle status according to a preset speed threshold, generate a queue outline, and obtain a sparse queue outline image by combining data observability information; S3. Construct a depth image prior model, using the sparse queue outline image as input, and iteratively optimize the depth image prior model through a composite loss function including reconstruction loss, traffic wave physical loss, and traffic light control loss to repair and obtain a complete queue outline image; S4. Based on the time series changes of the spatial units corresponding to the stop line in the complete queue outline image, identify the traffic light status at each moment, merge time periods with consecutive identical traffic light statuses, and calculate the corresponding red light duration, green light duration, and signal cycle duration; The construction of the depth image prior model in step S3 includes: constructing a depth image prior generator network. ,in The trainable parameters of the network are represented by the generator network, which employs a symmetric encoder-decoder structure and is composed of... Each downsampling stage and It consists of an upsampling stage that is symmetrical to it in terms of spatial scale; in the encoder, the first... The input features of each downsampling stage are represented as follows: Its number of channels Defined as: ;in, Basic number of channels, The upper limit of the channel; each downsampling stage includes a convolutional block with a stride of 1 that maintains the spatial resolution. And convolutional blocks with a stride of 2 that halve the spatial resolution and increase the number of channels : ; ;in, This represents a convolution with a stride of 1 that maintains the spatial resolution. This represents a convolution with a stride of 2 that halves the spatial resolution, where k represents the kernel size, s represents the stride, BN represents batch normalization, and LeakyReLU represents the linear rectified unit activation function with leakage correction; it also increases the number of channels to... ; the Each downsampling stage can be written as: In the decoder, each upsampling stage first performs a 2x nearest neighbor interpolation upsampling on the input features, followed by two convolutional blocks with a stride of 1. 、 Perform feature refinement and channel recovery: ;in, For the features corresponding to the decoding stage, Upsample represents the nearest neighbor interpolation upsampling operation; this continues until the highest resolution features are obtained. ;pass The convolutional layer maps the highest resolution features to a single-channel reconstruction result, and then obtains a normalized output through the Sigmoid function. : ;in, This represents a convolution operation with a kernel size of 1×1 and a stride of 1. Represents the Sigmoid activation function; for the normalized output Threshold binarization is performed to obtain the repaired queue contour matrix. The input P of the generator network in It is composed of the sparse queue contour matrix P observed in practice and random noise δ with added regularization perturbation: ;in, and These represent the mean and standard deviation of the disturbance noise, respectively.

2. The method according to claim 1, characterized in that, Step S2 includes: collecting the vehicle's operating status at spatial cell i and sampling time t within the spatiotemporal grid, and recording the corresponding average vehicle speed. ,in Indicates the spatial unit number, Indicates the sampling time number; based on a preset speed threshold. The average speed of the vehicle Binarization is performed to characterize the vehicle's driving or stationary state at the corresponding spatial cell and sampling time, thus constructing the vehicle state matrix. The vehicle state matrix Vehicle state variables corresponding to each spatial unit and each sampling time. Composition, the vehicle state variables Defined as: ;in, This indicates that the vehicles are stationary or in a queue. This indicates that the vehicle is in motion; based on the vehicle state matrix. Construct a binary queue profile to characterize the vehicle queuing evolution process at the target entrance lane.

3. The method according to claim 2, characterized in that, In constructing the vehicle state matrix In the process, an observability matrix M corresponding one-to-one with the spatiotemporal grid is further constructed, wherein the observability matrix M consists of observability markers corresponding to each spatial unit and each sampling time. The components are used to indicate whether the vehicle state at the corresponding spatial unit and sampling time was actually collected; the observability markers The definition of is: ;in, This indicates that the vehicle state at the corresponding spatial unit and sampling time was actually collected. This indicates that the corresponding vehicle state has not been collected; based on the observability matrix M and the vehicle state matrix Construct the sparse queue contour matrix obtained from actual data collection. , is represented as: Among them, the symbol This represents element-wise multiplication. This represents the observational disturbance term used to characterize vehicle detection errors, communication delays, or sensor noise.

4. The method according to claim 1, characterized in that, The composite loss function described in step S3 consists of a weighted sum of three parts: reconstruction loss, traffic wave physical loss, and traffic light control loss. Its expression is: For intersections using fixed timing, the total model loss is... Defined as: ; For intersections using dynamic timing, the total model loss Defined as: ;in, To reconstruct the loss, Traffic wave physical losses and traffic light control losses include Duration range of loss For time consistency loss in fixed timing scenarios and Addressing the loss of time smoothness in dynamic timing scenarios; 、 、 、 、 These are the non-negative weighting coefficients corresponding to each loss term.

5. The method according to claim 4, characterized in that, The reconstruction loss The expression is: ;in, Indicates the total number of spatial units. Indicates the total number of time units. Represents the mask matrix In spacetime unit<j, t> The element at that location, with a value of 1 or 0, indicates whether the cell has been observed. The repaired queue contours, representing the output of the depth image prior network, are represented in the cells.<j, t> The state value at that location; This represents the original binary queue profile obtained based on actual observation data in the cell.<j, t> The state value at the location; ||·|| represents calculating the square of the difference between the two.

6. The method according to claim 4, characterized in that, The physical loss of traffic waves It consists of four parts: traffic wave propagation loss, waiting time loss, formation boundary loss, and dissipation boundary loss. composition: ;in, 、 For traffic wave propagation loss, To compensate for the loss of waiting time, To form boundary loss, To dissipate boundary losses, the expression for the traffic wave propagation loss is: ; ;in, Indicates the total number of spatial units; The total number of time units is represented by ; j represents the index of the spatial unit; t represents the index of the time unit. This indicates the repaired queue outline in the spacetime unit.<j, t> The state value at this location can be either 0 or 1. Indicates spatially adjacent downstream units<j+1, t> The state value at that location; Indicates the unit of time that is adjacent to the previous time.<j, t-1> The state value at the specified point; the max(0, ·) function represents the value when the calculated value within the parentheses is positive, otherwise it is 0; the expression for the waiting time loss is: The expression for the boundary loss is: ;in, This represents the propagation step size in the spatial dimension, determined based on the wave formation velocity. This represents the propagation step size in the time dimension, determined based on the wave formation velocity. This indicates that the same spatial location j is earlier in time. The state value at each moment; This indicates that it is spatially closer to the downstream than position j. Each unit, earlier in time The state value at each time step; the expression for the dissipation boundary loss is: ;in, This represents the propagation step size in the spatial dimension, determined based on the dissipation wave velocity. This represents the propagation step size in the time dimension, determined based on the dissipation wave velocity. This represents the state value at the same spatial location j, which is one time later in time. This indicates that it is spatially closer to the downstream than position j. Each unit, earlier in time The state value at each moment.

7. The method according to claim 4, characterized in that, The calculation of the traffic light control loss includes duration range loss, time consistency loss for fixed control scenarios, and time smoothness loss for dynamic control scenarios; wherein, the duration range loss Including the range of red light duration loss And the loss of green light duration : ; ; ;in, 、 These represent the estimated red light duration and green light duration for the nth cycle, respectively. 、 These represent the total number of red light cycles and the total number of green light cycles detected, respectively. 、 These are the preset lower and upper limits for the red light duration; 、 These are the preset reasonable lower and upper limits for green light duration; max(0, ·) represents the positive value function; the time consistency loss for fixed control scenarios... The calculation includes: for the nth period, constructing an eigenvector based on the estimated timing parameters. where norm(·) denotes normalization. The nth period represents the period length; a connectivity constraint matrix W is constructed to constrain clustering to not span time-discontinuous periods, where the elements... exist The value is 1 if the condition is met, and 0 otherwise; hierarchical clustering is performed based on the Ward connectivity criterion and connectivity constraints, and the inter-cluster distance is defined as: ;in and There are two clusters. and For each, the number of cycles it contains. and For each of their respective feature mean vectors, The L2 norm is represented; the optimal number of clusters is determined by the silhouette coefficient SI. ,in and For the preset range of cluster numbers, The silhouette coefficient is used to calculate the consistency measure within each cluster. Clusters: ; ;in, and Let the red light duration and green light duration be the values ​​for the nth cycle within the kth cluster. and The true average of the red and green light durations for all cycles within this cluster; the final time consistency loss is: ;in, The mean absolute deviation; the time smoothness loss for dynamic control scenarios. The calculations include: calculating the difference in duration between adjacent periods: ; Where n = 1, 2, ..., N-1, and N is the total number of detected cycles. 、 Represent the difference in duration between adjacent cycles of red and green lights, respectively; calculate the time smoothness loss. : ; ; ;in, and These represent the smoothness loss for red light duration and the smoothness loss for green light duration, respectively. and These are the preset upper limits for the red light duration and green light duration, respectively.

8. The method according to claim 1, characterized in that, Step S4 includes: extracting the spatial units corresponding to the stop line based on the repaired complete queue contour image. The state values ​​at a given point constitute a time series data; this time series data is then mapped to the signal light state at each time step, where, when When, it is mapped to a red light state; when When, it is mapped to a green light state, where The system defines the vehicle state variables corresponding to each spatial unit and each sampling time. It merges temporally adjacent and identical traffic light states to form multiple continuous traffic light segments with the same state. Among the merged traffic light segments, it identifies complete red and green light segments and calculates the corresponding red and green light durations based on the number of time units each segment lasts.

9. The method according to claim 1, characterized in that, The evaluation process also includes the following steps: using Mean Absolute Error (MAE) and Mean Absolute Percentage Error (MAPE) as evaluation metrics; for any given duration parameter type, including red light duration, green light duration, or total cycle duration, the MAE and MAPE are calculated as follows: Let the actual duration of the nth evaluation sample be... The corresponding model estimate is The total number of samples evaluated is ,but: ; The method for constructing the evaluation samples varies depending on the timing scenario: In a fixed timing scenario, the signal period is first clustered to obtain K timing scheme categories; then, the actual average duration of each cluster category and the average duration estimated by the model constitute the sample pairs. ,at this time Equal to the number of timing schemes obtained from clustering In dynamic timing scenarios, the actual duration of each signal cycle and the model-estimated duration are directly used to form sample pairs. ,at this time It equals the total number of signal cycles.

Citation Information

Patent Citations

  • Emergency vehicle priority and queuing optimization cooperative control method

    CN121214701A

  • Monocular unsupervised depth estimation method based on contextual attention mechanism

    US20210390723A1