Traffic signal joint sensing methods, devices, electronic equipment and storage media

CN122024208BActive Publication Date: 2026-08-14IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

这种“分而治之”的方法存在显著局限:首先,它忽略了二者之间丰富的上下文关联,导致在复杂光照、部分遮挡、恶劣天气或目标距离较远等挑战性场景下,任一任务的感知不确定性会单独放大,系统缺乏利用另一任务提供的可靠线索进行交叉验证与辅助推理的能力,整体感知鲁棒性不足

Benefits of technology

[0017]本发明提供的交通信号联合感知方法、装置、电子设备及存储介质,先获取待检测交通信号图像序列;然后,将待检测交通信号图像序列输入至联合感知模型中,得到联合感知模型输出的交通信号灯检测结果和交通标志检测结果。其中,联合感知模型是基于样本交通信号图像序列、样本交通信号图像序列对应的样本交通信号灯检测结果和样本交通标志检测结果、预设语义关联信息以及预设状态转移先验信息训练得到的;联合感知模型用于基于交通信号灯增强特征向量、交通标志增强特征向量以及校验后的交通信号灯状态进行联合感知;其中,交通信号灯增强特征向量和交通标志增强特征向量,是基于待检测交通信号图像序列和预设语义关联信息确定得到的;校验后的交通信号灯状态,是基于预设状态转移先验信息进行时序一致性校验得到的。本发明通过构建联合感知模型,利用预设语义关联信息确定交通信号灯与交通标志的增强特征向量,实现了特征级的跨类别交叉验证与辅助增强,从而提升了复杂场景下的整体感知鲁棒性;同时,利用预设状态转移先验信息对交通信号灯状态进行时序一致性校验,有效过滤了违背物理常识的单帧感知跳变。本发明将语义交互特征与时序校验状态深度融合于单一模型中进行联合感知决策,不仅避免了多套独立模型并行带来的算力浪费,有效降低了系统的部署复杂度与计算开销,更从感知源头消除了独立输出引发的语义歧义,最终输出经过交叉验证与逻辑关联的高置信度检测结果,从而为下游的规划与控制模块提供更为准确、可靠、语义统一的交通信号灯与交通标志检测结果,提升了整个自动驾驶系统的性能与安全水平。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024208B_ABST
    Figure CN122024208B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, electronic device, and storage medium for joint perception of traffic signals, relating to the field of computer vision technology. The method includes: inputting a sequence of traffic signal images to be detected into a joint perception model to obtain detection results for traffic lights and traffic signs; the model is trained based on sample traffic signal image sequences, the detection results of sample traffic lights and traffic signs, preset semantic association information, and preset state transition prior information, and is used for joint perception based on enhanced feature vectors of traffic lights and traffic signs and verified traffic light states; wherein, the enhanced feature vectors of traffic lights and traffic signs are determined based on the sequence of traffic signal images to be detected and preset semantic association information; the verified traffic light states are obtained by performing temporal consistency verification based on preset state transition prior information. This invention can improve the accuracy and reliability of traffic signal detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method, apparatus, electronic device and storage medium for joint perception of traffic signals. Background Technology

[0002] In urban autonomous driving scenarios, traffic lights (also known as red and green lights) and traffic signs are the core basis for vehicles to make compliant passage decisions. Therefore, achieving accurate and stable detection of traffic lights and traffic signs is a very important task in autonomous driving scenarios, providing key information for subsequent safe and compliant driving. Specifically, traffic lights provide dynamic real-time passage instructions, such as red / yellow / green light state switching; while traffic signs provide static, long-term effective rule constraints, such as "stop and yield," "no left turn," and "speed limit 60." From a driving semantic perspective, the two do not exist in isolation, but rather have a natural strong coupling and mutual interpretation relationship in semantic logic, spatial layout, and temporal behavior. For example, the semantics of "red light" usually defines a complete stopping instruction together with "stop line" and "stop and yield" signs; the effective passage direction of "green light" needs to be specifically constrained by "straight / left turn arrows" or lane guidance signs; and the warning meaning of "yellow light" can be mutually reinforced with signs such as "caution pedestrians" and "yield."

[0003] However, existing autonomous driving perception systems generally treat traffic light recognition and traffic sign detection as two independent tasks, typically employing parallel or serial perception architectures. This "divide and conquer" approach has significant limitations: First, it ignores the rich contextual relationships between the two tasks, leading to amplified uncertainty in either task's perception under challenging scenarios such as complex lighting, partial occlusion, inclement weather, or long target distances. The system lacks the ability to utilize reliable cues from the other task for cross-validation and assisted reasoning, resulting in insufficient overall perception robustness. Second, when independent perception results are output to the downstream planning and control module, decision hesitation or conflict may arise due to time asynchrony, confidence conflicts, or semantic ambiguity (e.g., a green light is on but a "no entry" sign is detected simultaneously), introducing unnecessary safety risks and performance bottlenecks to the autonomous driving system. Furthermore, two independent models also introduce computational resource redundancy and increased system complexity.

[0004] Therefore, as advanced autonomous driving places increasing demands on the reliability, real-time performance, and decision-making friendliness of perception systems, the industry urgently needs a new perception framework that can explicitly model and utilize the multi-level and multi-dimensional contextual dependencies between traffic lights and traffic signs to improve the performance and safety of the entire autonomous driving system. Summary of the Invention

[0005] This invention provides a traffic signal joint perception method, device, electronic device, and storage medium, which deeply integrates semantic interaction features and temporal verification status into a single model for joint perception and decision-making. This provides downstream planning and control modules with more accurate, reliable, and semantically unified traffic light and traffic sign detection results, thereby improving the performance and safety level of the entire autonomous driving system.

[0006] This invention provides a traffic signal joint sensing method, comprising: Acquire the image sequence of the traffic signal to be detected; The traffic signal image sequence to be detected is input into the joint perception model to obtain the traffic light detection results and traffic sign detection results output by the joint perception model. The joint perception model is trained based on a sequence of sample traffic signal images, the detection results of sample traffic lights and sample traffic signs corresponding to the sequence of sample traffic signal images, preset semantic association information, and preset state transition prior information. The joint perception model is used to perform joint perception based on the enhanced feature vectors of traffic lights, enhanced feature vectors of traffic signs, and the verified state of traffic lights. The enhanced feature vectors of the traffic lights and the enhanced feature vectors of the traffic signs are determined based on the traffic signal image sequence to be detected and the preset semantic association information; the verified traffic light state is obtained by performing temporal consistency verification based on the preset state transition prior information.

[0007] According to a traffic signal joint sensing method provided by the present invention, the step of inputting the traffic signal image sequence to be detected into a joint sensing model to obtain the traffic light detection results and traffic sign detection results output by the joint sensing model includes: The feature extraction layer of the joint sensing model is used to extract multi-scale feature maps corresponding to each frame of the traffic signal image sequence to be detected. The feature map processing layer of the joint perception model processes the multi-scale feature map to obtain traffic light feature vectors and traffic sign feature vectors. Through the semantic interaction layer of the joint perception model, based on the preset semantic association information, semantic interaction enhancement processing is performed on the traffic light feature vector and the traffic sign feature vector respectively to obtain the traffic light enhanced feature vector and the traffic sign enhanced feature vector. Through the temporal verification layer of the joint sensing model, based on the preset state transition prior information, the temporal consistency verification of the predicted state of the traffic light is performed to obtain the verified state of the traffic light. The traffic light detection head of the joint perception model determines the traffic light detection result based on the traffic light enhanced feature vector and the verified traffic light state. Similarly, the traffic sign detection head of the joint perception model determines the traffic sign detection result based on the traffic sign enhanced feature vector.

[0008] According to a traffic signal joint sensing method provided by the present invention, the multi-scale feature map is processed through the feature map processing layer of the joint sensing model to obtain traffic light feature vectors and traffic sign feature vectors, including: The candidate box generation network shared in the feature map processing layer of the joint perception model is used to extract candidate boxes from the multi-scale feature map to obtain a traffic light candidate box set and a traffic sign candidate box set; wherein the traffic light candidate box set and the traffic sign candidate box set are generated based on their respective anchor box parameters. By performing region of interest alignment operations, traffic light feature vectors corresponding to each candidate box in the traffic light candidate box set and traffic sign feature vectors corresponding to each candidate box in the traffic sign candidate box set are extracted from the multi-scale feature map.

[0009] According to a traffic signal joint sensing method provided by the present invention, the method involves performing semantic interaction enhancement processing on the traffic light feature vector and the traffic sign feature vector based on the preset semantic association information through the semantic interaction layer of the joint sensing model, thereby obtaining enhanced traffic light feature vectors and enhanced traffic sign feature vectors, including: The semantic interaction layer of the joint perception model is used to obtain a semantic association matrix constructed based on the preset semantic association information. The semantic association matrix is ​​used to characterize the association strength between each traffic light state and each traffic sign category. Based on the category distribution probability of the traffic light candidate box set and the semantic association matrix, the first target semantic attention weight is calculated, and based on the category distribution probability of the traffic sign candidate box set and the semantic association matrix, the second target semantic attention weight is calculated. Based on the first target semantic attention weight, the traffic light feature vector processed by the network mapping is weighted and summed to obtain the first semantic context vector, and the first semantic context vector is fused with the traffic sign feature vector to obtain the traffic sign enhanced feature vector; Based on the second target semantic attention weight, the traffic sign feature vector processed by network mapping is weighted and summed to obtain the second semantic context vector, and the second semantic context vector is fused with the traffic light feature vector to obtain the traffic light enhanced feature vector.

[0010] According to a traffic signal joint sensing method provided by the present invention, the step of performing a temporal consistency verification on the predicted state of traffic lights based on the preset state transition prior information through the temporal verification layer of the joint sensing model to obtain the verified traffic light state includes: The state transition probability matrix is ​​obtained by using the temporal verification layer of the joint sensing model and constructing it based on the preset state transition prior information. The state transition probability matrix is ​​used to characterize the probability of different traffic light state transitions. Obtain a historical state memory queue for the same traffic light instance, wherein the historical state memory queue stores the historical state detection results of the same traffic light instance in historical traffic signal image frames; Based on the state transition probability matrix and the historical state memory queue, the temporal support score corresponding to the candidate state of the traffic light is calculated, and the original predicted state and its original state prediction distribution corresponding to the current traffic light image frame are obtained. The original state prediction distribution is weighted and fused with the temporal support score to obtain a smooth state confidence distribution. The verified traffic light state is determined based on the smooth state confidence distribution.

[0011] According to a traffic signal joint sensing method provided by the present invention, before processing the multi-scale feature map through the feature map processing layer of the joint sensing model to obtain the traffic light feature vector and the traffic sign feature vector, the method further includes: Based on the feature alignment layer of the joint sensing model, the target deep feature map in the multi-scale feature map is temporally aligned to obtain the aligned target deep feature map. The aligned target deep feature map and shallow feature map are fused to obtain an updated multi-scale feature map; wherein, the shallow feature map is the feature map in the multi-scale feature map other than the target deep feature map; The feature map processing layer of the joint perception model processes the multi-scale feature map to obtain traffic light feature vectors and traffic sign feature vectors, including: The updated multi-scale feature map is processed through the feature map processing layer of the joint perception model to obtain the traffic light feature vector and the traffic sign feature vector.

[0012] According to a traffic signal joint sensing method provided by the present invention, before determining the traffic signal detection result based on the traffic signal enhanced feature vector and the verified traffic signal state using a traffic signal detection head of the joint sensing model, and before determining the traffic sign detection result based on the traffic sign enhanced feature vector using a traffic sign detection head of the joint sensing model, the method further includes: The aligned target deep feature map is globally pooled through the weight adjustment layer of the joint perception model to obtain a scene-level semantic feature vector. Based on the scene-level semantic feature vector, calculate the first gating weight corresponding to the traffic light detection task and the second gating weight corresponding to the traffic sign detection task. Based on the first gating weight, the traffic light enhancement feature vector is weighted and modulated to obtain the modulated traffic light enhancement feature vector. Based on the second gating weight, the traffic sign enhancement feature vector is weighted and modulated to obtain the modulated traffic sign enhancement feature vector. The traffic light detection head using the joint perception model determines the traffic light detection result based on the enhanced feature vector of the traffic light and the verified state of the traffic light, and the traffic sign detection head using the joint perception model determines the traffic sign detection result based on the enhanced feature vector of the traffic sign, including: The traffic light detection head of the joint sensing model determines the traffic light detection result based on the modulated traffic light enhancement feature vector and the verified traffic light state. Similarly, the traffic sign detection head of the joint sensing model determines the traffic sign detection result based on the modulated traffic sign enhancement feature vector.

[0013] According to a traffic signal joint sensing method provided by the present invention, the joint sensing model is trained using a preset loss function, which is determined based on traffic light detection loss, traffic sign detection loss, semantic consistency loss and temporal smoothing loss. The semantic consistency loss is determined based on the preset semantic association information, the confidence level of the sample traffic light status prediction, and the confidence level of the sample traffic sign category prediction. The temporal smoothing loss is determined based on the preset state transition prior information and the sample state prediction distribution corresponding to adjacent image frames in the sample traffic signal image sequence.

[0014] The present invention also provides a traffic signal joint sensing device, comprising: The image sequence acquisition module is used to acquire the image sequence of the traffic signal to be detected; The detection result acquisition module is used to input the traffic signal image sequence to be detected into the joint perception model to obtain the traffic light detection result and traffic sign detection result output by the joint perception model; The joint perception model is trained based on a sequence of sample traffic signal images, the detection results of sample traffic lights and sample traffic signs corresponding to the sequence of sample traffic signal images, preset semantic association information, and preset state transition prior information. The joint perception model is used to perform joint perception based on the enhanced feature vectors of traffic lights, enhanced feature vectors of traffic signs, and the verified state of traffic lights. The enhanced feature vectors of the traffic lights and the enhanced feature vectors of the traffic signs are determined based on the traffic signal image sequence to be detected and the preset semantic association information; the verified traffic light state is obtained by performing temporal consistency verification based on the preset state transition prior information.

[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the traffic signal joint sensing method as described in any of the preceding claims.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the traffic signal joint sensing method as described in any of the preceding claims.

[0017] The present invention provides a traffic signal joint sensing method, apparatus, electronic device, and storage medium. First, a sequence of traffic signal images to be detected is acquired. Then, the sequence of traffic signal images is input into a joint sensing model to obtain the traffic light detection results and traffic sign detection results output by the joint sensing model. The joint sensing model is trained based on a sample traffic signal image sequence, the corresponding sample traffic light detection results and sample traffic sign detection results, preset semantic association information, and preset state transition prior information. The joint sensing model is used for joint sensing based on enhanced feature vectors of traffic lights, enhanced feature vectors of traffic signs, and verified traffic light states. The enhanced feature vectors of traffic lights and traffic signs are determined based on the traffic signal image sequence to be detected and preset semantic association information. The verified traffic light states are obtained by performing temporal consistency verification based on preset state transition prior information. This invention constructs a joint perception model and utilizes pre-defined semantic association information to determine enhanced feature vectors for traffic lights and signs. This achieves feature-level cross-category cross-validation and auxiliary enhancement, thereby improving the overall perception robustness in complex scenarios. Simultaneously, it uses pre-defined state transition prior information to perform temporal consistency verification of traffic light states, effectively filtering out single-frame perception jumps that violate physical common sense. This invention deeply integrates semantic interaction features and temporal verification states into a single model for joint perception decision-making. This not only avoids the computational waste caused by multiple independent models running in parallel, effectively reducing system deployment complexity and computational overhead, but also eliminates semantic ambiguity caused by independent outputs from the perception source. The final output is a high-confidence detection result that has undergone cross-validation and logical association, thus providing downstream planning and control modules with more accurate, reliable, and semantically consistent traffic light and sign detection results, improving the performance and safety level of the entire autonomous driving system. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 This is one of the flowcharts of the traffic signal joint sensing method provided by the present invention.

[0020] Figure 2 This is the second flowchart of the traffic signal joint sensing method provided by the present invention.

[0021] Figure 3This is the third flowchart of the traffic signal joint sensing method provided by the present invention.

[0022] Figure 4 This is the fourth flowchart of the traffic signal joint sensing method provided by the present invention.

[0023] Figure 5 This is the fifth flowchart of the traffic signal joint sensing method provided by the present invention.

[0024] Figure 6 This is a schematic diagram of the traffic signal joint sensing device provided by the present invention.

[0025] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0027] In urban autonomous driving scenarios, traffic lights and signs are the core basis for vehicles to make compliant passage decisions. Therefore, achieving accurate and stable detection of traffic lights and signs is a crucial task in autonomous driving scenarios, providing key information for subsequent safe and compliant driving. Specifically, traffic lights provide dynamic real-time passage instructions, such as red / yellow / green light switching; while traffic signs provide static, long-term effective rule constraints, such as "stop and yield," "no left turn," and "speed limit 60." From a driving semantic perspective, the two do not exist in isolation, but rather have a natural strong coupling and mutual interpretation relationship in semantic logic, spatial layout, and temporal behavior. For example, the semantics of a "red light" usually defines a complete stopping instruction together with "stop line" and "stop and yield" signs; the effective passage direction of a "green light" needs to be specifically constrained by "straight / left turn arrows" or lane guidance signs; and the warning meaning of a "yellow light" can be mutually reinforced with signs such as "caution pedestrians" and "yield."

[0028] Currently, in autonomous driving environmental perception systems for urban roads, the mainstream engineering implementations and academic research typically employ the following two technical solutions for processing static traffic signals, namely traffic lights and traffic signs.

[0029] The first approach is a dual-model parallel processing method, a relatively traditional and widely used deployment method. The system independently deploys and maintains two completely separate detection models: a high-performance general-purpose target detection model, specifically trained for detecting and recognizing traffic lights, whose output includes the location, color status (red, yellow, green), and arrow direction of the traffic lights; and another model specifically for detecting and classifying various traffic signs. Both models independently receive camera signals at the same time, perform forward inference in parallel, and after simple timestamp alignment, the results are directly output to the downstream fusion planning and decision-making module. This approach has significant drawbacks: first, the independent parameters of the two models result in high computational resource consumption and prevent computational reuse; second, complete "information isolation" between the models means that when one model experiences false detections or missed detections due to interference such as occlusion or backlighting, the other model cannot provide any effective contextual clues for correction, and the system's robustness depends entirely on the performance limit of a single model.

[0030] The second approach is a multi-task learning architecture with shared backbone features. To improve computational efficiency and implicitly promote feature learning, a shared backbone network is used, such as ResNet (Residual Network), EfficientNet, or Vision Transformer. This backbone network is responsible for extracting general, rich, multi-level feature maps from the input image. Based on these shared features, two structurally independent task-specific heads are connected: one responsible for traffic light detection and state classification, and the other for traffic sign detection and type recognition. Although this approach reduces computation through parameter sharing and may bring some positive transfer to low-level feature learning, its task interaction is limited to shallow or mid-level features of the backbone network. At the high-level semantic decoding and decision-making level, the two task heads still operate independently. The traffic light detection head cannot actively use the semantics of the "stop and yield" sign to confirm the credibility of the red light state, and the traffic sign detection head cannot prioritize "straight arrows" over "no left turn" signs based on the context of a "green light." This loose coupling method fails to model the deep semantic logic and spatial relationships between the two tasks. It is essentially an improvement on "feature parallelization and task divide-and-conquer" and cannot achieve true collaborative reasoning.

[0031] In summary, existing autonomous driving perception systems generally treat traffic light recognition and traffic sign detection as two independent tasks, typically employing parallel or serial perception architectures. This "divide and conquer" approach has significant limitations: First, it ignores the rich contextual relationships between the two tasks, leading to amplified uncertainty in either task under challenging scenarios such as complex lighting, partial occlusion, inclement weather, or long target distances. The system lacks the ability to utilize reliable cues from the other task for cross-validation and assisted reasoning, resulting in insufficient overall perception robustness. Second, when independent perception results are output to the downstream planning and control module, decision hesitation or conflict may arise due to time asynchrony, confidence conflicts, or semantic ambiguity (e.g., a green light is on but a "no entry" sign is detected simultaneously), introducing unnecessary safety risks and performance bottlenecks to the autonomous driving system. Furthermore, two independent models also introduce computational resource redundancy and increased system complexity.

[0032] Therefore, existing technical solutions have significant bottlenecks in the reliability, consistency, and security of traffic signal perception results when dealing with long-tail scenarios such as complex urban intersections, severe weather, and extreme lighting. As advanced autonomous driving places increasing demands on the reliability, real-time performance, and decision-making friendliness of perception systems, there is an urgent need for a solution that can achieve deep collaborative perception of traffic lights and traffic signs.

[0033] Based on the above analysis, this invention proposes a traffic signal joint sensing method, device, electronic device, and storage medium, which are described below in conjunction with... Figures 1-7 Describe it.

[0034] Figure 1 This is one of the flowcharts of the traffic signal joint sensing method provided by the present invention, such as... Figure 1 As shown, the traffic signal joint sensing method includes steps S110 and S120.

[0035] Step S110: Obtain the image sequence of the traffic signal to be detected.

[0036] First, the image sequence of the traffic signal to be detected is acquired. In autonomous driving scenarios, this sequence typically consists of T consecutive frames of images captured in real time by the vehicle's forward-facing camera.

[0037] It should be noted that, considering the challenges of occlusion, motion blur, or difficulty in identifying extremely small targets in autonomous driving, this embodiment of the invention acquires multiple consecutive frames as a sequence of traffic signal images to be detected, which then serves as input to the model. By introducing temporal information in this way, the detection capability of the current frame can be enhanced by utilizing the context of preceding and following frames.

[0038] Step S120: Input the traffic signal image sequence to be detected into the joint perception model to obtain the traffic light detection results and traffic sign detection results output by the joint perception model.

[0039] The joint perception model is trained based on a sequence of sample traffic signal images, the detection results of sample traffic lights and sample traffic signs corresponding to the sequence of sample traffic signal images, preset semantic association information, and preset state transition prior information. The joint perception model is used to perform joint perception based on the enhanced feature vectors of traffic lights, enhanced feature vectors of traffic signs, and the verified state of traffic lights. The traffic light enhancement feature vector and the traffic sign enhancement feature vector are determined based on the traffic signal image sequence to be detected and the preset semantic association information; the verified traffic light state is obtained by performing temporal consistency verification based on the preset state transition prior information.

[0040] The training process of the joint perception model is as follows: First, a sample set containing a large number of real driving scenarios is constructed, namely, the sample traffic signal image sequence, and it is manually labeled to obtain the corresponding sample traffic light detection results (such as the border and category of traffic lights) and sample traffic sign detection results (such as the border and category of speed limit signs and yield signs). At the same time, two important physical and common sense rules are injected into the model: (1) Pre-set semantic association information, which refers to the logical co-occurrence rules between different traffic signals (including traffic lights and traffic signs) defined by traffic rules in the physical world that are collected in advance. For example, the "red light" and the "stop and yield" sign have a high semantic association. The specific implementation of this pre-set semantic association information can be a semantic association matrix, which transforms this human logic into a quantified weight value. (2) Pre-set state transition prior information, which refers to the physical switching rules followed by traffic lights in the time dimension that are collected in advance. For example, "green light → yellow light → red light" is a legal transition, while "green → yellow → green" (no red light transition) and "red → green" (skipping the yellow light) are state transitions that violate traffic rules. The specific implementation of this pre-defined state transition prior information can be a state transition probability matrix, which encodes the legality probability of each type of state transition. Then, based on the aforementioned sample set, the initial joint sensing model is trained to obtain the trained joint sensing pattern.

[0041] Furthermore, during model training, in addition to calculating the basic detection error, semantic consistency loss and temporal smoothing loss are also calculated based on the above two types of information, forcing the model to learn feature expressions that conform to physical and common sense logic, and finally obtaining a well-trained joint perception model.

[0042] Furthermore, after training the joint perception model, it is deployed on the onboard computing platform of the autonomous vehicle.

[0043] After obtaining the traffic signal image sequence to be detected, the traffic signal image sequence to be detected is input into the joint sensing model.

[0044] The joint perception model first performs cross-category interactive enhancement processing on the feature dimension based on the sequence of traffic signal images to be detected and pre-defined semantic association information, resulting in enhanced feature vectors for traffic lights and traffic signs. Specifically, it first extracts feature vectors for traffic lights and traffic signs from the sequence of traffic signal images to be detected. Then, based on pre-defined semantic association information, it performs semantic interactive enhancement processing on the feature vectors for traffic lights and traffic signs respectively, resulting in enhanced feature vectors for traffic lights and traffic signs.

[0045] Simultaneously, a temporal consistency check is performed based on preset state transition prior information to obtain the checked traffic light state. Specifically, the original predicted state and its original state prediction distribution are extracted from the current traffic light image frame in the traffic light image sequence to be detected, combined with the state data in the historical state memory queue, and substituted into the preset state transition prior information for temporal consistency check, to obtain a smooth and physically consistent checked traffic light state.

[0046] Finally, joint perception is performed based on the enhanced feature vectors of traffic lights, enhanced feature vectors of traffic signs, and the verified state of traffic lights. The final output is the detection results of traffic lights and traffic signs for the current traffic signal image frame. The traffic light detection results include the state / category (e.g., red, yellow, green), location (e.g., bounding box), and confidence level of the traffic lights. The traffic sign detection results include the category, location (e.g., bounding box), and confidence level of the traffic signs.

[0047] The traffic signal joint sensing method provided in this invention first acquires a sequence of traffic signal images to be detected; then, the sequence of traffic signal images to be detected is input into a joint sensing model to obtain the traffic light detection results and traffic sign detection results output by the joint sensing model. The joint sensing model is trained based on a sample traffic signal image sequence, the corresponding sample traffic light detection results and sample traffic sign detection results, preset semantic association information, and preset state transition prior information. The joint sensing model is used for joint sensing based on enhanced feature vectors of traffic lights, enhanced feature vectors of traffic signs, and verified traffic light states. The enhanced feature vectors of traffic lights and traffic signs are determined based on the sequence of traffic signal images to be detected and preset semantic association information; the verified traffic light states are obtained by performing temporal consistency verification based on preset state transition prior information. This invention constructs a joint perception model and utilizes pre-defined semantic association information to determine enhanced feature vectors for traffic lights and signs. This achieves feature-level cross-category cross-validation and auxiliary enhancement, thereby improving the overall perception robustness in complex scenarios. Simultaneously, it uses pre-defined state transition prior information to perform temporal consistency verification of traffic light states, effectively filtering out single-frame perception jumps that violate physical common sense. This invention deeply integrates semantic interaction features and temporal verification states into a single model for joint perception decision-making. This not only avoids the computational waste caused by multiple independent models running in parallel, effectively reducing system deployment complexity and computational overhead, but also eliminates semantic ambiguity caused by independent outputs from the perception source. The final output is a high-confidence detection result that has undergone cross-validation and logical association, thus providing downstream planning and control modules with more accurate, reliable, and semantically consistent traffic light and sign detection results, improving the performance and safety level of the entire autonomous driving system.

[0048] Based on any of the above embodiments Figure 2 This is the second flowchart of the traffic signal joint sensing method provided by the present invention, as shown below. Figure 2 As shown, step S120 includes: step S121, step S122, step S123, step S124 and step S125.

[0049] Step S121: Extract multi-scale feature maps corresponding to each frame of the traffic signal image sequence to be detected through the feature extraction layer of the joint sensing model.

[0050] The sequence of traffic signal images to be detected is input into the feature extraction layer of the joint sensing model. Through the feature extraction layer of the joint sensing model, multi-scale feature maps corresponding to each frame of the sequence of traffic signal images to be detected are extracted.

[0051] Assuming the input is continuous TFrame traffic signal image That is, using the current frame (the first frame) T (Frame) Traffic signal image and its preceding frames T-1 Frames of historical traffic signal images constitute a total T A sequence of traffic signal images of frames, denoted as the traffic signal image sequence to be detected, wherein, R Represents the set of real numbers. H Indicates the height of the image. W This indicates the width of the image. Specifically, this invention does not limit the number of consecutive frames. T and image resolution .

[0052] The feature extraction layer is implemented using a shared backbone network. The structure of the backbone network is not specifically limited and can be a network architecture such as ResNet or Swin Transformer (Shifted Window Transformer). Each frame of the traffic signal image is processed by the feature extraction layer, and the corresponding output for each frame is given. M Multi-scale feature map of layers, denoted as .in, M Generally, integers between 3 and 5 are used for deep features. It contains high-level semantic information, while shallow features This preserves richer spatial information, and this multi-layer feature output lays the foundation for subsequent detection of targets of different sizes.

[0053] Step S122: The multi-scale feature map is processed through the feature map processing layer of the joint perception model to obtain the traffic light feature vector and the traffic sign feature vector.

[0054] Multi-scale feature maps are input into the feature map processing layer of the joint perception model. The feature map processing layer of the joint perception model performs candidate box extraction on the multi-scale feature maps to obtain the candidate box sets for traffic lights and traffic signs.

[0055] Then, through region of interest alignment, traffic light feature vectors corresponding to each candidate box in the traffic light candidate box set are extracted from the multi-scale feature map. At the same time, traffic sign feature vectors corresponding to each candidate box in the traffic sign candidate box set are extracted.

[0056] Step S123: Through the semantic interaction layer of the joint perception model, based on the preset semantic association information, semantic interaction enhancement processing is performed on the traffic light feature vector and the traffic sign feature vector respectively to obtain the traffic light enhanced feature vector and the traffic sign enhanced feature vector.

[0057] Pre-defined semantic association information refers to the logical co-occurrence patterns among different traffic signals (including traffic lights and traffic signs) defined by traffic rules in the physical world, which are collected in advance. For example, the "red light" and the "stop and yield" sign have a high semantic association.

[0058] The feature vectors of traffic lights and traffic signs are input into the semantic interaction layer of the joint perception model. Through this semantic interaction layer, based on the preset semantic association information, the feature vectors of traffic lights are subjected to semantic interaction enhancement processing to obtain enhanced feature vectors of traffic lights; at the same time, the feature vectors of traffic signs are subjected to semantic interaction enhancement processing to obtain enhanced feature vectors of traffic signs.

[0059] Specifically, through the semantic interaction layer of the joint perception model, a semantic association matrix constructed based on preset semantic association information is obtained. This semantic association matrix characterizes the association strength between the states of each traffic light and the categories of each traffic sign. Based on the category distribution probability of the traffic light candidate box set and the semantic association matrix, a first target semantic attention weight is calculated. Similarly, based on the category distribution probability of the traffic sign candidate box set and the semantic association matrix, a second target semantic attention weight is calculated. Based on the first target semantic attention weight, the traffic light feature vectors processed by the network mapping are weighted and summed to obtain a first semantic context vector. This first semantic context vector is then fused with the traffic sign feature vectors to obtain an enhanced traffic sign feature vector. Based on the second target semantic attention weight, the traffic sign feature vectors processed by the network mapping are weighted and summed to obtain a second semantic context vector. This second semantic context vector is then fused with the traffic light feature vectors to obtain an enhanced traffic light feature vector. The specific execution process can be found in the following embodiment, which will not be elaborated here.

[0060] Step S124: Through the temporal verification layer of the joint sensing model, based on the preset state transition prior information, the predicted state of the traffic light is verified for temporal consistency, and the verified state of the traffic light is obtained.

[0061] Preset state transition prior information refers to the physical switching rules that traffic lights follow in the time dimension, which are obtained in advance. For example, "green light → yellow light → red light" is a legal transition, while "green → yellow → green" (no red light transition) and "red → green" (skipping the yellow light) are state transitions that violate traffic rules.

[0062] The sequence of traffic signal images to be detected and the previously extracted set of traffic light candidate boxes are input into the temporal verification layer of the joint sensing model. Based on preset state transition prior information, the temporal verification layer performs temporal consistency verification on the predicted state of the traffic lights, resulting in the verified traffic light state. Specifically, the predicted state of the traffic lights can be the original predicted state corresponding to the current traffic signal image frame (i.e., the last frame of the sequence) in the sequence of traffic signal images to be detected.

[0063] Specifically, through the temporal verification layer of the joint sensing model, a state transition probability matrix constructed based on preset state transition prior information is obtained. This state transition probability matrix represents the probability of state transitions for different traffic lights. A historical state memory queue for the same traffic light instance is obtained, storing the historical state detection results of the same traffic light instance in historical traffic signal image frames. Based on the state transition probability matrix and the historical state memory queue, the temporal support score corresponding to the candidate state of the traffic light is calculated, and the original predicted state and its original state prediction distribution corresponding to the current traffic signal image frame are obtained. The original state prediction distribution and the temporal support score are weighted and fused to obtain a smoothed state confidence distribution. Based on the smoothed state confidence distribution, the verified traffic light state is determined. The specific execution process can be referred to in the following embodiment, which will not be elaborated here.

[0064] Step S125: Using the traffic light detection head of the joint perception model, the traffic light detection result is determined based on the traffic light enhanced feature vector and the verified traffic light state. Similarly, using the traffic sign detection head of the joint perception model, the traffic sign detection result is determined based on the traffic sign enhanced feature vector.

[0065] The traffic sign enhancement feature vector is input into the traffic sign detection head of the joint perception model, and the traffic sign detection result is output. The traffic light enhancement feature vector and the corrected traffic light state are input into the traffic light detection head of the joint perception model, and the traffic light detection result is output.

[0066] The traffic signal joint sensing method provided in this invention decouples the complex joint sensing task into well-defined hierarchical components. Low-level feature sharing significantly reduces computational overhead; while the collaborative work of the upper-level semantic interaction layer and temporal verification layer achieves dual verification of spatial logic and temporal physics, effectively filtering out various false positives (false detections) and false negatives (missed detections), thereby improving the accuracy of the detection results.

[0067] Based on any of the above embodiments, step S122 includes: step S1221 and step S1222.

[0068] Step S1221: The candidate box generation network shared in the feature map processing layer of the joint perception model is used to perform candidate box extraction processing on the multi-scale feature map to obtain a traffic light candidate box set and a traffic sign candidate box set; wherein the traffic light candidate box set and the traffic sign candidate box set are generated based on the corresponding anchor box parameters.

[0069] To better meet the real-time efficiency requirements of autonomous driving scenarios, this embodiment of the invention abandons the traditional approach of generating candidate boxes for traffic lights and traffic signs independently. Instead, it adopts a shared lightweight candidate box generation network to generate candidate boxes for both types of targets simultaneously based on unified features, which significantly reduces inference latency.

[0070] Furthermore, in order to adapt to the significant differences in shape and size between traffic lights and traffic signs, in this embodiment of the invention, two sets of anchor frame parameters with different scales and aspect ratios are preset for the RPN (Region Proposal Network): one set is optimized for the typical approximate rectangular or circular outline of traffic lights, and the other set is optimized for the common rectangular (such as square) or triangular shapes of traffic signs.

[0071] The multi-scale feature maps are input into the RPN network, which outputs two independent sets of candidate regions: a set of traffic light candidate boxes. Traffic sign candidate box set ,in, Indicates the first i Candidate boxes for traffic lights, Indicates the first j Candidate boxes for traffic signs, m This indicates the number of candidate boxes for traffic lights. n This indicates the number of candidate boxes for traffic signs.

[0072] It should be noted that by inputting multi-scale feature maps, it is possible to adapt to targets of different sizes, thereby ensuring the accuracy of subsequent feature vector extraction.

[0073] Step S1222: Through region of interest alignment operation, extract the traffic light feature vectors corresponding to each candidate box in the traffic light candidate box set and the traffic sign feature vectors corresponding to each candidate box in the traffic sign candidate box set from the multi-scale feature map.

[0074] After generating candidate bounding boxes for traffic lights and traffic signs, RoIAlign (Region of Interest Align) operation is further used to extract features of each candidate bounding box from the multi-scale feature map.

[0075] Specifically, the RoIAlign operation is used to extract the feature vectors corresponding to each candidate box in the traffic light candidate box set from the multi-scale feature map, denoted as the traffic light feature vector; at the same time, the feature vectors corresponding to each candidate box in the traffic sign candidate box set are extracted, denoted as the traffic sign feature vector.

[0076] Based on the above, the final feature map processing layer output consists of two sets of candidate boxes aligned to the same high-level semantic feature space and their feature vectors, denoted as... and ,in, Indicates the first i The feature vector of a traffic light Indicates the first j The feature vector of a traffic sign.

[0077] It should be noted that although each candidate box is extracted from feature maps of different resolutions, after the RoIAlign operation, it goes through a shared lightweight network (usually a small fully connected layer or a 1×1 convolutional layer) to map the features of all candidate boxes to the same dimension (e.g., 256-dimensional or 512-dimensional), so that the two types of feature vectors are in the same high-level semantic space, which is convenient for subsequent detection head processing.

[0078] The traffic signal joint sensing method provided in this invention, by sharing a lightweight RPN network and combining independent anchor box parameters and RoIAlign operation, can effectively solve the shape difference problem in traffic light and traffic sign detection while maintaining real-time performance, and provides high-quality, spatially aligned candidate regions and features for subsequent fine classification and regression.

[0079] Based on any of the above embodiments Figure 3 This is the third flowchart of the traffic signal joint sensing method provided by the present invention, as shown below. Figure 3 As shown, step S123 includes: step S1231, step S1232, step S1233 and step S1234.

[0080] Step S1231: Obtain the semantic association matrix constructed based on the preset semantic association information through the semantic interaction layer of the joint perception model. The semantic association matrix is ​​used to characterize the association strength between each traffic light state and each traffic sign category.

[0081] Considering that human drivers, when understanding traffic scenarios, do not identify individual signals in isolation, but rather associate and cross-verify the status of traffic lights with the meanings of surrounding traffic signs. For example, when observing a "red light," they anticipate and pay more attention to "stop and yield" signs; while when proceeding at a "green light," they actively seek out directional signs such as "go straight" or "turn left" to confirm their right of way. Therefore, in this embodiment of the invention, through a computable mechanism, this semantic-level prior knowledge is embedded into a neural network, enabling real-time, semantic-based information interaction and enhancement between the two perception tasks.

[0082] Specifically, two types of objects are defined first: traffic lights. Traffic signs ,in, and The lights are red, yellow, and green, respectively. N This indicates the number of traffic sign types. A semantic association matrix is ​​constructed based on pre-defined semantic association information (which can be determined based on traffic regulations and common sense). Among them, matrix elements Indicates the status of traffic lights Traffic sign categories The strength of semantic association between them.

[0083] In one implementation, the semantic association strength can be initialized to binary (e.g., 0 or 1) to represent “irrelevant” or “relevant”.

[0084] In another implementation, the semantic association strength can also be set as a continuous probability value or weight, representing the degree of closeness of the association.

[0085] It should be noted that the semantic association matrix can be constructed by analyzing co-occurrence patterns in large-scale datasets or by being manually defined by domain experts, and the embodiments of the present invention do not impose specific limitations.

[0086] Step S1232: Calculate the first target semantic attention weight based on the category distribution probability of the traffic light candidate box set and the semantic association matrix, and calculate the second target semantic attention weight based on the category distribution probability of the traffic sign candidate box set and the semantic association matrix.

[0087] Considering that there may be multiple traffic light and traffic sign candidate boxes in a certain frame of the image to be processed, it is necessary to calculate the semantic association strength between each traffic light candidate box and the traffic sign box.

[0088] Specifically, assuming for any traffic light candidate box i Its predicted class distribution probability is This indicates the probability that the light is red, yellow, or green; for any traffic sign candidate box j Its predicted class distribution probability is , which represents the probability that it belongs to each type of traffic sign.

[0089] For traffic light candidate boxes in the image i Calculate its relationship with traffic sign candidate boxes j The semantic attention weight is denoted as the first target semantic attention weight. The specific calculation formula is as follows: ; in, Indicates the semantic attention weights of the first target. Indicates traffic sign candidate boxes j The formula essentially uses the probability of traffic lights belonging to each category as weights to predict the category of the semantic association matrix. A A weighted average is taken from the corresponding row vectors to obtain a comprehensive correlation coefficient. .

[0090] Similarly, based on the category distribution probability and semantic association matrix of the traffic sign candidate box set, the semantic attention weight of each traffic sign candidate box to the traffic light candidate box is calculated, and denoted as the second target semantic attention weight.

[0091] Step S1233: Based on the first target semantic attention weight, the traffic light feature vector processed by network mapping is weighted and summed to obtain a first semantic context vector, and the first semantic context vector is fused with the traffic sign feature vector to obtain the traffic sign enhanced feature vector.

[0092] After calculating the first target semantic attention weight, the first target semantic attention weight is used. The traffic sign feature vector is modulated and fused to obtain the enhanced traffic sign feature vector, denoted as the enhanced traffic sign feature vector. During modulation and fusion, either feature weighted summation or channel attention mechanisms can be used.

[0093] In one implementation, a candidate bounding box for a traffic sign is first generated based on the first target semantic attention weight and the traffic light feature vector processed by the network mapping. j The semantic context vector, denoted as the first semantic context vector, is as follows: ; in, Represents the first semantic context vector. Indicates traffic light candidate boxes i The corresponding traffic light feature vector, This indicates that the traffic signal feature vector is processed through an MLP (Multi-Layer Perceptron). A nonlinear transformation is performed to map the data to a new feature space that is better suited for interacting with traffic signs.

[0094] Then, the first semantic context vector is compared with the traffic sign candidate boxes. j The traffic sign feature vector corresponding to itself The feature vectors of traffic signs are then fused to obtain the enhanced feature vectors. Fusion methods include, but are not limited to, splicing fusion or additive fusion.

[0095] Step S1234: Based on the second target semantic attention weight, the traffic sign feature vector processed by network mapping is weighted and summed to obtain the second semantic context vector, and the second semantic context vector is fused with the traffic light feature vector to obtain the traffic light enhanced feature vector.

[0096] Similarly, traffic sign feature vectors can be reverse-flowed to enhance traffic light feature vectors. First, based on the second target semantic attention weights and the traffic sign feature vectors processed by the network mapping, a semantic context vector for traffic light candidate boxes is generated, denoted as the second semantic context vector. Then, the second semantic context vector is fused with the traffic light feature vector to obtain the enhanced traffic light feature vector, denoted as the enhanced traffic light feature vector.

[0097] The traffic signal joint perception method provided in this invention transforms preset semantic association information into a quantized semantic association matrix, calculates the target semantic attention weight by combining the candidate box category distribution probability, and then performs weighted summation and feature fusion on the mapped feature vector to obtain an enhanced feature vector. Through this method, global semantic context constraints are forcibly injected into the independent features of local targets. This enables cross-task feature modulation of targets with local visual ambiguity caused by occlusion or strong light interference, effectively suppressing the activation response of logically mutually exclusive category combinations (such as green light status and stop sign), and enhancing the feature expression of high-frequency co-occurrence category combinations. This effectively overcomes local visual feature ambiguity caused by environmental interference such as strong light reflection and foliage occlusion in single-frame images, significantly reducing the false detection rate caused by logical conflicts, and ensuring strict traffic logic consistency and high robustness of the joint perception system's output results in complex scenarios.

[0098] Based on any of the above embodiments Figure 4 This is the fourth flowchart of the traffic signal joint sensing method provided by the present invention, as shown below. Figure 4As shown, step S124 includes: step S1241, step S1242, step S1243, step S1244 and step S1245.

[0099] Step S1241: Obtain the state transition probability matrix constructed based on the preset state transition prior information through the temporal verification layer of the joint sensing model. The state transition probability matrix is ​​used to characterize the probability of state transition of different traffic lights.

[0100] Traffic lights, as dynamic traffic control devices, follow strict temporal logic and periodic patterns in their state changes. At standard urban intersections, a typical signal cycle is usually "red → red + yellow (optional) → green → yellow → red," with each state having a defined lower limit for duration (e.g., green light at least 15 seconds, yellow light typically 3-5 seconds). However, existing traffic light detection methods based on single-frame images are highly susceptible to instantaneous interference, leading to unreasonable state transitions in consecutive frames, such as "green → yellow → green" (no red light transition) or "red → green" (skipping the yellow light), which violate traffic rules. To address this issue, this invention proposes a temporal consistency verification scheme based on a memory mechanism and state transition priors. By fusing historical state information of traffic lights with physically feasible state transition constraints, the original traffic light detection results of the current traffic signal image frame are smoothed, corrected, and reweighted with confidence, thereby outputting stable traffic light detection results that conform to real-world traffic logic.

[0101] Based on the above analysis, considering that the state changes of traffic lights are not random but follow a fixed cycle and sequence, this embodiment of the invention introduces preset state transition prior information to verify the timing consistency of traffic lights.

[0102] Furthermore, to encode the pre-defined state transition prior information, this embodiment of the invention introduces a learnable or rule-initialized state transition probability matrix. In, its matrix elements Indicates the status of traffic lights i Switch to traffic light status j Prior probability or rationality weight.

[0103] It should be noted that the state transition probability matrix can be used as a fixed parameter, dynamically loading different prior information based on map information (such as intersection type), or it can be designed as a learnable parameter, allowing the model to automatically learn the state transition patterns of traffic lights on actual roads from massive amounts of driving data.

[0104] Step S1242: Obtain the historical state memory queue for the same traffic light instance. The historical state memory queue stores the historical state detection results of the same traffic light instance in historical traffic signal image frames.

[0105] To fully utilize the historical state information of traffic lights, in this embodiment of the invention, at the current time t, a historical state memory queue Q with a fixed capacity (e.g., K frames) is maintained to cache the state detection results (denoted as historical state detection results) for the same traffic light instance within the most recent historical traffic signal image frames (e.g., K-1 frames). These historical state detection results include the historical detection state and its confidence level. Each element in the historical state memory queue Q is a triple. Where K represents the number of stored frames in the historical state memory queue. , indicating in tk The historical detection status of this traffic light instance at all times; , representing the confidence level corresponding to the historical detection status; It is the timestamp of the corresponding traffic signal image frame.

[0106] Furthermore, this historical state memory queue Q is updated according to the FIFO (First In First Out) principle. When a traffic light target is successfully tracked and matched in the current frame, its latest detection result is updated. The oldest test result will be added to the queue, while the oldest test result will be removed. This queue forms the basis of the short-term memory for temporal reasoning.

[0107] Step S1243: Based on the state transition probability matrix and the historical state memory queue, calculate the temporal support score corresponding to the candidate state of the traffic light, and obtain the original predicted state and its original state prediction distribution corresponding to the current traffic light image frame.

[0108] Step S1244: The original state prediction distribution and the temporal support score are weighted and fused to obtain a smooth state confidence distribution.

[0109] Step S1245: Determine the verified traffic light state based on the smoothed state confidence distribution.

[0110] After obtaining the state transition probability matrix B and the historical state memory queue Q, at each current time t, for each detected traffic light target in the current traffic signal image frame, the following smooth decision-making process is executed: First, based on the state transition probability matrix B and the historical state memory queue Q, calculate the degree of historical support that each traffic light candidate state can obtain, i.e., the sequential support score. For traffic light candidate states ( =0,1,2), its temporal support score is the weighted support of historical state detection results, and the calculation formula is: ,in, Z It is a normalization factor, where K represents the number of stored frames in the history state memory queue. From historical detection status Transition to candidate state The prior probability, This is the confidence level of historical detection states. This calculation means that the more reasonable the transition path formed by the current predicted state and the historical detection states, the higher the confidence level. The higher the value, the more reliable the historical state itself ( The higher the value, the higher the timing support score it receives.

[0111] Simultaneously, the original state prediction result is acquired. Specifically, the traffic light prediction result of the current traffic light image frame is obtained from the traffic light detection head, and is denoted as the original state prediction result. The original state prediction result includes the original predicted state. and its corresponding original state prediction distribution .

[0112] Then, the original state prediction distribution is weighted and fused with the temporal support score to obtain the smoothed state confidence distribution. That is, the smoothed state confidence distribution. ,in It is an adjustable hyperparameter used to balance the weight between current predictions and historical consistency.

[0113] Finally, based on the smooth state confidence distribution, the verified traffic light state is determined. The verified traffic light state is as follows: This state and their corresponding confidence levels max( ) , that is As the output of this embodiment of the invention, it is also used to update the historical state memory queue Q.

[0114] The traffic signal joint perception method provided in this invention significantly reduces the false detection rate and state flickering frequency of the model in complex lighting and partial occlusion scenarios without increasing the computational load of front-end visual feature extraction through a temporal verification layer. Simultaneously, the verified traffic signal light states strictly conform to the physical laws of traffic operation, improving the smoothness and safety of autonomous vehicles when making start-stop decisions at intersections.

[0115] Based on any of the above embodiments Figure 5This is the fifth flowchart of the traffic signal joint sensing method provided by the present invention, as shown below. Figure 5 As shown, before step S122, there are steps S126 and S127.

[0116] Step S126: Based on the feature alignment layer of the joint sensing model, perform temporal alignment on the target deep feature map in the multi-scale feature map to obtain the aligned target deep feature map.

[0117] Considering that the position of the same real-world object in an image can change due to vehicle or camera movement within consecutive frames, directly stitching or temporally fusing the features of these frames would result in spatial misalignment of the features of the same object, making it difficult for the network to effectively utilize temporal information. Therefore, in order to effectively fuse temporal information and compensate for target displacement caused by vehicle movement, this embodiment of the invention performs explicit temporal alignment of the high-level semantic features of consecutive frames (i.e., the sequence of traffic signal images to be detected).

[0118] Specifically, the target deep feature map in the multi-scale feature map is input into the feature alignment layer of the joint sensing model. The target deep feature map is temporally aligned through the feature alignment layer to obtain the aligned temporal feature sequence.

[0119] Among them, the target deep feature map can be the deepest feature map, i.e., the th layer. M Layer feature map It can also be a feature map of the next deeper layer, i.e., the first layer. M -1 layer feature map Preferably, it is the first M Layer feature map The following explanation will use this as an example.

[0120] In one implementation, temporal alignment can be achieved using a deformable convolutional module, which learns adaptive spatial offsets to semantically align features from the previous frame with features from the current frame, thereby obtaining an aligned target deep feature map. .

[0121] Temporal alignment ensures that features extracted from consecutive frames corresponding to the same scene element have spatial consistency, providing stable input for subsequent temporal analysis.

[0122] Step S127: The aligned target deep feature map and shallow feature map are fused to obtain an updated multi-scale feature map; wherein, the shallow feature map is the feature map in the multi-scale feature map other than the target deep feature map.

[0123] After obtaining the aligned target deep feature map, the aligned target deep feature map is fused with the aforementioned unaligned shallow feature map in a top-down manner using the Feature Pyramid Network (FPN) or its variant architecture.

[0124] Specifically, the aligned target deep feature map is first upsampled to match the spatial resolution of the shallow feature map, and then element-wise added or channel-wise concatenated with the corresponding shallow feature map. After the above multi-level fusion, an updated multi-scale feature map is obtained.

[0125] At this point, step S122 includes step S1221.

[0126] Step S1221: The updated multi-scale feature map is processed through the feature map processing layer of the joint perception model to obtain the traffic light feature vector and the traffic sign feature vector.

[0127] Next, the updated multi-scale feature map is processed through the feature map processing layer of the joint perception model to obtain traffic light feature vectors and traffic sign feature vectors, and then subsequent steps are executed. The specific execution process can be referred to the above embodiment, and will not be repeated here.

[0128] The traffic signal joint perception method provided in this invention effectively avoids the enormous computational overhead caused by complex spatial transformations of high-resolution shallow features by performing temporal alignment calculations only on the deep feature maps of targets with the lowest resolution. This significantly reduces the computational complexity of the model and meets the millisecond-level real-time requirements of autonomous driving scenarios. Simultaneously, through a top-down feature fusion mechanism, the aligned stable semantics and motion compensation information from the aligned deep features are progressively passed down to the shallow feature maps. This results in enhanced temporal consistency across all levels of the final multi-scale feature maps, improving the stability of long-distance, small-scale target detection in consecutive frames. It overcomes the problems of single-frame target flickering and missed detection caused by vehicle motion, comprehensively improving the robustness of the joint perception system.

[0129] Based on any of the above embodiments, before step S125, the method further includes steps S128, S129, S130, and S131.

[0130] It should be noted that the execution order of steps S130 and S131 is not important and they can be executed in parallel.

[0131] Step S128: Through the weight adjustment layer of the joint perception model, the aligned target deep feature map is globally pooled to obtain a scene-level semantic feature vector.

[0132] In traditional multi-task learning architectures, shared backbone network features are fed into all subtask heads. However, in complex and ever-changing real-world driving scenarios, the relative importance of traffic lights and traffic signs is not constant. For example, at urban intersections, traffic lights are the absolute dominant factor in traffic decisions; while on highways or in tunnels, traffic lights are absent, and traffic signs become crucial for perception. Based on the above analysis, this invention introduces a lightweight adaptive gating mechanism to generate adaptive gating weights for traffic light detection and traffic sign detection tasks in real time.

[0133] With the aligned target deep feature map The input is fed into the weight adjustment layer of the joint sensing model. Through the weight adjustment layer of the joint sensing model, the aligned target deep feature map is processed. Perform global pooling to extract scene-level semantic features. ,in, This indicates global average pooling.

[0134] Step S129: Based on the scene-level semantic feature vector, calculate the first gating weight corresponding to the traffic light detection task and the second gating weight corresponding to the traffic sign detection task.

[0135] Then, a lightweight, learnable network is used to generate dynamic gating weights for the traffic light detection task based on scene-level semantic feature vectors, denoted as the first gating weight. Simultaneously, dynamic gating weights corresponding to the traffic sign detection task are generated, denoted as the second gating weights. .

[0136] Furthermore, + =1.

[0137] Step S130: Based on the first gating weight, the traffic light enhancement feature vector is weighted and modulated to obtain the modulated traffic light enhancement feature vector.

[0138] Step S131: Based on the second gating weight, the traffic sign enhancement feature vector is weighted and modulated to obtain the modulated traffic sign enhancement feature vector.

[0139] After calculating the first and second gating weights, the traffic light enhancement feature vector is weighted and modulated based on the first gating weight to obtain the modulated traffic light enhancement feature vector. Simultaneously, the traffic sign enhancement feature vector is weighted and modulated based on the second gating weight to obtain the modulated traffic sign enhancement feature vector.

[0140] At this time, step S125 includes: step S1251.

[0141] Step S1251: Using the traffic light detection head of the joint sensing model, the traffic light detection result is determined based on the modulated traffic light enhancement feature vector and the verified traffic light state. Similarly, using the traffic sign detection head of the joint sensing model, the traffic sign detection result is determined based on the modulated traffic sign enhancement feature vector.

[0142] The modulated traffic light enhanced feature vector and the verified traffic light state are input into the traffic light detection head of the joint sensing model to perform traffic light detection and obtain the traffic light detection result output by the traffic light detection head.

[0143] Simultaneously, the modulated traffic sign enhanced feature vector is input into the traffic sign detection head of the joint perception model to perform traffic sign detection, and the traffic sign detection result output by the traffic sign detection head is obtained.

[0144] By dynamically generating differentiated gating weights, adaptive weighted allocation of features within the multi-task network is achieved, realizing feature splitting. On the one hand, it can directly suppress the feature response of the corresponding branch in scenarios where the probability of a specific target appearing is extremely low, effectively avoiding false positives caused by background noise such as ambient light sources or reflectors. On the other hand, it can simultaneously amplify the feature representation capabilities of tasks with higher decision priority, comprehensively improving the accuracy and environmental robustness of target detection results while strictly ensuring the high real-time performance and low inference latency of autonomous driving algorithms.

[0145] The traffic signal joint sensing method provided in this embodiment of the invention, through the above-mentioned dynamic task adaptive mechanism, can significantly reduce the false detection rate in specific scenarios with extremely low additional computational overhead and improve the overall efficiency of multi-task sensing.

[0146] Based on any of the above embodiments, the joint perception model is trained using a preset loss function, which is determined based on traffic light detection loss, traffic sign detection loss, semantic consistency loss, and temporal smoothing loss.

[0147] The semantic consistency loss is determined based on the preset semantic association information, the confidence level of the sample traffic light status prediction, and the confidence level of the sample traffic sign category prediction. The temporal smoothing loss is determined based on the preset state transition prior information and the sample state prediction distribution corresponding to adjacent image frames in the sample traffic signal image sequence.

[0148] To address the strong semantic and spatial correlation between traffic light detection and traffic sign detection, this invention proposes a collaborative perception loss function. Based on traditional detection loss, it introduces semantic consistency constraints and temporal smoothness constraints. The preset loss function is as follows: ; in, Indicates the total loss; and These represent the loss for traffic light detection and the loss for traffic sign detection, respectively. Both of these are basic detection loss functions. This represents the semantic consistency loss. Indicates the time-series smoothing loss; and This indicates the preset weighting coefficient.

[0149] Furthermore, semantic consistency loss aims to transform the inter-category logical relationships modeled by the semantic interaction layer into an optimizable training objective. Specifically, it directly penalizes semantically conflicting detection results that simultaneously appear with high confidence, forcing the model to output results that conform to traffic rule logic from the perspective of the loss function. Its calculation formula is as follows: ; Where A represents the semantic association matrix; This indicates an indicator function, which is 1 when the condition within the brackets [ ] is true, and 0 otherwise; and These are the category confidence scores for the i-th traffic light and the j-th traffic sign predicted by the model, respectively.

[0150] Furthermore, the temporal smoothing loss corresponds to the optimization objective of the temporal consistency verification module. It constrains the model's prediction distribution across consecutive frames to conform to the physical laws of traffic light state transitions, thereby achieving smooth and reasonable output in the time dimension and avoiding instantaneous jumps that violate state machine logic. Its calculation formula is as follows: ; in, Is the model in t The probability distribution of the final state predicted for a given traffic light instance in the current traffic signal image frame; This indicates that the parameters within the parentheses are considered as a categorical distribution; This represents the state transition prior matrix, which encodes the probability of transitioning from the previous state to the next state; KL (Kullback-Leibler) divergence, also known as relative entropy, is used to measure the difference between two probability distributions. This represents the expected distribution of the current traffic signal image frame state, calculated based on the predicted distribution and state transition matrix of the previous frame.

[0151] Temporal smoothing loss encourages the model to predict the distribution of the current traffic signal image frame. Compared with the reasonable expected distribution calculated based on the prediction of the previous frame and the state transition prior. Maintaining consistency across frames. Minimizing this KL divergence means the model is trained such that its predicted state changes must conform to physical temporal logic. For example, if the previous frame predicted a "red light" with a high probability, then the prediction for the current frame should focus on either "remaining red" or "changing to yellow," rather than suddenly predicting a "green light" with a high probability. This effectively suppresses state jitter caused by noise in a single frame that violates temporal logic.

[0152] Furthermore, the basic detection loss function follows the standard form of mainstream object detection frameworks. In one implementation, the traffic light detection loss can be determined based on the sample traffic light prediction results and sample traffic light detection results, and the traffic sign detection loss can be determined based on the sample traffic sign prediction results and sample traffic sign detection results, specifically characterized by cross-entropy loss.

[0153] The traffic signal joint perception method provided in this invention introduces semantic consistency and temporal smoothness constraints on top of traditional detection loss. The semantic consistency loss term, by numerically penalizing the confidence of mutually exclusive output categories, forces the network to actively suppress feature combinations that violate traffic rules during backpropagation, thereby effectively reducing the cross-class false detection rate in complex scenarios and ensuring the logical consistency of the joint perception results for the two types of targets. The temporal smoothness loss term, by minimizing the divergence between the current observed distribution and the historical expected distribution, forces the model to internalize the objective state transition rules of traffic lights, thus significantly suppressing state jumps caused by single-frame visual noise or instantaneous occlusion, ultimately comprehensively improving the output stability of the joint perception model in dynamic continuous driving environments.

[0154] The traffic signal joint sensing device provided in the embodiments of the present invention is described below. The traffic signal joint sensing device described below can be referred to in correspondence with the traffic signal joint sensing method described above.

[0155] Figure 6 This is a schematic diagram of the traffic signal joint sensing device provided by the present invention, as shown below. Figure 6 As shown, the device includes an image sequence acquisition module 610 and a detection result acquisition module 620; wherein: Image sequence acquisition module 610 is used to acquire the image sequence of traffic signals to be detected; The detection result acquisition module 620 is used to input the traffic signal image sequence to be detected into the joint perception model to obtain the traffic light detection result and traffic sign detection result output by the joint perception model; The joint perception model is trained based on a sequence of sample traffic signal images, the detection results of sample traffic lights and sample traffic signs corresponding to the sequence of sample traffic signal images, preset semantic association information, and preset state transition prior information. The joint perception model is used to perform joint perception based on the enhanced feature vectors of traffic lights, enhanced feature vectors of traffic signs, and the verified state of traffic lights. The enhanced feature vectors of the traffic lights and the enhanced feature vectors of the traffic signs are determined based on the traffic signal image sequence to be detected and the preset semantic association information; the verified traffic light state is obtained by performing temporal consistency verification based on the preset state transition prior information.

[0156] The traffic signal joint sensing device provided in this embodiment of the invention first acquires a sequence of traffic signal images to be detected; then, it inputs the sequence of traffic signal images to be detected into a joint sensing model to obtain the traffic light detection results and traffic sign detection results output by the joint sensing model. The joint sensing model is trained based on a sample traffic signal image sequence, the corresponding sample traffic light detection results and sample traffic sign detection results, preset semantic association information, and preset state transition prior information. The joint sensing model is used for joint sensing based on enhanced feature vectors for traffic lights, enhanced feature vectors for traffic signs, and verified traffic light states. The enhanced feature vectors for traffic lights and traffic signs are determined based on the sequence of traffic signal images to be detected and preset semantic association information; the verified traffic light states are obtained by performing temporal consistency verification based on preset state transition prior information. This invention constructs a joint perception model and utilizes pre-defined semantic association information to determine enhanced feature vectors for traffic lights and signs. This achieves feature-level cross-category cross-validation and auxiliary enhancement, thereby improving the overall perception robustness in complex scenarios. Simultaneously, it uses pre-defined state transition prior information to perform temporal consistency verification of traffic light states, effectively filtering out single-frame perception jumps that violate physical common sense. This invention deeply integrates semantic interaction features and temporal verification states into a single model for joint perception decision-making. This not only avoids the computational waste caused by multiple independent models running in parallel, effectively reducing system deployment complexity and computational overhead, but also eliminates semantic ambiguity caused by independent outputs from the perception source. The final output is a high-confidence detection result that has undergone cross-validation and logical association, thus providing downstream planning and control modules with more accurate, reliable, and semantically consistent traffic light and sign detection results, improving the performance and safety level of the entire autonomous driving system.

[0157] According to a traffic signal joint sensing device provided by the present invention, the detection result acquisition module 620 includes: The feature extraction unit is used to extract multi-scale feature maps corresponding to each frame of the traffic signal image sequence to be detected through the feature extraction layer of the joint sensing model; The feature map processing unit is used to process the multi-scale feature map through the feature map processing layer of the joint perception model to obtain traffic light feature vectors and traffic sign feature vectors. The semantic interaction unit is used to perform semantic interaction enhancement processing on the traffic light feature vector and the traffic sign feature vector respectively through the semantic interaction layer of the joint perception model, based on the preset semantic association information, to obtain the traffic light enhanced feature vector and the traffic sign enhanced feature vector. The timing verification unit is used to perform timing consistency verification on the predicted state of the traffic light through the timing verification layer of the joint sensing model, based on the preset state transition prior information, to obtain the verified traffic light state. The result acquisition unit is used to determine the traffic light detection result based on the traffic light enhancement feature vector and the verified traffic light state through the traffic light detection head of the joint perception model, and to determine the traffic sign detection result based on the traffic sign enhancement feature vector through the traffic sign detection head of the joint perception model.

[0158] According to a traffic signal joint sensing device provided by the present invention, the feature map processing unit is specifically used for: The candidate box generation network shared in the feature map processing layer of the joint perception model is used to extract candidate boxes from the multi-scale feature map to obtain a traffic light candidate box set and a traffic sign candidate box set; wherein the traffic light candidate box set and the traffic sign candidate box set are generated based on their respective anchor box parameters. By performing region of interest alignment operations, traffic light feature vectors corresponding to each candidate box in the traffic light candidate box set and traffic sign feature vectors corresponding to each candidate box in the traffic sign candidate box set are extracted from the multi-scale feature map.

[0159] According to a traffic signal joint sensing device provided by the present invention, the semantic interaction unit is specifically used for: The semantic interaction layer of the joint perception model is used to obtain a semantic association matrix constructed based on the preset semantic association information. The semantic association matrix is ​​used to characterize the association strength between each traffic light state and each traffic sign category. Based on the category distribution probability of the traffic light candidate box set and the semantic association matrix, the first target semantic attention weight is calculated, and based on the category distribution probability of the traffic sign candidate box set and the semantic association matrix, the second target semantic attention weight is calculated. Based on the first target semantic attention weight, the traffic light feature vector processed by the network mapping is weighted and summed to obtain the first semantic context vector, and the first semantic context vector is fused with the traffic sign feature vector to obtain the traffic sign enhanced feature vector; Based on the second target semantic attention weight, the traffic sign feature vector processed by network mapping is weighted and summed to obtain the second semantic context vector, and the second semantic context vector is fused with the traffic light feature vector to obtain the traffic light enhanced feature vector.

[0160] According to a traffic signal joint sensing device provided by the present invention, the timing verification unit is specifically used for: The state transition probability matrix is ​​obtained by using the temporal verification layer of the joint sensing model and constructing it based on the preset state transition prior information. The state transition probability matrix is ​​used to characterize the probability of different traffic light state transitions. Obtain a historical state memory queue for the same traffic light instance, wherein the historical state memory queue stores the historical state detection results of the same traffic light instance in historical traffic signal image frames; Based on the state transition probability matrix and the historical state memory queue, the temporal support score corresponding to the candidate state of the traffic light is calculated, and the original predicted state and its original state prediction distribution corresponding to the current traffic light image frame are obtained. The original state prediction distribution is weighted and fused with the temporal support score to obtain a smooth state confidence distribution. The verified traffic light state is determined based on the smooth state confidence distribution.

[0161] According to a traffic signal joint sensing device provided by the present invention, the detection result acquisition module 620 further includes: The feature alignment unit is used to perform temporal alignment of the target deep feature map in the multi-scale feature map based on the feature alignment layer of the joint sensing model, so as to obtain the aligned target deep feature map. The feature map fusion unit is used to fuse the aligned target deep feature map with the shallow feature map to obtain an updated multi-scale feature map; wherein, the shallow feature map is the feature map in the multi-scale feature map other than the target deep feature map; The feature map processing unit is specifically used for: The updated multi-scale feature map is processed through the feature map processing layer of the joint perception model to obtain the traffic light feature vector and the traffic sign feature vector.

[0162] According to a traffic signal joint sensing device provided by the present invention, the detection result acquisition module 620 further includes: A global pooling unit is used to perform global pooling on the aligned target deep feature map through the weight adjustment layer of the joint perception model to obtain a scene-level semantic feature vector. The weight calculation unit is used to calculate the first gating weight corresponding to the traffic light detection task and the second gating weight corresponding to the traffic sign detection task based on the scene-level semantic feature vector. The first modulation unit is used to perform weighted modulation on the traffic light enhancement feature vector based on the first gate weight to obtain the modulated traffic light enhancement feature vector. The second modulation unit is used to perform weighted modulation on the traffic sign enhancement feature vector based on the second gating weight to obtain the modulated traffic sign enhancement feature vector. The result acquisition unit is specifically used for: The traffic light detection head of the joint sensing model determines the traffic light detection result based on the modulated traffic light enhancement feature vector and the verified traffic light state. Similarly, the traffic sign detection head of the joint sensing model determines the traffic sign detection result based on the modulated traffic sign enhancement feature vector.

[0163] According to a traffic signal joint sensing device provided by the present invention, the joint sensing model is trained using a preset loss function, which is determined based on traffic light detection loss, traffic sign detection loss, semantic consistency loss and temporal smoothing loss. The semantic consistency loss is determined based on the preset semantic association information, the confidence level of the sample traffic light status prediction, and the confidence level of the sample traffic sign category prediction. The temporal smoothing loss is determined based on the preset state transition prior information and the sample state prediction distribution corresponding to adjacent image frames in the sample traffic signal image sequence.

[0164] It should be noted that the traffic signal joint sensing device provided in this embodiment of the invention can implement all the method steps implemented in the above traffic signal joint sensing method embodiment and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.

[0165] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communications bus 740. The processor 710 can call logic instructions in the memory 730 to execute the traffic signal joint perception method provided in the above embodiments. The method includes: acquiring a traffic signal image sequence to be detected; inputting the traffic signal image sequence to be detected into a joint perception model to obtain traffic light detection results and traffic sign detection results output by the joint perception model; wherein, the joint perception model is trained based on a sample traffic signal image sequence, sample traffic light detection results and sample traffic sign detection results corresponding to the sample traffic signal image sequence, preset semantic association information, and preset state transition prior information; the joint perception model is used to perform joint perception based on traffic light enhancement feature vectors, traffic sign enhancement feature vectors, and verified traffic light states; wherein, the traffic light enhancement feature vectors and traffic sign enhancement feature vectors are determined based on the traffic signal image sequence to be detected and the preset semantic association information; the verified traffic light states are obtained by performing time sequence consistency verification based on the preset state transition prior information.

[0166] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0167] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the traffic signal joint sensing method provided in the above embodiments, the method comprising: acquiring a sequence of traffic signal images to be detected; The traffic signal image sequence to be detected is input into a joint perception model to obtain the traffic light detection results and traffic sign detection results output by the joint perception model. The joint perception model is trained based on a sample traffic signal image sequence, the corresponding sample traffic light detection results and sample traffic sign detection results, preset semantic association information, and preset state transition prior information. The joint perception model is used for joint perception based on enhanced feature vectors for traffic lights, enhanced feature vectors for traffic signs, and verified traffic light states. The enhanced feature vectors for traffic lights and traffic signs are determined based on the traffic signal image sequence to be detected and the preset semantic association information. The verified traffic light states are obtained by performing temporal consistency verification based on the preset state transition prior information.

[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0169] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary high-resource hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for joint sensing of traffic signals, characterized in that, include: Acquire the image sequence of the traffic signal to be detected; The traffic signal image sequence to be detected is input into the joint perception model to obtain the traffic light detection results and traffic sign detection results output by the joint perception model. The joint perception model is trained based on a sequence of sample traffic signal images, the detection results of sample traffic lights and sample traffic signs corresponding to the sequence of sample traffic signal images, preset semantic association information, and preset state transition prior information. The preset state transition prior information is the physical switching rule followed by traffic lights in the time dimension. The joint perception model is used to perform joint perception based on the enhanced feature vectors of traffic lights, enhanced feature vectors of traffic signs, and the verified state of traffic lights. The enhanced feature vectors of the traffic lights and the enhanced feature vectors of the traffic signs are determined based on the traffic signal image sequence to be detected and the preset semantic association information; the verified traffic light state is obtained by performing temporal consistency verification based on the preset state transition prior information.

2. The traffic signal joint sensing method according to claim 1, characterized in that, The step of inputting the traffic signal image sequence to be detected into the joint perception model to obtain the traffic light detection results and traffic sign detection results output by the joint perception model includes: The feature extraction layer of the joint sensing model is used to extract multi-scale feature maps corresponding to each frame of the traffic signal image sequence to be detected. The feature map processing layer of the joint perception model processes the multi-scale feature map to obtain traffic light feature vectors and traffic sign feature vectors. Through the semantic interaction layer of the joint perception model, based on the preset semantic association information, semantic interaction enhancement processing is performed on the traffic light feature vector and the traffic sign feature vector respectively to obtain the traffic light enhanced feature vector and the traffic sign enhanced feature vector. Through the temporal verification layer of the joint sensing model, based on the preset state transition prior information, the temporal consistency verification of the predicted state of the traffic light is performed to obtain the verified state of the traffic light. The traffic light detection head of the joint perception model determines the traffic light detection result based on the traffic light enhanced feature vector and the verified traffic light state. Similarly, the traffic sign detection head of the joint perception model determines the traffic sign detection result based on the traffic sign enhanced feature vector.

3. The traffic signal joint sensing method according to claim 2, characterized in that, The feature map processing layer of the joint perception model processes the multi-scale feature map to obtain traffic light feature vectors and traffic sign feature vectors, including: The candidate box generation network shared in the feature map processing layer of the joint perception model is used to extract candidate boxes from the multi-scale feature map to obtain a traffic light candidate box set and a traffic sign candidate box set; wherein the traffic light candidate box set and the traffic sign candidate box set are generated based on their respective anchor box parameters. By performing a region of interest alignment operation, the traffic light feature vectors corresponding to each candidate box in the traffic light candidate box set and the traffic sign feature vectors corresponding to each candidate box in the traffic sign candidate box set are extracted from the multi-scale feature map.

4. The traffic signal joint sensing method according to claim 3, characterized in that, The semantic interaction layer of the joint perception model, based on the preset semantic association information, performs semantic interaction enhancement processing on the traffic light feature vector and the traffic sign feature vector respectively, to obtain enhanced traffic light feature vector and enhanced traffic sign feature vector, including: The semantic interaction layer of the joint perception model is used to obtain a semantic association matrix constructed based on the preset semantic association information. The semantic association matrix is ​​used to characterize the association strength between each traffic light state and each traffic sign category. Based on the category distribution probability of the traffic light candidate box set and the semantic association matrix, the first target semantic attention weight is calculated, and based on the category distribution probability of the traffic sign candidate box set and the semantic association matrix, the second target semantic attention weight is calculated. Based on the first target semantic attention weight, the traffic light feature vector processed by network mapping is weighted and summed to obtain the first semantic context vector, and the first semantic context vector is fused with the traffic sign feature vector to obtain the traffic sign enhanced feature vector; Based on the second target semantic attention weight, the traffic sign feature vector processed by network mapping is weighted and summed to obtain the second semantic context vector, and the second semantic context vector is fused with the traffic light feature vector to obtain the traffic light enhanced feature vector.

5. The traffic signal joint sensing method according to claim 2, characterized in that, The step of using the temporal verification layer of the joint sensing model to perform temporal consistency verification on the predicted state of the traffic lights based on the preset state transition prior information, and obtaining the verified traffic light state, includes: The state transition probability matrix is ​​obtained through the temporal verification layer of the joint sensing model, which is constructed based on the preset state transition prior information. The state transition probability matrix is ​​used to characterize the probability of different traffic light state transitions. Obtain a historical state memory queue for the same traffic light instance, wherein the historical state memory queue stores the historical state detection results of the same traffic light instance in historical traffic signal image frames; Based on the state transition probability matrix and the historical state memory queue, the temporal support score corresponding to the candidate state of the traffic light is calculated, and the original predicted state and its original state prediction distribution corresponding to the current traffic light image frame are obtained. The original state prediction distribution is weighted and fused with the temporal support score to obtain a smooth state confidence distribution. The verified traffic light state is determined based on the smooth state confidence distribution.

6. The traffic signal joint sensing method according to any one of claims 2 to 5, characterized in that, Before the feature map processing layer of the joint perception model processes the multi-scale feature map to obtain the traffic light feature vector and traffic sign feature vector, the method further includes: Based on the feature alignment layer of the joint sensing model, the target deep feature map in the multi-scale feature map is temporally aligned to obtain the aligned target deep feature map. The aligned target deep feature map and shallow feature map are fused to obtain an updated multi-scale feature map; wherein, the shallow feature map is the feature map in the multi-scale feature map other than the target deep feature map; The feature map processing layer of the joint perception model processes the multi-scale feature map to obtain traffic light feature vectors and traffic sign feature vectors, including: The updated multi-scale feature map is processed through the feature map processing layer of the joint perception model to obtain the traffic light feature vector and the traffic sign feature vector.

7. The traffic signal joint sensing method according to claim 6, characterized in that, Before the traffic light detection head using the joint perception model determines the traffic light detection result based on the enhanced feature vector of the traffic light and the verified state of the traffic light, and before the traffic sign detection head using the joint perception model determines the traffic sign detection result based on the enhanced feature vector of the traffic sign, the process further includes: The aligned target deep feature map is globally pooled through the weight adjustment layer of the joint perception model to obtain a scene-level semantic feature vector. Based on the scene-level semantic feature vector, calculate the first gating weight corresponding to the traffic light detection task and the second gating weight corresponding to the traffic sign detection task. Based on the first gating weight, the traffic light enhancement feature vector is weighted and modulated to obtain the modulated traffic light enhancement feature vector. Based on the second gating weight, the traffic sign enhancement feature vector is weighted and modulated to obtain the modulated traffic sign enhancement feature vector. The traffic light detection head using the joint perception model determines the traffic light detection result based on the enhanced feature vector of the traffic light and the verified state of the traffic light, and the traffic sign detection head using the joint perception model determines the traffic sign detection result based on the enhanced feature vector of the traffic sign, including: The traffic light detection head of the joint sensing model determines the traffic light detection result based on the modulated traffic light enhancement feature vector and the verified traffic light state. Similarly, the traffic sign detection head of the joint sensing model determines the traffic sign detection result based on the modulated traffic sign enhancement feature vector.

8. The traffic signal joint sensing method according to any one of claims 1 to 5, characterized in that, The joint perception model is trained using a preset loss function, which is determined based on traffic light detection loss, traffic sign detection loss, semantic consistency loss, and temporal smoothing loss. The semantic consistency loss is determined based on the preset semantic association information, the confidence level of the sample traffic light status prediction, and the confidence level of the sample traffic sign category prediction. The temporal smoothing loss is determined based on the preset state transition prior information and the sample state prediction distribution corresponding to adjacent image frames in the sample traffic signal image sequence.

9. A traffic signal joint sensing device, characterized in that, include: The image sequence acquisition module is used to acquire the image sequence of the traffic signal to be detected; The detection result acquisition module is used to input the traffic signal image sequence to be detected into the joint perception model to obtain the traffic light detection result and traffic sign detection result output by the joint perception model; The joint perception model is trained based on a sequence of sample traffic signal images, the detection results of sample traffic lights and sample traffic signs corresponding to the sequence of sample traffic signal images, preset semantic association information, and preset state transition prior information. The preset state transition prior information is the physical switching rule followed by traffic lights in the time dimension. The joint perception model is used to perform joint perception based on the enhanced feature vectors of traffic lights, enhanced feature vectors of traffic signs, and the verified state of traffic lights. The enhanced feature vectors of the traffic lights and the enhanced feature vectors of the traffic signs are determined based on the traffic signal image sequence to be detected and the preset semantic association information; the verified traffic light state is obtained by performing temporal consistency verification based on the preset state transition prior information.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the traffic signal joint sensing method as described in any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the traffic signal joint sensing method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Traffic sign recognition method and device, equipment and medium

    CN115273032A

  • Map generation method and device, equipment and storage medium

    CN117928574A