Target identification method and device, equipment, storage medium and product
By setting target selection rules and a weakly supervised training mechanism with pseudo-labels, the problem of low traffic light recognition accuracy in traditional methods is solved, achieving accurate target recognition and vehicle behavior guidance in complex traffic scenarios, and improving the system's generalization ability.
Patent Information
- Application Number
- CN202511525602.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-24
AI Technical Summary
Traditional target recognition methods are difficult to adapt flexibly in dynamic traffic scenarios, especially in scenarios with multiple overlapping lights, semantic ambiguity, or long distances. Traffic light recognition accuracy is low, perception generalization ability is poor, and there is a lack of traffic light behavior guidance, making it difficult for vehicles to accurately determine which traffic light to follow.
By setting target selection rules, a target candidate set is selected from the feature map. Combining the correlation of autonomous vehicle driving behavior, a semantic alignment mechanism between autonomous vehicle intent and candidate targets is established. A weakly supervised training mechanism with pseudo-labels is adopted, and rule scores and intent matching scores are fused to generate pseudo-label data. A fusion classifier is then trained to identify the set targets.
It improves the accuracy and generalization ability of target recognition in complex environments, solves the problem of difficult dataset construction, and enhances the robustness and accuracy of traffic light recognition.
Smart Images

Figure CN120997785A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and in particular to a target recognition method, apparatus, device, storage medium, and product. Background Technology
[0002] Urban intersections, as critical nodes with high traffic density and complex traffic rules, directly impact traffic safety and system efficiency through their perception and decision-making capabilities. In typical urban roads, small targets such as traffic lights, signs, and traffic cones are widely distributed in important locations such as intersections, construction zones, and ramp boundaries. Although these targets occupy only a tiny pixel area in the image, they carry crucial road traffic information.
[0003] Traditional target recognition methods struggle to adapt flexibly to dynamic traffic scenarios, especially in situations with overlapping traffic lights, semantic ambiguity, or long distances. These methods exhibit low accuracy in target traffic light recognition and poor perceptual generalization. Furthermore, traditional target recognition methods lack effective integration of traffic lights with behavioral guidance. Particularly in complex urban intersections, where numerous traffic lights exist and suffer from severe obstruction, small size, and long distances, the system often struggles to accurately determine which traffic light a vehicle should follow.
[0004] It is evident that improving the accuracy of traffic light recognition is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] This application provides a target recognition method, apparatus, device, storage medium, and product to at least solve the problem in the related art of accurately determining which traffic light a vehicle should follow.
[0006] This application provides a target recognition method, including: Based on the set target selection rules, a target candidate set is selected from the feature map set; the target selection rules are set according to the correlation between the set target and the autonomous vehicle driving behavior; the target candidate set contains each candidate target and its corresponding rule score; The feature information of each candidate target in the target candidate set is semantically aligned with the vehicle's intent vector to obtain a fused feature vector; The visual representation features of each candidate target and the vehicle's intent vector are analyzed to obtain the intent matching score of each candidate target. Based on the intent matching score and rule score of each candidate target, pseudo-label data is constructed; the pseudo-label data includes the predicted probability that each candidate target belongs to the set target. Based on the pseudo-label data and the recognition results output by the fusion classifier in analyzing the fused feature vector, the parameters of the fusion classifier are adjusted to obtain a trained fusion classifier, which can then be used to identify the set target.
[0007] This application also provides a target recognition device, including a filtering unit, an alignment unit, a matching unit, a construction unit, and an adjustment unit; The filtering unit is used to filter out a set of candidate targets from the feature map set according to the set target filtering rules; wherein, the target filtering rules are set based on the correlation between the set target and the autonomous vehicle driving behavior; the target candidate set contains each candidate target and its corresponding rule score; The alignment unit is used to semantically align the feature information of each candidate target in the target candidate set with the vehicle's intent vector to obtain a fused feature vector; The matching unit is used to analyze the visual representation features of each candidate target and the vehicle's intent vector to obtain the intent matching score of each candidate target. The construction unit is used to construct pseudo-label data based on the intent matching score and rule score of each candidate target; wherein, the pseudo-label data includes the predicted probability that each candidate target belongs to the set target; The adjustment unit is used to adjust the parameters of the fusion classifier based on the pseudo-label data and the recognition results output by the fusion classifier in analyzing the fusion feature vector, so as to obtain a trained fusion classifier, which can be used to identify the set target.
[0008] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the target recognition methods described above.
[0009] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described target recognition methods.
[0010] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described target recognition methods.
[0011] This application uses a set of target candidates to filter from a feature map set according to a set of defined target selection rules. These rules are based on the correlation between the defined target and the vehicle's driving behavior. The target candidate set includes each candidate target and its corresponding rule score. By setting the target selection rules, irrelevant targets in the feature map set can be eliminated. To increase the accuracy of target selection, the feature information of each candidate target in the target candidate set can be semantically aligned with the vehicle's intent vector to obtain a fused feature vector. The visual representation features of each candidate target and the vehicle's intent vector are analyzed to obtain the intent matching score for each candidate target. Based on the intent matching score and rule score of each candidate target, pseudo-label data is constructed, which includes the predicted probability that each candidate target belongs to the defined target. Based on the pseudo-label data and the recognition results output by the fusion classifier in analyzing the fused feature vector, the parameters of the fusion classifier are adjusted to obtain a trained fusion classifier, which can then be used to identify the defined target. In this application, based on the selection of a candidate target set, a semantic alignment mechanism between the vehicle's intent and the candidate targets is established through intent guidance to address the problem of discriminating between multiple candidate targets in the same scene that are related to the vehicle's behavior but have different priorities. A pseudo-label-based weakly supervised training mechanism fuses the rule scores and intent matching scores of candidate targets to generate a pseudo-label distribution. By weakly supervising the fusion classifier, the difficulties in constructing datasets and the scarcity of true labels in object detection and recognition tasks are solved. Furthermore, the pseudo-label-based weakly supervised training mechanism improves the generalization ability of the fusion classifier in complex environments. The trained fusion classifier can accurately identify the designated targets. Attached Figure Description
[0012] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart illustrating a target recognition method provided for an embodiment of this application; Figure 2 A flowchart illustrating a method for selecting a target candidate set from a feature map set, provided in an embodiment of this application; Figure 3 This application provides a schematic diagram of a target traffic light selection process. Figure 4 This is a schematic diagram of the structure of a target recognition device provided in an embodiment of this application. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0015] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0016] While significant progress has been made in traffic light recognition and behavioral decision guidance, current solutions still have significant shortcomings in real-world urban road environments. Firstly, in traffic light recognition, complex urban intersections often present multiple traffic light targets, which are small, distant, and obstructed. Current methods struggle to accurately identify the target traffic light by considering the vehicle's intent—the traffic light the vehicle should follow. Furthermore, attention is frequently diverted to multiple irrelevant light states during traffic light recognition, leading to misjudgments and delayed responses. Secondly, regarding the coordination between perception and decision-making modules, most system modules are fragmented, lacking effective information sharing mechanisms. This results in a lack of intent and behavioral guidance for the detection and recognition of the primary target.
[0017] To enhance the spatial understanding capabilities of perception systems, bird's-eye view (BEV) has become the mainstream form of perception modeling. BEV constructs a scale-consistent, geometrically aligned spatial semantic representation by projecting the views of multiple cameras onto a unified ground plane. Despite BEV's advantages in spatial semantic representation, an effective and systematic solution remains lacking for efficiently and reliably extracting key traffic light states from BEV feature maps and driving behavior prediction.
[0018] Therefore, this application provides a target recognition method, apparatus, device, storage medium, and product. By setting target selection rules based on the correlation between targets and autonomous vehicle driving behavior, a candidate target set is selected from the feature map set according to the set rules, thereby eliminating irrelevant targets in the feature map set. Based on this, a semantic alignment mechanism between the autonomous vehicle's intent and candidate targets is established through intent guidance to solve the problem of discriminating multiple candidate targets in the same scene that are related to the autonomous vehicle's behavior but have different priorities. A pseudo-label-based weakly supervised training mechanism fuses the rule scores and intent matching scores of candidate targets to generate a pseudo-label distribution. By weakly supervising the fusion classifier, the problems of difficult dataset construction and scarce real labels in target detection and recognition tasks are solved. Furthermore, the pseudo-label-based weakly supervised training mechanism improves the generalization ability of the fusion classifier in complex environments, enabling accurate identification of the set targets using the trained fusion classifier.
[0019] The small targets mentioned in this application refer to targets that are significantly smaller than the average target size in the image space or BEV space, and are prone to losing semantic structure in the feature map, making them difficult to detect. Small targets typically include, but are not limited to, structural units that are small in size, have weak texture, and are frequently occluded, such as traffic lights, traffic cones, road signs, warning signs, and temporary obstacles. Traffic lights will be used as an example in the following descriptions.
[0020] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] Figure 1 A flowchart of a target recognition method provided for embodiments of this application includes: S101: Select a set of target candidates from the feature map set according to the set target selection rules.
[0022] Among them, the target selection rules can be set based on the correlation between the set target and the driver's behavior.
[0023] In practical applications, feature maps can be extracted from bird's-eye view feature maps. Bird's-eye view feature maps, also known as BEV feature maps, are feature maps obtained by encoding multi-camera images through bird's-eye view encoding.
[0024] The target candidate set contains each candidate target and its corresponding rule score.
[0025] In the embodiments of this application, the description is based on identifying the traffic lights that the vehicle should follow. The target is the target traffic light, which refers to the traffic light that the vehicle should follow.
[0026] To effectively narrow the search space, eliminate irrelevant lights, and improve the recognition efficiency of target traffic lights, target filtering rules can be set based on the correlation between traffic lights and driving behavior.
[0027] Target filtering rules can include type filtering and spatial relationship judgment. By comprehensively considering the type of traffic light and its spatial relationship with the vehicle, interfering targets unrelated to the vehicle's behavior can be quickly eliminated. A feasible filtering method can be found in [reference needed]. Figure 2 The details of its introduction will not be repeated here.
[0028] S102: Semantically align the feature information of each candidate target in the target candidate set with the vehicle's intent vector to obtain a fused feature vector.
[0029] Based on the selection of target candidate sets, a semantic alignment mechanism between the vehicle's intent and candidate traffic lights is established through an intent guidance module to address the problem of distinguishing between multiple candidate lights that are related to the vehicle's behavior but have different priorities in the same scenario. Priority refers to the degree of influence of multiple candidate lights on the vehicle's behavior, with the traffic light that the vehicle needs to follow during driving having the highest priority.
[0030] To obtain a more expressive intent vector for the vehicle, the vehicle's trajectory point sequence is used as input. The encoder in a deep learning model (Transformer) based on a multi-head self-attention mechanism obtains the vehicle's intent vector. By concatenating the feature information of each candidate target with the vehicle's intent vector, a fused feature vector is obtained.
[0031] In practical implementation, the trajectory point sequence of the vehicle can be extracted through a path planning system or a multi-mode trajectory prediction module: .
[0032] In the trajectory encoder, to obtain a more expressive intent vector, the trajectory point sequence G of the vehicle is used as input, and the vehicle intent embedding vector is obtained through the Transformer encoder. The operation performed by the Transformer encoder can be denoted as IntentEncoder: Among them, F intent This represents the vehicle's intention vector.
[0033] The process of constructing an IntentEncoder is as follows: First, the trajectory point sequence G is mapped into an embedding vector sequence. ;in, , This represents the spatial location of the candidate target, and LinearEmbedding represents linear embedding.
[0034] By adding location encoding and injecting location information, the model can perceive the temporal order of points. ;in, This indicates the position code.
[0035] The input is fed into a multi-layer Transformer Encoder for self-attention modeling to capture dependencies within the sequence. Finally, a fixed-dimensional intent vector is generated through pooling operations or by taking the end token. .
[0036] Extracting feature information from each candidate target in the candidate target set can include visual representation features. Cosine value of direction angle The relative distance between the candidate target and the vehicle Among them, P ego Indicates the spatial position of the vehicle.
[0037] The vehicle's intent vector is fused with the feature information of each candidate target to obtain the intent vector of each candidate target.
[0038] ; where z i This represents the fusion vector obtained by splicing, which combines candidate light information and intent vector.
[0039] Then, the intent vector of the candidate lights is generated: ; in, The linear mapping weight matrix will be used to fuse the vector z. i Projected onto the query space, Let be the intention vector of the i-th candidate light.
[0040] Based on the key-value pair information at each location in the bird's-eye view feature map, the intent vectors of each candidate target are aligned with the intent vector of the vehicle to obtain a fused feature vector.
[0041] To enhance the alignment between candidate lights and the vehicle's intent, a cross-attention mechanism is employed to guide the perception layer to focus on regions relevant to the current intent, and a fused feature vector is calculated. ; ; Among them, F i Represents the fused feature vector. This represents the intent vector of the i-th candidate light. This represents the key at each position in the BEV feature map. This represents the value at each position in the corresponding BEV feature map.
[0042] Cross-Attention can employ a multi-head mechanism, allowing different intent subspaces to participate in attention computation separately.i This reflects the correlation between the candidate lights and the vehicle's intention vector.
[0043] S103: Analyze the visual representation features of each candidate target and the vehicle's intent vector to obtain the intent matching score for each candidate target.
[0044] In the specific implementation, the visual representation features of each candidate target and the vehicle's intent vector can be concatenated to obtain a concatenated vector; the concatenated vector is analyzed using multiple fully connected layers, and the analysis results are mapped to intent matching scores.
[0045] For each traffic light L i Its intent matching score Visual representation features from its BEV feature map and the vehicle's intent vector get: ; in, Indicating concatenation operations or other feature fusion methods, the Sigmoid function maps its output to... An interval represents the probability of a match.
[0046] S104: Construct pseudo-label data based on the intent matching score and rule score of each candidate target.
[0047] The pseudo-label data includes the predicted probability that each candidate target belongs to the set target.
[0048] In this embodiment, the intent matching score and rule score of each candidate target can be probabilistically normalized to obtain the predicted probability that each candidate target belongs to the set target. That is, the pseudo-label data is a normalized probability distribution.
[0049] To further improve the generalization ability of the model, this application proposes a rule-based scoring method. Intent matching score The fusion pseudo-label generation strategy uses pseudo-labels derived from a normalized probability distribution. ; in, This represents the predicted probability that candidate target i belongs to the set target. It's a weighting hyperparameter, which can be set to... N is the number of targets.
[0050] S105: Based on the pseudo-label data and the recognition results output by the fusion classifier in analyzing the fusion feature vector, the parameters of the fusion classifier are adjusted to obtain a trained fusion classifier, so as to use the trained fusion classifier to identify the set target.
[0051] After obtaining the fused feature vector, the fused feature vector can be input into the fusion classifier to determine the probability that each candidate target belongs to the set target.
[0052] Taking candidate traffic lights as an example, a fusion classifier performs joint inference on all candidate traffic lights, outputting the confidence score of each candidate traffic light as the "target traffic light". First, all fused feature vectors, i.e., the intent-guided perception output F, are then processed. i The input is fed into a fusion classifier, which outputs the probability that each candidate traffic light is the target traffic light. : ; In this context, Classifier consists of one or more fully connected layers, and Softmax transforms the classification results into normalized probabilities.
[0053] The fusion classifier supports information exchange and semantic comparison between different candidate traffic lights, improving the robustness of discrimination under complex intersection conditions such as multiple mixed lights and occlusion interference.
[0054] In practice, a cross-entropy loss function can be constructed based on the probability that each candidate target belongs to the set target and the predicted probability that each candidate target belongs to the set target. The parameters of the fusion classifier can be adjusted using the cross-entropy loss function to obtain a trained fusion classifier.
[0055] The fusion classifier is trained using the following cross-entropy loss function: ; Among them, L cls This represents the cross-entropy loss function.
[0056] In this embodiment, a pseudo-label-based weakly supervised training mechanism weighted and fused the rule scores and intent matching scores of candidate lights to generate a pseudo-label distribution. By weakly supervising the training of the fusion classifier, the problem of difficulty in constructing traffic light datasets and the scarcity of true labels in object detection and recognition tasks is solved. Furthermore, the pseudo-label-based weakly supervised training mechanism improves the model's generalization ability in unstructured intersections / complex environments. By limiting non-differentiable rules to the candidate light set selection and pseudo-label generation stages, and using weak supervision to train differentiable models, the problem of non-differentiable rules directly interfering with gradient propagation is effectively avoided, facilitating end-to-end training of the system guided by planning.
[0057] As can be seen from the above technical solution, a target candidate set is selected from the feature map set according to the set target selection rules. The target selection rules are set based on the correlation between the set target and the vehicle's driving behavior. The target candidate set includes each candidate target and its corresponding rule score. By setting the target selection rules, irrelevant targets in the feature map set can be excluded. To increase the accuracy of target selection, the feature information of each candidate target in the target candidate set can be semantically aligned with the vehicle's intent vector to obtain a fused feature vector. The visual representation features of each candidate target and the vehicle's intent vector are analyzed to obtain the intent matching score of each candidate target. Based on the intent matching score and rule score of each candidate target, pseudo-label data is constructed; the pseudo-label data contains the predicted probability that each candidate target belongs to the set target. Based on the pseudo-label data and the recognition results output by the fusion classifier in analyzing the fused feature vector, the parameters of the fusion classifier are adjusted to obtain a trained fusion classifier, which can then be used to identify the set target. In this application, based on the selection of a candidate target set, a semantic alignment mechanism between the vehicle's intent and the candidate targets is established through intent guidance to address the problem of discriminating between multiple candidate targets in the same scene that are related to the vehicle's behavior but have different priorities. A pseudo-label-based weakly supervised training mechanism fuses the rule scores and intent matching scores of candidate targets to generate a pseudo-label distribution. By weakly supervising the fusion classifier, the difficulties in constructing datasets and the scarcity of true labels in object detection and recognition tasks are solved. Furthermore, the pseudo-label-based weakly supervised training mechanism improves the generalization ability of the fusion classifier in complex environments. The trained fusion classifier can accurately identify the designated targets.
[0058] Figure 2 A flowchart of a method for filtering a target candidate set from a feature map set, provided in an embodiment of this application, is included. S201: Extract feature map sets from bird's-eye view feature maps using parallel branch networks.
[0059] Among them, the bird's-eye view feature map is a feature map obtained by encoding multiple camera images through bird's-eye view encoding.
[0060] The feature map set may include the feature map set corresponding to the first target and the feature map set corresponding to the second target, wherein the number of pixels of the first target is greater than the number of pixels of the second target.
[0061] In a specific implementation, the first branch network can be used to extract features from the bird's-eye view feature map to obtain the first target feature map set corresponding to each first target; the second branch network can be used to extract features from the bird's-eye view feature map to obtain the second target feature map set corresponding to each second target; the first target feature map set and the second target feature map set can be fused to obtain the feature map set.
[0062] The first target feature map set contains the first structured information corresponding to each first target, and the second target feature map set contains the second structured information corresponding to each second target.
[0063] To facilitate a clear distinction between the first and second targets, the first target can be referred to as the large-size target, or simply the large target; the second target can be referred to as the small-size target, or simply the small target. The large target may include vehicles, lane markings, etc. The small target may include traffic lights and traffic signs. In this application's embodiments, vehicles and traffic lights are used as examples for description.
[0064] To improve the detection accuracy of small targets, two parallel target detection heads can be used to extract the structured information of small and large targets respectively.
[0065] This application employs two parallel detection heads to extract structured information for small and large targets respectively. These two detection heads can be referred to as the small target detection head (Head_Small) and the large target detection head (Head_Large). The Head_Large is used to detect larger target objects, while the Head_Small is used to detect small targets with fine-grained textures.
[0066] It should be noted that BEV feature maps can have different hierarchical structures in different model architectures. In a backbone structure with multi-scale BEV output capability, the two detector heads mentioned above can connect to feature layers of different resolutions from the BEV encoder. Head_Large connects to low-resolution feature layers, such as layers four and five, providing a larger receptive field and suitable for perceiving large targets. Head_Small connects to high-resolution feature layers, such as layers two and three, preserving spatial details and facilitating the modeling of small targets.
[0067] The detection head is a neural network module deployed within the software model. In the BEV perception system, it receives BEV feature maps as input and uses multiple branch networks to extract structured information about targets at a specific scale. For ease of distinction, the branch network corresponding to the large target detection head can be called the first branch network, and the branch network corresponding to the small target detection head can be called the second branch network.
[0068] The bird's-eye view feature map contains multiple first targets, each with its corresponding first structured information. The set of all first structured information is called the first target feature map set. The bird's-eye view feature map contains multiple second targets, each with its corresponding second structured information. The set of all second structured information is called the second target feature map set.
[0069] In practical applications, there may be conflicting outputs between large target detection heads and small target detection heads, meaning that the first target detected by the large target detection head and the second target detected by the small target detection head may have overlapping regions.
[0070] In this embodiment, a region attention suppression mechanism is introduced to handle conflicts in the output results of the target detection head for different sizes.
[0071] In the spatial dimension, the system constructs a target region with the center of each detected target as the center and a set radius to detect spatial conflicts between the first and second targets. When the predicted target centers of two detectors simultaneously fall within an overlapping region, the system performs differentiated processing based on the target category.
[0072] Taking traffic lights as an example, for targets that simultaneously fall within a certain overlapping area, the target is classified as a traffic light target, and only the results from the small target detection head are retained. For non-traffic light targets, a non-maximum suppression algorithm (Soft-NMS) is introduced to perform confidence-weighted fusion. By attenuating the confidence, redundant boxes are effectively suppressed, reducing the misclassification of small targets. At the same time, it avoids the accidental deletion of valid targets in multi-target close-range scenes, which is particularly suitable for dense target detection scenes, such as traffic lights or areas where vehicles gather.
[0073] If the two target regions do not overlap spatially, the feature map sets of the first target and the feature map set of the second target are directly combined as the feature map set.
[0074] In complex intersection scenarios, especially when multiple traffic lights coexist, some are obscured, or have inconsistent directions, current traffic light recognition methods often suffer from inaccurate target identification, missed detections, and false detections. Particularly for detecting distant, small traffic lights, current methods lack dedicated optimization mechanisms, resulting in low detection accuracy and impacting the accuracy of subsequent decision-making and planning. This application addresses this issue by designing two parallel detection heads to construct independent traffic light channels in the BEV feature map. By combining information such as the height, direction angle, and relative distance of the traffic lights to the vehicle, accurate small target detection is achieved, improving the robustness and accuracy of target traffic light recognition.
[0075] S202: Based on the target type of each second target in the feature map set, select the first candidate set belonging to the set target from the feature map set.
[0076] The first candidate set contains structured information about the target set in the feature map set.
[0077] Taking a traffic light as an example, the structured information corresponding to the second target in the feature map set that belongs to the traffic light is used as the first candidate set.
[0078] For example, the current ego coordinates are... The vehicle's heading angle is The selected traffic lights must serve motor vehicles and can be identified by traffic light type labels parsed from BEV feature maps or sensor supplementary structures. ; in, This represents a function for determining the type of traffic light, used to classify traffic light I. i Does it fall under the category of "motor vehicle lights"?
[0079] S203: Based on the spatial relationship between each second target and the vehicle, select the second candidate set belonging to the set target from the feature map set.
[0080] Spatial relationships can include positional relationships and angular relationships.
[0081] In practical implementation, the positional relationship between each second target in the feature map set and the vehicle's driving area can be used to filter out a candidate set of second targets belonging to the vehicle's driving area from the feature map set. Then, the angular relationship between each second target in the feature map set and the vehicle can be used to filter out a candidate set of directions for second targets matching the vehicle's direction from the feature map set.
[0082] For the selection of region candidate sets, the angle between each second target and the vehicle can be determined based on the deviation between the spatial position of each second target in the feature map set and the position of the vehicle, as well as the unit vector of the vehicle's forward direction. The feature map set corresponding to the number of second targets with the smallest angle to the vehicle in the feature map set is selected as the region candidate set.
[0083] The regional candidate set contains structured information about traffic lights within the vehicle's driving area.
[0084] In practical applications, the selected traffic light must fall within the near-field area of the vehicle's forward travel path. This is achieved by projecting the traffic light position vector onto the unit vector of the vehicle's forward direction. Above, if the spatial position of the traffic light is The location of the bicycle is The forward unit vector of the vehicle is ,but ;in, This indicates the angle between the traffic light and the vehicle.
[0085] According to The traffic lights are sorted by their values. A positive value indicates that the light is in front of the vehicle. The smaller the value, the closer the traffic light is to the vehicle.
[0086] In practical applications, all can be retained. And the k lights with the smallest values, i.e. ; in, This function determines whether the projected traffic light is in front of the vehicle; a positive value indicates that the light is in front of the vehicle.
[0087] For the selection of direction candidate sets, the direction difference between each second target and the vehicle can be determined based on the orientation angle of each second target in the feature map set and the current direction of the vehicle. The feature map set corresponding to the second target in the feature map set whose direction difference is less than the set angle threshold is used as the direction candidate set.
[0088] The direction candidate set contains structured information about traffic lights that are in the same direction as the vehicle's current direction or whose direction difference is less than a set angle threshold.
[0089] In practical applications, the directional angle of traffic lights Should be in the current direction of the vehicle By maintaining a consistent angle or ensuring the angle difference does not exceed a set threshold, traffic lights facing other road traffic directions can be removed. The difference between the traffic light direction and the vehicle's direction is defined as: ;in, This represents the difference in direction between the i-th traffic light and the vehicle.
[0090] ; in, A function to determine whether the direction of a traffic light is consistent with the direction of a vehicle. For the set angle threshold, satisfy This means the traffic light is considered to be aligned with the direction of the vehicle. Considering that in some road conditions the traffic light is not directly aligned with the stop line at the intersection, it can be... Set as .
[0091] S204: Normalize the first and second candidate sets to determine the rule scores corresponding to each second objective.
[0092] The second candidate set includes a region candidate set and a direction candidate set. In a specific implementation, the type candidate set, region candidate set, and direction candidate set can be normalized according to the following formula: ;in, This represents the rule score corresponding to the second objective.
[0093] In this embodiment, by quantizing the three types of rules into a weighted scoring function, the output is a ranking value instead of a discrete label. Compared with traditional binary rule judgment, the scoring function retains the ranking information, which is convenient for the learning module to access. It supports the construction of weakly supervised pseudo-labels and the gradient propagation between rules and deep learning models.
[0094] S205: Use the feature map set corresponding to the second target with the highest rule score as the target candidate set.
[0095] In practical applications, it can be retained The top k highest-ranking traffic lights are used as the target candidate set, denoted as . This set serves as the input basis for subsequent modules.
[0096] In this embodiment, to balance the accuracy and real-time performance of large and small target detection in complex environments, two parallel detection heads are used to extract structured information of small and large targets from the bird's-eye view feature map, respectively. This enables independent branch modeling for targets of different scales, improving the overall system detection performance and hardware execution efficiency. Furthermore, rule-based filtering eliminates light entities irrelevant to the vehicle's current decision, reducing the search space and enhancing the system's real-time performance and interpretability.
[0097] After obtaining the trained fusion classifier, if a new feature map set is acquired, a new set of target candidates can be selected from the new feature map set according to the set target selection rules. The feature information of each new candidate target in the new target candidate set is semantically aligned with the vehicle's intent vector to obtain the target fusion feature vector. The trained fusion classifier is then used to analyze the target fusion feature vector to determine the probability that each new candidate target belongs to the set target; the new candidate target with the highest probability is selected as the target associated with the vehicle.
[0098] Figure 3This document provides a flowchart illustrating the selection process for a target traffic light, where the target traffic light refers to the traffic light that the vehicle must follow during its journey. Based on the operations involved in target traffic light selection, the entire implementation process can be divided into different functional modules, including a parallel target detection head, a rule filtering module, an intent guidance module, a fusion classifier, trajectory prediction, and a trajectory encoder. The parallel target detection head extracts structured information for small and large targets from the bird's-eye view feature map. To avoid class conflicts or misjudgments caused by multiple outputs, the parallel target detection head fuses the first and second target feature maps to obtain a feature map set. The rule filtering module excludes traffic lights in the feature map set that are irrelevant to the vehicle, thereby determining a candidate traffic light set, which includes each candidate traffic light and its corresponding rule score. Trajectory prediction extracts the vehicle's trajectory point sequence. The trajectory encoder encodes the vehicle's trajectory point sequence into an intent vector. The intent guidance module semantically aligns the structured information of the candidate traffic lights with the vehicle's intent vector to obtain a fused feature vector. The fusion classifier analyzes the probability that each candidate traffic light in the fused feature vector belongs to the target traffic light, and selects the candidate traffic light with the highest probability as the target traffic light associated with the vehicle. The output of the fusion classifier influences the vehicle's driving path, thus generating a new sequence of trajectory points. During the training phase of the fusion classifier, a pseudo-label generation module is included to generate pseudo-labels based on candidate light rule scores and intent matching scores. Weakly supervised training based on these pseudo-labels is used to train the fusion classifier. Using the trained fusion classifier, the target traffic light can be identified relatively accurately.
[0099] The target recognition scheme proposed in this application can be directly deployed with edge computing products. The model deployment adopts a modular deployment strategy, including structural units such as a perception backbone (BEV encoder), a small target detection head, a rule filtering module, and an intent guidance module. These units undergo graph fusion and quantization processing using edge-side inference optimization tools to compress the model size and improve inference efficiency. During deployment, the model receives image input from multiple vehicle-mounted cameras, generates a BEV feature map after front-end perception processing, and then extracts small target information and identifies the target traffic light using the structure proposed in this application. The target traffic light status and position are written as structured results to the edge cache for real-time reading by the downstream path planning and control module. The overall system inference latency is controlled within 50ms, meeting the perception response requirements of edge deployment.
[0100] Figure 4 A schematic diagram of the structure of a target recognition device provided in an embodiment of this application includes a filtering unit 41, an alignment unit 42, a matching unit 43, a construction unit 44, and an adjustment unit 45; The filtering unit 41 is used to filter out a set of target candidates from the feature map set according to the set target filtering rules; wherein, the target filtering rules are set according to the correlation between the set target and the autonomous vehicle driving behavior; the target candidate set includes each candidate target and its corresponding rule score; Alignment unit 42 is used to semantically align the feature information of each candidate target in the target candidate set with the vehicle intent vector to obtain a fused feature vector; Matching unit 43 is used to analyze the visual representation features of each candidate target and the vehicle's intent vector to obtain the intent matching score of each candidate target. Construction unit 44 is used to construct pseudo-label data based on the intent matching score and rule score of each candidate target; wherein, the pseudo-label data includes the predicted probability that each candidate target belongs to the set target; The adjustment unit 45 is used to adjust the parameters of the fusion classifier based on the pseudo-label data and the recognition results output by the fusion classifier in analyzing the fusion feature vector, so as to obtain a trained fusion classifier, so as to use the trained fusion classifier to identify the set target.
[0101] In some embodiments, the filtering unit includes an extraction subunit, a type matching subunit, a spatial relationship matching subunit, a normalization subunit, and a filter subunit; The extraction subunit is used to extract a feature map set from the bird's-eye view feature map using a parallel branch network; wherein, the bird's-eye view feature map is a feature map obtained by encoding multiple camera images through bird's-eye view; the feature map set includes the feature map set corresponding to the first target and the feature map set corresponding to the second target, wherein the number of pixels of the first target is greater than the number of pixels of the second target; The type matching subunit is used to filter out the first candidate set belonging to the set target from the feature map set according to the target type to which each second target in the feature map set belongs; The spatial relationship matching subunit is used to filter out the second candidate set belonging to the set target from the feature map set based on the spatial relationship between each second target and the vehicle. The normalization subunit is used to normalize the first candidate set and the second candidate set in order to determine the rule score corresponding to each second objective. As a sub-unit, it is used to select the feature map set corresponding to the second target with the highest rule score as the target candidate set.
[0102] In some embodiments, the extraction subunit is used to extract features from the bird's-eye view feature map using the first branch network to obtain a first target feature map set corresponding to each first target; The second branch network is used to extract features from the bird's-eye view feature map to obtain the second target feature map set corresponding to each second target; The first target feature map set and the second target feature map set are fused to obtain a feature map set.
[0103] In some embodiments, the spatial relationship matching subunit is used to filter out a candidate set of regions belonging to the vehicle's driving area from the feature map set by utilizing the positional relationship between each second target in the feature map set and the vehicle's driving area. By utilizing the angular relationship between each second target in the feature map set and the vehicle, a candidate set of directions for the second targets that match the vehicle's direction is selected from the feature map set.
[0104] In some embodiments, the spatial relationship matching subunit is used to determine the angle between each second target and the vehicle based on the deviation between the spatial position of each second target in the feature map set and the position of the vehicle and the unit vector of the vehicle's forward direction. The feature map set corresponding to the second target with the smallest angle to the vehicle in the feature map set is used as the region candidate set.
[0105] In some embodiments, the spatial relationship matching subunit is used to determine the direction difference between each second target and the vehicle based on the orientation angle of each second target in the feature map set and the current direction of the vehicle. The feature map set corresponding to the second target whose direction difference in the feature map set is less than the set angle threshold is used as the direction candidate set.
[0106] In some embodiments, the alignment unit includes an encoding subunit, an extraction subunit, a first fusion subunit, and a second fusion subunit; The encoding subunit is used to encode the trajectory point sequence of the vehicle into the vehicle's intention vector; The extraction subunit is used to extract the feature information of each candidate target from the target candidate set; the feature information includes visual representation features, direction angle cosine value, and relative distance between the candidate target and the vehicle; The first fusion subunit is used to fuse the vehicle's intent vector with the feature information of each candidate target to obtain the intent vector of each candidate target. The second fusion subunit is used to align the intent vector of each candidate target with the intent vector of the vehicle based on the key-value pair information at each location in the bird's-eye view feature map, so as to obtain the fused feature vector.
[0107] In some embodiments, the matching unit includes a splicing subunit and an analysis subunit; The splicing subunit is used to splice the visual representation features of each candidate target and the vehicle's intent vector to obtain a spliced vector; The analysis subunit is used to analyze the spliced vector using multiple fully connected layers and map the analysis results to an intent matching score.
[0108] In some embodiments, the construction unit is used to perform probability normalization processing on the intent matching score and rule score of each candidate target to obtain the predicted probability that each candidate target belongs to the set target.
[0109] In some embodiments, the adjustment unit is used to input the fused feature vector into the fusion classifier to determine the probability that each candidate target belongs to a set target; Based on the probability that each candidate target belongs to the set target and the predicted probability that each candidate target belongs to the set target, a cross-entropy loss function is constructed. The parameters of the fusion classifier are adjusted using the cross-entropy loss function to obtain a trained fusion classifier.
[0110] In some embodiments, it further includes an analysis unit and an application unit; The filtering unit is used to filter out a new set of target candidates from the new set of feature maps according to the set of target filtering rules when a new set of feature maps is obtained. The alignment unit is used to semantically align the feature information of each new candidate target in the new target candidate set with the vehicle's intent vector to obtain the target fusion feature vector; The analysis unit is used to analyze the target fusion feature vector using a trained fusion classifier to determine the probability that each new candidate target belongs to the set target. As a unit, it is used to identify the new candidate target with the highest probability as the target associated with the vehicle.
[0111] For a description of the features in the embodiment corresponding to the target recognition device, please refer to the relevant description in the embodiment corresponding to the target recognition method, which will not be repeated here.
[0112] As can be seen from the above technical solution, a target candidate set is selected from the feature map set according to the set target selection rules. The target selection rules are set based on the correlation between the set target and the vehicle's driving behavior. The target candidate set includes each candidate target and its corresponding rule score. By setting the target selection rules, irrelevant targets in the feature map set can be excluded. To increase the accuracy of target selection, the feature information of each candidate target in the target candidate set can be semantically aligned with the vehicle's intent vector to obtain a fused feature vector. The visual representation features of each candidate target and the vehicle's intent vector are analyzed to obtain the intent matching score of each candidate target. Based on the intent matching score and rule score of each candidate target, pseudo-label data is constructed; the pseudo-label data contains the predicted probability that each candidate target belongs to the set target. Based on the pseudo-label data and the recognition results output by the fusion classifier in analyzing the fused feature vector, the parameters of the fusion classifier are adjusted to obtain a trained fusion classifier, which can then be used to identify the set target. In this application, based on the selection of a candidate target set, a semantic alignment mechanism between the vehicle's intent and the candidate targets is established through intent guidance to address the problem of discriminating between multiple candidate targets in the same scene that are related to the vehicle's behavior but have different priorities. A pseudo-label-based weakly supervised training mechanism fuses the rule scores and intent matching scores of candidate targets to generate a pseudo-label distribution. By weakly supervising the fusion classifier, the difficulties in constructing datasets and the scarcity of true labels in object detection and recognition tasks are solved. Furthermore, the pseudo-label-based weakly supervised training mechanism improves the generalization ability of the fusion classifier in complex environments. The trained fusion classifier can accurately identify the designated targets.
[0113] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described target recognition method embodiments.
[0114] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described target recognition method embodiments at runtime.
[0115] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0116] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described target recognition method embodiments.
[0117] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described target recognition method embodiments.
[0118] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0119] The above provides a detailed description of a target identification method, apparatus, device, storage medium, and product provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A target recognition method characterized by, The method comprises the following steps: screening a target candidate set from a feature map set according to a set target screening rule; wherein the target screening rule is set according to the relevance between the set target and the driving behavior of the ego vehicle; the target candidate set comprises each candidate target and the corresponding rule score of the candidate target; performing semantic alignment on the feature information of each candidate target in the target candidate set and the ego vehicle intention vector to obtain a fusion feature vector; analyzing the visual representation feature of each candidate target and the ego vehicle intention vector to obtain an intention matching score of each candidate target; constructing pseudo-label data based on the intention matching score of each candidate target and the rule score; wherein the pseudo-label data comprises the predicted probability of each candidate target belonging to the set target; adjusting the parameters of the fusion classifier according to the pseudo-label data and the recognition result output by the fusion classifier after analyzing the fusion feature vector, so as to obtain a trained fusion classifier, so as to identify the set target by using the trained fusion classifier.
2. The object recognition method of claim 1, wherein, Screening a target candidate set from a feature map set according to a set target screening rule comprises the following steps: extracting a feature map set from a bird's eye view feature map by using a parallel branch network; wherein the bird's eye view feature map is a feature map obtained by encoding a bird's eye view of multiple camera images; the feature map set comprises a feature map set corresponding to a first target and a feature map set corresponding to a second target, and the number of pixels of the first target is greater than that of the second target; screening a first candidate set belonging to a set target from the feature map set according to the target type to which each second target in the feature map set belongs; screening a second candidate set belonging to a set target from the feature map set according to the spatial relationship between each second target and the ego vehicle; normalizing the first candidate set and the second candidate set to determine the rule score corresponding to each second target; taking the feature map set corresponding to a set number of second targets with the highest rule score as the target candidate set.
3. The object recognition method of claim 2, wherein, Extracting a feature map set from a bird's eye view feature map by using a parallel branch network comprises the following steps: extracting features from the bird's eye view feature map by using a first branch network to obtain a first target feature map set corresponding to each first target; extracting features from the bird's eye view feature map by using a second branch network to obtain a second target feature map set corresponding to each second target; fusing the first target feature map set and the second target feature map set to obtain a feature map set.
4. The object recognition method of claim 2, wherein, Screening a second candidate set belonging to a set target from the feature map set according to the spatial relationship between each second target and the ego vehicle comprises the following steps: screening a region candidate set of second targets belonging to the driving area of the ego vehicle from the feature map set by using the positional relationship between each second target in the feature map set and the driving area of the ego vehicle; screening a direction candidate set of second targets matching the direction of the ego vehicle from the feature map set by using the angle relationship between each second target in the feature map set and the ego vehicle.
5. The object recognition method of claim 4, wherein, Screening a region candidate set of the second targets belonging to the self-vehicle driving region from the feature map set according to the positional relationship between each of the second targets in the feature map set and the self-vehicle driving region comprises: Determining the angle between each of the second targets and the self-vehicle according to the spatial position of each of the second targets in the feature map set and the positional deviation of the self-vehicle and the unit vector of the forward direction of the self-vehicle; Taking the feature map set corresponding to a set number of second targets with the smallest angle to the self-vehicle in the feature map set as the region candidate set.
6. The object recognition method of claim 4, wherein, Screening a direction candidate set of the second targets matching the direction of the self-vehicle from the feature map set according to the angle relationship between each of the second targets in the feature map set and the self-vehicle comprises: Determining the direction difference between each of the second targets and the self-vehicle according to the orientation angle of each of the second targets in the feature map set and the current direction of the self-vehicle; Taking the feature map set corresponding to the second targets with a direction difference less than a set angle threshold in the feature map set as the direction candidate set.
7. The object recognition method of claim 1, wherein, Aligning the feature information of each candidate target in the target candidate set with the self-vehicle intention vector to obtain a fusion feature vector comprises: Encoding the trajectory point sequence of the self-vehicle into a self-vehicle intention vector; Extracting the feature information of each candidate target from the target candidate set; wherein the feature information comprises visual representation features, direction angle cosine values, and the relative distance between the candidate target and the self-vehicle; Fusing the self-vehicle intention vector with the feature information of each candidate target to obtain an intention vector of each candidate target; Aligning the intention vector of each candidate target with the self-vehicle intention vector according to the key-value pair information of each position in the bird's eye view feature map to obtain a fusion feature vector.
8. The object recognition method of claim 1, wherein, Analyzing the visual representation features of each candidate target and the self-vehicle intention vector to obtain an intention matching score of each candidate target comprises: Splicing the visual representation features of each candidate target and the self-vehicle intention vector to obtain a spliced vector; Analyzing the spliced vector using multiple fully connected layers and mapping the analysis result as an intention matching score.
9. The object recognition method of claim 8, wherein, Based on the intention matching score of each candidate target and the rule score, constructing pseudo-label data comprises: Probability normalization processing the intention matching score of each candidate target and the rule score to obtain a prediction probability of each candidate target belonging to a set target.
10. The object recognition method of claim 9, wherein, According to the recognition result output by analyzing the fusion feature vector by the fusion classifier according to the pseudo-label data, adjusting the parameters of the fusion classifier to obtain a trained fusion classifier comprises: Inputting the fusion feature vector into the fusion classifier to determine the probability of each candidate target belonging to a set target; According to the probability of each candidate target belonging to a set target and the prediction probability of each candidate target belonging to a set target, constructing a cross-entropy loss function; Adjusting the parameters of the fusion classifier using the cross-entropy loss function to obtain a trained fusion classifier.
11. The object recognition method according to any one of claims 1 to 10, characterized in that, After obtaining the trained fusion classifier, further comprising: In the case of obtaining a new feature map set, a new target candidate set is screened from the new feature map set according to a set target screening rule; The feature information of each new candidate target in the new target candidate set is semantically aligned with the ego vehicle intention vector to obtain a target fusion feature vector; The target fusion feature vector is analyzed by using the trained fusion classifier to determine the probability of each new candidate target belonging to the set target; The new candidate target with the highest probability is taken as the target associated with the ego vehicle.
12. A target recognition device, characterized by The method comprises a screening unit, an alignment unit, a matching unit, a construction unit and an adjustment unit. The screening unit is configured to screen a target candidate set from a feature map set according to a set target screening rule; wherein the target screening rule is set according to the relevance between the set target and the ego vehicle driving behavior; and the target candidate set comprises each candidate target and its corresponding rule score. The alignment unit is configured to semantically align the feature information of each candidate target in the target candidate set with the ego vehicle intention vector to obtain a fusion feature vector. The matching unit is configured to analyze the visual representation features of each candidate target and the ego vehicle intention vector to obtain an intention matching score of each candidate target. The construction unit is configured to construct pseudo-label data based on the intention matching score and the rule score of each candidate target; wherein the pseudo-label data comprises the predicted probability of each candidate target belonging to the set target. The adjustment unit is configured to adjust the parameters of the fusion classifier according to the pseudo-label data and the recognition result output by analyzing the fusion feature vector by using the fusion classifier, so as to obtain a trained fusion classifier, so as to identify the set target by using the trained fusion classifier.
13. An electronic device, comprising: The computer readable storage medium stores a computer program, wherein the computer program is executed by the processor to implement the steps of the target recognition method according to any one of claims 1 to 11. The computer readable storage medium stores a computer program, wherein the computer program is executed by the processor to implement the steps of the target recognition method according to any one of claims 1 to 11. The computer program is executed by the processor to implement the steps of the target recognition method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, 15. A computer program product comprising a computer program, characterized in that,
Citation Information
Patent Citations
Target detection method and device based on scene semantics, equipment and storage medium
CN112966697A
Method and device for extracting implicit intention data from natural driving data set
CN116010862A
Visual relation identification method and device based on knowledge and data collaborative reasoning
CN117874697A
Traffic light detection method and device, vehicle and storage medium
CN117912280A
Emoji package retrieval method, electronic equipment and computer readable storage medium
CN118551068A