ATS-oriented semantic-driven single-image-based 3D multi-target detection method

By adopting a semantic-driven 3D multi-object detection method for single-image images based on ATS in 3D vision detection, combining natural language processing and multi-head self-attention mechanism, the problem of high-cost and insufficient semantic understanding in the existing technology is solved, and efficient and low-cost 3D multi-object detection is achieved.

CN120107845AActive Publication Date: 2025-06-06CHANGAN UNIV

Patent Information

Application Number
CN202510116192.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-06
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The prior art has problems such as high cost, high computing resource requirements and insufficient multi-objective semantic understanding in 3D vision detection, and it is difficult to promote and deploy in resource-constrained environments.

Method used

Using a semantic-driven 3D multi-object detection method for single-image images based on ATS, 3D bounding boxes and 2D projections are extracted through advanced 3D detectors, language features are extracted in combination with natural language processing technology, and effective correlation between language features and visual features is achieved through multi-head self-attention mechanism and cross-modal semantic alignment strategy.

Benefits of technology

It significantly lowers the technical threshold, improves the adaptability to complex traffic scenarios and the state perception of the carrier equipment, realizes efficient and low-cost 3D multi-object detection, and improves detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_3
    Figure SMS_3
  • Figure SMS_7
    Figure SMS_7
  • Figure SMS_8
    Figure SMS_8
Patent Text Reader

Abstract

The invention discloses an ATS-oriented semantic-driven single-image-based 3D multi-target detection method, and the method comprises the steps: processing an input RGB image, and extracting a 3D bounding box of each object in the image; generating all potential 2D projections of the 3D object in the scene; extracting keywords, phrases and semantic information thereof in the description to form feature information Pt representing language description; fusing the 2D image information and the 3D geometric information of the object to obtain a complete object representation fa; correlating the language features extracted by the language description with the detected 3D object, and capturing a semantic correspondence relationship between the text and the visual modality; and filtering the targets according to the generated matching scores to obtain all targets according with natural language description. According to the method, the accuracy and efficiency of retrieval and recognition are improved, higher recognition precision is realized, the calculation complexity is remarkably reduced, the precision and speed of cross-modal retrieval can be improved, and the accurate recognition and positioning capability of traffic events is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video detection technology, and in particular to a 3D multi-target detection method in a single image based on semantics-driven oriented ATS. Background Art

[0002] With the rapid development of artificial intelligence technology, how to make machines understand and associate natural language and visual information has become a key challenge in human-computer interaction and scene understanding. This technology has broad application potential in many fields and has promoted the intelligent development of various industries. In the field of autonomous driving, through natural language instructions, vehicles can identify and locate target objects on the road, optimize driving strategies, and improve driving safety and efficiency. Drivers can use voice commands to let vehicles identify obstacles, predict pedestrian behavior and plan routes, thereby achieving a higher level of autonomous driving. In the field of intelligent robots, robots can perform complex tasks according to natural language instructions, understand diverse contexts and respond. This technology can enhance the application scenarios of robots, such as home services, industrial manufacturing, etc., and improve work efficiency and flexibility. In addition, in intelligent traffic management, real-time monitoring and analysis technology can help management departments optimize traffic flow and reduce congestion and accidents. Through cross-modal data recognition, traffic conditions can be quickly transmitted, improving decision-making efficiency and optimizing resource allocation.

[0003] Existing detection technologies are mainly divided into two directions: 2D visual detection and 3D visual detection. However, existing technologies have many disadvantages, including:

[0004] (1) In terms of 2D visual detection, datasets mainly focus on 2D images. However, real-world scenes are inherently 3D, and 2D information alone is insufficient to capture object depth, spatial relationships, and structural geometry.

[0005] (2) 3D visual detection based on high-cost sensors such as LiDAR or millimeter-wave radar. Although these sensors can provide high-precision 3D point cloud data, they are expensive and require high computing resources. In resource-constrained environments (such as consumer devices or small robotic systems), these methods are difficult to promote and deploy in practice, limiting their wider applicability.

[0006] (3) Mono3DVG has explored language-guided monocular 3D detection, but its scope is limited to single object detection. However, in real scenes, there are usually multiple objects, and the semantic relationship between these objects is crucial for comprehensive scene understanding. Insufficient semantic understanding of multiple targets may lead to insufficient scene understanding capabilities, thus affecting the reliability of practical applications.

[0007] In view of the problem that the current 3D visual positioning method based on LiDAR is expensive and difficult to popularize, the present invention proposes an innovative multi-modal multi-target 3D perception technology based on ATS text guidance, which can be combined with natural language description to realize the accurate detection and positioning of multiple ATS carriers in a single image. The method extracts the 3D bounding box and its 2D projection of the ATS carrier through an advanced 3D detector, and extracts language features in combination with natural language processing technology. By fusing 2D image information with 3D geometric information, the method realizes a complete target representation. In addition, the present invention adopts a selective matching module, which fully realizes the effective association between language features and visual features through a multi-head self-attention mechanism and a cross-modal semantic alignment strategy, thereby significantly improving the detection accuracy and efficiency. This method gets rid of the dependence on expensive sensors, significantly reduces the technical threshold, and improves its adaptability to complex traffic scenes and the perception of the status of carriers. Especially in the fields of intelligent transportation, autonomous driving and augmented reality, it has broad application prospects.

[0008] Compared with the existing technology, the present invention demonstrates more outstanding multimodal fusion capabilities and cross-modal semantic understanding capabilities. It is an efficient, low-cost and easy-to-promote 3D target detection solution that can provide strong support for the perception and decision-making of autonomous transportation systems. Summary of the invention

[0009] In view of this, the present invention provides a 3D multi-target detection method in a single image based on semantic driving for ATS.

[0010] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0011] A 3D multi-target detection method based on semantic-driven single image for ATS, comprising the following steps:

[0012] Step 1: Process the input RGB image to extract the 3D bounding boxes of each object in the image; generate all potential 2D projections of the 3D object in the scene;

[0013] Step 2: By using natural language processing technology, extract the keywords, phrases and their semantic information in the description to form feature information P representing the language description. t ;

[0014] Step 3: Fuse the object’s 2D image information and 3D geometric information to obtain a complete object representation f a ;

[0015] Step 4: Associating the language features extracted from the language description with the detected 3D objects to capture the semantic correspondence between the text and visual modalities;

[0016] Step 5: Filter the targets according to the generated matching scores to obtain all targets that match the natural language description.

[0017] Preferably, the step three comprises the following steps:

[0018] Step 3.1: For each 3D object, use its corresponding 2D bounding box to crop the object area from the original RGB image and obtain the corresponding 2D image information;

[0019] Step 3.2: Using 2D image information, extract the visual features of 3D objects through a pre-trained network v , the size is 768 × the number of patches in the image, and the visual features f are processed by a multi-head self-attention mechanism. v Get fine features f' v , calculate the attention weights between different parts of the image features, effectively capturing the relationships between various regions within the object;

[0020] Step 3.3: For each 3D object, use its corresponding 3D geometric information to construct a text description through the designed fixed template, input it into the pre-trained language model, and encode it into a text embedding vector f t ;

[0021] Step 3.4: Under the guidance of image features, relevant semantic information is extracted from 3D text features and integrated with visual features to form a joint representation of object appearance and geometric attributes to obtain complete object information f a .

[0022] Preferably, the step 3.4 includes the following contents:

[0023] Adopt a dual-head attention mechanism to take the image feature f' v As the query Q, the As key K and value V, cross attention calculation is performed with the following formula:

[0024]

[0025] Where Q∈R L×D And K,V∈R 1×D ;

[0026] Use the above Q and K to calculate the query-key attention graph A tt , and aggregate the weight information V to obtain a visual and 3D text-aware query Q', as follows:

[0027]

[0028] Where D is the length of the feature vector, that is, the dimension size of each query, key or value vector, Q'∈R L×D .

[0029] Preferably, the step 4 comprises the following steps:

[0030] Step 4.1: Apply a bidirectional attention mechanism to achieve a preliminary fusion between language and object features, so that language and object features complement and enhance each other; language features guide the model to focus on the visual aspects of the object related to the description, while the visual features of the object enrich the semantic information of the language description. The language features P t and object features f a Used alternately as query, key, and value, the process is as follows:

[0031] O2T=MHCA(p t ,f a ,f a )T2O=MHCA(f a ,p t ,p t ) (3)

[0032] Among them, O2T∈R C×D And T2O∈R L×D ;

[0033] Step 4.2: The fused object and language features O2T and T2O are first concatenated to obtain x input Input into the module for adaptive fusion to effectively capture the interaction between cross-modal features. Input feature x input First, mix the channels through convolution, and then use the function for activation to get x mixed ; Then, the processed feature x mixed is input into the forward and backward modules, and S forward and S backward ,These modules work in parallel to capture contextual information from different directions in the feature sequence.

[0034] Preferably, the step 4.2 includes the following contents:

[0035] x input =Concat(MLP(T2O),MLP(O2T)) (4)

[0036] x mixed =SiLU(Conv 1x1 (x input )) (5)

[0037] S forward =SSM(x mixed ) (6)

[0038] S backward =Flip(SSM(Flip(x mixed ))) (7).

[0040] Preferably, the step five includes the following contents:

[0041] The binary cross entropy loss BCELoss is used to supervise the classification and matching performance; BCELoss measures the difference between the predicted probability of the target class and its true label, then calculates the cosine similarity between each object feature and the language feature, and inputs the similarity score into the contrastive loss function to improve the model's ability to distinguish between matched and unmatched objects.

[0042] Compared with the prior art, the present invention has achieved the following technical effects:

[0043] (1) The present invention improves the accuracy and efficiency of retrieval and recognition by integrating a large-scale intelligent traffic event recognition method based on a selective state space model with a cross-modal retrieval technology; compared with the existing cross-modal retrieval method, the present invention not only achieves higher recognition accuracy, but also significantly reduces the computational complexity;

[0044] (2) The present invention adopts a selective filtering module and a selective alignment module based on a state-space model to optimize and update the representation learned by the model, solving the problems of invalid semantic alignment and high computational complexity caused by global allocation in the traditional cross-modal attention mechanism;

[0045] (3) By performing implicit and selective fine-grained matching between images and texts, the model can effectively filter out irrelevant information and enhance the perception and alignment of key features, thereby improving the accuracy and speed of cross-modal retrieval and significantly improving the ability to accurately identify and locate traffic events. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present invention are described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0048] The present invention discloses a 3D multi-target detection method based on semantic-driven single image for ATS, comprising the following steps:

[0049] Step 1: The input RGB image is processed using an advanced 3D detector to extract the 3D bounding boxes of each object in the image; the detector generates all potential 2D projections of these 3D objects in the scene;

[0050] Step 2: Convert the natural language description into feature information that can be processed by the machine; by using natural language processing technology, extract the keywords, phrases and their semantic information in the description to form feature information P representing the language description t ;

[0051] Step 3: Fuse the object’s 2D image information and 3D geometric information to obtain a complete object representation f a ;

[0052] The specific implementation method is as follows:

[0053] Step 3.1: For each 3D object, use its corresponding 2D bounding box to crop the object area from the original RGB image and obtain the corresponding 2D image information;

[0054] Step 3.2: Using 2D image information, extract the visual features of 3D objects through a pre-trained network v , the size is 768 × the number of patches in the image, and the visual features f are processed by a multi-head self-attention mechanism. v Get fine features f' v , calculate the attention weights between different parts of the image features, effectively capturing the relationships between various regions within the object;

[0055] Enhance the expressiveness of features by focusing on key areas while suppressing irrelevant information, thereby improving the overall feature representation;

[0056] Step 3.3: For each 3D object, use its corresponding 3D geometric information to construct a text description through a designed fixed template. The text description is input into the pre-trained language model, which encodes it into a text embedding vector f t ;

[0057] Step 3.4: Under the guidance of image features, relevant semantic information is selectively extracted from 3D text features and fused with visual features to form a joint representation of object appearance and geometric attributes to obtain complete object information f a ;

[0058] The specific implementation method is as follows:

[0059] Adopt a dual-head attention mechanism to take the image feature f' v is regarded as the query (Q), and the As the key (K) and value (V) for cross attention calculation, the formula is as follows:

[0060]

[0061] Where Q∈R L×DAnd K,V∈R 1×D .

[0062] Use the above Q and K to calculate the query-key attention graph A tt , and aggregate the weight information V to obtain a visual and 3D text-aware query Q', as follows:

[0063]

[0064] Where D is the length of the feature vector, that is, the dimension size of each query, key or value vector, Q'∈R L×D ;

[0065] Step 4: Associate the language features extracted from the language description with the detected 3D objects to capture the semantic correspondence between the text and visual modalities;

[0066] The specific implementation method is as follows:

[0067] Step 4.1: For a given description, selectively focus on the relevant parts and apply a bidirectional attention mechanism to achieve a preliminary fusion between language and object features, so that language and object features complement and enhance each other; language features guide the model to focus on the visual aspects of the object related to the description, while the visual features of the object enrich the semantic information of the language description. The language feature P t and object features f a Used interchangeably as query, key, and value,

[0068] The process is as follows:

[0069] O2T=MHCA(p t ,f a ,f a )T2O=MHCA(f a ,p t ,p t ) (3)

[0070] Among them, O2T∈R C×D And T2O∈R L×D ;

[0071] Step 4.2: The fused object and language features (O2T and T2O) are first concatenated to obtain x input Input into the module for adaptive fusion to effectively capture the interaction between cross-modal features. Input feature x input First, the channels are mixed by convolution, and then the activation function is used to get x mixed ; Then, the processed feature x mixed is input into the forward and backward modules, and S forward and S backward,These modules work in parallel to capture contextual information from different directions in the feature sequence,

[0072] The process is as follows:

[0073] x input =Concat(MLP(T2O),MLP(O2T)) (4)

[0074] x mixed =SiLU(Conv 1x1 (x input )) (5)

[0075] S forward =SSM(x mixed ) (6)

[0076] S backward =Flip(SSM(Flip(x mixed ))) (7)

[0077] Step 5: Filter the targets according to the generated matching scores to obtain all targets that match the natural language description;

[0078] Binary cross entropy loss (BCELoss) is used to supervise classification and matching performance;

[0079] BCELoss measures the difference between the predicted probability of the target class and its true label, calculates the cosine similarity between each object feature and the language feature, and inputs the similarity score into the contrastive loss function to improve the model's ability to distinguish between matched and unmatched objects.

[0080] Example 1

[0081] Experimental conditions:

[0082] The experiments were conducted using the PyTorch deep learning framework on a platform equipped with an NVIDIA GeForce RTX 3090 GPU.

[0083] The dataset is divided into training set, validation set and test set in a ratio of 3:1:1.

[0084] The methods compared in the experiment are as follows:

[0085] One is a feature fusion module based on the attention mechanism, denoted as AFM in the experiment, which is used to perform feature fusion between 3D point cloud and visual features. Based on the self-attention mechanism, the model can automatically learn to weightedly fuse 3D point cloud data and 2D visual feature data, that is, the attention score will be calculated between the point cloud features and the image features, so as to determine which parts of the point cloud and image information are most important for target positioning and ignore irrelevant information.

[0086] One is a model based on the Transformer architecture, denoted as AFM3DVG-Transformer in the experiment. It is used to realize visual-guided target positioning on 3D point clouds. It solves the key problem of the fusion of 3D point clouds and natural language descriptions through self-attention and cross-modal relationship modeling, and establishes multi-scale and multi-relationship associations between point clouds and language. The model finally generates a segmentation result or positioning result of a target point cloud area.

[0087] One is a 3D object detection method based on natural language description, denoted as ScanRefer in the experiment. It uses natural language description to accurately locate the target object in the 3D point cloud, extracts point cloud features and language features respectively, and uses a cross-modal attention mechanism to effectively fuse point cloud and language features. The attention mechanism is used to generate fused multimodal features. The point cloud features of each object are enhanced, which can better represent the correlation with the language description, and achieve accurate positioning of the target object in complex 3D scenes.

[0088] One is a 3D visual detection model based on multi-view learning (Multi-View Learning) and Transformer architecture, denoted as Multi-View Transformer in the experiment, which solves the target positioning task in 3D point cloud scenes, captures the global and local features of the scene through multi-view projection, and uses Transformer to achieve intra-modal and inter-modal relationship modeling, dynamically learns the complex semantic relationship between language description and point cloud features, and multi-layer Transformer comprehensively captures the spatial relationship in the point cloud and the contextual information between objects.

[0089] The last one is a 3D visual detection method based on multimodal feature fusion, denoted as Multi3DRefer in the experiment. It accurately locates the natural language description to multiple target objects in the 3D scene, and performs point cloud feature extraction and language feature extraction respectively. It uses the multimodal feature fusion module to align the language description with the point cloud features through the cross-modal attention mechanism, and dynamically captures the association between the language description and multiple targets in the scene. It also designs a scene relationship modeling module, which combines the spatial position and semantic relationship between objects to construct a language-guided global relationship matrix, accurately models the contextual information between objects, and realizes target positioning and classification.

[0090] It can be seen that the present invention creates two datasets, MM3DRefer and MT3DRefer, which provide a large amount of natural language descriptions, corresponding 3D object annotations and multi-object relationship information, and solve the shortcomings of existing datasets in multi-object 3D vision basic tasks.

[0091] Experimental content:

[0092] According to the specific implementation of the present invention, the detection evaluation index of the test data sets in the data sets MM3DRefer and MT3DRefer are calculated, and compared with the indexes of the AFM method, AFM3DVG-Transformer method, ScanRefer method, Multi-ViewTransformer method and Multi3DRefer method. The results are shown in Tables 1 and 2, where ↑ means the higher the better, and ↓ means the lower the better.

[0093] Table 1 Detection evaluation indicators of MM3DRefer dataset

[0094]

[0095] Table 2 Detection evaluation indicators of MT3DRefer dataset

[0096]

[0097] As can be seen from Table 1 and Table 2, due to the selective fusion module and the selective interaction module based on the selective fusion and interaction mechanism adopted by the present invention, the synergy of these two modules enables the cross-modal matching architecture to more accurately capture the semantic correspondence between text descriptions and object visual features, and maintain a high basic performance even if there are object detection errors. Therefore, a very significant retrieval effect is achieved, verifying the advanced nature of the present invention.

[0098] The above description is only a preferred embodiment of the present invention and does not limit the technical scope of the present invention. Therefore, any slight modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A semantically driven 3D multi-target detection method for ATS in a single image, characterized in that: The following steps are involved: Step 1: Process the input RGB image to extract the 3D bounding boxes of each object in the image; generate all potential 2D projections of the 3D object in the scene; Step 2: By using natural language processing technology, extract the keywords, phrases and their semantic information in the description to form feature information P representing the language description t ; Step 3: Fuse the object’s 2D image information and 3D geometric information to obtain a complete object representation f a ; Step 4: Associate the language features extracted from the language description with the detected 3D objects to capture the semantic correspondence between the text and visual modalities; Step 5: Filter the targets according to the generated matching scores to obtain all targets that match the natural language description.

2. According to claim 1, a semantically driven 3D multi-target detection method for ATS in a single image, characterized in that: The step three comprises the following steps: Step 3.1: For each 3D object, use its corresponding 2D bounding box to crop the object area from the original RGB image and obtain the corresponding 2D image information; Step 3.2: Using 2D image information, extract the visual features of 3D objects through a pre-trained network v , the size is 768 × the number of patches in the image, and the visual features f are processed by a multi-head self-attention mechanism. v Get fine features f' v , calculate the attention weights between different parts of the image features, effectively capturing the relationships between various regions within the object; Step 3.3: For each 3D object, use its corresponding 3D geometric information to construct a text description through the designed fixed template, input it into the pre-trained language model, and encode it into a text embedding vector f t ; Step 3.4: Under the guidance of image features, relevant semantic information is extracted from 3D text features and integrated with visual features to form a joint representation of object appearance and geometric attributes to obtain complete object information f a .

3. The method for 3D multi-target detection in a single image based on semantics-driven ATS according to claim 2, characterized in that: The step 3.4 includes the following contents: Adopt a dual-head attention mechanism to take the image feature f′ v As the query Q, the As key K and value V, cross attention calculation is performed with the following formula: Where Q∈R L×D And K, V∈R 1×D ; Use the above Q and K to calculate the query-key attention graph A tt , and aggregate the weight information V to obtain a visual and 3D text-aware query Q′, as follows: Where D is the length of the feature vector, that is, the dimension size of each query, key or value vector, Q′∈R L×D .

4. The method for 3D multi-target detection in a single image based on semantics-driven ATS according to claim 1, characterized in that: The step 4 comprises the following steps: Step 4.1: Apply a bidirectional attention mechanism to achieve a preliminary fusion between language and object features, so that language and object features complement and enhance each other; language features guide the model to focus on the visual aspects of the object related to the description, while the visual features of the object enrich the semantic information of the language description. The language features P t and object features f a Used alternately as query, key, and value, the process is as follows: O2T=MHCA(p t ,f a ,f a )T2O=MHCA(f a ,p t ,p t ) (3) Among them, O2T∈R C×D And T2O∈R L×D ; Step 4.2: The fused object and language features O2T and T2O are first concatenated to obtain x input Input into the module for adaptive fusion to effectively capture the interaction between cross-modal features. Input feature x input First, mix the channels through convolution, and then use the function for activation to get x mixed ; Then, the processed feature x mixed are input into the forward and backward modules, and S forward and S backward ,These modules work in parallel to capture contextual information from different directions in the feature sequence.

5. The method for 3D multi-target detection in a single image based on semantics-driven ATS according to claim 4, characterized in that: The step 4.2 includes the following contents: x input =Concat(MLP(T2O),MLP(O2T)) (4) x mixed =SiLU(Conv 1x1 (x input )) (5) S forward =SSM(x mixed ) (6) S backward =Flip(SSM(Flip(x mixed ))) (7)。 6. The method for 3D multi-target detection in a single image based on semantics-driven ATS according to claim 1, characterized in that: The step five includes the following contents: The binary cross entropy loss BCELoss is used to supervise the classification and matching performance; BCELoss measures the difference between the predicted probability of the target class and its true label, then calculates the cosine similarity between each object feature and the language feature, and inputs the similarity score into the contrastive loss function to improve the model's ability to distinguish between matched and unmatched objects.

Citation Information

Patent Citations

  • Image description method based on adaptive enhanced self-attention network

    CN114677580A

  • Text anaphora video object segmentation method based on reference analysis and perception enhancement

    CN117079177A

  • Named entity recognition method based on comparative learning and multi-modal semantic interaction

    CN117574904A

  • Information guide target searching method based on cross-modal self-evolution knowledge generalization

    CN118170938A

  • Semantic information fused loopback detection method and system

    CN118570502A

Cited By

  • 3D shape segmentation method and system based on text driving and attention mechanism

    CN120747975A