Automatic driving space planning enhancement method based on visual marker

Through the visual marking method, combined with visual coding and large language model, the semantic separation problem between visual and language modalities in autonomous driving is solved, and the spatial understanding and decision-making reliability is achieved with higher accuracy, which is suitable for autonomous driving systems.

CN120298992AActive Publication Date: 2025-07-11SOUTH CHINA UNIV OF TECH

Patent Information

Application Number
CN202510359692.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing text-based coordinate method increases the cross-modal alignment complexity and expression error of the model, and the multimodal large language model cannot capture fine-grained spatio-temporal correlations when dealing with dynamic interactions in dense traffic flows, resulting in insufficient perception and decision-making reliability of the autonomous driving system in complex scenarios.

Method used

The visual marking method is adopted to generate visual marking images by detecting expert models, combine visual encoder and large language models to generate text output with visual markings, and realize precise replacement of object coordinates through coordinate mapping tables, improving the accuracy and semantic consistency of spatial understanding.

Benefits of technology

It significantly improves the analytical accuracy of object positions, motion states and interaction relationships in autonomous driving scenarios, enhances the reliability of decisions and the naturalness of planning, and improves the synchronization of perception and decision-making of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298992A_ABST
    Figure CN120298992A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving space planning enhancement method based on a visual marker. The method comprises the following steps: acquiring an original image and text input; processing the original image to obtain image features; processing the text input to obtain text features; generating a text output with a visual mark by using the image features and the text features; the text output with the visual marks is converted, and text output with coordinates is obtained; the accuracy and semantic consistency of spatial understanding in an automatic driving scene are remarkably improved, high synchronization of visual perception and semantic expression is achieved, and the problem of visual and language modal semantic segmentation in an existing method is effectively solved. The analysis precision of the object position, the motion state and the interaction relation in the automatic driving question and answer task is greatly improved, and the decision reliability and the planning naturalness in a complex driving scene can be remarkably enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of autonomous driving technology and artificial intelligence technology, and particularly to an enhanced method for autonomous driving space planning based on visual markers. Background Art

[0002] With the rapid development of autonomous driving technology, accurately perceiving the spatial relationships in complex traffic scenarios (such as object positioning, motion states, and interaction intentions) has become the core challenge for achieving safe decision-making and planning. As a key innovation in the fields of intelligent transportation systems, driverless vehicles, etc., it is necessary to real-time analyze the environmental semantics through multi-modal data and construct an interpretable driving logic. However, existing autonomous driving visual question answering (AD-VQA) methods based on multi-modal large language models (MLLMs) have significant bottlenecks in spatial information expression. Directly encoding coordinate information in text form will lead to semantic disconnection between visual representations and language descriptions, which not only increases the expression complexity of the model but also easily causes cross-modal alignment errors, ultimately affecting the reliability of perception, prediction, and planning tasks.

[0003] There are specific technical difficulties in the spatial understanding of autonomous driving technology. Generating spatial semantic descriptions directly from multi-modal inputs (visual images and text instructions) has ambiguity in cross-modal alignment. Due to the one-to-many dynamic mapping relationship between language description logic and visual coordinate features (such as "the vehicle on the left" may correspond to multiple candidate coordinate regions), this process is prone to over-generalization problems in scene parsing, manifested as fuzzy object positioning, misjudgment of motion states, and inaccurate description of spatial relationships.

[0004] Existing text coordinate-based methods can enhance the target positioning accuracy, but due to relying on the direct embedding of numerical coordinates, they lead to semantic disconnection between visual representations and language modalities, significantly increasing the cross-modal alignment complexity and expression errors of the model. At the same time, although methods based on multi-modal large language models have made breakthroughs in autonomous driving visual question answering tasks, they often fail to capture fine-grained spatio-temporal correlations when dealing with dynamic interactions in dense traffic flows (such as predicting vehicle cut-in intentions or planning pedestrian obstacle avoidance). In particular, when spatial understanding needs to be strictly synchronized with real-time decision-making logic (such as judging the priority of lane changes, multi-object risk grading), the outputs of existing methods often show a tendency of semantic disconnection or over-conservatism, resulting in damage to the spatio-temporal continuity of the planning link.

[0005] Therefore, to bridge the semantic gap between visual coordinate representations and language descriptions and improve the interpretability and robustness of spatial reasoning, it is of urgent significance to research and develop an enhanced framework for autonomous driving that integrates dual-grained visual cues. This can not only strengthen the environmental perception and decision-making reliability in complex scenarios but also promote the paradigm upgrade of autonomous driving systems from "perception-driven" to "cognition-driven", providing core technical support for human-vehicle collaboration and high-level intelligent driving. Summary of the Invention

[0006] (1) Technical Problems to be Solved

[0007] The present invention discloses a method for enhancing autonomous driving space planning based on visual markers, aiming to solve the problems that the existing text coordinate-based method relies on the direct embedding of numerical coordinates, increasing the cross-modal alignment complexity and expression error of the model; at the same time, the method based on the multimodal large language model fails to capture fine-grained spatio-temporal associations when dealing with dynamic interactions in dense traffic flows.

[0008] (2) Technical Solutions

[0009] The present invention discloses a method for enhancing autonomous driving space planning based on visual markers, including the following steps:

[0010] Step 1, obtain the original image and text input;

[0011] Step 2, process the original image to obtain image features;

[0012] Step 3, process the text input to obtain text features;

[0013] Step 4, generate a text output with visual markers using the image features and text features;

[0014] Step 5, transform the text output with visual markers to obtain a text output with coordinates.

[0015] Further, the process of processing the original image to obtain image features includes the following steps:

[0016] Step 201, use the original image to generate a visual marker image through a detection expert model;

[0017] Step 202, generate scene-level features using the original image and the visual marker image;

[0018] Step 203, generate instance-level features using the scene-level features;

[0019] Step 204, generate image features using the scene-level features and the instance-level features respectively.

[0020] Further, the specific steps for processing the text input to obtain text features are:

[0021] Extract text features from the text input in Step 1 through a token encoder.

[0022] Further, the generation of a text output with visual markers using the image features and text features includes the following steps:

[0023] Step 401: Input the text features and image features into the large language model;

[0024] Step 402: The large language model generates a text output with visual markers.

[0025] Further, the conversion of the text output with visual markers to obtain a text output with coordinates includes the following steps:

[0026] Step 501: Extract the object numbers in the text output with visual markers;

[0027] Step 502: Construct a coordinate mapping table, query the coordinate mapping table according to the object number, and obtain the object coordinates;

[0028] Step 503: Replace the object numbers in the text output with visual markers with the object coordinates;

[0029] Step 504: Output the text output with coordinates.

[0030] Further, the generation of the visual marker image using the original image through the detection expert model includes the following steps:

[0031] Step 2011: Use the detection expert model StreamPETR to identify traffic objects in the original image I according to the specified object categories;

[0032] Step 2012: The detection expert model generates detection masks for k traffic objects according to the identified k traffic objects, expressed as a binary mask set R = [r1, r2,..., r k , where r k ∈{0,1} H×W represents the k-th detection mask, H is the height of the mask, and W is the width of the mask;

[0033] Step 2013: Calculate the average centroid coordinate c k for each r k =(x k , y k ), x k represents the abscissa, and y k represents the ordinate;

[0034] Step 2014: In the original image, mark the corresponding index k at the centroid coordinate c k of each detected object;

[0035] Step 2015: Overlay a semi-transparent detection mask r k of the corresponding size on the original image I to describe the object boundary;

[0036] In step 2016, k detection masks and their corresponding indices are superimposed on the original image I to generate a visual marker image I m .

[0037] Furthermore, generating scene-level features using the original image and the visual marker image includes the following steps:

[0038] In step 2021, the parameters θ of the original visual encoder E are frozen, and a trainable copy E with parameters θ c is created c ;

[0039] In step 2022, the original image I passes through the original visual encoder E to obtain the encoding of the original image, and the visual marker image I m passes through the trainable copy E c to obtain the encoding of the visual marker image;

[0040] In step 2023, the encoding of the visual marker image passes through a zero-linear network Z with parameters θ Z to output the scene-level features of the visual marker image; among them, the weights and biases of the zero-linear network Z are both initialized to 0;

[0041] In step 2024, the scene-level features of the visual marker image and the elements of the encoding of the original image are accumulated to obtain the scene-level features y of the original image s , and the expression is:

[0042] y s = E(I; θ) + Z(E c (I m ; θ c )); Z );

[0043] where y s ∈R H×W×C , H is the height, W is the width, and C is the number of channels.

[0044] Furthermore, generating instance-level features using the scene-level features includes the following steps:

[0045] First, through the scene-level features y of the original image s and the detection mask r of the k-th object k , the detection mask r k is scaled to the same size as the scene-level features y s ;

[0046] Then, the mask average pooling MAP is used to obtain the instance-level features of the k-th object

[0047]

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0049] By using a detection expert model to annotate the original image, generating two types of fine-grained image features through a vision encoder and an MLP, and combining the text features generated by a token encoder from the text input, an action plan combining the image and text inputs is generated via a large language model, significantly improving the accuracy and semantic consistency of spatial understanding in the autonomous driving scenario, achieving a high degree of synchronization between visual perception and semantic expression, effectively solving the problem of semantic fragmentation between the visual and language modalities in the existing methods. It not only greatly improves the parsing accuracy of the object position, motion state and interaction relationship in the autonomous driving question and answer task, but also significantly enhances the decision-making reliability and planning naturalness in complex driving scenarios, and is more suitable for application fields such as autonomous driving systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is the overall flowchart of the present invention;

[0051] Figure 2 is the flowchart of processing the original image and obtaining image features in step 2 of the present invention.

[0052] Figure 3 is the flowchart of converting the text output with visual markers and obtaining the text output with coordinates in step 5 of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0053] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] For the sake of citation and clarity, the technical terms, abbreviations or acronyms used hereinafter are summarized and explained as follows:

[0055] Vision encoder: A model component that extracts key visual features from an image and performs efficient encoding.

[0056] MLP: A feedforward neural network composed of fully connected layers, which processes complex tasks through non-linear activation functions.

[0057] Large language model: A deep learning system trained based on a large amount of text data and capable of understanding and generating natural language.

[0058] Text Token Encoder: A component that splits text into basic units (tokens) and converts them into numerical vector representations for processing and analysis by machine learning models.

[0059] Reference Figure 1 , the present invention proposes an enhanced method for autonomous driving space planning based on visual markers, specifically including the following steps:

[0060] Step 1, obtain the original image and text input;

[0061] Step 2, process the original image to obtain image features;

[0062] Step 3, process the text input to obtain text features;

[0063] Step 4, use the image features and text features to generate a text output with visual markers;

[0064] Step 5, convert the text output with visual markers to obtain a text output with coordinates.

[0065] The following specifically illustrates the process of an enhanced method for autonomous driving space planning based on visual markers with examples:

[0066] Step 1, obtain the original image and text input, specifically including the following steps:

[0067] Step 101, the original image can be collected by a camera or other collection devices and processed by upstream algorithms to obtain the image input;

[0068] Step 102, the text input can be collected by the user through collection devices and processed by upstream algorithms to obtain the text input, or the system can directly input the text.

[0069] Step 2, process the original image to obtain image features, reference Figure 2 ;

[0070] Specifically, it includes the following algorithm processes:

[0071] Step 201, use the original image to generate a visual marker image through a detection expert model;

[0072] Step 202, use the original image and the visual marker image to generate scene-level features;

[0073] Step 203, use the scene-level features to generate instance-level features;

[0074] Step 204, use the scene-level features and instance-level features to generate image features respectively.

[0075] Specifically, in step 201, the process of generating a visual marker image from the original image using the detection expert model is as follows:

[0076] In step 2011, use the detection expert model StreamPETR to identify traffic objects (including but not limited to cars, trucks, and buses) in the original image I according to the specified object categories;

[0077] In step 2012, the detection expert model generates detection masks for the k traffic objects, represented as a binary mask set R = [r1, r2,..., r k , where r k ∈{0,1} H×W represents the k-th detection mask, H is the height of the mask, and W is the width of the mask.

[0078] In step 2013, calculate the average centroid coordinate c k =(x k ,y k ) for each r k , where x k represents the abscissa and y k represents the ordinate;

[0079] In step 2014, in the original image, mark the corresponding index k at the centroid coordinate c k of each detected object;

[0080] In step 2015, cover the original image I with a semi-transparent detection mask r k of the corresponding size to describe the object boundary;

[0081] In step 2016, the k detection masks and their corresponding indices will be superimposed on the original image I to generate a visual marker image I m .

[0082] In addition, if there is a new coordinate reference c new in the text input, an index k + 1 will be assigned to it, and a semi-transparent detection mask r k+1 of the corresponding size will be covered on the original image I to maintain the consistency of the visual and text modes.

[0083] Specifically, in step 202, the process of generating scene-level features from the original image and the visual marker image is as follows:

[0084] In step 2021, freeze the parameters θ of the original visual encoder E, and at the same time create a trainable copy E c with parameters θ c ;

[0085] In step 2022, the original image I passes through the original vision encoder E to obtain the encoding of the original image, and the visual token image I m obtains the encoding of the visual token image through the trainable copy E c ;

[0086] In step 2023, the encoding of the visual token image passes through a zero-linear network Z with parameters θ Z to output the scene-level features of the visual token image; among them, the weights and biases of the zero-linear network Z are both initialized to 0;

[0087] In step 2024, the scene-level features of the visual token image and the elements of the encoding of the original image are accumulated to obtain the scene-level features y of the original image s , and the expression is:

[0088] y s = E(I; θ) + Z(E c (I m ; θ c ); θ Z );

[0089] where y s ∈ R H×W×C , H is the height, W is the width, and C is the number of channels.

[0090] Specifically, the vision encoder used in step 202 is the InternViT-300M-448px model, the core of which is an improved structure based on VisionTransformer, supporting dynamic resolution input and pixel shuffling technology. In the initial stage of signal processing, the input image is cut into several 448×448 pixel tiles through a dynamic chunking strategy, and each tile is downsampled to 1 / 4 size through pixel shuffling. The 256 visual tokens corresponding to a single tile are input into the vision encoder. Through a multi-layer self-attention mechanism, the model extracts global scene-level features, combines the object masks generated by the detection algorithm, scales the masks to the same size as the scene features using bilinear interpolation, and then extracts local instance-level features through masked average pooling.

[0091] Specifically, in step 203, instance-level features are generated using the scene-level features, and the specific steps are as follows:

[0092] First, through the scene-level features y of the original image s and the detection mask r of the k-th object k , the detection mask r k is scaled to the same size as the scene-level features y s ;

[0093] Next, the masked average pooling MAP is used to obtain the instance-level features of the k-th object

[0094]

[0095] Specifically, the steps of generating image features using scene-level features and instance-level features in step 204 are as follows:

[0096] First, the scene-level features and instance-level features are concatenated, and then mapped to a 4096-dimensional hidden space aligned with the text features through MLP to obtain the image features.

[0097] In step 3, the text input is processed to obtain text features. The specific steps are as follows: The text features are extracted from the text input in step 1 through a token encoder.

[0098] In step 4, the image features and text features are used to generate a text output with visual markers. The specific steps are as follows:

[0099] Step 401, input the text features and image features into a large language model;

[0100] Step 402, the large language model generates a text output with visual markers.

[0101] Specifically, the large language model used in step 4 is the InternLM2.5-7B-Chat model, which includes the following structure:

[0102] First, InternLM2.5-7B-Chat is based on a multi-layer Transformer architecture and adopts a grouped query attention (GQA) mechanism, which groups query heads to share key / value heads, reducing the video memory occupancy while maintaining the 8K context window processing ability;

[0103] Secondly, for dynamic gating cross-modal interaction, dynamic gating cross-attention modules are inserted into the 16th, 24th, and 32nd layers of the Transformer. The visual features serve as key-value pairs (K / V), and the text features serve as queries (Q). The contribution degree of visual information is adjusted through learnable gating weights (in the range of 0-1) to prevent irrelevant image regions from interfering with text generation.

[0104] In step 5, the text output with visual markers is transformed to obtain a text output with coordinates, including the following steps, as shown Figure 3 :

[0105] Step 501, extract the object numbers in the text output with visual markers;

[0106] Step 502, construct a coordinate mapping table, query the coordinate mapping table according to the object numbers to obtain the object coordinates;

[0107] Step 503, replace the object numbers in the text output with visual markers with object coordinates;

[0108] Step 504, output the text output with coordinates.

[0109] Specifically, the process of extracting the object numbers in the text output with visual markers in Step 501:

[0110] The pre-trained detection expert model (such as StreamPETR) performs multi-object detection on the driving scene image, identifies traffic objects (vehicles, pedestrians, etc.) and generates unique digital numbers. The detection expert model first outputs a binary mask set R = [r1, r 2, ..., r k , and then assigns consecutive integer numbers (such as ID1, ID 2) to each detected object, and embeds the numbers into the text output with visual markers in text form.

[0111] Specifically, the process of querying the coordinate table according to the object number to obtain the object coordinates in Step 502:

[0112] Build a coordinate mapping table synchronously during the detection phase, where each object number corresponds to its geometric center coordinate c k =(x k , y k ). The coordinate table is stored in the form of key-value pairs (such as ID 1→(864.2, 468.3)). When the model generates text with numbers, parse the object numbers in the text through a string matching algorithm, and quickly obtain the corresponding coordinate values based on the hash table retrieval mechanism to achieve an accurate mapping from semantic labels to spatial positions.

[0113] Specifically, the process of replacing the object numbers in the text output with visual markers with object coordinates in Step 503:

[0114] Use pattern recognition technology based on regular expressions to locate all markers in the text in the form of "<ID k>", where k is an integer number. By traversing the replacement queue, dynamically replace each numbered marker with the corresponding coordinate string "<x k , y k >", while retaining the original text syntax structure. The coordinate format verification module is introduced during the replacement process to ensure that the numerical accuracy and coordinate range meet the image space resolution constraints.

[0115] Specifically, the process of outputting the text with coordinates in Step 504:

[0116] The sequence of the replaced text is input into the language model for grammar correction and format standardization, and finally a spatial description text that conforms to the natural language specification is generated. The output text strictly follows the triple structure of "<object identifier (average centroid), coordinates, camera view>" (such as "<c1,1020.0,515.8,CAM_FRONT>"), and at the same time, the consistency between the spatial description and the context semantics is ensured through the attention mechanism enhancement module.

[0117] Those of ordinary skill in the art will recognize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0118] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0119] The technical features of the above embodiments can be combined arbitrarily. At the same time, the labels of each step are not used to restrict the sequence between steps. As long as there is no strict sequence constraint between steps, their order is allowed to be adjusted and transformed. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0120] The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An enhanced method for autonomous driving space planning based on visual markers, characterized in that: Including the following steps: Step 1, obtain the original image and text input; Step 2, process the original image to obtain image features; Step 3, process the text input to obtain text features; Step 4, generate a text output with visual markers using the image features and text features; Step 5, transform the text output with visual markers to obtain a text output with coordinates.

2. The enhanced method for autonomous driving space planning based on visual markers according to claim 1, wherein: The processing of the original image to obtain image features includes the following steps: Step 201, use the original image to generate a visual marker image through a detection expert model; Step 202, generate scene-level features using the original image and the visual marker image; Step 203, generate instance-level features using the scene-level features; Step 204, generate image features using the scene-level features and the instance-level features respectively.

3. A method for enhancing autonomous driving space planning based on visual markers according to claim 1, characterized in that: The specific steps for processing the text input to obtain text features are: Extract text features from the text input in Step 1 through a token encoder.

4. A method for enhancing autonomous driving space planning based on visual markers according to claim 1, characterized in that: The generation of a text output with visual markers using the image features and text features includes the following steps: Step 401, input the text features and image features into a large language model; Step 402, the large language model generates a text output with visual markers.

5. The enhanced method for autonomous driving space planning based on visual markers according to claim 1, characterized in that: The transformation of the text output with visual markers to obtain a text output with coordinates includes the following steps: Step 501, extract the object numbers in the text output with visual markers; Step 502, construct a coordinate mapping table, query the coordinate mapping table according to the object numbers, and obtain the object coordinates; Step 503, replace the object numbers in the text output with visual markers with the object coordinates; Step 504, output the text output with coordinates.

6. The enhanced method for autonomous driving space planning based on visual markers according to claim 2, wherein: The use of the original image to generate a visual marker image through a detection expert model includes the following steps: Step 2011, use the detection expert model StreamPETR to identify traffic objects in the original image I according to the specified object categories; Step 2012, the detection expert model generates detection masks for the k traffic objects recognized, represented as a binary mask set R = [r1, r2,..., r k , where r k ∈ {0, 1} H×W represents the k-th detection mask, H is the height of the mask, and W is the width of the mask; Step 2013, for each r k calculate its average centroid coordinate c k =(x k , y k ), where x k represents the abscissa and y k represents the ordinate; Step 2014, in the original image, mark the corresponding index k at the centroid coordinate c of each detected object k ; Step 2015, overlay a semi-transparent detection mask r of corresponding size on the original image I k to describe the object boundary; In step 2016, k detection masks and their corresponding indices are superimposed on the original image I to generate a visual marker image I m .

7. The enhanced method for autonomous driving space planning based on visual markers according to claim 6, characterized in that: The generation of scene-level features using the original image and the visual marker image includes the following steps: Step 2021, freeze the parameters θ of the original visual encoder E, and at the same time create a trainable copy E with parameters θ c ; c ; In step 2022, the original image I is encoded by the original vision encoder E to obtain the encoding of the original image, and the visual marker image I m is encoded by the trainable copy E c to obtain the encoding of the visual marker image; Step 2023, the encoding of the visual marker image passes through a zero-linear network Z with parameters θ Z to output the scene-level features of the visual marker image; wherein, the weights and biases of the zero-linear network Z are both initialized to 0; In step 2024, the scene-level features of the visual marker image and each element of the encoding of the original image are accumulated to obtain the scene-level features y of the original image s , and the expression is: y s = E(I; θ) + Z(E c (I m ; θ c ); θ Z ); where y s ∈R H×W×C , H is the height, W is the width, and C is the number of channels.

8. A method for enhancing autonomous driving space planning based on visual markers according to claim 7, characterized in that: The generation of instance-level features using the scene-level features includes the following steps: First, through the scene-level feature y of the original image s and the detection mask r of the k-th object k , scale the detection mask r k to the same size as the scene-level feature y s ; Next, use masked average pooling (MAP) to obtain the instance-level features of the k-th object

Citation Information

Patent Citations

  • Character image recognition method and device based on global semantics and computer equipment

    CN114445832A

  • Model training method and device, equipment, storage medium and program product

    CN118378633A

  • Visual question and answer data enhancement method and device, equipment and storage medium

    CN119128118A

  • Automated transformation of information from images to textual representations, and applications therefor

    US20240362197A1

  • Visual question answering method and apparatus, electronic device and storage medium

    WO2024164616A1

Cited By

  • Automatic driving space planning enhancement method based on semantic region focusing

    CN121119149A

  • An automatic driving space planning enhancement method based on semantic region focusing

    CN121119149B