Robot Fine-grained Target Localization Method and Device Based on Conditional Multimodal Hints
Through multiple cross-coding of images and text and cross-attention calculations, combined with early and late fusion, the problems of low training efficiency and insufficient cross-modal fusion information representation capabilities in robot visual positioning methods are solved, and the efficiency and accuracy of robot refined target positioning are improved.
Patent Information
- Application Number
- CN202510008088.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-01-03
AI Technical Summary
In the prior art, robot visual positioning methods have shortcomings in training efficiency and model robustness, and it is difficult to achieve refined target positioning, especially cross-modal fusion information representation capabilities, which cannot meet the robot's refined target positioning needs.
The method based on conditional multimodal cues is adopted to cross-code images and texts to generate target visual features and language features, and to combine early and late fusion to achieve precise fine-grained target positioning of the robot.
The efficiency and accuracy of robots' refined target positioning are improved, and precise fine-grained target positioning can be achieved according to free form of language expression.
Smart Images

Figure CN119417906B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a robot fine-grained target positioning method and device based on conditional multi-modal prompts. Background Art
[0002] With the rapid development of artificial intelligence and computer vision technology, visual positioning technology has made remarkable progress and is widely used in fields such as autonomous driving, intelligent manufacturing, robotics, and drone navigation.
[0003] Through visual positioning technology, a robot can more naturally understand and respond to a user's operation. This naturalness is not only reflected in the accurate capture of the user's actions by the robot, but also in the in-depth understanding of the user's intentions by the robot. For example, a robot can capture the movement trajectory of a patient through visual positioning technology and provide personalized movement assistance according to the patient's rehabilitation needs.
[0004] In related technologies, a two-stage (including candidate generation and cross-modal matching) visual positioning model is usually adopted to explore more effective cross-modal interactions, or the best-matched candidate is selected in an interpretable reasoning manner to achieve object detection and positioning. However, the two-stage visual positioning model is a serial architecture, and the model training efficiency is limited. Moreover, it is overly dependent on the training effect of the candidate generation stage, resulting in low model robustness. When using a one-stage visual positioning method for object visual positioning, first, two independent encoders are used to extract corresponding language features and visual features respectively, and then the two types of features are cross-modally fused through an aggregation module. The representational ability of the fused features is limited, and only coarse-grained object positioning (such as outputting a positioning box) can be achieved, which is difficult to meet the requirements of the robot for fine-grained object positioning (such as outputting pixel-level coordinates). Summary of the Invention
[0005] The present invention provides a robot fine-grained target positioning method and device based on conditional multi-modal prompts, which are used to solve the defects that when the existing two-stage visual positioning method is used for robot fine-grained target positioning, the training efficiency is limited and it depends on the training effect of the first stage, resulting in low positioning efficiency and low model robustness. When using a one-stage visual positioning method, the cross-modal fusion information representation ability is low, resulting in the inability to meet the requirements of robot fine-grained target positioning, and improving the efficiency and accuracy of robot fine-grained target positioning.
[0006] The present invention provides a robot fine-grained target positioning method based on conditional multi-modal prompts, including:
[0007] Perform multiple cross - encodings on the image and text respectively to obtain target visual features and target language features. Among them, in each cross - encoding, determine the first prompt guidance according to the i - th visual feature obtained with the image as the initial input, and perform language encoding on the first prompt guidance and the i - th visual feature to obtain the (i + 1) - th language feature. Determine the second prompt guidance according to the i - th language feature obtained with the text as the initial input, and perform visual encoding on the second prompt guidance to obtain the (i + 1) - th visual feature. i is a positive integer greater than 0;
[0008] Map the target visual features and the target language features to the same space, and perform cross - attention calculation on the mapped visual features and the mapped language features to obtain new visual features and new language features, so that the robot can adjust its motion posture according to the position decoding result under the condition of performing position decoding on the new visual features and the new language features.
[0009] According to a robot fine - grained target positioning method based on conditional multi - modal prompts provided by the present invention, the target visual features include the visual features output by each cross - encoding, and the target language features include the language features output by the last cross - encoding;
[0010] Before mapping the target visual features and the target language features to the same space, the method further includes:
[0011] Reshape the visual features output by each cross - encoding respectively, and splice the reshaped visual features to obtain new target visual features.
[0012] According to a robot fine - grained target positioning method based on conditional multi - modal prompts provided by the present invention, the cross - attention calculation is represented by the following formula:
[0013] ;
[0014] Among them, is the multi - head cross - attention operation, is the feed - forward network; is the mapped visual feature, is the mapped language feature, is the new visual feature, is the new language feature.
[0015] According to a robot fine - grained target positioning method based on conditional multi - modal prompts provided by the present invention, the visual encoding is implemented through the following steps:
[0016] Perform visual encoding on the image or visual feature through a pre - trained Swin Transformer;
[0017] The language encoding is implemented through the following steps:
[0018] Perform language encoding on the text or language features through a BERT model.
[0019] According to a method for fine-grained target localization of a robot based on conditional multi-modal prompts provided by the present invention, the determination of the first prompt guidance according to the i-th visual feature obtained with an image as the initial input includes:
[0020] Project the i-th visual feature into the language feature space through the linear layer of the conditional multi-modal prompt generator to obtain the first prompt guidance;
[0021] The determination of the second prompt guidance according to the i-th language feature obtained with a text as the initial input includes:
[0022] Project the i-th language feature into the visual feature space through the linear layer to obtain the second prompt guidance.
[0023] According to a method for fine-grained target localization of a robot based on conditional multi-modal prompts provided by the present invention, after obtaining the new visual feature and the new language feature, the method further includes:
[0024] Perform pixel-level position decoding based on the new visual feature and the new language feature by a robot locator to obtain the position decoding result; wherein, the robot locator includes at least three upsampling decoders;
[0025] Convert the position decoding result according to the spatial coordinate system corresponding to the robotic arm of the robot to obtain positioning parameters, and adjust the motion posture of the robotic arm according to the positioning parameters.
[0026] The present invention also provides a device for fine-grained target localization of a robot based on conditional multi-modal prompts, including:
[0027] An encoding module for performing multiple cross encodings on an image and a text respectively to obtain a target visual feature and a target language feature; wherein, in each cross encoding, determine a first prompt guidance according to the i-th visual feature obtained with an image as the initial input, and perform language encoding on the first prompt guidance and the i-th visual feature to obtain the (i + 1)-th language feature; determine a second prompt guidance according to the i-th language feature obtained with a text as the initial input, and perform visual encoding on the second prompt guidance to obtain the (i + 1)-th visual feature; i is a positive integer greater than 0;
[0028] The feature calculation module maps the target visual feature and the target language feature to the same space, and performs cross-attention calculation on the mapped visual feature and the mapped language feature to obtain a new visual feature and a new language feature, so that the robot can adjust its motion posture according to the position decoding result under the condition of performing position decoding on the new visual feature and the new language feature.
[0029] According to a robot fine-grained target positioning device based on conditional multi-modal prompts provided by the present invention, the device further includes:
[0030] The decoding module is configured to, after obtaining the new visual feature and the new language feature, perform pixel-level position decoding based on a robot locator according to the new visual feature and the new language feature to obtain the position decoding result; wherein, the robot locator includes at least three upsampling decoders;
[0031] The motion adjustment module converts the position decoding result according to the space coordinate system corresponding to the robotic arm of the robot to obtain positioning parameters, and adjusts the motion posture of the robotic arm according to the positioning parameters.
[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the robot fine-grained target positioning method according to any one of the above.
[0033] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the robot fine-grained target positioning method according to any one of the above.
[0034] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the robot fine-grained target positioning method according to any one of the above.
[0035] The robot fine-grained target positioning method and device based on conditional multi-modal prompts provided by the present invention perform cross-encoding based on language encoding and visual encoding on images and texts respectively multiple times to obtain target visual features and target language features, and perform cross-attention calculation and decoding on the above two types of features to obtain fine-grained target positioning results, and then complete the adjustment of the robot's motion posture. Combining the advantages of early and late fusion, it can achieve precise fine-grained target positioning of the robot according to free-form language expressions, improving the efficiency and accuracy of robot fine-grained target positioning. Description of the Drawings
[0036] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0037] Figure 1 It is one of the schematic flowcharts of the robot fine target positioning method based on conditional multi-modal prompts provided by the present invention.
[0038] Figure 2 It is another schematic flowchart of the robot fine target positioning method based on conditional multi-modal prompts provided by the present invention.
[0039] Figure 3 It is the third schematic flowchart of the robot fine target positioning method based on conditional multi-modal prompts provided by the present invention.
[0040] Figure 4 It is the schematic structural diagram of the robot fine target positioning device based on conditional multi-modal prompts provided by the present invention.
[0041] Figure 5 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0042] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0043] The following will describe Figures 1 - 4 the robot fine target positioning method and device based on conditional multi-modal prompts of the present invention.
[0044] Figure 1 It is one of the schematic flowcharts of the robot fine target positioning method based on conditional multi-modal prompts provided by the present invention. As Figure 1 shown, the method includes the following steps:
[0045] Step 110: Perform multiple cross-codings on the image and the text respectively to obtain the target visual feature and the target language feature. Among them, in each cross-coding, determine the first prompt guidance according to the i-th visual feature obtained with the image as the initial input, and perform language coding on the first prompt guidance and the i-th visual feature to obtain the (i + 1)-th language feature; determine the second prompt guidance according to the i-th language feature obtained with the text as the initial input, and perform visual coding on the second prompt guidance to obtain the (i + 1)-th visual feature; i is a positive integer greater than 0.
[0046] In this step, the image and the text can be real-time acquired data. For example, capture images of the surrounding environment (such as people, animals, desks, chairs or other objects) through a camera sensor; collect voice data in the surrounding environment through a microphone, and convert the voice data into text data through speech recognition technology.
[0047] Specifically, a robot is equipped with a processor for executing the above-mentioned robot fine-grained target positioning method based on conditional multi-modal prompts. The robot is also equipped with a camera and a camera pan-tilt head composed of a damping-adjustable buffer hinge, and the camera pan-tilt head can adjust the pitch angle of the camera; the robot is also equipped with a microphone for acquiring voice data during human-robot interaction.
[0048] In this step, the image and the text can also be historical data uploaded by the server or data sent by remote personnel.
[0049] In this embodiment, the models for visual coding include but are not limited to Convolutional Neural Network (CNN), Autoencoder, Recurrent Neural Network (RNN), Transformer model, or Vision Transformer (ViT) model, etc.
[0050] In this embodiment, the models for language coding include but are not limited to Neural Network Language Model (NNLM), word embedding model, multi-modal language model, or pre-trained language model.
[0051] In this step, by performing multiple cross-codings on the image and the text respectively, generate language features with visual perception and visual features with language perception, realize comprehensive cross-perception early fusion, allow for a deeper exploration of cross-modal information, and generating language features with visual perception can reduce the influence of noise in the original language expression.
[0052] In this step, the first prompt guidance is the feature vector after the linear transformation of the visual feature. The first prompt guidance may include one or more feature vectors. One visual encoding outputs one visual feature, and one visual feature corresponds to one prompt guidance.
[0053] Figure 2 is the second flowchart of the robot fine-grained target localization method based on conditional multi-modal prompts provided by the present invention. In Figure 2 the shown embodiment, the picture v0 is input into the visual encoder M1, and the visual feature v1 is output. The text is input into the language encoder S1, and the language feature is output; the conditional multi-modal prompt generator converts v1 into the first prompt guidance regarding vision, and then inputs the first prompt guidance regarding vision and the language feature into the language encoder S2, and the language feature is output. And so on. After passing through four language encoders, the language feature with visual perception ( , and ) can be obtained; the conditional multi-modal prompt generator converts into the second prompt guidance regarding language, and then inputs the second prompt guidance regarding language into the visual encoder M2, and the visual feature v2 is output. And so on. After passing through four visual encoders, the visual feature with language perception (v2, v3 and v4) can be obtained.
[0054] In this embodiment, the target visual feature can be one or more of the above visual features with language perception, and the target language feature can also be one or more of the above language features with visual perception.
[0055] For example, the target visual features include v2, v3 and v4, and the target language feature is .
[0056] Step 120: Map the target visual feature and the target language feature to the same space, and perform cross-attention calculation on the mapped visual feature and the mapped language feature to obtain new visual features and new language features, so that the robot can adjust its motion posture according to the position decoding result under the condition of performing position decoding on the new visual features and the new language features.
[0057] In this step, the features of the two modalities of the target visual feature and the target language feature can be mapped to a shared semantic space, so that features of different modalities can be compared and interacted in the same space.
[0058] For example, the visual modality and the language modality can be aligned based on the Transformer model through the attention mechanism; a shared embedding layer can also be designed to convert visual features and language features into embedding vectors of the same dimension to ensure that the features of the two modalities are represented in the same space.
[0059] In Figure 2 the illustrated embodiment, the target visual features V and the target language features L are mapped to the same space to obtain the mapped visual features and the mapped language features , and then the multi-head cross-attention mechanism is used to perform cross-attention calculations on and respectively to obtain the new visual features and the new language features ; it should be noted that the multi-head cross-attention mechanism allows the model to simultaneously focus on different parts of the input information by using multiple attention heads in parallel. In the multi-head cross-attention mechanism, Q (Query), K (Key), and V (Value) are three important vectors used to calculate the attention weights and generate the output.
[0060] In this embodiment, after extracting features from the encoder, cross-modal interaction is further used to fuse cross-modal features; specifically, only the visual features are used, reshaped and concatenated to create the final visual input , where ; in addition, the final output of the language encoder is used as the language input , and these inputs are projected into a common space to obtain and ; where is the real number field, is the dimension number of the visual features, is the number of feature channels, is the spatial dimension; is for the data length of , and the number of channels of the language features is corresponding to the real number field; is for the data length of , and the number of feature channels is corresponding to the real number field.
[0061] In this embodiment, the cross-attention calculation is represented by the following formula:
[0062] ;
[0063] where is the multi-head cross-attention operation, is the feed-forward network; is the mapped visual feature, is the mapped language feature, is the new visual feature, is the new language feature.
[0064] In this embodiment, a locator is also installed on the robot. After the prompt-guided cross-modal interaction, the above new visual feature and the new language feature are finally output, and then the locator decodes the robot position according to and to obtain a pixel-level representation of the positioning result, that is, the positioning coordinates, and then determine the central position of the target object. The robotic arm grabs the target object according to the coordinates of the central position to complete the human-machine interaction process.
[0065] The robot fine-grained target positioning method based on conditional multi-modal prompting provided by the embodiment of the present invention performs multiple cross-codings based on language encoding and visual encoding on images and texts respectively to obtain target visual features and target language features, and performs cross-attention calculation and decoding on the above two types of features to obtain a fine-grained target positioning result, and then completes the adjustment of the robot's motion posture. Combining the advantages of early and late fusion, it can achieve precise fine-grained target positioning of the robot according to free-form language expressions, improving the efficiency and accuracy of robot fine-grained target positioning.
[0066] In some embodiments, the target visual feature includes the visual feature output by each cross-coding, and the target language feature includes the language feature output by the last cross-coding. Before mapping the target visual feature and the target language feature to the same space, the robot fine-grained target positioning method based on conditional multi-modal prompting further includes: reshaping the visual features output by each cross-coding respectively, and splicing the reshaped visual features to obtain a new target visual feature.
[0067] In this embodiment, the target visual features include v2, v3, and v4, which are visual features with language perception of different contents, and the target language feature is .
[0068] In this embodiment, before projecting the features to the same space, v2, v3, and v4 are reshaped respectively to unify the formats of visual features with different shapes or dimensions, and then the reshaped visual features are feature-spliced to obtain a visual feature including rich visual information and language information.
[0069] The method for fine-grained target localization of a robot based on conditional multi-modal prompts provided by the embodiments of the present invention reshapes the visual features output by each cross-encoding respectively, and splices the reshaped visual features, so as to enhance the representation ability of the target visual features, and further improve the accuracy of the fine-grained target localization result.
[0070] In some embodiments, the visual encoding is implemented through the following steps: the image or visual features are visually encoded by a pre-trained Swin Transformer.
[0071] In this embodiment, the pre-trained Swin Transformer adopts a hierarchical structure, which can extract features at different scales and has strong representation learning ability; each visual encoder is divided into four stages, and the output of the th visual stage is denoted as , where represents the number of channels of the feature map, and represents the spatial dimension (same as above ). Each language encoder is divided into four stages, and the output of the th language stage is denoted as , where is the maximum length of the expression, and represents the number of channels of the language features. Through the staged interaction between these features, two-way early fusion is achieved.
[0072] In this embodiment, the visual encoding of the image by the pre-trained Swin Transformer is implemented through the following steps:
[0073] (1) Preprocess the image data to be measured, such as resizing and normalizing, to ensure that the input image meets the input requirements of the pre-trained model.
[0074] (2) Select a trained Swin Transformer model according to the task requirements (such as classification, detection, segmentation, etc.), and input the preprocessed image into the pre-trained Swin Transformer model. The model outputs visual feature maps, which contain the feature information of the image at different visual field scales and positions.
[0075] In this embodiment, the language encoding is implemented through the following steps: the text or language features are linguistically encoded by a BERT model.
[0076] Specifically, the steps of linguistically encoding the text by the BERT model include:
[0077] (1)Preprocess the text to be detected by performing operations such as word segmentation, removing stop words, removing punctuation marks, and lowercasing to ensure that the input text meets the input requirements of the BERT model; in addition, for Chinese text, word segmentation is usually also required, and the Chinese word segmentation tool provided by BERT or other third-party word segmentation tools can be used.
[0078] (2)Select an appropriate BERT model version (such as BERT-base, BERT-large, etc.) according to the task requirements (such as text classification, named entity recognition, question answering system, etc.), and this model can be pre-trained on a large-scale corpus (such as Wikipedia, BooksCorpus, etc.); perform the following transformations on the preprocessed text: add special tokens (such as [CLS] and [SEP]), construct the input sequence, calculate the input mask and segment index, etc., to ensure that the text is converted into an input format that the BERT model can accept. Input the processed input sequence into the BERT model, and the model will output a series of hidden states, that is, language features, and these language features contain the feature information of the text at different positions and in different contexts.
[0079] The robot fine-grained target localization method based on conditional multi-modal prompts provided by the embodiments of the present invention can extract features containing rich visual information through visual encoding of images or visual features by a pre-trained Swin Transformer, and can extract features containing rich language information through language encoding of text or language features by a BERT model, providing high-quality feature data for subsequent cross-modal fusion.
[0080] In some embodiments, determining the first prompt guidance according to the i-th visual feature obtained with an image as the initial input includes: projecting the i-th visual feature into the language feature space based on the linear layer of the conditional multi-modal prompt generator to obtain the first prompt guidance.
[0081] In Figure 2 In the shown embodiment, the conditional prompt generator derives visual prompts from the visual features to guide the generation of language features in the next stage; these prompts are projected into the language feature space through the linear layer, and the projected prompts are connected to the language features to form the input for the subsequent stage of the language encoder; the conditional prompt generator uses pyramid average pooling to generate the first prompt guidance.
[0082] Determining the second prompt guidance according to the i-th language feature obtained with text as the initial input includes: projecting the i-th language feature into the visual feature space based on the linear layer to obtain the second prompt guidance.
[0083] In Figure 2In the illustrated embodiment, hints about the language are derived from the language features by a conditional hint generator to guide the generation of visual features in the next stage; these hints are projected into the visual feature space through a linear layer, and the projected hints form the input for subsequent stages of the language encoder; the conditional hint generator uses pyramid average pooling to generate a second hint guidance.
[0084] The robot fine-grained target localization method based on conditional multimodal hints provided by the embodiments of the present invention projects the i-th visual feature into the language feature space through a linear layer of the conditional multimodal hint generator to obtain a first hint guidance, and projects the i-th language feature into the visual feature space through a linear layer to obtain a second hint guidance, realizing cross-sensory early fusion, and thus improving the accuracy of subsequent fine-grained target localization.
[0085] In some embodiments, after obtaining new visual features and new language features, the robot fine-grained target localization method based on conditional multimodal hints further includes: performing pixel-level position decoding based on the robot locator according to the new visual features and new language features to obtain a position decoding result; wherein, the robot locator includes at least three upsampling decoders; converting the position decoding result according to the spatial coordinate system corresponding to the robotic arm of the robot to obtain localization parameters, and adjusting the motion posture of the robotic arm according to the localization parameters.
[0086] In this embodiment, the target locator is composed of three upsampling decoder modules, which can fully integrate multi-scale visual information and language information to obtain a pixel-level fine-grained prediction result.
[0087] In Figure 2 the illustrated embodiment, through the new visual features and the new language features ; performing upsampling decoding to obtain a pixel-level fine-grained prediction result .
[0088] In this embodiment, the processor of the robot processes the input image and text by executing the above-mentioned robot fine-grained target localization method based on conditional multimodal hints to generate a fine-grained localization result of the target object, and determines the central position of the target object. Then, the position of the target object in the image coordinate system is transformed through a rotation matrix to convert it into a three-dimensional position in the robotic arm coordinate system. Finally, the robotic arm grasps the target object according to this three-dimensional coordinate to complete the service process for people.
[0089] The robot fine-grained target positioning method based on conditional multi-modal cues provided by the embodiments of the present invention decodes the pixel-level position according to the new visual features and new language features based on the robot locator, obtains the pixel-level interactive object positioning result and its specific position, and then flexibly adjusts the motion posture of the robotic arm, improving the efficiency and accuracy of robot motion control.
[0090] Figure 3 FIG. 3 is a schematic flowchart of the robot fine-grained target positioning method based on conditional multi-modal cues provided by the present invention. In Figure 3 the illustrated embodiment, the voice data of the surrounding environment is received through the microphone of the robot, and the voice data is converted into text by the voice recognition module, and the image data of the surrounding environment is obtained through the camera; the text and the image data are input into the robot fine-grained language-vision positioning model based on conditional multi-modal cues (obtained by training the robot fine-grained target positioning method based on conditional multi-modal cues) to generate the fine-grained positioning result of the target object, and then the position of the target object is transformed to the coordinate system of the robotic arm through coordinate transformation for the robotic arm to perform related operations.
[0091] Next, the robot fine-grained target positioning device based on conditional multi-modal cues provided by the present invention will be described. The robot fine-grained target positioning device described below and the robot fine-grained target positioning method based on conditional multi-modal cues described above can be referred to each other correspondingly.
[0092] Figure 4 FIG. 4 is a schematic structural diagram of the robot fine-grained target positioning device based on conditional multi-modal cues provided by the present invention. As Figure 4 shown, the robot fine-grained target positioning device based on conditional multi-modal cues includes: an encoding module 410 and a feature calculation module 420.
[0093] The encoding module 410 is configured to perform multiple cross-encodings on the image and the text respectively to obtain target visual features and target language features; wherein, in each cross-encoding, a first cue guidance is determined according to the i-th visual feature obtained with the image as the initial input, and the first cue guidance and the i-th visual feature are subjected to language encoding to obtain the (i + 1)-th language feature; a second cue guidance is determined according to the i-th language feature obtained with the text as the initial input, and the second cue guidance is subjected to visual encoding to obtain the (i + 1)-th visual feature; i is a positive integer greater than 0;
[0094] The feature calculation module 420 maps the target visual feature and the target language feature to the same space, and performs cross-attention calculation on the mapped visual feature and the mapped language feature to obtain a new visual feature and a new language feature, so that the robot can adjust the motion posture according to the position decoding result under the condition of performing position decoding on the new visual feature and the new language feature.
[0095] The robot fine-grained target positioning device provided by the embodiment of the present invention performs cross-encoding based on language encoding and visual encoding on the image and text respectively for multiple times to obtain the target visual feature and the target language feature, performs cross-attention calculation and decoding on the above two types of features to obtain the fine-grained target positioning result, and then completes the adjustment of the robot's motion posture. It combines the advantages of early and late fusion, can achieve precise fine-grained target positioning of the robot according to free-form language expressions, and improves the efficiency and accuracy of the robot's fine-grained target positioning.
[0096] In some embodiments, the robot fine-grained target positioning device based on conditional multi-modal prompts further includes: a decoding module, configured to perform pixel-level position decoding based on the robot locator according to the new visual feature and the new language feature after obtaining the new visual feature and the new language feature, to obtain a position decoding result; wherein, the robot locator includes at least three upsampling decoders; a motion adjustment module, which converts the position decoding result according to the space coordinate system corresponding to the robot's manipulator to obtain positioning parameters, and adjusts the motion posture of the manipulator according to the positioning parameters.
[0097] The robot fine-grained target positioning device provided by the embodiment of the present invention performs pixel-level position decoding based on the robot locator according to the new visual feature and the new language feature to obtain the pixel-level interactive object positioning result and its specific position, and then realizes the flexible adjustment of the manipulator's motion posture, improving the efficiency and accuracy of the robot's motion control.
[0098] Figure 5 is a schematic structural diagram of the electronic device provided by the present invention, as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute a refined target positioning method for a robot based on conditional multi-modal prompts. The method includes: performing multiple cross-codings on the image and the text respectively to obtain a target visual feature and a target language feature; wherein, in each cross-coding, a first prompt guidance is determined according to the i-th visual feature obtained with the image as the initial input, and the first prompt guidance and the i-th visual feature are subjected to language coding to obtain the (i + 1)-th language feature; a second prompt guidance is determined according to the i-th language feature obtained with the text as the initial input, and the second prompt guidance is subjected to visual coding to obtain the (i + 1)-th visual feature; i is a positive integer greater than 0; the target visual feature and the target language feature are mapped to the same space, and cross-attention calculation is performed on the mapped visual feature and the mapped language feature to obtain a new visual feature and a new language feature, so that the robot can adjust the motion posture according to the position decoding result under the condition of performing position decoding on the new visual feature and the new language feature.
[0099] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0100] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the robot refined target positioning method based on conditional multi-modal prompts provided by the above-mentioned various methods. The method includes: performing multiple cross-codings on an image and text respectively to obtain a target visual feature and a target language feature; wherein, in each cross-coding, a first prompt guidance is determined according to the i-th visual feature obtained with the image as the initial input, and the first prompt guidance and the i-th visual feature are subjected to language coding to obtain the (i + 1)-th language feature; a second prompt guidance is determined according to the i-th language feature obtained with the text as the initial input, and the second prompt guidance is subjected to visual coding to obtain the (i + 1)-th visual feature; i is a positive integer greater than 0; the target visual feature and the target language feature are mapped to the same space, and cross-attention calculation is performed on the mapped visual feature and the mapped language feature to obtain a new visual feature and a new language feature, so that the robot can adjust its motion posture according to the position decoding result under the condition of performing position decoding on the new visual feature and the new language feature.
[0101] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the robot refined target positioning method based on conditional multi-modal prompts provided by the above-mentioned various methods. The method includes: performing multiple cross-codings on an image and text respectively to obtain a target visual feature and a target language feature; wherein, in each cross-coding, a first prompt guidance is determined according to the i-th visual feature obtained with the image as the initial input, and the first prompt guidance and the i-th visual feature are subjected to language coding to obtain the (i + 1)-th language feature; a second prompt guidance is determined according to the i-th language feature obtained with the text as the initial input, and the second prompt guidance is subjected to visual coding to obtain the (i + 1)-th visual feature; i is a positive integer greater than 0; the target visual feature and the target language feature are mapped to the same space, and cross-attention calculation is performed on the mapped visual feature and the mapped language feature to obtain a new visual feature and a new language feature, so that the robot can adjust its motion posture according to the position decoding result under the condition of performing position decoding on the new visual feature and the new language feature.
[0102] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A fine-grained target localization method for robots based on conditional multi-modal prompts, characterized in that, Including: Performing multiple cross-codings on the image and text respectively to obtain target visual features and target language features; wherein, in each cross-coding, a first prompt guidance is determined according to the i-th visual feature obtained with the image as the initial input, and the first prompt guidance and the i-th language feature obtained with the text as the initial input are subjected to language coding to obtain the (i + 1)-th language feature; a second prompt guidance is determined according to the i-th language feature, and the second prompt guidance is subjected to visual coding to obtain the (i + 1)-th visual feature; i is a positive integer greater than 0; the target visual features include multiple visually-perceptive visual features; Mapping the target visual features and the target language features to the same space, and performing cross-attention calculation on the mapped visual features and the mapped language features to obtain new visual features and new language features, for the robot to adjust the motion posture according to the position decoding result under the condition of performing position decoding on the new visual features and the new language features; the target visual features include the visual features output by each cross-coding, and the target language features include the language features output by the last cross-coding; Before mapping the target visual features and the target language features to the same space, the method further includes: Reshaping the visual features output by each cross-coding respectively, and splicing the reshaped visual features to obtain new target visual features.
2. The robot fine-grained target localization method based on conditional multi-modal prompts according to claim 1, wherein The cross-attention calculation is represented by the following formula: ; Among them, is the multi-head cross-attention operation, is the feed-forward network; is the mapped visual feature, is the mapped language feature, is the new visual feature, is the new language feature.
3. The method for fine-grained target localization of a robot based on conditional multi-modal prompts according to claim 1, wherein The visual coding is implemented through the following steps: Performing visual coding on the image or visual features through a pre-trained Swin Transformer; The language coding is implemented through the following steps: Performing language coding on the text or language features through a BERT model.
4. The method for refined target positioning of a robot based on conditional multi-modal prompts according to claim 1, wherein The determining the first prompt guidance according to the i-th visual feature obtained with the image as the initial input includes: Projecting the i-th visual feature to the language feature space through a linear layer of a conditional multi-modal prompt generator to obtain the first prompt guidance; The determining the second prompt guidance according to the i-th language feature obtained with the text as the initial input includes: Projecting the i-th language feature to the visual feature space through the linear layer to obtain the second prompt guidance.
5. The method for fine-grained target localization of a robot based on conditional multi-modal prompts according to claim 1, wherein After obtaining the new visual features and the new language features, the method further includes: Performing pixel-level position decoding on the new visual features and the new language features based on a robot locator to obtain the position decoding result; wherein, the robot locator includes at least three upsampling decoders; Converting the position decoding result according to the spatial coordinate system corresponding to the robot's manipulator to obtain positioning parameters, and adjusting the motion posture of the manipulator according to the positioning parameters.
6. A robot refined target positioning device based on conditional multi-modal prompts, characterized in that, Including: An encoding module for performing multiple cross-encodings on images and texts respectively to obtain target visual features and target language features; wherein, in each cross-encoding, a first prompt guidance is determined according to the i-th visual feature obtained with the image as the initial input, and the first prompt guidance and the i-th language feature obtained with the text as the initial input are subjected to language encoding to obtain the (i + 1)-th language feature; a second prompt guidance is determined according to the i-th language feature, and the second prompt guidance is subjected to visual encoding to obtain the (i + 1)-th visual feature; i is a positive integer greater than 0; the target visual features include multiple visually-perceptive language features; A feature calculation module for mapping the target visual features and the target language features to the same space, and performing cross-attention calculation on the mapped visual features and the mapped language features to obtain new visual features and new language features, for the robot to adjust the motion posture according to the position decoding result under the condition of performing position decoding on the new visual features and the new language features; The target visual features include the visual features output by each cross-encoding, and the target language features include the language features output by the last cross-encoding; The feature calculation module is further configured to: Before mapping the target visual features and the target language features to the same space, reshape the visual features output by each cross-encoding respectively, and splice the reshaped visual features to obtain new target visual features.
7. The robot refined target positioning device based on conditional multi-modal prompts according to claim 6, characterized in that, The device further includes: A decoding module for, after obtaining the new visual features and the new language features, performing pixel-level position decoding based on a robot locator according to the new visual features and the new language features to obtain the position decoding result; wherein, the robot locator includes at least three upsampling decoders; A motion adjustment module for converting the position decoding result according to the spatial coordinate system corresponding to the robot's robotic arm to obtain positioning parameters, and adjusting the motion posture of the robotic arm according to the positioning parameters.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the robot fine-grained target positioning method based on conditional multi-modal prompts as described in any one of claims 1 to 5.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the robot fine-grained target positioning method based on conditional multi-modal prompts as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Vision-language-action joint modeling-based disordered scene target object capturing method
CN115861596A