Image processing and model training method, device and equipment

By splicing images and language instructions from the robot's perspective and using visual language models for prediction, the problem of low coupling between the robot's target recognition and trajectory planning in a dynamic environment is solved, achieving more efficient image processing and human-computer interaction.

CN120612495APending Publication Date: 2025-09-09BEIJING GALBOT AI CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510709767.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In the existing technology, when a robot performs embodied visual target following tasks in a dynamic environment, the coupling between the target recognition and trajectory planning models is low, resulting in error accumulation and difficulty in effectively processing language description instructions in complex scenarios.

Method used

By obtaining image sequences and language command information from the robot's perspective, converting them into visual tokens and language tokens, and splicing them based on the task type, the visual language model is used for prediction to achieve image processing tasks.

Benefits of technology

It improves the robot's target recognition and trajectory planning capabilities in dynamic environments, achieves a higher level of human-computer interaction and intelligence, enhances the ability to collaboratively process multimodal information, and improves the accuracy and efficiency of image processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612495A_ABST
    Figure CN120612495A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing and model training method, device and equipment. The method comprises the following steps: acquiring an image sequence of a robot view angle and language instruction information; the language instruction information is used for representing the task type of the image processing task; determining a visual token corresponding to at least one frame of image in the image sequence and a language token corresponding to the language instruction information; based on the task type of the image processing task, splicing the language token and the visual token to obtain a spliced token; performing prediction processing on the spliced tokens to obtain predicted tokens; and performing an image processing task on the prediction token.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer deep learning technology, and in particular to an image processing and model training method, device and equipment. Background Art

[0002] The embodied visual object following task requires a robot to follow a verbally described object in a dynamic environment using only its own visual observations. This task is inherently challenging, requiring both accurate object recognition and effective trajectory planning under severe occlusion and high scene dynamics. High scene dynamics refers to scenes with a large number of dynamically changing elements, such as scenes with significant brightness changes or scenes containing a large number of interactive, dynamic objects.

[0003] In related technologies, embodied visual object following tasks are typically performed by decoupling object recognition and trajectory planning into a detection model and a planning model, respectively. However, the coupling between the detection and planning models is weak, leading to error accumulation between the recognition and planning models. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide at least one image processing and model training method, device and equipment.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides an image processing method, the method comprising:

[0007] Acquire an image sequence and language instruction information from the robot's perspective; the language instruction information is used to characterize the task type of the image processing task;

[0008] Determining a visual token corresponding to at least one frame of the image sequence, and a language token corresponding to the language instruction information;

[0009] Based on the task type of the image processing task, the language token and the visual token are spliced ​​to obtain a spliced ​​token;

[0010] Performing prediction processing on the spliced ​​tokens to obtain predicted tokens;

[0011] The image processing task is performed on the predicted token.

[0012] In some embodiments, determining the visual tokens corresponding to at least one frame of image in the image sequence includes: when the task type represents a trajectory planning task, determining the first visual token of the historical image in the image sequence and the second visual token of the current image in the image sequence; when the task type represents a question-answering task, determining the third visual token corresponding to at least one frame of image in the image sequence; wherein the resolution of the second visual token is greater than the resolution of the first visual token and the resolution of the third visual token.

[0013] In some embodiments, the language token and the visual token are spliced ​​based on the task type of the image processing task, including: when the task type represents a trajectory planning task, the first visual token, the second visual token, the language token and the preset token are spliced; the preset token is used to represent that the image processing task includes a trajectory planning task; when the task type represents a question-answering task, the third visual token and the language token are spliced.

[0014] In some embodiments, performing the image processing task on the prediction token includes: determining the task type of the image processing task based on the spliced ​​token corresponding to the prediction token; and performing image processing on the prediction token based on the network model corresponding to the task type.

[0015] In some embodiments, the image processing task is a trajectory planning task; the network model corresponding to the trajectory planning task includes a diffusion motion model; the image processing of the prediction token based on the network model corresponding to the task type includes: using the diffusion motion model, based on the prediction token and a first trajectory set, determining at least two predicted trajectories and the scores corresponding to the at least two predicted trajectories respectively; the first trajectory set is obtained by clustering multiple sample trajectories in a sample trajectory set; the method also includes: determining the predicted trajectory with the highest score among the at least two predicted trajectories as the target trajectory.

[0016] In some embodiments, the method further includes: adding noise data to each trajectory in the first trajectory set to obtain a second trajectory set; and using the diffusion motion model to determine at least two predicted trajectories and scores corresponding to the at least two predicted trajectories based on the prediction tokens and the first trajectory set, includes: using the diffusion motion model to process the prediction tokens and the second trajectory set to obtain at least two predicted trajectories and scores corresponding to the at least two predicted trajectories.

[0017] In a second aspect, an embodiment of the present application provides a model training method, the method comprising:

[0018] Acquire at least two sample image sequences and sample language instructions corresponding to the at least two sample image sequences respectively; different sample language instructions correspond to different image processing tasks;

[0019] Determining, by using an image processing model, sample visual tokens corresponding to the at least two sample image sequences, and sample language tokens corresponding to the at least two sample language instructions;

[0020] Based on the task type of the image processing task, performing splicing processing on at least two of the sample language tokens and at least two of the sample visual tokens to obtain at least two spliced ​​sample tokens;

[0021] Using an image processing model, performing prediction processing on the at least two spliced ​​sample tokens to obtain at least two predicted sample tokens; performing corresponding image processing tasks on the at least two predicted sample tokens to obtain at least two image processing results;

[0022] Determining losses corresponding to different task types based on the at least two image processing results;

[0023] Based on the losses corresponding to the different task types, the network parameters of the image processing model are adjusted.

[0024] In some embodiments, the method further includes: obtaining a sample trajectory set; performing clustering processing on multiple sample trajectories in the sample trajectory set to obtain a first trajectory set; adding noise data to each trajectory in the first trajectory set to obtain a third trajectory set; when the image processing task is a trajectory planning task, performing corresponding image processing tasks on the at least two predicted sample tokens respectively to obtain at least two image processing results, including: determining at least two predicted sample trajectories and the scores corresponding to the at least two predicted sample trajectories respectively based on the predicted sample tokens corresponding to the trajectory planning task and the third trajectory set.

[0025] In some embodiments, the third trajectory set includes a first sample trajectory and a second sample trajectory; the first sample trajectory is annotated with a first label score, and the second sample trajectory is annotated with a second label score; determining the losses corresponding to different task types based on the at least two image processing results includes: determining the similarity between the at least two predicted sample trajectories and the first sample trajectory respectively; determining the predicted sample trajectory with the greatest similarity among the at least two predicted sample trajectories as the target predicted sample trajectory; determining the first trajectory loss based on the target predicted sample trajectory and the first sample trajectory; determining the second trajectory loss based on the scores corresponding to the at least two predicted sample trajectories, the first label score and the second label score; determining the third trajectory loss based on the first trajectory loss and the second trajectory loss.

[0026] In some embodiments, adjusting the network parameters of the image processing model based on the losses corresponding to the different task types includes: adjusting the network parameters of the diffusion action model in the image processing model for performing trajectory planning tasks based on the third trajectory loss; adjusting the network parameters of the large language model in the image processing model for performing question-answering tasks based on the text loss; adjusting the network parameters of the token generation model in the image processing model for generating visual tokens based on the third trajectory loss and the text loss, and adjusting the network parameters of the visual language model in the image processing model for obtaining predicted sample tokens.

[0027] In a third aspect, an embodiment of the present application provides an image processing device, comprising:

[0028] A first acquisition unit is configured to acquire an image sequence and language instruction information from the robot's perspective; the language instruction information is used to characterize a task type of the image processing task;

[0029] A first determining unit is configured to determine a visual token corresponding to at least one frame of the image sequence, and a language token corresponding to the language instruction information;

[0030] A first splicing unit is configured to splice the language token and the visual token based on the task type of the image processing task to obtain a spliced ​​token;

[0031] A prediction unit, configured to perform prediction processing on the concatenated tokens to obtain a predicted token;

[0032] The first processing unit is configured to perform the image processing task on the prediction token.

[0033] In a fourth aspect, an embodiment of the present application provides a model training device, comprising:

[0034] A second acquisition unit is configured to acquire at least two sample image sequences and sample language instructions corresponding to the at least two sample image sequences; different sample language instructions correspond to different image processing tasks;

[0035] a second determining unit, configured to determine, by using an image processing model, sample visual tokens corresponding to the at least two sample image sequences, and sample language tokens corresponding to the at least two sample language instructions;

[0036] A second splicing unit is configured to splice at least two of the sample language tokens and at least two of the sample visual tokens based on a task type of the image processing task to obtain at least two spliced ​​sample tokens;

[0037] A second processing unit is configured to perform prediction processing on the at least two spliced ​​sample tokens using an image processing model to obtain at least two predicted sample tokens; and perform corresponding image processing tasks on the at least two predicted sample tokens to obtain at least two image processing results.

[0038] a loss determining unit, configured to determine losses corresponding to different task types based on the at least two image processing results;

[0039] An adjustment unit is used to adjust the network parameters of the image processing model based on the losses corresponding to the different task types.

[0040] In a fifth aspect, an embodiment of the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.

[0041] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, some or all of the steps in the above method are implemented.

[0042] In a seventh aspect, an embodiment of the present application provides a computer program product, comprising a computer program or instructions, which implement some or all of the steps in the above method when executed by a processor.

[0043] In the embodiment of the present application, first, the image sequence and language instruction information from the robot's perspective are obtained and converted into visual tokens and language tokens; secondly, these tokens are spliced ​​based on the task type to generate an input representation corresponding to the task type (i.e., the spliced ​​tokens); then the different spliced ​​tokens are processed in the same way to obtain prediction tokens; finally, the corresponding image processing task is executed. In this way, on the one hand, the robot can be driven to perform the corresponding task through language instructions, which realizes a higher level of human-computer interaction and improves the robot's ease of use and intelligence; on the other hand, the visual information and language information are fused and processed, which improves the multimodal information coordination ability and enhances the model's understanding and adaptability to complex tasks; on the other hand, for different image processing tasks (such as trajectory planning tasks and question-answering tasks), the spliced ​​tokens are processed using the same prediction method, so that the models performing different image processing tasks can be tightly coupled together. On the other hand, through the guidance mechanism of the task type, the model can adaptively switch different processing flows, thereby improving the accuracy and efficiency of the image processing task.

[0044] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0046] Figure 1 Schematic diagram of the implementation process of an image processing method provided in the embodiment of the present application Figure 1 ;

[0047] Figure 2 Schematic diagram of the implementation process of an image processing method provided in the embodiment of the present application Figure 2 ;

[0048] Figure 3 Schematic diagram of the implementation process of an image processing method provided in the embodiment of the present application Figure 3 ;

[0049] Figure 4 Schematic diagram of the implementation process of a model training method provided in the embodiment of the present application Figure 1 ;

[0050] Figure 5 Schematic diagram of the implementation process of a model training method provided in the embodiment of the present application Figure 2 .

[0051] Figure 6Schematic diagram of the image processing model provided in the embodiment of this application Figure 1 ;

[0052] Figure 7 Schematic diagram of the image processing model provided in the embodiment of this application Figure 2 ;

[0053] Figure 8 A schematic diagram of the architecture of the diffusion action model provided in an embodiment of the present application;

[0054] Figure 9 A schematic diagram of embodied visual tracking data provided by an embodiment of the present application;

[0055] Figure 10 A schematic diagram of video-based question-answering data provided in an embodiment of the present application;

[0056] Figure 11 A schematic diagram of a scene of an image processing method provided in an embodiment of the present application;

[0057] Figure 12 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application;

[0058] Figure 13 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;

[0059] Figure 14 A schematic diagram of a hardware entity of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this application are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0061] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0062] The terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing this application only and are not intended to limit this application.

[0064] The embodied visual object following task requires a robot to follow a verbally described object in a dynamic environment using only its own visual observations. This task is inherently challenging as it requires both accurate object recognition and efficient trajectory planning under severe occlusion and high scene dynamics.

[0065] Related Technology 1: This solution is primarily used in scenarios such as drone tracking. Users manually select a target in the observed image. The system then uses traditional control models such as Kalman filtering to perform path planning and motion control based on the target's position in the image, completing autonomous tracking of the target.

[0066] Related Technology 2: The solution achieves modular separation of perception and decision-making in the system architecture. The perception module detects and locates objects based on a basic visual model, outputting candidate bounding boxes or segmentation masks. Subsequently, the reinforcement learning-based planning module uses the perception results as input to learn and execute the optimal motion strategy for continuous tracking of the target.

[0067] As can be seen, traditional visual following algorithms rely on manual annotation of target locations and are unable to support direct input of semantic commands. Related technologies often adopt a modular design, decoupling target recognition and motion strategy planning into two serial modules, which can easily lead to cumulative errors between perception and decision-making. Following methods in related technologies generally lack strong recognition capabilities for verbally described targets. Some methods require manual target specification or only have coarse-grained target recognition capabilities, making it difficult to parse and execute detailed instructions such as "Follow the woman in red" in complex scenarios.

[0068] In order to solve the above technical problems, the embodiment of the present application provides an image processing method, which can be applied to various electronic devices, including but not limited to fixed devices and / or mobile devices. For example, the fixed devices include but are not limited to: personal computers (PCs), or servers, etc. The servers can be cloud servers or ordinary servers. The mobile devices include but are not limited to: one or more of mobile phones, tablet computers or wearable devices. Figure 1 As shown, the image processing method includes steps S101 to S105:

[0069] Step S101 : obtaining an image sequence from the robot's perspective and language instruction information; the language instruction information is used to characterize the task type of the image processing task.

[0070] Here, the image sequence typically consists of consecutive image frames captured by the robot's onboard camera or other visual sensors, reflecting a first-person perspective of the robot's current environment. Language instructions can include natural language descriptions provided by the user through voice input, keyboard input, or touchscreen input, such as "Please identify and track the person in red" or "Please answer what you see." This language instruction not only specifies the task objective but also implies the task type (such as tracking, recognition, question-and-answer, etc.).

[0071] In an embodiment of the present application, after the robot collects an image sequence through its own camera or other visual sensors, the collected image sequence can be sent to the server, so that the server obtains the image sequence from the robot's perspective.

[0072] In an embodiment of the present application, when a user interacts with a robot through voice, the robot can parse the received voice data to obtain language instruction information, and then send the language instruction information to the server, or send the received voice data to the server, and then the server parses the voice data to obtain language instruction information.

[0073] In some embodiments, when the user inputs language instruction information through the terminal, the terminal can send the language instruction information to the server. Exemplarily, the terminal can be a browser or a software application such as a robot application.

[0074] In the embodiments of this application, the role of language instruction information is to guide the model to select the appropriate processing branch. For example, when the language instruction information contains keywords such as tracking, following, and navigation, the system will identify the task as a tracking task and execute the image processing task corresponding to the tracking branch. If the instruction is "What can you see?" or "Please describe what you see," it will be determined as a recognition or question-answering task, and the image processing task corresponding to the recognition branch will be executed.

[0075] Step S102: determining the visual tokens corresponding to at least one frame of image in the image sequence and the language tokens corresponding to the language instruction information.

[0076] Here, a visual token is used to represent the visual features of an image block, which contains the semantic information and spatial information of the image block.

[0077] In an embodiment of the present application, for each frame of image, the image can be segmented to obtain multiple image blocks, and then the visual features of each image block are extracted, and the visual features of the image block are projected into a predetermined space, thereby obtaining visual tokens corresponding to the multiple image blocks of each frame of image. For example, each frame of image can be segmented to obtain multiple image blocks of size 224×224, and then each image block is input into a pre-trained visual encoder (such as EVA-CLIP) to obtain the visual features of each image block, and then the visual features of each image block are input into a cross-modal projector to obtain the visual token corresponding to each image block.

[0078] A language token refers to the smallest semantic unit of language instruction information, which can be obtained by performing operations such as word segmentation and embedding on the language instruction information.

[0079] In an embodiment of the present application, the language instruction information can be tokenized by a word segmenter to obtain a language token of the language instruction information. For example, when the language instruction information is "Please identify and track the person wearing red clothes", the word segmenter can decompose the language instruction information into multiple sub-word units, such as: please / identify / and / track / the / person / wearing / red / clothes / , where each sub-word unit will be converted into a language token.

[0080] Step S103 : Based on the task type of the image processing task, the language token and the visual token are spliced ​​together to obtain a spliced ​​token.

[0081] In the embodiment of the present application, different task types of image processing tasks correspond to different spliced ​​tokens. In other words, the splicing method of language tokens and visual tokens is determined by the task type of the image processing task.

[0082] It is understandable that in the embodiment of the present application, it is necessary to execute the corresponding image processing task according to the task type of the image processing task represented by the language instruction information. The processing objects of different image processing tasks are language tokens and visual tokens. Therefore, in order to be able to identify the task type of the current image processing task through input parameters (i.e., language tokens and visual tokens), it is necessary to perform corresponding splicing processing on the language tokens and visual tokens according to different task types.

[0083] In some embodiments, splicing methods corresponding to different task types can be predetermined. During the inference phase, after determining the task type of the image processing task, the corresponding splicing method can be determined based on the current task type. The visual tokens and language tokens are then spliced ​​based on the splicing method. In this way, the task type of the current image processing task can be identified through the spliced ​​tokens, and the spliced ​​tokens can be input into the model of the corresponding processing branch to execute the corresponding image processing task.

[0084] Step S104: performing prediction processing on the concatenated tokens to obtain predicted tokens.

[0085] Here, the predicted token may be the token with the highest probability of appearing next.

[0086] In the embodiment of the present application, the same processing method, i.e., prediction processing, is used for the spliced ​​tokens corresponding to different task types. In other words, the same prediction processing is required before executing different image processing tasks. In this way, the models performing different image processing tasks can be tightly coupled together.

[0087] In an embodiment of the present application, a pre-trained video-based visual language model (VLM) can be used to predict the next token with the highest probability of occurrence after the spliced ​​tokens, i.e., the predicted token, using an autoregressive method. In other words, the input parameters of the visual language model include the spliced ​​tokens of image processing tasks of different task types.

[0088] Step S105: performing the image processing task on the prediction token.

[0089] In an embodiment of the present application, the task type of the image processing task to be performed can be first determined based on the prediction token, and then the prediction token can be input into the network model corresponding to the task type, so that the corresponding image processing task can be performed through the network model.

[0090] In the embodiment of the present application, first, the image sequence and language instruction information from the robot's perspective are obtained and converted into visual tokens and language tokens; secondly, these tokens are spliced ​​based on the task type to generate an input representation corresponding to the task type (i.e., the spliced ​​tokens); then the different spliced ​​tokens are processed in the same way to obtain prediction tokens; finally, the corresponding image processing task is executed. In this way, on the one hand, the robot can be driven to perform the corresponding task through language instructions, which realizes a higher level of human-computer interaction and improves the robot's ease of use and intelligence; on the other hand, the visual information and language information are fused and processed, which improves the multimodal information coordination ability and enhances the model's understanding and adaptability to complex tasks; on the other hand, for different image processing tasks (such as trajectory planning tasks and question-answering tasks), the spliced ​​tokens are processed using the same prediction method, so that the models performing different image processing tasks can be tightly coupled together. On the other hand, through the guidance mechanism of the task type, the model can adaptively switch different processing flows, thereby improving the accuracy and efficiency of the image processing task.

[0091] In some embodiments, the “determining visual tokens corresponding to at least one frame of the image sequence” in step S102 may include:

[0092] Based on the task type of the image processing task, visual tokens corresponding to at least one frame of image in the image sequence are determined.

[0093] Here, different image processing tasks have different task types and corresponding visual token resolutions.

[0094] It is understandable that the resolution of the visual token determines the processing accuracy of the image processing task. The larger the resolution of a single visual token, the more image blocks a single image is divided into, so the model's processing accuracy for the image will also be higher. The smaller the resolution of a single visual token, the fewer image blocks a single image is divided into, so the model's processing accuracy for the image will also be lower. Therefore, in order to improve the efficiency of the model in performing image processing tasks, for image processing tasks that do not require high precision (such as question-answering tasks), visual tokens with lower resolutions can be determined; for image processing tasks that require high precision (such as trajectory planning tasks), visual tokens with higher resolutions can be determined.

[0095] In some embodiments, the step S102 of “determining, based on the task type of the image processing task, the visual tokens corresponding to at least one frame of the image sequence” may be implemented by steps S1021 and S1022:

[0096] Step S1021 , when the task type represents a trajectory planning task, determining a first visual token of a historical image in the image sequence and a second visual token of a current image in the image sequence.

[0097] In embodiments of the present application, the image sequence captured by the robot includes historical images and a current image. The current image may refer to the image corresponding to the current timestamp in the image sequence, and the historical image may refer to the image corresponding to a previous timestamp in the image sequence. In some embodiments, the historical image may refer to all images in the image sequence except the current image. In other embodiments, the historical image may refer to images in the image sequence that are a preset number of frames prior to the current image, where the preset number is less than or equal to the total number of frames in the image sequence.

[0098] For example, the image sequence is O T ={x1,…,x T}, where T is the number of frames, x T is the image corresponding to the current timestamp, and the historical image can be {x1,…,x T-1}; When the preset number is K, the historical image can also be {x T-K ,…,x T-1}, where K is less than T. For example, K may be 32.

[0099] In an embodiment of the present application, the resolution of the second visual token is greater than the resolution of the first visual token. It is understandable that in the case where the task type represents a trajectory planning task, it is necessary to determine the trajectory of the robot following the target object in the image based on the current image and the historical image. Since the core of the trajectory planning task is to continuously update the trajectory of the robot following the target object in the image based on the real-time status of the target object, the real-time status can be obtained from the current image of the target object. Therefore, the importance of the current image is greater than the importance of the historical image, so the resolution of the second visual token of the current image is greater than the resolution of the first visual token of the historical image. In other words, the length of the second visual token is greater than the length of the first visual token of the historical image, that is, the number of image blocks segmented by the current image is greater than the number of image blocks segmented by a single historical image.

[0100] In an embodiment of the present application, when the task type represents a trajectory planning task, fine-grained tokens can be determined based on the current image and coarse-grained tokens can be determined based on the historical image, thereby improving image processing efficiency without reducing image processing accuracy.

[0101] Step S1022 , when the task type represents a question-answering task, determining a third visual token corresponding to at least one frame of image in the image sequence.

[0102] In the embodiment of the present application, at least one frame of image in the image sequence may refer to all images in the image sequence, or may refer to a portion of images including the current image.

[0103] For example, the image sequence is O T ={x1,…,x T}, where T is the number of frames, x T For the current image, at least two frames of images may include {x1,…,x T}, and can also include {x T-N ,…,x T}.

[0104] In this embodiment of the present application, the resolution of the third visual token is smaller than that of the second visual token. It is understandable that the number of image features required for the question-answering task is smaller than the number of image features required for the trajectory planning task. Therefore, in order to increase the processing speed of the model, a third visual token with a smaller resolution can be determined.

[0105] In some embodiments, the resolution of the third visual token may be less than the resolution of the first visual token.

[0106] It is understandable that the core of the question-answering task is to extract high-level semantic information from the image (such as object categories, scene relationships, action attributes, etc.) to answer questions such as "what is it" and "where is it", which requires recognition through the global features of the image rather than detailed features. For example, when answering "Is there a dog in the image", the image processing model only needs to determine whether the overall semantic features of "dog" (such as a combination of limbs, tail, hair, etc.) exist in the image, without paying attention to the texture of the dog's whiskers or the specific shape of its paws". For the trajectory planning task, detailed features need to be extracted from the image. Therefore, for the question-answering task, global features need to be extracted from large blocks of the image, and for the trajectory planning task, detailed features need to be extracted from small blocks of the image. Therefore, the number of image blocks segmented by the image corresponding to the question-answering task is smaller than the number of image blocks segmented by the historical image corresponding to the trajectory planning task, that is, the length of the third visual token is smaller than the length of the first visual token of the historical image, and the resolution of the third visual token is smaller than the resolution of the first visual token.

[0107] In some embodiments, the resolution of the third visual token may be the same as the resolution of the first visual token.

[0108] It is understandable that if the question-answering task involves fine-grained visual reasoning (such as "What color is a cat's eye?" or "The pattern on the chair"), the number of tiles needs to be increased to capture local features, so the image needs to be divided into smaller sizes. Therefore, the number of image tiles divided by the image corresponding to the question-answering task can also be equal to the number of image tiles divided by the historical image corresponding to the trajectory planning task, that is, the length of the third visual token is equal to the length of the first visual token of the historical image, and the resolution of the third visual token is equal to the resolution of the first visual token.

[0109] In some embodiments, when the visual token is determined by extracting visual features of an image, since the visual features and the visual token have the same resolution, the resolutions of the first visual token, the second visual token, and the third visual token can be expressed by the following formula (1):

[0110]

[0111] in, is the visual feature corresponding to the first visual token, with a resolution of 64×C, where C is the embedding dimension; are the visual features corresponding to the second visual token and the third visual token, with a resolution of 4×C.

[0112] In the embodiment of this application, when the task type is trajectory planning, high-resolution visual tokens are used to enhance target recognition accuracy; while in question-answering tasks, low-resolution visual tokens are used to reduce computational overhead. In this way, on the one hand, by setting the resolution of visual tokens differently, it is possible to balance recognition accuracy and resource consumption in different tasks; on the other hand, this design enables the model to more flexibly respond to diverse task requirements and improve overall processing efficiency.

[0113] In some embodiments, as Figure 2 As shown, the above step S103 can be implemented through steps S201 and S202:

[0114] Step S201, when the task type represents a trajectory planning task, the first visual token, the second visual token, the language token and the preset token are spliced ​​to obtain a spliced ​​token; the preset token is used to represent that the image processing task includes a trajectory planning task.

[0115] Step S202 , when the task type represents a question-answering task, concatenate the third visual token and the language token to obtain a concatenated token.

[0116] Here, the preset token is used to identify the task category corresponding to the currently executed image processing task. Exemplarily, the preset token can be a specific string or a numerical code.

[0117] In an embodiment of the present application, when the image processing task is a trajectory planning task, a preset token can be added to the first visual token, the second visual token, and the language token to obtain a spliced ​​token corresponding to the trajectory planning task. When the image processing task is a question-answering task, no preset token is added, that is, only the third visual token and the language token are spliced ​​to obtain a spliced ​​token. Then, the spliced ​​token is predicted to obtain a predicted token. Finally, before executing the corresponding image processing task on the predicted token, the image processing process of the predicted token can be determined by the spliced ​​token corresponding to the predicted token.

[0118] In some embodiments, as Figure 3 As shown, the above step S105 can be implemented through steps S301 and S302:

[0119] Step S301: Determine the task type of the image processing task based on the concatenated tokens corresponding to the predicted tokens.

[0120] In the embodiment of the present application, the predicted token is obtained by predicting the spliced ​​tokens using the visual language model, so the input data corresponding to the output of the predicted token by the visual language model (i.e., the spliced ​​tokens) can be determined. The task type of the image processing task is determined by whether there is a preset token in the spliced ​​tokens.

[0121] Exemplarily, token 1 is a token with a preset token added, and token 2 is a token without a preset token added. The visual language model predicts token 1 to obtain predicted token 1, and the visual language model predicts token 2 to obtain predicted token 2. Because predicted token 1 is determined based on token 1, and token 1 is a token with a preset token added, it is possible to determine that the image processing process related to the trajectory planning task is performed on predicted token 1. Similarly, predicted token 2 is determined based on token 2, and token 2 is a token without a preset token added, it is possible to determine that the image processing process related to the question and answer task is performed on predicted token 2.

[0122] Step S302: performing image processing on the prediction token based on a network model corresponding to the task type.

[0123] In an embodiment of the present application, when the task type is a trajectory planning task, the prediction token can be input into the network model corresponding to the trajectory planning task, so as to perform the trajectory planning task on the prediction token; when the task type is a question and answer task, the prediction token can be input into the network model corresponding to the question and answer task, so as to perform the question and answer task on the prediction token.

[0124] In the embodiment of the present application, by analyzing the spliced ​​tokens, the task type can be identified and the corresponding network model can be called for processing. In this way, dynamic judgment of the task type is achieved, and the flexibility and applicability of image processing are enhanced.

[0125] In some embodiments, the image processing task is a trajectory planning task; the network model corresponding to the trajectory planning task includes a diffusion motion model; the above step S302 can be implemented by step S11, and the above method can also be implemented by step S12:

[0126] Step S11 : using the diffusion motion model, based on the prediction token and a first trajectory set, determining at least two predicted trajectories and scores corresponding to the at least two predicted trajectories respectively; the first trajectory set is obtained by clustering multiple sample trajectories in the sample trajectory set.

[0127] Step S12: Determine the predicted trajectory with the highest score among the at least two predicted trajectories as the target trajectory.

[0128] Here, the diffusion action model is a deep learning model for trajectory generation. The first trajectory set is used to provide an initial rough trajectory for the diffusion action model, so that the diffusion action model can output at least two predicted trajectories and scores corresponding to at least two predicted trajectories based on the initial rough trajectory. Among them, the score refers to the confidence or relevance evaluation value obtained for each predicted trajectory. The score reflects the degree of match between the predicted trajectory and the current task requirements (i.e., the prediction token). A high score means that the predicted trajectory is more in line with the optimal solution in the current situation and has a higher priority.

[0129] Exemplarily, the diffusion action model may be a diffusion transformer.

[0130] In an embodiment of the present application, the diffusion action model can be trained jointly using a sample image sequence, a sample language instruction, and a sample trajectory set. The sample trajectory set contains a true sample trajectory corresponding to the sample image sequence and the sample language instruction, and the label score of the true sample trajectory is 1, while the label scores of the other sample trajectories in the sample trajectory set are 0. Thus, through the above training, the diffusion action model can output the score of each predicted trajectory based on the prediction token and the first trajectory set. That is, during both the training phase and the inference phase, the sample trajectory set is used to output the corresponding trajectory.

[0131] It is understood that when generating trajectories, diffusion models in related art start with initial noise and generate trajectories corresponding to the model input data (e.g., various sensory data) through a continuous diffusion process. In other words, diffusion models in related art do not generate trajectories based on a preset set of trajectories. This increases the number of model iterations and reduces the efficiency of the model's output trajectory.

[0132] In some embodiments, the above method may also be implemented through step S21, and correspondingly, the above step S3021 may be implemented through step S22:

[0133] Step S21 : adding noise data to each trajectory in the first trajectory set to obtain a second trajectory set.

[0134] In the embodiment of the present application, adding noise data to the trajectory refers to introducing random perturbations at each trajectory point to generate a new trajectory that is similar to the original trajectory but has slight deviations. This can increase the diversity of the trajectory and enhance the robustness of the model. For example, the noise data can be Gaussian noise.

[0135] In an embodiment of the present application, noise data may be added to each track point in each track. In some embodiments, noise data may also be added to some track points in each track. For example, the noise data may be Gaussian noise.

[0136] In some embodiments, the complexity of the robot's environment can be determined based on a sequence of images from the robot's perspective. When the robot's environment is highly complex, noise data with a larger amplitude can be added to the trajectory; when the robot's environment is less complex, noise data with a smaller amplitude can be added to the trajectory. For example, when the robot's environment has many obstacles, the robot's environment is determined to be highly complex; when the robot's environment has few obstacles, the robot's environment is determined to be less complex.

[0137] Step S22: Process the prediction token and the second trajectory set using a diffusion motion model to obtain at least two predicted trajectories and scores corresponding to the at least two predicted trajectories.

[0138] In the embodiment of the present application, the diffusion motion model can perform denoising on each trajectory in the second trajectory set to obtain a denoised trajectory, namely the predicted trajectory.

[0139] It is understandable that, because the denoising process of the diffusion motion model changes the original trajectories in the second trajectory set (ie, the trajectories in the first trajectory set), the predicted trajectories ultimately output by the diffusion motion model are not the trajectories in the first trajectory set.

[0140] In some embodiments, the above step S22 can be implemented by formula (2):

[0141]

[0142] Among them, A θ (·) represents the diffusion action model, is the second trajectory set, For the prediction token, is the score of at least two predicted trajectories, is at least two predicted trajectories.

[0143] The present application also provides a model training method, such as Figure 4 As shown, the model training method can be implemented through steps S401 to S406:

[0144] Step S401 : obtaining at least two sample image sequences and sample language instructions corresponding to the at least two sample image sequences; different sample language instructions correspond to different image processing tasks.

[0145] In the embodiment of the present application, different sample image sequences correspond to image processing tasks of different task types. In other words, the image processing model is trained using sample image sequences and sample language instructions corresponding to image processing tasks of different task types.

[0146] In an embodiment of the present application, when the image processing task is a trajectory planning task, the sample image sequence includes image sequences of different people in various scenes. The different people have different genders and clothing. At the same time, sample language instructions corresponding to the sample image sequences can be generated.

[0147] In practical applications, different avatars can be generated using a human body model with random body shapes and randomly sampled texture maps. A visual language model can then be used to obtain textual descriptions of the avatars. The avatars can then be assigned a series of target points that must be reached in sequence, and their walking / non-walking states can be set, with walking speeds randomly sampled from the natural human walking speed range. Finally, a sequence of sample images can be obtained as the avatars move.

[0148] In some embodiments, at least two sample image sequences may correspond to image processing tasks of the same task type. That is, the same sample image sequence may be used for training image processing tasks of different task types.

[0149] In practical applications, large-scale person-text datasets can be used to construct sample image sequences corresponding to question-answering tasks. Image processing models are required to identify or describe individuals in videos containing random combinations of human subjects and background scenes. Each sample is generated from at least one randomly selected image of a person and different backgrounds, accompanied by a text description detailing each person's attributes, relative spatial position, and whether they represent the same identity. In addition to human recognition samples, open-world image samples can also be provided. These samples enhance the image processing model's ability to recognize open-world objects.

[0150] In the embodiment of the present application, at least two sample language instructions corresponding to the task type of the image processing task may be generated for at least two sample image sequences.

[0151] For example, when the object in the sample image sequence is a man wearing a green shirt and black pants, for the trajectory planning task, the sample language instruction may be "follow the man wearing a green shirt and black pants", and for the question-answering task, the sample language instruction may be "the color of the man's clothes in the video" or "can you describe the person you see".

[0152] Step S402: Using an image processing model, determine sample visual tokens corresponding to the at least two sample image sequences, and sample language tokens corresponding to the at least two sample language instructions.

[0153] In an embodiment of the present application, the image processing model includes a visual encoder and a cross-modal projector. At least two sample images in a sample image sequence are input into the visual encoder to obtain visual features of the sample images, and then the visual features are input into the cross-modal projector to obtain sample visual tokens corresponding to the at least two sample image sequences.

[0154] In the embodiment of the present application, the image processing model further includes a word segmenter, which inputs at least two sample language instructions into the word segmenter to obtain sample language tokens.

[0155] Step S403 : Based on the task type of the image processing task, splicing is performed on at least two of the sample language tokens and at least two of the sample visual tokens to obtain at least two spliced ​​sample tokens.

[0156] In the embodiment of the present application, the token splicing process of the model training process differs from the token splicing process of the model inference stage in that: when the task type represents a trajectory planning task, the preset token does not participate in the token splicing process. Specifically, when the task type represents a trajectory planning task, the token splicing process of the model training process includes: splicing the sample visual tokens of the sample historical images in the sample image sequence, the sample visual tokens of the sample current image in the sample image sequence, and the sample language tokens.

[0157] In an embodiment of the present application, when the task type represents a question-answering task, the token splicing process of the model training process is the same as the token splicing process of the model inference stage.

[0158] Step S404: using an image processing model, perform prediction processing on the at least two spliced ​​sample tokens to obtain at least two predicted sample tokens; and perform corresponding image processing tasks on the at least two predicted sample tokens to obtain at least two image processing results.

[0159] In an embodiment of the present application, the image processing model includes a visual language model, and the visual language model can be used to perform prediction processing on the two spliced ​​sample tokens to obtain at least two predicted sample tokens.

[0160] In an embodiment of the present application, the image processing model includes a diffusion action model and a large language model. After determining at least two predicted sample tokens, the predicted sample token corresponding to the trajectory planning task can be input into the diffusion action model to perform the trajectory planning task to obtain a predicted sample trajectory; the predicted sample token corresponding to the question and answer task can be input into the large language model to perform the question and answer task to obtain answer information corresponding to the sample language instruction.

[0161] Step S405: Determine the losses corresponding to different task types based on the at least two image processing results.

[0162] In an embodiment of the present application, when the image processing task is a trajectory planning task, the image processing result can be a predicted sample trajectory and a score of the predicted sample trajectory. The loss corresponding to the trajectory planning task can be determined by comparing the predicted sample trajectory with the true sample trajectory, and / or the loss corresponding to the trajectory planning task can be determined by comparing the score of the predicted sample trajectory with the label score of the sample trajectory.

[0163] In an embodiment of the present application, when the image processing task is a question-answering task, the image processing result may be answer information, and the loss corresponding to the question-answering task may be determined by comparing the answer information with the description information corresponding to the sample image sequence.

[0164] Step S406: Adjust the network parameters of the image processing model based on the losses corresponding to the different task types.

[0165] In an embodiment of the present application, image processing tasks of different task types are processed using the same model. For example, in step S402, when performing different image processing tasks, the corresponding sample visual tokens are determined through the visual encoder and the cross-modal projector; in step S404, when performing different image processing tasks, the corresponding predicted sample tokens are determined through the visual language model; therefore, the network parameters of the visual encoder, cross-modal projector and visual language model can be jointly adjusted through the losses corresponding to different task types.

[0166] For the models corresponding to different image processing tasks, for example, when the image processing task is a trajectory planning task, the diffusion action model outputs the predicted sample trajectory and the score of the predicted sample trajectory. Therefore, the network parameters of the diffusion action model can be adjusted by the loss corresponding to the trajectory planning task; when the image processing task is a question-answering task, the answer information is output through the large language model. Therefore, the network parameters of the large language model can be adjusted by the loss corresponding to the question-answering task.

[0167] In some embodiments, in order to improve the efficiency of model training, the network parameters of the visual encoder can be frozen, that is, the network parameters of the visual encoder are not adjusted.

[0168] In the embodiment of the present application, the network parameters of the image processing model are adjusted jointly by the losses corresponding to different task types. In this way, the image processing model can integrate image processing tasks of different task types into a unified training framework, thereby enhancing the coupling between different models within the image processing model.

[0169] In some embodiments, as Figure 5 As shown, the above method can also be implemented through steps S501 to S503, and the above step S405 can be implemented through step S504:

[0170] Step S501: Obtain a sample trajectory set.

[0171] In an embodiment of the present application, the trajectories of different virtual characters in various scenarios can be generated through a pre-built embodied visual tracking simulator.

[0172] Step S502 : performing clustering processing on multiple sample trajectories in the sample trajectory set to obtain a first trajectory set.

[0173] In an embodiment of the present application, a preset clustering algorithm may be used to cluster multiple sample trajectories in the sample trajectory set to obtain a first trajectory set. Exemplarily, the preset clustering algorithm may include at least one of the following: a K-means algorithm, a K-medoids algorithm, and a DBSCAN algorithm.

[0174] It is understandable that, through clustering, representative sample trajectories in the sample trajectory set can be determined from the multiple sample trajectories, thereby forming the first trajectory set. This can improve the efficiency of model training.

[0175] In some embodiments, the first trajectory set may be reused in the model inference stage, so as to determine the predicted trajectory through the first trajectory set.

[0176] Step S503: adding noise data to each trajectory in the first trajectory set to obtain a third trajectory set.

[0177] In the embodiments of the present application, the noise addition processing method in the model training process can refer to the noise addition processing method in the model inference stage.

[0178] Step S504 : determining at least two predicted sample trajectories and scores corresponding to the at least two predicted sample trajectories based on the predicted sample token corresponding to the trajectory planning task and the third trajectory set.

[0179] In the embodiment of the present application, the processing method of the predicted sample tokens and the third trajectory set during the model training process can refer to the processing method of the predicted tokens and the second trajectory set in the model reasoning stage.

[0180] In some embodiments, the third trajectory set includes a first sample trajectory and a second sample trajectory; the first sample trajectory is annotated with a first label score, and the second sample trajectory is annotated with a second label score; the above step S405 is implemented by steps S31 to S35:

[0181] Step S31 : determining the similarity between the at least two predicted sample trajectories and the first sample trajectory.

[0182] Step S32 : determining the predicted sample trajectory with the greatest similarity among the at least two predicted sample trajectories as the target predicted sample trajectory.

[0183] Here, the first sample trajectory can be a pre-set true trajectory in the third trajectory set, or a trajectory close to the true trajectory. The first sample trajectory can be used as a positive sample, and the second sample trajectory in the third trajectory set other than the first sample trajectory can be used as a negative sample. The first label score of the first sample trajectory is greater than the second label score of the second sample trajectory. For example, the first label score can be set to 1, and the second label score can be set to 0.

[0184] In the embodiment of the present application, by determining the similarity between at least two predicted sample trajectories and the first sample trajectory, the gap between each predicted sample trajectory and the positive sample can be obtained, thereby obtaining the predicted sample trajectory closest to the positive sample.

[0185] For example, the similarity may include at least one of the following: Euclidean distance, Hausdorff distance, Fréchet distance, and trajectory edit distance (TED).

[0186] Step S33: determining a first trajectory loss based on the target predicted sample trajectory and the first sample trajectory.

[0187] Here, the first trajectory loss may be a mean squared error loss (MSE).

[0188] In some embodiments, the first trajectory loss can be implemented by formula (3):

[0189]

[0190] in, is at least two predicted sample trajectories, s i is the fraction of sample trajectories in the third trajectory set, τ gt Since the second label score of the second sample trajectory in the third trajectory set is 0, it is only necessary to determine the trajectory loss corresponding to the target prediction sample trajectory in at least two prediction sample trajectories.

[0191] Step S34: determining a second trajectory loss based on the scores corresponding to the at least two predicted sample trajectories, the first label score, and the second label score.

[0192] Here, the second trajectory loss may be a binary cross-entropy loss (BCE).

[0193] In some embodiments, the second trajectory loss can be implemented by formula (4):

[0194]

[0195] in, is the score corresponding to at least two predicted sample trajectories, s i is the first label score and the second label score.

[0196] Step S35 : determining a third trajectory loss based on the first trajectory loss and the second trajectory loss.

[0197] In the embodiment of the present application, the sum of the first trajectory loss and the second trajectory loss may be determined as the third trajectory loss.

[0198] In some embodiments, a weighted sum of the first trace loss and the second trace loss may be performed to obtain a third trace loss.

[0199] In some embodiments, the third trace loss can be implemented by formula (5):

[0200]

[0201] Among them, L track is the third trajectory loss, and λ is the balancing parameter.

[0202] In some embodiments, the above step S406 is implemented through steps S41 to S43:

[0203] Step S41: Based on the third trajectory loss, adjust the network parameters of the diffusion motion model used for the trajectory planning task in the image processing model.

[0204] Step S42: adjusting the network parameters of the language model used for the question-answering task in the image processing model based on the text loss.

[0205] Step S43: Based on the third trajectory loss and the text loss, adjust the network parameters of the token generation model in the image processing model for generating visual tokens, and adjust the network parameters of the visual language model in the image processing model for obtaining predicted sample tokens.

[0206] Here, the token generation model can be the above-mentioned cross-modal projector.

[0207] In an embodiment of the present application, when adjusting the corresponding model based on the corresponding loss, the network parameters of each model can be adjusted when the corresponding loss meets the preset range or the number of model iterations is equal to the preset number.

[0208] In the embodiment of this application, for different task modules, the corresponding network parameters are optimized according to the corresponding losses. In this way, on the one hand, by independently optimizing the parameters of each sub-module, the model can maintain good performance across different tasks; on the other hand, by jointly optimizing the shared parameters across tasks, the overall coordination ability of the model is enhanced, improving the comprehensive performance of multi-task processing.

[0209] The following describes the application of the image processing and model training methods provided in the embodiments of the present application in actual scenarios:

[0210] Embodied visual tracking is a fundamental capability of embodied AI, enabling intelligent agents to track specific targets in dynamic environments using only first-person perspective (egocentric vision). This task is inherently challenging because it relies on two tightly coupled skills: (1) target recognition, which is the ability to accurately identify and distinguish targets; and (2) trajectory planning, which is the ability to determine the optimal action for effective tracking. The coordination between recognition and planning is particularly demanding in challenging conditions such as severe occlusion or highly dynamic scenes.

[0211] Related technologies typically decouple recognition and trajectory planning into detection and planning models, respectively. However, these methods are limited to category-level tracking in relatively open areas. This is because their loosely coupled design leads to error accumulation between the recognition and planning models. For example, incorrect recognition can lead to planning errors.

[0212] In order to achieve synergy between target recognition and trajectory planning, an embodiment of the present application provides a TrackVLA (Vision-Language-Action) model, which is characterized by a unified framework that integrates target recognition and trajectory planning. Specifically, the two tasks use the same token encoding and large language model (LLM) forward propagation mechanism to predict the next token, while the decoding depends on the specific task. For recognition tasks, TrackVLA uses a language modeling head (i.e., the large language model in the above embodiment) to decode text responses. For planning tasks, TrackVLA uses an anchor-based diffusion head (i.e., the diffusion action model in the above embodiment) to generate waypoint trajectories. The two tasks are jointly trained to optimize TrackVLA to achieve tight coupling between recognition and planning.

[0213] like Figure 6As shown, TrackVLA 601 (the image processing model in the above embodiment) includes a pretrained video-based visual language model (Pretrained Video-based VLM) 6011 (the visual language model in the above embodiment), an action model 6012 (the diffuse action model in the above embodiment), and a language model 6013 (the large language model in the above embodiment). TrackVLA 601 can be trained using embodied visual tracking samples 602 and open-world recognition samples 603. The pretrained video-based visual language model 6011 can input its output data into the action model 6012 and the language model 6013 to perform corresponding tasks. The action model 6012 can output a trajectory 60121, and the language model 6013 can output a description 60131, i.e., "The girl you observed is wearing a red top and black pants."

[0214] In such Figure 6 In the left part of the figure, when TrackVLA 601 obtains the natural language instruction "Follow the man in the black shirt and black pants", it can track the target based on the image sequence from the robot's perspective (i.e., image ① to image ④), thereby completing robust tracking.

[0215] In such Figure 6 In the middle part, when TrackVLA 601 receives the natural language instruction "Follow the man in the dark gray T-shirt and black pants", it can track the target based on the image sequence from the robot's perspective (i.e., image ⑤ to image ⑧), thereby completing long-term tracking.

[0216] In such Figure 6 In the right part of the image, TrackVLA 601 is able to achieve cross-domain generalization, meaning it can complete tracking in different scenarios. For example, it can complete tracking targets when receiving natural language instructions such as "walk behind the robot dog," "chase the man skiing," or "chase the Lego figure."

[0217] In the embodiment of the present application, at each timestamp T, a natural language instruction I (the language instruction information in the above embodiment) describing the appearance of a specific target is given, as well as a first-person RGB observation value O consisting of a series of frames. T ={x1,…,x T}(image sequence from the robot's perspective in the above embodiment), x TRefers to the image frame captured by the robot at timestamp T. The agent (the aforementioned TrackVLA) must output the next action aT∈A={v,ω} to continuously track the target in an unknown environment. The action space A contains the agent's linear velocity v and angular velocity ω. The mission is considered successful if the agent can maintain a suitable following distance (1 to 3 meters) from the target while facing the target.

[0218] like Figure 7 As shown, TrackVLA includes a visual encoder 703, a cross-modal projector 704, a video-based pre-trained visual language model 6011, a diffusion transformer 705 (the diffusion action model in the above embodiment), and a language modeling head 706 (the large language model in the above embodiment). The processing tasks of TrackVLA include embodied visual tracking 701 and visual question answering 702.

[0219] For embodied visual tracking 701, the input data of the visual encoder 703 includes historical observations 7011 (historical images in the image sequence in the above embodiment) and current observations 7012 (the current image in the image sequence in the above embodiment). The first-person RGB observation value O T ={x1,…,x T x in} T For the current observation 7012, {x1,…,x T-1} is a historical observation 7011. The visual encoder 703 can extract visual features Where N is the number of patches (set to 256), C represents the embedding dimension, V1: T means {x1,…,x T} respectively correspond to the visual features.

[0220] In the embodiment of the present application, the above-mentioned visual features have two resolution scales, which can be expressed by formula (6):

[0221]

[0222] in, Provides fine-grained observations, while Providing coarse-grained observations. To achieve the best balance between token length and performance, fine-grained features are used for the current observation 7012 to enhance object recognition, while coarse-grained features are used for historical observations 7011 and input data for visual question answering 702.

[0223] In this embodiment, to ensure the consistency of inference speed during the tracking process, a sliding window mechanism is used to retain only the latest k frames (k is 32). For embodied visual tracking, the visual features corresponding to each frame are sorted by time to obtain a visual feature sequence. The visual feature sequences corresponding to the historical observation 7011 and the current observation 7012 are constructed as follows: in, is the visual feature corresponding to historical observation 7011 (i.e., coarse-grained feature), is the visual feature corresponding to the current observation 7012 (ie, the fine-grained feature).

[0224] like Figure 7 As shown, the visual encoder 703 outputs visual features, which are input to the cross-modal projector 704. The cross-modal projector 704 can be P(·)(a 2-layer MLP), which is used to project the visual features into the latent space of the large language model (i.e., the video-based pre-trained visual language model). The processing of the cross-modal projector can be expressed by formula (7):

[0225]

[0226] in, is a visual token, V T For visual features.

[0227] like Figure 7 As shown, cross-modal projector 704 outputs visual tokens 7041 corresponding to historical observation 7011 and visual tokens 7042 corresponding to current observation 7012. Visual tokens 7041 include 4×k tokens, and visual tokens 7042 include 1×64 tokens. Because the visual features corresponding to current observation 7012 are fine-grained features, while the visual features corresponding to historical observation 7011 are coarse-grained features, the number of visual tokens 7042 is greater than the number of visual tokens 7041.

[0228] In the embodiment of the present application, it is necessary to splice the visual tokens and the language tokens and input them into the video-based pre-trained visual language model. Figure 7 As shown, the visual token 7041 corresponding to the historical observation 7011, the visual token 7042 corresponding to the current observation 7012, the [Track] token 7044 and the language token 7045 are spliced ​​together, and the spliced ​​tokens are input into the video-based pre-trained visual language model 6011 to obtain the predicted token.

[0229] Because the [Track] token exists in the concatenated tokens corresponding to the predicted token, the predicted token is input to the diffusion transformer 705 and the track 7051 is output, thereby completing the embodied visual tracking 701.

[0230] For visual question answering 702 , the input data for visual encoder 703 includes video stream 7021 . During the inference phase, video stream 7021 can be the same as historical observation 7011 and current observation 7012 , or it can be a frame of image from historical observation 7011 and / or current observation 7012 . Visual encoder 703 extracts visual features from video stream 7021 , where these visual features are coarse-grained features.

[0231] In the embodiment of the present application, the visual feature sequence corresponding to the video stream 7021 is constructed as follows: The visual features corresponding to video stream 7021 are input to cross-modal projector 704 to obtain visual tokens 7043. Visual tokens 7043 include 4×T tokens. Because the visual features corresponding to video stream 7021 are coarse-grained features, the number of visual tokens 7043 is smaller than the number of visual tokens 7042.

[0232] In the embodiment of the present application, it is necessary to splice the visual tokens and the language tokens and input them into the video-based pre-trained visual language model. Figure 7 As shown, the visual token 7043 and the language token 7045 corresponding to the video stream 7021 are spliced ​​together, and the spliced ​​tokens are input into the video-based pre-trained visual language model 6011 to obtain the predicted token.

[0233] Because the [Track] token does not exist in the concatenated token corresponding to the predicted token, the predicted token is input into the language modeling head 706. The language modeling head 706 decodes the predicted token into words in the vocabulary in an autoregressive manner to obtain the text description "The man on the left is wearing XX", thereby completing the recognition visual question answering 702.

[0234] In this embodiment, the diffusion action model (i.e., diffusion transformer) is an anchor-based diffusion model that performs denoising operations starting from predefined anchor points to generate navigation waypoints. The predefined anchor points provide an initial rough trajectory, significantly reducing the number of denoising iterations required and increasing inference speed by up to 5 times compared to traditional vanilla diffusion policies.

[0235] In an embodiment of the present application, the training method of the diffusion motion model includes the following steps:

[0236] Step S801: Collect all trajectories from the training data and obtain a set of trajectories by using the K-means clustering algorithm.

[0237] Where M represents the number of trajectory points. Each trajectory after clustering Represents a representative trajectory pattern of a robot. ω is the number of waypoints in each trajectory.

[0238] Step S802: perturb each trajectory point with Gaussian noise to create a noisy trajectory set.

[0239] Step S803: noisy trajectory set and prediction tokens As a diffusion action model A θ (.) Input data and output: denoised trajectory and the corresponding trajectory classification scores

[0240] In the embodiment of the present application, the processing process of the diffusion action model can be expressed by formula (8):

[0241]

[0242] like Figure 8 As shown, the predicted token 801 and the noisy trajectory set 802 are projected into the space of the diffusion action model 705 through a multilayer perceptron (MLP). Then the diffusion action model 705 outputs parameters, and the multilayer perceptron projects the parameters output by the diffusion action model 705 into the original space to obtain the trajectory classification score 803 and the denoised trajectory 804.

[0243] Step S804: adjusting network parameters of the diffusion motion model based on the loss.

[0244] In the embodiment of the present application, for each sample, the closest true trajectory τ gt The anchor point trajectory is marked as a positive sample (S nearest =1), and all other anchor point trajectories are marked as negative samples (S else = 0). Then the tracking loss is determined by jointly optimizing the trajectory regression loss and the score prediction loss. The tracking loss can be expressed by formula (9):

[0245]

[0246] Among them, L track To track losses, is the score prediction loss, is the trajectory regression loss, and λ is the balancing parameter.

[0247] In the embodiment of the present application, the diffusion transformer (DiT) is used for denoising, and the anchor-based diffusion strategy only requires two steps of denoising. The overall training loss L of TrackVLA is defined as the tracking loss L track With the text prediction loss L text The weighted combination of L = L track +αL text , where α is the equilibrium parameter.

[0248] In an embodiment of the present application, the training data of TrackVLA includes embodied visual tracking data and video-based question-answering data.

[0249] In an embodiment of the present application, embodied visual tracking data is used to generate humanoid avatars and natural human behaviors through an embodied visual tracking simulator. The humanoid avatar can be initialized using a human body model with a random body shape and randomly sampled texture maps. A textual description of the avatar is then obtained using a visual language model. Each avatar is then assigned a series of target points that must be reached in sequence and is assigned a walking / non-walking state. The walking speed is randomly sampled from the natural human walking speed range of [1.0 m / s-1.5 m / s].

[0250] In an embodiment of the present application, TrackVLA can also be subjected to an embodied visual tracking benchmark test, that is, the embodied visual tracking capability of TrackVLA can be comprehensively evaluated through the embodied visual tracking benchmark test. The visual tracking benchmark test includes a training set and a test set, each of which includes multiple different scenarios, each of which corresponds to a different humanoid avatar and scene environment, and each avatar has a corresponding description. In order to comprehensively evaluate the performance of TrackVLA in different scenarios, it can be divided into three subtask categories according to increasing difficulty, and each subtask contains multiple training scenarios and multiple test scenarios.

[0251] like Figure 9 As shown, the training set and the test set include multiple humanoid avatars 901, each of which has a description 902, such as "this man is wearing a blue shirt and brown pants", "this woman is wearing a red leather jacket and black pants", and "this boy is wearing a light gray suit jacket and a white shirt". Then, based on the training set and the test set, the embodied visual tracking benchmark is tested on three task categories, including single target tracking 903, interference tracking 904, and blur tracking 905:

[0252] For Single Target Tracking (STT) 903, the basic following capability of TrackVLA is evaluated through simple instructions such as “Follow this person / man / woman”.

[0253] For Distracted Tracking (DT) 904, TrackVLA’s target recognition capability is evaluated by giving a fine-grained description of the target (e.g., “follow the light-skinned man in a black suit and a white belt”).

[0254] For Ambiguous Tracking (AT) 905, TrackVLA’s ability to identify the correct target is evaluated by giving intentionally ambiguous instructions (e.g., “Follow the first person you see”) in the presence of identical-looking distractors.

[0255] In the embodiment of the present application, the video-based question-answering data includes human recognition samples and open-world samples. For human recognition samples, a large-scale person-text dataset is used to construct human recognition samples, requiring TrackVLA to identify or describe individuals in videos containing randomly combined human subjects and background scenes. Each sample is generated by placing 1 to 3 randomly selected human images on different backgrounds, and is accompanied by a text description detailing each person's attributes, relative spatial position, and whether they represent the same identity. Open-world samples are used to enhance TrackVLA's ability to identify open-world targets.

[0256] As shown in 10, the video-based question-answering data includes a person re-identification dataset 1002, each sample has a random background 1001, and each sample has a corresponding question-answer description, for example: "Q: Can you describe the person you see? A: I see a lady, she is wearing a black top and dark pants, and she is carrying a black bag", "Q: Can you describe all the people in the video? A: The man on the left is wearing a black top and gray pants, and the man on the right is wearing a pink shirt and white pants", "Q: Can you tell whether all the individuals in the video are the same person? A: Yes, the two people in the video are actually the same person, wearing a black top and blue jeans".

[0257] In an embodiment of the present application, when training the model, the ratio of embodied visual tracking data to video-based question-answering data can be 1:1.

[0258] In some embodiments, TrackVLA follows a two-stage training process:

[0259] In the first stage, the projector of the visual encoder is trained using a large amount of image-text data to align the visual embedding space with the latent space of a large language model.

[0260] In the second stage, the visual projector, a large language model, and an action model are jointly trained using the mixed training data. During training, the diffusion plan of the action model is truncated to a maximum of 50 steps, for a total of 1000 steps, to diffuse the trajectory anchors, which only introduces a small amount of noise.

[0261] In some embodiments, during the inference process of TrackVLA, each input frame is resized to 224×224 and fed into the visual encoder. Figure 2 After extracting the tokens output by the projector on the left, we organize these tokens according to the task type.

[0262] For the embodied visual tracking task, a special [Track] tag is added before the instruction token, and only a single-step autoregression is performed with the large language model. The last layer hidden state of the LLM output is then passed to the action model. 10 steps of diffusion are applied to the trajectory anchors in 1000 steps, and 2 steps of denoising are performed using DDIM to generate a set of predicted trajectories and corresponding score vectors. The trajectory corresponding to the anchor with the highest score is selected as the final output. For the visual question answering (VQA) task, following the standard autoregressive decoding process of the LLM, the language modeling head re-tokenizes the predicted tokens into text answers.

[0263] In the embodiments of the present application, the architecture of the above-mentioned action model includes multilayer perceptrons (MLPs) with 3 and 6 layers, respectively, and diffusion transformers of different sizes. The hidden state dimensions of the two MLPs are set to 1024 and 4096, respectively. The basic diffusion transformer is configured with a depth of 12, a hidden state size of 768, and 12 attention heads, while the small diffusion transformer uses a depth of 6, a hidden state size of 384, and 4 attention heads.

[0264] like Figure 11 As shown, TrackVLA 601 is deployed in the cloud. The robot collects images from its own perspective, forming observation results 1101. Observation results 1101 are then sent to the cloud via the robot's communication module. TrackVLA 601 deployed in the cloud generates a trajectory 1102 based on observation results 1101. The cloud then sends trajectory 1102 to the robot. After receiving the trajectory, the robot uses a pure tracking algorithm, combined with its posture information, to perform closed-loop control of its linear and angular velocity, enabling it to accurately follow the trajectory. The robot also utilizes lidar point cloud data and implements an elastic band algorithm for obstacle avoidance.

[0265] Figure 12 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application is shown in FIG. Figure 12 As shown, the image processing device 1200 includes: a first acquisition unit 1201, a first determination unit 1202, a first splicing unit 1203, a prediction unit 1204 and a first processing unit 1205, wherein:

[0266] The first acquisition unit 1201 is used to acquire an image sequence from the robot's perspective and language instruction information; the language instruction information is used to represent the task type of the image processing task;

[0267] A first determining unit 1202 is configured to determine a visual token corresponding to at least one frame of the image sequence, and a language token corresponding to the language instruction information;

[0268] A first splicing unit 1203 is configured to splice the language token and the visual token based on the task type of the image processing task to obtain a spliced ​​token;

[0269] The prediction unit 1204 is used to perform prediction processing on the concatenated tokens to obtain predicted tokens;

[0270] The first processing unit 1205 is configured to perform the image processing task on the prediction token.

[0271] In some embodiments, the first determination unit 1202 is further used to determine the first visual token of the historical image in the image sequence and the second visual token of the current image in the image sequence when the task type represents a trajectory planning task; and to determine the third visual token corresponding to at least one frame of image in the image sequence when the task type represents a question-answering task; wherein the resolution of the second visual token is greater than the resolution of the first visual token and the resolution of the third visual token.

[0272] In some embodiments, the first splicing unit 1203 is also used to splice the first visual token, the second visual token, the language token and the preset token when the task type represents a trajectory planning task; the preset token is used to represent that the image processing task includes a trajectory planning task; and when the task type represents a question-answering task, the third visual token and the language token are spliced.

[0273] In some embodiments, the first processing unit 1205 is used to determine the task type of the image processing task based on the spliced ​​token corresponding to the prediction token; and perform image processing on the prediction token based on the network model corresponding to the task type.

[0274] In some embodiments, the image processing task is a trajectory planning task; the network model corresponding to the trajectory planning task includes a diffusion motion model; in some embodiments, the first processing unit 1205 is used to use the diffusion motion model to determine at least two predicted trajectories and the scores corresponding to the at least two predicted trajectories based on the prediction token and the first trajectory set; the first trajectory set is obtained by clustering multiple sample trajectories in the sample trajectory set; the image processing unit also includes a trajectory determination unit, which is further used to determine the predicted trajectory with the highest score among the at least two predicted trajectories as the target trajectory.

[0275] In some embodiments, the image processing unit further includes a first noise adding unit, which is used to add noise data to each trajectory in the first trajectory set to obtain a second trajectory set; a first processing unit 1205 is used to use a diffusion motion model to process the prediction token and the second trajectory set to obtain at least two predicted trajectories and scores corresponding to the at least two predicted trajectories.

[0276] Figure 13 A schematic diagram of the structure of a model training device provided in an embodiment of the present application is shown in FIG. Figure 13 As shown, the model training device 1300 includes: a second acquisition unit 1301, a second determination unit 1302, a second splicing unit 1303, a second processing unit 1304, a loss determination unit 1305 and an adjustment unit 1306, wherein:

[0277] The second acquisition unit 1301 is configured to acquire at least two sample image sequences and sample language instructions corresponding to the at least two sample image sequences; different sample language instructions correspond to different image processing tasks;

[0278] A second determining unit 1302 is configured to determine, by using an image processing model, sample visual tokens corresponding to the at least two sample image sequences, and sample language tokens corresponding to the at least two sample language instructions;

[0279] A second splicing unit 1303 is configured to splice at least two of the sample language tokens and at least two of the sample visual tokens based on the task type of the image processing task to obtain at least two spliced ​​sample tokens;

[0280] The second processing unit 1304 is configured to perform prediction processing on the at least two spliced ​​sample tokens using an image processing model to obtain at least two predicted sample tokens; and perform corresponding image processing tasks on the at least two predicted sample tokens to obtain at least two image processing results.

[0281] A loss determining unit 1305 is configured to determine losses corresponding to different task types based on the at least two image processing results;

[0282] The adjustment unit 1306 is used to adjust the network parameters of the image processing model based on the losses corresponding to the different task types.

[0283] In some embodiments, the model training device further includes a trajectory processing unit, which is used to obtain a sample trajectory set; perform clustering processing on multiple sample trajectories in the sample trajectory set to obtain a first trajectory set; add noise data to each trajectory in the first trajectory set to obtain a third trajectory set; and a second processing unit 1304 is used to determine at least two predicted sample trajectories and scores corresponding to the at least two predicted sample trajectories based on the predicted sample token corresponding to the trajectory planning task and the third trajectory set.

[0284] In some embodiments, the third trajectory set includes a first sample trajectory and a second sample trajectory; the first sample trajectory is annotated with a first label score, and the second sample trajectory is annotated with a second label score; the loss determination unit 1305 is used to determine the similarity between the at least two predicted sample trajectories and the first sample trajectory respectively; the predicted sample trajectory with the greatest similarity among the at least two predicted sample trajectories is determined as the target predicted sample trajectory; based on the target predicted sample trajectory and the first sample trajectory, a first trajectory loss is determined; based on the scores corresponding to the at least two predicted sample trajectories, the first label score and the second label score, a second trajectory loss is determined; based on the first trajectory loss and the second trajectory loss, a third trajectory loss is determined.

[0285] In some embodiments, the adjustment unit 1306 is used to adjust the network parameters of the diffusion action model in the image processing model for performing the trajectory planning task based on the third trajectory loss; adjust the network parameters of the large language model in the image processing model for performing the question-answering task based on the text loss; adjust the network parameters of the token generation model in the image processing model for generating visual tokens based on the third trajectory loss and the text loss, and adjust the network parameters of the visual language model in the image processing model for obtaining predicted sample tokens.

[0286] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to perform the methods described in the above method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0287] It should be noted that, in the embodiment of the present application, if the above-mentioned image processing method and model training method are implemented in the form of a software function module and sold or used as an independent product, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.

[0288] An embodiment of the present application provides a computer device including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.

[0289] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.

[0290] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.

[0291] The present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be implemented in hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK).

[0292] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the description of the method embodiments of this application for understanding.

[0293] Figure 14 A schematic diagram of a hardware entity of a computer device in an embodiment of the present application is shown in FIG. Figure 14 As shown, the hardware entity of the computer device 1400 includes: a processor 1401, a communication interface 1402 and a memory 1403, wherein:

[0294] The processor 1401 generally controls the overall operation of the computer device 1400 , and the overall operation may be to implement the image processing method and model training method provided in the embodiments of the present application.

[0295] The communication interface 1402 enables the computer device 1400 to communicate with other terminals or servers through a network.

[0296] The memory 1403 is configured to store instructions and applications executable by the processor 1401 and to cache data to be processed or processed by the processor 1401 and various modules in the computer device 1400 (e.g., image data, audio data, voice communication data, and video communication data). This can be implemented using flash memory (FLASH) or random access memory (RAM). Data can be transmitted between the processor 1401, the communication interface 1402, and the memory 1403 via a bus 1404.

[0297] An embodiment of the present application provides a computer storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the image processing method and model training method of any of the above embodiments.

[0298] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0299] The processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It is understood that the electronic device that implements the functions of the processor may also be other electronic devices, which are not specifically limited in the embodiments of the present application.

[0300] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0301] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.

[0302] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0303] The above are only implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the protection scope of the present application.

Claims

1. An image processing method, characterized in that: The method comprises: Acquire an image sequence and language instruction information from the robot's perspective; the language instruction information is used to characterize the task type of the image processing task; Determining a visual token corresponding to at least one frame of the image sequence, and a language token corresponding to the language instruction information; Based on the task type of the image processing task, the language token and the visual token are spliced ​​to obtain a spliced ​​token; Performing prediction processing on the spliced ​​tokens to obtain predicted tokens; The image processing task is performed on the predicted token.

2. The method according to claim 1, characterized in that The determining of the visual tokens corresponding to at least one frame of the image sequence includes: In a case where the task type represents a trajectory planning task, determining a first visual token of a historical image in the image sequence and a second visual token of a current image in the image sequence; In a case where the task type represents a question-answering task, determining a third visual token corresponding to each of at least one frame of the image sequence; The resolution of the second visual token is greater than the resolution of the first visual token and the resolution of the third visual token.

3. The method according to claim 2, characterized in that The splicing processing of the language token and the visual token based on the task type of the image processing task includes: In the case where the task type represents a trajectory planning task, the first visual token, the second visual token, the language token, and a preset token are concatenated; the preset token is used to represent that the image processing task includes a trajectory planning task; In a case where the task type represents a question-answering task, the third visual token and the language token are concatenated.

4. The method according to any one of claims 1 to 3, characterized in that The performing the image processing task on the prediction token includes: Determining a task type of the image processing task based on the concatenated tokens corresponding to the predicted tokens; When the task type characterizes the image processing task as a trajectory planning task, determining at least two predicted trajectories and scores corresponding to the at least two predicted trajectories based on the prediction tokens and a first trajectory set using a diffusion motion model; the first trajectory set is obtained by clustering multiple sample trajectories in the sample trajectory set; The method further comprises: The predicted trajectory with the highest score among the at least two predicted trajectories is determined as the target trajectory.

5. The method according to claim 4, characterized in that The method further comprises: adding noise data to each trajectory in the first trajectory set to obtain a second trajectory set; The determining, using the diffusion motion model, at least two predicted trajectories and scores corresponding to the at least two predicted trajectories based on the prediction token and the first trajectory set, includes: The prediction token and the second trajectory set are processed using a diffusion motion model to obtain at least two predicted trajectories and scores corresponding to the at least two predicted trajectories.

6. A model training method, characterized in that: The method comprises: Acquire at least two sample image sequences and sample language instructions corresponding to the at least two sample image sequences respectively; different sample language instructions correspond to different image processing tasks; Determining, by using an image processing model, sample visual tokens corresponding to the at least two sample image sequences, and sample language tokens corresponding to the at least two sample language instructions; Based on the task type of the image processing task, performing splicing processing on at least two of the sample language tokens and at least two of the sample visual tokens to obtain at least two spliced ​​sample tokens; Using an image processing model, performing prediction processing on the at least two spliced ​​sample tokens to obtain at least two predicted sample tokens; performing corresponding image processing tasks on the at least two predicted sample tokens to obtain at least two image processing results; Determining losses corresponding to different task types based on the at least two image processing results; Based on the losses corresponding to the different task types, the network parameters of the image processing model are adjusted.

7. The method according to claim 6, characterized in that The method further comprises: Get a set of sample trajectories; performing clustering processing on a plurality of sample trajectories in the sample trajectory set to obtain a first trajectory set; adding noise data to each trajectory in the first trajectory set to obtain a third trajectory set; In the case where the image processing task is a trajectory planning task, performing corresponding image processing tasks on the at least two predicted sample tokens respectively to obtain at least two image processing results includes: At least two predicted sample trajectories and scores corresponding to the at least two predicted sample trajectories are determined based on the predicted sample token corresponding to the trajectory planning task and the third trajectory set.

8. The method according to claim 7, characterized in that The third trajectory set includes a first sample trajectory and a second sample trajectory; the first sample trajectory is annotated with a first label score, and the second sample trajectory is annotated with a second label score; and determining the losses corresponding to different task types based on the at least two image processing results includes: Determining similarities between the at least two predicted sample trajectories and the first sample trajectory respectively; Determining the predicted sample trajectory with the greatest similarity among the at least two predicted sample trajectories as the target predicted sample trajectory; Determining a first trajectory loss based on the target predicted sample trajectory and the first sample trajectory; Determining a second trajectory loss based on the scores corresponding to the at least two predicted sample trajectories, the first label score, and the second label score; A third trajectory loss is determined based on the first trajectory loss and the second trajectory loss.

9. An image processing device, characterized in that: The device comprises: A first acquisition unit is configured to acquire an image sequence and language instruction information from the robot's perspective; the language instruction information is used to characterize a task type of the image processing task; A first determining unit is configured to determine a visual token corresponding to at least one frame of the image sequence, and a language token corresponding to the language instruction information; A first splicing unit is configured to splice the language token and the visual token based on the task type of the image processing task to obtain a spliced ​​token; A prediction unit, configured to perform prediction processing on the concatenated tokens to obtain a predicted token; The first processing unit is configured to perform the image processing task on the prediction token.

10. A computer device, wherein: include: a memory for storing executable instructions; A processor, configured to implement the method of any one of claims 1 to 5, or implement the method of any one of claims 6 to 8, when executing the executable instructions stored in the memory.

Citation Information

Cited By

  • Model training method, model prediction method and related device

    CN121614876A

  • Model training method, model prediction method and related devices

    CN121614876B