Decision-making method, decision-making device and computing equipment
By generating multimodal instructions through semantic alignment, the shortcomings of pure language descriptions in complex tasks are addressed, the task coverage and generalization capabilities of the execution device are improved, and clearer task execution is achieved.
Patent Information
- Application Number
- CN202410612967.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies struggle to design effective and universal decision-making methods for complex and highly dynamic scenarios, resulting in insufficient task coverage and generalization capabilities of execution devices, especially the inability to clearly describe and eliminate the ambiguity of instructions caused by purely linguistic descriptions.
Semantically align task information with environmental and observation information to generate multimodal commands to control execution devices. Execute corresponding tasks through these multimodal commands, thereby improving task coverage and generalization capabilities.
It achieves a clear description of complex tasks, eliminates instruction ambiguity, and enhances the action instruction mapping capability and policy execution generalization capability of the execution device.
Smart Images

Figure CN120975122A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a decision method, a decision device and a computing device. BACKGROUND
[0002] Embodied intelligence is an intelligent system that can understand, reason and interact with the physical world, which requires embodied entities to perform tasks in the physical environment and adapt to environmental changes. Embodied intelligence requires embodied entities to make planning decisions and executable actions corresponding to tasks according to current perceived environmental information and internal information, and to adaptively adjust according to environmental changes to complete a series of complex tasks. Therefore, how to design an effective and universal decision method is a problem to be solved. SUMMARY
[0003] To solve the above problems, the embodiments of the present application provide a decision method, which can convert tasks represented in various forms into multi-modal instructions, and use the multi-modal instructions as a medium to control an execution device to execute corresponding tasks, thereby improving the task coverage and generalization ability of the execution device. In addition, the present application also provides a decision device and a computing device corresponding to the decision method.
[0004] To this end, the embodiments of the present application adopt the following technical solutions:
[0005] In a first aspect, the embodiments of the present application provide a decision method, comprising: receiving task information of a target task; the task information is represented in one or more of text, image and voice; the target task includes at least one subtask, and a subtask refers to a partial task that constitutes the target task; performing semantic alignment on the task information, environmental information and first observation information to obtain feature information of a first subtask; the environmental information refers to surrounding environment data of an execution device; the first observation information refers to a plurality of objects around the execution device at a first time and states of the plurality of objects; the at least one subtask includes the first subtask; generating at least one first multi-modal instruction based on the feature information of the first subtask and the first observation information; the at least one first multi-modal instruction is used to convert the execution device to a first state, so that the states of the plurality of objects in the observation information collected next time are target states after the first subtask is completed.
[0006] In this embodiment, after receiving task information of a target task in any form, the method can perform semantic alignment on the task information, environment information and observation information, and then convert the obtained feature information into a multi-modal instruction, which can clearly describe each sub-task in the task, solve the problem that pure language cannot describe in a complex task, and eliminate instruction ambiguity. After converting the task into a multi-modal instruction, the method can control the execution device to execute the corresponding sub-task by using the multi-modal instruction as a medium, thereby improving the task coverage and generalization ability of the execution device.
[0007] In an embodiment, the generating at least one first multi-modal instruction based on the feature information of the first sub-task and the first observation information specifically includes: generating a target state of the target object according to the feature information of the first sub-task and the first observation information; the target object refers to an object to be controlled by the first sub-task, and the plurality of objects includes the target object; generating a skill action instruction according to the target state of the target object and the first observation information; and generating the at least one first multi-modal instruction according to the target state of the target object, the skill action instruction and the first observation information.
[0008] In this embodiment, the method can generate a target state of a target object according to the feature information and observation information of a sub-task, and then generate a multi-modal instruction according to the target state of the target object and the observation information, so as to realize the generation of a serial multi-modal instruction. This process can reuse a large amount of action-free label data, such as human video data, and can reduce the difficulty of data acquisition.
[0009] In an embodiment, the generating at least one first multi-modal instruction based on the feature information of the first sub-task and the first observation information specifically includes: obtaining feature information of a skill action and feature information of a target state of a target object according to the feature information of the first sub-task; generating the target state of the target object according to the feature information of the target state of the target object and the first observation information; and generating the at least one first multi-modal instruction according to the target state of the target object, a skill action instruction and the first observation information; the skill action instruction is obtained by decoding the feature information of the skill action.
[0010] In this embodiment, the method can directly output skill action instructions and target states according to the feature information of sub-tasks, realize parallel generation of multi-modal instructions, solve the problem that pure language cannot describe in complex tasks, and eliminate instruction ambiguity. Compared with the inverse solution of skill action by using observation and target state in the serial generation of multi-modal instructions, the method simultaneously generates target states and skill action instructions in the process of parallel generation of multi-modal instructions. Therefore, the generation of skill action instructions does not need to rely on the target state, so the method can generate other control information such as the moving speed, direction, angle, etc. of skill action when generating skill action instructions, so that the skill action instructions are more detailed and comprehensive. In addition, the parallel multi-modal instructions can have a wider instruction representation capability and can cover more comprehensive embodied tasks through large-scale multi-modal instruction data training.
[0011] In an embodiment, the method further comprises: converting the at least one first multi-modal instruction into a plurality of action instructions; the action instructions are instructions executable by the execution device.
[0012] In this embodiment, the method can convert multi-modal instructions into action instructions that can be recognized by the execution device, so that the execution device can clearly understand the operation object and operation target of the action instructions, thereby effectively enhancing the mapping ability of the action instructions and the generalization ability of the strategy execution.
[0013] In an embodiment, the conversion of the at least one first multi-modal instruction into a plurality of action instructions specifically comprises: encoding the at least one first multi-modal instruction to obtain a feature vector of the at least one first multi-modal instruction; and converting the feature vector of the at least one first multi-modal instruction into the plurality of action instructions.
[0014] In an implementation, the semantic alignment of the task information, the environment information, and the first observation information obtains feature information of a first subtask, specifically including: encoding the task information, the environment information, and the first observation information respectively to obtain a feature vector of the task information, a feature vector of the environment information, and a feature vector of the first observation information; fusing the feature vector of the task information, the feature vector of the environment information, a feature vector of initial observation information, and the feature vector of the first observation information to obtain a first feature vector and a second feature vector; the first feature vector represents an aligned and fused feature vector between the task information and the initial observation information; the second feature vector represents an aligned and fused feature vector between the task information and the first observation information; and calculating the correlation between the feature vector of the task information, the feature vector of the environment information, the feature vector of the first observation information, the first feature vector, and the second feature vector to obtain the feature information of the first subtask.
[0015] In an implementation, the method further includes: receiving second observation information; the second observation information refers to observation information collected after the first subtask is completed; performing semantic alignment on the task information, the environment information, and the second observation information to obtain second subtask feature information; the at least one subtask includes the second subtask; generating at least one second multi-modal instruction based on the feature information of the second subtask and the second observation information; and the at least one second multi-modal instruction is used to convert the execution device to a second state, so that the state of the plurality of objects in the next collected observation information is the target state after the second subtask is completed.
[0016] In an implementation, before the receiving second observation information, the method further includes: receiving feedback information sent by the execution device; and the feedback information is used to obtain the second observation information in a case where it is determined that the execution device completes an action operation corresponding to the at least one first multi-modal instruction.
[0017] In a second aspect, the embodiments of the present application provide a decision device, comprising: a semantic alignment module configured to receive task information of a target task; the task information is represented in one or more of text, image, and voice; the target task comprises at least one subtask; and the semantic alignment module is configured to perform semantic alignment on the task information, environment information, and first observation information to obtain feature information of a first subtask; the environment information refers to surrounding environment data of an execution device; the first observation information refers to a plurality of objects around the execution device at a first time and states of the plurality of objects; the at least one subtask comprises the first subtask; and an instruction generation module configured to generate at least one first multi-modal instruction based on the feature information of the first subtask and the first observation information; the at least one first multi-modal instruction is used to convert the execution device to a first state, so that states of the plurality of objects in observation information collected next time are target states after the first subtask is completed.
[0018] In an embodiment, the instruction generation module is specifically configured to generate a target state of a target object according to the feature information of the first subtask and the first observation information; the target object refers to an object to be controlled by the first subtask, and the plurality of objects comprises the target object; generate a skill action instruction according to the target state of the target object and the first observation information; and generate the at least one first multi-modal instruction according to the target state of the target object, the skill action instruction, and the first observation information.
[0019] In an embodiment, the instruction generation module is specifically configured to obtain feature information of a skill action and feature information of a target state of a target object according to the feature information of the first subtask; generate the target state of the target object according to the feature information of the target state of the target object and the first observation information; and generate the at least one first multi-modal instruction according to the target state of the target object, a skill action instruction, and the first observation information; the skill action instruction is obtained by decoding the feature information of the skill action.
[0020] In an embodiment, the decision device further comprises an instruction tracking module configured to convert the at least one first multi-modal instruction into a plurality of action instructions; the action instruction refers to an instruction executable by the execution device.
[0021] In an embodiment, the instruction tracking module is specifically configured to encode the at least one first multi-modal instruction to obtain a feature vector of the at least one first multi-modal instruction; and convert the feature vector of the at least one first multi-modal instruction into the plurality of action instructions.
[0022] In an implementation, the semantic alignment module is specifically configured to encode the task information, the environment information, and the first observation information respectively to obtain a feature vector of the task information, a feature vector of the environment information, and a feature vector of the first observation information; fuse the feature vector of the task information, the feature vector of the environment information, a feature vector of initial observation information, and the feature vector of the first observation information to obtain a first feature vector and a second feature vector; the first feature vector represents an aligned and fused feature vector between the task information and the initial observation information; the second feature vector represents an aligned and fused feature vector between the task information and the first observation information; and calculate correlations between the feature vector of the task information, the feature vector of the environment information, the feature vector of the first observation information, the first feature vector, and the second feature vector to obtain feature information of the first subtask.
[0023] In an implementation, the semantic alignment module is further configured to receive second observation information; the second observation information refers to observation information collected after the first subtask is completed; perform semantic alignment on the task information, the environment information, and the second observation information to obtain second subtask feature information; the at least one subtask includes the second subtask; and the instruction generation module is further configured to generate at least one second multi-modal instruction based on the second subtask feature information and the second observation information; the at least one second multi-modal instruction is used to convert the execution device to a second state, so that a state of the plurality of objects in the observation information collected next time is a target state after the second subtask is completed.
[0024] In an implementation, the semantic alignment module is further configured to, before the second observation information is received, receive feedback information sent by the execution device; and the feedback information is used to acquire the second observation information in a case where it is determined that the execution device completes an action operation corresponding to the at least one first multi-modal instruction.
[0025] In a third aspect, an embodiment of the present application provides a computing device, including: at least one memory; and at least one processor configured to execute instructions stored in the memory to cause the computing device to perform the embodiments of the first aspect.
[0026] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium including computer program instructions, when the computer program instructions are executed by a computing device, the computing device performs the embodiments of the first aspect.
[0027] In a fifth aspect, the embodiments of the present application provide a computer program product comprising instructions, characterized in that the computer program product stores instructions, which, when executed by a computing device, cause the computing device to implement various possible implementations of the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0028] The drawings needed to be used in the description of the embodiments or prior art are briefly introduced as follows.
[0029] Figure 1 FIG. 1 is a structural schematic diagram of an intelligent control system provided in the embodiments of the present application;
[0030] Figure 2 FIG. 2 is a structural schematic diagram of a decision device provided in the embodiments of the present application;
[0031] Figure 3 FIG. 3 is a specific structural schematic diagram of a semantic alignment module provided in the embodiments of the present application;
[0032] FIG. 4(a) is a process schematic diagram of a series type generation of a multi-modal instruction provided in the embodiments of the present application;
[0033] FIG. 4(b) is a process schematic diagram of a parallel type generation of a multi-modal instruction provided in the embodiments of the present application;
[0034] Figure 5 FIG. 5 is a specific structural schematic diagram of an instruction generation module provided in the embodiments of the present application;
[0035] Figure 6 FIGS. 6(a)-6(e) are process schematic diagrams of the execution of a task by a decision device provided in the embodiments of the present application;
[0036] Figure 7 FIG. 7 is a flow schematic diagram of a decision method provided in the embodiments of the present application;
[0037] Figure 8 FIG. 8 is a structural schematic diagram of a decision device provided in the embodiments of the present application. DETAILED DESCRIPTION
[0038] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.
[0039] The term “and / or” in the present document is a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that there are three cases of A alone, A and B together, and B alone. The symbol “ / ” in the present document represents an or relationship of the associated objects, for example, A / B represents A or B.
[0040] The terms “first” and “second” and the like in the description and claims herein are used for distinguishing between similar objects, not for describing a particular sequential order. For example, the first response message and the second response message are used for distinguishing between different response messages, not for describing a particular sequential order of the response messages.
[0041] In the embodiments of the present application, the words “exemplary” and “for example” are used to mean serving as an example or illustration. Any embodiment or design described herein as “exemplary” or “for example” should not be construed as preferred or advantageous over other embodiments or designs. Rather, use of the words “exemplary” and “for example” is intended to present concepts in a concrete manner.
[0042] In the description of the embodiments of the present application, unless otherwise specified, “a plurality of” means two or more, for example, a plurality of processing units means two or more processing units, and the like; a plurality of elements means two or more elements, and the like.
[0043] Before introducing the technical solutions protected by the present application, several professional terms related to the technical solutions protected by the present application are explained in advance, which are as follows:
[0044] Cross attention is a kind of attention mechanism, which is usually used to process multi-modal information or the association between multiple input sequences. In cross attention, each input sequence has a query, a key and a value, and by calculating the association between the query and the keys of other sequences, a weighted value vector is obtained to represent the degree of association between different sequences. Cross attention has been widely used in natural language processing, image processing and speech processing tasks, which can help the model better understand and process the relationship between different modalities or inputs.
[0045] Diffusion model is a kind of statistical model used to infer, represent and simulate complex data distribution. Diffusion model is based on the propagation process between data points, which estimates the probability distribution of the entire data set by propagating the information of each data point to the neighboring data points. Diffusion model has a wide range of applications in image generation, image denoising, image segmentation and other tasks, which can improve the understanding and generation ability of the model to data.
[0046] A multilayer perceptron (MLP) is a deep learning model composed of multiple layers of neurons, with each layer fully connected to the next. An MLP typically consists of an input layer, several hidden layers, and an output layer, learning complex patterns in input data through multiple nonlinear transformations. Each neuron performs nonlinear processing on the input signal through an activation function, enabling the MLP to learn and approximate complex nonlinear function relationships. MLPs are commonly used for tasks such as classification, regression, and pattern recognition, and are one of the most fundamental models in deep learning.
[0047] Multimodal instructions refer to instructions or information that can handle inputs from multiple sensory modalities, such as vision, hearing, touch, etc. In the field of human-computer interaction, multimodal instructions can be instructions input by users through voice, gestures, touch, etc., and the system can perform corresponding operations or provide feedback based on these instructions. This technology can improve the efficiency and convenience of user interaction with the system, allowing users to communicate and control devices more naturally. Multimodal instructions are widely used in smartphones, smart speakers, virtual reality, and other fields, providing users with more flexible and intuitive interaction methods.
[0048] The decision-making method in the related art is a rule-based planning method and a data-driven planning method. The rule-based planning method relies on a large amount of domain knowledge, so it lacks universality, making it difficult to make reasonable decisions in complex and highly dynamic scenarios. The data-driven planning method generally breaks down tasks into language-describable subtasks, so it cannot be applied to tasks that cannot be clearly and accurately described in language, such as desk tidying, tableware arrangement, and clothes folding. Since the intermediate state of a task that cannot be described in language is not clear in pure language description, the ambiguity brought by pure language description has a great impact on the underlying strategy execution, resulting in reduced performance and generalization ability of the decision-making method.
[0049] To solve the defects in the related art, an embodiment of the present application designs a decision-making method. After receiving task information in any form, the method can perform semantic alignment on the task information, environmental information, and observation information, and then convert the obtained feature information into multimodal instructions. This can clearly describe each subtask in the task, solve the problem of pure language description in complex tasks, and eliminate instruction ambiguity. After converting the task into multimodal instructions, the method can use the multimodal instructions as a medium to control the execution device to perform the corresponding task, thereby improving the task coverage and generalization ability of the execution device. The method can convert the multimodal instructions into action instructions that the execution device can recognize, so that the execution device can clearly understand the operation object and operation target of the action instructions, thereby effectively enhancing the mapping ability of the action instructions and the generalization ability of the strategy execution.
[0050] Figure 1 Fig. 1 is a structural schematic diagram of an intelligent control system according to an embodiment of the present application. As shown in Fig. 1, the intelligent control system 100 can be divided into a perception layer 110, a decision layer 120 and a control layer 130 according to the execution function. The intelligent control system 100 can be deployed on at least one computing device, which can be a smart robot (i.e., an embodied entity), a server, a computer or the like. Figure 1
[0051] In the embodiment of the present application, the perception layer 110 can receive data uploaded by various sensors to obtain environmental information such as temperature, position, surrounding environment image and the like. The perception layer 110 can receive instructions input by a user uploaded by a screen, a microphone and the like to obtain task information of a target task input by the user. The perception layer 110 can receive images uploaded by a specific camera to obtain various objects and states of the objects around the execution device to obtain observation information. After receiving the environmental information, the task information and the observation information provided by the perception layer 110, the decision layer 120 can perform task decomposition based on the environmental information, the task information and the observation information, decompose the target task into a plurality of subtasks, implement planning of the task, and then issue action instructions corresponding to the plurality of subtasks to the control layer 130. After receiving the action instructions corresponding to the respective subtasks in sequence, the control layer 130 performs corresponding operations to implement efficient operation of the intelligent control system 100.
[0052] The execution device refers to a device that executes a target task. For example, the intelligent control system 100 is deployed on only one embodied entity, and the execution device can refer to the embodied entity. For another example, the intelligent control system 100 is deployed on a computer and a mechanical hand, and the computer can control the mechanical hand to execute a skill action, and the execution device refers to the mechanical hand. For another example, the intelligent control system 100 is deployed on a smart phone and a sweeping robot, and the smart phone can control the sweeping robot to sweep, and the execution device refers to the sweeping robot.
[0053] Figure 2 Fig. 2 is a structural schematic diagram of a decision device according to an embodiment of the present application. As shown in Fig. 2, the decision device 200 can be divided into a perception layer 210, a decision layer 220 and a control layer 230 according to the execution function. The decision device 200 can be deployed on at least one computing device, which can be a smart robot (i.e., an embodied entity), a server, a computer or the like. Figure 2 As shown, the decision device 200 can be divided into a semantic alignment module 210, an instruction generation module 220, an instruction tracking module 230, and a feedback module 240 according to the execution function. The decision device 200 can be an application program, software code, etc., which can be deployed on the decision layer 120 described above to execute. That is, the decision device 200 can be deployed on at least one computing device, such as a physical entity, a computer, a server, a portable notebook computer, a tablet computer, a smart phone, etc. The decision device 200 can be deployed on a cloud server. If the decision device 200 is deployed on a cloud service, a designer can use a local device to call the cloud server to use the decision device 200 to complete the corresponding task.
[0054] The semantic alignment module 210 is used to process and align different modalities of environment information, observation information, and task information after receiving the environment information, observation information, and task information, and extract key features in subtasks in the target task, so as to obtain feature information of the subtasks. The semantic alignment module 210 can use cross-attention, large language model, etc. algorithm for semantic alignment.
[0055] Semantic alignment refers to matching and corresponding similar or related content in different data sources or different languages, so as to establish a semantic association between these data. In the embodiments of the present application, semantic alignment refers to matching and corresponding multi-modal data such as text, image, voice, etc. so as to establish a semantic association between different modalities of data.
[0056] The environment information can be in the form of text, image, voice, etc. and refers to the surrounding environment data of the execution device, which can generally include temperature data, positioning data, images of the surrounding environment, depth data, etc.
[0057] The observation information can be in the form of text, image, depth, etc. and refers to the states of each object and each object around the execution device at a certain time. Generally, the observation information is in the form of an image. For example, when the computing device is a physical entity, the observation information refers to the image collected by the camera of the physical entity, which includes each object around the physical entity and the state of each object. In the embodiments of the present application, the perception layer 110 can periodically obtain the observation information, and then send the observation information obtained each time to the semantic alignment module 210, so that the semantic alignment module 210 can generate the feature information of each subtask in turn.
[0058] The task information for the target task can be represented in the form of text, images, voice, etc., referring to the user's current control of the execution device to switch to a set state. For example, the user can operate on the screen to generate the target task. Another example is that the user can speak into a microphone, which records the user's speech to generate the target task. Yet another example is that the user can input text via a keyboard to generate the target task.
[0059] The execution of a task may be complex or involve multiple hardware components working together. Therefore, after receiving a target task, the system can break it down into at least one subtask. Each subtask is a part of the task consisting of one or more steps that the target task must execute. For example, if the target task is "Give me the phone on the table," it can be broken down into the subtasks of "picking up the phone," "moving the phone from the table to the user's hand," and "putting down the phone." By breaking the target task down into at least one subtask, the system can execute the target task according to these subtasks, thereby reducing the difficulty of executing the target task.
[0060] For example, such as Figure 3 As shown, the semantic alignment module 210 may include a visual encoder, a text encoder, a multimodal encoder, a fusion module, a first cross-attention module, and a second cross-attention module. After receiving environmental information in image form and current observation information, the visual encoder can convert the image into corresponding feature representations, obtaining feature vectors for the environmental information and the current observation information, facilitating subsequent information processing and analysis. After receiving environmental information in text form and current observation information, the text encoder can convert the text into corresponding feature representations, obtaining feature vectors for the environmental information and the current observation information, facilitating subsequent information processing and analysis. After receiving environmental information in language form, the speech encoder can convert the language into corresponding feature representations, obtaining feature vectors for the environmental information, facilitating subsequent information processing and analysis. After receiving task information, the multimodal encoder can convert different forms of task information into corresponding feature representations to obtain feature vectors for the task information, facilitating subsequent information processing and analysis.
[0061] After obtaining the feature vector of the environment information, the feature vector of the observation information and the feature vector of the task information, the fusion module can fuse the feature vector of the environment information, the feature vector of the initial observation information, the feature vector of the current observation information and the feature vector of the task information to obtain a first feature vector Z0 and a second feature vector Zi. The first feature vector Z0 represents the feature vector of the alignment fusion between the current task and the initial observation information. The second feature vector Zi represents the feature vector of the alignment fusion between the current task and the current observation information. The initial observation information refers to the observation information received for the first time after receiving the task information, that is, the objects and the states of the objects around the execution device at the initial moment. The current observation information refers to the observation information received at the current moment, that is, the objects and the states of the objects around the execution device at the current moment.
[0062] After obtaining the first feature vector Z0 and the second feature vector Zi, the first cross attention can calculate the correlation between the first feature vector Z0 and the second feature vector Zi, and express it in a weighted feature to obtain the initial feature vector of the target object of the subtask control. The cross attention can perform alignment fusion on the first feature vector Z0 and the second feature vector Zi, so that the output feature vector better represents the relationship between different inputs.
[0063] The second cross attention can calculate the correlation between the initial feature vector of the target object, the feature vector of the environment information, the feature vector of the current observation information and the feature vector of the task information, and express it in a weighted feature to obtain the final feature vector of the target object. The final feature vector represents the feature information of the subtask The feature information indicates the clues of which subtask the target task has reached and which subtask should be reached next. The feature information will be injected into the subsequent diffusion model later, as an accurate semantic guide to guide the generation of the next subtask.
[0064] Optionally, the semantic alignment module 210 can use a cross attention to calculate the correlation between the feature vector of the task information, the feature vector of the environment information, the feature vector of the current observation information, the first feature vector and the second feature vector, and express it in a weighted feature to obtain the feature information of the subtask.
[0065] The instruction generation module 220 is configured to generate at least one multi-modal instruction according to the feature information of the subtask and the current observation information. In the embodiment of the present application, the instruction generation module 220 generates the target state of the target object and the skill action instruction in the process of generating the multi-modal instruction. The instruction generation module 220 can generate the multi-modal instruction in two different ways according to the order of generating the target state and the skill action instruction. One way is to generate the target state first and then generate the skill action instruction. The other way is to generate the target state and the skill action instruction at the same time. Therefore, the instruction generation module 220 can divide the first generation module 221 and the second generation module 222 according to the execution function. The functions of the two generation modules will be introduced below.
[0066] As shown in FIG. 4(a), the first generation module 221 is configured to generate the target state of the target object controlled by the subtask according to the feature information of the subtask and the current observation information. In the embodiment of the present application, the first generation module 221 can use a diffusion model to generate the target state of the target object on the basis of the current observation information, taking the feature information of the subtask as a conditional guidance signal.
[0067] Exemplarily, taking the current observation information in the form of an image as an example, the current observation information is a current observation image. The diffusion model can use a frame splicing method to splice the current observation image and a standard Gaussian noise image together to obtain a spliced image, thereby providing a visually continuous feature information for the denoising process. The diffusion model can convert the spliced image into a set of data points, and then recover the clear structure in the initial image through the information propagation process between the data points to generate a denoised image. The diffusion model can train the denoised image using the feature information of the target object to generate the target state of the target object. At this time, the target state is in the form of an image.
[0068] The first generation module 221 is configured to generate the skill action instruction required for the target object to reach the target state from the current observation information according to the target state of the target object and the current observation information. The first generation module 221 is configured to combine the target state of the target object, the current observation information and the skill action instruction together to construct at least one multi-modal instruction. The multi-modal instruction is used to convert the execution device to a specified state, so that the state of the target object in the next observation information collected is the target state.
[0069] Exemplarily, taking the observation information in the form of an image as an example, the observation information is an observation image. The target state of the target object generated by the first generation module 221 is in the form of an image, which can be a target image. As shown in FIG. 4(b), the first generation module 221 generates the target image of the target object according to the feature information of the subtask and the current observation image. The first generation module 221 generates the skill action instruction according to the target image and the current observation image. Figure 5As shown, after the first generation module 221 receives the target image and the current observation image, the key image and the current observation image can be input into the MLP. Each layer of neurons of the MLP performs a non-linear transformation on the input target image and current observation image through an activation function, and passes the result to the next layer, so as to output the skill action instruction required for the target object to reach the target image from the current observation image. The first generation module 221 can combine the target image, the current observation image and the skill action instruction together, and use the skill action instruction to associate the target image with the current observation image, to obtain at least one multi-modal instruction for the target object to reach the target image from the current observation image. The multi-modal instruction can more comprehensively describe the requirements of the task or action, and also provide more information and context to help the system more accurately understand and perform the task.
[0070] In the embodiments of the present application, the decision device 200 can implement serial generation of multi-modal instructions based on the first generation module 221, which can reuse a large amount of actionless label data, such as human video data, and can reduce the difficulty of data acquisition. After the first generation module 221 generates the multi-modal instruction, the multi-modal instruction can be used to control the execution equipment to perform the corresponding task, thereby improving the task coverage ability and generalization ability of the execution equipment.
[0071] As shown in FIG. 4(b), the second generation module 222 is configured to generate a skill action instruction required for a target object to reach a target state from a state in current observation information according to feature information of a subtask and the current observation information. Exemplarily, the second generation module 222 can include a transformer model, a text encoder and a diffusion model. After the transformer model receives the feature information of the subtask, the feature information of the subtask can be encoded to obtain feature information of a skill action and feature information of a target object controlled by the subtask. The text decoder can decode the feature information of the skill action to obtain the skill action instruction. The diffusion model can use the feature information of the target object as a conditional guidance signal to generate a target state of the target object based on the current observation information. The diffusion model can combine the target state of the target object, the current observation information and the skill action instruction together to construct at least one multi-modal instruction.
[0072] In an embodiment of the present application, the second generation module 222 can directly output skill action instructions and target states according to the feature information of the sub-tasks, implement parallel generation of multi-modal instructions, solve the problem that pure language cannot describe in complex tasks, and eliminate instruction ambiguity. Compared with the use of observation and target state in the reverse solution of skill action in the serial generation of multi-modal instructions, the second generation module 222 simultaneously generates target states and skill action instructions in the process of parallel generation of multi-modal instructions. Therefore, the generation of skill action instructions does not need to rely on the target state, so the second generation module 222 can generate other control information such as the moving speed, direction, angle, etc. of the skill action when generating the skill action instructions, so that the skill action instructions are more detailed and comprehensive. In addition, the second generation module 222 is trained by large-scale multi-modal instruction data, and the parallel multi-modal instruction can have a wider instruction representation capability and can cover a more comprehensive embodied task.
[0073] The instruction tracking module 230 is configured to, after obtaining the at least one multi-modal instruction, convert the at least one multi-modal instruction into a corresponding feature representation by using a multi-modal encoder to obtain a feature vector of the at least one multi-modal instruction. The instruction tracking module 230 can convert the feature vector of the at least one multi-modal instruction into a plurality of action instructions executable by the execution device by using a transformer model. The computing device can convert the target object from the state in the current observation information to the target state based on the action instructions.
[0074] In an embodiment of the present application, the instruction tracking module 230 can convert the multi-modal instruction into an action instruction recognizable by the execution device, so that the execution device can clearly understand the operation object and operation target of the action instruction, thereby effectively enhancing the mapping capability of the action instruction and the action and improving the generalization capability of the strategy execution.
[0075] The feedback module 240 can receive feedback information sent by the execution device or observation information collected next time, and determine the execution situation of the task according to the received information. The feedback information is used to indicate that the execution device completes the plurality of action instructions. The execution device does not execute the action instruction means that the state of the target object is not converted to the target state. The execution device executes all the action instructions means that the state of the target object is the target state.
[0076] In one case, when the feedback module 240 determines that the execution device does not execute the received action instruction, the instruction tracking module 230 can be instructed to reissue the unexecuted action instruction.
[0077] In another case, when the feedback module 240 determines that the execution device cannot execute the received action instruction, the user can be informed through the screen, the loudspeaker, etc.
[0078] In another case, the feedback module 240 determines that the execution device has completed all the action instructions, and can inform the user through the screen, speaker, etc. component, or instruct the perception layer 110 to collect the next observation information, and instruct the semantic alignment module 210, the instruction generation module 220 and the instruction tracking module 230 to process the next subtask in the task. In this way, all subtasks in the task are completed.
[0079] The implementation process of the technical solutions protected by the present application will be described below with an embodiment as an example.
[0080] Suppose the user inputs a task "arrange four files neatly on the right upper corner of the desktop" in the form of pure text. After the decision device 200 receives the task information, the initial observation information is obtained. As shown in Figure 6 (a), the four files are placed in various positions on the desktop in disorder. For the convenience of subsequent description, the "leftmost file" is defined as "file 1", the "second left file" is defined as "file 2", the "second right file" is defined as "file 3", and the "rightmost file" is defined as "file 4".
[0081] The semantic alignment module 210 of the decision device 200 performs semantic alignment on the task information, the initial observation information and the environment information after receiving the task information, the initial observation information and the environment information, and obtains the feature information of the target state of file 4. The instruction generation module 220 of the decision device 200 generates the target state of file 4 in the process of generating the multi-modal instruction, as shown in Figure 6 (b). The instruction tracking module 230 of the decision device 200 converts the multi-modal instruction into a plurality of action instructions, and then issues the plurality of action instructions to the execution device. After receiving the plurality of action instructions, the execution device performs corresponding operations to move file 4 from the right lower corner of the desktop to the right upper corner, that is, to move file 4 to the position as shown in Figure 6 (b).
[0082] When the feedback module 240 of the decision device 200 receives the observation information at the current time as shown in Figure 6 (b), the semantic alignment module 210, the instruction generation module 220 and the instruction tracking module 230 can be instructed to generate action instructions again. At this time, the target state of file 2 generated by the instruction generation module 220 is as shown in Figure 6 (c). The plurality of action instructions issued by the instruction tracking module 230 to the execution device moves file 2 from the left middle position of the desktop to the right upper corner, that is, moves file 2 to the position as shown in Figure 6 (c).
[0083] When the feedback module 240 of the decision device 200 receives the observation information at the current time as shown inFigure 6 (c) shown, the semantic alignment module 210, the instruction generation module 220 and the instruction tracking module 230 can be instructed to generate action instructions again. At this time, the target state of file 1 generated by the instruction generation module 220 is as shown in Figure 6 (d). The instruction tracking module 230 issues a plurality of action instructions to the execution device, so that the execution device moves file 1 from the lower left of the desktop to the upper right, that is, moves file 1 to the position as shown in Figure 6 (d).
[0084] The feedback module 240 of the decision device 200 receives the observation information at the current time as shown in Figure 6 (d). The semantic alignment module 210, the instruction generation module 220 and the instruction tracking module 230 can be instructed to generate action instructions again. At this time, the target state of file 3 generated by the instruction generation module 220 is as shown in Figure 6 (e). The instruction tracking module 230 issues a plurality of action instructions to the execution device, so that the execution device moves file 3 from the lower left of the desktop to the upper right, that is, moves file 3 to the position as shown in Figure 1-6 (e).
[0085] The implementation process of the technical solution of the protection in Figure 7 will be introduced in the following flow manner.
[0086] Figure 7 A flowchart of a decision method provided in an embodiment of the present application. As shown in Figure 8 , the method can be executed by the above-mentioned decision device 200, and the specific implementation process is as follows:
[0087] Step S701, receiving task information of a target task.
[0088] Step S702, performing semantic alignment on the task information, the environment information and the first observation information to obtain feature information of a first subtask.
[0089] The above steps S701-S702 can be executed by the semantic alignment module 210 in the decision device 200.
[0090] The semantic alignment module 210 can also receive environment information and first observation information in the process of receiving task information. The task information of the target task can be in the form of text, image, voice, etc., and refers to the control of the user to execute the device to switch to the set state. The target task can include at least one subtask. Each subtask refers to a part of the task that is one or more steps of the target task. The environment information can be in the form of text, image, voice, etc., and refers to the surrounding environment data of the execution device. The observation information can be in the form of text, image, depth, etc., and refers to the state of each object and each object around the execution device at a certain time. The first observation information refers to the observation information received at the current time.
[0091] After receiving the environment information, observation information, and task information, the semantic alignment module 210 can use text, image, voice, etc. Multimodal encoding technology to process and align the environment information, observation information, and task information of different modalities, extract the key features in the first subtask of the target task, and obtain the feature information of the first subtask. The semantic alignment module 210 can use cross attention, large language model, etc. Algorithm for semantic alignment.
[0092] In one embodiment, the semantic alignment module 210 can use a visual encoder to convert the image-form environment information and the first observation information into corresponding feature representations, respectively, to obtain the feature vector of the environment information and the feature vector of the first observation information. The semantic alignment module 210 can use a text encoder to convert the text-form environment information and the first observation information into corresponding feature representations, respectively, to obtain the feature vector of the environment information and the feature vector of the first observation information. The semantic alignment module 210 can use a speech encoder to convert the speech-form environment information into a corresponding feature representation, to obtain the feature vector of the first observation information. The semantic alignment module 210 can use a text encoder to convert the task information of different forms into a corresponding feature representation, to obtain the feature vector of the task information.
[0093] The semantic alignment module 210 can fuse the feature vector of the environment information, the feature vector of the first observation information, and the feature vector of the task information to obtain the first feature vector and the second feature vector. The semantic alignment module 210 can use cross attention to calculate the correlation between the feature vector of the task information, the feature vector of the environment information, the initial observation information, the feature vector of the first observation information, the first feature vector, and the second feature vector, and to obtain a weighted feature representation to generate the feature information of the first subtask.
[0094] Step S703, based on the feature information of the first subtask and the first observation information, at least one first multimodal instruction is generated.
[0095] The step S703 can be executed by the instruction generation module 220 in the decision device 200.
[0096] In one scheme, the instruction generation module 220 generates the multi-modal instruction in series. Specifically:
[0097] The instruction generation module 220 can utilize the diffusion model to generate the target state of the target object controlled by the first sub-task based on the current observation information, taking the feature information of the first sub-task as the conditional guidance signal. The instruction generation module 220 is configured to generate the skill action instruction required for the target object to reach the target state from the state in the current observation information according to the target state of the target object and the current observation information. The instruction generation module 220 is configured to combine the target state, the current observation information and the skill action instruction together to construct at least one multi-modal instruction. The multi-modal instruction is configured to convert the execution device to a specified state so that the state of the target object in the next observation information collected is the target state.
[0098] In another scheme, the instruction generation module 220 generates the multi-modal instruction in parallel. Specifically:
[0099] After receiving the feature information of the first sub-task, the instruction generation module 220 can decode the feature information of the first sub-task to obtain the feature information of the skill action and the feature information of the target object controlled by the first sub-task. The instruction generation module 220 can decode the feature information of the skill action to obtain the skill action instruction. The instruction generation module 220 can generate the target state of the target object based on the current observation information, taking the feature information of the first sub-task as the conditional guidance signal. The instruction generation module 220 can combine the target state, the current observation information and the skill action instruction together to construct at least one multi-modal instruction.
[0100] In step S704, the at least one first multi-modal instruction is converted into a plurality of action instructions.
[0101] The step S704 can be executed by the instruction tracking module 230 in the decision device 200.
[0102] The instruction tracking module 230 is configured to, after obtaining the at least one multi-modal instruction, convert the at least one multi-modal instruction into a corresponding feature representation by using a multi-modal encoder to obtain a feature vector of the at least one multi-modal instruction. The instruction tracking module 230 can convert the feature vector of the at least one multi-modal instruction into a plurality of action instructions executable by the execution device by using a transformer model. The computing device can convert the target object from the state in the current observation information to the target state based on the action instruction.
[0103] The feedback module 240 can receive feedback information sent by the execution device or next observation information collected, and determine the execution of the task according to the received information. The feedback information is used to indicate that the execution device completes the plurality of action instructions. The execution device does not execute the action instruction, which means that the state of the target object does not change to the target state. The execution device executes all the action instructions, which means that the state of the target object is the target state.
[0104] In one case, when the feedback module 240 determines that the execution device does not execute the received action instruction, the instruction tracking module 230 can be instructed to reissue the unexecuted action instruction. In another case, when the feedback module 240 determines that the execution device cannot execute the received action instruction, the user can be informed through the screen, speaker, and other components. In another case, when the feedback module 240 determines that the execution device executes all the action instructions, the user can be informed through the screen, speaker, and other components, or the perception layer 110 can be instructed to collect the next observation information, and the semantic alignment module 210, the instruction generation module 220, and the instruction tracking module 230 can be instructed to process the next subtask in the task. In this way, all subtasks in the task can be completed.
[0105] In the embodiments of the present application, after receiving task information in any form, the task information can be semantically aligned with environmental information and observation information, and then the obtained feature information can be converted into multi-modal instructions, which can clearly describe each subtask in the task, solve the problem that pure language cannot describe complex tasks, and eliminate instruction ambiguity. After the task is converted into multi-modal instructions, the execution device can be controlled to execute the corresponding task through the multi-modal instructions, thereby improving the task coverage and generalization ability of the execution device.
[0106] Compared with the prior art, which uses pure language or application programming interface (API) to disassemble and describe tasks, it cannot be applied to language-described tasks. The technical solution protected by the present application introduces multi-modal description schemes such as text, images, and voice, which not only applies to language-described tasks, but also applies to some tasks that cannot be described by language, so that the technical solution protected by the present application can cover more tasks.
[0107] Compared with the prior art, which uses pure language or application programming interface (API) to disassemble and describe tasks, it cannot be applied to language-described tasks. The technical solution protected by the present application introduces multi-modal description schemes such as text, images, and voice, which not only applies to language-described tasks, but also applies to some tasks that cannot be described by language, so that the technical solution protected by the present application can cover more tasks.
[0108] Figure 8 A structural schematic diagram of a computing device provided in the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the computing device includes a perception layer 110, a semantic alignment module 210, an instruction generation module 220, an instruction tracking module 230, and a feedback module 240.Figure 8 As shown, the computing device 800 includes a bus 810, a processor 820, a memory 830, and a communication interface 840. The processor 820, the memory 830, and the communication interface 840 communicate with each other through the bus 810. The computing device 800 can be a server, a computer, a robot, etc. It should be understood that the number of processors and memories in the computing device 800 is not limited.
[0109] The bus 810 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 In the figure, only one line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus. The bus 810 can include a path for transmitting information between various components (e.g., the processor 820, the memory 830, the communication interface 840) of the computing device 800.
[0110] The processor 820 can be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.
[0111] The memory 830 can include a volatile memory (e.g., a random access memory (RAM)), and can also include a non-volatile memory (e.g., a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD)).
[0112] The memory 830 stores executable program codes, and the processor 820 executes the executable program codes to respectively implement the functions of the aforementioned modules, such as the semantic alignment module 210, the instruction generation module 220, the instruction tracking module 230, and the feedback module 240, etc., so as to implement the decision method. That is, the memory 830 stores instructions for executing the decision method.
[0113] Alternatively, the memory 830 stores executable code that, when executed by the processor 820, implements the functionality of the aforementioned modules to implement the decision-making method. That is, the memory 830 stores instructions for implementing the decision-making method.
[0114] The communication interface 840 uses a transceiver module such as, but not limited to, a network interface card, a transceiver, and the like to enable communication between the computing device 800 and other devices or communication networks.
[0115] In an embodiment, the communication interface 840 is configured to receive task information of a target task. The task information is represented in one or more of text, image, voice, and the like. The target task includes at least one subtask. The processor 820 is configured to perform semantic alignment on the task information, environment information, and first observation information to obtain feature information of a first subtask. The environment information refers to surrounding environment data of the execution device. The first observation information refers to a plurality of objects in the surroundings of the execution device at a first time and states of the plurality of objects. The at least one subtask includes the first subtask. The processor 820 is configured to generate at least one first multi-modal instruction based on the feature information of the first subtask and the first observation information. The at least one first multi-modal instruction is configured to convert the execution device to a first state such that a state of the plurality of objects in observation information collected next time is a target state after completion of the first subtask.
[0116] In an embodiment, the instruction generation module is specifically configured to generate a target state of a target object based on the feature information of the first subtask and the first observation information. The target object refers to an object to be controlled by the first subtask, and the plurality of objects includes the target object. A skill action instruction is generated based on the target state of the target object and the first observation information. At least one first multi-modal instruction is generated based on the target state of the target object, the skill action instruction, and the first observation information.
[0117] In an embodiment, the processor 820 is specifically configured to obtain feature information of a skill action and feature information of the target state of the target object based on the feature information of the first subtask. The processor 820 is specifically configured to generate the target state of the target object based on the feature information of the target state of the target object and the first observation information. The processor 820 is specifically configured to generate at least one first multi-modal instruction based on the target state of the target object, the skill action instruction, and the first observation information. The skill action instruction is obtained by decoding the feature information of the skill action.
[0118] In an embodiment, the processor 820 is configured to convert the at least one first multi-modal instruction into a plurality of action instructions. The action instruction refers to an instruction executable by the execution device.
[0119] In an embodiment, the processor 820 is specifically configured to encode the at least one first multi-modal instruction to obtain a feature vector of the at least one first multi-modal instruction. The processor 820 is specifically configured to convert the feature vector of the at least one first multi-modal instruction into a plurality of action instructions.
[0120] In an embodiment, the processor 820 is specifically configured to respectively encode the task information, the environment information, and the first observation information to obtain a feature vector of the task information, a feature vector of the environment information, and a feature vector of the first observation information. The processor 820 is specifically configured to fuse the feature vector of the task information, the feature vector of the environment information, the feature vector of the initial observation information, and the feature vector of the first observation information to obtain a first feature vector and a second feature vector. The first feature vector represents an aligned and fused feature vector between the task information and the initial observation information. The second feature vector represents an aligned and fused feature vector between the task information and the first observation information. The processor 820 is specifically configured to calculate a correlation between the feature vector of the task information, the feature vector of the environment information, the feature vector of the first observation information, the first feature vector, and the second feature vector to obtain feature information of the first sub-task.
[0121] In an embodiment, the communication interface 840 is further configured to receive second observation information. The second observation information refers to observation information collected after the first sub-task is completed. The processor 820 is further configured to perform semantic alignment on the task information, the environment information, and the second observation information to obtain second sub-task feature information. The at least one sub-task includes the second sub-task. The processor 820 is further configured to generate at least one second multi-modal instruction based on the feature information of the second sub-task and the second observation information. The at least one second multi-modal instruction is used to convert the execution device to a second state, so that a plurality of objects in the observation information collected next time are in a target state after the second sub-task is completed.
[0122] In an embodiment, the communication interface 840 is further configured to receive feedback information sent by the execution device before receiving the second observation information. The feedback information is used to obtain the second observation information in a case where it is determined that the execution device completes an action operation corresponding to the at least one first multi-modal instruction.
[0123] The embodiments of the present application also provide a computer readable storage medium including computer program instructions, when the computer program instructions are executed by a computing device, the computing device performs any one of the methods described in the embodiments of the present application and the corresponding description. Figure 7 and any one of the methods described in the embodiments of the present application and the corresponding description.
[0124] The embodiments of the present application also provide a computer program product including instructions, characterized in that the computer program product stores instructions, when the instructions are executed by a computing device, the computing device implements any one of the methods described in the embodiments of the present application and the corresponding description.Figure 2 any of the methods recited in the corresponding description.
[0125] Those skilled in the art can realize the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be realized in electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software depends on the specific application and design constraints. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present application.
[0126] Furthermore, various aspects or features of the embodiments disclosed herein can be realized using any combination of methods, apparatuses, or articles of manufacture. The term "article of manufacture" as used herein is intended to encompass a computer program accessible from any computer-readable device, carrier, or media. For example, computer-readable media can include but are not limited to magnetic storage devices (e.g., hard disk, floppy disk, magnetic strips, etc.), optical disks (e.g., compact disk (CD), digital versatile disk (DVD), etc.), smart cards, and flash memory devices (e.g., EPROM, card, stick, or key drive, etc.). Additionally, various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine- readable medium" can include, without being limited to, wireless channels and various other media capable of storing, containing, and / or carrying instruction and / or data.
[0127] In the above-described embodiments, The decision device 200 in the foregoing embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, the decision device 200 can be implemented in the form of a computer program product in whole or in part. The computer program product includes at least one computer instruction. When the computer program instruction is loaded and executed on a computer, the flow or function described in the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center through a wired (for example, coaxial cable, optical fiber, digital subscriber line) or wireless (for example, infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with at least one available medium. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD) or a semiconductor medium (for example, SSD) and the like.
[0128] It should be understood that the size of the sequence number of each process described above does not mean the order of execution in various embodiments of the embodiments of the present application. The execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0129] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0130] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0131] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0132] The functions, if implemented in the form of software functional units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application can be embodied in the form of software products, and the computer software products are stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or an access network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, and various storage program codes.
[0133] The above is only a specific implementation of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the embodiments of the present application, which should be covered within the protection scope of the embodiments of the present application.
Claims
1. A decision-making method, characterized in that, include: Receive task information for the target task; the task information is represented in one or more forms, such as text, image, and voice. The target task includes at least one sub-task; The task information, environmental information, and first observation information are semantically aligned to obtain the feature information of the first subtask; the environmental information refers to the surrounding environment data of the execution device; the first observation information refers to multiple objects around the execution device and the state of the multiple objects at a first moment; the at least one subtask includes the first subtask; Based on the feature information of the first subtask and the first observation information, at least one first multimodal instruction is generated; The at least one first multimodal instruction is used to switch the execution device to a first state so that the state of the plurality of objects in the next acquired observation information is the target state after the completion of the first sub-task.
2. The method according to claim 1, characterized in that, The step of generating at least one first multimodal instruction based on the feature information of the first subtask and the first observation information specifically includes: Based on the feature information of the first subtask and the first observation information, the target state of the target object is generated; the target object refers to the object to be controlled by the first subtask, and the plurality of objects includes the target object; Based on the target state of the target object and the first observation information, a skill action command is generated; Based on the target state of the target object, the skill action instruction, and the first observation information, the at least one first multimodal instruction is generated.
3. The method according to claim 1, characterized in that, The step of generating at least one first multimodal instruction based on the feature information of the first subtask and the first observation information specifically includes: Based on the feature information of the first subtask, the feature information of the skill action and the feature information of the target state of the target object are obtained; The target state of the target object is generated based on the feature information of the target state of the target object and the first observation information; Based on the target state of the target object, the skill action instruction, and the first observation information, at least one first multimodal instruction is generated; the skill action instruction is obtained by decoding the feature information of the skill action.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: The at least one first multimodal instruction is converted into multiple action instructions; the action instructions refer to instructions that the execution device can execute.
5. The method according to claim 4, characterized in that, The step of converting the at least one first multimodal instruction into multiple action instructions specifically includes: The at least one first multimodal instruction is encoded to obtain the feature vector of the at least one first multimodal instruction; The feature vector of the at least one first multimodal instruction is converted into the plurality of action instructions.
6. The method according to any one of claims 1-5, characterized in that, The semantic alignment of the task information, environmental information, and first observation information to obtain the feature information of the first sub-task specifically includes: The task information, the environmental information, and the first observation information are encoded respectively to obtain the feature vectors of the task information, the environmental information, and the first observation information. The feature vectors of the task information, the environment information, the initial observation information, and the first observation information are fused to obtain a first feature vector and a second feature vector; the first feature vector represents the alignment and fusion feature vector between the task information and the initial observation information; the second feature vector represents the alignment and fusion feature vector between the task information and the first observation information. The feature vectors of the task information, the environmental information, the first observation information, and the correlation between the first feature vector and the second feature vector are calculated to obtain the feature information of the first subtask.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: Receive the second observation information; the second observation information refers to the observation information collected after the first subtask is completed; Semantic alignment is performed on the task information, the environmental information, and the second observation information to obtain the second subtask feature information; the at least one subtask includes the second subtask. Based on the feature information of the second sub-task and the second observation information, at least one second multimodal instruction is generated; the at least one second multimodal instruction is used to switch the execution device to a second state so that the state of the multiple objects in the next collected observation information is the target state after the completion of the second sub-task.
8. The method according to claim 7, characterized in that, Before receiving the second observation information, the method further includes: The system receives feedback information sent by the execution device; the feedback information is used to obtain the second observation information when it is determined that the execution device has completed the action operation corresponding to the at least one first multimodal instruction.
9. A decision-making device, characterized in that, include: The semantic alignment module is used to receive task information of the target task; the task information is represented in one or more forms, such as text, image, and speech. The target task includes at least one sub-task; as well as The task information, environmental information, and first observation information are semantically aligned to obtain the feature information of the first subtask; the environmental information refers to the surrounding environment data of the execution device; the first observation information refers to multiple objects around the execution device and the state of the multiple objects at a first moment; the at least one subtask includes the first subtask; The instruction generation module is used to generate at least one first multimodal instruction based on the feature information of the first subtask and the first observation information; The at least one first multimodal instruction is used to switch the execution device to a first state so that the state of the plurality of objects in the next acquired observation information is the target state after the completion of the first sub-task.
10. The apparatus according to claim 9, characterized in that, The instruction generation module is specifically used to generate the target state of the target object based on the feature information of the first subtask and the first observation information; the target object refers to the object to be controlled by the first subtask, and the plurality of objects includes the target object; Based on the target state of the target object and the first observation information, a skill action command is generated; Based on the target state of the target object, the skill action instruction, and the first observation information, the at least one first multimodal instruction is generated.
11. The apparatus according to claim 9, characterized in that, The instruction generation module is specifically used to obtain the feature information of the skill action and the feature information of the target state of the target object based on the feature information of the first sub-task. The target state of the target object is generated based on the feature information of the target state of the target object and the first observation information; Based on the target state of the target object, the skill action instruction, and the first observation information, at least one first multimodal instruction is generated; the skill action instruction is obtained by decoding the feature information of the skill action.
12. The apparatus according to any one of claims 9-11, characterized in that, Also includes: The instruction tracking module is used to convert the at least one first multimodal instruction into multiple action instructions; the action instructions refer to instructions that the execution device can execute.
13. The apparatus according to claim 12, characterized in that, The instruction tracking module is specifically used to encode the at least one first multimodal instruction to obtain the feature vector of the at least one first multimodal instruction; The feature vector of the at least one first multimodal instruction is converted into the plurality of action instructions.
14. The apparatus according to any one of claims 9-13, characterized in that, The semantic alignment module is specifically used to encode the task information, the environment information, and the first observation information respectively to obtain the feature vector of the task information, the feature vector of the environment information, and the feature vector of the first observation information; The feature vectors of the task information, the environment information, the initial observation information, and the first observation information are fused to obtain a first feature vector and a second feature vector; the first feature vector represents the alignment and fusion feature vector between the task information and the initial observation information; the second feature vector represents the alignment and fusion feature vector between the task information and the first observation information. The feature vectors of the task information, the environmental information, the first observation information, and the correlation between the first feature vector and the second feature vector are calculated to obtain the feature information of the first subtask.
15. The apparatus according to any one of claims 9-14, characterized in that, The semantic alignment module is also used to receive second observation information; the second observation information refers to the observation information collected after the first subtask is completed. Semantic alignment is performed on the task information, the environmental information, and the second observation information to obtain the second subtask feature information; the at least one subtask includes the second subtask. The instruction generation module is further configured to generate at least one second multimodal instruction based on the feature information of the second sub-task and the second observation information; the at least one second multimodal instruction is configured to switch the execution device to a second state so that the state of the multiple objects in the next collected observation information is the target state after the completion of the second sub-task.
16. The apparatus according to claim 15, characterized in that, The semantic alignment module, before receiving the second observation information, is further used for... The system receives feedback information sent by the execution device; the feedback information is used to obtain the second observation information when it is determined that the execution device has completed the action operation corresponding to the at least one first multimodal instruction.
17. A computing device, characterized in that, include: At least one memory; At least one processor, the processor being configured to execute instructions stored in memory to cause the computing device to perform the method as described in any one of claims 1-8.
18. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a computing device, perform the method as described in any one of claims 1-8.
19. A computer program product containing instructions, characterized in that, The computer program product stores instructions that, when executed by a computing device, cause the computing device to perform the method according to any one of claims 1-8.