Control method based on artificial intelligence AI and related device

By combining embodied intelligence agents and diffusion networks, action sequences are generated and executed, solving the problem of unsatisfactory performance of diffusion strategies in complex environments and achieving more efficient and accurate action sequence generation and execution.

CN121928532APending Publication Date: 2026-04-28HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-10-25
Publication Date
2026-04-28

Smart Images

  • Figure CN121928532A_ABST
    Figure CN121928532A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a control method based on artificial intelligence AI and a related device.The method comprises the steps that multiple pieces of key information in first information and a target task are input into a smart Agent with the body, a key information path is generated, the first information is used for describing the state of an operation scene, and the target task is used for executing the target task; the key information path is used for representing the process of completing the target task in the operation scene; and inputting the key information path into a control network to obtain an action sequence needing to be executed for completing the target task. By adopting the embodiment of the invention, the action sequence with a better execution effect can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence (AI), and in particular to a control method and related device based on artificial intelligence (AI). Background Technology

[0002] Diffusion strategies are probabilistic control methods commonly used in robotics and automation systems. Their core idea is to simulate and predict robot actions and trajectories using probability distributions. In practical applications, diffusion strategies generate optimal or near-optimal action sequences by considering possible actions and their consequences. Robotic arm control problems often involve action execution, and diffusion strategies can plan an optimal or near-optimal action sequence for a given environment. However, in complex environments, with rapid dynamic changes, and where real-time response is required, the effectiveness of diffusion strategies in planning action sequences is not ideal. Summary of the Invention

[0003] This application discloses a control method and related device based on artificial intelligence (AI), which can improve the effect of planned action sequences.

[0004] In a first aspect, embodiments of this application provide a control method based on artificial intelligence (AI), the method comprising:

[0005] Multiple key pieces of information and target tasks from the first information are input into the embodied intelligent agent to generate a key information path, wherein the first information is used to describe the state of the work scenario, and the key information path is used to characterize the process of completing the target task in the work scenario;

[0006] The key information path is input into the control network to obtain the sequence of actions required to complete the target task.

[0007] In the above method, before predicting the action sequence for completing the target task through a control network (such as a diffusion network), an embodied intelligence agent first analyzes the initial information of the work scenario where the target task is located to obtain the key information path in the initial information. This key information path is then input into the control network for processing to generate the action sequence. Because the embodied intelligence agent performs preprocessing, meaning the input to the control network is processed information of higher quality, the action sequence predicted by the control network performs better in completing the target task. In other words, if the control network lacks prior knowledge, has insufficient understanding and generalization ability, and cannot possess common sense, human-like perception, and actions, then the introduction of an embodied intelligence agent effectively compensates for this problem. Furthermore, if the embodied intelligence agent cannot perform functions such as few-shot learning, multi-capability coverage, autonomous interaction with the environment, and autonomous evolutionary learning, then the control network (such as a diffusion network) can compensate for these shortcomings.

[0008] In conjunction with the first aspect, one possible implementation of the first aspect also includes:

[0009] The system controls the first object in the work scenario to execute the action sequence. In other words, in addition to generating the action sequence, it can also control the first object to execute the first sequence to complete the target task, thus realizing the integration of action sequence generation and execution.

[0010] In conjunction with the first aspect, or any of the possible implementations of the first aspect described above, in yet another possible implementation of the first aspect, the key information path includes one or more path segments. Key information at the start position of each path segment is used as the initial state of the control network, and key information at the end position of each path segment is used as the final state of the control network. The control network generates a corresponding action for each path segment using each path segment as input. The actions corresponding to the multiple path segments are used to form a sequence of actions required to complete the target task. It can be understood that dividing the key information path into path segments is equivalent to the control network being able to generate action sequences with smaller-granularity inputs, achieving more refined generation, and the generated action sequences have higher accuracy during execution.

[0011] In conjunction with the first aspect, or any of the above-mentioned possible implementations of the first aspect, another possible implementation of the first aspect further includes:

[0012] The feature information of the first information is obtained, wherein the feature information includes two-dimensional feature information and / or three-dimensional point cloud information;

[0013] The feature information is analyzed using a visual model to obtain several key pieces of information from the first information.

[0014] This is understandable; the key information in the initial information is extracted through model analysis.

[0015] In conjunction with the first aspect, or any of the above-mentioned possible implementations of the first aspect, another possible implementation of the first aspect further includes:

[0016] The system queries a first database for second information that has a preset similarity relationship with the first information, wherein the first database stores multiple pieces of information and key information corresponding to each of the multiple pieces of information.

[0017] Multiple key pieces of information in the first information are determined based on multiple key pieces of information in the second information.

[0018] This is understandable. A first database is pre-established to provide a basis for memory. When the first piece of information is input later, similar second information is directly searched from the first database, and the key information of the first information is determined based on the key information of the second information.

[0019] In conjunction with the first aspect, or any of the above possible implementations of the first aspect, in another possible implementation of the first aspect, the step of inputting multiple key pieces of information and the target task from the first information into the embodied intelligent agent to generate a key information path includes:

[0020] The first model predicts multiple key pieces of information in the first information, wherein the first model is trained based on multiple first training data, and the first training data includes feature information of the third information and key information in the third information.

[0021] The second model predicts the key information path composed of the multiple key information items, wherein the second model is trained based on multiple second training data, and the second training data includes the key information in the fourth information and the key information path corresponding to the fourth information.

[0022] It is understandable that here, a first model and a second model are trained. Then, when it is necessary to determine the key information and key information path of the first information, the first information is directly input into the first model and the second model to obtain the key information and key information path respectively. This method is highly efficient, and its prediction effect is usually good when the training data of the first model and the second model is sufficient.

[0023] In conjunction with the first aspect, or any of the above possible implementations of the first aspect, in another possible implementation of the first aspect, the step of inputting multiple key pieces of information and the target task from the first information into the embodied intelligent agent to generate a key information path includes:

[0024] The key information path corresponding to the first information is predicted by the third model, wherein the third model is trained based on multiple third training data, including the feature information of the fifth information, the key information in the fifth information, and the key information path corresponding to the fifth information.

[0025] It is understandable that here a third model is trained. Then, when it is necessary to determine the key information path of the first information, the first information is directly input into the third model to obtain the key information path. This method is more efficient, and its prediction effect is usually better when the training data of the third model is sufficient.

[0026] In conjunction with the first aspect, or any of the above-mentioned possible implementations of the first aspect, another possible implementation of the first aspect further includes:

[0027] Obtain effect parameters, wherein the effect parameters are used to characterize the effect of completing the target task according to the action sequence, and the effect parameters are used to calibrate the key information generated by the embodied intelligent agent next time.

[0028] It is understandable that the execution effect of the action sequence generated by the control network can be evaluated to obtain effect parameters. These effect parameters can then be used to optimize the key information paths input to the control network, thereby improving the effect of the action sequence generated by the control network.

[0029] In conjunction with the first aspect, or any of the above possible implementations of the first aspect, in yet another possible implementation of the first aspect, the input of the control network further includes the state of a first object executing the operation sequence, the state of the first object being used to constrain the action sequence generated by the control network.

[0030] It is understandable that, considering that the first object is the main body that directly executes the action sequence, the state of the first object itself usually affects the effect of the first object executing the action sequence. Therefore, taking the state of the first object as an input to the control network can enable the control network to generate an action sequence that is more in line with the operation requirements of the first object, and the effect will be better when the first object executes the action sequence later.

[0031] In conjunction with the first aspect, or any of the above possible implementations of the first aspect, in yet another possible implementation of the first aspect, the input to the control network further includes preference guidance parameters, which are used to constrain the action sequence generated by the control network, and which are used to characterize one or more requirements of speed, safety, and accuracy.

[0032] It is understandable that by setting preference guidance parameters, the generated action sequences can be specifically designed to meet the user's specific needs in a certain aspect.

[0033] In conjunction with the first aspect, or any of the above possible implementations of the first aspect, in yet another possible implementation of the first aspect, the control network includes a diffusion control network.

[0034] In conjunction with the first aspect, or any of the above possible implementations of the first aspect, in yet another possible implementation of the first aspect, the first information includes one or more of the following: image, video, point cloud, speech, heatmap, and depth map.

[0035] Secondly, embodiments of this application provide a control device, the device including units for implementing all or part of the steps in the method described in the first aspect or any possible implementation of the first aspect.

[0036] Thirdly, embodiments of this application provide a control device, which includes a processor and a memory, wherein the memory is used to store a computer program, and the processor calls the computer program to implement the method described in the first aspect or any possible implementation of the first aspect.

[0037] Fourthly, embodiments of this application provide a computer-readable storage medium for storing a computer program that, when run on a processor, implements the method described in the first aspect or any possible implementation of the first aspect. Attached Figure Description

[0038] The accompanying drawings used in the embodiments of this application are described below.

[0039] Figure 1 This is a schematic diagram illustrating the principle of evolution from Internet AI to Embedded AI, provided in an embodiment of this application.

[0040] Figure 2 This is a schematic diagram of the architecture of a control system provided in an embodiment of this application;

[0041] Figure 3This is a flowchart illustrating a control method based on artificial intelligence (AI) provided in an embodiment of this application;

[0042] Figure 4 This is a flowchart illustrating a key information path generation method provided in an embodiment of this application;

[0043] Figure 5 This is a flowchart illustrating another method for generating key information paths provided in this application embodiment;

[0044] Figure 6 This is a flowchart illustrating another method for generating key information paths provided in this application embodiment;

[0045] Figure 7 This is a flowchart illustrating another method for generating key information paths provided in this application embodiment;

[0046] Figure 8 This is a schematic diagram of the processing flow of a diffusion control network based on an embodied intelligent agent provided in an embodiment of this application;

[0047] Figure 9 This is a schematic diagram of the processing flow of another diffusion control network based on an embodied intelligent agent provided in the embodiments of this application;

[0048] Figure 10 This is a schematic diagram of the processing flow of another diffusion control network based on an embodied intelligent agent provided in the embodiments of this application;

[0049] Figure 11 This is a schematic diagram of the structure of a control device provided in an embodiment of this application;

[0050] Figure 12 This is a schematic diagram of the structure of a control device provided in an embodiment of this application. Detailed Implementation

[0051] The embodiments of this application are described below with reference to the accompanying drawings.

[0052] First, the relevant technologies of this application will be introduced.

[0053] (1) From Internet AI to Embodied AI

[0054] Internet AI primarily refers to the use of the massive amounts of data and computing resources available on the internet to train and deploy artificial intelligence (AI) models. These models typically rely on static datasets for learning and are built using methods such as supervised learning, unsupervised learning, and reinforcement learning; examples include large language models (LLMs).

[0055] Embodied AI (E-AI) is a novel approach in the field of AI that emphasizes the interaction between intelligent agents and their physical environment. E-AI posits that intelligent agents should be able to perceive their environment, take action, remember experiences, and learn from them to achieve true intelligence. The core elements of embodied AI include: 1. Perception—the agent can perceive its environment; 2. Autonomous Planning—the agent can plan the most rational path steps; 3. Decision Making—to solve the current step, the agent determines the best action by evaluating the potential consequences of different action plans; 4. Action—the agent can interact with and change the environment. Embodied AI typically has a "brain" responsible for task planning and reasoning. This part usually refers to the agent's high-level decision-making and planning capabilities, similar to the human brain. It handles complex problem-solving, abstract thinking, long-term planning, and reasoning. In embodied AI, this typically involves understanding the context of the task, setting goals, planning strategies to achieve those goals, and predicting the consequences of actions. This "brain" part enables the agent to search decision trees, use known information for logical reasoning, and generate action plans to achieve the goal while performing tasks. Typically, the cerebellum is responsible for perception and motor control. The cerebellum, similar to the human cerebellum, is responsible for the agent's perception and motor control. It involves acquiring data from sensors, processing this data to understand the environmental state, maintaining balance, and coordinating movement. In embodied intelligence, this includes real-time perception of the environment, dynamic planning, reaction control, and fine motor execution. The cerebellum ensures that the agent can accurately execute actions planned by the brain and react quickly to unexpected situations.

[0056] Figure 1 This illustrates the evolution from Internet AI to Embodied AI.

[0057] (2) Control Strategy

[0058] Model Predictive Control (MPC) is a model-based control method that optimizes current control decisions by predicting the system's behavior over a future period. MPC is suitable for systems with complex dynamics and constraints, such as chemical processes, power systems, and autonomous vehicles.

[0059] Imitation learning is a method of training an agent by imitating the behavior of experts. It is often used for tasks that are difficult to solve directly through reinforcement learning, such as scenarios that require precise imitation of human behavior.

[0060] Reinforcement learning (RL) is a method of learning optimal policies through interaction with the environment. The agent learns through trial and error to maximize cumulative rewards. In embodied intelligence, RL can be used to train robots to perform complex tasks such as walking, grasping, and manipulation.

[0061] (3) Generative decision-making model

[0062] Generative decision models are models capable of generating new data or decision sequences, typically used to predict future behaviors or states. In embodied intelligence, generative decision models can be used to generate action sequences for agents, predict environmental changes, and optimize decision-making strategies. Available models include Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), sequence generation models (such as Transformers), Diffusion Models, and Generative Flow Networks (GFlowNets).

[0063] See Figure 2 , Figure 2 This is a schematic diagram of the architecture of a control system provided in an embodiment of this application. The control system includes an intelligent agent 201 and a control network 202. The control system may include a single device or multiple devices (e.g., a device cluster composed of multiple devices). The control system may be deployed locally or in the cloud; this embodiment of the application does not impose any limitations. The control system is used to control a first object, which may be a device or tool on the control system, an external device or tool attached to the control system, or a product independent of the control system. The first object is controlled by the control system and completes corresponding tasks under the control of the control system to achieve a predetermined goal.

[0064] The first object can be a robotic arm, robot, workshop equipment, smart home equipment (e.g., refrigerator, television, air conditioner, electricity meter, etc.), vehicle equipment (e.g., car, bicycle, electric vehicle, airplane, ship, etc.), handheld device (e.g., mobile phone, tablet computer, PDA, etc.), or other controllable equipment or devices, etc., and the embodiments of this application do not limit it.

[0065] Agent 201 can be an agent of embodied intelligence. Embodied intelligence (E-AI) has been introduced previously and will not be repeated here.

[0066] The control network 202 can be a diffusion network (also known as a diffusion model), a generative adversarial network (GAN), etc. Figure 2 The control network 202 is illustrated as a diffusion network.

[0067] The control network (such as a Diffusion network) 202 based on an agent 201 provided in this application embodiment is a control scheme that combines deep learning, multimodal perception (optional), and autonomous decision-making, and may include one or more of the following components:

[0068] Perception and Key Information Extraction: Real-time perceived data is processed (e.g., through models such as the Dual-Stage Implicit Object-Oriented Network (DINOv2) or the Visual Language Model (VLM)) to extract key information and object feature information from the initial information (e.g., the initial image). Then, an embodied intelligence agent is used to label the key information in the initial information; for example, if the key information is key points, then the key points may include the object's edges, center, or specific recognition features.

[0069] Path planning: This involves using an embodied intelligence agent to plan the path for critical information. Constraints during path planning can include avoiding obstacles, optimizing action sequences to improve efficiency, and reducing energy consumption, and the specific constraints are set according to the actual application scenario.

[0070] Input to the control network: The feature and key information from the initial information are used as input to the control network (such as a Diffusion model). This control network can be a generative model that generates data through progressive denoising to produce action sequences.

[0071] Action sequence generation: The control network generates a series of possible action sequences based on the current state (also called the initial state, such as the 3D initial coordinates) and the target task (also called the target final state, such as the 3D end coordinates). The control network evaluates the potential effect of each action sequence and selects the optimal action sequence for output.

[0072] Execution and Feedback: After selecting the optimal action sequence, it is passed to the first object (such as a robotic arm) for execution, and feedback information is collected during the execution. Execution is carried out in M ​​steps, with the strategy adjusted at each step based on the execution results and environmental feedback, until the target task is completed.

[0073] Autonomous evolution: Over time, control networks can optimize their performance through autonomous evolution, improving the accuracy and efficiency of task completion through continuous learning and optimization.

[0074] Such agent-based control networks (such as Diffusion networks) enable primary objects (such as robots) to make effective decisions and take actions in complex environments, improving their performance in embodied intelligence tasks. Optionally, by combining visual language models and Diffusion networks, this control network can achieve highly autonomous and adaptive intelligent behavior across a variety of tasks.

[0075] Figure 2 This illustration uses a diffusion network as an example of a control system. Figure 2 As shown, the input to the diffusion network includes initial states and target final states. In this embodiment, the key information path may include multiple path segments. The key information at the start position of each path segment is used as the initial states of the control network, and the key information at the end position of each path segment is used as the target final states of the diffusion network. Optionally, the observation part of the diffusion network processes the initial states and target final states of the input, and then outputs the processing result to a real-time diffusion controller with guidance gradient. Under the constraints of a collision-free cost function (or possibly other functions), an action sequence is generated.

[0076] The generation of the action sequence may require K iterations. Optionally, the iteration process can incorporate the following loss function: max θ E τ~D [logp θ [x0(τ)|y(τ))], where τ is the key information path, D is the sampling data distribution of the key information path, x0(τ) is the parameter about the key information path, y(τ) is the condition that needs to be achieved for reasoning based on the key information path, p is the probability function, and θ is the neural network parameter.

[0077] Then, the first object (such as a robotic arm) executes the sequence of actions to complete the target task. Optionally, this sequence of actions can be fed back to the observation section to optimize its performance. The execution principle of the above control system will be discussed below in conjunction with... Figure 3 The method embodiments shown will be described in detail.

[0078] Please see Figure 3 , Figure 3 This is a flowchart illustrating a control method based on artificial intelligence (AI) provided in an embodiment of this application. The method can be implemented based on the mentioned control system and includes, but is not limited to, the following steps:

[0079] Step S301: Input multiple key information and target tasks from the first information into the embodied intelligent agent to generate key information paths.

[0080] The first information is used to describe (or reflect, or characterize) the state of the work scene. For example, the first information includes one or more of the following: image (e.g., 2D image), video, point cloud (e.g., 3D point cloud), language, heat map, and depth map. For instance, when the first information includes an image, the image can be an image of the work scene captured by an image sensor (e.g., camera, depth camera). In this work scene, one or more of the following can be included: a first object (e.g., a robotic arm), a second object that the first object needs to operate (e.g., clothing), and the environment in which the first and second objects are located. Therefore, the image (i.e., the first information) includes these information.

[0081] Correspondingly, these multiple key pieces of information are the more important information within the first set of information. This key information can be obtained through manual annotation, algorithmic extraction, or other methods. Optionally, this key information can be key points. After obtaining the key information, the multiple key pieces of information and the target task are input into the embodied intelligent agent to generate a key information path. The specific nature of the target task is not limited here; for example, it could be "Please help me design a path for folding clothes," "Suturing a wound," "Processing a round cup," etc. The key information path is used to characterize the process of completing the target task in the work scenario, and the key information represents the key nodes in this process.

[0082] The path mentioned in the embodiments of this application can also be referred to as a trajectory.

[0083] To facilitate understanding, the following lists several methods for obtaining key information from the first piece of information, as well as the paths to obtain this key information:

[0084] Method 1: Obtain the feature information of the first information, wherein the feature information includes two-dimensional feature information (such as features in a 2D image) and / or three-dimensional point cloud information (such as features in a 3D point cloud); then, analyze the feature information using a visual model to obtain multiple key pieces of information from the first information, in order to... Figure 4 Taking the illustrated process as an example, the first information is a 2D image and / or a 3D point cloud. The 2D image is used to find the feature information of the closest image through the RAG model and input it into the visual model for processing. The 3D point cloud is used to input the information obtained after passing through the point encoder into the visual model for processing. The visual model analyzes the input feature information to obtain key information (such as key points or function detection). Then, the key information and the target task (such as a language task or instruction) are input into the visual language model. Under the constraint of the target task, a key information path (such as a key point path) is generated based on the key information.

[0085] Optionally, the visual model may include a Segment Anything Model (SAM) and GraspNet. The SAM segments the first information, generating a segmentation mask based on given cues (such as text descriptions or specific points in an image). GraspNet is a convolutional neural network for grasping detection, capable of detecting grasping poses in images in real time. The segmentation mask is input into GraspNet to detect key information in the first information. In one case, GraspNet not only generates key information but also key information paths. For example, the key information generated by GraspNet includes an initial 6D grasping pose and a target 6D grasping pose. These initial and target 6D grasping poses are the two key pieces of information. The initial 6D grasping pose corresponds to the starting state, and the target 6D grasping pose corresponds to the ending state, thus implicitly containing the concept of a path. In this case, the visual model essentially performs both... Figure 4 The generation of key information also includes the generation of key information paths. In other words, the actions performed by the visual language model are completed by the visual model. Understandably, in this case, the target task can be used as the input of the visual model to constrain the generation of key information paths.

[0086] Optionally, the above segmentation mask process may not exist, meaning that the input when GraspNet generates key information does not include the segmentation mask but includes the first information; or, when a segmentation mask exists, the input when GraspNet generates key information may include the segmentation mask but not the first information; how to implement it can be set as needed, and is not limited here.

[0087] Method Two: Query second information from a first database that satisfies a preset similarity relationship with the first information. The first database stores multiple pieces of information and their corresponding key information. Then, determine multiple key information from the first information based on the key information in the second information. Since the key information in the second information is already stored in the first database, it can be directly retrieved. This key information can then be directly used as the key information in the first information, or it can be transformed or processed accordingly to obtain the key information in the first information. Optionally, the first and second information can be aligned first, and then the key information in the first information can be determined based on the key information in the second information.

[0088] Subsequently, the embodied intelligent agent generates a key information path based on multiple key pieces of information and the target task in the first information.

[0089] like Figure 5 The diagram illustrates the architecture of a vision agent based on RAG (Retrieval-Augmented Generation). Taking a folding clothes scenario as an example, the first piece of information is the first image, the second piece of information is the second image, multiple pieces of information are multiple images, and key information is key points. Before using the first database, the first stage (memory) can be completed. For example, this can be achieved using a vision agent based on Retrieval-Augmented Generation (RAG). After acquiring images through an image sensor, image features (such as visual features, category, location, orientation, etc.) and key information describing objects are extracted from the images using networks such as the Visual Language Model (VLM). The VLM model can encode the image content into a high-dimensional feature vector, providing a foundation for subsequent matching and retrieval. Then, the user can label key points (and possibly key point paths) based on the extracted image features and key information describing objects. These key points may include the boundaries, center points, or other significant features of objects in the image, while the key point paths describe the movement paths of objects or parts of objects over time. This approach allows for the acquisition of key points from multiple images. These images and their corresponding key points are then stored in a first database (i.e., a memory-evolved database) for retrieval and inference tasks. Over time, the Vision Agent can continuously learn and optimize the first database, improving the accuracy of matching image features through autonomous evolution.

[0090] After determining the key information in the first image, the second stage (reasoning) begins. This involves first extracting image features (such as visual features, category, location, orientation, etc.) and key object descriptions from the first image using networks like VLM. This activates the data stored in the first database (i.e., the memory evolution database). The Vision Agent uses this data to more accurately match new image features. Specifically, it matches the image features and key object descriptions of the first image with the image features (such as visual features, category, location, orientation, etc.) and key object descriptions of multiple images stored in the first database (e.g., retrieving the top N categories and image features). The image that meets the preset similarity criteria is then selected as the second image. Optionally, meeting the preset similarity criteria could mean the most similar image, or one of the top N most similar images that also meets other constraints (e.g., the same scene). The preset similarity criteria can be set as needed and are not limited here; they simply need to reflect a strong similarity between the two images in one or more dimensions.

[0091] Then, the first image is subjected to an affine transformation and visually aligned with the second image. Then, multiple key points in the second image are mapped onto the first image to obtain multiple key points of the first image.

[0092] This approach can generate more reasonable key information (and may also include the corresponding key information path), achieving fully autonomous generation and addressing the impact of overly detailed prompts and significant human bias on the generated results.

[0093] Next, the fourth model in the embodied intelligence agent (such as LLM GPT4o) generates keypoint paths based on multiple keypoints in the first image and the target task. For example, the target task is "Combining the given image path information of the folded shorts example ##, and the given shorts image and corresponding keypoint ## the object to be inferred ##, please predict the keypoint path of the folded shorts." Then, the fifth model in the embodied intelligence agent (such as the LLM summary model) summarizes the keypoint paths and outputs summary information, such as "Combining the above conclusions and the image estimation information ##, the most likely (unique) path for # the object to be inferred ## is given: Considering the example information of the folded shorts, the most reasonable path inference is from keypoint 5 to keypoint 6." It should be noted that the specific case mentioned here is only an example, and the actual scenario and result may be different.

[0094] Method 3: A first model predicts multiple key pieces of information from the first information. This first model is trained using multiple sets of first training data, including feature information of the third information and key information within the third information. Optionally, when predicting multiple key pieces of information, the input to the first model may include other information besides the first information, such as input constraint information. Additionally, a second model predicts the key information path composed of the multiple key pieces of information. This second model is trained using multiple sets of second training data, including key information in the fourth information and the corresponding key information path. Optionally, when predicting the key information path, the input to the second model may include other information besides the multiple key pieces of information, such as input constraint information.

[0095] like Figure 6 The diagram illustrates an embodied intelligent agent architecture based on VLM and DINOv2, using a clothes-folding scenario as an example. The first piece of information is the first image, the third is the third image, the fourth is the fourth image, and the key information is the key points. Before using the first and second models, they are trained. For example, the first stage (memory) involves performing the following operations: After acquiring images through an image sensor, image features (such as visual features, category, location, orientation, etc.) and key object descriptions are extracted using a visual language model (VLM) or similar network. The VLM model encodes image content into high-dimensional feature vectors, providing a foundation for subsequent matching and retrieval. Then, the user can label key points and key point paths based on the extracted image features and key object descriptions. These key points may include the boundaries, center points, or other significant features of objects in the image, while the key point paths describe the movement path of objects or parts of objects over time. Using this method, key points from multiple images can be obtained, and then these images, along with their corresponding key points and key point paths, are stored in a path dataset for model training. For example, a first model is trained using the third image and its corresponding keypoints from the path dataset, and a second model is trained using the fourth image, its corresponding keypoints, and the keypoint path from the same dataset. Optionally, the third and fourth images can be the same image or different images. Furthermore, training the first and second models may utilize other information, such as user input constraints.

[0096] After the first and second models are trained, image features (such as visual features, category, location, orientation, etc.) and key information describing objects in the first image can be extracted using networks such as the Visual Language Model (VLM). These are then input into the first model (such as the DINOv2 keypoint model) to obtain multiple keypoints in the first image. The input at this time may also include user-input constraints (e.g., user: "Please help me design a path for folding clothes"). After obtaining multiple keypoints, these keypoints are input into the second model (such as the path VLM model) to obtain the keypoint path (e.g., from keypoint 5 to keypoint 6). The input at this time may also include user-input constraints (e.g., user: "Given an image of shorts and the corresponding keypoint ##, infer the object's ##, please infer the keypoint path of the shorts").

[0097] In this embodiment, DINOv2 is a self-supervised visual Transformer model that can learn visual features from a large number of unlabeled images and simultaneously perform pre-training of a visual language model (VLM).

[0098] Method four involves predicting the key information path corresponding to the first information using a third model. This third model is trained using multiple third training data sets, including feature information of the fifth information (e.g., item feature information), key information within the fifth information (e.g., item key information), and the key information path corresponding to the fifth information. Optionally, other training data may be available during the training of the third model, such as user-inputted tasks (or instructions). Accordingly, when predicting the key information path using the third model, in addition to inputting the first information, it may also be necessary to input the user-inputted tasks (or instructions) into the third model.

[0099] like Figure 7As shown, this is a distributed real-time control architecture based on Vision-Agent embodied intelligence Diffusion. Taking a folding clothes scenario as an example, the first information is the first image, the fifth information is the fifth image, and the key information is the key points. Before using the third model, it is trained first. For example, the following operations are performed: after acquiring images through an image sensor, key points and image features (such as visual features, category, position, orientation, etc.) and object description key information are extracted from the images through networks such as the Visual Language Model (VLM). The key point information and image features are used to understand and identify important elements in the image, such as objects, scenes, or actions. Then, the user can label the key points and key point paths based on the extracted key points, image features, and object description key information. These key points may include the boundaries, center points, or other salient features of objects in the image, while the key point paths describe the movement path of objects or parts of objects over time. In this way, key points and corresponding key point paths of multiple images can be obtained. Then, multiple images and their corresponding key points and key point paths are stored in a path dataset for model pre-training. For example, a third model (which could be a Visual Modeling Library) can be pre-trained using the fifth image in the path dataset, along with its corresponding keypoints and keypoint paths. This pre-trained third model (e.g., a VLM) can better understand and process new, unseen image data. Pre-training is a common practice in deep learning, allowing the model to learn general feature representations before handling a specific task. During pre-training, the third model may be trained using a large amount of image data, which may include a wide variety of objects and scenes. In this way, the model can learn rich visual features, which will be used to match new image data in subsequent inference stages.

[0100] Optionally, training the third model may also utilize other information, such as user input constraints. This training process can be considered as training an end-to-end embodied VLM agent based on a dataset of keypoints and keypoint paths.

[0101] After the third model is trained, image features (such as visual features, category, position, orientation, etc.) and key information describing objects in the first image can be extracted using a visual language model (VLM) or similar network. This information is then input into the third model (such as the DINOv2 keypoint model) to obtain the keypoint path of the first image. Optionally, when predicting the keypoint path of the first image using the third model, a target task, such as the instruction "fold shorts," needs to be input as a constraint when generating the keypoint path. Alternatively, the key information path output from the third model can be converted into 6D grasping information, which includes the object's position and orientation. This 6D grasping information can then be used as the key information path required when the first object performs a grasping task. The 6D grasping information can then be used as input to a control network (such as a Diffusion network).

[0102] Step S302: Input the key information path into the control network to obtain the sequence of actions required to complete the target task.

[0103] Specifically, the control network is a network that can plan action sequences according to the path. For example, it can be a diffusion network, also known as a diffusion control network, diffusion model, or diffusion control model.

[0104] In this embodiment of the application, the key information path includes one or more path segments. The key information at the start position of each path segment is used as the initial state of the control network, and the key information at the end position of each path segment is used as the target final state (also called the final state or target state) of the control network. The control network is used to generate the corresponding action for each path segment with each path segment as input. The actions corresponding to the multiple path segments are used to form the action sequence that needs to be executed to complete the target task.

[0105] In this embodiment of the application, the key information can be key points, which can be 3D coordinates or 6D coordinates, such as... Figure 4 As shown, in a path segment, a keypoint represented by the 3DOF XYZ position state is used as the initial state, and a keypoint represented by the 3DOF target is used as the target's final state. The initial state and the target's final state are used together as the input to the diffusion network. Alternatively, in a path segment, a keypoint represented by the 6DOF XYZ+Euler angle state may be used as the initial state, and a keypoint represented by the 6DOF target may be used as the target's final state. The initial state and the target's final state are used together as the input to the diffusion network.

[0106] For example, the observation part of a diffusion network processes the initial input states and the target final states, then outputs the processing result to a real-time diffusion controller with guidance gradients. Under the constraints of a collision-free cost function (or possibly other functions), it generates an action sequence. The real-time diffusion controller with guidance gradients can include Feature-wise Linear Modulation (FiLM), Convolutional Neural Networks (CNNs), and Transformers. FiLM is a mechanism to enhance the representational power of neural networks by dynamically adjusting the activation functions of network layers to improve the network's response to input data. CNNs and Transformers are two commonly used deep learning models that can be used to process visual and sequential data. After denoising using FiLM, CNNs, and Transformers, the output is an action sequence that performs the target task. Combining FiLM, CNN, and Transformer can improve the performance of a primary object (such as a robot) when performing a target task. Optionally, information related to the target task can also be used as input parameters for the diffusion network to generate action sequences.

[0107] Optionally, this action sequence can be fed back to the observation part to optimize its performance. Specifically, K iterations can be set to achieve the optimal state for the action sequence; for example, by introducing the gradient descent formula:

[0108]

[0109] Where x is the parameter related to the first object (e.g., a robotic arm) after gradient descent (which may include the key information path mentioned earlier), γ is the parameter related to the first object (e.g., a robotic arm) before gradient descent (which may include the key information path mentioned earlier), and γ is the learning rate. This involves finding the gradient, where E(x) is the loss function. A more ideal action sequence can be obtained using the gradient descent formula.

[0110] Optional, such as Figure 8 The diagram illustrates a processing flow of a diffusion control network based on an embodied intelligent agent, including a key point path (a type of key information path) generation process 801, an action sequence generation process 802, and a feedback optimization network 803, wherein:

[0111] The keypoint path generation process 801 includes: acquiring 2D images and depth information of the work scene using a camera; inputting the 2D images and depth information into a point cloud generation network to obtain a 3D point cloud; inputting the 2D images and text-based instructions (or tasks) into an intelligent agent network, which, together with an obstacle avoidance perception network, can generate keypoints in the 2D images, where the obstacle avoidance perception network's input includes information about the work scene provided by the camera; fusing the keypoint information in the 2D images with the preceding 3D point cloud to obtain 3D keypoint information, such as 3D keypoint coordinates; inputting the 2D images and their keypoints into a path decision agent network to obtain the order of these keypoints; with the 3D coordinates and order of the keypoints, the keypoint path of these 3D keypoints can be determined. Subsequently, the 3D keypoint coordinates involved in the 3D keypoint path are processed by a 3D-to-6D (i.e., 3Dof -> 6Dof) generation network to obtain 6D keypoint coordinates, with each path segment corresponding to a 6D start coordinate and a 6D end coordinate. Optionally, the 6D keypoint coordinates in this embodiment can be represented as (x,y,z)+(qx,qy,qz,qw), where (x,y,z) corresponds to the position coordinates and (qx,qy,qz,qw) corresponds to the rotation angle.

[0112] In the action sequence generation process 802: the 6D starting coordinates are used as the initial state of the diffusion network (i.e., the 6Dof diffusion control network), and the 6D ending coordinates are used as the target final state of the diffusion network. After processing by the diffusion network, the corresponding action sequence can be obtained.

[0113] The feedback optimization network 803 includes: substituting the action sequence into the control code of the first object (such as a robotic arm), executing the code to control the movement of the robotic arm, and then using a perception network to sense whether the movement of the robotic arm has successfully achieved the expected goal (which can be set as needed or generated or judged by the network itself). If the expected goal is successfully achieved, the next step is executed, and the next action sequence is substituted into the robotic arm control code to continue controlling the movement of the robotic arm. If the expected goal is not achieved, the difference (or loss) between the effect of the robotic arm movement and the expected goal is determined by an error estimation network. This difference can be represented by an effect parameter. Therefore, the effect parameter is used to characterize the effect of completing the target task according to the action sequence. The effect parameter is used to calibrate the key information generated by the embodied intelligent agent in the next iteration; for example, updating the key information generated in the next iteration (e.g., updating the 6D coordinates of key points) based on the effect parameter.

[0114] Optional, such as Figure 9 The diagram illustrates another processing flow of a diffusion control network based on an embodied intelligent agent, including a key point path (a type of key information path) generation process 801, an action sequence generation process 802, a feedback optimization network 803, and a distance calculation network 804. The key point path generation process 801, action sequence generation process 802, and feedback optimization network 803 have been previously described. The design principle of the distance calculation network 804 is as follows: The control system needs to control a first object (such as a robotic arm) to execute an action sequence. Therefore, the distance between the first object (such as the robotic arm) and the key points also affects the generation of the action sequence. For example, the closer the distance, the slower the action sequence can be executed; the farther the distance, the faster the action sequence can be executed. Therefore, the state of the first object (such as the robotic arm) is obtained, and the distance between the first object (such as the robotic arm) and the key points is determined based on this state. This distance information is then input into the diffusion network as one of the input parameters for generating the action sequence, thus resulting in a more accurate generated action sequence.

[0115] Optional, such as Figure 10The diagram illustrates another processing flow of a diffusion control network based on an embodied intelligent agent, including a key point path (a type of key information path) generation process 801, an action sequence generation process 802, a feedback optimization network 803, a distance calculation network 804, and a preference guidance network 805. The key point path generation process 801, action sequence generation process 802, feedback optimization network 803, and distance calculation network 804 have been previously described. The design principle of the preference guidance network 805 is as follows: When generating action sequences, there may be corresponding requirements in terms of speed, safety, accuracy, and generalization. Therefore, a preference guidance network can be designed to meet these requirements. This preference guidance network is configured with parameters that constrain speed, safety, accuracy, and generalization. Under the guidance of this preference guidance network 805, action sequences that meet specific requirements can be generated.

[0116] Optional, such as Figure 7 As shown, the input to the diffusion control network can also be adjusted. Some of the previously mentioned implementations use the key information (such as key points) at the beginning of the path segment as the initial state of the diffusion network and the key information (such as key points) at the end of the path segment as the target final state of the diffusion network. In another alternative, the key information (such as key points) at the beginning of the path segment can be used as the initial state of the diffusion network, and one piece of information or one point between the key information at the beginning and the key information at the end of the path segment can be used as the target final state of the diffusion network.

[0117] Optionally, for each path segment described above, the control network (such as a Diffusion network) may output a series of possible action sequences, and then select the optimal action sequence from these sequences for execution according to a corresponding strategy. This can be done by evaluating the potential effect of each action sequence and its suitability for the target task. The selected optimal action sequence is used for execution by the first object (such as a robotic arm).

[0118] Optionally, deep learning, based on control networks (such as diffusion networks), can reduce model complexity and enable the model to learn from few samples through policy integration. By providing a complete target distribution, the model can achieve generalization and multiple capabilities. By learning and generating new action spaces, it can solve the problems of interacting with unfamiliar environments and autonomously evolving new capabilities.

[0119] Step S303: Control the first object in the work scenario to execute the action sequence.

[0120] In one alternative approach, after obtaining the action sequence through the preceding steps, the control system controls the first object in the work scenario (such as a robotic arm) to execute the action sequence. This process may be performed M times until the target task is finally completed, where M is a positive integer.

[0121] In another alternative approach, after the control system obtains the action sequence through the preceding steps, it sends the action sequence to other devices, which then control the first object (such as a robotic arm) to execute the action sequence. This process may be repeated M times until the target task is finally completed.

[0122] Optionally, as mentioned earlier, the critical information path includes multiple path segments. An action sequence is generated for each path segment. Generally, the target task is only completed after all the action sequences corresponding to each path segment are executed. For example, if the critical information path includes M path segments, then M expected action sequences will be obtained. Executing these M expected action sequences completes the target task. Optionally, after each action sequence is executed, the execution result and environmental information are fed back to the input of the control network, or used to optimize input parameters (such as the critical information path) to improve the effect of the next generated action sequence, thus better achieving the target task.

[0123] exist Figure 3 In the described method, before predicting the action sequence for completing the target task using a control network (such as a diffusion network), an embodied intelligence agent first analyzes the initial information of the work scenario where the target task is located to obtain the key information path within the initial information. This key information path is then input into the control network for processing to generate the action sequence. Because the embodied intelligence agent performs preprocessing, meaning the input to the control network is processed information of higher quality, the action sequence predicted by the control network performs better in completing the target task. In other words, if the control network lacks prior knowledge, has insufficient understanding and generalization ability, and cannot possess common sense, human-like perception, and actions, the introduction of an embodied intelligence agent effectively compensates for this problem. Furthermore, if the embodied intelligence agent cannot perform functions such as few-shot learning, multi-capability coverage, autonomous interaction with the environment, and autonomous evolutionary learning, then the control network (such as a diffusion network) can compensate for these shortcomings.

[0124] The methods of the embodiments of this application have been described in detail above, and the apparatus of the embodiments of this application is provided below.

[0125] Please see Figure 11 , Figure 11This is a schematic diagram of the structure of a control device provided in an embodiment of this application. The control device can be the control system mentioned above or a component or module in the control system. The software module here can be in hardware or software form. The control device 110 can include a first generation unit 1101 and a second generation unit 1102. The detailed description of each unit is as follows.

[0126] The first generation unit 1101 is used to input multiple key information and target task from the first information into the embodied intelligent agent to generate a key information path, wherein the first information is used to describe the state of the work scenario, and the key information path is used to characterize the process of completing the target task in the work scenario.

[0127] The second generation unit 1102 is used to input the key information path into the control network to obtain the action sequence required to complete the target task.

[0128] In one alternative implementation, the control device 110 further includes:

[0129] A control unit is used to control the first object in the work scenario to execute the action sequence.

[0130] In another alternative implementation, the key information path includes one or more path segments. The key information at the start position of each path segment is used as the initial state of the control network, and the key information at the end position of each path segment is used as the final state of the control network. The control network is used to generate a corresponding action for each path segment as input, and the actions corresponding to the multiple path segments are used to form a sequence of actions to be executed to complete the target task.

[0131] In another alternative implementation, the control device 110 further includes:

[0132] The first acquisition unit is used to acquire feature information of the first information, wherein the feature information includes two-dimensional feature information and / or three-dimensional point cloud information;

[0133] The third generation unit is used to analyze the feature information through a visual model to obtain multiple key pieces of information from the first information.

[0134] In another alternative implementation, the control device 110 further includes:

[0135] The query unit is used to query second information that satisfies a preset similarity relationship with the first information from the first database, wherein the first database stores multiple pieces of information and key information corresponding to each of the multiple pieces of information;

[0136] A determining unit is configured to determine multiple key pieces of information in the first information based on multiple key pieces of information in the second information.

[0137] In another alternative implementation, regarding the step of inputting multiple key pieces of information and the target task from the first information into the embodied intelligent agent to generate a key information path, the first generation unit 1101 is specifically used for:

[0138] The first model predicts multiple key pieces of information in the first information, wherein the first model is trained based on multiple first training data, and the first training data includes feature information of the third information and key information in the third information.

[0139] The second model predicts the key information path composed of the multiple key information items, wherein the second model is trained based on multiple second training data, and the second training data includes the key information in the fourth information and the key information path corresponding to the fourth information.

[0140] In another alternative implementation, regarding the step of inputting multiple key pieces of information and the target task from the first information into the embodied intelligent agent to generate a key information path, the first generation unit 1101 is specifically used for:

[0141] The key information path corresponding to the first information is predicted by the third model, wherein the third model is trained based on multiple third training data, including the feature information of the fifth information, the key information in the fifth information, and the key information path corresponding to the fifth information.

[0142] In another alternative implementation, the control device 110 further includes:

[0143] The second acquisition unit is used to acquire effect parameters, wherein the effect parameters are used to characterize the effect of completing the target task according to the action sequence, and the effect parameters are used to calibrate the key information generated by the embodied intelligent agent next time.

[0144] In another alternative implementation, the input to the control network further includes the state of a first object that performs the operation sequence, the state of the first object constraining the action sequence generated by the control network.

[0145] In another alternative implementation, the input to the control network further includes preference guidance parameters, which are used to constrain the action sequence generated by the control network and represent one or more requirements of speed, safety, and accuracy.

[0146] In another alternative implementation, the control network includes a diffusion control network.

[0147] In another alternative implementation, the first information includes one or more of the following: image, video, point cloud, speech, heatmap, and depth map.

[0148] It should be noted that the implementation and beneficial effects of each unit can be found by referring to [the relevant documentation / reference]. Figure 3 The corresponding description of the method embodiments shown.

[0149] Please see Figure 12 , Figure 12 This is a schematic diagram of the structure of a control device 120 provided in an embodiment of this application. The control device can be the control system mentioned above or a component or module in the control system. The control device 120 includes a processor 1201 and a memory 1202, and optionally may also include a communication interface 1203. The processor 1201, memory 1202 and communication interface 1203 are interconnected through a bus.

[0150] The memory 1202 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), and is used for related computer programs and data. The communication interface 1203 is used for receiving and sending data.

[0151] Processor 1201 can be one or more central processing units (CPUs). If processor 1201 is a CPU, the CPU can be a single-core CPU or a multi-core CPU. Processor 1201 can also be other forms of processing units with computing capabilities, such as graphics processing units (GPUs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), etc.

[0152] The processor 1201 in the control device 120 is used to read the computer program code stored in the memory 1202 and perform the following operations:

[0153] Multiple key pieces of information and target tasks from the first information are input into the embodied intelligent agent to generate a key information path, wherein the first information is used to describe the state of the work scenario, and the key information path is used to characterize the process of completing the target task in the work scenario;

[0154] The key information path is input into the control network to obtain the sequence of actions required to complete the target task.

[0155] In one possible implementation, the processor is further configured to:

[0156] Control the first object in the work scenario to execute the action sequence.

[0157] In another possible implementation, the key information path includes one or more path segments, where the key information at the start position of each path segment is used as the initial state of the control network, and the key information at the end position of each path segment is used as the final state of the control network; the control network is used to generate a corresponding action for each path segment as input, and the actions corresponding to the multiple path segments are used to form a sequence of actions to be executed to complete the target task.

[0158] In yet another possible implementation, the processor is further configured to:

[0159] Obtain feature information of the first information, wherein the feature information includes two-dimensional feature information and / or three-dimensional point cloud information;

[0160] The feature information is analyzed using a visual model to obtain several key pieces of information from the first information.

[0161] In yet another possible implementation, the processor is further configured to:

[0162] The system queries a first database for second information that has a preset similarity relationship with the first information, wherein the first database stores multiple pieces of information and key information corresponding to each of the multiple pieces of information.

[0163] Multiple key pieces of information in the first information are determined based on multiple key pieces of information in the second information.

[0164] In another possible implementation, regarding the step of inputting multiple key pieces of information and the target task from the first information into the embodied intelligent agent to generate a key information path, the processor is specifically used for:

[0165] The first model predicts multiple key pieces of information in the first information, wherein the first model is trained based on multiple first training data, and the first training data includes feature information of the third information and key information in the third information.

[0166] The second model predicts the key information path composed of the multiple key information items, wherein the second model is trained based on multiple second training data, and the second training data includes the key information in the fourth information and the key information path corresponding to the fourth information.

[0167] In another possible implementation, regarding the step of inputting multiple key pieces of information and the target task from the first information into the embodied intelligent agent to generate a key information path, the processor is specifically used for:

[0168] The key information path corresponding to the first information is predicted by the third model, wherein the third model is trained based on multiple third training data, including the feature information of the fifth information, the key information in the fifth information, and the key information path corresponding to the fifth information.

[0169] In yet another possible implementation, the processor is further configured to:

[0170] Obtain effect parameters, wherein the effect parameters are used to characterize the effect of completing the target task according to the action sequence, and the effect parameters are used to calibrate the key information generated by the embodied intelligent agent next time.

[0171] In another possible implementation, the input to the control network further includes the state of a first object that performs the operation sequence, the state of the first object constraining the action sequence generated by the control network.

[0172] In another possible implementation, the input to the control network further includes preference guidance parameters, which are used to constrain the action sequence generated by the control network and represent one or more requirements of speed, safety, and accuracy.

[0173] In yet another possible implementation, the control network includes a diffusion control network.

[0174] In another possible implementation, the first information includes one or more of the following: image, video, point cloud, speech, heatmap, and depth map.

[0175] It should be noted that the implementation and beneficial effects of each operation can be found by referring to [the relevant documentation / reference]. Figure 3 The corresponding description of the method embodiments shown.

[0176] This application embodiment also provides a chip system, the chip system including at least one processor, a memory, and interface circuitry, the memory, the transceiver, and the at least one processor being interconnected via circuitry, the at least one memory storing a computer program; when the computer program is executed by the processor, it implements... Figure 3 The method flow is shown.

[0177] This application also provides a computer-readable storage medium storing a computer program that, when run on a processor, implements... Figure 3 The method flow is shown.

[0178] This application also provides a computer program product that, when run on a processor, implements... Figure 3 The method flow is shown.

[0179] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program using computer program-related hardware. The computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing computer program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A control method based on artificial intelligence (AI), characterized in that, include: Multiple key pieces of information and target tasks from the first information are input into the embodied intelligent agent to generate a key information path, wherein the first information is used to describe the state of the work scenario, and the key information path is used to characterize the process of completing the target task in the work scenario; The key information path is input into the control network to obtain the sequence of actions required to complete the target task.

2. The method according to claim 1, characterized in that, Also includes: Control the first object in the work scenario to execute the action sequence.

3. The method according to claim 1 or 2, characterized in that, The key information path includes one or more path segments. The key information at the start position of each path segment is used as the initial state of the control network, and the key information at the end position of each path segment is used as the final state of the control network. The control network is used to generate a corresponding action for each path segment as input, and the actions corresponding to the multiple path segments are used to form a sequence of actions to be executed to complete the target task.

4. The method according to any one of claims 1-3, characterized in that, Also includes: Obtain feature information of the first information, wherein the feature information includes two-dimensional feature information and / or three-dimensional point cloud information; The feature information is analyzed using a visual model to obtain several key pieces of information from the first information.

5. The method according to any one of claims 1-3, characterized in that, Also includes: The system queries a first database for second information that has a preset similarity relationship with the first information, wherein the first database stores multiple pieces of information and key information corresponding to each of the multiple pieces of information. Multiple key pieces of information in the first information are determined based on multiple key pieces of information in the second information.

6. The method according to any one of claims 1-3, characterized in that, The step of inputting multiple key pieces of information and the target task from the first information into the embodied intelligent agent to generate a key information path includes: The first model predicts multiple key pieces of information in the first information, wherein the first model is trained based on multiple first training data, and the first training data includes feature information of the third information and key information in the third information. The second model predicts the key information path composed of the multiple key information items, wherein the second model is trained based on multiple second training data, and the second training data includes the key information in the fourth information and the key information path corresponding to the fourth information.

7. The method according to any one of claims 1-3, characterized in that, The step of inputting multiple key pieces of information and the target task from the first information into the embodied intelligent agent to generate a key information path includes: The key information path corresponding to the first information is predicted by the third model, wherein the third model is trained based on multiple third training data, including the feature information of the fifth information, the key information in the fifth information, and the key information path corresponding to the fifth information.

8. The method according to any one of claims 1-7, characterized in that, Also includes: Obtain effect parameters, wherein the effect parameters are used to characterize the effect of completing the target task according to the action sequence, and the effect parameters are used to calibrate the key information generated by the embodied intelligent agent next time.

9. The method according to any one of claims 1-8, characterized in that, The input to the control network also includes the state of a first object that executes the operation sequence, the state of the first object being used to constrain the action sequence generated by the control network.

10. The method according to any one of claims 1-9, characterized in that, The input to the control network also includes preference guidance parameters, which are used to constrain the action sequence generated by the control network and represent one or more requirements of speed, safety, and accuracy.

11. The method according to any one of claims 1-10, characterized in that, The control network includes a diffusion control network.

12. The method according to any one of claims 1-11, characterized in that, The first information includes one or more of the following: image, video, point cloud, speech, heat map, and depth map.

13. A control device, characterized in that, The apparatus includes units for implementing all or part of the steps in the method according to any one of claims 1-12.

14. A control device, characterized in that, The method includes a processor and a memory, wherein the memory is used to store a computer program, and the processor invokes the computer program to implement the method according to any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program that, when run on a processor, implements the method according to any one of claims 1-12.