Intelligent agent control method and device, equipment and storage medium
By constructing a two-layer intelligent agent control method and combining environmental characteristics and user intentions to generate control signals, the problem of uncontrollable end-to-end model behavior is solved, and the stability and controllability of the intelligent agent's actions are achieved.
Patent Information
- Application Number
- CN202510656786.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-19
AI Technical Summary
Existing end-to-end deep learning models lack explicit constraints in the process of intelligent agent action generation, resulting in uncontrollable behavior, especially in dynamic environments where dangerous actions are prone to occur.
A two-layer architecture is constructed with a separation of the basic capability layer and the user command control layer. The first action sequence of the intelligent agent is generated through environmental feature data, and the control signal is generated in combination with the user command intention information to regulate the action sequence to achieve stability and controllability.
The stability of the agent's actions and the controllability of its behavior are improved, ensuring the safety and reliability of the agent's behavior in dynamic environments.
Smart Images

Figure CN120669576A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent body technology, and in particular to a control method, device, equipment and storage medium for an intelligent body. Background Art
[0002] With the rapid development of artificial intelligence (AI) in areas such as robotic control, autonomous driving, and human-computer interaction, intelligent agents such as robots, robotic arms, end effectors, intelligent vehicle terminals, and human-computer interaction systems are becoming increasingly prevalent. To achieve the core goal of AI technology—that is, intelligent decision-making and behavior generation for agents in dynamic environments—related technologies typically rely on end-to-end deep learning models, training them using large-scale datasets to directly map environmental inputs to action outputs. However, because end-to-end models rely on black-box mapping and lack explicit constraints on the action generation process, they result in poor model output stability, making the agents prone to uncontrollable behavior. Summary of the Invention
[0003] The embodiments of the present application propose a control method, device, equipment and storage medium for an intelligent agent, which is beneficial for not only considering the constraints of the environment on the intelligent agent's actions during the action generation process of the intelligent agent, but also integrating the explicit constraints of the user's intentions on the intelligent agent's actions, thereby improving the stability of the intelligent agent's actions and further enhancing the controllability of the intelligent agent's behavior.
[0004] In a first aspect, an embodiment of the present application provides a method for controlling an intelligent agent, the method comprising:
[0005] generating a first action sequence of the agent based on first environmental feature data of the environment in which the agent is located;
[0006] In response to a received user instruction, performing intent analysis on the user instruction to generate instruction intent information;
[0007] generating a control signal based on the instruction intention information, the first environmental feature data, and the first action sequence;
[0008] regulating the first action sequence based on the control signal to generate a second action sequence of the agent;
[0009] Based on the second action sequence, the agent is controlled to perform an action.
[0010] In a second aspect, an embodiment of the present application provides a control device for an intelligent agent, comprising:
[0011] an action sequence generating module, configured to generate a first action sequence of the agent based on first environmental feature data of the environment in which the agent is located;
[0012] An intent parsing module, configured to, in response to a received user instruction, perform intent parsing on the user instruction and generate instruction intent information;
[0013] a control signal generating module, configured to generate a control signal based on the instruction intention information, the first environmental feature data, and the first action sequence;
[0014] a control module, configured to control the first action sequence based on the control signal to generate a second action sequence for the agent;
[0015] An action execution module is used to control the agent to perform actions based on the second action sequence.
[0016] In a third aspect, an embodiment of the present application provides an intelligent body, comprising: one or more processors; a memory on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the control method of the intelligent body as described in the first aspect.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the control method of the intelligent agent as described in the first aspect.
[0018] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the control method of the intelligent agent as described in the first aspect.
[0019] In an embodiment of the present application, a first action sequence of the intelligent agent is generated based on first environmental feature data of the environment in which the intelligent agent is located; in response to a received user instruction, the user instruction is parsed for intent to generate instruction intention information; a control signal is generated based on the instruction intention information, the first environmental feature data and the first action sequence; the first action sequence is regulated based on the control signal to generate a second action sequence of the intelligent agent; and the intelligent agent is controlled to perform actions based on the second action sequence. In this way, by constructing a two-layer architecture including a basic capability layer and a user instruction control layer, that is, the basic capability layer generates the first action sequence based on the first environmental feature data, and the user instruction control layer parses the user intention information through the user instruction, and generates a control signal based on the user intention information, the first environmental feature data and the first action sequence, and controls the first action sequence through the control signal, so that in the process of generating the action of the intelligent agent, not only the constraints of the environment on the action of the intelligent agent are considered, but also the explicit constraints of the user intention on the action of the intelligent agent are integrated, thereby improving the stability of the action of the intelligent agent and thus enhancing the controllability of the behavior of the intelligent agent. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1A schematic flow chart of an embodiment of the control method of an intelligent agent provided by the present application;
[0021] Figure 2 This is another flowchart of an embodiment of the method for controlling an intelligent body provided by the present application;
[0022] Figure 3 This is another flowchart of an embodiment of the method for controlling an intelligent body provided by the present application;
[0023] Figure 4 This is another flowchart of an embodiment of the method for controlling an intelligent body provided by the present application;
[0024] Figure 5 This is another flowchart of an embodiment of the method for controlling an intelligent body provided by the present application;
[0025] Figure 6 This is another flowchart of an embodiment of the method for controlling an intelligent body provided by the present application;
[0026] Figure 7 This is another flowchart of an embodiment of the method for controlling an intelligent body provided by the present application;
[0027] Figure 8 This is another flowchart of an embodiment of the method for controlling an intelligent body provided by the present application;
[0028] Figure 9 This is another flowchart of an embodiment of the method for controlling an intelligent body provided by the present application;
[0029] Figure 10 A schematic diagram of a flow chart of an embodiment of a control device for an intelligent body provided by the present application;
[0030] Figure 11 A schematic structural diagram of an embodiment of the intelligent agent provided in this application. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the technical solution of the present application, the technical solution provided by the present application is described in detail below with reference to the accompanying drawings.
[0032] Example embodiments will be described more fully hereinafter with reference to the accompanying drawings, but the described example embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the scope of this application to those skilled in the art.
[0033] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0034] The terms used herein are used only to describe specific embodiments and are not intended to limit this application. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, they specify the presence of features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof.
[0035] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0036] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present application, and will not be interpreted as having an idealized or overly formal meaning, unless clearly defined in the examples of the present application.
[0037] In related technologies, in order to achieve the core goal of artificial intelligence technology, that is, to realize the intelligent decision-making and behavior generation of intelligent agents (such as robots, robotic arms, end effectors, intelligent vehicle terminals, and human-computer interaction systems) in dynamic environments, it is usually necessary to rely on end-to-end deep learning models, and directly map environmental inputs to action outputs through large-scale data set training models. For example, some autonomous driving systems use multimodal sensor fusion and deep neural networks to generate driving decisions.
[0038] However, because end-to-end models rely on black-box mapping and lack explicit constraints on the action generation process, the model output is unstable, making the agent prone to uncontrollable behavior. For example, robots, manipulators, and end-effector grasping strategies based on reinforcement learning may suddenly generate dangerous actions (such as high-speed collisions with objects) under environmental disturbances, posing safety risks.
[0039] Based on this, the embodiments of the present application provide a control method, apparatus, device and storage medium for an intelligent agent, which generates a first action sequence of the intelligent agent based on first environmental feature data of the environment in which the intelligent agent is located; in response to a received user instruction, performs intent analysis on the user instruction to generate instruction intention information; generates a control signal based on the instruction intention information, the first environmental feature data and the first action sequence; controls the first action sequence based on the control signal to generate a second action sequence of the intelligent agent; and controls the intelligent agent to perform an action based on the second action sequence. In this way, by constructing a two-layer architecture including a basic capability layer and a user instruction control layer, that is, the basic capability layer generates the first action sequence based on the first environmental feature data, and the user instruction control layer parses the user intention information through the user instruction, and generates a control signal based on the user intention information, the first environmental feature data and the first action sequence, and controls the first action sequence through the control signal, so that in the process of generating the action of the intelligent agent, not only the constraints of the environment on the intelligent agent's action are considered, but also the explicit constraints of the user intention on the intelligent agent's action are integrated, thereby improving the stability of the intelligent agent's action and thereby enhancing the controllability of the intelligent agent's behavior.
[0040] The embodiments of the present application are further described below with reference to the accompanying drawings.
[0041] See Figure 1 , is a flow chart of a control method of an intelligent agent provided in an embodiment of the present application. The intelligent agent can be any electronic device based on artificial intelligence technology, including a robot, a robotic arm, an end effector, an intelligent vehicle terminal or a human-computer interaction system. Figure 1 As shown, the control method of the intelligent agent includes the following steps S101 to S105.
[0042] Step S101: Generate a first action sequence of the agent based on first environmental feature data of the environment in which the agent is located.
[0043] Before the above step S101, the intelligent agent can perceive the environment in which the intelligent agent is located through the environmental sensing device, obtain sensor data that can characterize the environmental state of the environment in which the intelligent agent is located, and form first environmental feature data based on the sensor data to realize the construction of the environmental cognition basis of the intelligent agent.
[0044] The aforementioned environmental sensing device can be any device or apparatus capable of sensing and obtaining an environmental state that is used to represent the environment in which the agent resides. Specifically, the environmental sensing device can be a multi-source sensing device that can perceive the environmental state of the agent's environment through multiple aspects, such as visual perception, tactile perception, and proprioception. In other words, the multi-source sensing device can sense the agent's environment through a stereoscopic vision unit, a tactile sensor array, and a proprioception unit, generating multimodal sensing data.
[0045] It should be noted that the above-mentioned environmental sensing device can be set in the intelligent body, that is, the environmental sensing device is a component of the intelligent body. In this case, the intelligent body directly collects the sensor data through its environmental sensing device; or, the environmental sensing device can also be a device independently set outside the intelligent body. In this case, the environmental sensing device perceives the sensor data and transmits it to the intelligent body in real time. This is not limited here.
[0046] The aforementioned sensor data may be any data capable of characterizing the environmental state of the agent's environment, and may include the agent's own data, data on other objects or subjects, and data on the ambient space, etc. Specifically, when the aforementioned environmental sensing device is a multi-source sensing device, the aforementioned multimodal sensor data may include multi-dimensional information such as the agent itself, the three-dimensional spatial position (position coordinates and attitude quaternions or other mathematical expressions) of the target object, surface physical properties (friction coefficient, material stiffness), and environmental constraints (obstacle distance, operating space boundaries).
[0047] The above-mentioned formation of the first environmental characteristic data based on the sensor data can be to directly use the sensor data as the above-mentioned first environmental characteristic data; or, the intelligent agent can obtain the above-mentioned first environmental characteristic data after preprocessing the above-mentioned sensor data. For example, the intelligent agent can perform at least one of timestamp alignment, noise filtering and dimensionality reduction on the received environmental data (more specifically, perform timestamp alignment, noise filtering and dimensionality reduction on the received environmental data) to form the above-mentioned first environmental characteristic data.
[0048] In some implementations, before step S101, the following steps may also be included:
[0049] Acquire multimodal sensory data of the agent’s environment;
[0050] The multimodal sensing data is input into a second encoder, and the second encoder extracts redundant information of the multimodal sensing data based on multi-layer nonlinear transformation to generate first environmental feature data, wherein the data dimension of the first environmental feature data is lower than the data dimension of the multimodal sensing data.
[0051] In this embodiment, the second encoder can be used to extract redundant information from multimodal sensing data based on multi-layer nonlinear transformations, convert high-dimensional multimodal sensing data into low-dimensional first environmental feature data, and achieve data dimensionality reduction, thereby reducing the complexity of data processing in intelligent agent behavior control and improving the response speed of intelligent agent behavior control.
[0052] The above-mentioned second encoder can be any encoder that can extract redundant information in high-dimensional data based on multi-layer nonlinearity, and the encoder can be a deep neural network, etc.
[0053] To ensure the effectiveness of the second encoder in data dimensionality reduction, before the multimodal sensor data is input into the second encoder, the initial encoder may be trained based on preset training data to obtain the second encoder. The training data may be sample data generated solely in a simulation environment.
[0054] Alternatively, in some embodiments, before inputting the multimodal sensor data into the second encoder, the method further includes:
[0055] Obtaining a first training sample set generated in a simulation environment;
[0056] Training an initial encoder based on the first training sample set to obtain an encoder to be migrated;
[0057] Obtaining a second training sample set generated in a real environment;
[0058] The encoder to be transferred is trained based on the second training sample set to obtain a second encoder.
[0059] In this embodiment, after obtaining the encoder to be migrated through training with sample data in a simulation environment, the intelligent agent can further train the encoder to be migrated with real sample data, thereby achieving migration fine-tuning of the encoder to be migrated, so that the trained second encoder has the ability to adapt to the actual physical environment.
[0060] The sample data in the simulation environment refers to data generated in a simulated environment. This data is based on the simulation and modeling of the real-world environment. For example, in an autonomous driving system, the simulation environment can simulate different traffic conditions, weather conditions, and road types to generate a large amount of driving scenario data. This data, including information such as the vehicle's position, speed, and surrounding obstacles, is used to train the autonomous driving model, enabling it to make accurate driving decisions in the simulated environment.
[0061] The aforementioned real-world sample data refers to data collected in actual, unsimulated environments. This data originates directly from the real world and possesses a higher degree of authenticity and complexity. For example, in a smart home system, real-world sample data might include user device operation records and environmental sensor data (such as temperature and humidity). This data reflects real-life user behavior patterns and environmental changes, and is used to train models to better adapt to real-world usage scenarios.
[0062] It should be noted that the above-mentioned pre-processing of the sensor data by the second encoder may be performed on all the sensor data; or, it may be performed on only part of the sensor data.
[0063] In some embodiments, the second encoder is configured with constraint information of at least one environmental factor; the multimodal sensing data includes: first feature data of at least one environmental factor, and second feature data of other environmental factors except the at least one environmental factor;
[0064] Extracting redundant information from the multimodal sensing data based on a multi-layer nonlinear transformation using a second encoder to generate first environmental feature data includes:
[0065] By using the constraint information through the second encoder, redundant information in the second feature data is extracted through multi-layer nonlinear transformation to obtain third feature data;
[0066] First environmental characteristic data including the first characteristic data and the third characteristic data is generated.
[0067] In this embodiment, by pre-configuring constraint information in the second encoder, the second encoder only extracts redundant information from the second feature data other than the data related to the constraint information during the dimensionality reduction process of the multimodal sensor data. This can achieve dimensionality reduction of the data while retaining the integrity of the required information, thereby improving the data processing efficiency in the intelligent agent behavior control while ensuring the execution accuracy of the intelligent agent behavior control.
[0068] The constraint information may include at least one pre-specified environmental factor that is associated with the environmental state of the environment in which the agent is located. Specifically, the constraint information may include an environmental factor that has a decisive influence on the behavior planning of the agent in executing the behavior.
[0069] Among them, the execution behavior of the intelligent agent may include any behavior that the intelligent agent can perform. For example, in the grasping action control of intelligent agents such as robots, robotic arms or end effectors, the constraint information may include key environmental factors such as contact surface friction characteristics and obstacle spatial distribution, so that the first environmental feature data includes key parameters such as refined environmental space information, target object attributes and operation constraints, which can not only reduce the amount of data in the sensor data and significantly reduce the computing load of subsequent modules, but also improve the adaptability of intelligent agents such as robots, robotic arms or end effectors to complex working conditions such as lighting changes and partial occlusion in grasping actions.
[0070] In the above step S101, after the intelligent agent obtains the above first environmental feature data, the intelligent agent may input the first environmental feature data into a preset action engine, and the preset action engine outputs the above first action sequence.
[0071] The first action sequence may include a series of actions performed by the agent over a continuous period of time. Specifically, for controlling grasping behavior of an agent such as a robot, robotic arm, or end effector, the first action sequence may include the action parameters of multiple parts (e.g., fingers, wrist, and palm) of the agent performing the grasping behavior over a continuous period of time, in at least one degree of freedom.
[0072] For example, in an intelligent agent such as a robot, a robotic arm, or an end effector, the first action sequence may include action parameters of 3-4 finger degrees of freedom, action parameters of 2-3 wrist degrees of freedom, and action parameters of 1-2 palm degrees of freedom, etc. in continuous time.
[0073] The above-mentioned preset action engine can be a pre-trained neural network model, the input of which is environmental feature data, and the output is an action sequence. Its core function is to convert the compressed environmental features into a basic action sequence that conforms to the laws of physics. For example, for intelligent bodies such as robots, robotic arms or end effectors, the preset action engine can map abstract environmental representations into executable motion parameters and output compound action instructions (i.e., action sequences) containing end trajectories, joint angles and contact force thresholds.
[0074] It should be noted that the above-mentioned preset action engine can adopt an encoder-decoder architecture. The encoder in the preset action engine can perform multi-level abstract processing on the input environmental features (i.e., the first environmental feature data), extract the core elements that affect the generation of action (such as obstacle distribution status, target grasping point posture, etc.), and form a feature vector that characterizes the potential laws of the motion pattern; the decoder in the preset action engine reconstructs the robot joint space motion trajectory based on the feature vector, and ensures that the output action meets the dynamic feasibility through the physical constraint layer.
[0075] The first action sequence may be an action sequence directly output by the preset action engine. Alternatively, in some embodiments, the first action sequence of the agent is generated based on the first environmental feature data of the environment in which the agent is located, including:
[0076] Inputting first environmental feature data of the environment in which the agent is located into a preset action engine, and the preset action engine outputs an initial action sequence;
[0077] Simulating and executing each action in the initial action sequence in the digital twin system to verify whether each action meets the first verification condition;
[0078] Actions that meet the first verification condition in the initial action sequence are eliminated to generate a first action sequence.
[0079] In this embodiment, after the preset action engine outputs the initial action sequence based on the first environmental feature data, the intelligent agent can further simulate the execution of each action in the initial action sequence in the digital twin system, and eliminate the actions in the initial action sequence that meet the first verification conditions to generate a first action sequence, so that the intelligent agent's action execution can be controlled through action verification.
[0080] The above-mentioned digital twin system can be configured in the above-mentioned intelligent agent, which can enable the intelligent agent to simulate the execution of corresponding actions (which can be called "virtual actions") in a virtual environment based on the execution of each action parameter in the initial action sequence, and then the intelligent agent can verify the simulated execution of each action according to the preset first verification condition.
[0081] The first verification condition can be any preset condition for eliminating some actions from the initial action sequence (also referred to as "eliminating some action parameters from the initial action sequence"). Specifically, for intelligent agents such as robots, manipulators, or end effectors, the first verification condition can include at least one of the following: a dangerous action that causes a collision or exceeds a torque limit; a smooth action with continuous acceleration; or an action in which the contact force changes beyond a preset range during the action.
[0082] In some implementations, the first verification condition can include: a dangerous action that causes a collision or excessive torque; a smooth action with continuous acceleration; and an action in which the contact force changes beyond a preset range during the action. This allows for a three-level verification of the action sequence output by the action engine, further enhancing the controllability of the agent's action execution.
[0083] To ensure the effectiveness of the preset action engine in the action sequence generation process, before inputting the first environmental feature data of the agent's environment into the preset action engine, the initial action engine can be trained based on preset training data to obtain the preset action engine. The training data can be sample data generated only in a simulation environment.
[0084] In some embodiments, before inputting the first environmental feature data of the environment in which the agent is located into the preset action engine, the method further includes:
[0085] Obtaining a third training sample set generated in a simulation environment;
[0086] Training the initial action engine based on the third training sample set to obtain the action engine to be transferred;
[0087] Obtaining a fourth training sample set generated in a real environment;
[0088] The action engine to be transferred is trained based on the fourth training sample set to obtain a preset action engine.
[0089] In this embodiment, after obtaining the action engine to be transferred through training with sample data in a simulation environment, the intelligent agent can further train the action engine to be transferred through real sample data, thereby achieving fine-tuning of the migration of the action engine to be transferred, so that the preset action engine obtained through training has the ability to adapt to the actual physical environment.
[0090] It should be noted that in addition to migrating and fine-tuning the trained motion engine through sample data generated in a real environment, the sample data in the third training sample set can also be perturbed. For example, noise parameters (such as sensor errors, material parameter perturbations, etc.) can be injected into the simulation environment through domain randomization technology to generate negative samples, thereby enhancing the preset motion engine's ability to adapt to real-world uncertainties.
[0091] In the above step S101 , the generation of the first action sequence may only depend on the above first environmental feature data.
[0092] Alternatively, in some embodiments, when the first environmental characteristic data is obtained by processing multimodal sensor data based on the second encoder, the second encoder is configured with a memory network component.
[0093] After the second encoder extracts redundant information from the multimodal sensing data based on multi-layer nonlinear transformation to generate the first environmental feature data, the method may further include:
[0094] Obtaining, by a second encoder, second environmental feature data recorded by the memory network component, where the second environmental feature data is environmental feature data generated at a historical moment, where the historical moment is a moment before the moment when the first environmental feature data was generated;
[0095] Generating a first action sequence of the agent based on first environmental feature data of the environment in which the agent is located, including:
[0096] Based on the first environment feature data and the second environment feature data, a first action sequence of the intelligent agent is generated.
[0097] In this embodiment, the intelligent agent can jointly generate the first action sequence based on the first environmental feature data at the current moment and the second environmental feature data at the historical moment, so that the generated first action sequence is more accurate.
[0098] The above historical moment can be any time point or period other than the current moment when the first environmental feature data is collected. For example, the agent can obtain the environmental feature data obtained when it performed an action within 2 minutes before the current moment as the above second environmental feature data, and so on.
[0099] The above-mentioned generation of the first action sequence of the intelligent agent based on the first environmental feature data and the second environmental feature data can be a weighted fusion of the first environmental feature data and the second environmental feature data to generate the fused environmental feature data, and the fused environmental feature data is input into the preset action engine to obtain the first action sequence.
[0100] It should be noted that, when performing weighted fusion of the first environmental feature data and the second environmental feature data, the weights corresponding to the first environmental feature data and the second environmental feature data can be preset according to actual needs. For example, the weights corresponding to the first environmental feature data and the second environmental feature data can be preset to 0.8 and 0.2, respectively, and so on.
[0101] Alternatively, in some embodiments, when the intelligent agent obtains the second environmental feature data generated at a historical moment through the memory network component, the first environmental feature data based on the environment in which the intelligent agent is located is used to generate the first action sequence of the intelligent agent, which may include: determining the second environmental feature data that matches the first environmental feature data; and determining the historical action sequence output in the preset action engine history based on the second environmental feature data that matches the first environmental feature data as the first action sequence, thereby realizing the backtracking of historical actions.
[0102] Of course, in some embodiments, when the agent obtains second environmental characteristic data generated at a historical moment through the memory network component, the method may further include generating environmental change trend prediction information based on the first environmental characteristic data and the second environmental characteristic data. This environmental change trend prediction information is used to predict the changing trend of the agent's environment. In this way, forward-looking decision support can be provided for the agent's behavioral control.
[0103] In the present application, after step S101, the method may further include: displaying third information, the third information being used to indicate the first action sequence and / or information about the generation process of the first action sequence. In this way, by displaying the third information, the first action sequence and its generation process can be visualized.
[0104] In some implementations, displaying the third information may include:
[0105] displaying third information and a second adjustment control, where the third information is used to indicate the first action sequence and / or information about the generation process of the first action sequence;
[0106] The above method may further include: adjusting the first action sequence in response to a second operation on a second adjustment control, wherein the second operation is based on a third information input.
[0107] In this way, the second operation can be input into the second control space based on the third information, and the first action sequence can be adjusted through the second operation, thereby realizing the generation process of the online control of the first action sequence.
[0108] The second operation may be any operation for inputting into the second control space based on the third information. For example, the second operation may be at least one of voice input, touch input, and air gesture input.
[0109] Step S102: In response to the received user instruction, perform intent analysis on the user instruction to generate instruction intent information.
[0110] In this step, when the intelligent agent receives the user instruction, it can parse the user instruction according to the preset intention parsing model to generate instruction intention information.
[0111] The above-mentioned instruction intention information can be any information that can represent the behavior that the user expects the intelligent agent to perform. Specifically, the above-mentioned instruction intention information can include expected action parameters. For example, if the user instruction includes the text "lower the Z axis by 0.1mm and rotate the X axis by 0.5°", then the expected action parameters can be (lower the Z axis by 0.1mm and rotate the X axis by 0.5°), etc.
[0112] It should be noted that the above-mentioned intelligent agent receives user instructions, performs intent analysis on user instructions, and generates instruction intention information, which can be executed simultaneously with the above-mentioned step S101, or before step S101, or after step S101, which is not limited here.
[0113] The user instruction can be any instruction input by the user to instruct the agent to perform an action, and there is an association between the user instruction and the first environmental feature data. For example, during the grasping action of an agent such as a robot, a robotic arm, or an end effector, the user instruction can be "Please handle the frosted glass on the left table with care, and be careful not to knock over the soy sauce bottle next to it." In this case, the first environmental feature data can include data such as the target object being a high-legged frosted glass (e.g., fragile, with a surface friction coefficient of 0.3), a soy sauce bottle in a tilted state 10 cm to the right, and a table surface that is partially slippery due to water stains.
[0114] The user instructions may be instructions of a single mode or may be composed of instructions of multiple modes, wherein the instructions of multiple modes may include at least two of natural language instructions, action teaching instructions, and structured instructions.
[0115] In the case where the above-mentioned user instructions only include instructions of a single modality, the intention parsing model can parse the user intention of the instructions of the modality based on the instruction parsing mechanism corresponding to the modality to obtain the above-mentioned instruction intention information.
[0116] In the case where the above-mentioned user instructions include instructions of multiple modalities, the intention parsing model can parse the instructions of each modality separately to obtain the user intentions corresponding to the instructions of various modalities, and fuse the user intentions corresponding to the instructions of multiple modalities to obtain the above-mentioned instruction intention information.
[0117] In some embodiments, the user instructions include instructions in multiple modalities.
[0118] like Figure 2 As shown, the above step S102 may include:
[0119] Step S1021: For each of the multiple modalities, based on the instruction parsing mechanism corresponding to the modality, perform intent parsing on the instruction of the modality to obtain the intent feature vector of the modality;
[0120] Step S1022: Map the intention feature vectors of multiple modalities into the same semantic space, and dynamically extract the key instruction components of the intention feature vectors in the semantic space through a multi-head attention mechanism to generate instruction intention information.
[0121] In this embodiment, the intention feature vectors corresponding to instructions of multiple modalities can be fused through a multi-head attention mechanism, so that the generated instruction intention information can better reflect the user's intention.
[0122] The above-mentioned intention parsing model may be pre-configured with instruction parsing mechanisms corresponding to the above-mentioned multiple modalities, so that when a user instruction is received, the instruction parsing mechanism corresponding to each modality can be used to perform intent parsing separately for the instruction of that modality, so as to ensure the accuracy of the parsing of the intention feature vectors of the instructions of each modality.
[0123] Specifically, the above-mentioned instruction parsing mechanism based on the corresponding modality performs intent parsing on the modal instruction to obtain the modal intent feature vector, including at least one of the following:
[0124] When multiple modal instructions include natural language instructions, the core components of the natural language instructions are extracted through part-of-speech tagging and dependency syntax analysis. The operation object attributes corresponding to the core components of the sentences are searched in the domain knowledge graph. Based on the operation object attributes, the intent feature vector of the modality corresponding to the natural language instruction is generated.
[0125] When multi-modal instructions include action teaching instructions, a spatiotemporal convolutional network is used to extract the trajectory features of the teaching actions in the action teaching instructions. Natural language instructions that match the trajectory features are searched in the shared embedding space, and an intent feature vector corresponding to the matching natural language instructions is generated.
[0126] In the case where instructions of multiple modalities include structured instructions, the structured instructions are converted into standard control parameter templates through syntax tree parsing, and the intention feature vector of the modality corresponding to the structured instructions is generated based on the standard control parameter template.
[0127] In this way, accurate analysis of the feature vectors of natural language instructions, action teaching instructions and structured instructions can be achieved respectively.
[0128] Furthermore, the above-mentioned user instructions include natural language instructions, action teaching instructions and structured instructions. The above-mentioned instruction parsing mechanism based on the corresponding modality performs intent parsing on the modal instructions to obtain the modal intention feature vector, including: extracting the core components of the natural language instructions through part-of-speech tagging and dependency syntax analysis, searching the operation object attributes corresponding to the core components of the sentences in the domain knowledge graph, and generating the intention feature vector of the modality corresponding to the natural language instructions based on the operation object attributes; using a spatiotemporal convolutional network to extract the trajectory features of the teaching action in the action teaching instruction, searching for natural language instructions that match the trajectory features in the shared embedding space, and generating the intention feature vector corresponding to the matching natural language instructions; converting the structured instructions into a standard control parameter template through syntax tree parsing, and generating the intention feature vector of the modality corresponding to the structured instructions based on the standard control parameter template. In this way, it can be ensured that when inputting instructions of multiple modalities, the user intention can be accurately parsed.
[0129] In the above step S102, the above instruction intention information can be directly output by the intention analysis model, that is, the instruction intention information is completely dependent on the user instruction received at the current moment; or, the intention analysis model can also refer to the user's historical instructions to output instruction intention information.
[0130] In some implementations, the above-mentioned intent analysis of the user instruction and generation of instruction intent information may include:
[0131] Get historical instructions associated with user instructions;
[0132] Based on the gating weights corresponding to the user instructions and the historical instructions, the user instructions and the historical instructions, the instruction intention information is generated.
[0133] In this way, the intelligent agent can generate the above-mentioned instruction intention information based on the user instructions and related historical instructions received at the current moment, so as to trace back the user's historical instructions, and then make the parsed user intention information more accurately reflect the user's true intention.
[0134] The above historical instructions associated with the user history may be any instructions input by the user before the current moment that are associated with the above user instructions. For example, the historical instructions may be instructions input a preset number of times before the current moment.
[0135] The above-mentioned command intention information is generated based on the gating weights corresponding to the user instructions and historical instructions, the user instructions and the historical instructions respectively. The intention feature vectors corresponding to the user instructions and the historical instructions are obtained respectively through the intention parsing model, and the intention feature vectors corresponding to the user instructions and the historical instructions are weightedly fused through the gating weights corresponding to the user instructions and the historical instructions respectively, and the intention feature vector obtained after fusion is determined as the above-mentioned intention command information.
[0136] It should be noted that the gating weights corresponding to the above-mentioned user instructions and historical instructions can be set according to actual needs. For example, the gating weights corresponding to the above-mentioned user instructions and historical instructions can be preset to be 0.7, 0.3, and so on.
[0137] Alternatively, in other embodiments, performing intent analysis on a user instruction to generate instruction intent information may include: obtaining historical instructions associated with the user instruction; and determining the historical intent information associated with the historical instructions as the instruction intent information. This can improve the efficiency of parsing and obtaining instruction intent information.
[0138] It should be noted that the acquisition of the above-mentioned historical instructions and historical intention information can also be achieved by presetting a memory network component related to the user instructions in the intelligent agent and using the memory network component, which will not be elaborated here.
[0139] Step S103: Generate a control signal based on the instruction intention information, the first environmental feature data and the first action sequence.
[0140] In this step, after obtaining the instruction intention information, the first environmental feature data and the first action sequence, the intelligent agent inputs the instruction intention information, the first environmental feature data and the first action sequence into a preset control signal generation model to generate a control signal.
[0141] The control signal generation model can be a pre-trained neural network-based model, the input of which is instruction intention information, first environmental feature data and first action sequence, and the corresponding control signal is obtained by the predicted output of the model.
[0142] In some embodiments, as Figure 3 As shown, the above step S103 includes:
[0143] Step S1031: inputting instruction intention information, first environmental feature data, and a first action sequence into a first controller, the first controller including a first encoder and a first decoder;
[0144] Step S1032: constructing a conditional latent space based on the instruction intention information and the first environmental feature data through the first encoder;
[0145] Step S1033: reparameterize the first action sequence in the conditional latent space through the first decoder to generate latent variables associated with the first action sequence;
[0146] Step S1034: Generate a control signal based on the latent variable and the conditional information of the conditioned latent space through the first decoder.
[0147] In this embodiment, in the first controller, the first action sequence may be taken as the object to be regulated, and the instruction intention information and the first environmental feature data may be jointly modeled as the basis for adjustment, that is, the first action sequence may be reparameterized in the conditional latent space constructed based on the instruction intention information and the first environmental feature data, and a control signal may be generated based on the latent variables and the conditional information of the conditional latent space, so that the generated control signal is more appropriate, thereby further improving the accuracy of the intelligent agent behavior control.
[0148] The above-mentioned first controller can be a pre-trained neural network model, that is, the first controller can include an encoding network (that is, a first encoder) and a decoding network (that is, a first decoder), and during the training process, the network parameters of the encoding network and the decoding network can be adjusted respectively according to the loss value between the control signal output by the first controller and the expected output signal of the first controller (that is, the labeled signal in the training sample), so as to obtain the first controller.
[0149] In some embodiments, the first controller is trained based on a target loss, and the target loss is calculated as follows:
[0150] Calculating a reconstruction sub-loss, where the reconstruction sub-loss reflects the consistency between the control signal output by the first controller and the expected output signal of the first controller;
[0151] Calculating a latent space loss, where the latent space loss reflects consistency between a first distribution of a control signal output by the first controller and a second distribution of a desired output signal of the first controller;
[0152] A smoothing loss is calculated, where the smoothing loss is used to suppress sudden jumps in the control signal output by the first controller.
[0153] In this embodiment, during the training process of the first controller, the reconstruction loss, latent space loss and smoothing loss can be used as the training optimization targets of the first controller, so that the model training process adopts a multi-objective optimization strategy, making the control signal output by the model more accurate, thereby further improving the accuracy of the intelligent agent behavior control.
[0154] Of course, the first controller can also be trained based on any one of the two losses: reconstruction loss, latent space loss and smoothing loss, which is not limited here.
[0155] From the above, it can be seen that the above-mentioned instruction intention information may include expected action parameters. In the case where the above-mentioned instruction intention information includes expected action parameters, the above-mentioned intelligent agent may directly rely on the expected action parameters to generate the above-mentioned control signal.
[0156] Alternatively, in some embodiments, Figure 4 As shown, the above step S103 may include:
[0157] Step S1035: abstractly express the expected action parameters and generate abstract functional features corresponding to the action parameters;
[0158] Step S1036: Generate a control signal based on the abstract functional feature, the first environmental feature data and the first action sequence.
[0159] In this embodiment, abstract functional characteristics can be generated by abstractly expressing the expected action parameters, and control signals can be generated based on the abstract functional characteristics, the first environmental feature data and the first action sequence, thereby realizing the abstraction of the control correction strategy (i.e., the control signal).
[0160] The above-mentioned abstract functional features can be any features of the abstract function used to describe the expected action parameters. Specifically, the abstract functional features can include task-oriented spatial features and high-level semantic features. That is, the above-mentioned abstract expression of the expected action parameters to generate the abstract functional features corresponding to the action parameters can include at least one of the following:
[0161] Mapping the desired motion parameters to the task-oriented feature space, for example, mapping the joint space motion parameters (such as angles and angular velocities) to the task-oriented feature space (such as end-point accuracy, energy efficiency, and safety margin);
[0162] The high-level semantic feature dimensions of the expected action parameters are converted to obtain the high-level semantic features corresponding to the expected action parameters. For example, the control signal is allowed to act on high-level semantic feature dimensions such as "grasping stability coefficient" and "path smoothness factor".
[0163] It should be noted that the above-mentioned abstract expression of the expected action parameters to generate abstract functional features corresponding to the action parameters can be determined by the mapping relationship between preset action parameters and abstract functional features to determine the abstract functional features corresponding to the expected action parameters.
[0164] In some embodiments, after the agent generates a control signal based on the instruction intent information, the first environmental characteristic data, and the first action sequence, the method further includes: displaying first information indicating the control signal and / or information about the generation process of the control signal. This can achieve visualization of the control signal and / or the generation process of the control signal.
[0165] The displaying of the first information may be displaying the first information only in the display interface so that the user can view the first information.
[0166] In some embodiments, after generating the control signal based on the instruction intention information, the first environmental feature data, and the first action sequence, the method further includes:
[0167] Displaying first information and a first adjustment control, where the first information is used to indicate the control signal and / or information about the generation process of the control signal;
[0168] In response to a first trigger operation on a first adjustment control, the control signal is adjusted, wherein the first trigger operation is based on a first information input.
[0169] In this embodiment, by displaying the first information and the first adjustment space, the user can input the first trigger operation based on the first adjustment control while viewing the first information, thereby achieving timely adjustment of the control signal.
[0170] The first adjustment control may be an adjustable control displayed simultaneously with the first information on the display interface, and when the user adjusts the adjustable control, the first information may also change in accordance with the input of the adjustment operation. Moreover, the adjustable control may be displayed in any form, for example, a floating ball or a progress bar.
[0171] The first information may be any information indicating the control signal and / or the generation process of the control signal. Specifically, the first information may include at least one of the following:
[0172] Dynamic heat map, used to show the influence weight of each control dimension in the control signal on the end effector posture of the intelligent agent;
[0173] A three-dimensional force diagram showing the direction and strength of the mechanical action of the action correction strategy generated by the regulatory signal;
[0174] The decision recording module is used to record relevant information of the action correction strategy generated by the control signal. For example, the decision recording module may include a historical control decision tree for recording the logical relationship of the key correction process in the action correction strategy.
[0175] For example, during the grasping control process of an intelligent agent such as a robot, robotic arm, or end effector, when a control signal indicating a connector deflection angle θ > 1° is detected, the historical control decision tree can prioritize activating the "contact surface self-alignment" control dimension (e.g., with a weight of 68%), followed by triggering the "flexible impedance adjustment" dimension (e.g., with a weight of 29%), and so on. This transparent display can shorten debugging cycles and improve the efficiency of abnormal condition diagnosis.
[0176] Specifically, the first information may include the dynamic heat map, the three-dimensional force line diagram, and the decision recording module.
[0177] In some embodiments, after generating the control signal based on the instruction intent information, the first environmental characteristic data, and the first action sequence, the method further includes: outputting second information, where the second information is prompt information generated based on the control signal strength of the control signal, and the second information includes at least one of the following:
[0178] Information indicating the strength of a regulatory signal;
[0179] Predictive information indicating the changing trend of regulatory signal strength;
[0180] Information indicating a control risk of controlling the first action sequence according to the control signal strength.
[0181] In this embodiment, prompt information for quantitatively evaluating the intensity of the control signal can be constructed, so that the user can intervene in the control signal in a timely manner according to the prompt information.
[0182] It should be noted that the above prompt information can be any information that can generate prompts for the user, and its output method can be configured according to actual needs.
[0183] For example, an indicator light may be set, where when the control signal strength satisfies |Ct|≤0.3, the indicator light displays green; when the control signal strength satisfies 0.3<|Ct|≤0.7, the indicator light displays yellow; and when the control signal strength satisfies |Ct|>0.7, the indicator light displays red; or, a strength trend prediction curve may be displayed, through which the changing trend of the control signal strength in the future (such as within the next 3 seconds) may be shown; or, a voice prompt may be output, where when the control signal strength satisfies 0.3<|Ct|≤0.7 for 3 consecutive seconds, a "pay attention to the correction amplitude" warning may be issued, and so on.
[0184] More specifically, the second information includes: information indicating the strength of the control signal; prediction information indicating the changing trend of the control signal strength; and information indicating the control risk of controlling the first action sequence according to the control signal strength.
[0185] Step S104: regulating the first action sequence based on the control signal to generate a second action sequence of the agent.
[0186] In this step, after the agent obtains the control signal, it can input the control signal into a preset action sequence control model, and the action sequence control model directly outputs the second action sequence. The action sequence model can be a prediction model obtained through neural network training.
[0187] In some embodiments, as Figure 5 As shown, the above step S104 includes:
[0188] Step S1041: inputting the control signal, the first action sequence, and the first environmental feature data into an action integrator, wherein the action integrator includes a feature fusion network and an action optimization network;
[0189] Step S1042: Generate a joint feature vector by dynamically associating the spatiotemporal features among the control signal, the first action sequence, and the first environmental feature data through a feature fusion network based on a temporal attention mechanism;
[0190] Step S1043: Generate an action correction value based on the joint feature vector through the action optimization network;
[0191] Step S1044: Correct the first action sequence based on the action correction amount through the action optimization network.
[0192] In this embodiment, the action integrator can dynamically associate the spatiotemporal features between the control signal, the first action sequence and the first environmental feature data based on the temporal attention mechanism to generate a joint feature vector, generate an action correction amount based on the joint feature vector, and correct the first action sequence based on the action correction amount, thereby ensuring the accuracy of the correction to the first action sequence.
[0193] The feature fusion network may be a network layer that dynamically associates the spatiotemporal features of the control signal, the first action sequence, and the first environmental feature data based on a temporal attention mechanism to generate a joint feature vector. For example, the network layer may include multiple feature extraction networks and a fusion network, wherein the multiple feature extraction networks respectively extract the spatiotemporal features of the control signal, the first action sequence, and the first environmental feature data, and the fusion feature network fuses the spatiotemporal features of the control signal, the first action sequence, and the first environmental feature data based on its attention mechanism to obtain a joint feature vector.
[0194] It should be noted that in the process of the above-mentioned action integrator dynamically associating the control signal, the first action sequence and the first environmental feature data to generate a joint feature vector based on the temporal attention mechanism, the above-mentioned action integrator can encode the following core information: the original motion pattern of the basic action, the correction requirements proposed by the control signal and the boundary conditions of the environmental state constraints, thereby providing a reliable input benchmark for subsequent decision-making.
[0195] The motion optimization network can output motion corrections based on the joint feature vector. For example, for an intelligent agent such as a robot, robotic arm, or end effector, the motion optimization network can output motion corrections for each degree of freedom at the corresponding motion part based on the joint feature vector.
[0196] The above-mentioned correction of the first action sequence based on the action correction amount through the action optimization network can be based on each action parameter in the action sequence, and the action parameters can be directly increased or decreased based on the corresponding action parameter amount in the action correction amount.
[0197] In some embodiments, the action integrator further comprises a dynamic weight decision network, and the first action sequence comprises a first action parameter of a target action part of the agent at a target degree of freedom;
[0198] like Figure 6 As shown, the above step S104 further includes:
[0199] Step S1045: Generate a first correction weight of the target action part in the target degree of freedom based on the joint feature vector through a dynamic weight decision network;
[0200] The above step S1044 includes:
[0201] The first action parameter is corrected based on the action correction amount and the first correction weight through the action optimization network to obtain the second action parameter, wherein the second action sequence includes the second action parameter.
[0202] In this embodiment, the first action parameter of the target action part in the target degree of freedom can be corrected based on the action correction amount and the first correction weight of the target action part in the target degree of freedom, so that different degrees of freedom of different action parts can be corrected with different amplitudes, so that the correction of each action is performed independently.
[0203] The above-mentioned method of generating the first correction weight of the target action part in the target degree of freedom based on the joint feature vector may be a method of pre-setting a correspondence between the feature vector and the correction weights of different degrees of freedom of different actions, and determining the correction weight of the target action part in the target degree of freedom corresponding to the above-mentioned joint feature vector as the above-mentioned first correction weight according to the correspondence; or, the dynamic weight decision network may be a trained prediction network, and the prediction network may input the joint feature vector to predict the first correction weight of the target action part in the target degree of freedom.
[0204] For example, for different joint eigenvectors, different action parts, and different degrees of freedom, the dynamic weight decision network can dynamically determine the corresponding weights in the interval [0,1] as the above-mentioned first correction weights, such as determining the weights of the end effector translational degree of freedom, the wrist rotational degree of freedom, etc.
[0205] In some embodiments, generating the first correction weight of the target action part in the target degree of freedom based on the joint feature vector includes at least one of the following:
[0206] When the initial correction weight is less than or equal to a preset weight threshold, determining the initial correction weight as the first correction weight, wherein the initial correction weight is the weight of the target action part in the target degree of freedom and is generated by the dynamic weight decision network based on the joint feature vector;
[0207] When the initial correction weight is less than or equal to the preset weight threshold, the preset weight threshold is determined as the first correction weight of the target degree of freedom to obtain the second action parameter.
[0208] In this embodiment, the intelligent agent can generate the initial correction weight of the target action part in the target degree of freedom based on the joint feature vector, and determine the initial correction weight as the first correction weight when the initial correction weight is less than or equal to the preset weight threshold, and determine the preset weight threshold as the first correction weight when the initial correction weight is greater than the preset weight threshold, thereby avoiding the correction weight being too high and realizing the constraint on the first correction weight.
[0209] It should be noted that the preset weight threshold can be a weight value pre-set based on actual needs and can be fixed; alternatively, the preset weight threshold can be dynamically adjusted based on the first environmental characteristic data. That is, before generating the first corrected weight for the target action part in the target degree of freedom based on the joint feature vector, the method further includes determining the preset weight threshold corresponding to the first environmental characteristic data. For example, when a wet surface is detected, the weight threshold is lowered compared to a dry surface, and so on.
[0210] Of course, the intelligent agent can also determine the above-mentioned first correction weight based on other information. For example, it can be dynamically adjusted based on the information of the stage of execution of each action (the startup period focuses on basic actions and has a larger weight; while the emergency control period focuses on environmental constraints and has a smaller weight, etc.), or it can be dynamically adjusted based on key correction needs (such as avoiding sudden obstacles, etc.).
[0211] The above-mentioned correction of the first action parameter based on the action correction amount and the first correction weight to obtain the second action parameter can be a fixed weight corresponding to the first action sequence set in the action optimization network. The second action parameter can be the product of the first action parameter and the fixed weight, and the sum of the product of the action correction amount and the first correction weight.
[0212] In some embodiments, as Figure 7 As shown, the above step S1044 may include:
[0213] Step S10441: Calculate a first product between the first action parameter and the first correction weight, and calculate a second product between the action correction amount and the second correction weight, where the second correction weight is 1 minus the first correction weight;
[0214] Step S10442: Determine the sum of the first product and the second product as the second motion parameter after the target degree of freedom is corrected.
[0215] In this embodiment, the first correction weight can be used to simultaneously adjust the weights corresponding to the first action parameter and the action correction amount, so that the determined second action parameter is more reasonable.
[0216] It should be noted that the above-mentioned regulation of the first action sequence based on the control signal to generate the second action sequence of the intelligent agent can be to correct the first action sequence only based on the action correction amount corresponding to each degree of freedom under each action and the first correction weight.
[0217] In some embodiments, as Figure 8 As shown, before the above step S104, the following steps are also included:
[0218] Step S106: obtaining the allowed action correction value associated with the first action sequence;
[0219] The above step S104 includes:
[0220] Based on the control signal and the minimum of the allowed action correction amount, the first action sequence is regulated to generate a second action sequence of the intelligent agent.
[0221] In this embodiment, in the regulation of the first motion sequence, the maximum motion correction amount of the first motion sequence can be limited by the regulation signal and the allowed motion correction amount, thereby ensuring stability in the motion adjustment.
[0222] The above-mentioned obtaining of the allowed action correction value associated with the first action sequence may be that a fixed allowed action correction value is preset in the intelligent agent, and the fixed allowed action correction value is used for any action sequence.
[0223] Alternatively, in some embodiments, Figure 9 As shown, before the above step S106, the following steps are also included:
[0224] Step S107: Identify the operation scenario type of the first action sequence;
[0225] Step S108: Determine a safety factor corresponding to the second action sequence according to the operation scenario type;
[0226] The above step S106 may include:
[0227] An allowable action correction amount associated with the first action sequence is determined according to the safety factor and the control signal strength of the control signal.
[0228] In this embodiment, different safety factors can be determined according to the operating scenario types of different action sequences, and the allowable action correction amount associated with the first action sequence can be determined based on the safety factor and the control signal strength of the control signal, so that the corresponding allowable action correction amount can be determined for different operating scenarios, further improving the safety of the action correction.
[0229] The above-mentioned determination of the safety factor corresponding to the second action sequence based on the operation scenario type can be that the intelligent body has a preset association relationship between different scenario types and correction amounts, and the correction amount associated with the operation scenario type of the above-mentioned first action sequence is determined as the allowable action correction amount based on the association relationship.
[0230] For example, for special scenarios (such as sudden obstacle avoidance), a larger allowable action correction amount can be set, and the correction limit can be automatically relaxed to meet emergency needs; while for routine operation scenarios, a relatively small allowable action correction amount is set, and strict constraints are maintained to ensure action stability.
[0231] The above-mentioned determination of the allowable action correction amount associated with the first action sequence based on the safety factor and the control signal strength of the control signal can be to set the allowable action correction amount to be less than or equal to the product of the safety factor and the control signal strength, that is, the allowable action correction amount ≤ safety factor × control signal strength.
[0232] Step S105: Based on the second action sequence, control the agent to perform actions.
[0233] In this step, after obtaining the second action sequence, the agent can be directly controlled to perform actions based on the second action sequence.
[0234] Alternatively, in some embodiments, before controlling the intelligent agent to perform actions based on the second action sequence, the above-mentioned method may also include: performing simulated actions based on the second action sequence through the digital twin system to verify whether each action meets the second verification condition; eliminating the actions that meet the second verification condition in the initial action sequence, and adjusting the second action sequence.
[0235] The adjustment to the second action sequence may be to adjust the first correction weight and re-adjust the action parameters in the first action sequence based on the adjusted first correction weight and the action correction amount.
[0236] The above-mentioned second verification condition can be the same as the above-mentioned first verification condition, that is, the above-mentioned second verification condition can include at least one of the following: a dangerous action that causes a collision or an excessive torque; a smooth action with continuous acceleration; an action in which the contact force changes during the action exceeds a preset range.
[0237] In some embodiments, after controlling the agent to perform actions based on the second action sequence, the method further includes:
[0238] generating a supplementary training sample based on the first environmental feature data and the second action sequence;
[0239] The preset action engine is trained based on the supplementary training samples to update the preset action engine.
[0240] In this embodiment, supplementary training samples can be generated through the first environmental feature data and the second action sequence, and the preset action engine can be retrained based on the supplementary training samples, thereby realizing a closed loop of the data link, updating the preset action engine in the form of a data flywheel, reducing the required data requirements, and improving sample utilization efficiency.
[0241] In an embodiment of the present application, a first action sequence of the intelligent agent is generated based on first environmental feature data of the environment in which the intelligent agent is located; in response to a received user instruction, the user instruction is parsed for intent to generate instruction intention information; a control signal is generated based on the instruction intention information, the first environmental feature data and the first action sequence; the first action sequence is regulated based on the control signal to generate a second action sequence of the intelligent agent; and the intelligent agent is controlled to perform actions based on the second action sequence. In this way, by constructing a two-layer architecture including a basic capability layer and a user instruction control layer, that is, the basic capability layer generates the first action sequence based on the first environmental feature data, and the user instruction control layer parses the user intention information through the user instruction, and generates a control signal based on the user intention information, the first environmental feature data and the first action sequence, and controls the first action sequence through the control signal, so that in the process of generating the action of the intelligent agent, not only the constraints of the environment on the action of the intelligent agent are considered, but also the explicit constraints of the user intention on the action of the intelligent agent are integrated, thereby improving the stability of the action of the intelligent agent and thus enhancing the controllability of the behavior of the intelligent agent.
[0242] The control method of the intelligent agent provided in the embodiment of the present application can be executed by the control device of the intelligent agent. In the embodiment of the present application, the control device of the intelligent agent provided in the embodiment of the present application is described by taking the control method of the intelligent agent executed by the control device of the intelligent agent as an example.
[0243] See Figure 10 , is a schematic diagram of the structure of the control device of the intelligent body provided in the embodiment of the present application. Figure 10 As shown, the control device 1000 of the intelligent body of the present application includes:
[0244] An action sequence generating module 1001 is configured to generate a first action sequence of the agent based on first environmental feature data of the environment in which the agent is located;
[0245] The intention analysis module 1002 is used to analyze the intention of the user instruction in response to the received user instruction and generate instruction intention information;
[0246] A control signal generating module 1003 is configured to generate a control signal based on the instruction intention information, the first environmental feature data, and the first action sequence;
[0247] a control module 1004, configured to control the first action sequence based on the control signal to generate a second action sequence of the agent;
[0248] The action execution module 1005 is used to control the agent to perform actions based on the second action sequence.
[0249] In some embodiments, the control signal generation module 1003 is specifically configured to:
[0250] Inputting the instruction intention information, the first environmental feature data and the first action sequence into a first controller, wherein the first controller includes a first encoder and a first decoder;
[0251] constructing, by the first encoder, a conditional latent space based on the instruction intention information and the first environmental feature data;
[0252] reparameterizing the first action sequence in the conditional latent space to generate latent variables associated with the first action sequence by the first decoder;
[0253] A control signal is generated by the first decoder based on the latent variable and conditional information of the conditioned latent space.
[0254] In some embodiments, the first controller is trained based on a target loss, and the target loss is calculated as follows:
[0255] Calculating a reconstruction sub-loss, where the reconstruction sub-loss reflects the consistency between the control signal output by the first controller and the expected output signal of the first controller;
[0256] Calculating a latent space loss, where the latent space loss reflects consistency between a first distribution of a control signal output by the first controller and a second distribution of a desired output signal of the first controller;
[0257] A smoothing loss is calculated, where the smoothing loss is used to suppress sudden jumps in the control signal output by the first controller.
[0258] In some embodiments, the apparatus 1000 further includes:
[0259] The allowed action correction value acquisition module is used to acquire the allowed action correction value associated with the first action sequence.
[0260] The control module 1004 may be specifically configured to:
[0261] Based on the control signal and the minimum of the allowed action correction amount, the first action sequence is regulated to generate a second action sequence of the intelligent agent.
[0262] In some embodiments, the apparatus 1000 further includes:
[0263] an operation scenario type identification module, configured to identify the operation scenario type of the first action sequence;
[0264] A safety factor determination module is used to determine the safety factor corresponding to the second action sequence according to the operation scenario type.
[0265] The allowable action correction amount acquisition module can be specifically used to:
[0266] The allowable action correction amount associated with the first action sequence is determined according to the safety factor and the control signal strength of the control signal.
[0267] In some embodiments, the instruction intention information includes expected action parameters.
[0268] The control signal generating module 1003 may be specifically configured to:
[0269] Abstractly expressing the expected action parameters and generating abstract functional features corresponding to the action parameters;
[0270] A control signal is generated based on the abstract functional feature, the first environmental feature data and the first action sequence.
[0271] In some embodiments, the apparatus 1000 further includes:
[0272] a first display module, configured to display first information and a first adjustment control, wherein the first information is configured to indicate the control signal and / or information about the generation process of the control signal;
[0273] A control signal adjustment module is configured to adjust the control signal in response to a first trigger operation on the first adjustment control, wherein the first trigger operation is based on the first information input.
[0274] In some embodiments, the first information includes at least one of the following:
[0275] A dynamic heat map, used to show the influence weight of each control dimension in the control signal on the end effector posture of the intelligent agent;
[0276] A three-dimensional force diagram showing the direction and intensity of the mechanical action of the action correction strategy generated by the control signal;
[0277] The decision recording module is used to record relevant information of the action correction strategy generated by the control signal.
[0278] In some embodiments, the apparatus 1000 further comprises:
[0279] An information output module is configured to output second information, where the second information is prompt information generated based on the control signal strength of the control signal, and the second information includes at least one of the following:
[0280] Information indicating the intensity of the regulatory signal;
[0281] Prediction information indicating a changing trend of the intensity of the regulatory signal;
[0282] Information used to indicate the control risk of controlling the first action sequence according to the control signal strength.
[0283] In some embodiments, the control module 1004 is specifically configured to:
[0284] Inputting the control signal, the first action sequence and the first environmental feature data into an action integrator, wherein the action integrator includes a feature fusion network and an action optimization network;
[0285] Dynamically associating the spatiotemporal features of the control signal, the first action sequence, and the first environmental feature data through the feature fusion network based on a temporal attention mechanism to generate a joint feature vector;
[0286] generating, by the motion optimization network, a motion correction based on the joint feature vector;
[0287] The first motion sequence is corrected based on the motion correction amount through the motion optimization network.
[0288] In some embodiments, the action integrator further includes a dynamic weight decision network, and the first action sequence includes first action parameters of a target action part of the agent at a target degree of freedom.
[0289] The apparatus 1000 may further include:
[0290] A correction weight generation module is used to generate a first correction weight of the target action part in the target degree of freedom based on the joint feature vector through the dynamic weight decision network.
[0291] The control module 1004 may be specifically configured to:
[0292] The first action parameter is corrected based on the action correction amount and the first correction weight by the action optimization network to obtain a second action parameter.
[0293] The second action sequence includes the second action parameter.
[0294] In some embodiments, the control module 1004 may be specifically configured to perform at least one of the following:
[0295] When the initial correction weight is less than or equal to a preset weight threshold, determining the initial correction weight as the first correction weight, wherein the initial correction weight is the weight of the target action part in the target degree of freedom and is generated by the dynamic weight decision network based on the joint feature vector;
[0296] In a case where the initial correction weight is less than or equal to a preset weight threshold, the preset weight threshold is determined as a first correction weight of the target degree of freedom.
[0297] In some embodiments, the control module 1004 may be specifically configured to:
[0298] Calculating a first product between the first action parameter and the first correction weight, and calculating a second product between the action correction amount and a second correction weight, where the second correction weight is 1 minus the first correction weight;
[0299] The sum of the first product and the second product is determined as a second motion parameter after correction of the target degree of freedom.
[0300] In some embodiments, the apparatus 1000 further comprises:
[0301] A sensor data acquisition module, configured to acquire multimodal sensor data of the environment in which the agent is located;
[0302] The environmental feature data generation module is used to input the multimodal sensing data into a second encoder, and extract redundant information of the multimodal sensing data based on multi-layer nonlinear transformation through the second encoder to generate the first environmental feature data.
[0303] In some embodiments, the apparatus 1000 further comprises:
[0304] A first training sample set acquisition module, configured to acquire a first training sample set generated in a simulation environment;
[0305] A first training module is configured to train an initial encoder based on the first training sample set to obtain an encoder to be migrated;
[0306] A second training sample set acquisition module, used to acquire a second training sample set generated in a real environment;
[0307] The second training module is configured to train the encoder to be migrated based on the second training sample set to obtain the second encoder.
[0308] In some embodiments, the second encoder is configured with constraint information of at least one environmental element; the multimodal sensing data includes: first feature data of the at least one environmental element, and second feature data of other environmental elements except the at least one environmental element.
[0309] The environmental feature data generation module can be specifically used to:
[0310] By using the constraint information through the second encoder, redundant information in the second feature data is extracted through multi-layer nonlinear transformation to obtain third feature data;
[0311] The first environmental characteristic data including the first characteristic data and the third characteristic data is generated.
[0312] In some embodiments, the second encoder is configured with a memory network component.
[0313] The environmental feature data generation module can also be used for:
[0314] Obtaining, by the second encoder, second environmental feature data recorded by the memory network component, where the second environmental feature data is environmental feature data generated at a historical moment, and the historical moment is a moment before a moment when the first environmental feature data is generated;
[0315] The action sequence generation module 1001 is specifically configured to:
[0316] Based on the first environmental feature data and the second environmental feature data, a first action sequence of the agent is generated.
[0317] In some embodiments, the user instructions include instructions in multiple modalities.
[0318] The intention analysis module 1002 can be specifically used to:
[0319] For each of the multiple modalities, based on the instruction parsing mechanism corresponding to the modality, perform intent parsing on the instruction of the modality to obtain an intent feature vector of the modality;
[0320] The intention feature vectors of multiple modalities are mapped into the same semantic space, and the key instruction components of the intention feature vectors are dynamically extracted in the semantic space through a multi-head attention mechanism to generate instruction intention information.
[0321] In some implementations, the user instructions include natural language instructions, action teaching instructions, and structured instructions.
[0322] The intention analysis module 1002 can be specifically used to:
[0323] Extracting the core components of the natural language instruction through part-of-speech tagging and dependency syntax analysis, searching the domain knowledge graph for the operation object attributes corresponding to the core components of the sentence, and generating the intention feature vector of the modality corresponding to the natural language instruction based on the operation object attributes;
[0324] Using a spatiotemporal convolutional network to extract trajectory features of the teaching action in the action teaching instruction, searching for natural language instructions matching the trajectory features in a shared embedding space, and generating an intent feature vector corresponding to the matching natural language instructions;
[0325] The structured instruction is converted into a standard control parameter template through syntax tree parsing, and an intention feature vector of the modality corresponding to the structured instruction is generated based on the standard control parameter template.
[0326] In some implementations, the intent parsing module 1002 may be specifically configured to:
[0327] Acquire historical instructions associated with the user instruction;
[0328] The instruction intention information is generated based on the gating weights corresponding to the user instruction and the historical instruction, the user instruction and the historical instruction respectively.
[0329] In some implementations, the action sequence generation module 1001 may be specifically configured to:
[0330] Inputting first environmental feature data of the environment in which the agent is located into a preset action engine, and having the preset action engine output an initial action sequence;
[0331] Simulating execution of each action in the initial action sequence in the digital twin system to verify whether each action meets the first verification condition;
[0332] Actions that meet the first verification condition in the initial action sequence are eliminated to generate the first action sequence.
[0333] In some embodiments, the first verification condition includes:
[0334] Dangerous actions that may cause collision or excessive torque;
[0335] It is a smooth action with continuous acceleration;
[0336] This refers to an action in which the contact force changes beyond the preset range during the action.
[0337] In some embodiments, the apparatus 1000 further comprises:
[0338] A third training sample set acquisition module, configured to acquire a third training sample set generated in a simulation environment;
[0339] A third training module is used to train the initial action engine based on the third training sample set to obtain an action engine to be migrated;
[0340] A fourth training sample set acquisition module, configured to acquire a fourth training sample set generated in a real environment;
[0341] The fourth training module is used to train the action engine to be transferred based on the fourth training sample set to obtain the preset action engine.
[0342] In some embodiments, the apparatus 1000 further comprises:
[0343] a second display module, configured to display third information and a second adjustment control, wherein the third information is configured to indicate the first action sequence and / or information about the generation process of the first action sequence;
[0344] An action sequence adjustment module is configured to adjust the first action sequence in response to a second operation on the second adjustment control, wherein the second operation is based on the third information input.
[0345] In some embodiments, the apparatus 1000 further comprises:
[0346] A supplementary training sample generating module, configured to generate a supplementary training sample based on the first environmental feature data and the second action sequence;
[0347] The action engine updating module is configured to train the preset action engine based on the supplementary training samples to update the preset action engine.
[0348] The control device 1000 of the intelligent body provided in the embodiment of the present application can execute the technical solution shown in the above-mentioned intelligent body control method embodiment. Its implementation principle and beneficial effects are similar and will not be repeated here.
[0349] The method for controlling an intelligent agent provided in the embodiment of the present application can be executed by a device for controlling the intelligent agent. In the embodiment of the present application, the method for controlling an intelligent agent performed by a device for controlling an intelligent agent is used as an example to illustrate the device for controlling an intelligent agent provided in the embodiment of the present application.
[0350] See Figure 11 , is a schematic diagram of the structure of an intelligent agent provided by an embodiment of the present application, the intelligent agent includes the above-mentioned bus routing node. Figure 11 As shown, the agent 300 includes:
[0351] one or more processors 1110;
[0352] The memory 1120 stores one or more programs. When the one or more programs are executed by the one or more processors 1110, the one or more processors 1110 implement the control method of the intelligent agent described in any of the above embodiments.
[0353] The memory 1120 is a non-transient network system that can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory 1120 may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 1120 may optionally include a memory 1120 remotely located relative to the processor 1110, and these remote memories 1120 may be connected to the processor 1110 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0354] The memory 1120 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1120 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1120 and is called by the processor 1110 to execute the methods of the embodiments of this application.
[0355] The processor 1110 can be implemented as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0356] In some embodiments, the agent further comprises:
[0357] Input / output interface, used to realize information input and output;
[0358] Communication interface, used to realize communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, Wi-Fi, Bluetooth, etc.);
[0359] A bus that transmits information between various components of the device (e.g., processor 1110, memory 1120, input / output interfaces, and communication interfaces);
[0360] The processor 1110 , the memory 1120 , the input / output interface, and the communication interface can be communicatively connected to each other within the device via a bus.
[0361] An embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the control method of the intelligent agent described in any of the above embodiments.
[0362] One embodiment of the present application also provides a computer program product, including a computer program or computer instructions, which are stored in a computer-readable storage medium. The processor of a computer device reads the computer program or computer instructions from the computer-readable storage medium, and the processor executes the computer program or computer instructions, so that the computer device executes the control method of the intelligent body described in any of the above embodiments.
[0363] The system architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.
[0364] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0365] Those skilled in the art will appreciate that all or some of the steps and systems in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those skilled in the art, the term computer storage media is included in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data) and is volatile and non-volatile, removable, and non-removable. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage, or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0366] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0367] The above description of some embodiments of the present application with reference to the accompanying drawings does not limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention shall be within the scope of the present application.
Claims
1. A method for controlling an intelligent agent, characterized in that: include: generating a first action sequence of the agent based on first environmental feature data of the environment in which the agent is located; In response to a received user instruction, performing intent analysis on the user instruction to generate instruction intent information; generating a control signal based on the instruction intention information, the first environmental feature data, and the first action sequence; regulating the first action sequence based on the control signal to generate a second action sequence of the agent; Based on the second action sequence, the agent is controlled to perform actions.
2. The method according to claim 1, characterized in that The generating of a control signal based on the instruction intention information, the first environmental feature data, and the first action sequence includes: Inputting the instruction intention information, the first environmental feature data and the first action sequence into a first controller, wherein the first controller includes a first encoder and a first decoder; constructing, by the first encoder, a conditional latent space based on the instruction intention information and the first environmental feature data; reparameterizing the first action sequence in the conditional latent space to generate latent variables associated with the first action sequence by the first decoder; A control signal is generated by the first decoder based on the latent variable and conditional information of the conditioned latent space.
3. The method according to claim 2, characterized in that The first controller is trained based on a target loss, and the target loss is calculated as follows: Calculating a reconstruction sub-loss, where the reconstruction sub-loss reflects the consistency between the control signal output by the first controller and the expected output signal of the first controller; Calculating a latent space loss, where the latent space loss reflects consistency between a first distribution of a control signal output by the first controller and a second distribution of a desired output signal of the first controller; A smoothing loss is calculated, where the smoothing loss is used to suppress sudden jumps in the control signal output by the first controller.
4. The method according to claim 1, wherein Before regulating the first action sequence based on the control signal to generate the second action sequence of the agent, the method further includes: obtaining an allowable action correction value associated with the first action sequence; The regulating the first action sequence based on the regulation signal to generate a second action sequence of the agent includes: Based on the control signal and the minimum of the allowed action correction amount, the first action sequence is regulated to generate a second action sequence of the intelligent agent.
5. The method according to claim 1, wherein The instruction intention information includes expected action parameters, The generating of a control signal based on the instruction intention information, the first environmental feature data, and the first action sequence includes: Abstractly expressing the expected action parameters and generating abstract functional features corresponding to the action parameters; A control signal is generated based on the abstract functional feature, the first environmental feature data and the first action sequence.
6. The method according to claim 1, characterized in that The regulating the first action sequence based on the regulation signal to generate a second action sequence of the agent includes: Inputting the control signal, the first action sequence and the first environmental feature data into an action integrator, wherein the action integrator includes a feature fusion network and an action optimization network; Dynamically associating the spatiotemporal features of the control signal, the first action sequence, and the first environmental feature data through the feature fusion network based on a temporal attention mechanism to generate a joint feature vector; generating, by the motion optimization network, a motion correction based on the joint feature vector; The first motion sequence is corrected based on the motion correction amount through the motion optimization network.
7. The method according to claim 6, characterized in that The action integrator further includes a dynamic weight decision network, wherein the first action sequence includes first action parameters of a target action part of the agent at a target degree of freedom; Before correcting the first action sequence based on the action correction amount using the action optimization network, the method further includes: generating, by the dynamic weight decision network, a first correction weight of the target action part at the target degree of freedom based on the joint feature vector; The step of modifying the first action sequence based on the action modification amount by the action optimization network includes: The first action parameter is corrected based on the action correction amount and the first correction weight by the action optimization network to obtain a second action parameter. The second action sequence includes the second action parameter.
8. The method according to claim 7, characterized in that The step of correcting the first action parameter based on the action correction amount and the first correction weight by the action optimization network to obtain the second action parameter includes: Calculating a first product between the first action parameter and the first correction weight, and calculating a second product between the action correction amount and a second correction weight, where the second correction weight is 1 minus the first correction weight; The sum of the first product and the second product is determined as a second motion parameter after correction of the target degree of freedom.
9. The method according to claim 1, characterized in that The user instructions include instructions of multiple modes. The step of performing intent analysis on the user instruction in response to the received user instruction to generate instruction intent information includes: For each of the multiple modalities, based on the instruction parsing mechanism corresponding to the modality, perform intent parsing on the instruction of the modality to obtain an intent feature vector of the modality; The intention feature vectors of multiple modalities are mapped into the same semantic space, and the key instruction components of the intention feature vectors are dynamically extracted in the semantic space through a multi-head attention mechanism to generate instruction intention information.
10. A control device for an intelligent agent, characterized in that: include: an action sequence generating module, configured to generate a first action sequence of the agent based on first environmental feature data of the environment in which the agent is located; An intent parsing module, configured to, in response to a received user instruction, perform intent parsing on the user instruction and generate instruction intent information; a control signal generating module, configured to generate a control signal based on the instruction intention information, the first environmental feature data, and the first action sequence; a control module, configured to control the first action sequence based on the control signal to generate a second action sequence for the agent; An action execution module is used to control the agent to perform actions based on the second action sequence.
Citation Information
Patent Citations
Intelligent agent control method and device, computer equipment and storage medium
CN117391181A
Intelligent agent behavior determination method, computer equipment and storage medium
CN117828039A
Robot operation creation device and program
JP2016215292A