Robot control method, device and equipment based on physical constraint embedding, and medium
By constructing a physical knowledge graph and fusing multimodal features, action prediction sequences are generated, solving the problem of inaccurate robot control in existing technologies. This enables a deep understanding and real-time perception of the physical environment, improving the accuracy of robot control and decision-making capabilities for complex tasks.
Patent Information
- Application Number
- CN202511188889.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Existing VLA technology systems neglect the objective laws and constraints of the physical world when dealing with physical environment interaction tasks, resulting in motion planning that does not conform to reality. Furthermore, they lack adaptability to complex physical environments and the ability to fuse multimodal information, making it difficult to achieve accurate robot control.
By collecting multimodal data, visual, linguistic, and action feature vectors are extracted, a physical knowledge graph is constructed and vectorized, and physically enhanced visual features are generated. A variational autoencoder is used to generate target features containing dynamic physical information, which are then input into a prediction network to generate action prediction sequences, which are then converted into control commands for the robot.
It achieves a deep understanding and real-time perception of the physical environment, improves the accuracy of robot control and decision-making ability for complex tasks, ensures that motion planning conforms to physical laws, and reduces the probability of task failure.
Smart Images

Figure CN120862691B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a robot control method, apparatus, device, and medium based on physical constraint embedding. Background Technology
[0002] In the existing VLA (Vision-Language-Action) technology system, the model has significant shortcomings when dealing with tasks involving physical environment interaction.
[0003] First, most current motion planning models often neglect objective laws and constraints in the physical world, such as gravity, friction, and collision rules. Taking robots commonly used in financial service halls and medical operating rooms as an example, traditional models may plan motion paths that do not conform to physical reality. For instance, the robot might directly pass through obstacles to carry objects, or the planned grasping action might be unable to support the weight of the object, causing it to fall. This significantly reduces the model's practicality in real-world physical scenarios.
[0004] Secondly, while some existing methods attempt to integrate physical knowledge into the model, most employ simple, hard-coded rules, lacking adaptability to dynamic changes in complex physical environments. These hard-coded physical rules cannot flexibly address differences in physical parameters across different scenarios; when factors such as the material, shape, and placement of objects in the environment change, the model struggles to make appropriate adjustments. Furthermore, existing methods of integrating physical knowledge with visual and linguistic modalities are not sufficiently in-depth, failing to effectively utilize multimodal information for a comprehensive understanding and reasoning about the physical environment, resulting in insufficient decision-making ability when handling complex physical tasks. Summary of the Invention
[0005] In view of the above, it is necessary to provide a robot control method, device, equipment and medium based on physical constraint embedding, which aims to solve the problem of inaccurate robot control caused by the lack of physical knowledge constraints.
[0006] A robot control method based on physical constraint embedding, the robot control method based on physical constraint embedding includes:
[0007] In response to control commands to the target robot, visual data, language data, and motion data of the target robot are collected as multimodal data;
[0008] Feature extraction is performed on the multimodal data to obtain a visual feature vector corresponding to the visual data, a language feature vector corresponding to the language data, and a motion feature vector corresponding to the motion data;
[0009] A physical knowledge graph is constructed, and the physical knowledge graph is vectorized to obtain physical knowledge feature vectors;
[0010] Construct a scene physics graph based on the visual feature vectors and the physical knowledge graph;
[0011] Generate physically enhanced visual features based on the language feature vectors and the scene physical graph;
[0012] The language feature vector, the physical knowledge feature vector, and the physical enhancement visual features are concatenated to obtain multimodal features that integrate physical knowledge.
[0013] Based on the variational autoencoder, target features containing dynamic physical information are generated according to the multimodal features, the visual feature vector, and the action feature vector;
[0014] The target features are input into a pre-built prediction network to obtain an action prediction sequence;
[0015] The predicted motion sequence is converted into motion commands, and the motion commands are used to control the target robot.
[0016] A robot control device based on physical constraint embedding, the robot control device based on physical constraint embedding includes:
[0017] The acquisition unit is used to acquire the visual data, language data and motion data of the target robot as multimodal data in response to the control command of the target robot;
[0018] The extraction unit is used to extract features from the multimodal data to obtain a visual feature vector corresponding to the visual data, a language feature vector corresponding to the language data, and an action feature vector corresponding to the action data.
[0019] A construction unit is used to construct a physical knowledge graph and to vectorize the physical knowledge graph to obtain physical knowledge feature vectors.
[0020] The construction unit is also used to construct a scene physical graph based on the visual feature vector and the physical knowledge graph;
[0021] The generation unit is used to generate physically enhanced visual features based on the language feature vector and the scene physical graph;
[0022] The splicing unit is used to splice the language feature vector, the physical knowledge feature vector, and the physical enhancement visual features to obtain multimodal features that integrate physical knowledge.
[0023] The generation unit is further configured to generate target features containing dynamic physical information based on the variational autoencoder, according to the multimodal features, the visual feature vector, and the action feature vector;
[0024] An input unit is used to input the target features into a pre-built prediction network to obtain an action prediction sequence;
[0025] A control unit is used to convert the motion prediction sequence into motion instructions and use the motion instructions to control the target robot.
[0026] A computer device, the computer device comprising:
[0027] Memory, storing at least one instruction; and
[0028] The processor executes instructions stored in the memory to implement the robot control method based on physical constraint embedding.
[0029] A computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to implement the robot control method based on physical constraint embedding.
[0030] As can be seen from the above technical solutions, this invention can extract visual feature vectors, language feature vectors, and action feature vectors from multimodal data to achieve a unified representation of multimodal data; it performs vectorization processing on the physical knowledge graph, thereby transforming abstract physical rules into machine-recognizable data structures; it constructs a scene physical graph based on visual feature vectors and the physical knowledge graph, generates physically enhanced visual features based on language feature vectors and the scene physical graph, and concatenates language feature vectors, physical knowledge feature vectors, and physically enhanced visual features to achieve deep fusion of physical knowledge with visual and language modalities; based on a variational autoencoder, it generates target features containing dynamic physical information based on multimodal features, visual feature vectors, and action feature vectors, improving the real-time perception capability of dynamic changes in the physical environment; it inputs the target features into a prediction network to obtain action prediction sequences, and converts the action prediction sequences into control commands for the robot, thereby achieving accurate control of the robot based on physical constraint embedding. Attached Figure Description
[0031] Figure 1 This is a flowchart of a preferred embodiment of the robot control method based on physical constraint embedding of the present invention.
[0032] Figure 2 This is a functional block diagram of a preferred embodiment of the robot control device based on physical constraint embedding of the present invention.
[0033] Figure 3This is a schematic diagram of the structure of a computer device that implements a preferred embodiment of the robot control method based on physical constraint embedding according to the present invention. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0035] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the robot control method based on physical constraint embedding of the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.
[0036] The robot control method based on physical constraint embedding is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0037] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.
[0038] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0039] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0040] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0041] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0042] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).
[0043] S10, in response to the control command to the target robot, collect the visual data, language data and motion data of the target robot as multimodal data.
[0044] In this embodiment, the target robot can be an intelligent interactive robot in a financial service hall, an auxiliary robot in a medical operating room, or a handling robot in a logistics scenario, etc.
[0045] In this embodiment, the control command can be automatically triggered when the target robot starts up, so as to achieve comprehensive control over the working process of the target robot.
[0046] In this embodiment, the visual data can be collected by a video acquisition device (such as a camera) deployed in the environment where the target robot is located, the language data can be collected by an audio acquisition device (such as a microphone) deployed in the environment where the target robot is located, and the motion data can be collected by a sensor (such as an infrared sensor) deployed in the environment where the target robot is located.
[0047] Of course, the video acquisition device, the audio acquisition device, and the sensor can also be deployed on the target robot.
[0048] S11, perform feature extraction on the multimodal data to obtain a visual feature vector corresponding to the visual data, a language feature vector corresponding to the language data, and an action feature vector corresponding to the action data.
[0049] In this embodiment, in order to perform targeted processing on data of different modalities, it is also necessary to extract various data features under the required modality.
[0050] Specifically, the step of extracting features from the multimodal data to obtain a visual feature vector corresponding to the visual data, a language feature vector corresponding to the language data, and a motion feature vector corresponding to the motion data includes:
[0051] The visual feature extraction model is used to extract features from the visual data to obtain the visual feature vector; wherein, the visual feature extraction model is constructed based on the Convolutional Neural Network (CNN) and the Vision Transformer (ViT) network architecture, the CNN is used to extract low-level visual features, and the Vision Transformer network architecture is used to capture global semantic associations based on a self-attention mechanism.
[0052] The language data is converted into context-dependent language features using a pre-trained language model, while retaining the physics-related semantic information in the language data, to obtain the language feature vector;
[0053] The dynamic dependencies of actions in the action data are captured using a Long Short-Term Memory (LSTM) network to obtain the action feature vector.
[0054] The pre-trained language model can be a BERT (Bidirectional Encoder Representation from Transformers) model, etc.
[0055] For example, when a surgical robot in a medical setting performs an instrument transfer task, it can extract the visual features (shape, weight) of the surgical instruments, the features of the doctor's voice commands (for example, handing over hemostats), and the features of historical transfer actions.
[0056] For example, in financial scenarios, when intelligent counter robots are handling the task of classifying bills, they can extract the visual features (size, material), language command features (putting large bills into the safe), and historical grasping action features of the bills.
[0057] The above embodiments enable unified feature representation of visual, linguistic, and action modalities, providing standardized input for subsequent processing.
[0058] S12, construct a physical knowledge graph and perform vectorization processing on the physical knowledge graph to obtain physical knowledge feature vectors.
[0059] In this embodiment, physical knowledge is first constructed, namely the physical knowledge graph and related features.
[0060] Specifically, the construction of the physical knowledge graph and the vectorization of the physical knowledge graph to obtain physical knowledge feature vectors include:
[0061] Obtain environmental data of the environment in which the target robot is located;
[0062] To acquire physical concepts, physical laws, and the interaction relationships between objects;
[0063] The physical knowledge graph is constructed based on the environmental data, the physical concepts, the physical laws, and the interaction relationships between objects.
[0064] The physical knowledge graph is converted into a vector using graph embedding technology to obtain the physical knowledge feature vector.
[0065] The physical concepts mentioned may include, but are not limited to, mass, velocity, force, coefficient of friction, etc.
[0066] The physical laws mentioned may include, but are not limited to, Newton's laws of motion and the law of conservation of energy.
[0067] The interaction relationships between objects may include, but are not limited to, collision, support, and adhesion.
[0068] The physical concepts, physical laws, and object interaction relationships can be obtained from trusted documents and platforms. For example, these can be collected from electronic textbooks, the history and cases of physics, and trusted experimental platforms.
[0069] Of course, the physical concepts, physical laws, and object interaction relationships can also be information uploaded by the user. For example, the physical concepts, physical laws, and object interaction relationships can be uploaded by experts.
[0070] The environmental data may include, but is not limited to, obstacles in the environment, light in the environment, and the location of objects in the environment.
[0071] The graph embedding technology mentioned above can include TransE (Translating Embeddings for Modeling Multi-relational Data) models, etc. Based on the idea of translation, the TransE model can embed entities and relations in a knowledge graph into a low-dimensional vector space. It is simple, efficient, and computationally inefficient, and trains quickly on large-scale knowledge graphs.
[0072] For example, in medical settings, when a surgical robot performs instrument transfer tasks, the physical knowledge graph can cover parameters such as instrument material (metal, plastic) and human tissue support force, thereby supporting the generation of subsequent precise movements.
[0073] For example, in financial scenarios, when intelligent counter robots are handling the task of classifying bills, they can embed physical parameters such as bill quality and desktop friction through a physical knowledge graph, thereby providing a basis for subsequent action planning.
[0074] Through the above embodiments, after constructing a physical knowledge graph, graph embedding technology is further used to realize the numerical representation of physical knowledge, so that abstract physical rules can be directly calculated and invoked by the model.
[0075] S13, Construct a scene physics graph based on the visual feature vector and the physical knowledge graph.
[0076] In this embodiment, the scene physics diagram is used to accurately depict the physical interaction relationships between objects, providing a scene physics logic basis for subsequent action planning.
[0077] In this embodiment, constructing the scene physical graph based on the visual feature vector and the physical knowledge graph includes:
[0078] Identify the objects and object attributes in the visual feature vector;
[0079] The physical knowledge graph is updated based on the object and its attributes to obtain the scene physical graph.
[0080] Among them, object detection algorithms (such as YOLO (You Only Look Once)) and semantic segmentation techniques (such as U-Net (U-Net Convolutional Neural Network)) can be used to identify objects (such as tables and cups) and their attributes (shape, size, material) in the scene from the visual feature vector.
[0081] The scene physics graph can add object nodes to the original physics knowledge graph, and use physical relationships as the edges of the object nodes, such as "the cup is supported by the table" and "the box is in friction with the ground".
[0082] Through the above embodiments, the physical knowledge graph can be expanded based on the visual feature vector, thereby providing a foundation for accurate subsequent action planning.
[0083] S14, Generate physically enhanced visual features based on the language feature vector and the scene physical graph.
[0084] To further improve the accuracy of motion planning, this embodiment performs physical enhancement of features based on the scene physics graph.
[0085] In this embodiment, generating physically enhanced visual features based on the language feature vector and the scene physical graph includes:
[0086] The scene physical graph is processed using a graph neural network (GNN) to obtain the visual features;
[0087] In the process of processing the scene physical graph using the graph neural network, the physical relationship features between objects are propagated through a message passing mechanism;
[0088] In the process of processing the scene physical graph using the graph neural network, physical semantic information is extracted from the language feature vector, and the weights of relevant nodes in the scene physical graph are enhanced through the attention mechanism and the physical semantic information.
[0089] For example, the physical semantic information may include "heavy objects" or "smooth surfaces".
[0090] For example, when enhancing the weights of relevant nodes in the scene's physical graph through attention mechanisms and the aforementioned physical semantic information, the weights of nodes corresponding to "heavy objects" can be increased.
[0091] S15, the language feature vector, the physical knowledge feature vector, and the physical enhancement visual features are concatenated to obtain multimodal features that integrate physical knowledge.
[0092] In this embodiment, after concatenating the language feature vector, the physical knowledge feature vector, and the physical enhancement visual features, the concatenated features can be input into a fully connected layer for mapping, thereby obtaining multimodal features that integrate physical knowledge.
[0093] Through the above embodiments, it is possible to achieve deep integration of physical knowledge with visual and linguistic modalities.
[0094] In this embodiment, the language feature vector, the physical knowledge feature vector, and the physically enhanced visual features can be directly concatenated sequentially. Alternatively, different weights can be assigned to each of the language feature vector, the physical knowledge feature vector, and the physically enhanced visual features, and then weighted and fused according to the corresponding weights. The specific concatenation method can be flexibly selected according to the actual needs of the scenario.
[0095] S16, Based on the Variational Autoencoder (VAE), generate target features containing dynamic physical information according to the multimodal features, the visual feature vector, and the action feature vector.
[0096] In this embodiment, the step of generating target features containing dynamic physical information based on the variational autoencoder, the multimodal features, the visual feature vector, and the action feature vector includes:
[0097] The visual feature vector and the action feature vector are input into the variational autoencoder, and the probability distribution of dynamic physical parameters is learned through the variational autoencoder.
[0098] The predicted dynamic physical parameters are obtained by sampling from the probability distribution;
[0099] The predicted dynamic physical parameters are encoded into feature vectors to obtain intermediate feature vectors;
[0100] The intermediate feature vector is concatenated with the multimodal feature to obtain the target feature.
[0101] The visual feature vector may include real-time scene information.
[0102] The action feature vector may include the robot's preceding grasping action, etc.
[0103] The dynamic physical parameters may include the object's speed, environmental friction, and the object's deformation coefficient.
[0104] Through the above embodiments, it is possible to perceive dynamic changes in the physical environment in real time (such as changes in the sliding speed of objects and changes in material hardness), breaking through the limitations of traditional hard-coded physical rules. Furthermore, the probability distribution modeling of dynamic physical parameters also improves the model's adaptability to uncertain environments (such as the fluctuation of friction between objects of different materials).
[0105] S17, the target features are input into a pre-built prediction network to obtain an action prediction sequence.
[0106] In this embodiment, the step of inputting the target features into a pre-built prediction network to obtain an action prediction sequence includes:
[0107] Introduce physical constraints as bias terms;
[0108] After the target features are input into the prediction network, a multi-head attention mechanism is used to mine the deep association between multimodal information and physical knowledge, and the deep association is adjusted according to the bias term to obtain the action prediction sequence.
[0109] The prediction network can be a network with reasoning capabilities, such as the Transformer network architecture.
[0110] The bias term is used to constrain attention weights. For example, if predicting the action of "grabbing an object", Newton's second law can be used as a bias term to assign higher attention weights to features related to "object mass" and "grabbing force"; if "object stacking" is involved, the weights of features related to "contact area" and "object center of gravity" can be increased based on the constraint of "support relationship".
[0111] The above embodiments ensure that the generated motion prediction sequence strictly conforms to physical laws, reducing the probability of task failure due to violations of physical constraints, such as the problem of objects falling due to insufficient gripping force.
[0112] S18, the motion prediction sequence is converted into motion instructions, and the motion instructions are used to control the target robot.
[0113] In this embodiment, after the motion prediction sequence is output through multi-layer Transformer processing, it can be transformed into specific motion instructions (such as robot joint angles, motion trajectories, force magnitudes, etc.) through fully connected layers and activation functions (such as ReLU). This enhances the collaborative reasoning ability of multimodal information and physical knowledge, thereby improving the decision-making accuracy of complex tasks.
[0114] For example, in financial scenarios, when intelligent robots are moving items, they can combine multimodal data such as "the weight of the item (physical parameter)," "the verbal command 'handle smoothly'," and "the flatness of the ground in vision" with physical constraints to plan a path with low acceleration and high gripping force, thereby avoiding slippage or collision during the handling of items.
[0115] For example, in medical settings, minimally invasive surgical robots can combine multimodal data such as "tissue elasticity coefficient (dynamic parameter)," "doctor's instruction to 'suturing gently'," and "tissue thickness in vision" with physical constraints to adjust the insertion depth and force of the suture needle, thereby avoiding tissue tearing (in accordance with the physical characteristics of human tissue).
[0116] As can be seen from the above technical solutions, this invention can extract visual feature vectors, language feature vectors, and action feature vectors from multimodal data to achieve a unified representation of multimodal data; it performs vectorization processing on the physical knowledge graph, thereby transforming abstract physical rules into machine-recognizable data structures; it constructs a scene physical graph based on visual feature vectors and the physical knowledge graph, generates physically enhanced visual features based on language feature vectors and the scene physical graph, and concatenates language feature vectors, physical knowledge feature vectors, and physically enhanced visual features to achieve deep fusion of physical knowledge with visual and language modalities; based on a variational autoencoder, it generates target features containing dynamic physical information based on multimodal features, visual feature vectors, and action feature vectors, improving the real-time perception capability of dynamic changes in the physical environment; it inputs the target features into a prediction network to obtain action prediction sequences, and converts the action prediction sequences into control commands for the robot, thereby achieving accurate control of the robot based on physical constraint embedding.
[0117] like Figure 2 The diagram shown is a functional block diagram of a preferred embodiment of the robot control device based on physical constraint embedding of the present invention. The robot control device 11 based on physical constraint embedding includes a data acquisition unit 110, an extraction unit 111, a construction unit 112, a generation unit 113, a splicing unit 114, an input unit 115, and a control unit 116. The module / unit referred to in this invention is a series of computer program segments that can be executed by a processor and perform a fixed function, and which are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0118] The acquisition unit 110 is used to acquire the visual data, language data and motion data of the target robot as multimodal data in response to the control command of the target robot.
[0119] In this embodiment, the target robot can be an intelligent interactive robot in a financial service hall, an auxiliary robot in a medical operating room, or a handling robot in a logistics scenario, etc.
[0120] In this embodiment, the control command can be automatically triggered when the target robot starts up, so as to achieve comprehensive control over the working process of the target robot.
[0121] In this embodiment, the visual data can be collected by a video acquisition device (such as a camera) deployed in the environment where the target robot is located, the language data can be collected by an audio acquisition device (such as a microphone) deployed in the environment where the target robot is located, and the motion data can be collected by a sensor (such as an infrared sensor) deployed in the environment where the target robot is located.
[0122] Of course, the video acquisition device, the audio acquisition device, and the sensor can also be deployed on the target robot.
[0123] The extraction unit 111 is used to extract features from the multimodal data to obtain a visual feature vector corresponding to the visual data, a language feature vector corresponding to the language data, and an action feature vector corresponding to the action data.
[0124] In this embodiment, in order to perform targeted processing on data of different modalities, it is also necessary to extract various data features under the required modality.
[0125] Specifically, the extraction unit 111 performs feature extraction on the multimodal data to obtain a visual feature vector corresponding to the visual data, a language feature vector corresponding to the language data, and a motion feature vector corresponding to the motion data, including:
[0126] The visual feature extraction model is used to extract features from the visual data to obtain the visual feature vector; wherein, the visual feature extraction model is constructed based on the Convolutional Neural Network (CNN) and the Vision Transformer (ViT) network architecture, the CNN is used to extract low-level visual features, and the Vision Transformer network architecture is used to capture global semantic associations based on a self-attention mechanism.
[0127] The language data is converted into context-dependent language features using a pre-trained language model, while retaining the physics-related semantic information in the language data, to obtain the language feature vector;
[0128] The dynamic dependencies of actions in the action data are captured using a Long Short-Term Memory (LSTM) network to obtain the action feature vector.
[0129] The pre-trained language model can be a BERT (Bidirectional Encoder Representation from Transformers) model, etc.
[0130] For example, when a surgical robot in a medical setting performs an instrument transfer task, it can extract the visual features (shape, weight) of the surgical instruments, the features of the doctor's voice commands (for example, handing over hemostats), and the features of historical transfer actions.
[0131] For example, in financial scenarios, when intelligent counter robots are handling the task of classifying bills, they can extract the visual features (size, material), language command features (putting large bills into the safe), and historical grasping action features of the bills.
[0132] The above embodiments enable unified feature representation of visual, linguistic, and action modalities, providing standardized input for subsequent processing.
[0133] The construction unit 112 is used to construct a physical knowledge graph and perform vectorization processing on the physical knowledge graph to obtain physical knowledge feature vectors.
[0134] In this embodiment, physical knowledge is first constructed, namely the physical knowledge graph and related features.
[0135] Specifically, the construction unit 112 constructs a physical knowledge graph and performs vectorization processing on the physical knowledge graph to obtain physical knowledge feature vectors, including:
[0136] Obtain environmental data of the environment in which the target robot is located;
[0137] To acquire physical concepts, physical laws, and the interaction relationships between objects;
[0138] The physical knowledge graph is constructed based on the environmental data, the physical concepts, the physical laws, and the interaction relationships between objects.
[0139] The physical knowledge graph is converted into a vector using graph embedding technology to obtain the physical knowledge feature vector.
[0140] The physical concepts mentioned may include, but are not limited to, mass, velocity, force, coefficient of friction, etc.
[0141] The physical laws mentioned may include, but are not limited to, Newton's laws of motion and the law of conservation of energy.
[0142] The interaction relationships between objects may include, but are not limited to, collision, support, and adhesion.
[0143] The physical concepts, physical laws, and object interaction relationships can be obtained from trusted documents and platforms. For example, these can be collected from electronic textbooks, the history and cases of physics, and trusted experimental platforms.
[0144] Of course, the physical concepts, physical laws, and object interaction relationships can also be information uploaded by the user. For example, the physical concepts, physical laws, and object interaction relationships can be uploaded by experts.
[0145] The environmental data may include, but is not limited to, obstacles in the environment, light in the environment, and the location of objects in the environment.
[0146] The graph embedding technology mentioned above can include TransE (Translating Embeddings for Modeling Multi-relational Data) models, etc. Based on the idea of translation, the TransE model can embed entities and relations in a knowledge graph into a low-dimensional vector space. It is simple, efficient, and computationally inefficient, and trains quickly on large-scale knowledge graphs.
[0147] For example, in medical settings, when a surgical robot performs instrument transfer tasks, the physical knowledge graph can cover parameters such as instrument material (metal, plastic) and human tissue support force, thereby supporting the generation of subsequent precise movements.
[0148] For example, in financial scenarios, when intelligent counter robots are handling the task of classifying bills, they can embed physical parameters such as bill quality and desktop friction through a physical knowledge graph, thereby providing a basis for subsequent action planning.
[0149] Through the above embodiments, after constructing a physical knowledge graph, graph embedding technology is further used to realize the numerical representation of physical knowledge, so that abstract physical rules can be directly calculated and invoked by the model.
[0150] The construction unit 112 is also used to construct a scene physical graph based on the visual feature vector and the physical knowledge graph.
[0151] In this embodiment, the scene physics diagram is used to accurately depict the physical interaction relationships between objects, providing a scene physics logic basis for subsequent action planning.
[0152] In this embodiment, the construction unit 112 constructs a scene physical graph based on the visual feature vector and the physical knowledge graph, including:
[0153] Identify the objects and object attributes in the visual feature vector;
[0154] The physical knowledge graph is updated based on the object and its attributes to obtain the scene physical graph.
[0155] Among them, object detection algorithms (such as YOLO (You Only Look Once)) and semantic segmentation techniques (such as U-Net (U-Net Convolutional Neural Network)) can be used to identify objects (such as tables and cups) and their attributes (shape, size, material) in the scene from the visual feature vector.
[0156] The scene physics graph can add object nodes to the original physics knowledge graph, and use physical relationships as the edges of the object nodes, such as "the cup is supported by the table" and "the box is in friction with the ground".
[0157] Through the above embodiments, the physical knowledge graph can be expanded based on the visual feature vector, thereby providing a foundation for accurate subsequent action planning.
[0158] The generation unit 113 is used to generate physically enhanced visual features based on the language feature vector and the scene physical graph.
[0159] To further improve the accuracy of motion planning, this embodiment performs physical enhancement of features based on the scene physics graph.
[0160] In this embodiment, the generation unit 113 generates physically enhanced visual features based on the language feature vector and the scene physical graph, including:
[0161] The scene physical graph is processed using a graph neural network (GNN) to obtain the visual features;
[0162] In the process of processing the scene physical graph using the graph neural network, the physical relationship features between objects are propagated through a message passing mechanism;
[0163] In the process of processing the scene physical graph using the graph neural network, physical semantic information is extracted from the language feature vector, and the weights of relevant nodes in the scene physical graph are enhanced through the attention mechanism and the physical semantic information.
[0164] For example, the physical semantic information may include "heavy objects" or "smooth surfaces".
[0165] For example, when enhancing the weights of relevant nodes in the scene's physical graph through attention mechanisms and the aforementioned physical semantic information, the weights of nodes corresponding to "heavy objects" can be increased.
[0166] The splicing unit 114 is used to splice the language feature vector, the physical knowledge feature vector, and the physical enhancement visual features to obtain multimodal features that integrate physical knowledge.
[0167] In this embodiment, after concatenating the language feature vector, the physical knowledge feature vector, and the physical enhancement visual features, the concatenated features can be input into a fully connected layer for mapping, thereby obtaining multimodal features that integrate physical knowledge.
[0168] Through the above embodiments, it is possible to achieve deep integration of physical knowledge with visual and linguistic modalities.
[0169] In this embodiment, the language feature vector, the physical knowledge feature vector, and the physically enhanced visual features can be directly concatenated sequentially. Alternatively, different weights can be assigned to each of the language feature vector, the physical knowledge feature vector, and the physically enhanced visual features, and then weighted and fused according to the corresponding weights. The specific concatenation method can be flexibly selected according to the actual needs of the scenario.
[0170] The generation unit 113 is further configured to generate target features containing dynamic physical information based on the variational autoencoder (VAE), according to the multimodal features, the visual feature vector, and the action feature vector.
[0171] In this embodiment, the generation unit 113, based on a variational autoencoder, generates target features containing dynamic physical information according to the multimodal features, the visual feature vector, and the action feature vector, including:
[0172] The visual feature vector and the action feature vector are input into the variational autoencoder, and the probability distribution of dynamic physical parameters is learned through the variational autoencoder.
[0173] The predicted dynamic physical parameters are obtained by sampling from the probability distribution;
[0174] The predicted dynamic physical parameters are encoded into feature vectors to obtain intermediate feature vectors;
[0175] The intermediate feature vector is concatenated with the multimodal feature to obtain the target feature.
[0176] The visual feature vector may include real-time scene information.
[0177] The action feature vector may include the robot's preceding grasping action, etc.
[0178] The dynamic physical parameters may include the object's speed, environmental friction, and the object's deformation coefficient.
[0179] Through the above embodiments, it is possible to perceive dynamic changes in the physical environment in real time (such as changes in the sliding speed of objects and changes in material hardness), breaking through the limitations of traditional hard-coded physical rules. Furthermore, the probability distribution modeling of dynamic physical parameters also improves the model's adaptability to uncertain environments (such as the fluctuation of friction between objects of different materials).
[0180] The input unit 115 is used to input the target features into a pre-constructed prediction network to obtain an action prediction sequence.
[0181] In this embodiment, the input unit 115 inputs the target features into a pre-built prediction network to obtain an action prediction sequence, including:
[0182] Introduce physical constraints as bias terms;
[0183] After the target features are input into the prediction network, a multi-head attention mechanism is used to mine the deep association between multimodal information and physical knowledge, and the deep association is adjusted according to the bias term to obtain the action prediction sequence.
[0184] The prediction network can be a network with reasoning capabilities, such as the Transformer network architecture.
[0185] The bias term is used to constrain attention weights. For example, if predicting the action of "grabbing an object", Newton's second law can be used as a bias term to assign higher attention weights to features related to "object mass" and "grabbing force"; if "object stacking" is involved, the weights of features related to "contact area" and "object center of gravity" can be increased based on the constraint of "support relationship".
[0186] The above embodiments ensure that the generated motion prediction sequence strictly conforms to physical laws, reducing the probability of task failure due to violations of physical constraints, such as the problem of objects falling due to insufficient gripping force.
[0187] The control unit 116 is used to convert the motion prediction sequence into motion instructions and use the motion instructions to control the target robot.
[0188] In this embodiment, after the motion prediction sequence is output through multi-layer Transformer processing, it can be transformed into specific motion instructions (such as robot joint angles, motion trajectories, force magnitudes, etc.) through fully connected layers and activation functions (such as ReLU). This enhances the collaborative reasoning ability of multimodal information and physical knowledge, thereby improving the decision-making accuracy of complex tasks.
[0189] For example, in financial scenarios, when intelligent robots are moving items, they can combine multimodal data such as "the weight of the item (physical parameter)," "the verbal command 'handle smoothly'," and "the flatness of the ground in vision" with physical constraints to plan a path with low acceleration and high gripping force, thereby avoiding slippage or collision during the handling of items.
[0190] For example, in medical settings, minimally invasive surgical robots can combine multimodal data such as "tissue elasticity coefficient (dynamic parameter)," "doctor's instruction to 'suturing gently'," and "tissue thickness in vision" with physical constraints to adjust the insertion depth and force of the suture needle, thereby avoiding tissue tearing (in accordance with the physical characteristics of human tissue).
[0191] As can be seen from the above technical solutions, this invention can extract visual feature vectors, language feature vectors, and action feature vectors from multimodal data to achieve a unified representation of multimodal data; it performs vectorization processing on the physical knowledge graph, thereby transforming abstract physical rules into machine-recognizable data structures; it constructs a scene physical graph based on visual feature vectors and the physical knowledge graph, generates physically enhanced visual features based on language feature vectors and the scene physical graph, and concatenates language feature vectors, physical knowledge feature vectors, and physically enhanced visual features to achieve deep fusion of physical knowledge with visual and language modalities; based on a variational autoencoder, it generates target features containing dynamic physical information based on multimodal features, visual feature vectors, and action feature vectors, improving the real-time perception capability of dynamic changes in the physical environment; it inputs the target features into a prediction network to obtain action prediction sequences, and converts the action prediction sequences into control commands for the robot, thereby achieving accurate control of the robot based on physical constraint embedding.
[0192] like Figure 3 The diagram shown is a schematic representation of the structure of a computer device that implements a preferred embodiment of the robot control method based on physical constraint embedding according to the present invention.
[0193] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a robot control program based on physical constraints.
[0194] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.
[0195] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are adaptable to this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.
[0196] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal and external storage units of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as code for robot control programs embedded based on physical constraints, but also to temporarily store data that has been output or will be output.
[0197] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., executing robot control programs based on physical constraints) and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.
[0198] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes these applications to implement the steps in the various embodiments of the robot control method based on physical constraint embedding described above, for example... Figure 1 The steps are shown.
[0199] For example, the computer program can be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units can be a series of computer-readable instruction segments capable of performing specific functions, which describe the execution process of the computer program in the computer device 1. For example, the computer program can be divided into a data acquisition unit 110, an extraction unit 111, a construction unit 112, a generation unit 113, a splicing unit 114, an input unit 115, and a control unit 116.
[0200] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute portions of the robot control method based on physical constraint embedding described in the various embodiments of this invention.
[0201] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.
[0202] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.
[0203] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.
[0204] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0205] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 3 The bus is represented by only one straight line, but this does not mean that there is only one bus or one type of bus. The bus is configured to enable communication between the memory 12 and at least one processor 13, etc.
[0206] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0207] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish a communication connection between the computer device 1 and other computer devices.
[0208] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.
[0209] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0210] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0211] Combination Figure 1 The memory 12 in the computer device 1 stores multiple instructions to implement a robot control method based on physical constraint embedding, and the processor 13 can execute the multiple instructions to achieve:
[0212] In response to control commands to the target robot, visual data, language data, and motion data of the target robot are collected as multimodal data;
[0213] Feature extraction is performed on the multimodal data to obtain a visual feature vector corresponding to the visual data, a language feature vector corresponding to the language data, and a motion feature vector corresponding to the motion data;
[0214] A physical knowledge graph is constructed, and the physical knowledge graph is vectorized to obtain physical knowledge feature vectors;
[0215] Construct a scene physics graph based on the visual feature vectors and the physical knowledge graph;
[0216] Generate physically enhanced visual features based on the language feature vectors and the scene physical graph;
[0217] The language feature vector, the physical knowledge feature vector, and the physical enhancement visual features are concatenated to obtain multimodal features that integrate physical knowledge.
[0218] Based on the variational autoencoder, target features containing dynamic physical information are generated according to the multimodal features, the visual feature vector, and the action feature vector;
[0219] The target features are input into a pre-built prediction network to obtain an action prediction sequence;
[0220] The predicted motion sequence is converted into motion commands, and the motion commands are used to control the target robot.
[0221] Specifically, the processor 13's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0222] It should be noted that all data involved in this case was legally obtained. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0223] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0224] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0225] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0226] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0227] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0228] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0229] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this invention can also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.
[0230] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A robot control method based on physical constraint embedding, characterized in that, The robot control method based on physical constraint embedding includes: In response to control commands to the target robot, visual data, language data, and motion data of the target robot are collected as multimodal data; Feature extraction is performed on the multimodal data to obtain a visual feature vector corresponding to the visual data, a language feature vector corresponding to the language data, and a motion feature vector corresponding to the motion data; A physical knowledge graph is constructed, and the physical knowledge graph is vectorized to obtain physical knowledge feature vectors. This includes: acquiring environmental data of the environment in which the target robot is located; acquiring physical concepts, physical laws, and object interaction relationships; constructing the physical knowledge graph based on the environmental data, the physical concepts, the physical laws, and the object interaction relationships; and converting the physical knowledge graph into vectors using graph embedding technology to obtain the physical knowledge feature vectors. The physical concepts include mass, velocity, force, or coefficient of friction. Construct a scene physics graph based on the visual feature vectors and the physical knowledge graph; Generating physically enhanced visual features based on the language feature vector and the scene physical graph includes: processing the scene physical graph using a graph neural network to obtain the visual features; wherein, during the process of processing the scene physical graph using the graph neural network, physical relationship features between objects are propagated through a message passing mechanism; wherein, during the process of processing the scene physical graph using the graph neural network, physical semantic information is extracted from the language feature vector, and the weights of relevant nodes in the scene physical graph are enhanced through an attention mechanism and the physical semantic information. The language feature vector, the physical knowledge feature vector, and the physical enhancement visual features are concatenated to obtain multimodal features that integrate physical knowledge. Based on a variational autoencoder, a target feature containing dynamic physical information is generated according to the multimodal features, the visual feature vector, and the action feature vector. The method includes: inputting the visual feature vector and the action feature vector into the variational autoencoder and learning a probability distribution of dynamic physical parameters through the variational autoencoder; sampling the predicted dynamic physical parameters from the probability distribution; encoding the predicted dynamic physical parameters into a feature vector to obtain an intermediate feature vector; and concatenating the intermediate feature vector with the multimodal features to obtain the target feature. The target features are input into a pre-constructed prediction network to obtain an action prediction sequence, including: introducing physical constraints as bias terms; after the target features are input into the prediction network, a multi-head attention mechanism is used to mine the deep association between multimodal information and physical knowledge, and the deep association is adjusted according to the bias terms to obtain the action prediction sequence; The predicted motion sequence is converted into motion commands, and the motion commands are used to control the target robot.
2. The robot control method based on physical constraint embedding as described in claim 1, characterized in that, The step of extracting features from the multimodal data to obtain a visual feature vector corresponding to the visual data, a language feature vector corresponding to the language data, and a motion feature vector corresponding to the motion data includes: The visual feature extraction model is used to extract features from the visual data to obtain the visual feature vector; wherein, the visual feature extraction model is constructed based on a convolutional neural network and a visual Transformer network architecture, the convolutional neural network is used to extract low-level visual features, and the visual Transformer network architecture is used to capture global semantic associations based on a self-attention mechanism. The language data is converted into context-dependent language features using a pre-trained language model, while retaining the physics-related semantic information in the language data, to obtain the language feature vector; The dynamic dependencies of actions in the action data are captured using a long short-term memory network to obtain the action feature vector.
3. The robot control method based on physical constraint embedding as described in claim 1, characterized in that, The step of constructing a scene physical graph based on the visual feature vector and the physical knowledge graph includes: Identify the objects and object attributes in the visual feature vector; The physical knowledge graph is updated based on the object and its attributes to obtain the scene physical graph.
4. A robot control device based on physical constraint embedding, characterized in that, The robot control device based on physical constraint embedding includes: The acquisition unit is used to acquire the visual data, language data and motion data of the target robot as multimodal data in response to the control command of the target robot; The extraction unit is used to extract features from the multimodal data to obtain a visual feature vector corresponding to the visual data, a language feature vector corresponding to the language data, and an action feature vector corresponding to the action data. A construction unit is used to construct a physical knowledge graph and vectorize the physical knowledge graph to obtain physical knowledge feature vectors. This includes: acquiring environmental data of the target robot's environment; acquiring physical concepts, physical laws, and object interaction relationships; constructing the physical knowledge graph based on the environmental data, the physical concepts, the physical laws, and the object interaction relationships; and converting the physical knowledge graph into vectors using graph embedding technology to obtain the physical knowledge feature vectors. The physical concepts include mass, velocity, force, or coefficient of friction. The construction unit is also used to construct a scene physical graph based on the visual feature vector and the physical knowledge graph; A generation unit is configured to generate physically enhanced visual features based on the language feature vector and the scene physical graph, including: processing the scene physical graph using a graph neural network to obtain the visual features; wherein, during the process of processing the scene physical graph using the graph neural network, physical relationship features between objects are propagated through a message passing mechanism; wherein, during the process of processing the scene physical graph using the graph neural network, physical semantic information is extracted from the language feature vector, and the weights of relevant nodes in the scene physical graph are enhanced through an attention mechanism and the physical semantic information. The splicing unit is used to splice the language feature vector, the physical knowledge feature vector, and the physical enhancement visual features to obtain multimodal features that integrate physical knowledge. The generation unit is further configured to generate target features containing dynamic physical information based on the variational autoencoder, according to the multimodal features, the visual feature vector, and the action feature vector, including: inputting the visual feature vector and the action feature vector into the variational autoencoder, and learning the probability distribution of dynamic physical parameters through the variational autoencoder; sampling the predicted dynamic physical parameters from the probability distribution; encoding the predicted dynamic physical parameters into feature vectors to obtain intermediate feature vectors; and concatenating the intermediate feature vectors with the multimodal features to obtain the target features. An input unit is used to input the target features into a pre-constructed prediction network to obtain an action prediction sequence, including: introducing physical constraints as bias terms; after inputting the target features into the prediction network, mining the deep association between multimodal information and physical knowledge through a multi-head attention mechanism, and adjusting the deep association according to the bias terms to obtain the action prediction sequence; A control unit is used to convert the motion prediction sequence into motion instructions and use the motion instructions to control the target robot.
5. A computer device, characterized in that, The computer device includes: Memory, storing at least one instruction; and The processor executes instructions stored in the memory to implement the robot control method based on physical constraint embedding as described in any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the robot control method based on physical constraint embedding as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Vision-language-action joint modeling-based disordered scene target object capturing method
CN115861596A
Visual language navigation method and device based on scene fusion knowledge and medium
CN116242359A