Robot control method and device combining scene and strategy model

By combining scene and strategy model methods, obtaining robot environment data and external instructions, generating and adjusting action sequences, the problem of lack of flexibility and accuracy of robot control in the prior art is solved, and more efficient task adaptation and control accuracy are achieved.

CN120056140AActive Publication Date: 2025-05-30CHUANGXIN QIZHI (BEIJING) TECH CO LTD +1

Patent Information

Application Number
CN202510560369.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-05-30
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing robot control methods lack flexibility and accuracy, cannot adapt to complex task scenarios, and require artificial re-analysis and setting of control strategies, resulting in excessive human and material consumption.

Method used

By obtaining the robot's three-dimensional environmental point cloud data and pose data, analyzing it in combination with external instructions, an initial action sequence is generated, and adjusting it according to the pose data, a target action sequence is generated, and finally the robot control is performed.

Benefits of technology

It improves the flexibility and accuracy of robot control, can adapt to different task scenarios, reduce human intervention, and reduce manpower and material consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120056140A_ABST
    Figure CN120056140A_ABST
Patent Text Reader

Abstract

The invention discloses a robot control method and device combining a scene and a strategy model, and relates to the field of robot control, and the method comprises the steps: obtaining the first scene data of a to-be-controlled robot; wherein the first scene data comprises three-dimensional environment point cloud data and robot pose data; obtaining a first external instruction, and analyzing the first external instruction in combination with the three-dimensional environment point cloud data to obtain a target task feature; inputting the target task features into a pre-training strategy model to generate an initial action sequence, and adjusting the initial action sequence according to the robot pose data to obtain a target action sequence; wherein the pre-training strategy model is obtained based on historical task data training of a to-be-controlled robot; and controlling a to-be-controlled robot according to the target action sequence. The flexibility and accuracy of robot control can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robot control, and in particular to a robot control method and device combining a scenario and a strategy model. Background Art

[0002] In the existing robot control methods, the preliminary task analysis work is generally manually analyzed and disassembled by relevant staff to obtain the corresponding specific parameters, and the task process and control strategy are preset according to these specific parameters to achieve control of the robot. However, when the robot's task environment, task objectives or task content changes, this traditional method based on task pre-analysis and pre-setting must be manually analyzed and the corresponding control strategy must be set again. It lacks flexibility and cannot adapt to complex task scenarios, and will cause unnecessary consumption of manpower and material resources. At the same time, the robot lacks perception of external commands during task execution, and cannot adjust task actions according to external commands, lacking corresponding flexibility. Therefore, how to improve the flexibility and accuracy of robot control is still one of the problems that need to be solved urgently in the existing technology. Summary of the invention

[0003] The present application provides a robot control method and device that combines scenario and strategy models to solve the technical problem that existing robot control methods lack flexibility and accuracy.

[0004] According to a first aspect of the embodiments of the present application, a robot control method combining a scenario and a strategy model is provided, comprising: Acquire first scene data of the robot to be controlled; wherein the first scene data includes three-dimensional environment point cloud data and robot posture data; Acquire a first external instruction, and analyze the first external instruction in combination with the three-dimensional environment point cloud data to obtain target task features; wherein the first external instruction is obtained by receiving a natural language input from a user; Input the target task feature into a pre-trained strategy model to generate an initial action sequence, and adjust the initial action sequence according to the robot posture data to obtain a target action sequence; wherein the pre-trained strategy model is trained based on historical task data of the robot to be controlled; The robot to be controlled is controlled according to the target action sequence.

[0005] This application first obtains the first scene data of the robot to be controlled, combines the first external instruction obtained by parsing the three-dimensional environmental point cloud data in the first scene data to obtain the target task feature, then inputs the target task feature into the pre-trained policy model to generate an initial action sequence, and then combines the robot pose data in the first scene data to adjust the initial action sequence to obtain the target action sequence, and finally controls according to the target action sequence. Compared with the prior art, this application obtains the first scene data and the first external instruction during the task execution based on the perception of the external environment, and improves the flexibility of subsequent robot control through the comprehensiveness of data types. At the same time, different from the prior art where the task process and control strategy are set manually, this application generates an initial action sequence based on the pre-trained policy model and adjusts it in combination with the robot pose data to obtain the target action sequence, improving the adaptability to different task scenarios, thereby improving the flexibility of subsequent robot control and the matching degree with different task scenarios, and thus improving the accuracy of subsequent robot control.

[0006] In some embodiments of this application, the obtaining of the first scene data of the robot to be controlled specifically includes: Based on radar technology, obtain the initial point cloud data of the robot to be controlled, and based on a binocular camera, obtain the environmental image data of the robot to be controlled; According to the environmental image data, combine the pose setting parameters of the robot to be controlled to obtain an initial rotation matrix; According to the initial rotation matrix, based on the feature point extraction method, construct a three-dimensional model of the initial point cloud data and align it to obtain the three-dimensional environmental point cloud data; According to the three-dimensional environmental point cloud data, combine the initial rotation matrix to map and obtain the robot pose data.

[0007] This application obtains the initial point cloud data and environmental image data based on radar technology and a binocular camera, then combines the pose setting parameters to obtain the initial rotation matrix, then obtains the three-dimensional environmental point cloud data based on the feature point extraction method, and finally combines the three-dimensional environmental point cloud data and the initial rotation matrix to obtain the robot pose data. Through the perception and processing of the external environmental data, it is possible to obtain the first scene data that better matches the current task scenario, making the subsequent initial action sequence and target action sequence better match the current task scenario, thereby improving the accuracy of robot control.

[0008] In some embodiments of this application, the parsing of the first external instruction by combining the three-dimensional environmental point cloud data to obtain the target task feature specifically includes: Parse the first external instruction based on a preset large language model to obtain a first task instruction, and extract features from the first task instruction based on a self-attention encoder to obtain a first task feature; Find multiple target task points in the three-dimensional environmental point cloud data according to the first task feature; Based on the multiple target task points and in combination with the first task feature, obtain a target task feature.

[0009] In this application, first, the first external instruction is parsed based on a preset large language model to obtain a first task instruction, and then feature extraction is performed based on a self-attention encoder to obtain a first task feature, which can convert the first external instruction in natural language format into a first task feature that is convenient for a computer to understand and process, facilitating subsequent processing; then, multiple target task points in the three-dimensional environmental point cloud data are found according to the first task feature, and a target task feature is obtained in combination with the first task feature, which can combine the features of multiple target task points and the first task feature to obtain a more accurate target task feature, so that the initial action sequence generated according to the target task feature is more accurate, improving the accuracy of controlling the robot.

[0010] In some embodiments of this application, the adjusting the initial action sequence according to the robot pose data to obtain a target action sequence specifically includes: Map the initial action sequence into the three-dimensional environmental point cloud data according to the robot pose data to obtain a mapped action point cloud and a task action path; Determine multiple unreachable action points in the task action path according to the three-dimensional environmental point cloud data and the mapped action point cloud; Find the nearest reachable points corresponding to the multiple unreachable action points respectively, and based on the nearest reachable points, correct the task action path to obtain a target action path; Adjust the initial action sequence according to the target action path to obtain a target action sequence.

[0011] In this application, first, the initial action sequence is mapped into the three-dimensional environmental point cloud data according to the robot pose data to obtain a mapped action point cloud and a task action path, and then multiple unreachable action points in the task action path are determined, and their nearest reachable points are found respectively, and then the target action path is corrected, which can eliminate and correct the unreasonable action points in the initial action sequence generated by the policy model, so as to obtain a more reasonable target action path, and then the initial action sequence is reversely adjusted according to the target action path to obtain a more accurate target action sequence, thereby improving the accuracy of controlling the robot.

[0012] In some embodiments of the present application, controlling the robot to be controlled according to the target action sequence specifically includes: Determining a task pose sequence according to the target action sequence, and calculating a target joint angle sequence corresponding to the task pose sequence based on an inverse kinematics algorithm; Determining a drive signal sequence according to the target joint angle sequence in combination with a preset action control instruction of the robot to be controlled; Controlling the robot to be controlled according to the drive signal sequence.

[0013] In the present application, first, a task pose sequence is determined according to a target action sequence, and a target joint angle sequence is calculated based on an inverse kinematics algorithm. Furthermore, in combination with a preset action control instruction, the obtained drive signal sequence can be made to match the current target action sequence, avoiding blind control, thereby improving the accuracy when controlling the robot.

[0014] According to the second aspect of the embodiments of the present application, there is provided a robot control device combining a scenario and a policy model, including a scenario data acquisition module, a task feature analysis module, an action sequence generation module, and a robot control module; The scenario data acquisition module is used to acquire first scenario data of the robot to be controlled; wherein, the first scenario data includes three-dimensional environmental point cloud data and robot pose data; The task feature analysis module is used to acquire a first external instruction, and analyze the first external instruction in combination with the three-dimensional environmental point cloud data to obtain target task features; wherein, the first external instruction is obtained by receiving a natural language input of a user; The action sequence generation module is used to input the target task features into a pre-trained policy model to generate an initial action sequence, and adjust the initial action sequence according to the robot pose data to obtain a target action sequence; wherein, the pre-trained policy model is trained based on historical task data of the robot to be controlled; The robot control module is used to control the robot to be controlled according to the target action sequence.

[0015] In some embodiments of the present application, the scenario data acquisition module includes an initial data acquisition unit, a rotation matrix acquisition unit, a three-dimensional point cloud construction unit, and a robot pose mapping unit; The initial data acquisition unit is used to acquire initial point cloud data of the robot to be controlled based on radar technology, and acquire environmental image data of the robot to be controlled based on a binocular camera; The rotation matrix acquisition unit is used to obtain an initial rotation matrix according to the environmental image data in combination with pose setting parameters of the robot to be controlled; The three-dimensional point cloud construction unit is used to construct a three-dimensional model of the initial point cloud data based on the feature point extraction method according to the initial rotation matrix and perform alignment to obtain three-dimensional environmental point cloud data; The robot pose mapping unit is used to map and obtain robot pose data according to the three-dimensional environmental point cloud data in combination with the initial rotation matrix.

[0016] In some embodiments of the present application, the task feature analysis module includes an instruction parsing and extraction unit, a target task point search unit, and a task feature acquisition unit; The instruction parsing and extraction unit is used to parse the first external instruction based on a preset large language model to obtain a first task instruction, and perform feature extraction on the first task instruction based on a self-attention encoder to obtain a first task feature; The target task point search unit is used to search for a plurality of target task points in the three-dimensional environmental point cloud data according to the first task feature; The task feature acquisition unit is used to obtain a target task feature based on the plurality of target task points in combination with the first task feature.

[0017] In some embodiments of the present application, the action sequence generation module includes an action sequence mapping unit, an unreachable point determination unit, an action path correction unit, and an action sequence adjustment unit; The action sequence mapping unit is used to map the initial action sequence into the three-dimensional environmental point cloud data according to the robot pose data to obtain a mapped action point cloud and a task action path; The unreachable point determination unit is used to determine a plurality of unreachable action points in the task action path according to the three-dimensional environmental point cloud data and the mapped action point cloud; The action path correction unit is used to respectively search for the nearest reachable points corresponding to the plurality of unreachable action points, and correct the task action path based on the nearest reachable points to obtain a target action path; The action sequence adjustment unit is used to adjust the initial action sequence according to the target action path to obtain a target action sequence.

[0018] In some embodiments of the present application, the robot control module includes a joint angle calculation unit, a drive signal determination unit, and a robot control unit; The joint angle calculation unit is used to determine a task pose sequence according to the target action sequence, and calculate a target joint angle sequence corresponding to the task pose sequence based on an inverse kinematics algorithm; The driving signal determining unit is configured to determine a driving signal sequence according to the target joint angle sequence and in combination with a preset action control instruction of the robot to be controlled. The robot control unit is configured to control the robot to be controlled according to the driving signal sequence.

[0019] In this application, first, the first scene data of the robot to be controlled is obtained, and the target task feature is obtained by combining the first external instruction parsed from the three-dimensional environmental point cloud data in the first scene data. Then, the target task feature is input into the pre-trained policy model to generate an initial action sequence. Next, the initial action sequence is adjusted in combination with the robot pose data in the first scene data to obtain the target action sequence. Finally, the robot is controlled according to the target action sequence. Compared with the prior art, this application obtains the first scene data and the first external instruction during the task execution based on the perception of the external environment, and improves the flexibility of subsequent robot control through the comprehensiveness of the data type. At the same time, different from the prior art where the task process and control strategy are set manually, this application generates the initial action sequence based on the pre-trained policy model and adjusts it in combination with the robot pose data to obtain the target action sequence, improving the adaptability to different task scenarios, thereby improving the flexibility of subsequent robot control and the matching degree with different task scenarios, and thus improving the accuracy of subsequent robot control. Description of the Drawings

[0020] Figure 1 : A flowchart showing a robot control method combining a scene and a policy model according to some embodiments of this application; Figure 2 : A module structure diagram of a robot control device combining a scene and a policy model according to some embodiments of this application. Detailed Embodiments

[0021] The following details the embodiments of this application. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by combining the drawings are exemplary and are only used to explain some embodiments of this application, and cannot be understood as a limitation of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments shown in this application without creative efforts shall fall within the protection scope of this application.

[0022] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, unless otherwise clearly and specifically defined, "multiple" and "several" mean two or more.

[0023] In the existing robot control methods, the robot is usually controlled by manually completing the task analysis work and presetting the task process and control strategy. However, this traditional method based on task pre-analysis and pre-setting requires manual re-analysis and setting of the corresponding control strategy, which lacks flexibility and cannot adapt to complex task scenarios, and will cause unnecessary consumption of manpower and material resources. In addition, the robot lacks the perception of external instructions during the execution of the task, and cannot adjust the task action according to the external instructions, lacking the corresponding flexibility. Therefore, how to improve the flexibility and accuracy of robot control is still one of the problems that need to be solved in the existing technology.

[0024] Based on the above technical background, please refer to Figure 1 The embodiment of the present application provides a robot control method combining a scenario and a strategy model, including steps S101 to S104, each of which is as follows: Step S101: Acquire first scene data of a robot to be controlled; wherein the first scene data includes three-dimensional environment point cloud data and robot posture data.

[0025] In certain embodiments of the present application, the obtaining of first scene data of the robot to be controlled specifically includes: Based on radar technology, the initial point cloud data of the robot to be controlled is obtained, and based on a binocular camera, the environmental image data of the robot to be controlled is obtained; According to the environmental image data, parameters are set in combination with the posture of the robot to be controlled to obtain an initial rotation matrix; According to the initial rotation matrix and based on the feature point extraction method, a three-dimensional model of the initial point cloud data is constructed and aligned to obtain three-dimensional environment point cloud data; The robot posture data is mapped based on the three-dimensional environment point cloud data in combination with the initial rotation matrix.

[0026] In certain embodiments of the present application, the initial rotation matrix is ​​obtained based on the environmental image data in combination with the posture setting parameters of the robot to be controlled, specifically: the first posture is obtained based on the environmental image data in combination with the posture setting parameters; the initial rotation matrix is ​​calculated based on the coordinate system where the first posture is located in combination with a preset coordinate system.

[0027] This application obtains initial point cloud data and environmental image data based on radar technology and binocular cameras, then combines pose setting parameters to obtain an initial rotation matrix, and then obtains three-dimensional environmental point cloud data based on a feature point extraction method. Finally, it combines the three-dimensional environmental point cloud data and the initial rotation matrix to obtain robot pose data. By perceiving and processing external environmental data, it is possible to obtain first scene data that better matches the current task scenario, making the subsequent obtained initial action sequence and target action sequence better match the current human scenario, thereby improving the accuracy of controlling the robot.

[0028] Step S102: Obtain a first external instruction, and combine the three-dimensional environmental point cloud data to parse the first external instruction to obtain target task features; wherein, the first external instruction is obtained by receiving the natural language input of the user.

[0029] In some embodiments of the present application, the combining the three-dimensional environmental point cloud data to parse the first external instruction to obtain target task features specifically includes: Based on a preset large language model, parse the first external instruction to obtain a first task instruction, and based on a self-attention encoder, perform feature extraction on the first task instruction to obtain first task features; According to the first task features, find multiple target task points in the three-dimensional environmental point cloud data; Based on the multiple target task points, combine the first task features to obtain target task features.

[0030] In some embodiments of the present application, the preferred solutions of the large language model include but are not limited to GPT, DeepSeek, Wenxin Yiyan, iFlytek Spark, Tongyi Qianwen, or an open-source model fine-tuned by Lora.

[0031] This application first parses the first external instruction based on a preset large language model to obtain a first task instruction, and then performs feature extraction based on a self-attention encoder to obtain first task features, which can convert the first external instruction in natural language format into first task features that are convenient for computers to understand and process, facilitating subsequent processing; then, according to the first task features, find multiple target task points in the three-dimensional environmental point cloud data and combine the first task features to obtain target task features, which can combine the features of multiple target task points and the first task features to obtain more accurate target task features, thereby making the initial action sequence generated according to the target task features more accurate and improving the accuracy of controlling the robot.

[0032] Step S103: Input the target task features into the pre-trained policy model to generate an initial action sequence, and adjust the initial action sequence according to the robot pose data to obtain a target action sequence; wherein, the pre-trained policy model is trained based on the historical task data of the robot to be controlled.

[0033] In some embodiments of the present application, the adjusting the initial action sequence according to the robot pose data to obtain a target action sequence specifically includes: According to the robot pose data, map the initial action sequence into the three-dimensional environmental point cloud data to obtain a mapped action point cloud and a task action path; Determine multiple unreachable action points in the task action path according to the three-dimensional environmental point cloud data and the mapped action point cloud; Respectively find the nearest reachable points corresponding to the multiple unreachable action points, and based on the nearest reachable points, correct the task action path to obtain a target action path; Adjust the initial action sequence according to the target action path to obtain a target action sequence.

[0034] In some embodiments of the present application, when finding the nearest reachable points corresponding to the unreachable action points and correcting the task action path based on the nearest reachable points, both are based on a preset path-finding algorithm, and the path-finding algorithm includes but is not limited to depth-first search algorithm, breadth-first search algorithm, Dijkstra algorithm, A* algorithm, and simulated annealing algorithm.

[0035] The present application first maps the initial action sequence into the three-dimensional environmental point cloud data according to the robot pose data to obtain a mapped action point cloud and a task action path, then determines multiple unreachable action points in the task action path, and respectively finds their nearest reachable points, and then corrects to obtain a target action path, which can eliminate and correct the unreasonable action points in the initial action sequence generated by the policy model, so as to obtain a more reasonable target action path, and then reversely adjust the initial action sequence according to the target action path to obtain a more accurate target action sequence, thereby improving the accuracy of controlling the robot.

[0036] Step S104: Control the robot to be controlled according to the target action sequence.

[0037] In some embodiments of the present application, the controlling the robot to be controlled according to the target action sequence specifically includes: According to the target action sequence, determine a task pose sequence, and calculate a target joint angle sequence corresponding to the task pose sequence based on the inverse kinematics algorithm; Determine a driving signal sequence according to the target joint angle sequence and in combination with a preset action control instruction of the robot to be controlled; Control the robot to be controlled according to the driving signal sequence.

[0038] In this application, first determine a task pose sequence according to a target action sequence, calculate a target joint angle sequence based on an inverse kinematics algorithm, and then in combination with a preset action control instruction, the obtained driving signal sequence can be made to match the current target action sequence, avoiding blind control, thereby improving the accuracy when controlling the robot.

[0039] Compared with the prior art, in this application, first obtain first scene data of the robot to be controlled, and in combination with a first external instruction obtained by parsing the three-dimensional environmental point cloud data in the first scene data, obtain a target task feature, then input the target task feature into a pre-trained policy model to generate an initial action sequence, and then adjust the initial action sequence in combination with the robot pose data in the first scene data to obtain a target action sequence, and finally control according to the target action sequence. Compared with the prior art, this application obtains the first scene data and the first external instruction in the process of executing a task according to the perception of the external environment, and improves the flexibility of subsequent control of the robot through the comprehensiveness of the data type; at the same time, different from the prior art where the task process and control strategy are set artificially, this application generates an initial action sequence based on a pre-trained policy model and adjusts it in combination with the robot pose data to obtain a target action sequence, improving the adaptability to different task scenarios, thereby improving the flexibility of subsequent control of the robot and the matching degree with different task scenarios, thereby improving the accuracy of subsequent control of the robot.

[0040] Corresponding to the foregoing method, please refer to Figure 2 This application embodiment provides a robot control device combining a scene and a policy model, including a scene data acquisition module 210, a task feature parsing module 220, an action sequence generation module 230, and a robot control module 240; The scene data acquisition module 210 is configured to acquire first scene data of the robot to be controlled; wherein, the first scene data includes three-dimensional environmental point cloud data and robot pose data; The task feature parsing module 220 is configured to acquire a first external instruction and, in combination with the three-dimensional environmental point cloud data, parse the first external instruction to obtain a target task feature; wherein, the first external instruction is obtained by receiving a natural language input from a user; The action sequence generation module 230 is configured to input the target task features into a pre-trained policy model to generate an initial action sequence, and adjust the initial action sequence according to the robot pose data to obtain a target action sequence; wherein, the pre-trained policy model is trained based on the historical task data of the robot to be controlled; The robot control module 240 is configured to control the robot to be controlled according to the target action sequence.

[0041] In some embodiments of the present application, the scene data acquisition module 210 includes an initial data acquisition unit, a rotation matrix acquisition unit, a three-dimensional point cloud construction unit, and a robot pose mapping unit; The initial data acquisition unit is configured to acquire initial point cloud data of the robot to be controlled based on radar technology, and acquire environmental image data of the robot to be controlled based on a binocular camera; The rotation matrix acquisition unit is configured to obtain an initial rotation matrix according to the environmental image data in combination with the pose setting parameters of the robot to be controlled; The three-dimensional point cloud construction unit is configured to construct a three-dimensional model of the initial point cloud data and align it based on a feature point extraction method according to the initial rotation matrix to obtain three-dimensional environmental point cloud data; The robot pose mapping unit is configured to map and obtain robot pose data according to the three-dimensional environmental point cloud data in combination with the initial rotation matrix.

[0042] In some embodiments of the present application, the task feature parsing module 220 includes an instruction parsing and extraction unit, a target task point searching unit, and a task feature acquisition unit; The instruction parsing and extraction unit is configured to parse the first external instruction based on a preset large language model to obtain a first task instruction, and perform feature extraction on the first task instruction based on a self-attention encoder to obtain first task features; The target task point searching unit is configured to search for a plurality of target task points in the three-dimensional environmental point cloud data according to the first task features; The task feature acquisition unit is configured to obtain target task features based on the plurality of target task points in combination with the first task features.

[0043] In some embodiments of the present application, the action sequence generation module 230 includes an action sequence mapping unit, an unreachable point determination unit, an action path correction unit, and an action sequence adjustment unit; The action sequence mapping unit is configured to map the initial action sequence into the three-dimensional environmental point cloud data according to the robot pose data to obtain a mapped action point cloud and a task action path; The unreachable point determination unit is configured to determine a plurality of unreachable action points in the task action path according to the three-dimensional environmental point cloud data and the mapped action point cloud; The action path correction unit is configured to separately find the nearest reachable points corresponding to the plurality of unreachable action points, and correct the task action path based on the nearest reachable points to obtain a target action path; The action sequence adjustment unit is configured to adjust the initial action sequence according to the target action path to obtain a target action sequence.

[0044] In some embodiments of the present application, the robot control module 240 includes a joint angle calculation unit, a drive signal determination unit, and a robot control unit; The joint angle calculation unit is configured to determine a task pose sequence according to the target action sequence, and calculate a target joint angle sequence corresponding to the task pose sequence based on an inverse kinematics algorithm; The drive signal determination unit is configured to determine a drive signal sequence according to the target joint angle sequence in combination with a preset action control instruction of the robot to be controlled; The robot control unit is configured to control the robot to be controlled according to the drive signal sequence.

[0045] In the present application, first, the first scene data of the robot to be controlled is obtained, and the target task feature is obtained by combining the first external instruction parsed from the three-dimensional environmental point cloud data in the first scene data. Then, the target task feature is input into the pre-trained policy model to generate an initial action sequence. Next, the initial action sequence is adjusted in combination with the robot pose data in the first scene data to obtain a target action sequence. Finally, the robot is controlled according to the target action sequence. Compared with the prior art, in the present application, the first scene data and the first external instruction in the task execution process are obtained according to the perception of the external environment, and the flexibility of subsequent robot control is improved through the comprehensiveness of the data type; at the same time, different from the prior art where the task process and control strategy are set artificially, in the present application, the initial action sequence is generated based on the pre-trained policy model and adjusted in combination with the robot pose data to obtain the target action sequence, improving the adaptability to different task scenarios, thereby improving the flexibility of subsequent robot control and the matching degree with different task scenarios, and thus improving the accuracy of subsequent robot control.

[0046] It should be understood that the device provided by the embodiments of the present application corresponds to the foregoing method. A robot control device combining a scene and a policy model provided by the embodiments of the present application can implement a robot control method combining a scene and a policy model provided by any one of the embodiments of the present application.

[0047] Adaptively, an embodiment of the present application further provides a computer device and a computer-readable storage medium.

[0048] The computer device includes: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; Wherein, when the processor executes the computer program, a robot control method combining a scenario and a policy model of the present application is implemented.

[0049] The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by the processor to execute a robot control method combining a scenario and a policy model of the present application.

[0050] The above are partial embodiments of the present application. The purpose, technical solutions, and beneficial effects of the present application have been further described in detail. It should be clear that the above partial embodiments of the present application should not be construed as a limitation of the present application. In particular, for those skilled in the art, any changes, modifications, equivalent replacements, and variations made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A robot control method combining scenario and strategy model, characterized in that: include: Acquire first scene data of the robot to be controlled; wherein the first scene data includes three-dimensional environment point cloud data and robot posture data; Acquire a first external instruction, and analyze the first external instruction in combination with the three-dimensional environment point cloud data to obtain target task features; wherein the first external instruction is obtained by receiving a natural language input from a user; Input the target task feature into a pre-trained strategy model to generate an initial action sequence, and adjust the initial action sequence according to the robot posture data to obtain a target action sequence; wherein the pre-trained strategy model is trained based on historical task data of the robot to be controlled; The robot to be controlled is controlled according to the target action sequence.

2. A robot control method combining scenario and strategy model according to claim 1, characterized in that: The step of obtaining the first scene data of the robot to be controlled specifically includes: Based on radar technology, the initial point cloud data of the robot to be controlled is obtained, and based on a binocular camera, the environmental image data of the robot to be controlled is obtained; According to the environmental image data, in combination with the posture setting parameters of the robot to be controlled, an initial rotation matrix is ​​obtained; According to the initial rotation matrix and based on the feature point extraction method, a three-dimensional model of the initial point cloud data is constructed and aligned to obtain three-dimensional environment point cloud data; The robot posture data is mapped based on the three-dimensional environment point cloud data in combination with the initial rotation matrix.

3. A robot control method combining scenario and strategy model according to claim 1, characterized in that: The first external instruction is analyzed in combination with the three-dimensional environment point cloud data to obtain target task features, specifically including: Based on a preset large language model, the first external instruction is parsed to obtain a first task instruction, and based on a self-attention encoder, a feature of the first task instruction is extracted to obtain a first task feature; According to the first task feature, searching for a plurality of target task points in the three-dimensional environment point cloud data; Based on the multiple target task points and in combination with the first task feature, a target task feature is obtained.

4. A robot control method combining scenario and strategy model according to claim 1, characterized in that: The adjusting the initial action sequence according to the robot posture data to obtain the target action sequence specifically includes: According to the robot posture data, mapping the initial action sequence to the three-dimensional environment point cloud data to obtain a mapping action point cloud and a task action path; Determining a plurality of unreachable action points in the task action path according to the three-dimensional environment point cloud data and the mapped action point cloud; Find the nearest reachable points corresponding to the multiple unreachable action points respectively, and based on the nearest reachable points, correct the task action path to obtain the target action path; According to the target action path, the initial action sequence is adjusted to obtain a target action sequence.

5. The robot control method combining scenario and strategy model according to claim 1, characterized in that: The controlling the robot to be controlled according to the target action sequence specifically includes: According to the target action sequence, a task posture sequence is determined, and a target joint angle sequence corresponding to the task posture sequence is calculated based on an inverse kinematics algorithm; Determine a drive signal sequence according to the target joint angle sequence and in combination with preset motion control instructions of the robot to be controlled; The robot to be controlled is controlled according to the drive signal sequence.

6. A robot control device combining a scenario and a strategy model, characterized in that: It includes scene data acquisition module, task feature analysis module, action sequence generation module and robot control module; The scene data acquisition module is used to acquire first scene data of the robot to be controlled; wherein the first scene data includes three-dimensional environment point cloud data and robot posture data; The task feature parsing module is used to obtain a first external instruction, and parse the first external instruction in combination with the three-dimensional environment point cloud data to obtain a target task feature; wherein the first external instruction is obtained by receiving a natural language input from a user; The action sequence generation module is used to input the target task features into the pre-trained strategy model to generate an initial action sequence, and adjust the initial action sequence according to the robot posture data to obtain a target action sequence; wherein the pre-trained strategy model is trained based on the historical task data of the robot to be controlled; The robot control module is used to control the robot to be controlled according to the target action sequence.

7. A robot control device combining scenario and strategy model according to claim 6, characterized in that: The scene data acquisition module includes an initial data acquisition unit, a rotation matrix acquisition unit, a three-dimensional point cloud construction unit and a robot posture mapping unit; The initial data acquisition unit is used to acquire initial point cloud data of the robot to be controlled based on radar technology, and to acquire environmental image data of the robot to be controlled based on a binocular camera; The rotation matrix acquisition unit is used to obtain an initial rotation matrix according to the environmental image data and in combination with the posture setting parameters of the robot to be controlled; The three-dimensional point cloud construction unit is used to construct a three-dimensional model of the initial point cloud data and align it according to the initial rotation matrix based on the feature point extraction method to obtain three-dimensional environment point cloud data; The robot posture mapping unit is used to map the robot posture data according to the three-dimensional environment point cloud data in combination with the initial rotation matrix.

8. A robot control device combining scenario and strategy model according to claim 6, characterized in that: The task feature parsing module includes an instruction parsing and extracting unit, a target task point searching unit and a task feature acquiring unit; The instruction parsing and extraction unit is used to parse the first external instruction based on a preset large language model to obtain a first task instruction, and to extract features of the first task instruction based on a self-attention encoder to obtain a first task feature; The target task point search unit is used to search for a plurality of target task points in the three-dimensional environment point cloud data according to the first task feature; The task feature acquisition unit is used to obtain a target task feature based on the multiple target task points and in combination with the first task feature.

9. A robot control device combining scenario and strategy model according to claim 6, characterized in that: The action sequence generation module includes an action sequence mapping unit, an unreachable point determination unit, an action path correction unit and an action sequence adjustment unit; The action sequence mapping unit is used to map the initial action sequence to the three-dimensional environment point cloud data according to the robot posture data to obtain a mapping action point cloud and a task action path; The inaccessible point determination unit is used to determine a plurality of inaccessible action points in the task action path according to the three-dimensional environment point cloud data and the mapped action point cloud; The action path correction unit is used to respectively find the nearest reachable points corresponding to the multiple unreachable action points, and based on the nearest reachable points, correct the task action path to obtain a target action path; The action sequence adjustment unit is used to adjust the initial action sequence according to the target action path to obtain a target action sequence.

10. A robot control device combining scenario and strategy model according to claim 6, characterized in that: The robot control module includes a joint angle calculation unit, a drive signal determination unit and a robot control unit; The joint angle calculation unit is used to determine the task posture sequence according to the target action sequence, and calculate the target joint angle sequence corresponding to the task posture sequence based on the inverse kinematics algorithm; The drive signal determination unit is used to determine the drive signal sequence according to the target joint angle sequence in combination with the preset action control instructions of the robot to be controlled; The robot control unit is used to control the robot to be controlled according to the drive signal sequence.

Citation Information

Patent Citations

  • Method for opening robot controller bottom layer position instruction interface

    CN107065682A

  • Autonomous pose measurement method based on SLAM technology

    CN112902953A

  • SLAM autonomous navigation method and device of mobile robot

    CN115200588A

  • Man-machine cooperation assembly method and device, terminal and storage medium

    CN118386226A

  • Unmanned vehicle and unmanned vehicle control method and device

    CN119105500A

Cited By

  • Robot motion control method based on digital twinning and electronic equipment

    CN120552084A

  • A robot motion control method and electronic device based on digital twinning

    CN120552084B

  • Task processing method for robot with body based on multi-modal perception and robot system

    CN121650021A

  • Embodied robot task processing method based on multi-modal perception and robot system

    CN121650021B

  • Humanoid robot control method based on generative model and reinforcement learning and related equipment

    CN121857307A