Multi-strategy based robot control method and device, electronic equipment and storage medium

CN122807903APending Publication Date: 2026-09-25CHONGQING PHOENIX TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611144969.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

上述各个控制策略之间的软件框架通常互不兼容,难以统一维护

Benefits of technology

[0020]关于第二方面至第四方面中任意方面的有益效果可以参见第一方面的有益效果的描述,为避免重复,对此不再赘述。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807903A_ABST
    Figure CN122807903A_ABST
Patent Text Reader

Abstract

The application provides a robot control method and device based on multiple strategies, electronic equipment and storage medium, and relates to the technical field of robot control. The method is applied to a controller of a robot. The controller comprises a state machine and multiple state subclass objects. Different state subclass objects correspond to different control strategies. The steps of the method comprise: calling a running function defined by a base class interface through the state machine; executing a built-in running function in a first state subclass object based on a first state subclass object pointed to by a state pointer of the state machine, so as to read a current state vector in a first buffer and obtain a control parameter for controlling the robot; and the current state vector is a state vector of a latest frame. The built-in running function in the first state subclass object is used to encapsulate control parameter determination logic corresponding to a first control strategy. The application can realize unified maintenance of originally incompatible control strategies in the same framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot control technology, and in particular to a multi-strategy-based robot control method, device, electronic device, and storage medium. Background Technology

[0002] There are various control strategies for legged robot motion control, such as traditional model control, reinforcement learning control, and mimic learning control. Traditional model control controls the robot by manually setting fixed control parameters (e.g., control parameters corresponding to multiple postures), offering stability and interpretability, but struggling to implement complex dynamic gaits. Reinforcement learning (RL) control relies on simulation training, transferring simulation data to real-world implementation, and excels at handling complex dynamic gaits. Mimic learning (ML) control relies on motion capture data to reproduce fixed actions. The software frameworks of these different control strategies are often incompatible, making unified maintenance difficult. Summary of the Invention

[0003] The purpose of this application is to provide a robot control method, device, electronic device, and storage medium based on multiple strategies to solve the above-mentioned technical problems.

[0004] Firstly, a multi-strategy-based robot control method is provided, applied to a robot controller. The controller includes a state machine and multiple state subclass objects, with different state subclass objects corresponding to different control strategies. The method includes: The state machine is used to call the execution function defined in the base class interface; Based on the first state subclass object pointed to by the state pointer of the state machine, the running function built into the first state subclass object is executed to read the current state vector in the first buffer and obtain the control parameters used to control the robot. The current state vector is the state vector of the latest frame, and the running function built into the first state subclass object is used to encapsulate the control parameter determination logic corresponding to the first control strategy.

[0005] Optionally, the method further includes: In response to a target switching instruction, the exit function defined by the base class interface is called through the state machine; wherein, the target switching instruction is used to indicate switching to the second control strategy; Based on the first state subclass object pointed to by the state pointer of the state machine, the exit function built into the first state subclass object is executed to perform the exit operation of the first state subclass object; Set the state pointer to point to the second state subclass object corresponding to the second control strategy; The state machine calls the entry function defined in the base class interface. Based on the second state subclass object pointed to by the state pointer of the state machine, the entry function built into the second state subclass object is executed to perform the initialization operation of the second state subclass object.

[0006] Optionally, the response to the target switching instruction further includes: Obtain the initial switching instruction input by the user, the initial switching instruction being used to indicate switching to the second control strategy; The system uses the built-in running function in the first state subclass object and the initial switching instruction to determine whether to switch to the second control strategy; wherein, the built-in running function in the first state subclass object is also used to encapsulate the state switching logic corresponding to the first control strategy. If so, then generate the target switching instruction.

[0007] Optionally, the target switching instruction includes an identifier of the second control strategy, and the step of calling the exit function defined by the base class interface through the state machine in response to the target switching instruction includes: In response to the target switching instruction, the state machine calls the state check function defined by the base class interface; Based on the first state subclass object pointed to by the state pointer of the state machine, the state check function built into the first state subclass object is executed to obtain the identifier of the first control strategy; Based on the identifier of the first control strategy, the identifier corresponding to the second control strategy, and the preset strategy identifier switching order, determine whether to switch to the second control strategy; If so, the exit function defined in the base class interface is called through the state machine.

[0008] Optionally, when the first control strategy is a neural network-based control strategy, the running function built into the first state subclass object performs the following steps: Based on the current state vector, determine the current observation vector required by the first control strategy; According to the first-in-first-out caching method, the current observation vector is written into the second buffer. The second buffer is used to maintain multiple frames of observation vectors, including the current observation vector and at least one frame of historical observation vectors before the current observation vector. Based on the data assembly mode of the first control strategy, the corresponding assembly strategy is called to stitch the multi-frame observation vectors together to obtain stitched data; wherein, the data assembly mode of the first control strategy is recorded in the strategy configuration file of the first control strategy. The spliced ​​data is input into the neural network corresponding to the first control strategy to obtain the control parameters.

[0009] Optionally, based on the data assembly mode of the first control strategy, the corresponding assembly strategy is invoked to stitch the multi-frame observation vectors together to obtain stitched data, including: When the data assembly mode is the first mode, the first assembly strategy is invoked to concatenate the multi-frame observation vectors according to the arrangement order from the current frame to the historical frames to obtain the first concatenated data; When the data assembly mode is the second mode, the second assembly strategy is invoked to concatenate the multi-frame observation vectors according to the arrangement order from the historical frame to the current frame, so as to obtain the second concatenated data; When the data assembly mode is the third mode, the third assembly strategy is invoked. For each observation mode in the multi-frame observation vector, the vectors are spliced ​​in the order from the historical frame to the current frame to obtain the vector to be spliced ​​for each observation mode. The vectors to be spliced ​​for each observation mode are then spliced ​​to obtain the third spliced ​​data.

[0010] Optionally, the controller has a first thread for calling the run function, and executing the run function built into the first state subclass object includes: Within the target thread, the built-in execution function in the first state subclass object is executed; wherein, if the first control strategy is a neural network-based control strategy, the target thread is a pre-built second thread; otherwise, it is the first thread. The first thread and the second thread are asynchronous, and the execution frequency of the first thread is greater than that of the second thread; and, after obtaining the control parameters, the process further includes: Within the target thread, the control parameters are written into the first buffer; Within the first thread, the latest control parameters are read from the first buffer and sent to the robot's actuator.

[0011] Secondly, a multi-strategy-based robot control device is also provided, applied to a robot controller. The controller includes a state machine and multiple state subclass objects, with different state subclass objects corresponding to different control strategies. The device includes: The first calling module is used to call the execution function defined by the base class interface through the state machine; The first execution module is used to execute the built-in running function in the first state subclass object based on the state pointer of the state machine, so as to read the current state vector in the first buffer and obtain the control parameters for controlling the robot. The current state vector is the state vector of the latest frame, and the running function built into the first state subclass object is used to encapsulate the control parameter determination logic corresponding to the first control strategy.

[0012] Optionally, the device further includes: The second calling module is used to call the exit function defined by the base class interface through the state machine in response to the target switching instruction; wherein, the target switching instruction is used to indicate switching to the second control strategy; The second execution module is used to execute the exit function built into the first state subclass object based on the state pointer of the state machine, so as to perform the exit operation of the first state subclass object; The pointer switching module is used to point the state pointer to the second state subclass object corresponding to the second control strategy; The third calling module is used to call the entry function defined by the base class interface through the state machine; The third execution module is used to execute the entry function built into the second state subclass object based on the state pointer of the state machine, so as to perform the initialization operation of the second state subclass object.

[0013] Optionally, the device further includes: The acquisition module is used to acquire the initial switching instruction input by the user, the initial switching instruction being used to indicate switching to the second control strategy; The judgment module is used to determine whether to switch to the second control strategy by using the running function built into the first state subclass object and the initial switching instruction; wherein, the running function built into the first state subclass object is also used to encapsulate the state switching logic corresponding to the first control strategy; The instruction generation module is used to generate the target switching instruction if the condition is met.

[0014] Optionally, the target switching instruction includes an identifier of the second control strategy, and the second calling module is specifically used for: In response to the target switching instruction, the state machine calls the state check function defined by the base class interface; Based on the first state subclass object pointed to by the state pointer of the state machine, the state check function built into the first state subclass object is executed to obtain the identifier of the first control strategy; Based on the identifier of the first control strategy, the identifier corresponding to the second control strategy, and the preset strategy identifier switching order, determine whether to switch to the second control strategy; If so, the exit function defined in the base class interface is called through the state machine.

[0015] Optionally, when the first control strategy is a neural network-based control strategy, the first execution module is specifically used for: Based on the current state vector, determine the current observation vector required by the first control strategy; According to the first-in-first-out caching method, the current observation vector is written into the second buffer. The second buffer is used to maintain multiple frames of observation vectors, including the current observation vector and at least one frame of historical observation vectors before the current observation vector. Based on the data assembly mode of the first control strategy, the corresponding assembly strategy is called to stitch the multi-frame observation vectors together to obtain stitched data; wherein, the data assembly mode of the first control strategy is recorded in the strategy configuration file of the first control strategy. The spliced ​​data is input into the neural network corresponding to the first control strategy to obtain the control parameters.

[0016] Optionally, the first execution module is further configured to: When the data assembly mode is the first mode, the first assembly strategy is invoked to concatenate the multi-frame observation vectors according to the arrangement order from the current frame to the historical frames to obtain the first concatenated data; When the data assembly mode is the second mode, the second assembly strategy is invoked to concatenate the multi-frame observation vectors according to the arrangement order from the historical frame to the current frame, so as to obtain the second concatenated data; When the data assembly mode is the third mode, the third assembly strategy is invoked. For each observation mode in the multi-frame observation vector, the vectors are spliced ​​in the order from the historical frame to the current frame to obtain the vector to be spliced ​​for each observation mode. The vectors to be spliced ​​for each observation mode are then spliced ​​to obtain the third spliced ​​data.

[0017] Optionally, the controller includes a first thread for calling the execution function, and the first execution module is further configured to: Within the target thread, the built-in execution function in the first state subclass object is executed; wherein, if the first control strategy is a neural network-based control strategy, the target thread is a pre-constructed second thread, otherwise, it is the first thread, the first thread is asynchronous with the second thread, and the execution frequency of the first thread is greater than the execution frequency of the second thread; The device further includes: A caching module is used to write the control parameters into the first buffer within the target thread; The reading module is used to read the latest control parameters from the first buffer within the first thread and send the latest control parameters to the actuator of the robot.

[0018] Thirdly, an electronic device is also provided, including a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program to implement any of the methods described above.

[0019] Fourthly, a computer-readable storage medium is also provided, the computer-readable storage medium storing a computer program that, when executed by at least one processor, implements the method as described above.

[0020] For the beneficial effects of any of the second to fourth aspects, please refer to the description of the beneficial effects of the first aspect. To avoid repetition, it will not be repeated here.

[0021] This application abstracts control strategies into independent states managed by a state machine. The execution functions built into the state subclass objects encapsulate the control parameter determination logic for the corresponding control strategy. When the controller is running, the execution function defined in the base class interface is called through the state machine. This automatically matches the first state subclass object pointed to by the current state pointer and executes the control parameter determination logic corresponding to the execution function built into the first state subclass object. In this way, by implementing control strategies as state subclass objects, a unified abstraction and standardized management of heterogeneous control strategies is achieved. By calling the base class interface and adjusting the state pointer, the corresponding control strategy can be executed, enabling the unified maintenance of originally incompatible control strategies within the same framework. Attached Figure Description

[0022] Figure 1 A flowchart illustrating the multi-strategy-based robot control method provided in this application embodiment; Figure 2 A schematic diagram of the structure of a multi-strategy-based robot control device provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0024] The terms "first" and "second" in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the term "comprising" and any variations thereof are intended to cover non-exclusive protection. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. The term "multiple" in this application can mean at least two, for example, two, three, or more, and the embodiments of this application do not impose limitations.

[0025] This application provides a multi-strategy-based robot control method, device, electronic device, and storage medium to solve the problem that the software frameworks of various control strategies are usually incompatible and difficult to maintain in a unified manner.

[0026] The above method is applied to a robot controller, which includes a state machine and multiple state subclass objects. Different state subclass objects correspond to different control strategies, and the control strategy is managed as a state of the state machine. Specifically, the state machine can manage the "entering a state - switching states - exiting a state" process, and further, it can manage the "entering a control strategy - switching control strategies - exiting a control strategy" process. The state pointer of the state machine points to the state subclass object that should be executed. By calling the execution function defined in the base class interface, the execution function implemented in the state subclass object pointed to by the state pointer can be called, thereby executing the control parameter determination logic of the corresponding control strategy.

[0027] For example, the state machine described above can be a finite state machine (FSM).

[0028] The following is a detailed description of the multi-strategy-based robot control method according to embodiments of this application.

[0029] See Figure 1 , Figure 1 This is a flowchart illustrating a multi-strategy-based robot control method provided in this application, as shown below. Figure 1 As shown, the method includes: Step 101: Call the runtime function defined in the base class interface through the state machine.

[0030] Among them, the state subclass is a class definition that inherits the above base class interface and overrides the execution function to implement a single control strategy logic.

[0031] A state subclass object is a runtime entity generated after the corresponding state subclass is instantiated. It is used to carry the control parameter determination logic of the corresponding control strategy and execute the specific control process.

[0032] The above-mentioned running function can be represented as onRun. During each cycle of the control parameters, i.e. each control cycle of the robot, the running function defined by the above base class interface can be called.

[0033] Step 102: Based on the state pointer of the state machine pointing to the first state subclass object, execute the built-in running function in the first state subclass object to read the current state vector in the first buffer and obtain the control parameters used to control the robot.

[0034] After calling the execution function of the base class interface mentioned above, based on the polymorphic feature, the first state subclass object bound to the state pointer can be automatically matched, and the logic of the execution function inside that object can be executed.

[0035] The execution function of the first state subclass object first reads the robot's latest frame state vector, which is pre-stored in the first buffer. This state vector includes the robot's sensor data, specifically joint data (i.e., motor data corresponding to joint positions) and pose data. The pose data can be acquired by an Inertial Measurement Unit (IMU).

[0036] For example, the above state vector It can be represented as: ; in, These represent the motor's rotation angle, speed, and torque, respectively. These represent pose, angular velocity, and linear velocity, respectively.

[0037] The running function implemented by the first state subclass object can determine the logical output of the aforementioned control parameters based on the control parameters of the first control strategy.

[0038] Among the robot's control strategies, there is a first type of control strategy that outputs fixed control parameters without needing to combine them with the current state vector; and there is also a second type of control strategy that needs to combine them with the current state vector to output dynamic control parameters.

[0039] For the first type of control strategy, which does not require the current state vector to be used to output control parameters, the current state vector can be used as input to other operational logic of the control strategy. For example, the running function can encapsulate state switching logic, and the current state vector can be used to determine whether a state switch can be performed.

[0040] The execution functions implemented by each state subclass object uniformly read the real-time state vector stored in the first buffer, which can standardize the data input process and improve the framework's versatility.

[0041] For example, the first type of control strategy mentioned above includes passive mode control strategy and force-based fixed posture control strategy, while the second type of control strategy includes motion control strategy based on reinforcement learning (RL) and mimic control strategy based on shadow following.

[0042] Among them, the passive mode control strategy outputs fixed control parameters, which can be null values, indicating that no active control is performed on the robot.

[0043] Force-based fixed posture control strategies include force-based fixed standing posture control strategies and force-based fixed prone posture control strategies. The fixed standing posture control strategy outputs fixed control parameters corresponding to the standing posture, used to control the robot to maintain a standing posture; the fixed prone posture control strategy outputs fixed control parameters corresponding to the prone posture, used to control the robot to maintain a prone posture.

[0044] Motion control strategies based on reinforcement learning require determining the current observation vector by combining the current state vector. The current observation vector is the input data required by the neural network corresponding to the reinforcement learning.

[0045] The shadow-following imitation learning control strategy also needs to combine the current state vector to determine the current observation vector, which serves as the input data required by the neural network corresponding to the imitation learning.

[0046] This application's embodiments abstract control strategies into independent states managed by a state machine, encapsulating the logic of each control strategy (including control parameter determination logic) into corresponding state subclasses. Each state subclass inherits the same base class interface and overrides its execution function. After being instantiated as a state subclass object, each state subclass object has a built-in execution function for the corresponding control strategy. During program execution, the state machine calls the execution function defined in the base class interface, automatically matching the first state subclass object pointed to by the current state pointer, and executing the control logic corresponding to the built-in execution function of the first state subclass object. Therefore, this application's embodiments, based on a state machine, achieve unified abstraction and standardized management of heterogeneous control strategies, allowing previously incompatible control strategies to be incorporated into a unified framework for maintenance.

[0047] The state machine switches control strategies by changing the state subclass object executed by switching state pointers. Therefore, in some embodiments, the method further includes: In response to a target switching instruction, the exit function defined in the base class interface is called via the state machine; the target switching instruction is used to indicate switching to the second control strategy. Based on the state pointer of the state machine pointing to the first state subclass object, execute the built-in exit function in the first state subclass object to perform the exit operation of the first state subclass object; Set the state pointer to the second state subclass object corresponding to the second control strategy; The state machine calls the entry function defined in the base class interface. Based on the state pointer of the state machine pointing to the second state subclass object, the built-in entry function in the second state subclass object is executed to perform the initialization operation of the second state subclass object.

[0048] The exit function mentioned above can be represented as (onExit), and the entry function mentioned above can be represented as (onEnter).

[0049] For example, the exit operation described above includes releasing relevant resources. For instance, for a neural network-based control strategy, which typically requires outputting control parameters using multi-frame observation vectors, a second buffer and a second thread need to be created during its inference process (the description of the second buffer and the second thread can be found in the following embodiments, and will not be repeated here to avoid repetition). Therefore, the exit operation also includes safely stopping and destroying the current second thread, as well as cleaning up the second buffer.

[0050] For example, the initialization operation described above, for non-neural network-based control strategies, typically does not require the use of multi-frame observation vectors to output control parameters. The initialization operation includes loading relevant configurations, such as loading relevant fixed control parameters. For neural network-based control strategies, which typically require the use of multi-frame observation vectors to output control parameters, the initialization operation also includes launching an asynchronous second thread and loading the strategy configuration file to initialize the second buffer and load the corresponding Open Neural Network Exchange (ONNX) model file.

[0051] The strategy configuration file mentioned above includes the model name M, the observation dimension D of the single-frame observation vector, the logical joint degrees of freedom N (i.e., the number of logical joints of the robot), and the history length H (i.e., the number of frames of the observation vector).

[0052] The second buffer of the corresponding size can be initialized using D, N, and H in the policy configuration file. Based on the model name in the policy configuration file, the corresponding Open Neural Network Exchange (ONNX) model file can be loaded to obtain the neural network for subsequent inference.

[0053] In this implementation, upon receiving a target switching instruction, the state machine schedules according to a unified lifecycle process: first, it executes the exit operation of the current state subclass object; then, it switches the state pointer to point to the second state subclass object corresponding to the second control strategy; and finally, it executes the built-in entry function of the second state subclass object to complete initialization. Through this scheduling logic, seamless hot switching between control strategies can be achieved.

[0054] The aforementioned target switching command can be an initial switching command input by the user, meaning that the control strategy will be switched immediately upon receiving the initial switching command.

[0055] The aforementioned target switching instruction can be the target switching instruction after the first state subclass object is determined; that is, the switching instruction output when the first state subclass object determines that switching can be performed based on the initial switching instruction. This helps to improve the security of control strategy switching.

[0056] Based on this, in some embodiments, prior to responding to the target switching instruction, the method further includes: Obtain the initial switching command input by the user, which is used to indicate switching to the second control strategy; The system uses the built-in execution function in the first state subclass object and the initial switching instruction to determine whether to switch to the second control strategy; the built-in execution function in the first state subclass object is also used to encapsulate the state switching logic corresponding to the first control strategy. If so, then generate a target switching instruction.

[0057] The aforementioned initial switching command can be a command issued by the user through input devices such as gamepad buttons or keyboard.

[0058] For example, the initial switching instruction can be appended to the aforementioned state vector for unified reading by the running functions. Based on this, the aforementioned state vector can be represented as: ; in, This indicates the aforementioned initial switching command.

[0059] For each state subclass object of the control strategy, its built-in execution function also includes the corresponding state switching logic. In this way, it is possible to determine whether the control strategy can be switched based on the read current state vector. For example, if the robot's posture is stable according to the current state vector, the control strategy can be switched; if the robot's posture is unstable according to the current state vector, the control strategy cannot be switched.

[0060] In this implementation, by setting it as described above, upon receiving an initial switching command input by the user, further determination is made based on the state subclass object to determine whether to switch, which helps to improve the security of control strategy switching.

[0061] In some embodiments, the target switching instruction includes an identifier of a second control strategy. In response to the target switching instruction, an exit function defined by the base class interface is called via a state machine, including: In response to a target switching instruction, the state check function defined in the base class interface is called through the state machine; Based on the state pointer of the state machine pointing to the first state subclass object, the built-in state check function in the first state subclass object is executed to obtain the identifier of the first control strategy; Based on the identifier of the first control strategy and the identifier corresponding to the second control strategy, as well as the preset strategy identifier switching order, determine whether to switch to the second control strategy; If so, the exit function defined in the base class interface is called through the state machine.

[0062] The aforementioned state check function can be represented as `checkState`. Calling `checkState` returns an identifier for the first control policy. For example, the identifier can be an enumeration value. There is a one-to-one correspondence between control policies and state subclass objects; therefore, the identifier of the first control policy can also be understood as the identifier of the first state subclass object.

[0063] The aforementioned strategy identifier switching order is used to indicate which control strategies can be switched from the current control strategy. For example, the aforementioned strategy identifier switching order may be, for instance, switching from the identifier of the fixed prone posture control strategy to the identifier of the fixed standing posture control strategy, and switching from the identifier of the fixed standing posture control strategy to the identifier of any other control strategy, etc.

[0064] In this implementation, by further combining the policy identifier switching order to determine whether to switch to the second control policy, the security of policy switching can be further improved.

[0065] When the first control strategy is a neural network-based control strategy, it is necessary to determine the observation vector based on the state vector, and then concatenate the observation vectors of multiple frames to obtain the concatenated data, which is then input into the neural network to obtain the output control parameters.

[0066] Learning-based control strategies (such as reinforcement learning or imitation learning mentioned above) typically require concatenating the observation vectors from the past H frames (e.g., H=5) into the input of the neural network to provide temporal context. , Let represent the observation vector at time t. When training the aforementioned neural network, the order in which the observation vectors are assembled differs across training platforms. In related technologies, modifying the source code logic for observation assembly is necessary for different assembly strategies. This not only increases the operational cost of switching assembly strategies but also easily leads to a decrease in the accuracy of the control parameters output by the neural network due to format mismatches.

[0067] Based on this, in some embodiments, when the first control strategy is a neural network-based control strategy, the built-in execution function in the first state subclass object performs the following steps: Read the current state vector from the first buffer. The current state vector is the state vector of the latest frame in the first buffer. Based on the current state vector, determine the current observation vector required by the first control strategy; Following the first-in-first-out caching method, the current observation vector is written to the second buffer. The second buffer is used to maintain multi-frame observation vectors, which include the current observation vector and at least one previous frame of historical observation vectors. Based on the data assembly mode of the first control strategy, the corresponding assembly strategy is called to stitch together the observation vectors of multiple frames to obtain stitched data; wherein, the data assembly mode of the first control strategy is recorded in the strategy configuration file of the first control strategy. The concatenated data is input into the neural network corresponding to the first control strategy to obtain the control parameters.

[0068] The first control strategy is a neural network-based control strategy, such as the RL-based control strategy mentioned above, or a control strategy based on imitation learning.

[0069] The control strategy based on the neural network has a corresponding strategy configuration file, which includes the aforementioned data assembly mode α. For example, the strategy configuration file may also include: model name M, dimension D of the single-frame observation vector, logical joint degrees of freedom N, and history length H.

[0070] The observation vector is the vector required as input to the first control strategy. The built-in execution function in the first state subclass object can determine the required observation vector from the current state vector.

[0071] Taking an RL-based control policy as an example, its observation vector Possible forms: The explanation of each parameter is shown in the table below.

[0072]

[0073] In the above scaling column, , , and All are scaling factors. This indicates that each dimension has an independent scaling factor. This is a coordinate transformation matrix used to transform the gravity vector in the world coordinate system. (That is, vertically downward) transform to the robot coordinate system. This indicates the angles of each joint when the robot is standing.

[0074] Total Observation Dimensions ,when hour, Stacking After frame history, the total input dimension is Provide about Timing context of ms.

[0075] The total observation dimension D, the number of logical joints N, and the history length H can all be recorded in the strategy configuration file for easy modification. Furthermore, when the first state subclass object is initialized, it facilitates the initialization of a second buffer of the corresponding size based on D, N, and H. The second buffer can be understood as a time-series observation buffer.

[0076] The above All can be determined through state vectors, meaning they can also be collected. , and will , Concatenate it to the state vector.

[0077] If the control parameters are written to the first buffer as described in the following embodiments, the control parameters of the previous frame can be... The state vector is concatenated to the current frame, and thus the above can be determined based on the state vector. Alternatively, it can be determined... Then, it is written to the second buffer, along with... The observation vector is obtained by splicing the vectors together.

[0078] After obtaining the current observation vector based on the current state vector, the current observation vector is written into the second buffer according to the First In First Out (FIFO) buffering method. The second buffer contains the H-frame observation vector of the current observation vector.

[0079] In this embodiment, the data assembly mode of the first control strategy is recorded in the corresponding strategy configuration file. In this way, if the neural network of the first control strategy changes, resulting in a change in the assembly strategy of its input vector, only the data assembly mode in the strategy configuration file needs to be modified, without modifying the source code. This helps to reduce the operation cost of switching assembly strategies and improve the output accuracy of the neural network.

[0080] In some embodiments, based on the data assembly mode of the first control strategy, the corresponding assembly strategy is invoked to stitch together multi-frame observation vectors to obtain stitched data, including: When the data assembly mode is in the first mode, the first assembly strategy is invoked, and the observation vectors of multiple frames are concatenated according to the order from the current frame to the historical frames to obtain the first concatenated data. The first concatenated data can be represented as a concatenated vector.

[0081] When the data assembly mode is in the second mode, the second assembly strategy is invoked to concatenate the observation vectors of multiple frames according to the order from historical frames to the current frame, resulting in the second concatenated data. The second concatenated data can be represented as a concatenated vector.

[0082] When the data assembly mode is in the third mode, the third assembly strategy is invoked. For each observation mode in the multi-frame observation vector, the vectors are concatenated according to the order from historical frames to the current frame to obtain the vector to be concatenated for each observation mode. The vectors to be concatenated for each observation mode are then concatenated to obtain the third concatenated data. The third concatenated data can be represented as a concatenation tensor.

[0083] The first mode can be represented as The second mode can be represented as The third mode can be represented as .

[0084] In the first mode, the corresponding first assembly strategy is a header appending type assembly method. That is, historical data in the buffer is shifted to higher addresses, and the current observation is filled to lower addresses. For example, suppose the concatenation vector in the second buffer... Current observation The update operation is as follows: The second buffer layout is as follows: As can be seen, under the first assembly strategy, historical observation vectors flow to the right for updates, with the latest observation occupying the head position of the buffer. The observation vectors are concatenated according to the order from the current frame to the historical frames. This first assembly strategy is typically applicable to neural networks trained with Isaac Gym.

[0085] In the second mode, the corresponding second assembly strategy is a tail-append type assembly method. That is, the historical observation vector in the buffer is shifted to a lower address, and the current observation vector is appended to a higher address. For example, suppose the concatenated vector in the second buffer... Current observation The update operation is as follows: The buffer layout is as follows: As can be seen, under the second assembly strategy, the order of observation vectors is rearranged and the append direction is opposite to that of the first assembly strategy. Historical observation vectors in the second buffer are shifted to lower addresses, and the current observation vector is appended to higher addresses. The observation vectors are concatenated according to the order from historical frames to the current frame. The second assembly strategy is typically applicable to neural networks trained with Mujoco.

[0086] In the case of the third mode, the corresponding third assembly strategy is a channel-independent buffer type assembly method. For example, suppose the observed modes in the second buffer... An observation mode refers to a logical grouping within the state vector, categorized by sensor type or physical meaning. Each group corresponds to an independent data channel. For example, the number of observation modes in the aforementioned RL-based control strategy... =6. The second buffer maintains an independent FIFO buffer for each observation mode. : ,in For modality Dimensions For a moment The Observations of the observed mode, The frames are stitched together according to the order from historical frames to the current frame. For example, in the above RL-based control strategy, , The final input concatenation tensor is constructed by concatenating all channels: ,in Indicates will Expanded column-wise into a one-dimensional vector. The third assembly strategy is typically applicable to neural networks with a multimodal Transformer architecture.

[0087] In this implementation, three data assembly modes are set. Based on the current data assembly mode, the corresponding assembly strategy can be invoked. Only the data assembly mode needs to be modified in the configuration file. The value of (0, 1 or 3) can be adapted to the input of different neural networks, which helps to reduce the operational cost of switching assembly strategies.

[0088] Control strategies based on neural network inference (ONNX Runtime forward propagation) typically take 1-5ms, while high-dynamic motion control of robots usually requires a control frequency of 500Hz (2ms control cycle). In related technologies, inference and control are often simply executed sequentially (inference complete → command issued → inference again → command issued again), which causes the control frequency to be constrained by the inference time, making it difficult to simultaneously meet the needs of high-frequency control and complex model inference.

[0089] Based on this, in some embodiments, the controller is provided with a first thread for calling the execution function, and step 102 includes: Within the target thread, the built-in execution function in the first state subclass object is executed; wherein, if the first control strategy is a neural network-based control strategy, the target thread is a pre-built second thread; otherwise, it is the first thread. The first thread and the second thread are asynchronous, and the execution frequency of the first thread is greater than that of the second thread; and, after step 102, the following is also included: Within the first thread, the latest control parameters are read from the first buffer and sent to the robot's actuators.

[0090] The first thread can be understood as the control thread, and the second thread as the inference thread. The first thread performs closed-loop control in each robot control cycle, while the second thread is only used for the neural network inference process.

[0091] When the first control strategy is a neural network-based control strategy, the initialization process of the first state subclass object also includes constructing the aforementioned asynchronous second thread, where the neural network inference process is performed. That is, within the second thread, data is assembled and concatenated, the ONNX Runtime engine is invoked to execute model inference, and control parameters are written to the first buffer after inference is completed.

[0092] Specifically, the control parameters can be the control parameters of the drive motors corresponding to each logic joint. These control parameters can be target values. After the first thread reads the latest control parameters, it can directly issue the control parameters, and the actuator can determine the corresponding control command based on the closed-loop control coefficients and the deviation between the target value and the actual value. Alternatively, within the first thread, after determining the corresponding control command based on the closed-loop control coefficients and the deviation between the target value and the actual value, the control command can be directly issued to the actuator.

[0093] The first and second threads can synchronize and share data using a mutex lock. While the second thread holds the lock, it only performs data copying (on a microsecond scale) to copy the current state vector from the first buffer. When the second thread does not hold the lock, it can perform neural network inference. While the first thread holds the lock, it only performs data copying (on a microsecond scale) to update the first buffer and read control parameters.

[0094] The execution frequency of the first thread is higher than that of the second thread. For example, the execution frequency of the first thread could be 500Hz (period of 2ms), and the execution frequency of the second thread could be 50Hz (period of 20ms). The first thread has higher priority than the second thread and reads the latest control parameters from the first buffer according to its control cycle, without waiting for the second thread to complete its inference. If the inference thread has not completed its inference in a certain cycle, the control thread will use the valid control parameters from the previous frame, or it will use the valid control parameters from the previous frame and perform a smooth transition through an interpolation algorithm.

[0095] In this implementation, the time-consuming neural network inference is isolated in the second thread by the above settings, which can reduce the blocking or delay of the first thread caused by its operation fluctuations. The first thread is responsible for writing the state vector into the first buffer, calling the state machine and reading the control parameters within its closed-loop control cycle. It does not need to perform neural network inference, and its control frequency is not constrained by the inference time, which is conducive to meeting the needs of high-frequency control and complex model inference at the same time.

[0096] This application also provides a control architecture, including: a hardware layer corresponding to L1, a communication layer corresponding to L2, an IO abstraction layer corresponding to L3, a control layer corresponding to L4, a finite state layer corresponding to L5, and an inference layer corresponding to L6.

[0097] The hardware layer corresponds to the robot body and includes drive motors for N logical joints, an IMU inertial measurement unit, a battery management system, and input devices for controlling the robot.

[0098] The communication layer is used for data forwarding, including forwarding uplink sensor data (LowState) and user input commands, as well as forwarding downlink control parameters or control instructions. The communication layer can be based on the ROS2 Topic publish / subscribe mechanism and forward data via reliable QoS.

[0099] The I / O abstraction layer encapsulates the underlying communication details, providing a unified bidirectional communication interface and event loop mechanism. The I / O abstraction layer provides two core primitives, including... and .in, : Send control parameters or control commands from the previous cycle, and receive status data from the current cycle to fill in the data. This operation is executed as the first step in the main loop of the control thread in each control cycle. : Start the event loop, enter blocking mode and wait for external messages (sensor data or user input) until a termination signal is received.

[0100] Taking the communication layer's interaction based on the ROS2 Topic publish / subscribe mechanism as an example, the IO abstraction layer initialization includes creating ROS2 node entities and registering status topic subscribers (topic name). ) and control of topic publishers (topic identifiers) The Quality of Service (QoS) is configured to retain the most recent 10 messages in a reliable transmission mode. The implementation includes publishing the instruction message for the calculation in the previous cycle. Read the latest sensor data from the shared buffer written by the subscription callback. Read user commands (such as initial switching instructions) and analog values ​​(such as handle command speed) from input devices.

[0101] The existence of the IO abstraction layer decouples the control layer from the communication layer by implementing new... and The adaptable subclasses, such as those for Lightweight Communications and Marshalling (LCM), Cyclone Data Distribution Service (Cyclone DDS), and serial communication, can be migrated to different communication platforms without modifying the control layer logic and state machine code.

[0102] The IO abstraction layer can be executed in the communication callback thread. It can be event-driven, triggered by the arrival of network packets or ROS2 Topic messages, and its priority can be the default priority. It is responsible for receiving and dumping the underlying hardware data. When the sensor data (LowState) published by the underlying communication arrives, the communication callback thread is immediately awakened, the packet is parsed and written into the system shared buffer (or lockless queue), providing the control layer with the latest state data source.

[0103] The control layer executes closed-loop control through the first thread at a fixed control cycle, including reading the state vector from the IO abstraction layer and writing it into the first buffer, calling the state machine, reading control parameters, and issuing control parameters or control commands to the IO abstraction layer.

[0104] A finite state layer contains a finite set of states. ( A state machine can be used to implement the "entry-transition-exit" of states. Each state corresponds to a control strategy, and each state independently maintains its control logic and lifecycle.

[0105] The inference layer, within the inference thread, executes the forward propagation of the neural network through the ONNX Runtime, and only runs when the control policy is a neural network-based control policy.

[0106] Legged robots often employ differentiated hardware designs for various applications, leading to code fragmentation in the control software. Different robot models may exhibit inconsistencies in joint degrees of freedom, fixed posture calibration, and zero-position calibration. Even within the same robot, the encoder reading symbols may differ due to the opposite mounting directions of the left and right leg motors, necessitating additional mapping at the drive layer. Hardware wiring methods also present challenges; for example, the physical CAN ID order of the joint drive motors in the Controller Area Network (CAN) bus often differs from the logical joint order expected by the algorithm layer, requiring the establishment of an additional mapping between CAN IDs and logical joints. Currently, new models rely on copying and modifying for adaptation, resulting in a continuously expanding codebase and increased maintenance and migration costs.

[0107] Based on this, the joint degrees of freedom, fixed posture calibration, zero-position calibration, and mapping relationships mentioned above can be written into the configuration file. In the event of subsequent changes, only the configuration file needs to be modified, and the controller can obtain the latest data by loading the configuration file.

[0108] For example, a hierarchical configuration strategy can be adopted to organize parameters of different dimensions in separate configuration files.

[0109] Robot model configuration. Used to achieve bidirectional mapping between logical joints and their corresponding physical CAN IDs. Includes three parameter groups, as follows: ;in, Total number of physical drives Representing logical joints The corresponding physical CAN identifier.

[0110] ;in, This indicates that the motor rotation direction of the logic joint is opposite to the positive rotation direction defined by the physical sensor; 1 indicates that the directions are the same.

[0111] ;in For logical joints The zero offset (in radians) is used to compensate for the fixed difference between the calibration zero position and the algorithm's default zero position.

[0112] Preset posture and PD control parameter configuration. Define three sets of preset joint angle vectors (in radians): The algorithm defaults to zero bits. : Fixed standing posture; : Fix the prone position. Define two sets of stiffness-damping parameter pairs: : PD parameters corresponding to the control parameters of the fixed attitude; : The PD parameters corresponding to the control parameters output by the neural network.

[0113] In addition, some basic configuration files can be set, which can include topic identifiers for commands and states. and and IMU topic identifier And so on. It also allows setting the aforementioned neural network-based strategy configuration file. The strategy configuration file includes the model name M, the observation dimension D of the single-frame observation vector, the joint degrees of freedom N, the history length H, and the data assembly mode α.

[0114] The above configuration file can be loaded at runtime. If there are any changes, only the configuration file needs to be modified, and the software program does not need to be recompiled.

[0115] The aforementioned bidirectional mapping can be used in the control layer. The input mapping of the control layer includes reading motor data based on the IO abstraction layer and converting it into data corresponding to the logical joints. The output mapping includes determining control parameters or control commands and converting them into data corresponding to the CAN ID. For ease of understanding, the following is an illustrative example.

[0116] Input mapping will physical sensor data space The joint measurements (i.e., the corresponding drive motor data) are converted into the algorithm's observation space. Logical joint values ​​in the context of physical joint position vectors: and velocity vector The logical joint values ​​are obtained through the following transformations: ; ; in For vectors The One portion, For vectors The One portion, For vectors The Each component. This indicates that the positive direction of the physical sensor is opposite to that defined by the algorithm, and is commonly seen in mirror mounting of symmetrical legs; Obtained through offline calibration, it compensates for mechanical installation deviations at the joint zero point.

[0117] Design principles of mapping matrices: mapping dimension This equals the logical joint degrees of freedom expected by the control strategy; The range of values ​​is ,in This represents the total number of physical drives.

[0118] Output mapping is the inverse transformation of input mapping, which transforms the algorithm space. Mapping control parameters back to physical space Given the target value of the logical joint output by the algorithm. The target value of the physical drive motor is obtained through inverse transformation: ; ; in and Logical joints The stiffness and damping values. Among them, the above... Includes fixed attitude corresponding and , Includes the output of the neural network and .

[0119] The output mapping is first subtracted by the offset. Restore the physical zero-position reference frame, then multiply by the direction sign. Matches the physically defined positive rotation direction. The input mapping and output mapping satisfy a reciprocal relationship: .

[0120] After the controller starts, it parses the above configuration file and... , , Loaded as a numerical array. The first thread in the control layer performs the mapping by directly indexing the array within a loop. The computational complexity of a single input or output mapping is O(n log n). And it only involves scalar multiplication and addition, affecting the control cycle. The time taken is less than 1 microsecond. After mapping, the current control strategy can be loaded based on the state of the finite state machine. or Combined with PD closed-loop control, output control commands.

[0121] In this embodiment, by using the above settings and the finite state machine of L5, hot switching of control strategies can be achieved. The asynchronous execution of the first thread of L4 and the second thread of L6 can reduce the interference of the second thread on the first thread, which is beneficial to simultaneously meeting the needs of high-frequency control and complex model inference. The use of the strategy configuration file to record the data assembly mode α can reduce the operation cost of switching assembly strategies and improve the output accuracy of the neural network. The use of the configuration file to record the robot model configuration can help adapt to more robot models and reduce the operation cost of switching robot hardware.

[0122] It should be understood that although the steps in the flowchart above are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart above may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0123] Based on the same inventive concept, please refer to Figure 2 As shown, this embodiment provides a multi-strategy-based robot control device applied to a robot controller. The controller includes a state machine and multiple state subclass objects, with different state subclass objects corresponding to different control strategies. The device includes: The first calling module 201 is used to call the running function defined by the base class interface through the state machine; The first execution module 202 is used to execute the running function built into the first state subclass object based on the state pointer of the state machine, so as to read the current state vector in the first buffer and obtain the control parameters for controlling the robot. The current state vector is the state vector of the latest frame, and the running function built into the first state subclass object is used to encapsulate the control parameter determination logic corresponding to the first control strategy.

[0124] Optionally, the device further includes: The second calling module is used to call the exit function defined by the base class interface through the state machine in response to the target switching instruction; wherein, the target switching instruction is used to indicate switching to the second control strategy; The second execution module is used to execute the exit function built into the first state subclass object based on the state pointer of the state machine, so as to perform the exit operation of the first state subclass object; The pointer switching module is used to point the state pointer to the second state subclass object corresponding to the second control strategy; The third calling module is used to call the entry function defined by the base class interface through the state machine; The third execution module is used to execute the entry function built into the second state subclass object based on the state pointer of the state machine, so as to perform the initialization operation of the second state subclass object.

[0125] Optionally, the device further includes: The acquisition module is used to acquire the initial switching instruction input by the user, the initial switching instruction being used to indicate switching to the second control strategy; The judgment module is used to determine whether to switch to the second control strategy by using the running function built into the first state subclass object and the initial switching instruction; wherein, the running function built into the first state subclass object is also used to encapsulate the state switching logic corresponding to the first control strategy; The instruction generation module is used to generate the target switching instruction if the condition is met.

[0126] Optionally, the target switching instruction includes an identifier of the second control strategy, and the second calling module is specifically used for: In response to the target switching instruction, the state machine calls the state check function defined by the base class interface; Based on the first state subclass object pointed to by the state pointer of the state machine, the state check function built into the first state subclass object is executed to obtain the identifier of the first control strategy; Based on the identifier of the first control strategy, the identifier corresponding to the second control strategy, and the preset strategy identifier switching order, determine whether to switch to the second control strategy; If so, the exit function defined in the base class interface is called through the state machine.

[0127] Optionally, when the first control strategy is a neural network-based control strategy, the first execution module 202 is specifically used for: Based on the current state vector, determine the current observation vector required by the first control strategy; According to the first-in-first-out caching method, the current observation vector is written into the second buffer. The second buffer is used to maintain multiple frames of observation vectors, including the current observation vector and at least one frame of historical observation vectors before the current observation vector. Based on the data assembly mode of the first control strategy, the corresponding assembly strategy is called to stitch the multi-frame observation vectors together to obtain stitched data; wherein, the data assembly mode of the first control strategy is recorded in the strategy configuration file of the first control strategy. The spliced ​​data is input into the neural network corresponding to the first control strategy to obtain the control parameters.

[0128] Optionally, the first execution module 202 is further configured to: When the data assembly mode is the first mode, the first assembly strategy is invoked to concatenate the multi-frame observation vectors according to the arrangement order from the current frame to the historical frames to obtain the first concatenated data; When the data assembly mode is the second mode, the second assembly strategy is invoked to concatenate the multi-frame observation vectors according to the arrangement order from the historical frame to the current frame, so as to obtain the second concatenated data; When the data assembly mode is the third mode, the third assembly strategy is invoked. For each observation mode in the multi-frame observation vector, the vectors are spliced ​​in the order from the historical frame to the current frame to obtain the vector to be spliced ​​for each observation mode. The vectors to be spliced ​​for each observation mode are then spliced ​​to obtain the third spliced ​​data.

[0129] Optionally, the controller is provided with a first thread for calling the running function, and the first execution module 202 is further configured to: Within the target thread, the built-in execution function in the first state subclass object is executed; wherein, if the first control strategy is a neural network-based control strategy, the target thread is a pre-constructed second thread, otherwise, it is the first thread, the first thread is asynchronous with the second thread, and the execution frequency of the first thread is greater than the execution frequency of the second thread; The device further includes: A caching module is used to write the control parameters into the first buffer within the target thread; The reading module is used to read the latest control parameters from the first buffer within the first thread and send the latest control parameters to the actuator of the robot.

[0130] It should be understood that, for the sake of brevity, some of the content described in the previous embodiments will not be repeated in this embodiment.

[0131] Based on the same inventive concept, embodiments of this application provide a computer device, which may be a server, and its internal structure diagram is as follows: Figure 3 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a multi-strategy-based robot control method.

[0132] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0133] Based on the same inventive concept, this application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it performs the following steps: The state machine is used to call the execution function defined in the base class interface; Based on the first state subclass object pointed to by the state pointer of the state machine, the running function built into the first state subclass object is executed to read the current state vector in the first buffer and obtain the control parameters used to control the robot. The current state vector is the state vector of the latest frame, and the running function built into the first state subclass object is used to encapsulate the control parameter determination logic corresponding to the first control strategy.

[0134] Optionally, when the processor executes a computer program, it performs the following steps: In response to a target switching instruction, the exit function defined by the base class interface is called through the state machine; wherein, the target switching instruction is used to indicate switching to the second control strategy; Based on the first state subclass object pointed to by the state pointer of the state machine, the exit function built into the first state subclass object is executed to perform the exit operation of the first state subclass object; Set the state pointer to point to the second state subclass object corresponding to the second control strategy; The state machine calls the entry function defined in the base class interface. Based on the second state subclass object pointed to by the state pointer of the state machine, the entry function built into the second state subclass object is executed to perform the initialization operation of the second state subclass object.

[0135] Optionally, when the processor executes a computer program, it performs the following steps: Obtain the initial switching instruction input by the user, the initial switching instruction being used to indicate switching to the second control strategy; The system uses the built-in running function in the first state subclass object and the initial switching instruction to determine whether to switch to the second control strategy; wherein, the built-in running function in the first state subclass object is also used to encapsulate the state switching logic corresponding to the first control strategy. If so, then generate the target switching instruction.

[0136] Optionally, the target switching instruction includes an identifier of the second control strategy, and the processor executes the following steps when executing the computer program: In response to the target switching instruction, the state machine calls the state check function defined by the base class interface; Based on the first state subclass object pointed to by the state pointer of the state machine, the state check function built into the first state subclass object is executed to obtain the identifier of the first control strategy; Based on the identifier of the first control strategy, the identifier corresponding to the second control strategy, and the preset strategy identifier switching order, determine whether to switch to the second control strategy; If so, the exit function defined in the base class interface is called through the state machine.

[0137] Optionally, when the first control strategy is a neural network-based control strategy, the processor executes the following steps when executing the computer program: determining the current observation vector required by the first control strategy based on the current state vector; According to the first-in-first-out caching method, the current observation vector is written into the second buffer. The second buffer is used to maintain multiple frames of observation vectors, including the current observation vector and at least one frame of historical observation vectors before the current observation vector. Based on the data assembly mode of the first control strategy, the corresponding assembly strategy is called to stitch the multi-frame observation vectors together to obtain stitched data; wherein, the data assembly mode of the first control strategy is recorded in the strategy configuration file of the first control strategy. The spliced ​​data is input into the neural network corresponding to the first control strategy to obtain the control parameters.

[0138] Optionally, when the processor executes a computer program, it performs the following steps: When the data assembly mode is the first mode, the first assembly strategy is invoked to concatenate the multi-frame observation vectors according to the arrangement order from the current frame to the historical frames to obtain the first concatenated data; When the data assembly mode is the second mode, the second assembly strategy is invoked to concatenate the multi-frame observation vectors according to the arrangement order from the historical frame to the current frame, so as to obtain the second concatenated data; When the data assembly mode is the third mode, the third assembly strategy is invoked. For each observation mode in the multi-frame observation vector, the vectors are spliced ​​in the order from the historical frame to the current frame to obtain the vector to be spliced ​​for each observation mode. The vectors to be spliced ​​for each observation mode are then spliced ​​to obtain the third spliced ​​data.

[0139] Optionally, the controller has a first thread for calling the execution function, and the processor performs the following steps when executing the computer program: Within the target thread, the built-in execution function in the first state subclass object is executed; wherein, if the first control strategy is a neural network-based control strategy, the target thread is a pre-constructed second thread; otherwise, it is the first thread. The first thread and the second thread are asynchronous, and the execution frequency of the first thread is greater than that of the second thread; and, after obtaining the control parameters, the processor further implements the following steps when executing the computer program: Within the target thread, the control parameters are written into the first buffer; Within the first thread, the latest control parameters are read from the first buffer and sent to the robot's actuator.

[0140] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0141] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0142] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A multi-strategy-based robot control method, characterized in that, A controller for a robot, comprising a state machine and multiple state subclass objects, each corresponding to a different control strategy, the method comprising: The state machine is used to call the execution function defined in the base class interface; Based on the first state subclass object pointed to by the state pointer of the state machine, the running function built into the first state subclass object is executed to read the current state vector in the first buffer and obtain the control parameters used to control the robot. The current state vector is the state vector of the latest frame, and the running function built into the first state subclass object is used to encapsulate the control parameter determination logic corresponding to the first control strategy.

2. The multi-strategy-based robot control method according to claim 1, characterized in that, The method further includes: In response to a target switching instruction, the exit function defined by the base class interface is called through the state machine; wherein, the target switching instruction is used to indicate switching to the second control strategy; Based on the first state subclass object pointed to by the state pointer of the state machine, the exit function built into the first state subclass object is executed to perform the exit operation of the first state subclass object; Set the state pointer to point to the second state subclass object corresponding to the second control strategy; The state machine calls the entry function defined in the base class interface. Based on the second state subclass object pointed to by the state pointer of the state machine, the entry function built into the second state subclass object is executed to perform the initialization operation of the second state subclass object.

3. The multi-strategy-based robot control method according to claim 2, characterized in that, The response prior to the target switching command also includes: Obtain the initial switching instruction input by the user, the initial switching instruction being used to indicate switching to the second control strategy; The system uses the built-in running function in the first state subclass object and the initial switching instruction to determine whether to switch to the second control strategy; wherein, the built-in running function in the first state subclass object is also used to encapsulate the state switching logic corresponding to the first control strategy. If so, then generate the target switching instruction.

4. The multi-strategy-based robot control method according to claim 2 or 3, characterized in that, The target switching instruction includes an identifier of the second control strategy. The step of responding to the target switching instruction by calling the exit function defined in the base class interface through the state machine includes: In response to the target switching instruction, the state machine calls the state check function defined by the base class interface; Based on the first state subclass object pointed to by the state pointer of the state machine, the state check function built into the first state subclass object is executed to obtain the identifier of the first control strategy; Based on the identifier of the first control strategy, the identifier corresponding to the second control strategy, and the preset strategy identifier switching order, determine whether to switch to the second control strategy; If so, the exit function defined in the base class interface is called through the state machine.

5. The method according to claim 1, characterized in that, When the first control strategy is a neural network-based control strategy, the built-in running function in the first state subclass object performs the following steps: Based on the current state vector, determine the current observation vector required by the first control strategy; According to the first-in-first-out caching method, the current observation vector is written into the second buffer. The second buffer is used to maintain multiple frames of observation vectors, including the current observation vector and at least one frame of historical observation vectors before the current observation vector. Based on the data assembly mode of the first control strategy, the corresponding assembly strategy is called to stitch the multi-frame observation vectors together to obtain stitched data; wherein, the data assembly mode of the first control strategy is recorded in the strategy configuration file of the first control strategy. The spliced ​​data is input into the neural network corresponding to the first control strategy to obtain the control parameters.

6. The method according to claim 5, characterized in that, Based on the data assembly mode of the first control strategy, the corresponding assembly strategy is invoked to stitch the multi-frame observation vectors together to obtain stitched data, including: When the data assembly mode is the first mode, the first assembly strategy is invoked to concatenate the multi-frame observation vectors according to the arrangement order from the current frame to the historical frames to obtain the first concatenated data; When the data assembly mode is the second mode, the second assembly strategy is invoked to concatenate the multi-frame observation vectors according to the arrangement order from the historical frame to the current frame, so as to obtain the second concatenated data; When the data assembly mode is the third mode, the third assembly strategy is invoked. For each observation mode in the multi-frame observation vector, the vectors are spliced ​​in the order from the historical frame to the current frame to obtain the vector to be spliced ​​for each observation mode. The vectors to be spliced ​​for each observation mode are then spliced ​​to obtain the third spliced ​​data.

7. The multi-strategy-based robot control method according to claim 1 or 5, characterized in that, The controller has a first thread for calling the run function, and executing the run function built into the first state subclass object includes: Within the target thread, the built-in execution function in the first state subclass object is executed; wherein, if the first control strategy is a neural network-based control strategy, the target thread is a pre-constructed second thread; otherwise, it is the first thread. The first thread and the second thread are asynchronous, and the execution frequency of the first thread is greater than that of the second thread; and, After obtaining the control parameters, the process also includes: Within the target thread, the control parameters are written into the first buffer; Within the first thread, the latest control parameters are read from the first buffer and sent to the robot's actuator.

8. A robot control device based on multi-strategy, characterized in that, A controller for robots, comprising a state machine and multiple state subclass objects, each corresponding to a different control strategy, the device comprising: The first calling module is used to call the execution function defined by the base class interface through the state machine; The first execution module is used to execute the built-in running function in the first state subclass object based on the state pointer of the state machine, so as to read the current state vector in the first buffer and obtain the control parameters for controlling the robot. The current state vector is the state vector of the latest frame, and the running function built into the first state subclass object is used to encapsulate the control parameter determination logic corresponding to the first control strategy.

9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by at least one processor, implements the method as described in any one of claims 1-7.