Base station speed estimation method, base station speed estimation network training method, device, equipment, medium and program product
Patent Information
- Application Number
- CN202610733439.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-09-01
AI Technical Summary
[0004]本申请实施例提供一种基座速度估计方法、基座速度估计网络的训练方法、装置、计算机设备、计算机可读存储介质、计算机程序产品,以解决或缓解上面提出的一项或更多项技术问题
通过获取连续多个控制周期分别对应的本体感知观测数据,并使本体感知观测数据同时包括双轮足底盘子系统信号、机械臂子系统信号以及历史动作信号,可以将底盘运动状态、机械臂运动状态和前序控制动作纳入同一历史序列表示。进一步地,通过将本体感知历史序列输入基座速度估计网络得到基座线速度估计值,使基座线速度估计过程能够利用联合系统的历史状态变化信息,而不是仅基于单一时刻或单一子系统信号进行估计。由此,可以提高基座线速度估计与双轮足机械臂联合系统当前运动状态的匹配程度,为全身协调控制提供更适配的速度输入。
Smart Images

Figure CN122672301A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control technology, and in particular to a base velocity estimation method, a training method for a base velocity estimation network, an apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] Bicycle-legged robots can achieve movement and balance control through the combination of wheeled and legged structures. In scenarios where a bicycle-legged robot is equipped with a robotic arm to form a mobile operating platform, the movement of the robotic arm will change the center of mass, inertia, and chassis force state of the combined system. Base linear velocity is usually an important input for whole-body coordinated control, but in some combined systems, base linear velocity is difficult to measure directly by conventional body sensors. When the robotic arm performs acceleration / deceleration, spatial swinging, or object picking / releasing operations, the relationship between chassis attitude, wheel speed, joint state, and base linear velocity may change with the robotic arm configuration and load state, leading to deviations between the base velocity estimation and the actual motion state of the combined system. Therefore, improving the matching degree between base velocity estimation and the motion state of the combined system in a bicycle-legged robotic arm system has become one of the technical problems that need to be solved.
[0003] It should be noted that the above content is not necessarily prior art, nor is it intended to limit the scope of patent protection of this application. Summary of the Invention
[0004] This application provides a base velocity estimation method, a base velocity estimation network training method, an apparatus, a computer device, a computer-readable storage medium, and a computer program product to solve or alleviate one or more of the technical problems mentioned above.
[0005] One aspect of this application provides a base speed estimation method applied to a control device of a dual-wheeled manipulator system, the dual-wheeled manipulator system including a dual-wheeled chassis and a manipulator mounted on the dual-wheeled chassis, the method comprising: Acquire body perception observation data corresponding to multiple consecutive control cycles. The body perception observation data includes signals from the dual-wheel foot plate system, signals from the robotic arm subsystem, and historical motion signals. Based on the ontology perception observation data corresponding to the multiple consecutive control cycles, an ontology perception history sequence is generated. The body perception history sequence is input into the base velocity estimation network to obtain the base linear velocity estimate of the dual-wheeled manipulator joint system; The estimated base linear velocity is output, and the estimated base linear velocity is used for the whole-body coordinated control of the dual-wheeled legged robotic arm combined system.
[0006] Optionally, acquire the ontology-sensing observation data corresponding to multiple consecutive control cycles, including: Acquire the base angular velocity signal and gravity projection vector provided by the inertial measurement unit; Obtain the angular position deviation of the leg joints relative to the default posture; Acquire angular velocity signals from the leg joints and drive wheels; The dual-wheel footplate system signal is generated based on the base angular velocity signal, the gravity projection vector, the angular position deviation of the leg joint relative to the default posture, and the angular velocity signals of the leg joint and the drive wheel.
[0007] Optionally, acquire the ontology-sensing observation data corresponding to multiple consecutive control cycles, including: Obtain the angular position deviation of the robotic arm joints relative to the default posture; Obtain the angular velocity signals of the robotic arm joints; Obtain the end position of the robotic arm's end effector in the base alignment coordinate system; Obtain the end target position corresponding to the end effector of the robotic arm; Based on the end position and the end target position, an end position error is generated; The robotic arm subsystem signal is generated based on the angular position deviation of the robotic arm joint relative to the default posture, the angular velocity signal of the robotic arm joint, the end position, the end target position, and the end position error.
[0008] Optionally, acquire the ontology-sensing observation data corresponding to multiple consecutive control cycles, including: Obtain the full-body control motion vector corresponding to the previous control cycle; The whole-body control motion vector is used as the historical motion signal; The whole-body control motion vector includes chassis control motion and robotic arm control motion.
[0009] Optionally, based on the ontology-sensing observation data corresponding to the consecutive multiple control cycles, an ontology-sensing historical sequence is generated, including: According to the time sequence of the control cycle, acquire multiple body sensing observation data within a preset historical time window; The multiple ontology-sensing observation data are stitched together to generate the ontology-sensing historical sequence. The ontology perception history sequence is used to characterize the joint temporal change pattern between the signals of the dual-wheel footplate system and the signals of the robotic arm subsystem.
[0010] Optionally, the preset historical time window includes multiple continuous control cycles, wherein the ontology perception observation data corresponding to each control cycle is a preset number of dimensions; The window length of the preset historical time window satisfies: L·Δt≥τ_rise; Where L is the number of historical time steps, Δt is the sampling interval, and τ_rise is the rise time constant of the chassis linear velocity under the acceleration command; The preset historical time window is also used to cover the duration of the reaction force caused by the robotic arm's operation.
[0011] Optionally, the ontology perception history sequence is input into the base velocity estimation network to obtain the base linear velocity estimate of the dual-wheeled robotic arm joint system, including: The ontological perception history sequence is input into the base velocity estimation network of the multilayer perceptron structure; The base velocity estimation network outputs a three-dimensional base linear velocity estimate. The three-dimensional base linear velocity estimate includes estimates of the base linear velocity in three axial directions.
[0012] Optionally, after outputting the estimated base linear velocity, the method further includes: Acquire the body-sensing observation data and upper-level velocity commands corresponding to the current control cycle; Based on the estimated base linear velocity, the body perception observation data corresponding to the current control cycle, and the upper-level velocity command, an extended observation vector is generated. The extended observation vector is input into the whole-body control strategy network to obtain the whole-body control command; The combined system of the two-wheeled legged robotic arm is controlled based on the whole-body control commands.
[0013] Another aspect of this application provides a training method for a base velocity estimation network, applied to a training device, the method comprising: Acquire training samples, which include the proprioception history sequence of the dual-wheeled legged robotic arm joint system and the actual base linear velocity corresponding to the proprioception history sequence. The ontology perception history sequence is input into the base velocity estimation network to obtain the base linear velocity estimate; Gradient truncation is performed on the base linear velocity estimate used as the input policy network, and the input data of the policy network is generated based on the gradient-trunculated base linear velocity estimate. The base velocity estimation network is updated based on the estimation loss between the base linear velocity estimate without gradient truncation and the actual base linear velocity. Update the policy network based on the policy loss corresponding to the policy network; The gradient of the policy loss is not backpropagated to the base velocity estimation network.
[0014] Optionally, before obtaining training samples, the following steps are also included: A simulation physical model of the dual-wheeled legged robotic arm combined system is constructed, the simulation physical model including a dual-wheeled legged chassis model and a robotic arm model; Run the simulation environment that includes the simulated physical model; The training samples are obtained from the simulation environment.
[0015] Optionally, obtaining training samples includes: Collect body-sensing observation data corresponding to multiple control cycles from the simulation environment; Organize the ontology perception observation data within a preset historical time window in chronological order to generate the ontology perception historical sequence. The actual base linear velocity corresponding to the ontology perception history sequence is obtained from the simulation environment.
[0016] Optionally, before obtaining training samples, the following steps are also included: Configure domain randomization parameters; The combined system of a two-wheeled robotic arm in the simulation environment is randomized based on the domain randomization parameters. The domain randomization parameters include at least one of chassis parameters, joint parameters, sensor noise parameters, motion delay parameters, and external disturbance parameters.
[0017] Optionally, gradient truncation is performed on the base linear velocity estimate used as input to the policy network, and input data for the policy network is generated based on the gradient-trunculated base linear velocity estimate, including: Tensor separation processing is performed on the estimated base linear velocity to obtain the base linear velocity estimated value after gradient truncation; Acquire the body-sensing observation data and upper-level velocity commands corresponding to the current control cycle; The base linear velocity estimate after gradient truncation, the body perception observation data corresponding to the current control cycle, and the upper-layer velocity command are concatenated to generate an extended observation vector. The extended observation vector is used as input data for the policy network.
[0018] Optionally, the base velocity estimation network is updated based on the estimation loss between the base linear velocity estimate without gradient truncation and the actual base linear velocity, including: The mean square error loss is determined based on the base linear velocity estimate without gradient truncation and the actual base linear velocity. The network parameters of the pedestal velocity estimation network are updated based on the mean square error loss.
[0019] Optionally, updating the policy network based on the policy loss corresponding to the policy network includes: Based on the interaction results between the whole-body control commands output by the policy network and the simulation environment, the policy loss is determined; Update the network parameters of the policy network based on the policy loss; The strategy loss is determined based on the PPO pruning agent objective function.
[0020] Optionally, the method further includes: The base velocity estimation network is updated using the first optimizer based on the estimated loss; The policy network is updated using a second optimizer based on the policy loss. The first optimizer and the second optimizer work independently within the same training cycle.
[0021] Optionally, the method further includes: After training is complete, output the trained pedestal velocity estimation network; The trained base velocity estimation network is deployed to the control device of the dual-wheeled robotic arm joint system; The trained pedestal velocity estimation network has an input interface for receiving ontology-aware historical sequences and an output interface for outputting pedestal linear velocity estimates.
[0022] Another aspect of this application provides a base speed estimation device applied to a control device for a dual-wheeled legged robotic arm combined system. The dual-wheeled legged robotic arm combined system includes a dual-wheeled legged chassis and a robotic arm disposed on the dual-wheeled legged chassis. The device includes: The acquisition module is used to acquire body perception observation data corresponding to multiple consecutive control cycles. The body perception observation data includes signals from the dual-wheel foot plate system, signals from the robotic arm subsystem, and historical motion signals. The generation module is used to generate a body perception history sequence based on the body perception observation data corresponding to the consecutive multiple control cycles. The estimation module is used to input the ontology perception history sequence into the base velocity estimation network to obtain the base linear velocity estimate of the dual-wheeled manipulator joint system; The output module is used to output the estimated value of the base linear velocity, which is used for the whole-body coordination control of the dual-wheeled legged robotic arm combined system.
[0023] Another aspect of this application provides a training apparatus for a pedestal velocity estimation network, applied to a training device, the apparatus comprising: The sample acquisition module is used to acquire training samples, which include the proprioception history sequence of the dual-wheeled legged robotic arm joint system and the actual base linear velocity corresponding to the proprioception history sequence. The estimation module is used to input the ontology perception history sequence into the base velocity estimation network to obtain the base linear velocity estimate. The truncation module is used to perform gradient truncation processing on the base linear velocity estimate used as input to the policy network, and to generate the input data of the policy network based on the gradient-trunculated base linear velocity estimate. The first update module is used to update the base velocity estimation network based on the estimated loss between the base linear velocity estimate without gradient truncation and the actual base linear velocity. The second update module is used to update the policy network based on the policy loss corresponding to the policy network. The gradient of the policy loss is not backpropagated to the base velocity estimation network.
[0024] Another aspect of this application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method as described above.
[0025] Another aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described above.
[0026] Another aspect of this application provides a computer program product including a computer program that, when executed by a processor, implements the method described above.
[0027] The embodiments of this application employing the above-described technical solution may have the following advantages: By acquiring proprioceptive observation data corresponding to multiple consecutive control cycles, and ensuring that this data simultaneously includes signals from the dual-wheeled foot chassis system, the robotic arm subsystem, and historical motion signals, the chassis motion state, robotic arm motion state, and preceding control actions can be incorporated into a single historical sequence representation. Furthermore, by inputting the proprioceptive historical sequence into a base velocity estimation network to obtain a base linear velocity estimate, the base linear velocity estimation process can utilize historical state change information of the combined system, rather than relying solely on signals from a single moment or a single subsystem. This improves the matching degree between the base linear velocity estimate and the current motion state of the dual-wheeled foot robotic arm combined system, providing a more suitable velocity input for whole-body coordinated control.
[0028] Furthermore, when training the pedestal velocity estimation network, gradient truncation is performed on the pedestal linear velocity estimates used as input to the policy network. This prevents the gradient of the policy loss from being backpropagated to the pedestal velocity estimation network and allows the pedestal velocity estimation network to be updated based on the estimated loss. This reduces the impact of the policy loss on the velocity estimation task and makes the output of the pedestal velocity estimation network more stably correspond to the physical quantity of pedestal linear velocity. Attached Figure Description
[0029] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0030] Figure 1 The diagram schematically illustrates the operating environment of the base velocity estimation method according to an embodiment of this application; Figure 2 A flowchart illustrating a base velocity estimation method according to Embodiment 1 of this application is shown schematically. Figure 3 Schematic illustration Figure 2 Flowchart of the first sub-step of step S200; Figure 4 Schematic illustration Figure 2 Flowchart of the second sub-step of step S200; Figure 5 Schematic illustration Figure 2 Flowchart of the third sub-step in step S200; Figure 6 Schematic illustration Figure 2 Flowchart of the sub-steps in step S202; Figure 7 Schematic illustration Figure 2 Flowchart of extended processing related to steps S204 and S206; Figure 8 The flowchart illustrating the training method of the base velocity estimation network according to Embodiment 2 of this application is shown in the schematic diagram. Figure 9 Schematic illustration Figure 8 Flowchart of training sample acquisition before and related to step S800; Figure 10 Schematic illustration Figure 8 Flowchart of the sub-steps in step S804; Figure 11 Schematic illustration Figure 8 Training update flowchart related to steps S806 and S808; Figure 12 A block diagram of a base velocity estimation device according to Embodiment 3 of this application is shown schematically; Figure 13 A block diagram schematically illustrates a training apparatus for a base velocity estimation network according to Embodiment 4 of this application; and Figure 14 A schematic diagram of the hardware architecture of a computer device according to Embodiment 5 of this application is shown. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0032] It should be noted that the descriptions involving "first," "second," etc., in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0033] In the description of this application, it should be understood that the numerical labels before the steps do not indicate the order of the steps, but are only used to facilitate the description of this application and to distinguish each step, and therefore should not be construed as a limitation of this application.
[0034] It should be noted that if this application involves the collection, storage, use, transmission, and processing of equipment data, sensor data, simulation data, training data, or control data, each stage of these data processing must strictly comply with the laws, regulations, industry standards, and regulatory requirements of the data source country, the place of use, and relevant countries and regions to ensure the legality and compliance of data activities. When collecting, storing, using, transmitting, and processing related data, authorization, anonymization, encryption, access control, or other security measures may be adopted according to the actual scenario to ensure that the relevant data processing meets the corresponding data security requirements.
[0035] In this embodiment, the dual-wheeled legged robotic arm combined system can be a mobile operating platform formed by combining a dual-wheeled legged chassis and a robotic arm. The dual-wheeled legged chassis can include leg joints and drive wheels, and the robotic arm can be mounted on the base or other load-bearing structure of the dual-wheeled legged chassis. The base can be understood as the main body structure of the dual-wheeled legged chassis used to support the robotic arm and move in coordination with the leg structure and drive wheel structure. The base linear velocity can be understood as the linear velocity component of the base in a spatial coordinate system or a coordinate system related to the machine body. The base velocity estimation network can be understood as a neural network used to output the estimated value of the base linear velocity based on the body perception history sequence. Full-body coordinated control can be understood as the coordinated control of the joints, drive wheels, or actuators of the dual-wheeled legged chassis and the robotic arm.
[0036] like Figure 1 As shown, the operating environment of this embodiment may include a training device 10, a control device 20, a bipedal robotic arm combined system 30, a sensor assembly 40, and an execution assembly 50. The training device 10 can be used to train a base velocity estimation network and a whole-body control strategy network in a simulation environment, and deploy the trained base velocity estimation network to the control device 20. The control device 20 may be mounted on the bipedal robotic arm combined system 30, or may be communicatively connected to the bipedal robotic arm combined system 30. The sensor assembly 40 may include at least one of an inertial measurement unit, a joint encoder, and a wheel speed sensor, for providing body perception observation data to the control device 20. The execution assembly 50 may include at least one of a leg joint actuator, a drive wheel actuator, and a robotic arm joint actuator, for executing corresponding actions according to the whole-body control commands output by the control device 20.
[0037] The aforementioned devices, systems, components, and their quantities are exemplary; the quantity and types of devices can be adjusted in different scenarios or according to different needs.
[0038] Example 1 Figure 2A flowchart illustrating a base speed estimation method according to Embodiment 1 of this application is shown schematically. This base speed estimation method can be applied to the control equipment of a dual-wheeled robotic arm combined system, which may include a dual-wheeled chassis and a robotic arm mounted on the dual-wheeled chassis. Figure 2 As shown, the base velocity estimation method may include steps S200 to S206, wherein: S200 acquires body perception observation data corresponding to multiple consecutive control cycles. The body perception observation data includes signals from the dual-wheel foot plate system, signals from the robotic arm subsystem, and historical motion signals.
[0039] S202, based on the ontology perception observation data corresponding to multiple consecutive control cycles, generates an ontology perception historical sequence.
[0040] S204. Input the body perception history sequence into the base velocity estimation network to obtain the base linear velocity estimate of the dual-wheeled legged robotic arm joint system.
[0041] S206 outputs the base linear velocity estimate, which is used for the whole-body coordination control of the dual-wheeled legged robotic arm combined system.
[0042] In this embodiment, by acquiring proprioceptive observation data covering multiple consecutive control cycles of the dual-wheeled foot chassis system, the robotic arm subsystem, and historical motion signals, the control device can obtain information on the state changes of the combined system over a period of time. By organizing this state change information into a proprioceptive historical sequence and inputting it into the base velocity estimation network, the base velocity estimation network can output an estimated base linear velocity value based on the joint temporal change pattern of the chassis and the robotic arm. Therefore, the base linear velocity estimation process can simultaneously utilize the chassis's own motion state, the robotic arm's configuration changes, and preceding control motion information, thereby improving the matching degree between the base linear velocity estimation and the motion state of the dual-wheeled foot robotic arm combined system.
[0043] In some embodiments, the dual-wheeled legged robotic arm combined system can maintain base balance solely through the cooperation of the two drive wheels and leg structure. Therefore, the estimation error of its base linear velocity directly affects subsequent balance control and whole-body coordination control. Base linear velocity can serve as one of the state inputs to the whole-body control strategy network or the whole-body controller, characterizing the motion state of the dual-wheeled chassis in the current control cycle. By simultaneously incorporating signals from the dual-wheeled chassis system, the robotic arm subsystem, and historical motion signals into the estimation input, the velocity estimation process can no longer rely solely on the single chassis state but can instead combine the influence of the robotic arm configuration, the robotic arm motion state, and preceding control actions on the base motion for estimation.
[0044] The following combination Figure 2The steps in steps S200 to S206, as well as other optional steps, are described in detail.
[0045] In step S200, the control device can acquire body perception observation data corresponding to multiple consecutive control cycles. The body perception observation data includes signals from the dual-wheel foot plate system, signals from the robotic arm subsystem, and historical motion signals.
[0046] Specifically, a control cycle can correspond to the cycle in which the control device performs one state acquisition, estimation calculation, or control output. For each control cycle, the control device can acquire corresponding proprioceptive observation data from sensor components, control buffers, or control output records. Proprioceptive observation data can primarily originate from the joint system's own sensor or control state records, rather than relying on external visual positioning data, external positioning system data, or force / torque sensor data as necessary input. The dual-wheeled footplate system signals in the proprioceptive observation data can be used to characterize the posture of the dual-wheeled footplate, the motion states of the leg joints and drive wheels, etc.; the robotic arm subsystem signals can be used to characterize the robotic arm joints, the robotic arm end effector, and their operational states; and the historical motion signals can be used to characterize the control output in the previous control cycle or adjacent historical control cycles. By simultaneously acquiring these three types of signals, a complete data foundation can be provided for the subsequent generation of proprioceptive historical sequences.
[0047] In some implementations, the control device can write the collected body perception observation data into a historical cache according to the control cycle. The historical cache can store data from multiple control cycles in chronological order. When the current control cycle arrives, the control device can read the body perception observation data corresponding to the current control cycle and several historical control cycles from the historical cache for subsequent step S202 to generate a body perception historical sequence.
[0048] In optional embodiments, such as Figure 3 As shown, the ontology perception observation data corresponding to multiple consecutive control cycles are obtained, including: S300 acquires the base angular velocity signal and gravity projection vector provided by the inertial measurement unit.
[0049] S302, obtain the angular position deviation of the leg joint relative to the default posture.
[0050] S304, acquires the angular velocity signals of the leg joints and drive wheels.
[0051] S306 generates a dual-wheel footplate system signal based on the base angular velocity signal, gravity projection vector, angular position deviation of the leg joint relative to the default posture, and angular velocity signals of the leg joint and drive wheel.
[0052] The inertial measurement unit (IMU) can be mounted on the base of the dual-wheeled chassis or on a structure fixedly connected to the base, and is used to collect the angular velocity signal of the base. The gravity projection vector can be used to characterize the projection of the gravity direction into the relevant coordinate system of the machine body, thus reflecting the attitude-related information of the base. The angular position deviation of the leg joints relative to the default attitude can be used to characterize the offset state of the leg structure relative to the reference attitude. The angular velocity signals of the leg joints and drive wheels can be used to characterize the leg joint movement speed and wheel movement state. For example, the angular position deviation of the leg joints relative to the default attitude may include the angular position deviation of three joints in each of the left and right legs, and the angular velocity signals of the leg joints and drive wheels may include the angular velocity signals of six leg joints and the angular velocity signals of two drive wheels. In this way, the chassis system signal not only contains base attitude-related information but also dynamic information of the leg joints and drive wheels, enabling the subsequent estimation network to perform velocity estimation based on a more complete chassis motion state.
[0053] In optional embodiments, such as Figure 4 As shown, the ontology perception observation data corresponding to multiple consecutive control cycles are obtained, including: S400, obtain the angular position deviation of the robotic arm joints relative to the default posture.
[0054] S402, acquire the angular velocity signal of the robotic arm joint.
[0055] S404, obtain the end position of the robotic arm end effector in the base alignment coordinate system.
[0056] S406, Obtain the end target position corresponding to the end effector of the robotic arm.
[0057] S408 generates end position error based on end position and end target position.
[0058] S410 generates robotic arm subsystem signals based on the angular position deviation of the robotic arm joints relative to the default posture, the angular velocity signal of the robotic arm joints, the end-effector position, the end-effector target position, and the end-effector position error.
[0059] The angular position deviation of the robotic arm joints relative to the default posture can be used to characterize the current configuration of the robotic arm, and the angular velocity signal of the robotic arm joints can be used to characterize the motion state of the robotic arm. The base alignment coordinate system can be understood as a coordinate system corresponding to the base posture or base reference direction, used to uniformly describe the spatial position of the robotic arm end effector relative to the base. The end-effector target position can be used to characterize the target position of the robotic arm end effector, and the end-effector position error can be used to characterize the difference between the end-effector position and the end-effector target position. For example, the robotic arm may include six robotic arm joints, and the robotic arm subsystem signals may include six-dimensional robotic arm joint position deviation signals, six-dimensional robotic arm joint angular velocity signals, three-dimensional end-effector position, three-dimensional end-effector target position, and three-dimensional end-effector position error. In this way, the robotic arm subsystem signals can reflect the robotic arm configuration, motion state, end-effector position, and tracking state, enabling the subsequent estimation network to use the relevant states of the robotic arm to identify or characterize the impact of the robotic arm motion on the chassis state.
[0060] In some embodiments, the end-effector position, target end-effector position, and end-effector position error of the robotic arm end-effector in the base alignment coordinate system can collectively form an end-effector information channel. Specifically, the end-effector position can be used to characterize the current spatial state of the robotic arm end-effector relative to the base, the target end-effector position can be used to characterize the target spatial state corresponding to the current operation task of the robotic arm, and the end-effector position error can be used to characterize the tracking deviation of the robotic arm end-effector. When the end-effector position error is large, the robotic arm may be undergoing a large adjustment process; when the end-effector position changes abruptly, it may correspond to the robotic arm picking up or releasing a load. By incorporating the end-effector position, target end-effector position, and end-effector position error into the robotic arm subsystem signal, the base velocity estimation network can obtain input information related to the robotic arm's operational intent, tracking state, and potential reaction force disturbances, thereby improving the matching degree between velocity estimation and the movement operation state.
[0061] In optional embodiments, such as Figure 5 As shown, the ontology perception observation data corresponding to multiple consecutive control cycles are obtained, including: S500: Obtain the full-body control motion vector corresponding to the previous control cycle.
[0062] S502 uses the whole-body control motion vector as historical motion signals. The whole-body control motion vector includes chassis control motion and robotic arm control motion.
[0063] The full-body control motion vector corresponding to the previous control cycle can originate from the control commands output by the control device to the execution component in the previous control cycle, or it can originate from the control output buffer recorded in the control device. The chassis control motion can correspond to the control motions of the leg joints and drive wheels, while the robotic arm control motion can correspond to the control motions of the robotic arm joints. For example, the full-body control motion vector can be fourteen-dimensional, with the chassis control motion being eight-dimensional and the robotic arm control motion being six-dimensional. By incorporating historical motion signals into the body perception observation data, the subsequent estimation network can utilize the relationship between the preceding control input and the subsequent system response when estimating the base linear velocity, thereby reducing the information deficiency caused by relying solely on the instantaneous state of the sensors.
[0064] In step S202, the control device can generate a body perception historical sequence based on the body perception observation data corresponding to multiple consecutive control cycles.
[0065] Specifically, the control device can read multiple proprioceptive observation data points within a preset historical time window from the historical cache and organize these data points in chronological order. These multiple proprioceptive observation data points can be concatenated into a vector form or organized into a tensor form, as long as the data structure can preserve the temporal relationship across multiple consecutive control cycles and serve as input to the base velocity estimation network. The proprioceptive historical sequence can be used to simultaneously characterize the changes in base plate system signals, robotic arm subsystem signals, and historical motion signals over multiple control cycles.
[0066] In some implementations, if the current control cycle is t, the control device can acquire proprioceptive observation data corresponding to multiple control cycles such as t, t-1, t-2, etc., and splice them in a fixed order from early to late or from late to early. The splicing order can be kept consistent during training and deployment to ensure that the input dimension and temporal meaning received by the base velocity estimation network remain stable. In this way, changes in chassis state, changes in robotic arm configuration, changes in end effector state, and changes in historical actions can be encoded into the same historical input, facilitating the extraction of joint temporal features by the base velocity estimation network.
[0067] In optional embodiments, such as Figure 6 As shown, based on the ontology-sensing observation data corresponding to multiple consecutive control cycles, an ontology-sensing historical sequence is generated, including: The S600 acquires multiple body-sensing observation data within a preset historical time window according to the time sequence of the control cycle.
[0068] S602, multiple body-sensing observation data are stitched together to generate a body-sensing history sequence. This history sequence is used to characterize the joint temporal variation pattern between the signals of the dual-wheel footplate system and the robotic arm subsystem.
[0069] The preset historical time window can be used to determine the number of historical control cycles included in the ontology perception historical sequence. The joint temporal change pattern can be understood as the temporal relationship between the signals of the dual-wheeled foot-plate system and the robotic arm subsystem. For example, changes in the robotic arm joint speed may be temporally correlated with changes in base posture or wheel speed, and changes in the robotic arm end-effector position may be temporally correlated with changes in the stress state of the chassis. By preserving these temporal relationships, an input representation reflecting the arm-chassis coupling effect can be provided to the base speed estimation network.
[0070] In some embodiments, the ontology perception history sequence can implicitly contain arm-chassis coupling dynamics information. Specifically, the current configuration of the robotic arm can be characterized by the angular position deviation of the robotic arm joints relative to the default posture, and is used to reflect the impact of changes in the robotic arm configuration on the overall centroid position of the joint system; the robotic arm's motion speed can be characterized by the angular velocity signals of the robotic arm joints, and is used to reflect the disturbance of the chassis posture by the reaction torque generated by the robotic arm's acceleration and deceleration; the end-effector load state can be implicitly characterized by the joint change pattern between the robotic arm joint state, the base posture change, and the end effector state, and is used to reflect the impact of load changes on the system inertia. In this way, the base velocity estimation network can learn the different contributions of the chassis's own drive and the robotic arm coupling effect to the sensor signals in the ontology perception history sequence, thereby reducing the impact of the aliasing of these two signals on the base velocity estimation results.
[0071] In some embodiments, the joint temporal variation pattern may include at least one of the following: characteristic changes in the reaction torque generated by the acceleration of the robotic arm joints in the history of the base attitude angle; the temporal correlation between the rate of change of the chassis pitch angle and the angular acceleration of the robotic arm joints; and the temporal correlation between the rate of change of the drive wheel speed and the change of the robotic arm configuration. The base velocity estimation network can learn the mapping relationship between the above temporal correlations and the base linear velocity through the training process. By preserving the above temporal correlations in the same ontology-aware history sequence, the base velocity estimation network can compensate for the changes in chassis state caused by the robotic arm motion without establishing an explicit arm-chassis coupled dynamics model.
[0072] In some embodiments, the center of mass and moment of inertia of the combined system may undergo a step change during the process of the robotic arm picking up or releasing a load. Changes in the arm joint state, base posture, and velocity trends in the body perception history sequence can be used to characterize this load abrupt change process. Based on the above-mentioned historical change information, the base velocity estimation network can adjust the base linear velocity estimate in the transient early stage of the load abrupt change to reduce the impact of velocity estimation lag caused by the abrupt change in model parameters on the whole-body coordinated control.
[0073] In some embodiments, the pedestal velocity estimation network can have different dependence strengths on ontology-aware observation data at different time steps in the ontology-aware historical sequence. For example, ontology-aware observation data corresponding to more recent control periods can have a greater impact on the current pedestal linear velocity estimate, while ontology-aware observation data corresponding to more distant control periods can be used to provide historical references for velocity change trends or coupled perturbations. This dependence strength can be autonomously formed by the pedestal velocity estimation network based on training samples during training, without requiring explicit configuration of fixed decay weights during the deployment phase. This approach allows for the consideration of both current state response and historical trend information.
[0074] In an optional embodiment, the preset historical time window includes ten consecutive control cycles, and the ontology perception observation data corresponding to each control cycle is 55-dimensional, and the ontology perception historical sequence is 550-dimensional.
[0075] Specifically, the proprioceptive observation data corresponding to a single control cycle can include 20-dimensional signals from the dual-wheel footplate system, 21-dimensional signals from the robotic arm subsystem, and 14-dimensional historical motion signals. The proprioceptive observation data from ten consecutive control cycles, spliced together chronologically, can form a 550-dimensional proprioceptive historical sequence. By using observation data from multiple control cycles instead of a single control cycle, the base velocity estimation network can obtain the dynamic changes of the joint system over a period of time, making it more suitable for handling the effects of time-varying centroid, time-varying inertia, and reaction torque coupling caused by the robotic arm's motion.
[0076] In some embodiments, the ontological perception observation data corresponding to a single control cycle can be 55-dimensional. Among these, the dual-wheel footplate system signal can be 20-dimensional, accounting for approximately 36% of the single-step observation data; the robotic arm subsystem signal can be 21-dimensional, accounting for approximately 38% of the single-step observation data; and the historical motion signal can be 14-dimensional, accounting for approximately 26% of the single-step observation data. By ensuring that the robotic arm subsystem signal has a relatively sufficient information proportion in the single-step observation data, the base velocity estimation network can obtain information such as the robotic arm joint state, robotic arm motion state, end effector position, and end-effector tracking error during forward inference, reducing the risk of the robotic arm state being masked by the chassis signal.
[0077] In an optional embodiment, the window length of the preset historical time window satisfies: L·Δt≥τ_rise; Where L is the number of historical time steps, Δt is the sampling interval, and τ_rise is the rise time constant of the chassis linear velocity under the acceleration command; The preset historical time window is also used to cover the duration of the reaction force caused by the robotic arm's operation.
[0078] Specifically, L can represent the number of historical time steps involved in the stitching, and Δt can represent the sampling interval between two adjacent control cycles. By ensuring that L·Δt is not less than the rise time constant of the chassis linear velocity under acceleration commands, the ontology-aware historical sequence can cover the main dynamic processes of the chassis speed response. Furthermore, by ensuring that the preset historical time window covers the duration of the reaction force caused by the robotic arm's operation, the base speed estimation network can obtain a more complete time segment at the input end of the influence of the robotic arm operation on the chassis attitude, wheel speed, or joint state. For example, L can be 10, Δt can be 0.02 seconds, and L·Δt can be 0.2 seconds. Through the above window constraints, the risk of missing coupled dynamic information due to an excessively short historical window can be reduced.
[0079] In some embodiments, the number of historical time steps L can be 10, the sampling interval Δt can be 0.02 seconds, and the historical time window length L·Δt can be 0.2 seconds. The rise time constant τ_rise of the chassis linear velocity under a typical acceleration command can be 0.15 seconds, thus the 0.2-second historical time window can cover the main dynamic process of the chassis speed response. In some embodiments, the control cycle can include four physical simulation steps, each of which can be 0.005 seconds, thereby forming a 0.02-second control cycle. Through the above timing configuration, the historical windows of both the training and deployment phases can cover the continuous process of chassis speed response and robotic arm operation reaction force.
[0080] In step S204, the control device can input the body perception history sequence into the base velocity estimation network to obtain the base linear velocity estimate of the dual-wheeled legged robotic arm joint system.
[0081] Specifically, the base velocity estimation network can be a pre-trained neural network deployed to the control device. The control device can convert the ontology-aware historical sequence generated in step S202 into the input format required by the base velocity estimation network and call the base velocity estimation network for forward inference. The base velocity estimation network can output an estimated value of the base linear velocity. This estimated value can include the estimated linear velocity components of the base in different axial directions. During the training phase, the base velocity estimation network can learn the correspondence between the ontology-aware historical sequence and the actual base linear velocity; during the deployment phase, it can output the velocity estimation result based on the currently acquired ontology-aware historical sequence.
[0082] In some implementations, the base velocity estimation network can infer patterns from the combined changes in chassis and robotic arm signals within a historical sequence of propagation perception. For example, acceleration and deceleration of robotic arm joints may cause base attitude fluctuations, changes in arm configuration may alter the overall center of mass position, and changes in the end effector state may correspond to changes in load or target tracking state during the movement operation. The base velocity estimation network can utilize this historical change information to output an estimate of the base linear velocity.
[0083] In optional embodiments, such as Figure 7 As shown, the proprioception history sequence is input into the base velocity estimation network to obtain the base linear velocity estimate of the dual-wheeled robotic arm joint system, including: S700 inputs the ontology perception history sequence into the base velocity estimation network of the multilayer perceptron structure.
[0084] S702 outputs a three-dimensional estimate of the linear velocity of the base through a base velocity estimation network.
[0085] S704, the three-dimensional base linear velocity estimate includes estimates of the base linear velocity in three axial directions.
[0086] The multilayer perceptron structure can be used for nonlinear mapping of the body's perception history sequence. The three-dimensional base linear velocity estimate can correspond to the estimated components of the base linear velocity in the x, y, and z axes. By setting the output dimension to a three-dimensional result corresponding to the spatial components of the base linear velocity, the estimation result can have a clear physical quantity meaning and can be used as input for subsequent whole-body coordinated control.
[0087] In some embodiments, the base velocity estimation network may have a 550-dimensional input interface and a 3D output interface. The 550-dimensional input interface can be used to receive the ontology-aware historical sequence corresponding to ten consecutive control cycles, and the 3D output interface can be used to output the estimated values of the base linear velocity in three axis directions. In one specific embodiment, the total number of network parameters of the base velocity estimation network can be on the order of 174,000; under the corresponding computing environment, the inference time per frame can be on the order of 0.3 milliseconds. This inference time of less than 20 milliseconds in a control cycle is beneficial for the base velocity estimation network to complete the velocity estimation calculation within the control cycle. The above parameters are only the implementation method in a specific embodiment and are not intended to limit the structure and scale of the base velocity estimation network.
[0088] In an optional embodiment, the pedestal velocity estimation network includes two hidden layers with 256 and 128 neurons respectively, and the activation function of the pedestal velocity estimation network is ELU.
[0089] Specifically, the input dimension of the base velocity estimation network can correspond to the dimension of the ontology-aware historical sequence, for example, 550 dimensions; the output dimension can correspond to the dimension of the base linear velocity estimate, for example, three dimensions. Two hidden layers can perform feature transformation on the input historical observation data, and the ELU activation function can provide nonlinear expressive power. Through this network structure, the mapping from the ontology-aware historical sequence to the base linear velocity estimate can be achieved in control devices with a relatively simple network computation structure, facilitating deployment in computing devices requiring real-time control.
[0090] In step S206, the control device can output an estimated value of the base linear velocity, which is used for the whole-body coordination control of the dual-wheeled legged robotic arm combined system.
[0091] Specifically, the output base linear velocity estimate may include transmitting the base linear velocity estimate to the whole-body control strategy network, the whole-body controller, the state estimation module, or the control decision module. The base linear velocity estimate can be used as part of the current motion state of the combined system for subsequent generation of control commands to the chassis and the robotic arm. Since this base linear velocity estimate is generated based on chassis system signals, robotic arm subsystem signals, and historical motion signals, it can better reflect the motion state of the dual-wheeled legged robotic arm combined system during robotic arm operation.
[0092] In optional embodiments, such as Figure 7 As shown, after outputting the estimated base linear velocity, the method may further include: S706, acquires the body perception observation data and upper-level velocity command corresponding to the current control cycle.
[0093] S708 generates an extended observation vector based on the base linear velocity estimate, the body sensing observation data corresponding to the current control cycle, and the upper-level velocity command.
[0094] S710 inputs the extended observation vector into the whole-body control strategy network to obtain whole-body control commands.
[0095] S712 is a combined system of a dual-wheeled legged robotic arm controlled by full-body control commands.
[0096] The upper-level velocity command can be a velocity target provided by the upper-level task planner, teleoperated device, or motion control module. The extended observation vector can unify the current state, estimated velocity, and control target as input to the whole-body control strategy network. Based on the extended observation vector, the whole-body control strategy network can output control commands for the leg joints, drive wheels, and robotic arm joints. In this way, the base linear velocity estimate not only serves as the output result but can also further participate in the whole-body coordinated control, enabling the control command generation process to obtain velocity information that better matches the current state of the joint system.
[0097] In an optional embodiment, the whole-body control command is a fourteen-dimensional control command, which includes an eight-dimensional control command corresponding to the chassis system and a six-dimensional control command corresponding to the robotic arm subsystem.
[0098] Specifically, the eight-dimensional control commands corresponding to the chassis system can be used to control the leg joints and drive wheels, while the six-dimensional control commands corresponding to the robotic arm subsystem can be used to control the robotic arm joints. By dividing the whole-body control commands into parts corresponding to the chassis system and parts corresponding to the robotic arm subsystem, the same whole-body control strategy network output can be applied simultaneously to the dual-wheeled chassis and the robotic arm, thereby achieving coordinated control of the combined system.
[0099] In an optional embodiment, the body perception observation data is acquired by at least one of an inertial measurement unit, a joint encoder, and a wheel speed sensor; the generation of the base linear velocity estimate does not require data acquired by a vision sensor, an external positioning system, or a force / torque sensor as necessary input.
[0100] Specifically, the inertial measurement unit (IMU) can provide base angular velocity signals or attitude-related signals, the joint encoder can provide position or velocity information of the leg joints and robotic arm joints, and the wheel speed sensor can provide drive wheel speed information. By obtaining body perception observation data based on the aforementioned body perception sensors, the dependence of velocity estimation on external positioning environment, visual lighting conditions, or additional force sensors can be reduced. In some embodiments, even without inputting external visual positioning data to the base velocity estimation network, the control device can still generate a base linear velocity estimate based on the body perception observation data.
[0101] By using the above methods, the base speed estimation process can rely more on the state variables and historical control actions that the joint system itself can collect, thereby reducing the impact of changes in external environmental conditions on the speed estimation input link.
[0102] The first embodiment will be illustrated below with an application example.
[0103] S1, during the operation of the dual-wheeled robotic arm combined system, the control equipment operates the control cycle at a control frequency of 50 Hz.
[0104] S2, in each control cycle, the control device obtains the base angular velocity signal and gravity projection vector from the inertial measurement unit, obtains the angular position deviation and angular velocity signal of the leg joint and the robotic arm joint from the joint encoder, obtains the drive wheel angular velocity signal from the wheel speed sensor, and reads the whole-body control motion vector of the previous control cycle from the control buffer.
[0105] S3, the control device combines the chassis system signal, robotic arm subsystem signal and historical motion signal corresponding to the current control cycle into body perception observation data and writes it into the historical cache.
[0106] S4, the control device reads the body perception observation data of ten consecutive control cycles from the historical cache and splices them into a 550-dimensional body perception historical sequence in chronological order.
[0107] S5, the control device inputs the body perception history sequence into the base velocity estimation network, and outputs the three-dimensional base linear velocity estimate through the base velocity estimation network.
[0108] S6, the control device concatenates the base linear velocity estimate, the body perception observation data corresponding to the current control cycle, and the upper-level velocity command into an extended observation vector, and inputs the extended observation vector into the whole-body control strategy network.
[0109] S7, the whole-body control strategy network outputs fourteen-dimensional whole-body control commands, and the control device controls the leg joints, drive wheels and robotic arm joints to perform corresponding actions based on these whole-body control commands.
[0110] Through the above application examples, the control device can estimate the base linear velocity based on the body perception history sequence of the joint system when the robotic arm moves, the end effector position changes, or the chassis state changes, and use the estimation result for whole-body coordinated control, so that the base velocity estimation result can be better adapted to the dynamic operation process of the dual-wheeled legged robotic arm joint system.
[0111] Example 2 Figure 8 A flowchart illustrating a training method for a base velocity estimation network according to Embodiment 2 of this application is shown schematically. This training method can be applied to training devices. Figure 8 As shown, the training method may include steps S800 to S808, wherein: S800, acquire training samples, which include the proprioception history sequence of the dual-wheeled legged robotic arm joint system and the actual base linear velocity corresponding to the proprioception history sequence.
[0112] S802, input the ontology perception history sequence into the base velocity estimation network to obtain the base linear velocity estimate.
[0113] S804 performs gradient truncation on the base linear velocity estimate used as input to the policy network, and generates input data for the policy network based on the gradient-trunculated base linear velocity estimate.
[0114] S806 updates the base velocity estimation network based on the estimation loss between the base linear velocity estimate without gradient truncation and the true base linear velocity.
[0115] S808 updates the policy network based on the policy loss corresponding to the policy network. The gradient of the policy loss is not backpropagated to the base velocity estimation network.
[0116] In this embodiment, the training device trains the pedestal velocity estimation network and the policy network in the same training process. However, gradient truncation cuts off the path of policy loss back to the pedestal velocity estimation network, allowing the pedestal velocity estimation network to primarily update its parameters based on the velocity estimation error. This ensures that the output of the pedestal velocity estimation network retains its meaning corresponding to the physical quantity of pedestal linear velocity, while the policy network can learn control policies based on the gradient-trunked velocity estimates. Therefore, the impact of policy task loss on the pedestal velocity estimation network can be reduced, allowing the estimation network and policy network to maintain functional division within the same training cycle.
[0117] In step S800, the training device can acquire training samples, which include the proprioception history sequence of the dual-wheeled robotic arm joint system and the actual base linear velocity corresponding to the proprioception history sequence.
[0118] Specifically, training samples can come from a simulation environment. The simulation environment can provide the actual base linear velocity available during the training phase, which can serve as a supervisory signal for the base velocity estimation network. The training device can collect proprioceptive perception observation data corresponding to multiple control cycles and generate proprioceptive perception history sequences in a manner consistent with the deployment phase. By including both the proprioceptive perception history sequences and the actual base linear velocity in the training samples, the base velocity estimation network can learn the mapping relationship from the proprioceptive perception history sequences to the base linear velocity.
[0119] In optional embodiments, such as Figure 9 As shown, before obtaining training samples, the method may further include: S900, construct the simulation physical model of the dual-wheeled legged robotic arm combined system. The simulation physical model includes the dual-wheeled legged chassis model and the robotic arm model.
[0120] S902 runs a simulation environment that includes a simulated physical model.
[0121] S904, obtains training samples from the simulation environment.
[0122] Specifically, the training device can construct a joint physical model of a two-wheeled chassis and a robotic arm within a simulation environment. The two-wheeled chassis model can be used to simulate chassis posture, leg joint movement, and drive wheel motion, while the robotic arm model can be used to simulate the motion of robotic arm joints and end effectors. During simulation operation, data such as sensor observable states, control actions, and actual base linear velocities can be generated. By acquiring training samples from the simulation environment, supervised training data can be provided for the base velocity estimation network without relying on direct measurement of base linear velocities from a real robot.
[0123] In some embodiments, the simulation environment can be built based on a simulation physics engine and can run at a physics engine time step of 5 milliseconds. The training device can execute control cycles at a control frequency of 50 Hz and can run multiple simulation environments in parallel. In one specific embodiment, the number of parallel environments can be 4096, and the trajectory duration of each training trajectory can be 20 seconds. By running multiple simulation environments in parallel, training samples can be collected under different robotic arm configurations, chassis states, and disturbance conditions, thereby improving the coverage of the training samples on the joint system's operating state.
[0124] In some embodiments, the simulation physical model may include a two-wheeled chassis model and a six-degree-of-freedom robotic arm model. The two-wheeled chassis model may include a base, left and right leg links, leg joints, and drive wheels. In one specific embodiment, the base mass may be 10.0 kg, the total mass of a single leg link may be 5.65 kg, the total chassis mass may be 21.3 kg, and the wheel diameter may be 0.25 m. The six-degree-of-freedom robotic arm may be mounted on top of the chassis base, and its mounting offset relative to the base may include -0.016 m in the x-direction and 0.120 m in the z-direction. The total mass of the robotic arm may be 3.53 kg, and the total mass of the combined two-wheeled robotic arm system may be 24.83 kg. By configuring the above physical parameters in the simulation physical model, the training samples can cover the mass distribution, mounting position, and inertia characteristics corresponding to the actual combined system.
[0125] In an optional embodiment, the dual-wheeled foot chassis model includes six leg joints and two drive wheels, the robotic arm model includes six robotic arm joints, and the simulation physical model of the dual-wheeled foot robotic arm combined system includes fourteen degrees of freedom.
[0126] Specifically, the six leg joints in the dual-wheeled chassis model correspond to multiple joints in the left and right legs, and the two drive wheels can be used to simulate wheeled movement. The six robotic arm joints in the robotic arm model can be used to simulate the spatial motion of the robotic arm. By constructing a simulation physical model with fourteen degrees of freedom, the training device can simultaneously simulate the motion coupling relationship between the chassis and the robotic arm during the simulation phase, enabling the training samples to cover the typical operating states of the combined system.
[0127] In optional embodiments, such as Figure 9 As shown, obtaining training samples can include: S906 collects body perception observation data corresponding to multiple control cycles from the simulation environment.
[0128] S908 organizes the ontology perception observation data within a preset historical time window in chronological order to generate an ontology perception historical sequence.
[0129] S910 obtains the actual base linear velocity corresponding to the ontological perception historical sequence from the simulation environment.
[0130] Specifically, the training device can collect ontology-aware observation data in the simulation environment using the same observation construction method as the deployment side, and organize multiple observation data within a preset historical time window in chronological order to generate an ontology-aware historical sequence. The actual base linear velocity can be provided by the simulation environment to supervise the base velocity estimation network. By making the input construction method consistent with that of the deployment phase, the difference in input distribution between training and deployment can be reduced.
[0131] In some embodiments, the true pedestal linear velocity can be determined from privileged observation information provided by the simulation environment. For example, the true pedestal linear velocity may correspond to the first three-dimensional velocity components in the state observation vector of the simulation environment. This true pedestal linear velocity is not used as input during the deployment phase, but only during the training phase to compute the estimated loss of the pedestal velocity estimation network. By limiting the use of the true pedestal linear velocity to the training phase, reliance on real velocity quantities that cannot be directly measured can be avoided during the deployment phase.
[0132] In an optional embodiment, before obtaining training samples, the method may further include: S912, configure domain randomization parameters.
[0133] S914 performs randomization processing on a dual-wheeled robotic arm combined system in a simulation environment based on domain randomization parameters.
[0134] The domain randomization parameters include at least one of the following: chassis parameters, joint parameters, sensor noise parameters, motion delay parameters, and external disturbance parameters.
[0135] Specifically, chassis parameters may include at least one of the following: wheel-to-ground friction coefficient, coefficient of restitution, base added mass, base center of mass offset, and inertia scaling factor. Joint parameters may include at least one of the following: Kp scaling factor, Kd scaling factor, torque limiting scaling factor, and joint default position offset. Sensor noise parameters may include at least one of the following: inertial measurement unit bias, angular velocity noise, gravity projection noise, joint position noise, and joint velocity noise. Motion delay parameters can be used to simulate the delay between the control action output and its application to the simulated physical model. External disturbance parameters can be used to simulate external thrust or operational disturbances. Through domain randomization, training samples can cover state changes under different physical parameters, sensor errors, and disturbance conditions, giving the trained base velocity estimation network better environmental adaptability.
[0136] In some embodiments, the domain randomization parameters may include chassis parameters, joint parameters, sensor noise parameters, motion delay parameters, and external disturbance parameters. Chassis parameters may include wheel-to-ground friction coefficient, coefficient of restitution, base-added mass, base center of gravity offset, and inertia scaling factor. For example, the wheel-to-ground friction coefficient may be randomized within the range of [0.2, 1.6], the coefficient of restitution may be randomized within the range of [0.0, 1.0], the base-added mass may be randomized within the range of [-1.0, +3.0] kg, the base center of gravity offset may be randomized within the range of ±0.04 m in the X direction, ±0.03 m in the Y direction, and ±0.04 m in the Z direction, and the inertia scaling factor may be randomized within the range of [0.8, 1.2].
[0137] Furthermore, joint parameters may include a Kp scaling factor, a Kd scaling factor, a torque limit scaling factor, and a default joint position offset. For example, the Kp and Kd scaling factors can be randomized within the range of [0.8, 1.4], the torque limit scaling factor can be randomized within the range of [0.8, 1.2], and the default joint position offset can be randomized within the range of [-0.05, 0.05] radians. Sensor noise parameters may include angular velocity noise, gravity projection noise, joint position noise, and joint velocity noise. For example, the angular velocity noise can be 0.03 rad / s, the gravity projection noise can be 0.04, the joint position noise can be 0.008 rad, and the joint velocity noise can be 0.05 rad / s. Motion delay parameters can be randomized within the range of [0, 20] milliseconds. External disturbance parameters may include random thrust intervals and maximum thrust velocity variations; for example, the random thrust interval can be 7 seconds, and the maximum thrust velocity variation can be 1.0 m / s. Through the above-mentioned domain randomization process, the trained base velocity estimation network can maintain good estimation adaptability under different conditions of friction, mass distribution, joint response, sensor error and external disturbance.
[0138] In step S802, the training device can input the ontology perception history sequence into the base velocity estimation network to obtain the base linear velocity estimate.
[0139] Specifically, the training device can use the same or corresponding pedestal velocity estimation network structure as the deployment phase to process the ontology-aware historical sequence and output a pedestal linear velocity estimate. This pedestal linear velocity estimate can be used to calculate the estimated loss with the real pedestal linear velocity, and can also be used as part of the input to the policy network after gradient truncation.
[0140] In an optional embodiment, the pedestal velocity estimation network is a multilayer perceptron structure, which includes two hidden layers with 256 and 128 neurons respectively. The policy network is a multilayer perceptron structure, consisting of three hidden layers with 512, 256, and 128 neurons respectively.
[0141] Specifically, the base velocity estimation network can be used to perform a regression mapping from the ontology-perceived historical sequence to the base linear velocity estimate, while the policy network can be used to output whole-body control commands based on the extended observation vector. By setting up the base velocity estimation network and the policy network separately, the velocity estimation task and the control policy generation task can be distinguished from each other in terms of network structure, which facilitates independent updates in subsequent execution.
[0142] In step S804, the training device can perform gradient truncation on the base linear velocity estimate used as input to the policy network, and generate input data for the policy network based on the gradient-trunculated base linear velocity estimate.
[0143] Specifically, gradient truncation can be used to prevent policy loss from backpropagating to the pedestal velocity estimation network via the policy network input data. The training device can retain the pedestal linear velocity estimates without gradient truncation for loss calculation, while using the gradient-truncated pedestal linear velocity estimates as input data to construct the policy network. In this way, the training objective of the pedestal velocity estimation network can remain the velocity estimation error, rather than being directly driven by policy task rewards or policy loss.
[0144] In optional embodiments, such as Figure 10 As shown, gradient truncation is performed on the base linear velocity estimate used as input to the policy network, and the input data for the policy network is generated based on the gradient-trunculated base linear velocity estimate, including: S1000 performs tensor separation processing on the base linear velocity estimate to obtain the base linear velocity estimate after gradient truncation.
[0145] S1002, acquire the body perception observation data and upper-level speed command corresponding to the current control cycle.
[0146] S1004 concatenates the base linear velocity estimate after gradient truncation, the body sensing observation data corresponding to the current control cycle, and the upper-level velocity command to generate an extended observation vector.
[0147] S1006 uses the extended observation vector as input data for the policy network.
[0148] Tensor separation can be understood as separating the base linear velocity estimate from the backpropagation path of the base velocity estimation network in the computational graph. The extended observation vector can simultaneously include the current observation state, the estimated velocity, and the upper-layer velocity command. Through tensor separation, the policy network can generate control actions using the base linear velocity estimate, but the gradient corresponding to the policy loss will not be backpropagated to the base velocity estimation network, thus achieving decoupling of the training of the estimation network and the policy network.
[0149] In an optional embodiment, performing tensor separation processing on the pedestal linear velocity estimate may include calling the detach operation in a deep learning framework to obtain a gradient-trunculated pedestal linear velocity estimate. Specifically, the pedestal linear velocity estimate without gradient truncation retains the computational graph connection with the pedestal velocity estimation network and is used to calculate the estimated loss; the gradient-trunculated pedestal linear velocity estimate no longer feeds back the policy loss gradient to the pedestal velocity estimation network and is used to generate the input data for the policy network. In this way, the pedestal linear velocity estimate can be used simultaneously for supervised training and policy training in the same training loop, avoiding interference between the backpropagation paths of the two types of losses.
[0150] In step S806, the training device can update the base velocity estimation network based on the estimated loss between the base linear velocity estimate without gradient truncation and the true base linear velocity.
[0151] Specifically, the base linear velocity estimate without gradient truncation retains the computational graph connectivity of the base velocity estimation network, and therefore can be used to calculate the estimation loss and back-update the network parameters. The true base linear velocity can be derived from the simulation environment or privileged observations provided by the simulator. By updating the base velocity estimation network using the estimation loss, the network can be optimized for the base linear velocity estimation task.
[0152] In optional embodiments, such as Figure 11 As shown, the base velocity estimation network is updated based on the estimation loss between the base linear velocity estimate without gradient truncation and the true base linear velocity, including: S1100 determines the mean square error loss based on the base linear velocity estimate without gradient truncation and the actual base linear velocity.
[0153] S1102, update the network parameters of the base velocity estimation network based on mean square error loss.
[0154] Specifically, the mean squared error loss can be used to measure the difference between the estimated base linear velocity and the true base linear velocity. The smaller the mean squared error loss, the closer the velocity estimate output by the base velocity estimation network is to the true base linear velocity in the training samples. By updating the base velocity estimation network based on the mean squared error loss, the network can learn a physically meaningful velocity estimation mapping.
[0155] In an optional embodiment, the mean squared error loss can be expressed as: L_est = (1 / N) Σ ‖ _i - v_i^true‖²2. Where N is the number of training samples, v_est,i is the base linear velocity estimate without gradient truncation, and v_true,i is the actual base linear velocity provided by the simulation environment. By updating the base velocity estimation network based on the above mean square error loss, the training objective of the base velocity estimation network can be directly correlated with the base linear velocity estimation error.
[0156] In step S808, the training device can update the policy network based on the policy loss corresponding to the policy network.
[0157] Specifically, the training device can input extended observation vectors into a policy network, which then outputs whole-body control commands and applies these commands to a dual-wheeled robotic arm system within a simulated environment. The training device can determine the policy loss based on the state, reward, or interaction results fed back from the simulated environment and update the policy network accordingly. Since the base linear velocity estimate used as input to the policy network has undergone gradient truncation, the policy loss is not propagated back to the base velocity estimation network.
[0158] In optional embodiments, such as Figure 11 As shown, updating the policy network based on the policy loss corresponding to the policy network includes: S1104: Determine the policy loss based on the interaction results between the whole-body control commands output by the policy network and the simulation environment.
[0159] S1106, Update the network parameters of the policy network based on the policy loss. The policy loss is determined based on the PPO-pruned agent objective function.
[0160] Specifically, the PPO pruning agent objective function can be used to limit the policy update magnitude, enabling the policy network to perform relatively stable parameter updates during training. The full-body control commands output by the policy network can be applied to the dual-wheeled chassis model and the robotic arm model in the simulation environment. After executing the full-body control commands, the simulation environment can return interactive results for policy updates. By updating the policy network based on policy loss and blocking the influence of policy loss on the base velocity estimation network through gradient truncation, the estimation and control tasks can be optimized independently within the same training cycle.
[0161] In some embodiments, the policy loss can be determined based on the PPO pruning agent objective function. Training hyperparameters associated with the PPO pruning agent objective function may include pruning coefficients, discount factors, generalized advantage estimation parameters, entropy coefficients, target KL divergence, and maximum gradient norm. For example, the pruning coefficient can be 0.2, the discount factor can be 0.99, the generalized advantage estimation parameter can be 0.95, the entropy coefficient can be 0.01, the target KL divergence can be 0.01, and the maximum gradient norm can be 1.0. By setting these training hyperparameters, the update magnitude of the policy network can be constrained, and the policy network can maintain a relatively stable training process during simulation interactions.
[0162] In an optional embodiment, the method may further include: S1108 uses the first optimizer to update the base velocity estimation network based on the estimated loss.
[0163] S1110 uses the second optimizer to update the policy network based on policy loss.
[0164] The first optimizer and the second optimizer work independently within the same training cycle.
[0165] Specifically, a first optimizer can be used to update the network parameters of the pedestal velocity estimation network, and a second optimizer can be used to update the network parameters of the policy network. Both the first and second optimizers can be Adam optimizers. In some implementations, the first and second optimizers can perform several updates separately within the same training loop. By using optimizers to update the two networks separately, the pedestal velocity estimation network and the policy network can maintain the independence of their parameter update paths during training.
[0166] In an optional embodiment, the first optimizer can be an Adam optimizer with a learning rate of 1.0 × 10⁻³, β1 of 0.9, and β2 of 0.999; the second optimizer can also be an Adam optimizer with a learning rate of 1.0 × 10⁻³, and can be adaptively adjusted according to the policy training process. In each training iteration, the first optimizer can update the pedestal velocity estimation network 5 times, and the second optimizer can update the policy network 5 times. During training, four mini-batches can be used to update the parameters of the collected training data. By allowing the first and second optimizers to work independently within the same training cycle, the pedestal velocity estimation network and the policy network can be updated with parameters around the estimated loss and policy loss, respectively.
[0167] In an optional embodiment, the method may further include: S1112, after training is complete, outputs the trained pedestal velocity estimation network.
[0168] S1114, deploys the trained base velocity estimation network to the control device of the dual-wheeled legged robotic arm joint system.
[0169] The trained pedestal velocity estimation network has an input interface for receiving ontology-aware historical sequences and an output interface for outputting pedestal linear velocity estimates.
[0170] Specifically, the training device can output the trained base velocity estimation network parameters after completing a preset number of training rounds or meeting a preset training stop condition. The control device can load these network parameters and, during operation, generate an ontology-aware historical sequence as described in Embodiment 1 and call the base velocity estimation network to output the base linear velocity estimate. Through the corresponding settings of the input and output interfaces, the trained base velocity estimation network can be used as a velocity estimation module in the control device during the deployment phase.
[0171] In some embodiments, the trained base velocity estimation network can serve as a replaceable sensing component in a dual-wheeled robotic arm joint system. Specifically, the control device can provide the sensing component with a body perception history sequence via a 550-dimensional input interface and receive the base linear velocity estimate via a 3D output interface. The base linear velocity estimate can serve as one of the inputs to the whole-body control strategy network.
[0172] In other embodiments, while maintaining consistency with the output interface of the downstream whole-body control strategy network, other state estimation modules can be used to replace the base velocity estimation network. For example, these other state estimation modules may include a filter-based state estimation module or a temporal neural network-based state estimation module; in other configurations that allow the use of external sensing inputs, these other state estimation modules may also include a visual odometry-based state estimation module. This interface-based design reduces the impact on the downstream whole-body control strategy network when the velocity estimation module is replaced or upgraded.
[0173] Example 2 will be described below with reference to a training application example.
[0174] S1, The training equipment constructs a physical model of the dual-wheeled legged robotic arm combined system in a simulation environment. The physical model includes a dual-wheeled legged chassis model and a robotic arm model.
[0175] S2, the training device configuration domain randomization parameters, so that chassis parameters, joint parameters, sensor noise parameters, motion delay parameters or external disturbance parameters in the simulation environment change during training.
[0176] S3, train the device to initialize the base velocity estimation network and policy network.
[0177] S4. During the training iteration process, the training device collects ontology perception observation data from the simulation environment and organizes the ontology perception observation data within the preset historical time window into an ontology perception historical sequence.
[0178] S5, the training device inputs the ontology perception history sequence into the base velocity estimation network to obtain the base linear velocity estimate, and obtains the corresponding real base linear velocity from the simulation environment.
[0179] S6, the training device uses the base linear velocity estimate without gradient truncation and the true base linear velocity to calculate the estimation loss, and updates the base velocity estimation network based on the estimation loss.
[0180] S7, the training device performs tensor separation processing on the base linear velocity estimate used for the input policy network, and concatenates the gradient-trunculated base linear velocity estimate with the ontology perception observation data and upper-layer velocity command corresponding to the current control cycle into an extended observation vector.
[0181] S8, the training device will expand the observation vector input strategy network to obtain the whole-body control command, and apply the whole-body control command to the dual-wheeled legged robotic arm joint system in the simulation environment.
[0182] S9, the training device determines the policy loss based on the interaction results between the policy network and the simulation environment, and updates the policy network based on the policy loss.
[0183] S10, after completing the training, the training device outputs the trained base velocity estimation network and deploys the base velocity estimation network to the control device.
[0184] In some embodiments, the training process may include 15,000 training iterations. In each training iteration, data corresponding to 24 control cycles can be collected from each simulation environment. Taking 4096 parallel simulation environments as an example, the total number of training steps can be on the order of 1.47 × 10^9. By performing training on a large-scale simulation dataset, the base velocity estimation network can be exposed to various chassis states, robotic arm configurations, end-effector states, and disturbance conditions, thereby improving its adaptability to different operating states of the joint system.
[0185] Through the above application examples, the training device can train the pedestal velocity estimation network and the policy network in the same training cycle, and through gradient truncation, ensure that the policy loss does not affect the parameter update of the pedestal velocity estimation network, thus making the trained pedestal velocity estimation network more suitable for use as a pedestal linear velocity estimation module in the deployment phase.
[0186] Example 3 Figure 12 A block diagram of a base speed estimation device according to Embodiment 3 of this application is schematically shown. This base speed estimation device can be applied to the control equipment of a dual-wheeled robotic arm combined system, which may include a dual-wheeled chassis and a robotic arm mounted on the dual-wheeled chassis. The base speed estimation device may include an acquisition module 1202, a generation module 1204, an estimation module 1206, and an output module 1208.
[0187] The acquisition module 1202 can be used to acquire body perception observation data corresponding to multiple consecutive control cycles. The body perception observation data includes signals from the dual-wheel foot plate system, signals from the robotic arm subsystem, and historical motion signals.
[0188] The generation module 1204 can be used to generate a body perception historical sequence based on the body perception observation data corresponding to multiple consecutive control cycles.
[0189] The estimation module 1206 can be used to input the body perception history sequence into the base velocity estimation network to obtain the base linear velocity estimate of the dual-wheeled legged robotic arm joint system.
[0190] The output module 1208 can be used to output the base linear velocity estimate, which is used for the whole-body coordination control of the dual-wheeled legged robotic arm combined system.
[0191] In some embodiments, the acquisition module 1202 can also be used to acquire the base angular velocity signal and gravity projection vector provided by the inertial measurement unit, acquire the angular position deviation of the leg joints relative to the default posture, acquire the angular velocity signals of the leg joints and drive wheels, and generate a dual-wheel footplate system signal based on the above signals. The acquisition module 1202 can also be used to acquire the angular position deviation of the robotic arm joints relative to the default posture, the angular velocity signal of the robotic arm joints, the end effector position of the robotic arm end effector in the base alignment coordinate system, the end effector target position, and the end effector position error, and generate a robotic arm subsystem signal based on the above signals. The acquisition module 1202 can also be used to acquire the whole-body control motion vector corresponding to the previous control cycle and use the whole-body control motion vector as a historical motion signal.
[0192] In some implementations, the generation module 1204 can be used to acquire multiple body-sensing observation data within a preset historical time window according to the time sequence of the control cycle, and to stitch the multiple body-sensing observation data together to generate a body-sensing historical sequence. The estimation module 1206 can be used to input the body-sensing historical sequence into the base velocity estimation network of the multilayer perceptron structure, and output a three-dimensional base linear velocity estimate through the base velocity estimation network. The output module 1208 can also be used to provide the base linear velocity estimate to the whole-body control strategy network or the whole-body controller for generating whole-body control commands.
[0193] Example 4 Figure 13 A block diagram of a training apparatus for a base velocity estimation network according to Embodiment 4 of this application is schematically shown. This training apparatus can be applied to a training device. The training apparatus may include a sample acquisition module 1302, an estimation module 1304, a truncation module 1306, a first update module 1308, and a second update module 1310.
[0194] The sample acquisition module 1302 can be used to acquire training samples, which include the proprioception history sequence of the dual-wheeled legged robotic arm joint system and the actual base linear velocity corresponding to the proprioception history sequence.
[0195] The estimation module 1304 can be used to input the ontology perception history sequence into the base velocity estimation network to obtain the base linear velocity estimate.
[0196] The truncation module 1306 can be used to perform gradient truncation processing on the base linear velocity estimate used as input to the policy network, and generate input data for the policy network based on the gradient-trunculated base linear velocity estimate.
[0197] The first update module 1308 can be used to update the base velocity estimation network based on the estimated loss between the base linear velocity estimate without gradient truncation and the true base linear velocity.
[0198] The second update module 1310 can be used to update the policy network based on the policy loss corresponding to the policy network, wherein the gradient of the policy loss is not backpropagated to the base velocity estimation network.
[0199] In some implementations, the sample acquisition module 1302 can also be used to construct a simulation physical model of the dual-wheeled robotic arm joint system, run a simulation environment including the simulation physical model, and acquire training samples from the simulation environment. The truncation module 1306 can also be used to perform tensor separation processing on the base linear velocity estimate, acquire the body perception observation data and upper-level velocity command corresponding to the current control cycle, and concatenate the gradient-trunculated base linear velocity estimate, the body perception observation data and upper-level velocity command corresponding to the current control cycle to generate an extended observation vector.
[0200] Example 5 This application also provides a computer device. Figure 14 A schematic diagram of the hardware architecture of a computer device according to an embodiment of this application is shown. This computer device can be a training device or a control device. The computer device 1400 may include a processor 1402, a memory 1404, a communication interface 1406, and a bus 1408. The processor 1402, memory 1404, and communication interface 1406 can be connected via the bus 1408. The memory 1404 can store computer programs, and the processor 1402 can execute the computer programs stored in the memory 1404 to implement the steps in the method embodiments described above.
[0201] In some implementations, processor 1402 may include a central processing unit, graphics processing unit, neural network processor, digital signal processor, microcontroller, field-programmable gate array, or other processing unit capable of executing computer programs. Memory 1404 may include volatile memory and / or non-volatile memory. Communication interface 1406 may be used for communicative connections with sensor components, execution components, training devices, control devices, or other external devices.
[0202] When the computer device 1400 is used as a control device, the processor 1402 can execute a computer program to acquire proprioceptive observation data, generate proprioceptive historical sequences, call the pedestal velocity estimation network to obtain pedestal linear velocity estimates, and output the pedestal linear velocity estimates. When the computer device 1400 is used as a training device, the processor 1402 can execute a computer program to acquire training samples, train the pedestal velocity estimation network and the policy network, and output the trained pedestal velocity estimation network.
[0203] Example 6 This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps described in the above method embodiments.
[0204] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Of course, the computer-readable storage medium may include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is typically used to store the operating system and various application software installed on the computer device, such as the program code of the method in any of the above embodiments. In addition, the computer-readable storage medium can also be used to temporarily store various types of data that have been output or will be output.
[0205] Example 7 This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps described in the method embodiments above.
[0206] In some implementations, the computer program product can be deployed on a control device to enable the control device to execute the base velocity estimation method; it can also be deployed on a training device to enable the training device to execute the training method for the base velocity estimation network. The computer program product can be implemented in the form of a software package, firmware, executable file, model deployment file, or a combination thereof, as long as it enables the processor to perform the corresponding steps in the above method embodiments.
[0207] It should be noted that the modules, devices, media, and program products in the above embodiments correspond to the steps in the above method embodiments. The above embodiments can be referred to mutually; content not detailed in one embodiment can be found in the relevant descriptions in other embodiments. The above modules can be implemented using software, hardware, or a combination of both. The above module division is merely exemplary; in actual implementation, modules can be merged, split, or adjusted according to computing resources, control architecture, training deployment methods, or system integration requirements.
[0208] It should also be noted that the base velocity estimation network, whole-body control strategy network, simulation environment, control device, training device, sensor components, and execution components in the embodiments of this application can all be adjusted according to specific implementation conditions. As long as they can achieve the corresponding steps in the above method embodiments, they should be understood as falling within the implementable scope of the embodiments of this application.
Claims
1. A method for estimating base velocity, characterized in that, A control device for a dual-wheeled legged robotic arm combined system, the dual-wheeled legged robotic arm combined system including a dual-wheeled legged chassis and a robotic arm disposed on the dual-wheeled legged chassis, the method comprising: Acquire body perception observation data corresponding to multiple consecutive control cycles. The body perception observation data includes signals from the dual-wheel foot plate system, signals from the robotic arm subsystem, and historical motion signals. Based on the ontology perception observation data corresponding to the multiple consecutive control cycles, an ontology perception history sequence is generated. The body perception history sequence is input into the base velocity estimation network to obtain the base linear velocity estimate of the dual-wheeled manipulator joint system; The estimated base linear velocity is output, and the estimated base linear velocity is used for the whole-body coordinated control of the dual-wheeled legged robotic arm combined system.
2. The method according to claim 1, characterized in that, Acquire the ontology-sensing observation data corresponding to multiple consecutive control cycles, including: Acquire the base angular velocity signal and gravity projection vector provided by the inertial measurement unit; Obtain the angular position deviation of the leg joints relative to the default posture; Acquire angular velocity signals from the leg joints and drive wheels; The dual-wheel footplate system signal is generated based on the base angular velocity signal, the gravity projection vector, the angular position deviation of the leg joint relative to the default posture, and the angular velocity signals of the leg joint and the drive wheel.
3. The method according to claim 1, characterized in that, Acquire the ontology-sensing observation data corresponding to multiple consecutive control cycles, including: Obtain the angular position deviation of the robotic arm joints relative to the default posture; Obtain the angular velocity signals of the robotic arm joints; Obtain the end position of the robotic arm's end effector in the base alignment coordinate system; Obtain the end target position corresponding to the end effector of the robotic arm; Based on the end position and the end target position, an end position error is generated; The robotic arm subsystem signal is generated based on the angular position deviation of the robotic arm joint relative to the default posture, the angular velocity signal of the robotic arm joint, the end position, the end target position, and the end position error.
4. The method according to claim 1, characterized in that, Acquire the ontology-sensing observation data corresponding to multiple consecutive control cycles, including: Obtain the full-body control motion vector corresponding to the previous control cycle; The whole-body control motion vector is used as the historical motion signal; The whole-body control motion vector includes chassis control motion and robotic arm control motion.
5. The method according to claim 1, characterized in that, Based on the ontology-sensing observation data corresponding to the aforementioned multiple consecutive control cycles, an ontology-sensing historical sequence is generated, including: According to the time sequence of the control cycle, acquire multiple body sensing observation data within a preset historical time window; The multiple ontology-sensing observation data are stitched together to generate the ontology-sensing historical sequence. The ontology perception history sequence is used to characterize the joint temporal change pattern between the signals of the dual-wheel footplate system and the signals of the robotic arm subsystem.
6. The method according to claim 5, characterized in that, The preset historical time window includes multiple continuous control cycles, wherein the ontology perception observation data corresponding to each control cycle is a preset number of dimensions; The window length of the preset historical time window satisfies: L·Δt≥τ_rise; Where L is the number of historical time steps, Δt is the sampling interval, and τ_rise is the rise time constant of the chassis linear velocity under the acceleration command; The preset historical time window is also used to cover the duration of the reaction force caused by the robotic arm's operation.
7. The method according to claim 1, characterized in that, The proprioceptive historical sequence is input into the base velocity estimation network to obtain the base linear velocity estimate of the dual-wheeled robotic arm joint system, including: The ontological perception history sequence is input into the base velocity estimation network of the multilayer perceptron structure; The base velocity estimation network outputs a three-dimensional base linear velocity estimate. The three-dimensional base linear velocity estimate includes estimates of the base linear velocity in three axial directions.
8. The method according to claim 1, characterized in that, After outputting the estimated base linear velocity, the process also includes: Acquire the body-sensing observation data and upper-level velocity commands corresponding to the current control cycle; Based on the estimated base linear velocity, the body perception observation data corresponding to the current control cycle, and the upper-level velocity command, an extended observation vector is generated. The extended observation vector is input into the whole-body control strategy network to obtain the whole-body control command; The combined system of the two-wheeled legged robotic arm is controlled based on the whole-body control commands.
9. A training method for a base velocity estimation network, characterized in that, Applied to a training device, the method includes: Acquire training samples, which include the proprioception history sequence of the dual-wheeled legged robotic arm joint system and the actual base linear velocity corresponding to the proprioception history sequence. The ontology perception history sequence is input into the base velocity estimation network to obtain the base linear velocity estimate; Gradient truncation is performed on the base linear velocity estimate used as the input policy network, and the input data of the policy network is generated based on the gradient-trunculated base linear velocity estimate. The base velocity estimation network is updated based on the estimation loss between the base linear velocity estimate without gradient truncation and the actual base linear velocity. Update the policy network based on the policy loss corresponding to the policy network; The gradient of the policy loss is not backpropagated to the base velocity estimation network.
10. The method according to claim 9, characterized in that, Before obtaining training samples, the following steps are also included: A simulation physical model of the dual-wheeled legged robotic arm combined system is constructed, the simulation physical model including a dual-wheeled legged chassis model and a robotic arm model; Run the simulation environment that includes the simulated physical model; The training samples are obtained from the simulation environment.
11. The method according to claim 9, characterized in that, Obtain training samples, including: Collect body-sensing observation data corresponding to multiple control cycles from the simulation environment; Organize the ontology perception observation data within a preset historical time window in chronological order to generate the ontology perception historical sequence. The actual base linear velocity corresponding to the ontology perception history sequence is obtained from the simulation environment.
12. The method according to claim 9, characterized in that, Before obtaining training samples, the following steps are also included: Configure domain randomization parameters; The combined system of a two-wheeled robotic arm in the simulation environment is randomized based on the domain randomization parameters. The domain randomization parameters include at least one of chassis parameters, joint parameters, sensor noise parameters, motion delay parameters, and external disturbance parameters.
13. The method according to claim 9, characterized in that, Gradient truncation is performed on the base linear velocity estimate used as input to the policy network, and input data for the policy network is generated based on the gradient-trunculated base linear velocity estimate, including: Tensor separation processing is performed on the estimated base linear velocity to obtain the base linear velocity estimated value after gradient truncation; Acquire the body-sensing observation data and upper-level velocity commands corresponding to the current control cycle; The base linear velocity estimate after gradient truncation, the body perception observation data corresponding to the current control cycle, and the upper-layer velocity command are concatenated to generate an extended observation vector. The extended observation vector is used as input data for the policy network.
14. The method according to claim 9, characterized in that, The base velocity estimation network is updated based on the estimation loss between the base linear velocity estimate without gradient truncation and the actual base linear velocity, including: The mean square error loss is determined based on the base linear velocity estimate without gradient truncation and the actual base linear velocity. The network parameters of the pedestal velocity estimation network are updated based on the mean square error loss.
15. The method according to claim 9, characterized in that, Updating the policy network based on the policy loss corresponding to the policy network includes: Based on the interaction results between the whole-body control commands output by the policy network and the simulation environment, the policy loss is determined; Update the network parameters of the policy network based on the policy loss; The strategy loss is determined based on the PPO pruning agent objective function.
16. The method according to claim 9, characterized in that, The method further includes: The base velocity estimation network is updated using the first optimizer based on the estimated loss; The policy network is updated using a second optimizer based on the policy loss. The first optimizer and the second optimizer work independently within the same training cycle.
17. The method according to claim 9, characterized in that, The method further includes: After training is complete, output the trained pedestal velocity estimation network; The trained base velocity estimation network is deployed to the control device of the dual-wheeled robotic arm joint system; The trained pedestal velocity estimation network has an input interface for receiving ontology-aware historical sequences and an output interface for outputting pedestal linear velocity estimates.
18. A base velocity estimation device, characterized in that, A control device for a dual-wheeled legged robotic arm combined system, the dual-wheeled legged robotic arm combined system including a dual-wheeled legged chassis and a robotic arm disposed on the dual-wheeled legged chassis, the device comprising: The acquisition module is used to acquire body perception observation data corresponding to multiple consecutive control cycles. The body perception observation data includes signals from the dual-wheel foot plate system, signals from the robotic arm subsystem, and historical motion signals. The generation module is used to generate a body perception history sequence based on the body perception observation data corresponding to the consecutive multiple control cycles. The estimation module is used to input the ontology perception history sequence into the base velocity estimation network to obtain the base linear velocity estimate of the dual-wheeled manipulator joint system; The output module is used to output the estimated value of the base linear velocity, which is used for the whole-body coordination control of the dual-wheeled legged robotic arm combined system.
19. A training device for a base velocity estimation network, characterized in that, Applied to training equipment, the device includes: The sample acquisition module is used to acquire training samples, which include the proprioception history sequence of the dual-wheeled legged robotic arm joint system and the actual base linear velocity corresponding to the proprioception history sequence. The estimation module is used to input the ontology perception history sequence into the base velocity estimation network to obtain the base linear velocity estimate. The truncation module is used to perform gradient truncation processing on the base linear velocity estimate used as input to the policy network, and to generate the input data of the policy network based on the gradient-trunculated base linear velocity estimate. The first update module is used to update the base velocity estimation network based on the estimated loss between the base linear velocity estimate without gradient truncation and the actual base linear velocity. The second update module is used to update the policy network based on the policy loss corresponding to the policy network. The gradient of the policy loss is not backpropagated to the base velocity estimation network.
20. A computer device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 17.
21. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 17.
22. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 17.