Energy efficiency optimization motion control method and system based on deep reinforcement learning

The motion control of the wheel-leg hybrid robot is optimized by using a slip prediction model and an energy consumption difference scoring function based on deep reinforcement learning, which solves the problem of unnecessary mode switching caused by visual perception and achieves improved energy efficiency and enhanced control stability.

CN120742701AActive Publication Date: 2025-10-03CHANGSHU INSTITUTE OF TECHNOLOGY

Patent Information

Application Number
CN202511271980.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-10-03
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Existing wheel-leg hybrid robots rely on visual perception to switch modes in complex terrain, which leads to unnecessary leg movement intervention, increased energy consumption, decreased system energy efficiency and weakened control stability.

Method used

A slip prediction model based on deep reinforcement learning is used to combine terrain data, load data and operating status data to generate a slip probability map. Comprehensive decisions are made through energy consumption difference and scoring function to optimize motion mode switching.

Benefits of technology

Reduce unnecessary leg switching movements, improve the pertinence and energy efficiency of motion mode switching, reduce energy consumption, extend battery life, and improve system reliability and terrain adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120742701A_ABST
    Figure CN120742701A_ABST
Patent Text Reader

Abstract

The invention discloses an energy efficiency optimization motion control method and system based on deep reinforcement learning, and relates to the technical field of robot motion control, and the method comprises the steps: receiving load data, operation state data and topographic data of a robot in a current region in a wheel type mode, generating a slip probability graph through a slip prediction model, calculating an energy consumption difference value of switching to a leg type mode and a wheel type keeping mode in combination with a preset terrain-energy consumption mapping table, and generating a control instruction of keeping the wheel type mode or switching to the leg type mode by judging whether the maximum slippage probability in the slippage probability graph exceeds a preset probability threshold value or balancing the slippage risk and the energy consumption cost based on a preset scoring function; the method has the beneficial effects that the real slippage risk can be identified, and the situation that the complex appearance of the terrain is mistakenly judged as inevitable slippage is avoided, so that unnecessary leg type switching actions are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot motion control technology, and in particular to an energy-efficiency optimized motion control method and system based on deep reinforcement learning. Background Art

[0002] With the widespread promotion of mobile robots in application scenarios such as industrial inspection, disaster relief, and field exploration, robots need to operate autonomously for a long time in unknown or rapidly changing complex terrain environments. Traditional wheeled robots have become the mainstream platform as early as the mid-20th century due to their simple structure, low energy consumption, and fast speed. In the past decade, multi-degree-of-freedom bionic legged robots have gradually emerged. Their flexible, maneuverable, and adaptable characteristics have prompted researchers to organically integrate the wheeled and legged mechanisms and propose a wheel-legged composite mobile platform.

[0003] In existing technologies, wheel-leg hybrid robots typically rely on machine vision to perceive the terrain type ahead and proactively switch between wheeled and legged modes accordingly. This applies even to typical scenarios where tires are prone to slipping, such as steep slopes, loose soil, or slippery surfaces. The robot will proactively switch to legged mode upon sensing such terrain. This strategy can improve its terrain adaptability and task completion rate. However, legged motion is typically accompanied by higher energy consumption, typically several times that of wheeled motion per unit distance. Furthermore, frequent movement of leg joints, servos, or hydraulic components can easily lead to structural fatigue, thus affecting overall machine reliability. In reality, "complex terrain" does not necessarily mean "uncontrollable slippage." Some surfaces that appear to cause tire slippage can still maintain good friction, eliminating the need to switch to legged mode. In reality, the degree of slippage is often determined by a combination of factors, including terrain material and slope changes. Relying solely on visual perception to trigger mode switching can easily lead to unnecessary legged motion intervention, resulting in redundant switching, reduced system energy efficiency, and weakened control stability.

[0004] Therefore, an energy-efficiency optimization motion control method and system based on deep reinforcement learning are proposed. Summary of the Invention

[0005] In view of the above-mentioned state of the art, this application is proposed. The embodiments of this application provide an energy-efficient motion control method and system based on deep reinforcement learning, which can balance safety risks and energy costs and improve the energy efficiency of wheel-leg hybrid robots in complex terrain.

[0006] According to one aspect of the present application, an energy-efficiency optimization motion control method based on deep reinforcement learning is provided, comprising: receiving load data, operating status data and terrain data of the current area of ​​the robot in wheeled mode, wherein the current area is an area where the robot may slip; generating a slip probability map representing the slip risk at different positions on the current travel path of the robot according to the load data, operating status data and terrain data through a slip prediction model based on deep reinforcement learning; calculating the slip probability map of the robot in a preset terrain-energy consumption mapping table and the slip probability map according to the preset terrain-energy consumption mapping table and the slip probability map. The energy consumption difference between the first expected energy consumption for switching to the leg mode and the second expected energy consumption for maintaining the wheel mode on the current travel path; judging whether the maximum slip probability in the slip probability map exceeds a preset probability threshold, if so, generating an instruction to switch to the leg mode; otherwise, calculating a switching score that weighs the safety risk and energy consumption cost according to the maximum slip probability and the energy consumption difference through a preset scoring function; judging whether the switching score is greater than a preset scoring threshold, if so, generating an instruction to maintain the wheel mode, otherwise, generating an instruction to switch to the leg mode; and sending the instruction to the instruction execution end of the robot.

[0007] According to another aspect of the present application, an energy-efficiency optimization motion control system based on deep reinforcement learning is provided, including: a data receiving module for receiving load data, operating status data and terrain data of the current area of ​​the robot in wheeled mode, wherein the current area is an area where the robot may slip; a slip prediction module for generating a slip probability map representing the slip risk at different positions on the current travel path of the robot according to the load data, operating status data and terrain data through a slip prediction model based on deep reinforcement learning; an energy consumption calculation module for calculating the energy consumption of the robot in a preset terrain-energy consumption mapping table and the slip probability map. The energy consumption difference between the first expected energy consumption for switching to the leg mode and the second expected energy consumption for maintaining the wheel mode on the current travel path; a risk decision module, used to determine whether the maximum slip probability in the slip probability map exceeds a preset probability threshold, and if so, generates an instruction to switch to the leg mode; otherwise, a switching score that weighs the safety risk and energy consumption cost is calculated according to the maximum slip probability and the energy consumption difference through a preset scoring function; a mode decision module, used to determine whether the switching score is greater than a preset scoring threshold, and if so, generates an instruction to maintain the wheel mode, otherwise, generates an instruction to switch to the leg mode; an instruction sending module, used to send the instruction to the instruction execution end of the robot.

[0008] According to another aspect of the present application, an electronic device is provided, comprising a memory and a processor, wherein the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which implement the steps of the above-described method when executed by the processor.

[0009] According to another aspect of the present application, a computer storage medium is provided, on which computer executable instructions are stored. When the computer executable instructions are executed by a processor, the steps of the above method are implemented.

[0010] Compared with the existing technology, the energy-efficiency optimized motion control method and system based on deep reinforcement learning according to the embodiments of the present application can identify the real slip risk and avoid misjudging the "complex appearance" of the terrain as "inevitable slip", thereby reducing unnecessary leg switching movements. At the same time, the energy consumption difference and the scoring function are combined for comprehensive decision-making, which improves the targetedness and energy efficiency of the robot's motion mode switching and avoids the problems of redundant switching and high energy consumption caused by relying on visual judgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0012] Figure 1 This is a flow chart of the energy-efficiency optimized motion control method based on deep reinforcement learning of the present invention.

[0013] Figure 2 This is a block diagram of the energy-efficiency optimized motion control system based on deep reinforcement learning in the present invention.

[0014] Figure 3 This is a timing diagram of the energy-efficiency optimized motion control system based on deep reinforcement learning in the present invention.

[0015] Figure 4 The present invention is a block diagram of an electronic device. DETAILED DESCRIPTION

[0016] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0017] Application Overview

[0018] In the existing technology, wheel-leg hybrid robots usually rely on visual sensors to identify terrain types and trigger mode switching in real time when they detect steep slopes, loose soil or slippery ground, which may cause tire slippage. However, the complexity of terrain appearance is not completely positively correlated with the risk of slippage. Some seemingly complex surfaces can still maintain effective tire adhesion. Relying solely on visual perception to trigger mode switching can easily lead to unnecessary leg action intervention, resulting in redundant switching, reduced system energy efficiency and weakened control stability.

[0019] In order to solve the above problems, the inventors realized that the slip risk is determined by multiple factors such as terrain, robot load and operating status, and dynamic prediction is required. They further realized that the robot's mode switching decision must balance safety risks and energy costs. Simply relying on vision cannot quantify the correlation between slip probability and energy consumption. Therefore, they proposed combining a deep reinforcement learning model to predict the slip risk distribution, and realize multi-objective optimization decision-making through energy consumption difference calculation and scoring function.

[0020] Exemplary Methods

[0021] Figure 1 The diagram illustrates an energy-efficiency optimization motion control method based on deep reinforcement learning according to an embodiment of the present application, including: receiving load data, operating status data, and terrain data of a robot in a wheeled mode in a current area, where the current area is an area where the robot may slip; generating a slip probability map representing the slip risk at different locations on the robot's current travel path based on the load data, operating status data, and terrain data using a slip prediction model based on deep reinforcement learning; calculating the energy consumption difference between a first expected energy consumption for the robot switching to a legged mode and a second expected energy consumption for maintaining the wheeled mode on the current travel path based on a preset terrain-energy consumption mapping table and the slip probability map; determining whether the maximum slip probability in the slip probability map exceeds a preset probability threshold, and if so, generating an instruction to switch to the legged mode; otherwise, calculating a switching score that balances safety risk and energy cost based on the maximum slip probability and the energy consumption difference using a preset scoring function; determining whether the switching score is greater than a preset scoring threshold, and if so, generating an instruction to maintain the wheeled mode; otherwise, generating an instruction to switch to the legged mode; and sending the instruction to the robot's instruction execution terminal.

[0022] Terrain data refers to terrain point cloud data collected by LiDAR. Operational status data includes the robot's operating parameters, such as motor speed and tilt angle. Load data represents the robot's current load distribution. The slip probability map, a visual representation of the slip risk at each point along the robot's current path, can be a heat map or a two-dimensional matrix. The terrain-energy consumption mapping table is a database that stores the unit energy consumption of motion modes corresponding to different terrain types. For example, energy consumption benchmarks for terrains like gravel and mud are established through experimental measurements.

[0023] Specifically, when the robot enters a potential slip area in wheeled mode, it collects load data and operating status data in real time, combines the terrain point cloud data obtained by the lidar, and inputs the pre-trained slip prediction model. The model outputs the slip probability of each position point on the path, and then queries the preset database to obtain the unit energy consumption of the wheeled mode and legged mode corresponding to the current terrain. The total energy consumption difference of the entire path is calculated based on the slip probability. If the maximum probability output by the slip prediction model exceeds the threshold, it will directly switch to the leg mode. Otherwise, the balance risk and energy consumption switching score will be calculated through the preset scoring function, and the decision on whether to switch to the leg mode will be made based on the calculation result.

[0024] Through the above technical solution, this application realizes the dynamic trade-off between the quantitative assessment of slip risk and energy consumption cost. Compared with relying on machine vision for mode switching, it can reduce unnecessary leg action intervention. On the premise of ensuring movement safety, it can reduce energy consumption by maintaining the wheeled mode, extend the robot's endurance, and avoid mechanical wear caused by frequent switching, thereby improving system reliability and terrain adaptability.

[0025] The present application further proposes that before determining whether the maximum slip probability in the slip probability map exceeds a preset probability threshold, it also includes: obtaining the ambient temperature data of the current area; extracting the corresponding probability adjustment function based on a preset terrain-temperature range-probability adjustment function mapping table and terrain data and ambient temperature data; and adjusting the preset probability threshold based on the ambient temperature and the probability adjustment function.

[0026] Among them, ambient temperature data refers to the thermodynamic parameters of the surface area where the robot is located, collected by temperature sensors. Specifically, it can be implemented using infrared thermal imaging modules or contact temperature probes to reflect the physical state of the surface friction characteristics changing with temperature. The terrain-temperature range-probability adjustment function mapping table refers to a database that stores the probability adjustment rules corresponding to different terrain categories and temperature range combinations. Specifically, it can be implemented using a multi-dimensional lookup table structure to establish a quantitative model of the relationship between temperature and slip probability. The probability adjustment function refers to a mathematical relationship that dynamically corrects the preset probability threshold based on the temperature parameter. Specifically, it can be implemented using a piecewise linear interpolation or nonlinear regression model to eliminate the interference of temperature changes on the slip judgment benchmark value.

[0027] Specifically, when the robot enters a potential slip zone, it simultaneously collects surface temperature data for that area. By querying a terrain-temperature range-probability adjustment function mapping table, it matches the current terrain type with the temperature range and extracts the corresponding probability adjustment function. For example, in sandy terrain, when the temperature is between 30-50°C, the probability adjustment function can be set to reduce the original preset probability threshold by 5% to 15%. This adjusted probability threshold serves as a dynamic benchmark for slip risk assessment, avoiding misjudgments of slip risk caused by using the standard criteria at room temperature even when the surface friction coefficient decreases due to rising temperature.

[0028] Through the above technical solution, the present application realizes the adaptive adjustment of the slip probability threshold under different temperature environments, avoiding the missed judgment of slip risk caused by the decrease in friction coefficient in high temperature environment, and preventing the redundant mode switching caused by excessively high friction coefficient in low temperature environment, thereby improving the reliability of the wheel-leg composite robot's motion control and the energy efficiency optimization accuracy.

[0029] The present application further proposes that determining whether the switching score is greater than a preset score threshold also includes: extracting a corresponding score adjustment function based on a preset terrain-temperature range-score adjustment function mapping table and terrain data and ambient temperature data; and adjusting the preset score threshold based on the ambient temperature data and the score adjustment function.

[0030] The terrain-temperature range-score adjustment function mapping table is a data structure that stores the score adjustment functions corresponding to different terrain types and temperature range combinations. It can be implemented as a database or hash table and is used to dynamically modify the score judgment criteria based on the real-time ambient temperature. The score adjustment function is a mathematical expression that defines the relationship between temperature and the score threshold adjustment amount. It can be implemented as linear interpolation or a piecewise function and is used to quantify the impact of temperature changes on slip risk as a threshold adjustment parameter.

[0031] Specifically, based on the terrain type and the current temperature range, the corresponding scoring adjustment function is matched from the terrain-temperature range-scoring adjustment function mapping table. Assuming that the current terrain is loose sand and the temperature is in the 30-40°C range, the scoring adjustment function corresponding to this combination is selected, and the temperature value is substituted into the calculation to obtain the scoring threshold adjustment amount. The adjusted scoring threshold will be used for subsequent switching scoring judgments, so that the change in slip risk caused by the softening of tire materials in high temperature environments is incorporated into the decision logic.

[0032] Through the above technical solution, the present application can avoid the misjudgment of mode switching caused by fixed scoring thresholds in high or low temperature environments, achieve balanced optimization of energy consumption and safety in complex temperature scenarios, and further reduce the mechanical structure loss caused by redundant mode switching.

[0033] The present application further proposes that generating a slip probability map includes: dividing the current travel path into multiple location points for slip prediction; extracting terrain sub-region data corresponding to each location point based on the terrain area; outputting the slip probability of each location point based on the terrain sub-region data, load data, and operating status data through a slip prediction model; constructing a slip probability map based on multiple location points and corresponding slip probabilities, wherein the slip probability map is a two-dimensional matrix representing the spatial distribution of slip probabilities on the current path.

[0034] Location point partitioning refers to discretizing a continuous path into a finite number of prediction units. This can be achieved through equally spaced sampling or dynamic segmentation based on terrain mutation points. Discretization reduces model computational complexity and improves prediction resolution. Terrain sub-region data refers to the set of local terrain characteristic parameters corresponding to each location point. Specifically, after acquiring 3D point cloud data using lidar or depth cameras, grid-based segmentation algorithms are used to extract parameters such as slope, friction coefficient, and surface material type to characterize the impact of local terrain on slip.

[0035] Specifically, the current travel path is discretized into multiple location points, for example, a prediction point is set every 0.5 meters. For each location point, the sub-area terrain data such as the slope and surface roughness of the area where it is located is collected, and input into the slip prediction model together with the robot's current load weight, drive wheel speed and other operating status parameters. The model processes the input data through a multi-layer neural network and outputs the probability value of slippage at each location point. Finally, the slip probabilities of all location points on the path are arranged in spatial order into a two-dimensional matrix. The row and column indices of the matrix correspond to the path coordinates, and the matrix element values ​​represent the slip risk intensity, thereby forming a slip probability distribution map covering the entire path.

[0036] Through the above technical solution, the present application can realize the refined slip risk assessment of the robot's travel path, avoid the omission or misjudgment of local risks caused by global unified judgment, thereby reducing unnecessary mode switching actions, reducing system energy consumption and improving motion control stability.

[0037] This application further proposes the steps of calculating the energy consumption difference as follows: obtaining the path length of the current travel path in the current area ; Set the path length Discrete The length of each path unit is equal. ; Calculate the average slip probability of the robot's wheeled mode in each path unit according to the slip probability map ;According to the preset terrain-energy consumption mapping table, the energy consumption per unit distance of the robot in wheeled mode is obtained , and the energy consumption per unit distance of the current area in leg mode ,in, Represents the index of the terrain in the terrain-energy consumption mapping table; calculate the first expected energy consumption:

[0038] ;

[0039] Calculate the second expected energy consumption:

[0040] ;

[0041] in, Indicates the The index of the terrain type of the path unit, Indicates the The average slip probability of path units, is the preset slip compensation coefficient, is the energy consumption of the robot when switching from wheeled mode to legged mode. The energy consumption difference is calculated based on the first expected energy consumption and the second expected energy consumption: .

[0042] The path length can be calculated from the coordinate data output by the robot's path planning module and the boundary of the current area to determine the spatial range of energy consumption calculation. The current area can be obtained by lidar or depth camera. Here, the perception boundary limit can be set to prevent the current area from being too large, resulting in excessive perception costs. Average slip probability It refers to the mean value of the slip probability in each path unit. It can be obtained by the slip probability data of the corresponding position point in the slip probability graph, and is used to quantify the slip risk of the wheeled mode in this path section. It refers to the fixed energy consumption generated by the mechanical structure movement and control system during mode switching. It can be obtained through experimental measurement or historical data statistics, and is used to correct the total energy consumption prediction value of the leg mode.

[0043] Specifically, the path discretization process divides the continuous travel path into multiple independently calculated units, and the terrain type index of each unit is used to calculate the path. Match the corresponding unit energy consumption parameters from the terrain-energy consumption mapping table. In the first expected energy consumption calculation of the wheeled mode, the energy consumption of each path unit is calculated by the basic energy consumption. Superimposed slip compensation Composition, slip probability The higher the value, the greater the energy consumption compensation. The second expected energy consumption of the leg mode is directly added to the basic energy consumption of each unit. , and additionally take into account the fixed energy consumption of mode switching ,Through segmented cumulative calculation, it can more accurately reflect the difference in the impact of terrain changes on the energy consumption of the two motion modes, thereby obtaining a reliable energy consumption difference evaluation result.

[0044] Through the above technical solution, the present application can accurately quantify the incremental impact of slip risk on the energy consumption of the wheeled mode, and accurately calculate the fixed cost of mode switching, providing reliable energy consumption difference data support for the robot mode switching decision, thereby avoiding redundant switching or delayed switching problems caused by energy consumption prediction errors, and improving the robustness of the system energy efficiency optimization control.

[0045] This application further proposes a preset scoring function:

[0046] ;

[0047] in, represents the maximum slip probability, is the energy consumption difference, and is the preset normalized reference value constant, is the preset risk weight adjustment factor, is the preset energy consumption weight adjustment coefficient, and Respectively represent the importance of slip risk and the importance of energy consumption difference, satisfying .

[0048] in, and Specifically, the slip probability benchmark value and energy consumption benchmark value based on historical data statistics are used to eliminate the impact of different dimensions on the score calculation. and Specifically, this can be achieved by using a dynamic adjustment algorithm or experience value setting to adjust the priority of safety and economy according to mission requirements.

[0049] Specifically, the scoring function constructs a quantitative evaluation index by normalizing the maximum slip probability and the energy consumption difference and performing weighted summation. When the maximum slip probability exceeds the preset threshold, the mode switching is directly triggered. Otherwise, the slip risk and energy consumption cost are converted into comparable numerical forms. For example, when Set to 0.7 and When it is 0.3, it is more inclined to avoid slip risk; when is 0.4 and When it is 0.6, more attention is paid to energy saving.

[0050] This application further proposes: obtaining the actual energy consumption data of the robot passing through the current travel path in wheeled mode or legged mode; calculating the energy consumption deviation value based on the actual energy consumption data and the corresponding first expected energy consumption or second expected energy consumption; and adjusting the energy consumption mapping parameters of the wheeled mode or legged mode of the robot under the corresponding terrain in the terrain-energy consumption mapping table based on the energy consumption deviation value.

[0051] Actual energy consumption data refers to the energy consumption data collected in real time by sensors when the robot performs wheeled or legged motion. This data can be implemented using current sensors, voltage sensors, or power meters, and reflects the energy consumption level under real-world operating conditions. The energy consumption deviation value refers to the difference between expected and actual energy consumption. This can be achieved through subtraction or percentage error calculation and is used to quantify the model's prediction accuracy. The terrain-energy consumption mapping table is a database that stores the energy consumption parameters per unit distance for different terrain types and corresponding motion modes. This table can be implemented using a hash table or matrix structure and provides basic parameters for energy consumption prediction.

[0052] Specifically, after the robot completes its current path, the system uses embedded sensors to collect energy consumption data from the actual motion process, such as the number of current pulses in the motor in wheeled mode or the pressure curve of the hydraulic pump in legged mode. The system then compares the actual energy consumption data with the first or second expected energy consumption calculated during the mode switching decision phase. If the deviation exceeds a preset threshold, a parameter update mechanism is triggered. For example, in sandy terrain, if the actual energy consumption of the wheeled mode is 15% lower than the predicted value in the mapping table, the wheeled mode energy consumption parameters for sandy terrain in the mapping table are iteratively optimized using a gradient descent algorithm. This process is implemented through online learning, and parameter calibration is automatically completed after each path execution.

[0053] Through the above technical solution, this application solves the problem of mismatch between terrain-energy consumption mapping and the actual environment, improving the energy efficiency optimization accuracy of mode switching decisions. For example, in long-term field operations, as tire wear increases, the actual energy consumption of the wheeled mode may gradually deviate from the initial parameters. In this case, real-time calibration can avoid overly conservative switching strategies caused by model aging, thereby maintaining better energy utilization while ensuring safety.

[0054] Exemplary Systems

[0055] Figure 2~Figure 3The diagram shows an energy-efficiency optimized motion control system based on deep reinforcement learning according to an embodiment of the present application, including: a data receiving module for receiving load data, operating status data and terrain data of the current area of ​​the robot in wheeled mode, where the current area is an area where the robot may slip; a slip prediction module for generating a slip probability map representing the slip risk of different positions on the current travel path of the robot based on the load data, operating status data and terrain data through a slip prediction model based on deep reinforcement learning; an energy consumption calculation module for calculating the slip probability of the robot in the current travel path based on a preset terrain-energy consumption mapping table and the slip probability map. The first expected energy consumption of switching to the leg mode and the second expected energy consumption of maintaining the wheel mode on the path are used as the energy consumption difference; the risk decision module is used to determine whether the maximum slip probability in the slip probability map exceeds the preset probability threshold. If so, an instruction to switch to the leg mode is generated; otherwise, a switching score that weighs the safety risk and energy consumption cost is calculated according to the maximum slip probability and the energy consumption difference through a preset scoring function; the mode decision module is used to determine whether the switching score is greater than the preset scoring threshold. If so, an instruction to maintain the wheel mode is generated; otherwise, an instruction to switch to the leg mode is generated; an instruction sending module is used to send instructions to the instruction execution end of the robot.

[0056] In one example, before the risk decision module determines whether the maximum slip probability in the slip probability map exceeds a preset probability threshold, it also includes: obtaining the ambient temperature data of the current area; extracting the corresponding probability adjustment function based on a preset terrain-temperature range-probability adjustment function mapping table and terrain data and ambient temperature data; and adjusting the preset probability threshold based on the ambient temperature and the probability adjustment function.

[0057] In one example, the mode decision module determines whether the switching score is greater than a preset score threshold and further includes: extracting a corresponding score adjustment function based on a preset terrain-temperature range-score adjustment function mapping table and terrain data and ambient temperature data; and adjusting the preset score threshold based on the ambient temperature data and the score adjustment function.

[0058] In one example, the slip prediction module generates a slip probability map, including: dividing the current travel path into multiple location points for slip prediction; extracting terrain sub-region data corresponding to each location point based on the terrain area; using the slip prediction model, outputting the slip probability of each location point based on the terrain sub-region data, load data, and operating status data; and constructing a slip probability map based on multiple location points and corresponding slip probabilities, where the slip probability map is a two-dimensional matrix representing the spatial distribution of slip probabilities on the current path.

[0059] In one example, the energy consumption calculation module calculates the energy consumption difference by: obtaining the path length of the current travel path in the current area ; Set the path length Discrete The length of each path unit is equal. ; Calculate the average slip probability of the robot's wheeled mode in each path unit according to the slip probability map ;According to the preset terrain-energy consumption mapping table, the energy consumption per unit distance of the robot in wheeled mode is obtained , and the energy consumption per unit distance of the current area in leg mode ,in, Represents the index of the terrain in the terrain-energy consumption mapping table; calculate the first expected energy consumption:

[0060] ;

[0061] Calculate the second expected energy consumption:

[0062] ;

[0063] in, Indicates the The index of the terrain type of the path unit, Indicates the The average slip probability of path units, is the preset slip compensation coefficient, is the energy consumption of the robot when switching from wheeled mode to legged mode. The energy consumption difference is calculated based on the first expected energy consumption and the second expected energy consumption: .

[0064] In one example, the preset scoring function in the risk decision module is:

[0065] ;

[0066] in, represents the maximum slip probability, is the energy consumption difference, and is the preset normalized reference value constant, is the preset risk weight adjustment factor, is the preset energy consumption weight adjustment coefficient, and Respectively represent the importance of slip risk and the importance of energy consumption difference, satisfying .

[0067] In one example, the system also includes a feedback optimization module: obtaining actual energy consumption data of the robot passing through the current travel path in wheeled mode or legged mode; calculating an energy consumption deviation value based on the actual energy consumption data and the corresponding first expected energy consumption or second expected energy consumption; and adjusting the energy consumption mapping parameters of the robot's wheeled mode or legged mode under the corresponding terrain in the terrain-energy consumption mapping table based on the energy consumption deviation value.

[0068] Exemplary electronic devices

[0069] Figure 4 The figure shows an electronic device according to an embodiment of the present application. The electronic device can be the mobile device itself, or a stand-alone device independent of the mobile device, which can communicate with the mobile device to receive collected input signals from the mobile device and send the selected target driving behavior to the mobile device.

[0070] Figure 4 The figure shows a block diagram of an electronic device according to an embodiment of the present application.

[0071] like Figure 4 As shown, the electronic device includes one or more processors and memory.

[0072] The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0073] The memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor may execute the program instructions to implement the driving behavior decision-making method of each embodiment of the present application described above and / or other desired functions.

[0074] In one example, the electronic device may further include an input device and an output device, and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0075] Of course, to simplify, Figure 4 Only some of the components in the electronic device related to the present application are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.

[0076] Exemplary computer-readable media

[0077] An embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, causes the processor to execute the steps of the driving behavior decision-making method according to various embodiments of the present application described in the above “Exemplary Method” section of this specification.

[0078] Computer-readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or components, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0079] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this application are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of this application. In addition, the specific details disclosed above are merely illustrative and facilitating understanding, and are not restrictive. The above details do not limit this application to necessarily being implemented using the above specific details.

[0080] The block diagrams of the devices, devices, equipment, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0081] It should also be noted that in the apparatus, device, and method of the present application, each component or each step can be decomposed and / or recombined, and such decomposition and / or recombination should be regarded as equivalent solutions of the present application.

[0082] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0083] The above description has been provided for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. An energy-efficient motion control method based on deep reinforcement learning is applied to the mode switching control of a wheel-leg hybrid robot, characterized by: include: receiving load data, operating status data, and terrain data of a current area of ​​the robot in a wheeled mode, wherein the current area is an area where the robot may slip; generating, by a slip prediction model based on deep reinforcement learning, a slip probability map representing the slip risk at different locations on the current travel path of the robot according to the load data, the operating status data, and the terrain data; Calculating, based on a preset terrain-energy consumption mapping table and the slip probability map, an energy consumption difference between a first expected energy consumption of the robot switching to the legged mode and a second expected energy consumption of the robot maintaining the wheeled mode on the current travel path; determining whether the maximum slip probability in the slip probability map exceeds a preset probability threshold, and if so, generating an instruction to switch to the leg mode; otherwise, calculating a switching score that balances safety risk and energy cost based on the maximum slip probability and the energy consumption difference using a preset scoring function; determining whether the switching score is greater than a preset score threshold, and if so, generating an instruction to maintain the wheel mode; otherwise, generating an instruction to switch to the leg mode; Send the command to the command execution terminal of the robot.

2. The energy efficiency optimization motion control method based on deep reinforcement learning according to claim 1 is characterized in that: Before determining whether the maximum slip probability in the slip probability map exceeds a preset probability threshold, the method further includes: Obtaining ambient temperature data of the current area; Extracting a corresponding probability adjustment function according to a preset terrain-temperature range-probability adjustment function mapping table and the terrain data and the ambient temperature data; The preset probability threshold is adjusted according to the ambient temperature and the probability adjustment function.

3. The energy efficiency optimization motion control method based on deep reinforcement learning according to claim 2 is characterized in that: The determining whether the handover score is greater than a preset score threshold further includes: Extracting a corresponding score adjustment function according to a preset terrain-temperature range-score adjustment function mapping table and the terrain data and the ambient temperature data; The preset scoring threshold is adjusted according to the ambient temperature data and the scoring adjustment function.

4. The energy efficiency optimization motion control method based on deep reinforcement learning according to claim 1 is characterized in that: Generating a slip probability map representing slip risks at different positions on the current travel path of the robot includes: Dividing the current travel path into a plurality of position points for slip prediction; Extracting terrain sub-region data corresponding to each of the position points according to the terrain region; Outputting the slip probability of each of the position points according to the terrain sub-region data, the load data, and the operating status data through the slip prediction model; The slip probability map is constructed according to the plurality of position points and the corresponding slip probabilities. The slip probability map is a two-dimensional matrix representing the spatial distribution of the slip probabilities on the current path.

5. The energy efficiency optimization motion control method based on deep reinforcement learning according to claim 1 is characterized in that: The calculation steps of the energy consumption difference are: Get the path length of the current path within the current area ; The path length Discrete The length of each path unit is equal. ; Calculate the average slip probability of the wheeled mode of the robot in each path unit according to the slip probability map ; The energy consumption per unit distance of the robot passing through the current area in wheeled mode is obtained according to the preset terrain-energy consumption mapping table. , and the energy consumption per unit distance of the current area in leg mode ,in, Represents the index of the terrain in the terrain-energy consumption mapping table; Calculate the first expected energy consumption: ; Calculate the second expected energy consumption: ; in, Indicates the The index of the terrain type of the path unit, Indicates the The average slip probability of path units, is the preset slip compensation coefficient, The preset energy consumption of the robot when switching from the wheeled mode to the legged mode; The energy consumption difference is obtained by calculating the first expected energy consumption and the second expected energy consumption: 。 6. The energy efficiency optimization motion control method based on deep reinforcement learning according to claim 1 is characterized in that: The preset scoring function is specifically: ; Among them, the represents the maximum slip probability, is the energy consumption difference, and is the preset normalized reference value constant, is the preset risk weight adjustment factor, is the preset energy consumption weight adjustment coefficient, and Respectively represent the importance of slip risk and the importance of energy consumption difference, satisfying .

7. The energy efficiency optimization motion control method based on deep reinforcement learning according to claim 1 is characterized in that: Also includes: Acquiring actual energy consumption data of the robot passing through the current travel path in a wheeled mode or a legged mode; Calculating an energy consumption deviation value based on the actual energy consumption data and the corresponding first expected energy consumption or the second expected energy consumption; According to the energy consumption deviation value, the energy consumption mapping parameters of the wheeled mode or legged mode of the robot under the corresponding terrain in the terrain-energy consumption mapping table are adjusted.

8. Energy efficiency optimization motion control system based on deep reinforcement learning, characterized by: include: a data receiving module, configured to receive load data, operating status data, and terrain data of a current area in which the robot is in a wheeled mode, wherein the current area is an area in which the robot may slip; a slip prediction module, configured to generate a slip probability map representing slip risks at different locations on the current travel path of the robot based on the load data, the operating status data, and the terrain data using a slip prediction model based on deep reinforcement learning; an energy consumption calculation module, configured to calculate, based on a preset terrain-energy consumption mapping table and the slip probability map, an energy consumption difference between a first expected energy consumption of the robot switching to the legged mode and a second expected energy consumption of the robot maintaining the wheeled mode on the current travel path; a risk decision module, configured to determine whether the maximum slip probability in the slip probability map exceeds a preset probability threshold, and if so, generate an instruction to switch to the leg mode; otherwise, calculate a switching score that balances safety risk and energy cost based on the maximum slip probability and the energy consumption difference using a preset scoring function; a mode decision module, configured to determine whether the switching score is greater than a preset score threshold, and if so, generate an instruction to maintain the wheeled mode; otherwise, generate an instruction to switch to the legged mode; The instruction sending module is used to send the instruction to the instruction execution end of the robot.

9. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer storage medium having computer-executable instructions stored thereon, characterized in that: When the computer-executable instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multifunctional leg-and-wheel combination robot and multi-movement-mode intelligent switching method thereof

    CN103786806A

  • Robot capable of being switched between wheel mode and leg mode

    CN104527835A

  • Method and device for determining constraint relation data of wheel-legged robot and medium

    CN116834865A

  • Wheel-leg robot wheel-foot switching control method based on BP neural network

    CN116859975A

  • Biped Robot which transforms itself into wheel-typevehicle and its method of operation

    KR1020050088695A

Cited By

  • Robot motion balance control method and system

    CN121254886A