Network training method and device, electronic equipment and computer readable storage medium
By encoding and predicting the state data of intelligent devices and the distance to obstacles, an action network is trained to improve the accuracy of robot movement in complex dynamic scenarios, solving the problem of inaccurate robot movement and achieving more precise navigation and obstacle avoidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UBTECH ROBOTICS CORP LTD
- Filing Date
- 2025-12-02
- Publication Date
- 2026-04-21
AI Technical Summary
The robot's motion accuracy is not high in complex and dynamic scenarios.
By calling a pre-defined action network, the state data and historical action data of the smart device in multiple time frames are encoded. Action prediction is performed by combining obstacle distance sequences. The action network is trained based on reward information and loss values, and the action network is optimized to improve motion accuracy.
It improves the accuracy of smart devices in complex and dynamic scenarios, can accurately predict the operation actions at the current moment, and enhances navigation capabilities in complex environments.
Smart Images

Figure CN121902899A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot navigation, and more particularly to a network training method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] Humanoid robot navigation involves multiple disciplines, including robotics, computer science, control theory, sensor technology, and artificial intelligence. Research on humanoid robot navigation has promoted interdisciplinary integration. For example, path planning in robot navigation requires the application of mathematical optimization theory and algorithms, combined with the robot's kinematics and dynamics models, thus facilitating the integration of control theory and robotics. Humanoid robots have broad application prospects in service sectors such as healthcare and logistics. For instance, in hospitals, humanoid robots can assist medical staff in transporting medicines and equipment, and provide patient guidance services; in elderly care facilities, robots can accompany the elderly, providing daily living assistance and health monitoring. To achieve these service functions, robots need to possess the ability to navigate autonomously in complex environments.
[0003] In related technologies, the accuracy of robot movement in complex dynamic scenarios is not high. Summary of the Invention
[0004] This application provides a network training method, apparatus, electronic device, and computer-readable storage medium, which can improve the motion accuracy of intelligent devices in complex dynamic scenarios.
[0005] The technical solution of this application embodiment is implemented as follows: This application provides a network training method, the method comprising: A preset action network is invoked to encode the state data of the smart device in multiple time frames and the historical action data corresponding to each time frame to obtain a first vector. The historical action data is the action data in the previous time frame of the current time frame. Determine the distance sequence corresponding to each time frame, and encode the distance sequences corresponding to multiple time frames to obtain a second vector, wherein the distance sequence includes the distance between the smart device and each obstacle; Based on the first vector and the second vector, action prediction is performed to obtain the predicted action data for the current time frame; Reward information is determined based on the state data of the current time frame, the distance sequence, and the predicted action data, and the loss value is determined based on the reward information; The action network is trained based on the loss value to obtain a trained action network, which is then used to control the movement of the smart device.
[0006] This application provides a network training device, including: The first encoding module is used to call a preset action network to encode the state data of the smart device in multiple time frames and the historical action data corresponding to each time frame to obtain a first vector. The historical action data is the action data in the previous time frame of the time frame, and the multiple time frames include the current time frame. The second encoding module is used to determine the distance sequence corresponding to each time frame and to encode the distance sequences corresponding to multiple time frames to obtain a second vector, wherein the distance sequence includes the distance between the smart device and each obstacle; The action prediction module is used to perform action prediction based on the first vector and the second vector to obtain the predicted action data of the current time frame; The loss determination module is used to determine reward information based on the state data, distance sequence, and predicted action data of the current time frame, and to determine the loss value based on the reward information; The network training module is used to train the action network based on the loss value to obtain the trained action network, so as to control the movement of the smart device using the trained action network.
[0007] This application provides an electronic device, the electronic device comprising: Memory is used to store executable instructions or computer programs. The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the network training method provided in the embodiments of this application.
[0008] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs, which are executed by a processor to implement the network training method provided in this application.
[0009] This application provides a computer program product, including computer-executable instructions or a computer program, which, when executed by a processor, implements the network training method provided in this application.
[0010] The embodiments of this application have the following beneficial effects: By applying the embodiments of this application, a preset action network is invoked to encode the state data of the smart device across multiple time frames and the historical action data corresponding to each time frame, resulting in a first vector. The historical action data refers to the action data in the previous time frame, and the multiple time frames include the current time frame. Next, a distance sequence corresponding to each time frame is determined, and the distance sequence corresponding to multiple time frames is encoded to obtain a second vector. The distance sequence includes the distances between the smart device and various obstacles. Then, action prediction is performed based on the first and second vectors to obtain the predicted action data for the current time frame. This method can combine and encode multi-source information such as the state data of the smart device across multiple time frames, historical action data, and the distances between the smart device and various obstacles, making full use of the time-series characteristics of the data and improving the accuracy of action prediction. Then, reward information is determined based on the state data, distance sequence, and predicted action data of the current time frame, and a loss value is determined based on the reward information. The action network is then trained based on the loss value to obtain a trained action network, which is used to control the movement of the smart device. In this way, by adjusting and optimizing the action network using state data, distance sequences, and predicted action data, the action network can capture dynamic changes in the environment and accurately predict the current action. This improves the accuracy of smart device movement in complex dynamic scenarios when using the trained action network to control the movement of smart devices. Attached Figure Description
[0011] Figure 1 This is a schematic diagram illustrating the application mode of the network training method provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figure 3 This is a first flowchart illustrating the network training method provided in this application embodiment; Figure 4 This is a schematic diagram of the second process of the network training method provided in the embodiments of this application; Figure 5 This is a schematic diagram of the third process of the network training method provided in the embodiments of this application; Figure 6 This is a schematic diagram of the fourth process of the network training method provided in the embodiments of this application. Figure 7 This is a schematic diagram of the action network training process provided in an embodiment of this application; Figure 8 This is a schematic flowchart of the motion control method provided in the embodiments of this application; Figure 9 This is a schematic diagram of an obstacle measured by lidar rays in a grid map provided in an embodiment of this application; Figure 10This is a schematic diagram of the network structure provided in the embodiments of this application.
[0012] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0015] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0016] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0017] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0018] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for descriptive purposes only and is not intended to limit the scope of this application.
[0019] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0020] 1) Action Network: This is the decision-making module responsible for generating specific robot actions. It is a type of policy network and can output action data (such as movement direction, speed, turning angle, etc.) based on the current environmental state (such as sensor data, map information, target position, etc.).
[0021] 2) Evaluation Network: It is a predictive model that evaluates the long-term value of actions. It is an extension of the State Value Function and can quantify the potential benefits of the current state or action sequence (such as distance to the target, energy consumption, and risk level).
[0022] This application provides a network training method, apparatus, electronic device, and computer-readable storage medium, which can improve the motion accuracy of intelligent devices in complex dynamic scenarios.
[0023] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. These electronic devices can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or as servers. The following will describe exemplary applications when the device is implemented as a server.
[0024] See Figure 1 , Figure 1 This is a schematic diagram illustrating the application mode of the network training method provided in the embodiments of this application, for example. Figure 1 The system involves server 200, network 300, and terminal 400. Terminal 400 connects to server 200 through network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.
[0025] During network training, server 200 calls a preset action network to encode the state data of the intelligent device across multiple time frames and the historical action data corresponding to each time frame, obtaining a first vector. The historical action data refers to the action data in the previous time frame, and the multiple time frames include the current time frame. A distance sequence corresponding to each time frame is determined, and the distance sequences corresponding to multiple time frames are encoded to obtain a second vector, which includes the distances between the intelligent device and various obstacles. Action prediction is performed based on the first and second vectors to obtain the predicted action data for the current time frame. Reward information is determined based on the state data, distance sequence, and predicted action data of the current time frame, and a loss value is determined based on the reward information. The action network is trained based on the loss value to obtain the trained action network, which is then used to control the movement of the intelligent device. The intelligent device can be a humanoid robot used in navigation scenarios, i.e., terminal 400.
[0026] When controlling the movement of the smart device, server 200 invokes the trained motion network to predict actions, obtains predicted action data for the current time frame, and sends the predicted action data to terminal 400. For example, in a robot navigation scenario, terminal 400 performs real-time path planning and dynamic obstacle avoidance based on the predicted action data. Or, in a logistics and warehousing scenario, terminal 400 dynamically schedules logistics resources based on the predicted action data.
[0027] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may be a terminal or a server. Figure 2 The illustrated electronic device includes at least one processor 410, a memory 450, and at least one network interface 420. The various components of the electronic device are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2 The general labeled all buses as Bus System 440.
[0028] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor.
[0029] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.
[0030] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0031] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0032] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, for implementing various basic business functions and handling hardware-based tasks.
[0033] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including Bluetooth, WiFi, and Universal Serial Bus (USB).
[0034] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A network training device 455 stored in memory 450 is shown. It can be software in the form of programs and plug-ins, including the following software modules: a first encoding module 4551, a second encoding module 4552, an action prediction module 4553, a loss determination module 4554, and a network training module 4555. These modules are logically related and can therefore be arbitrarily combined or further divided according to the functions they implement. The functions of each module will be described below.
[0035] The network training method provided in this application will be described in conjunction with exemplary applications and implementations of the server devices provided in the embodiments of this application.
[0036] The following describes the network training method provided in the embodiments of this application. For example, in order to facilitate understanding of the network training method provided in the embodiments of this application, the method is described using a navigation scenario of a humanoid robot as an example.
[0037] As mentioned above, the electronic device implementing the network training method of this application embodiment can be a terminal, a server, or a combination of both. The following explanation uses an electronic device as a server as an example to illustrate the network training method provided in this application embodiment. See also... Figure 3 , Figure 3 This is a first flowchart illustrating the network training method provided in this application embodiment, which will be combined with... Figure 3 The steps shown are explained.
[0038] In step 301, a preset action network is invoked to encode the state data of the smart device in multiple time frames and the historical action data corresponding to each time frame to obtain the first vector.
[0039] Here, the preset action network can be a Deep Q-Network (DQN), a Proximal Policy Optimization (PPO), or something similar. The action network is the decision-making module responsible for generating the robot's specific actions. It is a type of policy network implementation and can output action data (such as movement direction, speed, and turning angle) based on the current environmental state (such as sensor data, map information, target location, etc.).
[0040] Intelligent devices can be humanoid robots, stationary robots, wheeled robots, etc. State data is used to characterize the state of the intelligent device itself. State data includes the base linear velocity, yaw rate, position error and orientation error of the device itself relative to the target position in the coordinate system of the intelligent device itself.
[0041] Example, state data Depend on It consists of two parts. Represents the linear velocity of the base in the coordinate system of the intelligent device body ( ) and yaw rate , This represents the position error of the smart device relative to the target position in its own coordinate system. and orientation error .
[0042] Historical action data refers to the action data in the previous time frame of a given time frame. Multiple time frames include multiple historical time frames and the current time frame. A time frame refers to a time interval. The state data of a smart device in each time frame and the corresponding historical action data for each time frame can be represented as follows: ,in To represent historical motion data, a preset motion network is invoked to encode the state data of multiple time frames and the historical motion data corresponding to each time frame, resulting in a first vector. For example, the first vector is represented as... .
[0043] Continue to refer to Figure 3 In step 302, the distance sequence corresponding to each time frame is determined, and the distance sequences corresponding to multiple time frames are encoded to obtain the second vector.
[0044] Here, the distance sequence includes the distances between the smart device and each obstacle. For example, the distance sequence corresponding to each time frame is represented as follows: Encoding the distance sequences corresponding to multiple time frames yields a second vector. For example, the second vector is represented as... ,in .
[0045] In some embodiments, determining the distance sequence corresponding to each time frame can be achieved through the following steps: within each time frame, controlling the smart device to emit multiple detection rays in a preset direction; for each detection ray, determining the obstacle distance corresponding to the detection ray; combining the obstacle distances corresponding to each detection ray to obtain the distance sequence corresponding to the time frame.
[0046] Here, the smart device carries a lidar sensor. Within each time frame, the device emits multiple detection rays in a preset direction. These detection rays, also known as lidar rays, have their azimuth intervals set in degrees. The ratio of 360 degrees to the interval is used to determine the number of detection rays. For example, the azimuth interval for each detection ray is... The number of detection rays is .
[0047] The detection rays that detect obstacles are designated as target detection rays. For each obstacle around the smart device, at least one target detection ray is identified that detects the obstacle. The length of this target detection ray intercepted by the obstacle is determined as the distance between the smart device and the obstacle, i.e., the obstacle distance corresponding to the target detection ray. For detection rays that do not detect obstacles, an infinite or fixed value can be defined as the obstacle distance corresponding to that ray. The obstacle distances corresponding to each detection ray are combined to obtain a distance sequence corresponding to a time frame.
[0048] In this embodiment, for each detection ray, the distance to the obstacle corresponding to the detection ray is determined, and the distances to obstacles corresponding to each detection ray are combined to obtain a distance sequence corresponding to the time frame. The distance sequence can reflect the distance information of obstacles around the smart device at the current moment, so that the distance sequence can be used for action prediction, thereby improving the navigation capability of the smart device in complex environments.
[0049] Continue to refer to Figure 3 In step 303, action prediction is performed based on the first vector and the second vector to obtain the predicted action data for the current time frame.
[0050] Here, the predicted motion data for the current time frame includes the linear velocity, yaw rate, and joint rotation speed of the intelligent device. Linear velocity is the instantaneous rate at which the intelligent device moves along a straight line or path, used to measure the robot's movement efficiency. Yaw rate is the angular velocity of the wheelset or joints when the intelligent device turns. Joint rotation speed is the instantaneous angular velocity of the motors of the intelligent device's joints (such as robotic arms or legs).
[0051] In some embodiments, see Figure 4 , Figure 4 This is a schematic diagram of the second process of the network training method provided in the embodiments of this application. Figure 3 Step 303 shown can be achieved through... Figure 4 Steps 3031 to 3034 are implemented, and will be explained in detail below.
[0052] In step 3031, the action data of the smart device in the historical time frame is obtained, and the action data is encoded to obtain the third vector.
[0053] Here, the motion data of the smart device in historical time frames is collected by sensors (such as accelerometers, joint position sensors, etc.) carried on the smart device. Multiple data such as motion phase signals, joint velocities, and overall velocities in the motion data are encoded to obtain the encoding results. The encoding results are then spliced and fused to obtain the third vector.
[0054] In step 3032, the first vector, the second vector, and the third vector are fused to obtain a fused vector.
[0055] Here, the first vector and the second vector are concatenated to obtain the first concatenated vector. For example, the first vector is... Second vector The concatenation process is performed, and the first concatenated vector is represented as follows: The first and third concatenation vectors are concatenated to merge them into a single vector. For example, the first concatenation vector... and the third vector After concatenation, the fused vector is represented as follows: .
[0056] In step 3033, temporal prediction is performed based on the fusion vector to obtain the hidden state data.
[0057] Here, the fused vector is input into a gated recurrent unit of a recurrent neural network to perform temporal prediction on the fused vector, obtaining the hidden state data. For example, the hidden state data is represented as follows: ,in Indicates a gated loop unit. Represents the fusion vector. This represents the hidden state data of the previous time step.
[0058] In step 3034, action prediction is performed based on the hidden state data to obtain the predicted action data for the current time frame.
[0059] Here, action sampling, action constraints, and clipping are performed based on the latent state data to achieve action prediction and obtain the predicted action data for the current time frame.
[0060] In some embodiments, see Figure 5 , Figure 5 This is a schematic diagram of the third process of the network training method provided in the embodiments of this application. Figure 4 Step 3034 shown can be achieved through... Figure 5 Steps 30341 to 30343 are implemented, and will be explained in detail below.
[0061] In step 30341, the probability distribution corresponding to the hidden state data is determined, and the hidden state data is sampled based on the probability distribution to obtain the first action data.
[0062] Here, an action network is used to perform network mapping on the hidden state data, and the probability distribution is output through the action network. The probability distribution is represented as follows: ,in The mean and variance of the probability distribution represent the probability distribution. Related. Based on the probability distribution, actions are sampled from the hidden state data to obtain the first action data. Action sampling is represented as... (mean is) Variance is ),in This represents the first action data.
[0063] In step 30342, based on preset motion constraint data, the first motion data is subjected to boundary constraint processing to obtain the second motion data.
[0064] Here, the preset motion constraint data includes minimum motion constraint value and maximum motion constraint value. Boundary constraint value is determined based on the minimum and maximum motion constraint values, and the first motion data is subjected to boundary constraint processing based on the boundary constraint value to obtain the second motion data.
[0065] In some embodiments, the boundary constraint processing of the first action data to obtain the second action data can be achieved through the following steps: determining a first boundary constraint value based on the difference between the minimum action constraint value and the maximum action constraint value; determining a second boundary constraint value based on the sum of the minimum action constraint value and the maximum action constraint value; and performing boundary constraint processing on the first action data based on the first boundary constraint value and the second boundary constraint value to obtain the second action data.
[0066] Here, the difference between the minimum and maximum motion constraint values is averaged to obtain the first boundary constraint value. The sum of the minimum and maximum motion constraint values is averaged to obtain the second boundary constraint value. The first boundary constraint value is multiplied by the first motion data to obtain the product, and the product is added to the second boundary constraint value to obtain the second motion data.
[0067] For example, the minimum action constraint value is The maximum action constraint value is First boundary constraint value for Second boundary constraint value for The first action data is The second action data is then represented as .
[0068] In this embodiment, the first action data is subjected to boundary constraint processing based on the first boundary constraint value and the second boundary constraint value to obtain the second action data. The boundary constraint processing can filter out outliers and unreasonable action data in the action data, so that the action network can focus more on learning effective action data features, thereby improving the accuracy of action prediction.
[0069] Continue to refer to Figure 5 In step 30343, the second action data is cropped based on the action constraint data to obtain the predicted action data for the current time frame.
[0070] Here, based on the minimum and maximum motion constraint values in the motion constraint data, the second motion data is cropped to obtain the predicted motion data for the current time frame. For example, the predicted motion data is determined using the following formula (1):
[0071] in, This represents the predicted action data. These represent the minimum and maximum motion constraint values, respectively. Indicates the first boundary constraint value , Indicates the second boundary constraint value .
[0072] In this embodiment, the first vector, the second vector, and the third vector are fused to obtain a fused vector. Temporal prediction is performed based on the fused vector to obtain hidden state data. Action prediction is then performed based on the hidden state data to obtain the predicted action data for the current time frame. This approach can introduce multi-frame observation stacking data and the memory mechanism of a recurrent neural network during action prediction, enabling temporal modeling and prediction of non-stationary environments. Furthermore, it takes into account the dependency between historical action data and predicted action data, thus improving the feasibility of predicting action data.
[0073] Continue to refer to Figure 3 In step 304, reward information is determined based on the state data, distance sequence, and predicted action data of the current time frame, and loss value is determined based on the reward information.
[0074] Here, the reward information includes navigation reward information and motion reward information, and a preset evaluation network is invoked to construct a loss value based on the reward information.
[0075] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of the fourth process of the network training method provided in the embodiments of this application. Figure 3 In step 304 shown, "determining reward information based on the state data, distance sequence, and predicted action data of the current time frame" can be achieved through... Figure 6 Steps 3041 to 3043 are implemented, and will be explained in detail below.
[0076] In step 3041, navigation reward information is determined based on the state data, distance sequence, and predicted action data of the current time frame.
[0077] Here, the state data of the current time frame includes the positional error between the current position and the target position of the smart device, the lateral velocity of the smart device, and the longitudinal velocity of the smart device.
[0078] In some embodiments, determining navigation reward information based on the state data, distance sequence, and predicted motion data of the current time frame can be achieved through the following steps: performing nonlinear processing on the position error of the current time frame to obtain the target approach reward value of the current time frame; normalizing the distance sequence of the current time frame to obtain the obstacle distance value of the current time frame; determining the orientation indication information of the current time frame, and performing consistency constraint processing on the lateral velocity, longitudinal velocity, and position error based on the orientation indication information to obtain the consistency reward value of the current time frame; determining the backtracking indication information of the current time frame, and performing exponential transformation processing on the lateral velocity based on the backtracking indication information to obtain the backtracking penalty value of the current time frame; performing motion smoothing analysis based on the predicted motion data of the current time frame and the motion data of the previous time frame to obtain the smoothness reward value of the current time frame; and combining the target approach reward value, obstacle distance value, consistency reward value, backtracking penalty value, and smoothness reward value to obtain the navigation reward information.
[0079] Here, the position error of the current time frame is the position error of the smart device's own position relative to the target position in the current time frame. The Euclidean distance corresponding to the position error of the current time frame is calculated using the Euclidean norm. The hyperbolic tangent function is used to perform nonlinear processing on the Euclidean distance corresponding to the position error and the preset adjustment coefficient to obtain the target proximity reward value of the current time frame. For example, the target proximity reward value is determined using the following formula (2):
[0080] in, This indicates that the target is close to the reward value. This represents the Euclidean distance corresponding to the position error. This indicates the preset adjustment coefficient.
[0081] Logarithmic averaging is performed on the distance sequence of the current time frame and the preset reference distance to achieve normalization, thereby obtaining the obstacle distance value of the current time frame. For example, the obstacle distance value is determined by the following formula (3):
[0082] in, Indicates the obstacle distance value. This represents the distance sequence of the current time frame. This indicates the preset reference distance.
[0083] In the current time frame, information indicating a positive lateral velocity of the smart device is defined as the first orientation indication information. The lateral and longitudinal velocities of the smart device are combined to form the volume coordinate velocity, and information indicating a positive dot product between the velocity vector corresponding to the volume coordinate velocity and the vector corresponding to the position error is defined as the second orientation indication information. The intersection of the first and second orientation indication information is defined as the orientation indication information for the current time frame. This orientation indication information is used to represent the orientation of the smart device as closely as possible to the line connecting its own position and the target position.
[0084] For example, the orientation indication information is represented as ,in, Indicates the first orientation indication information. This indicates the second orientation information.
[0085] Using a preset clipping threshold, the vector corresponding to the position error is normalized to obtain a normalized reference vector. The absolute value of the dot product between the velocity vector corresponding to the volume coordinate velocity and the normalized reference vector is determined, and the orientation indication information is used to perform consistency constraint processing on the absolute value of the dot product to obtain the consistency reward value of the current time frame. For example, the consistency reward value is determined by the following formula (4):
[0086] in, This represents the consistency reward value. This represents the velocity vector corresponding to the volume coordinate velocity. This represents the vector corresponding to the position error. Represents the normalized reference vector. Indicates the clipping threshold. This indicates the direction of the direction.
[0087] In the current time frame, information indicating that the horizontal speed of the smart device is negative is identified as the first back-back indication information, and information indicating that the absolute value of the horizontal speed is greater than the preset minimum speed is identified as the second back-back indication information. The first and second back-back indication information are combined to form the back-back indication information for the current time frame. The difference between the absolute value of the horizontal speed and the preset minimum speed is determined, and the difference is subjected to an exponential transformation using the back-back indication information to obtain the back-back penalty value for the current time frame. For example, the back-back penalty value is determined using the following formula (5):
[0088] in, This represents the penalty value for going back. This indicates a backward instruction message. This indicates the first backward instruction message. This indicates the second back navigation instruction. This indicates the preset minimum speed.
[0089] Subtract the body coordinate velocity from the predicted motion data of the current time frame and the body coordinate velocity from the motion data of the previous time frame to obtain the change in body coordinate velocity. Subtract the yaw angular velocity from the predicted motion data of the current time frame and the yaw angular velocity from the motion data of the previous time frame to obtain the change in yaw angular velocity. Perform exponential transformation on the Euclidean norm corresponding to the change in body coordinate velocity and the absolute value of the change in yaw angular velocity to achieve motion smoothing analysis and obtain the smoothness reward value of the current time frame. For example, the smoothness reward value is determined by the following formula (6):
[0090] in, Represents the smoothness reward value. The Euclidean norm representing the change in volumetric velocity. It represents the absolute value of the change in yaw rate.
[0091] Finally, by combining the target proximity reward value, obstacle distance value, consistency reward value, backtracking penalty value, and smoothness reward value, the navigation reward information can be obtained.
[0092] In this embodiment, navigation reward information is determined based on the state data, distance sequence, and predicted action data of the current time frame. The navigation reward information includes target approach reward value, obstacle distance value, consistency reward value, backtracking penalty value, and smoothness reward value. By comprehensively considering multiple factors such as target approach, obstacle distance, and speed orientation, the navigation reward information can be determined quickly. This enables the optimal action strategy to be found quickly based on the navigation reward information, which accelerates the convergence speed of action network training and improves the stability and security of navigation path planning based on action networks.
[0093] Continue to refer to Figure 6 In step 3042, motion reward information is determined based on the state data of the current time frame and the predicted motion data.
[0094] Here, the state data of the current time frame also includes the posture data of the smart device, which can be the actual angle between the smart device and the vertical direction (such as the direction of gravity). The predicted motion data includes the predicted joint data of the smart device and the predicted ground contact timing data of the smart device. The predicted joint data can be the predicted joint acceleration of the smart device, and the predicted ground contact timing data is the predicted ground contact timing data corresponding to the foot of the smart device.
[0095] In some embodiments, determining motion reward information based on state data and predicted motion data of the current time frame can be achieved through the following steps: adjusting the difference between the posture data of the current time frame and the preset expected posture data based on a first preset coefficient to obtain the posture reward value of the current time frame; adjusting the difference between the predicted joint data and the preset expected joint data based on a second preset coefficient to obtain the joint reward value of the current time frame; adjusting the difference between the predicted ground contact timing data and the preset expected ground contact timing data based on a third preset coefficient to obtain the ground contact reward value of the current time frame; and combining the posture reward value, joint reward value, and ground contact reward value of the current time frame to obtain motion reward information.
[0096] Here, when the attitude data is the actual angle between the smart device and the vertical direction (such as the direction of gravity), the preset expected attitude data is the expected angle between the smart device and the vertical direction. The difference between the actual angle and the expected angle is multiplied by a first preset coefficient to obtain the attitude reward value for the current time frame. For example, the attitude reward value is represented as... ,in This is represented as the first preset coefficient. This is expressed as the actual included angle. It is represented as the expected angle.
[0097] When the predicted joint data is the predicted joint acceleration of the smart device, the preset expected joint data is the expected joint acceleration of the smart device. The difference between the predicted joint acceleration and the expected joint acceleration is multiplied by a second preset coefficient to obtain the joint reward value for the current time frame. For example, the joint reward value is represented as... ,in This is represented as the second preset coefficient. Represented as the desired joint acceleration, This is represented as the predicted joint acceleration.
[0098] The ground contact bonus value for the current time frame is obtained by multiplying the difference between the predicted ground contact time series data and the preset expected ground contact time series data by a third preset coefficient. For example, the ground contact bonus value is represented as... ,in This is represented as the third preset coefficient. Represented as expected ground contact time series data, This is represented as predicted ground contact time series data.
[0099] In this embodiment, motion reward information is determined based on the state data and predicted action data of the current time frame. The motion reward information includes posture reward value, joint reward value and ground contact reward value. By comprehensively considering multiple factors such as posture, joints and ground contact to determine motion reward information, the motion network can be guided to learn more general motion strategies and quickly eliminate actions that are not conducive to posture stability, precise joint control or reasonable ground contact, so that the motion network can converge to the desired motion pattern more quickly.
[0100] Continue to refer to Figure 6 In step 3043, the navigation reward information and motion reward information are superimposed to obtain reward information.
[0101] Here, the navigation reward information and the motion reward information are added together to obtain the reward information. For example, the reward information is represented as follows: .in, Indicates navigation reward information, This indicates information about exercise rewards.
[0102] In some embodiments, determining the loss value based on reward information can be achieved through the following steps: calling a preset evaluation network to evaluate the hidden state data and obtain the predicted evaluation value of the current time frame; determining the evaluation deviation value based on the reward information, the predicted evaluation value of the current time frame, and the predicted evaluation value of the next time frame; estimating the preset dominance function based on the evaluation deviation value to obtain the dominance estimate value, and constructing the loss value based on the dominance estimate value and a preset threshold.
[0103] Here, the evaluation network is a predictive model that assesses the long-term value of actions. It's an extension of the State Value Function, capable of quantifying the potential benefits (such as distance to the target, energy consumption, and risk level) of current state data or action sequences. A pre-defined evaluation network is used to map the hidden state data to achieve state evaluation, and the network outputs the predicted evaluation value for the current time frame. For example, the predicted evaluation value is represented as... ,in This represents hidden state data.
[0104] Using preset parameter values, combined with reward information, the predicted evaluation value of the current time frame, and the predicted evaluation value of the next time frame, the evaluation deviation is calculated to obtain the evaluation deviation value. For example, the evaluation deviation value is represented as... ,in, Indicates reward information, This represents the predicted evaluation value for the current time frame. This represents the predicted evaluation value for the next time frame. This represents the preset parameter value, which can be set to 0.99.
[0105] The advantage function measures the degree of advantage of taking a certain action in a given state relative to the average policy; that is, how much more reward can be obtained by taking this action compared to the average policy. An estimated advantage is obtained by performing a backward recursive calculation over multiple time steps, combining a pre-defined advantage function and an evaluation bias value. For example, the estimated advantage is expressed as... ,in, This indicates the evaluation deviation value. and These are preset parameter values, which can be set to 0.99 and 0.95 respectively. This represents the advantage estimate for the next time step.
[0106] The preset threshold is a preset pruning threshold. A pruning function is constructed based on the preset action sampling ratio and the preset threshold. The loss value is constructed by combining the advantage estimate, the preset action sampling ratio, and the pruning function. The goal of the loss value is to determine the parameters of the policy network that maximizes the loss of the action network. This policy network is used to generate the probability distribution of taking a certain action in a certain state. For example, the loss value is determined by the following formula (7):
[0107] in, This represents the loss value of the action network. This represents the clipping function. Indicates the preset action sampling ratio. This indicates a preset threshold.
[0108] In this embodiment, a preset dominance function is estimated based on the evaluation deviation value to obtain the dominance estimate. A loss value is then constructed based on the dominance estimate and a preset threshold. This loss value, constructed based on the dominance estimate and the preset threshold, provides a more effective feedback signal for the training of the action network. It can guide the action network to learn the optimal action strategy more quickly, thereby accelerating the convergence speed of the action network training.
[0109] Continue to refer to Figure 3 In step 305, the action network is trained based on the loss value to obtain the trained action network, which is then used to control the movement of the smart device.
[0110] Here, the network parameters of the action network are updated as a whole based on the loss value to obtain the trained action network. The trained action network is then used to predict actions, and the obtained predicted action data is used to control the movement of the smart device.
[0111] Example, reference Figure 7 , Figure 7This is a schematic diagram of the action network training process provided in this application embodiment. A preset action network 320 is invoked to encode the state data of the smart device across multiple time frames and the historical action data 321 corresponding to each time frame, obtaining a first vector. The distance sequences 322 corresponding to multiple time frames are encoded to obtain a second vector, where the distance sequences include the distances between the smart device and each obstacle. Action prediction is performed based on the first and second vectors to obtain the predicted action data for the current time frame. Reward information is determined based on the state data, distance sequence, and predicted action data of the current time frame, and a loss value is determined based on the reward information. The action network is trained based on the loss value to obtain the trained action network.
[0112] In some embodiments, the network training method provided in this application can be applied in the medical field. A preset action network is invoked to encode the state data of a smart device across multiple time frames and the historical action data corresponding to each time frame, resulting in a first vector. The historical action data refers to the action data in the previous time frame, and the multiple time frames include the current time frame. Next, a distance sequence corresponding to each time frame is determined and encoded to obtain a second vector. The distance sequence includes the distances between the smart device and various obstacles. Action prediction is then performed based on the first and second vectors to obtain the predicted action data for the current time frame. This method combines multi-source information such as the state data of the smart device across multiple time frames, historical action data, and the distances between the smart device and various obstacles for fusion encoding, fully utilizing the time-series characteristics of the data and improving the accuracy of action prediction. Then, reward information is determined based on the state data, distance sequence, and predicted action data of the current time frame, and a loss value is determined based on the reward information. The action network is then trained based on the loss value to obtain a trained action network, which is used to control the movement of the smart device. In this way, by adjusting and optimizing the action network using state data, distance sequences, and predicted action data, the action network can capture dynamic changes in the environment and accurately predict the current action. This improves the accuracy of smart device movement in complex dynamic scenarios when using the trained action network to control the movement of smart devices.
[0113] The motion control method provided in the embodiments of this application is described below. See also: Figure 8 , Figure 8 This is a flowchart illustrating the motion control method provided in the embodiments of this application, which will be combined with... Figure 8 The steps shown are explained.
[0114] In step 330, the trained action network is invoked to encode the state data of the smart device in multiple time frames and the historical action data corresponding to each time frame to obtain the first target vector.
[0115] Here, historical motion data refers to the motion data in the previous time frame of a given time frame. Multiple time frames include multiple historical time frames and the current time frame. The state data of multiple time frames and the historical motion data corresponding to each time frame are encoded to obtain the first target vector.
[0116] In step 331, the distance sequence corresponding to each time frame is determined, and the distance sequences corresponding to multiple time frames are encoded to obtain the second target vector.
[0117] Here, within each time frame, the intelligent device is controlled to emit multiple detection rays in a preset direction. The obstacle distances corresponding to the detection rays in each time frame are determined, and the obstacle distances corresponding to each detection ray are combined to obtain a distance sequence for a time frame. The distance sequences corresponding to multiple historical time frames and the current time frame are encoded to obtain the second target vector.
[0118] In step 332, motion prediction is performed based on the first target vector and the second target vector to obtain the predicted motion data for the current time frame, and the movement of the smart device is controlled based on the predicted motion data.
[0119] Here, the motion data of the smart device across multiple historical time frames is encoded to obtain a third target vector. The first, second, and third target vectors are then fused to obtain a target fusion vector. Motion prediction is performed based on this target fusion vector to obtain the predicted motion data for the current time frame. This predicted motion data is then converted into motion control commands to control the movement of the smart device.
[0120] Based on steps 330 to 332, by calling the trained action network, action prediction is performed based on multi-source information such as the state data, historical action data, and distance sequence of the smart device in multiple time frames. This fully utilizes the time series characteristics of the data and improves the motion accuracy of the smart device in complex dynamic scenarios.
[0121] The following will describe an exemplary application of the network training method provided in the embodiments of this application in a humanoid robot navigation scenario.
[0122] Traditional mobile robot navigation typically employs a hierarchical system, generating a reference path through global path planning and then combining it with local trajectory tracking and obstacle avoidance algorithms for execution control. On wheeled platforms, this approach has demonstrated relatively mature performance. However, for humanoid robots, their high-dimensional, underactuated dynamics and strong dependence on stability make it difficult to balance efficiency and safety in complex and dynamic crowd environments. In contrast, gait generation based on Model Predictive Control (MPC) or trajectory optimization can ensure stable walking for humanoid robots. However, directly introducing high-dimensional sensory inputs (such as LiDAR, point clouds, or local maps) during real-time optimization results in enormous computational overhead. Furthermore, when coupling obstacle avoidance and navigation targets, it is prone to over-conservatism or getting trapped in local optima.
[0123] Deep reinforcement learning, within a long-term reward framework, can jointly optimize multiple objectives such as task success rate, safety, and efficiency, thereby enabling humanoid robots to possess greater adaptability. However, end-to-end reinforcement learning schemes still face challenges such as low sample acquisition efficiency, sparse rewards, poor training stability, safety risks during exploration, and significant discrepancies between simulation and reality in perceptual dynamics. Therefore, more practical research approaches in recent years tend to adopt hierarchical or collaborative design ideas, that is, ensuring walking stability at lower levels through reliable motion control priors, and achieving dynamic sub-objective selection, local obstacle avoidance, and decision optimization for compromises between advancing and retreating at higher levels.
[0124] This application proposes a network training method to address the problems existing in related technologies, which includes the following improvements compared to related technologies: By integrating structured multi-frame coding and recurrent memory mechanisms, a joint reward design for navigation and gait, and a hierarchically coupled training system, it can generate high-level speed commands (predicted action data in the above embodiments) with executability, smoothness, and high success rate in dynamic environments. At the same time, it effectively shortens the convergence time of network training and significantly improves the robustness from simulation to actual deployment.
[0125] In this embodiment, a multi-frame observation stacking and a recurrent neural network memory mechanism are introduced into the navigation strategy to systematically capture dynamic obstacles, historical speeds, and short-term planning context. This enables temporal modeling and prediction of non-stationary environments, thereby enhancing the adaptability of humanoid robots in complex dynamic scenarios.
[0126] In this embodiment, when constructing a reward system for the walking characteristics of a humanoid robot, the reward function for navigation tasks and the stability reward function for gait control are jointly modeled to avoid conflicts between high-level instructions and low-level stability objectives. Simultaneously, by imposing constraints on the consistency of forward speed and orientation, undesirable behaviors such as starting backward are effectively mitigated.
[0127] In this embodiment, an implicit representation of the low-level motion control strategy is explicitly introduced into the high-level navigation decision-making process, enabling the high-level layer to fully consider the executability of the low-level layer when generating velocity commands. This mechanism not only improves the executability of velocity commands but also accelerates the convergence process of network training and enhances the overall robustness from simulation to real deployment.
[0128] The network training process is described below. During each environment reset (the robot restarts from...) To begin a new round of exploration, it is necessary to set the start-end distance range, arrival accuracy, and obstacle density for different stages of robot learning, and reset the sampling step size for each iteration, the number of multi-frame observation states, and the multi-frame observation states. Historical speed command (Historical action data in the above embodiments), hidden states of the gated recurrent unit (GRU) of the recurrent neural network And the robot's initial attitude (yaw angle randomization reset).
[0129] Robot body state Depend on It consists of two parts. Represents the linear velocity of the base in the robot's body coordinate system ( ) and yaw rate , This represents the positional error of the self-position relative to the target position in the body coordinate system. and orientation error .
[0130] For robots Azimuth (like Vectorized ray projection is performed to obtain the distance sequence. Example, see reference. Figure 9 , Figure 9 This is a schematic diagram of LiDAR ray measurement of obstacles in a grid map provided in this application embodiment. The grid map includes a robot 401, obstacles, and multiple LiDAR rays, with ray length representing the distance between the robot and the obstacles. .
[0131] In each time frame (interval) Constructing single-frame observation data and update multi-frame stacked data. The historical speed commands and robot body states are encoded, and the encoding results are represented as follows: (The first vector in the above embodiments), where The distance sequence is encoded, and the encoding result is represented as follows: (The second vector in the above embodiments), where .Will and The vectors are concatenated to obtain a multimodal encoded fusion vector. .
[0132] Get hidden variables of the lower-level controller (The third vector in the above embodiments), for example, see reference Figure 10 , Figure 10 This is a schematic diagram of the network structure provided in an embodiment of this application. In the motion state coding layer 501, motion phase signals, joint velocities, overall velocities, gravity, hidden state data, and the latest joint data are concatenated and fused to obtain the hidden variables of the lower-level controller. The hidden variables are then... With fusion vector By concatenating the vectors, the target vector is obtained. (The fusion vector in the above embodiments) enables the higher layer to explicitly consider the low-level executability when generating speed commands, and utilizes the gated recurrent unit of the recurrent neural network to process the target vector. Perform time series modeling, and the result of the time series modeling (the hidden state data in the above embodiments) is represented as follows: This is to depict the dependence on speed trends, obstacle movement trends, and strategy history.
[0133] The temporal modeling results are mapped using a pre-defined actor network, and the output probability distribution is expressed as follows: The mean of the probability distribution is variance and Related. Action sampling is performed based on this probability distribution, and the action sampling is represented as... (mean is) Variance is ). For the output Affine transformation and clipping (replacing saturated activation to avoid boundary gradient collapse) are performed to obtain the predicted action data. For example, the predicted action data is determined using the following formula (1):
[0134] in, This represents the predicted action data. These represent the minimum and maximum motion constraint values, respectively. express , express .
[0135] exist[ During the high-level cycle, the low-level controller cycles. (Usually 0.01s or 0.005s) Execute low-level control, track motion data, and update the robot's body state in a closed loop.
[0136] A pre-defined critic network is used to map the time series modeling results, outputting predicted critic values, which are represented as follows: The reward function (the reward information in the above embodiments) includes navigation reward information. and sports reward information The reward function is expressed as Navigation reward information includes target proximity reward value, obstacle distance value, consistency reward value, backoff penalty value, and smoothness reward value.
[0137] For example, the target proximity reward value is determined by the following formula (2):
[0138] in, This indicates that the target is close to the reward value. This represents the Euclidean distance corresponding to the position error. This indicates the preset adjustment coefficient.
[0139] For example, the obstacle distance value is determined by the following formula (3):
[0140] in, Indicates the obstacle distance value. This represents the distance sequence of the current time frame. This indicates the preset reference distance.
[0141] For example, the consistency reward value is determined by the following formula (4):
[0142] in, This represents the consistency reward value. This represents the velocity vector corresponding to the volume coordinate velocity. This represents the vector corresponding to the position error. Indicates the clipping threshold. (This indicates that the robot is oriented towards a line that fits as closely as possible to the target position.)
[0143] For example, the backoff penalty value is determined by the following formula (5):
[0144] in, Indicates the penalty value for going back ( This indicates moving backward; in this case, the entire expression is negative, equivalent to a penalty.
[0145] For example, the smoothness reward value is determined by the following formula (6):
[0146] in, Represents the smoothness reward value. The Euclidean norm representing the change in volumetric velocity. Represents the absolute value of the change in yaw rate, motion smoothing difference. (The smoothness reward value indicates that the network output action for each action should not differ too much from the previous output action.)
[0147] The motion reward information includes posture and height, energy consumption and joint quality, and contact gait consistency (posture reward value, joint reward value, and ground contact reward value in the above embodiments). Posture and height include angular velocity penalties. The angle between the robot and the vertical direction (such as gravity). Base height deviation ( This indicates the height of a preset point on the robot's waist. Energy consumption and joint quality include penalty joint torque. acceleration Exceeding limits and frequent starts and stops are examples of issues. Contact gait consistency is based on the contact phase and gait template, penalizing inconsistent ground contact timing to encourage stable foot timing.
[0148] The reward function can be designed as follows: This allows the navigation item reward to contribute slightly more than the gait motion item reward, while maintaining the same order of magnitude, thus avoiding the robot's propulsion suppressing stability.
[0149] The speed command is output through the action network. For example, the loss value of the action network is determined by the following formula (7):
[0150] in, This represents the loss value of the action network. Indicates the advantage estimate , and These are the set parameters, which can be set to 0.99 and 0.95 respectively. Represents the reward function, It evaluates the network's output value.
[0151] The parameters of the evaluation network are updated using the following formula (8):
[0152] in, This represents the loss value used to evaluate the network.
[0153] If the robot reaches its destination, collides with an obstacle, or times out during the current learning phase, the process terminates, and the start-end point distance range, arrival accuracy, obstacle density, etc., are reset. The difficulty (start-end point distance, obstacle density) is increased stage by stage, checkpoints are saved periodically, and a scripted model is exported.
[0154] In the aforementioned humanoid robot navigation scenario, the introduction of multi-frame temporal perception and cyclic memory modeling effectively captures non-stationary environmental features, enhancing the prediction and avoidance capabilities of dynamic obstacles. This reduces collision risks and improves the success rate of navigation tasks. By constructing a joint reward system for navigation and motion control, higher-level decisions can explicitly perceive the executability of lower-level systems, avoiding conflicts between speed commands and motion stability. This significantly improves the smoothness of walking and control stability, achieving the goal of stable task execution in complex scenarios. The adoption of structured coding and parallel training strategies reduces reliance on online optimization, lowers computational overhead during system deployment, and improves real-time performance and system integration efficiency, thus meeting the needs of practical engineering applications.
[0155] The following description continues to illustrate the exemplary structure of the network training device 455 provided in the embodiments of this application as a software module. In some embodiments, such as... Figure 2 As shown, the software modules stored in the network training device 455 in the memory 450 may include: a first encoding module 4551, used to call a preset action network to encode the state data of the smart device in multiple time frames and the historical action data corresponding to each time frame to obtain a first vector, wherein the historical action data is the action data in the previous time frame, and the multiple time frames include the current time frame; a second encoding module 4552, used to determine the distance sequence corresponding to each time frame and encode the distance sequence corresponding to multiple time frames to obtain a second vector, wherein the distance sequence includes the distance between the smart device and each obstacle; an action prediction module 4553, used to perform action prediction based on the first vector and the second vector to obtain the predicted action data of the current time frame; a loss determination module 4554, used to determine reward information based on the state data, distance sequence and predicted action data of the current time frame, and determine the loss value based on the reward information; and a network training module 4555, used to train the action network based on the loss value to obtain the trained action network, so as to control the movement of the smart device using the trained action network.
[0156] In some embodiments, the second encoding module 4552 is further configured to control the smart device to emit multiple detection rays in a preset direction within each time frame; determine the obstacle distance corresponding to each detection ray; and combine the obstacle distances corresponding to multiple detection rays to obtain a distance sequence corresponding to the time frame.
[0157] In some embodiments, the multiple time frames also include historical time frames. The action prediction module 4553 is further configured to acquire action data of the smart device in the historical time frames, encode the action data to obtain a third vector, fuse the first vector, the second vector and the third vector to obtain a fused vector, perform temporal prediction based on the fused vector to obtain latent state data, and perform action prediction based on the latent state data to obtain the predicted action data of the current time frame.
[0158] In some embodiments, the action prediction module 4553 is further configured to determine the probability distribution corresponding to the hidden state data, and sample the hidden state data based on the probability distribution to obtain first action data; perform boundary constraint processing on the first action data based on preset action constraint data to obtain second action data; and perform cropping processing on the second action data based on the action constraint data to obtain predicted action data for the current time frame.
[0159] In some embodiments, the motion constraint data includes a minimum motion constraint value and a maximum motion constraint value. The motion prediction module 4553 is further configured to determine a first boundary constraint value based on the difference between the minimum and maximum motion constraint values; determine a second boundary constraint value based on the sum of the minimum and maximum motion constraint values; and perform boundary constraint processing on the first motion data based on the first and second boundary constraint values to obtain the second motion data.
[0160] In some embodiments, the loss determination module 4554 is further configured to determine navigation reward information based on the state data, distance sequence, and predicted action data of the current time frame; determine motion reward information based on the state data and predicted action data of the current time frame; and perform superposition processing on the navigation reward information and motion reward information to obtain reward information.
[0161] In some embodiments, the state data includes the position error between the current position and the target position of the smart device, the lateral velocity of the smart device, and the longitudinal velocity of the smart device. The loss determination module 4554 is further configured to perform nonlinear processing on the position error of the current time frame to obtain the target approach reward value of the current time frame; normalize the distance sequence of the current time frame to obtain the obstacle distance value of the current time frame; determine the orientation indication information of the current time frame, and perform consistency constraint processing on the lateral velocity, longitudinal velocity, and position error based on the orientation indication information to obtain the consistency reward value of the current time frame; determine the backtracking indication information of the current time frame, and perform exponential transformation processing on the lateral velocity based on the backtracking indication information to obtain the backtracking penalty value of the current time frame; perform motion smoothing analysis based on the predicted motion data of the current time frame and the motion data of the previous time frame to obtain the smoothness reward value of the current time frame; and combine the target approach reward value, obstacle distance value, consistency reward value, backtracking penalty value, and smoothness reward value to obtain navigation reward information.
[0162] In some embodiments, the state data further includes the posture data of the smart device, and the predicted motion data includes the predicted joint data and the predicted ground contact time series data of the smart device. The loss determination module 4554 is further configured to adjust the difference between the posture data of the current time frame and the preset expected posture data based on a first preset coefficient to obtain the posture reward value of the current time frame; adjust the difference between the predicted joint data and the preset expected joint data based on a second preset coefficient to obtain the joint reward value of the current time frame; adjust the difference between the predicted ground contact time series data and the preset expected ground contact time series data based on a third preset coefficient to obtain the ground contact reward value of the current time frame; and combine the posture reward value, joint reward value and ground contact reward value of the current time frame to obtain motion reward information.
[0163] In some embodiments, the loss determination module 4554 is further configured to invoke a preset evaluation network to evaluate the hidden state data and obtain the predicted evaluation value of the current time frame; determine the evaluation deviation value based on the reward information, the predicted evaluation value of the current time frame and the predicted evaluation value of the next time frame; perform a preset dominance function estimation based on the evaluation deviation value to obtain the dominance estimate value, and construct the loss value based on the dominance estimate value and a preset threshold.
[0164] This application provides a computer program product, which includes computer-executable instructions or a computer program stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions or computer program from the computer-readable storage medium and executes the computer-executable instructions or computer program, causing the electronic device to perform the network training method provided in this application.
[0165] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the network training method provided in this application. For example, ... Figure 3 The network training method is shown.
[0166] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0167] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0168] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0169] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0170] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A network training method, characterized in that, The method includes: A preset action network is invoked to encode the state data of the smart device in multiple time frames and the historical action data corresponding to each time frame to obtain a first vector. The historical action data is the action data in the previous time frame of the current time frame. Determine the distance sequence corresponding to each time frame, and encode the distance sequences corresponding to multiple time frames to obtain a second vector, wherein the distance sequence includes the distance between the smart device and each obstacle; Based on the first vector and the second vector, action prediction is performed to obtain the predicted action data for the current time frame; Reward information is determined based on the state data of the current time frame, the distance sequence, and the predicted action data, and the loss value is determined based on the reward information; The action network is trained based on the loss value to obtain a trained action network, which is then used to control the movement of the smart device.
2. The method according to claim 1, characterized in that, The multiple time frames also include historical time frames. The step of performing action prediction based on the first vector and the second vector to obtain the predicted action data for the current time frame includes: The action data of the smart device in the historical time frame is obtained, and the action data is encoded to obtain a third vector; The first vector, the second vector, and the third vector are fused together to obtain a fused vector. Temporal prediction is performed based on the fusion vector to obtain hidden state data; Action prediction is performed based on the hidden state data to obtain the predicted action data for the current time frame.
3. The method according to claim 2, characterized in that, The step of performing action prediction on the hidden state data to obtain the predicted action data for the current time frame includes: Determine the probability distribution corresponding to the hidden state data, and sample the hidden state data based on the probability distribution to obtain the first action data; Based on preset motion constraint data, the first motion data is subjected to boundary constraint processing to obtain the second motion data; The second action data is cropped based on the action constraint data to obtain the predicted action data for the current time frame.
4. The method according to claim 3, characterized in that, The motion constraint data includes a minimum motion constraint value and a maximum motion constraint value. The second motion data is obtained by performing boundary constraint processing on the first motion data based on the preset motion constraint data, including: The first boundary constraint value is determined based on the difference between the minimum action constraint value and the maximum action constraint value; The second boundary constraint value is determined based on the sum of the minimum action constraint value and the maximum action constraint value; Based on the first boundary constraint value and the second boundary constraint value, the first action data is subjected to boundary constraint processing to obtain the second action data.
5. The method according to claim 1, characterized in that, The determination of reward information based on the state data of the current time frame, the distance sequence, and the predicted action data includes: Navigation reward information is determined based on the status data of the current time frame, the distance sequence, and the predicted action data. Motion reward information is determined based on the status data of the current time frame and the predicted action data; The navigation reward information and the motion reward information are superimposed to obtain the reward information.
6. The method according to claim 5, characterized in that, The status data includes the positional error between the current position and the target position of the smart device, the lateral velocity of the smart device, and the longitudinal velocity of the smart device. The determination of navigation reward information based on the state data of the current time frame, the distance sequence, and the predicted action data includes: The position error of the current time frame is processed nonlinearly to obtain the target proximity reward value of the current time frame. The distance sequence of the current time frame is normalized to obtain the obstacle distance value of the current time frame; Determine the orientation indication information of the current time frame, and based on the orientation indication information, perform consistency constraint processing on the lateral velocity, the longitudinal velocity, and the position error to obtain the consistency reward value of the current time frame; Determine the backward indication information of the current time frame, and based on the backward indication information, perform exponential transformation processing on the lateral velocity to obtain the backward penalty value of the current time frame; Based on the predicted motion data of the current time frame and the motion data of the previous time frame, motion smoothing analysis is performed to obtain the smoothness reward value of the current time frame. The navigation reward information is obtained by combining the target proximity reward value, the obstacle distance value, the consistency reward value, the backtracking penalty value, and the smoothness reward value.
7. The method according to claim 5, characterized in that, The state data also includes the posture data of the smart device, and the predicted motion data includes the predicted joint data and the predicted ground contact time sequence data of the smart device. Determining motion reward information based on the state data of the current time frame and the predicted motion data includes: Based on a first preset coefficient, the difference between the attitude data of the current time frame and the preset desired attitude data is adjusted to obtain the attitude reward value of the current time frame. Based on the second preset coefficient, the difference between the predicted joint data and the preset expected joint data is adjusted to obtain the joint reward value of the current time frame; Based on a third preset coefficient, the difference between the predicted ground contact time series data and the preset expected ground contact time series data is adjusted to obtain the ground contact reward value of the current time frame; The motion reward information is obtained by combining the posture reward value, the joint reward value, and the ground contact reward value of the current time frame.
8. The method according to claim 2, characterized in that, Determining the loss value based on the reward information includes: A preset evaluation network is invoked to evaluate the hidden state data and obtain the predicted evaluation value of the current time frame. Based on the reward information, the predicted evaluation value of the current time frame, and the predicted evaluation value of the next time frame, the evaluation deviation value is determined. Based on the evaluation deviation value, a preset dominance function is estimated to obtain a dominance estimate, and a loss value is constructed based on the dominance estimate and a preset threshold.
9. The method according to any one of claims 1 to 8, characterized in that, Determining the distance sequence corresponding to each time frame includes: Within each time frame, the intelligent device is controlled to emit multiple detection rays in a preset direction; For each of the detection rays, determine the distance to the obstacle corresponding to that detection ray; The distances to obstacles corresponding to each of the detection rays are combined to obtain the distance sequence corresponding to the time frame.
10. A network training device, characterized in that, The device includes: The first encoding module is used to call a preset action network to encode the state data of the smart device in multiple time frames and the historical action data corresponding to each time frame to obtain a first vector. The historical action data is the action data in the previous time frame of the time frame, and the multiple time frames include the current time frame. The second encoding module is used to determine the distance sequence corresponding to each time frame and to encode the distance sequences corresponding to multiple time frames to obtain a second vector, wherein the distance sequence includes the distance between the smart device and each obstacle; The action prediction module is used to perform action prediction based on the first vector and the second vector to obtain the predicted action data of the current time frame; The loss determination module is used to determine reward information based on the state data, distance sequence, and predicted action data of the current time frame, and to determine the loss value based on the reward information; The network training module is used to train the action network based on the loss value to obtain the trained action network, so as to control the movement of the smart device using the trained action network.
11. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. A processor, configured to execute computer-executable instructions or computer programs stored in the memory, implements the network training method according to any one of claims 1 to 9.
12. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the network training method according to any one of claims 1 to 9.