Training method of humanoid robot energy consumption prediction model and energy optimization method
By performing feature fusion and matrix extraction on the static parameters and load data of the humanoid robot's motor, an energy consumption prediction model is constructed, which solves the problem of large prediction error of motor energy consumption in cross-scenario scenarios and realizes precise control of motor energy consumption and extension of battery life.
Patent Information
- Application Number
- CN202511530279.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies struggle to accurately control the motor energy consumption of humanoid robots across different scenarios, resulting in significant prediction errors and failing to meet the demand for precise control of motor energy consumption.
By acquiring static parameters of the humanoid robot's motor and load sample data from different scenario types, feature fusion and matrix feature extraction are performed to construct a humanoid robot energy consumption prediction model. The LSTM-CNN structure is used to extract temporal and multidimensional difference features, and the electric drive parameters are dynamically adjusted to optimize energy use.
It reduces the error in cross-scenario energy consumption prediction, improves the precise control capability of motor energy consumption, and extends the working endurance of humanoid robots.
Smart Images

Figure CN121502335A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent robot technology, and in particular to a training method and energy optimization method for a humanoid robot energy consumption prediction model. Background Technology
[0002] Humanoid robots, as a new generation of intelligent equipment, have been gradually applied in industrial production and household services. For example, in industrial settings, humanoid robots need to complete high-load, long-duration handling and assembly tasks (such as assembling automotive parts); while in household settings, they need to perform low-load, high-frequency cleaning and delivery tasks (such as sweeping and retrieving items). As the core power unit of humanoid robots, the energy consumption level of the electric drive system directly determines the robot's endurance and operating efficiency in various scenarios. Therefore, precise control of electric drive energy consumption has become a key focus for relevant personnel.
[0003] Currently, the relevant technologies typically involve using specialized equipment (such as power analyzers) to collect measured motor energy consumption data of humanoid robots in specific scenarios, and then constructing an energy consumption prediction model based on the collected measured motor energy consumption data to achieve energy consumption prediction. This method can only predict the energy consumption of humanoid robots in a single scenario, and its prediction error is relatively large when predicting across scenarios, making it difficult to meet the subsequent requirements for precise control of motor energy consumption.
[0004] Therefore, the problems with the relevant technologies still need to be solved and optimized. Summary of the Invention
[0005] The purpose of this invention is to at least partially solve one of the technical problems existing in the related art.
[0006] Therefore, one objective of this invention is to provide a training method and an energy optimization method for a humanoid robot energy consumption prediction model. The training method can provide a humanoid robot energy consumption prediction model that helps reduce the prediction error of humanoid robot motor energy consumption when predicting across scenarios, thereby effectively meeting the precise control requirements of subsequent motor energy consumption assessment. To achieve the above-mentioned technical objectives, the technical solutions adopted in the embodiments of this application include: In a first aspect, embodiments of this application provide a method for training a humanoid robot energy consumption prediction model, including: Acquire static parameter sample data of the humanoid robot motor and scene load sample data of several different scene types, as well as measured energy consumption sample data corresponding to each scene load sample data; Based on the static parameter sample data, feature fusion is performed on all the scene load sample data to obtain a fused feature matrix; Matrix feature extraction is performed on the fused feature matrix to obtain feature difference data. The feature difference data is used to characterize the differences in peak feature amplitude, temporal density, and energy distribution of scene load sample data under different scene types. Based on the feature difference data and all the measured energy consumption sample data, the parameters of the initialized humanoid robot energy consumption prediction model are updated to obtain a trained humanoid robot energy consumption prediction model.
[0007] In addition, the method according to the above embodiments of this application may also have the following additional technical features: Furthermore, in one embodiment of this application, the step of performing feature fusion on all the scene load sample data based on the static parameter sample data to obtain a fused feature matrix includes: The static parameter sample data is statically vectorized to obtain static feature vectors; Dynamic feature vectorization is performed on all the scene load sample data to obtain several dynamic feature vectors, each of which corresponds to a combination of the scene load sample data and the scene type; Based on the static feature vector, all the dynamic feature vectors are concatenated to obtain the fused feature matrix.
[0008] Furthermore, in one embodiment of this application, the step of extracting matrix features from the fused feature matrix to obtain feature difference data includes: Temporal feature extraction is performed on the fused feature matrix to obtain temporal dependency data, which records the temporal dependency relationship of each scenario load sample data. Multidimensional difference features are extracted from the time-series dependent data to obtain the feature difference data.
[0009] Furthermore, in one embodiment of this application, the step of extracting multidimensional difference features from the time-series dependent data to obtain the feature difference data includes: The time-dependent data is subjected to the first convolution dimension extraction to obtain the first feature map, and the first feature map records the peak feature amplitude difference of the scene load sample data under different scene types. The second convolution dimension is extracted from the first feature map to obtain the second feature map. The second feature map records the temporal density difference of scene load sample data under different scene types and the peak feature amplitude difference of scene load sample data under different scene types. The second feature map is subjected to pooling dimension extraction to obtain the feature difference data.
[0010] Further, in one embodiment of this application, the step of updating the parameters of the initialized humanoid robot energy consumption prediction model based on the feature difference data and all the measured energy consumption sample data to obtain a trained humanoid robot energy consumption prediction model includes: Energy consumption prediction is performed on the aforementioned feature difference data to obtain energy consumption prediction data; Based on the measured energy consumption sample data, the energy consumption prediction data is analyzed for prediction error to obtain prediction error information; Based on the prediction error information, the parameters of the initialized humanoid robot energy consumption prediction model are updated to obtain the trained humanoid robot energy consumption prediction model.
[0011] Furthermore, in one embodiment of this application, the prediction error information is functionally represented as:
[0012] in, This represents prediction error information; n is the total number of samples. For the first Scene weights for each sample; For the first One sample of measured energy consumption data; In the energy consumption prediction data, the first Predicted energy consumption for each sample.
[0013] Secondly, embodiments of this application provide an energy optimization method, including: Obtain the static parameters, scene information, and real-time load information of the humanoid robot's motor at the current time point, as well as the scene energy consumption threshold corresponding to the scene information; The static parameters, the scene information, and the real-time load information are input into the humanoid robot energy consumption prediction model described above to predict energy consumption and obtain the real-time energy consumption prediction value. Based on the scenario energy consumption threshold and the real-time energy consumption prediction value, the energy of the humanoid robot motor is optimized.
[0014] Thirdly, embodiments of this application provide a training system for a humanoid robot energy consumption prediction model, comprising: The first processing unit is used to acquire static parameter sample data of the humanoid robot motor, scene load sample data of several different scene types, and measured energy consumption sample data corresponding to each scene load sample data. The second processing unit is used to perform feature fusion on all the scene load sample data based on the static parameter sample data to obtain a fused feature matrix; The third processing unit is used to extract matrix features from the fused feature matrix to obtain feature difference data. The feature difference data is used to characterize the peak feature amplitude difference, temporal density difference, and energy distribution difference of scene load sample data under different scene types. The fourth processing unit is used to update the parameters of the initialized humanoid robot energy consumption prediction model based on the feature difference data and all the measured energy consumption sample data, so as to obtain the trained humanoid robot energy consumption prediction model.
[0015] Fourthly, embodiments of this application also provide an electronic device, including: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method described above.
[0016] Fifthly, embodiments of this application also provide a computer-readable storage medium storing a processor-executable program, which, when executed by the processor, is used to implement the above-described method.
[0017] The advantages and beneficial effects of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application: This application discloses a training method and energy optimization method for a humanoid robot energy consumption prediction model. The training method acquires static parameter sample data of the humanoid robot's motor, scene load sample data for several different scene types, and measured energy consumption sample data corresponding to each scene load sample data. Based on the static parameter sample data, feature fusion is performed on all scene load sample data to obtain a fused feature matrix. Matrix feature extraction is performed on the fused feature matrix to obtain feature difference data, which characterizes the differences in peak feature amplitude, temporal density, and energy distribution of scene load sample data under different scene types. Based on the feature difference data and all measured energy consumption sample data, the parameters of the initialized humanoid robot energy consumption prediction model are updated to obtain a trained humanoid robot energy consumption prediction model. This training method, by performing feature fusion and matrix feature extraction on scene load sample data of multiple scene types, enables the subsequently trained energy consumption prediction model to not only predict motor energy consumption across scenes but also reduces the prediction error of the energy consumption prediction model during cross-scene prediction, effectively meeting the subsequent requirements for precise control of motor energy consumption. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following description is provided with accompanying drawings of the relevant technical solutions in the embodiments of this application or the prior art. It should be understood that the accompanying drawings described below are only for the purpose of clearly illustrating some embodiments of the technical solutions in this application. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0019] Figure 1 A flowchart illustrating a training method for a humanoid robot energy consumption prediction model provided in an embodiment of this application; Figure 2 A schematic diagram of the framework of a training system for a humanoid robot energy consumption prediction model provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0020] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application. The step numbers in the following embodiments are set only for ease of explanation, and there is no limitation on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0022] Currently, relevant technologies typically involve using specialized equipment (such as power analyzers) to collect measured motor energy consumption data of humanoid robots in specific scenarios, and then constructing an energy consumption prediction model based on this data. This approach can only predict the energy consumption of a humanoid robot in a single scenario, failing to consider the differences in load characteristics across different scenarios (e.g., high torque, low frequency in industrial scenarios; low torque, high frequency in household scenarios). This results in single-scenario energy consumption prediction models generally exhibiting double-digit prediction errors across different scenarios (i.e., large prediction errors when predicting across scenarios), making it difficult to meet the subsequent requirements for precise control of motor energy consumption.
[0023] Furthermore, some related technologies train several network models for specific scenario types based on scenario load data for each scenario type, and then fuse these network models from all different scenario types to obtain the final cross-scenario energy consumption prediction model. This approach requires building a separate test environment for each scenario and acquiring energy consumption data from different scenario tasks, which is labor-intensive and requires a large amount of data.
[0024] Furthermore, some related technologies have electric drive parameters (such as PWM duty cycle, working mode, etc.) of humanoid robots when working in various scenarios that are mostly preset fixed values. These parameters cannot be dynamically adjusted according to real-time energy consumption prediction results, resulting in energy consumption redundancy often exceeding 25%, which significantly shortens the working endurance of humanoid robots.
[0025] It should be noted that the aforementioned related technologies are only used to assist in understanding the technical solutions of this application and do not mean that they belong to the publicly disclosed prior art.
[0026] In view of this, embodiments of this application provide a training method and an energy optimization method for a humanoid robot energy consumption prediction model. The training method performs feature fusion and matrix feature extraction on scene load sample data of multiple scene types. Specifically, it extracts temporal features and multi-dimensional difference features from the fused feature matrix. This allows the model to consider the differences in load characteristics of the humanoid robot motor under different scenes, thereby enabling the energy consumption prediction model obtained after training to not only predict motor energy consumption across scenes, but also reduce the prediction error of the energy consumption prediction model when predicting across scenes, effectively meeting the subsequent requirements for precise control of motor energy consumption.
[0027] Furthermore, this training method is based on the extraction of multidimensional differential features from time-dependent data. Specifically, through load pattern differentiation differences (amplitude differences, temporal density differences, and energy distribution differences) of different scenario types, it enables the energy consumption prediction model to assign "differentiated weights" to different scenarios (e.g., industrial scenarios emphasize high-power features, while household scenarios emphasize high-frequency features). This not only avoids the situation where the model uses the same set of feature weights to predict all scenarios (e.g., using the high torque feature weights of industrial scenarios to predict household scenarios), reducing cross-modal energy consumption prediction errors, but also reduces the amount of data required for model training, thus reducing workload.
[0028] Furthermore, this optimization method optimizes the energy of the humanoid robot's motor based on the scene's energy consumption threshold and real-time energy consumption prediction. It can dynamically adjust the electric drive parameters of the humanoid robot when it is working in the scene according to the real-time energy consumption prediction results, which is beneficial to improving the working endurance of the humanoid robot.
[0029] Reference Figure 1 In this embodiment of the application, a training method for a humanoid robot energy consumption prediction model includes: Step 110: Obtain static parameter sample data of the humanoid robot motor and scene load sample data of several different scene types, as well as measured energy consumption sample data corresponding to each scene load sample data. In this embodiment, the motor can be a drive motor in the humanoid robot's electric drive system. Static parameter sample data can include the motor's rated voltage, stator resistance, stator inductance, rated power, etc. Scene load sample data can include the motor's load torque (i.e., the torque required for the motor to drive the load), motor speed, and task duration (i.e., the duration of the load application) under the corresponding scene type (e.g., industrial handling scene, industrial assembly scene, household cleaning scene, household delivery scene, etc.) when the humanoid robot is working in each scene type. Measured energy consumption sample data can be the measured energy consumption collected when the humanoid robot is working in each scene type.
[0030] Step 120: Based on the static parameter sample data, perform feature fusion on all the scene load sample data to obtain a fused feature matrix; In this embodiment of the application, feature fusion can be achieved by fusing the sample features of static parameter sample data with the sample features of scene load sample data under each scene type to obtain a fused feature matrix.
[0031] In some embodiments, the step of performing feature fusion on all the scene load sample data based on the static parameter sample data to obtain a fused feature matrix includes: The static parameter sample data is statically vectorized to obtain static feature vectors; Dynamic feature vectorization is performed on all the scene load sample data to obtain several dynamic feature vectors, each of which corresponds to a combination of the scene load sample data and the scene type; Based on the static feature vector, all the dynamic feature vectors are concatenated to obtain the fused feature matrix.
[0032] In this embodiment, static feature vectorization can be achieved by standardizing static parameter sample data and then converting it into vector form to obtain static feature vectors. These static feature vectors can be vector forms of static features such as rated voltage, stator resistance, stator inductance, and rated power.
[0033] Understandably, for any given scenario load sample data, dynamic feature vectorization can be achieved by converting features such as load torque, motor speed, and task duration in the scenario load sample data into vector form, and then adding corresponding scenario codes (e.g., scenario code = 1 for industrial handling, scenario code = 2 for industrial assembly, scenario code = 3 for household cleaning, and scenario code = 4 for household delivery), thereby obtaining a dynamic feature vector corresponding to that scenario load sample data. Feature concatenation can be performed by concatenating (Conca) the static and dynamic feature vectors to obtain a matrix-form fused feature, denoted as the fused feature matrix.
[0034] Step 130: Extract matrix features from the fused feature matrix to obtain feature difference data. The feature difference data is used to characterize the differences in peak feature amplitude, temporal density, and energy distribution of scene load sample data under different scene types. In this embodiment of the application, matrix feature extraction can extract feature difference data of scene load sample data under different scene types in the fused feature matrix. The feature difference data includes peak feature amplitude difference, temporal density difference, and energy distribution difference.
[0035] In some embodiments, the step of extracting matrix features from the fused feature matrix to obtain feature difference data includes: Temporal feature extraction is performed on the fused feature matrix to obtain temporal dependency data, which records the temporal dependency relationship of each scenario load sample data. In this embodiment, temporal feature extraction can be achieved by inputting the fused feature matrix into the temporal extraction layer of the humanoid robot energy consumption prediction model. The temporal extraction layer is constructed based on a long short-term memory (LSTM) network, which may specifically include 128 hidden units and the activation function is the tanh function. The temporal extraction layer extracts the temporal dependency relationship of each scene load sample data in the fused feature matrix, and the temporal dependency data output by the temporal extraction layer is used as the temporal dependency data.
[0036] Multidimensional difference features are extracted from the time-series dependent data to obtain the feature difference data.
[0037] Further, the step of extracting multidimensional difference features from the time-series dependent data to obtain the feature difference data includes: The time-dependent data is subjected to the first convolution dimension extraction to obtain the first feature map, and the first feature map records the peak feature amplitude difference of the scene load sample data under different scene types. The second convolution dimension is extracted from the first feature map to obtain the second feature map. The second feature map records the temporal density difference of scene load sample data under different scene types and the peak feature amplitude difference of scene load sample data under different scene types. The second feature map is subjected to pooling dimension extraction to obtain the feature difference data.
[0038] In this embodiment, multidimensional difference feature extraction can be achieved by inputting the time-dependent data output by the time-series extraction layer into the scene feature layer of the humanoid robot energy consumption prediction model. The scene feature layer can be constructed based on a convolutional neural network (CNN). The CNN network is configured with a 3×3 small convolutional kernel, a 5×5 large convolutional kernel, and a 2×2 max pooling layer. The scene feature layer extracts the load feature differences of different scenes and outputs the feature difference data.
[0039] Understandably, the first convolutional dimension extraction can be achieved by inputting time-dependent data into a 3×3 small convolutional kernel. This 3×3 small convolutional kernel is used to capture local details and can accurately capture local abrupt changes in load data such as "torque peak, speed peak" (e.g., instantaneous high torque during industrial handling). This results in a first feature map that records the difference in peak feature amplitude. This difference in peak feature amplitude refers to the inherent difference between the "peak size" and "mean level" of the load sample data (load torque T and motor speed n) under different scenarios. One example of this difference in peak feature amplitude is shown in Table 1 below. Table 1
[0040] The second convolutional dimension extraction can be achieved by inputting the first feature map into a 5×5 large convolutional kernel. This 5×5 kernel is used to capture global temporal correlations. Building upon the "local peaks" extracted by the 3×3 convolution, it further captures global temporal density features such as the "frequency of peak occurrence and interval duration" (e.g., high-frequency peaks occurring every 3 seconds in household cleaning, and low-frequency peaks occurring every 20 seconds in industrial assembly). This results in a second feature map that records temporal density differences and peak feature amplitude differences. These temporal density differences refer to the differences in the "frequency" and "interval" of scene load sample data fluctuations over time in different scenarios. One example of this temporal density difference is shown in Table 2 below. Table 2
[0041] Pooling dimension extraction can be achieved by inputting the second feature map into a 2×2 max pooling layer, which then filters out the regions with the largest power contributions to obtain feature difference data. Here, the energy distribution difference refers to the difference in the "power distribution interval" corresponding to the scene load sample data under different scenarios. One example of this energy distribution difference is shown in Table 3 below. Table 3
[0042] It should be noted that, due to the different weights of the impact of varying load characteristics on energy consumption (e.g., "high-amplitude torque" accounts for 60% of the impact in industrial scenarios, while "high-frequency fluctuations" account for 50% in household scenarios), this embodiment uses a CNN network for multi-dimensional differential feature extraction. This allows the model to assign "differentiated weights" to different scenarios (e.g., emphasizing high-power features in industrial scenarios and high-frequency features in household scenarios). Furthermore, in practical applications, scenario load sample data often contains interference noise (e.g., small torque fluctuations caused by uneven ground). This embodiment uses a 2×2 max-pooling layer in the CNN layer to filter out the "core features with the largest energy contribution" (e.g., removing features <0.1N). The noise torque (m) makes subsequent features more accurate and further reduces prediction error. Specifically, experimental verification shows that the error can be further reduced by 1.2%-1.5% after removing noise.
[0043] Step 140: Based on the feature difference data and all the measured energy consumption sample data, update the parameters of the initialized humanoid robot energy consumption prediction model to obtain the trained humanoid robot energy consumption prediction model.
[0044] In this embodiment, parameter updates can be performed by combining a loss function and determining the target loss value corresponding to the feature difference data based on all measured energy consumption sample data. There are many commonly used loss functions, such as 0-1 loss function, squared loss function, absolute loss function, logarithmic loss function, and cross-entropy loss function, which can all be used as model loss functions and will not be elaborated upon here. In this embodiment, any one of these loss functions can be selected to determine the training loss value, such as the cross-entropy loss function. Based on the training target loss value, the parameters of the model are updated using a backpropagation algorithm and / or an optimization algorithm. After several iterations, a well-trained humanoid robot energy consumption prediction model can be obtained. The specific number of iterations can be preset, or training can be considered complete when the test set reaches the required accuracy.
[0045] In some embodiments, updating the parameters of the initialized humanoid robot energy consumption prediction model based on the feature difference data and all the measured energy consumption sample data to obtain a trained humanoid robot energy consumption prediction model includes: Energy consumption prediction is performed on the aforementioned feature difference data to obtain energy consumption prediction data; Based on the measured energy consumption sample data, the energy consumption prediction data is analyzed for prediction error to obtain prediction error information; Based on the prediction error information, the parameters of the initialized humanoid robot energy consumption prediction model are updated to obtain the trained humanoid robot energy consumption prediction model.
[0046] In this embodiment, energy consumption prediction can involve inputting feature difference data into a fully connected layer in a humanoid robot energy consumption prediction model, and outputting energy consumption prediction data through a fully connected layer containing several neurons. Prediction error analysis can be based on a prediction error function to determine the prediction error information between the measured energy consumption sample data and the predicted energy consumption data. The function representation of this prediction error information can be:
[0047] in, This represents prediction error information; n is the total number of samples. For the first The scene weight of a sample, which can be scene load sample data and / or static parameter sample data; For the first One sample of measured energy consumption data; In the energy consumption prediction data, the first Predicted energy consumption for each sample.
[0048] Understandably, parameter updates can be implemented by using the Adam optimization algorithm to adjust the forget gate weights of the LSTM layer, the convolution kernel parameters of the CNN layer, and the parameters of subsequent fully connected layers in the humanoid robot energy consumption prediction model when the prediction error information is greater than the prediction error threshold (e.g., 5%); or by stopping the iterative training of the model when the prediction error information is less than the prediction error threshold, and determining the current humanoid robot energy consumption prediction model as the well-trained humanoid robot energy consumption prediction model.
[0049] In this embodiment of the application, an energy optimization method includes: Obtain the static parameters, scene information, and real-time load information of the humanoid robot's motor at the current time point, as well as the scene energy consumption threshold corresponding to the scene information; The static parameters, the scene information, and the real-time load information are input into the humanoid robot energy consumption prediction model described above to predict energy consumption and obtain the real-time energy consumption prediction value. Based on the scenario energy consumption threshold and the real-time energy consumption prediction value, the energy of the humanoid robot motor is optimized.
[0050] In this embodiment of the application, in practical applications, the static parameters of the motor, scene information (such as industrial assembly scene or home delivery scene) and real-time load information can be obtained in real time through the sensors of the humanoid robot (such as voltage sensor, torque sensor, encoder, etc.); then the collected static parameters, scene information and real-time load information are input into the humanoid robot energy consumption prediction model, and the energy consumption value of the motor is predicted by the humanoid robot energy consumption prediction model to obtain the real-time energy consumption prediction value.
[0051] Understandably, energy optimization can begin by dynamically adjusting the electric drive system parameters based on a comparison between real-time energy consumption predictions and scenario energy consumption thresholds. Here, the scenario energy consumption threshold is calculated as (battery capacity × 0.8) / target runtime × scenario power coefficient. The battery capacity is the rated capacity of the battery configured for the humanoid robot (e.g., 2000Wh), the target runtime is the preset working time (e.g., 20 hours), and the scenario power coefficient can be flexibly set according to actual conditions. For example, the scenario power coefficient for industrial handling is 1.2, for industrial assembly it is 1.0, for household cleaning it is 0.5, and for household delivery it is 0.6. Then, if the real-time energy consumption prediction value is less than or equal to the scene energy consumption threshold, the PWM duty cycle of the control signal input to the motor can be maintained, and the original working mode of the motor can be maintained; or, if the real-time energy consumption prediction value is greater than the scene energy consumption threshold, the PWM duty cycle of the control signal input to the motor can be reduced, and the working mode of the motor can be switched to intermittent operation mode. The intermittent operation mode has a running cycle of 6 seconds in the household cleaning scenario, including 5 seconds of operation and 1 second of stop; while in the industrial assembly scenario, the running cycle of the intermittent operation mode is 12 seconds, including 10 seconds of operation and 2 seconds of stop.
[0052] It is worth mentioning that, compared to the existing CNN-LSTM structure (i.e., CNN first, LSTM second), the LSTM-CNN structure used for temporal feature extraction and multidimensional differential feature extraction in this application embodiment is a targeted improvement based on humanoid robot energy consumption prediction. Specifically: 1. In terms of structural order: The typical application logic of the CNN-LSTM structure is "first extract spatial / local features through CNN, and then process temporal correlation through LSTM". However, in the scenario of "electric drive energy consumption prediction", since the "temporal continuity" in the scenario load sample data is the core dynamic factor that determines energy consumption (such as "high-frequency torque fluctuations of 3-4 seconds / time" in the household cleaning scenario and "low-frequency torque fluctuations of 20-30 seconds / time" in the industrial assembly scenario), the CNN in the front of the CNN-LSTM structure will first perform convolution operation, which will cut the continuous temporal load data (such as the torque sequence within 1 minute) into "local feature blocks", destroying the "preceding and subsequent dependencies" of the temporal data (for example, cutting the continuous sequence of "fluctuation interval of 3 seconds" into fragmented features of "interval of 1 second"). This makes it impossible for the subsequent LSTM to fully capture the "true frequency and interval of load fluctuation", ultimately causing the temporal related energy consumption error (such as reactive power loss caused by high-frequency start and stop) to be unable to be accurately predicted.
[0053] In this application embodiment, the LSTM-CNN structure can completely preserve and extract the "temporal dependency" of the scene load sample data through the front-end LSTM, and then use the back-end CNN to extract scene differences based on structured temporal features, which fully matches the influence mechanism of "energy consumption = static baseline + dynamic temporal load contribution" in electric drive energy consumption prediction.
[0054] 2. Regarding differences in scene features: The initial CNN layer of the CNN-LSTM structure processes raw load data that has not undergone temporal structuring. This raw load data typically contains a mixture of "scene features" (such as high torque in industrial applications and high frequency in household applications) and "noise features" (such as minor torque fluctuations caused by uneven ground and sensor errors). While the initial CNN layer can extract local features, it cannot distinguish between "which features are scene-specific" and "which are noise." As a result, during subsequent LSTM processing, scene features are diluted by noise, leading to low feature discrimination across scenes (e.g., the inability to accurately distinguish between "low load in industrial applications" and "high load in household applications").
[0055] In this embodiment, the input to the post-CNN is structured features with clear time-series dependencies and preliminary noise filtering (such as the "domestic high-frequency fluctuation time series map" and "industrial high-torque stable time series map" output by LSTM), which allows the post-CNN to more accurately focus on "scene-specific load feature differences". 3. Regarding training costs: Since the CNN-LSTM structure first processes the raw data through a pre-CNN, and the temporal information in the raw data is unstructured, the features extracted by the pre-CNN have a phenomenon of "high redundancy and weak correlation" (for example, in the same scenario, the CNN may extract a large number of features such as "torque peak", "average speed" and "meaningless noise fluctuations"); the subsequent LSTM needs to process these redundant features, resulting in low efficiency of model parameter updates, slow convergence speed, and high training costs.
[0056] The LSTM-CNN structure in this embodiment filters out "noise features without temporal significance" (such as <0.1N) in advance by pre-processing LSTM to structure temporal features. According to experimental data, the small torque fluctuations of m can reduce the feature redundancy of the input CNN by more than 60%. At the same time, the temporal features output by LSTM are more correlated with "scene-energy consumption" (such as high-frequency temporal features directly related to household scene energy consumption). The CNN layer can quickly focus on core features, the model parameters are updated more efficiently, and the training cost is effectively reduced.
[0057] The following describes in detail, with reference to the accompanying drawings, a training system for a humanoid robot energy consumption prediction model according to an embodiment of this application.
[0058] Reference Figure 2 The present application proposes a training system for a humanoid robot energy consumption prediction model, including... The first processing unit 101 is used to acquire static parameter sample data of the humanoid robot motor, scene load sample data of several different scene types, and measured energy consumption sample data corresponding to each scene load sample data. The second processing unit 102 is used to perform feature fusion on all the scene load sample data according to the static parameter sample data to obtain a fused feature matrix; The third processing unit 103 is used to perform matrix feature extraction on the fused feature matrix to obtain feature difference data. The feature difference data is used to characterize the peak feature amplitude difference, temporal density difference and energy distribution difference of scene load sample data under different scene types. The fourth processing unit 104 is used to update the parameters of the initialized humanoid robot energy consumption prediction model based on the feature difference data and all the measured energy consumption sample data, so as to obtain a trained humanoid robot energy consumption prediction model.
[0059] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0060] Reference Figure 3 This application also provides an electronic device, including: At least one processor 201; At least one memory 202 is used to store at least one program; When the at least one program is executed by the at least one processor 201, the at least one processor 201 implements the method embodiment described above.
[0061] Similarly, it can be understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0062] This application also provides a computer-readable storage medium storing a program executable by a processor 201, which, when executed by the processor 201, is used to implement the above-described method embodiments.
[0063] Similarly, the content of the above method embodiments is applicable to the present computer-readable storage medium embodiments. The specific functions implemented by the present computer-readable storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0064] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0065] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.
[0066] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this application are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.
[0067] Furthermore, although this application is described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding this application. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional technology for an engineer. Therefore, those skilled in the art can implement the application set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of this application, which is determined by the full scope of the appended claims and their equivalents.
[0068] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0069] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0070] More specific examples (a non-exhaustive list) of computer-readable media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0071] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0072] In the foregoing description of this specification, the references to terms such as "one embodiment," "another embodiment," or "some embodiments," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0073] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.
[0074] The above is a detailed description of the preferred embodiments of this application, but this application is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A training method for a humanoid robot energy consumption prediction model, characterized in that, include: Acquire static parameter sample data of the humanoid robot motor and scene load sample data of several different scene types, as well as measured energy consumption sample data corresponding to each scene load sample data; Based on the static parameter sample data, feature fusion is performed on all the scene load sample data to obtain a fused feature matrix; Matrix feature extraction is performed on the fused feature matrix to obtain feature difference data. The feature difference data is used to characterize the differences in peak feature amplitude, temporal density, and energy distribution of scene load sample data under different scene types. Based on the feature difference data and all the measured energy consumption sample data, the parameters of the initialized humanoid robot energy consumption prediction model are updated to obtain a trained humanoid robot energy consumption prediction model.
2. The method according to claim 1, characterized in that, The step of fusing features on all scenario load sample data based on the static parameter sample data to obtain a fused feature matrix includes: The static parameter sample data is statically vectorized to obtain static feature vectors; Dynamic feature vectorization is performed on all the scene load sample data to obtain several dynamic feature vectors, each of which corresponds to a combination of the scene load sample data and the scene type; Based on the static feature vector, all the dynamic feature vectors are concatenated to obtain the fused feature matrix.
3. The method according to claim 1, characterized in that, The step of extracting matrix features from the fused feature matrix to obtain feature difference data includes: Temporal feature extraction is performed on the fused feature matrix to obtain temporal dependency data, which records the temporal dependency relationship of each scenario load sample data. Multidimensional difference features are extracted from the time-series dependent data to obtain the feature difference data.
4. The method according to claim 3, characterized in that, The step of extracting multidimensional difference features from the time-series dependent data to obtain the feature difference data includes: The time-dependent data is subjected to the first convolution dimension extraction to obtain the first feature map, and the first feature map records the peak feature amplitude difference of the scene load sample data under different scene types. The second convolution dimension is extracted from the first feature map to obtain the second feature map. The second feature map records the temporal density difference of scene load sample data under different scene types and the peak feature amplitude difference of scene load sample data under different scene types. The second feature map is subjected to pooling dimension extraction to obtain the feature difference data.
5. The method according to claim 1, characterized in that, The step of updating the parameters of the initialized humanoid robot energy consumption prediction model based on the feature difference data and all the measured energy consumption sample data to obtain a trained humanoid robot energy consumption prediction model includes: Energy consumption prediction is performed on the aforementioned feature difference data to obtain energy consumption prediction data; Based on the measured energy consumption sample data, the energy consumption prediction data is analyzed for prediction error to obtain prediction error information; Based on the prediction error information, the parameters of the initialized humanoid robot energy consumption prediction model are updated to obtain the trained humanoid robot energy consumption prediction model.
6. The method according to claim 5, characterized in that, The prediction error information is represented as a function of: in, This represents prediction error information; n is the total number of samples. For the first Scene weights for each sample; For the first One sample of measured energy consumption data; In the energy consumption prediction data, the first Predicted energy consumption for each sample.
7. An energy optimization method, characterized in that, include: Obtain the static parameters, scene information, and real-time load information of the humanoid robot's motor at the current time point, as well as the scene energy consumption threshold corresponding to the scene information; The static parameters, the scene information, and the real-time load information are input into the humanoid robot energy consumption prediction model as described in any one of claims 1-6 to perform energy consumption prediction and obtain the real-time energy consumption prediction value. Based on the scenario energy consumption threshold and the real-time energy consumption prediction value, the energy of the humanoid robot motor is optimized.
8. A training system for a humanoid robot energy consumption prediction model, characterized in that, include: The first processing unit is used to acquire static parameter sample data of the humanoid robot motor, scene load sample data of several different scene types, and measured energy consumption sample data corresponding to each scene load sample data. The second processing unit is used to perform feature fusion on all the scene load sample data based on the static parameter sample data to obtain a fused feature matrix; The third processing unit is used to extract matrix features from the fused feature matrix to obtain feature difference data. The feature difference data is used to characterize the peak feature amplitude difference, temporal density difference, and energy distribution difference of scene load sample data under different scene types. The fourth processing unit is used to update the parameters of the initialized humanoid robot energy consumption prediction model based on the feature difference data and all the measured energy consumption sample data, so as to obtain the trained humanoid robot energy consumption prediction model.
9. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method as described in any one of claims 1-7.
10. A computer-readable storage medium storing a processor-executable program, characterized in that, The processor-executable program, when executed by the processor, is used to implement the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Data-driven industrial robot energy consumption optimization method
CN110936382A
Non-intrusive load monitoring method based on transfer learning and self-attention feature fusion
CN116742795A
Vehicle energy consumption information prediction method, computer equipment and storage medium
CN118269991A
Airport energy balance management method, system, equipment and medium
CN120235415A
Non-intrusive power load decomposition method and system based on multi-modal feature learning
CN120822062A