Humanoid robot learning method based on frequency domain enhanced wavelet Transform model
By adopting the wavelet Transformer model based on frequency domain enhancement in humanoid robot learning methods, using wavelet transform to capture multi-scale information, the problem of difficulty in capturing frequency information in the prior art is solved, and the learning efficiency and task completion ability are significantly improved.
Patent Information
- Application Number
- CN202510020789.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively capture frequency information in complex and changing scenarios, resulting in humanoid robots losing key information during the learning process and making it difficult to realize the learning and operation of fine tasks.
Using the wavelet Transformer model based on frequency domain enhancement, frequency domain enhancement is introduced on EMA through the FE-EMA module, multi-scale information is captured using wavelet transform, and stronger robustness is shown in frequency domain information extraction.
It significantly improves the learning efficiency and task completion ability of humanoid robots, achieving richer feature expression and stronger robustness.
Smart Images

Figure CN120105049A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of humanoid robot learning, in particular to a humanoid robot learning method based on a frequency domain enhanced wavelet Transformer model. Background Art
[0002] Imitation learning has important theoretical and practical significance in the research of humanoid robots. Humanoid robots have always been the core direction of intelligent robot research due to their high degree of freedom, multi-redundancy motion characteristics and potential ability to perform delicate tasks in dynamic and complex environments. However, traditional rule-based or optimization-based control methods often have difficulty in dealing with high-dimensional state space and nonlinear dynamics problems in humanoid robot motion planning. Imitation learning provides an efficient and intuitive solution by directly learning the optimal action strategy or control model using human expert demonstration data. The introduction of imitation learning enables humanoid robots to complete tasks in a more human-like way, improving their applicability and scalability in the fields of human-machine collaboration, service robots, and human behavior assistance. Therefore, as a key technology that bridges human intelligence and robot intelligence, imitation learning has irreplaceable value in promoting humanoid robots to achieve autonomous learning and complex behavior generation.
[0003] Under the current technological background, visual image detection and time series prediction still face many challenges. This is because the information in the real world is distributed in complex and changeable scenes. The prediction results of existing mainstream prediction models deviate greatly from the true values. This phenomenon is mainly due to the inability of the model to fully capture the rich frequency information in the real data set, which causes humanoid robots to lose a lot of key information during the learning process and make it difficult to learn and operate fine tasks. Discrete wavelet transform (DWT), as a tool widely used in image processing, is known for its excellent performance in multi-resolution analysis. Although the application of DWT in image processing has matured, its potential in time series data modeling has not been fully explored and utilized. The traditional image detection and time series prediction methods are difficult to meet the technical bottleneck of the fine task learning needs of humanoid robots. Summary of the invention
[0004] In view of the above problems existing in the prior art, the present invention is proposed.
[0005] Therefore, the technical problem to be solved by the present invention is a problem.
[0006] To achieve the above object, the present invention provides the following technical solution: a humanoid robot learning method based on a frequency domain enhanced wavelet Transformer model, comprising:
[0007] Collect the RGB image dataset, active arm joint information dataset, slave arm joint information dataset, chassis linear angular velocity dataset and IMU linear angular velocity dataset of the humanoid robot;
[0008] Input the boom joint information dataset, chassis linear angular velocity dataset, IMU linear angular velocity dataset, and CLS classification head into the Transformer encoder module to obtain re-parameterized samples;
[0009] The RGB image dataset processed by the FE-EMA module, the IMU linear angular velocity dataset, the joint information dataset, and the re-parameterized samples are input into the wavelet Transformer model, and the forward propagation calculation is performed to obtain the predicted action sequence;
[0010] The loss value is calculated by the mean square error between the predicted action sequence and the actual action sequence, and the gradient is updated through back propagation to optimize the weight parameters of the wavelet Transformer model;
[0011] Obtain a real-time data set of a humanoid robot, and input the re-parameterized samples set to 0 into the trained model. At the same time, pass in the real-time observation data of the humanoid robot (the real-time observation data includes camera image data, joint information, and IMU linear angular velocity data set) to verify whether the actions of the humanoid robot can complete the previously trained tasks.
[0012] As a further solution of the present invention: the RGB image data set obtains image data through cameras respectively arranged on the head and left and right wrists of the humanoid robot.
[0013] As a further solution of the present invention: the active arm joint information dataset, the slave arm joint information dataset, the chassis linear angular velocity dataset and the IMU linear angular velocity dataset are obtained through ROS, and the slave arm joint information dataset, the chassis linear angular velocity dataset, the IMU linear angular velocity dataset and the CLS are input into the Transformer encoder module, and the step of obtaining the re-parameterized sample includes:
[0014] The slave arm joint information dataset, chassis linear angular velocity dataset, IMU linear angular velocity dataset and CLS are uniformly mapped to 512 dimensions through a linear layer for processing, and the action sequence is separately added with sinusoidal position encoding;
[0015] Input the processed data into the Transformer encoder module;
[0016] The data processed by the Transformer encoder module is trained by VAE to obtain the mean and variance of the sample distribution in the latent space and obtain the reparameterized samples.
[0017] As a further solution of the present invention: the action sequence includes an active arm joint information data set and a chassis linear angular velocity data set.
[0018] As a further solution of the present invention: the RGB image data set processed by the FE-EMA module, the IMU linear angular velocity data set, the active arm joint information data set, the slave arm joint information data set and the re-parameterized samples are input into the wavelet Transformer model, and the step of calculating and predicting the action sequence includes extracting the frequency domain features of the image from the RGB image data set through the FE-EMA module;
[0019] The processed image data set is input into the Transformer encoder module together with the IMU linear angular velocity data set, the active arm joint information data set, the slave arm joint information data set and the re-parameterized samples to obtain the time domain data;
[0020] Input the time domain data into the DWT module for processing and then output the frequency domain data;
[0021] At the same time, the time domain data and frequency domain data are fused into time-frequency data and input into the Transformer decoder module for calculation to generate a predicted action sequence.
[0022] As a further solution of the present invention: the steps of extracting the frequency domain features of the image by wavelet transform in the FE-EMA module are as follows:
[0023] Step 1: Expand and reconstruct; expand the feature map by group along the batch and channel dimensions:
[0024]
[0025] B is the batch size, C is the number of channels, and H and W are the height and width of the image respectively.
[0026] Step 2: Convolution operation realizes wavelet decomposition, and calculates cA, cH, cV, cD for each set of feature maps:
[0027]
[0028] Step 3: Fusion of high-frequency information. In the FE-EMA module, the high-frequency information HighFreq = cH + cV + cD is added to the low-frequency component cA to obtain enhanced features:
[0029] Enhanced_Wavelet=ResBlock(cA+HighFreq)
[0030] Step 4: Fusion of time domain and frequency domain features. A dynamic weight α is used in this module to control the fusion of time domain and frequency domain features:
[0031] Fused_X=α·X+(1-α)·Enhanced_Wavelet
[0032] The weight α is generated by global average pooling and a linear layer:
[0033] α=σ(W·GAP(X))
[0034] where σ(·) is the Sigmoid function.
[0035] As a further solution of the present invention: the step of inputting time domain data into the DWT module for processing and then outputting frequency domain data comprises:
[0036] Input time domain data and flatten the three-dimensional data into two-dimensional data;
[0037] Decompose input data into low-frequency and high-frequency components;
[0038] Through specific feature extraction modules, low-frequency and high-frequency components are processed separately;
[0039] The low-frequency and high-frequency component data are concatenated and feature fused through 1D convolution, and reconstructed back to the original shape through residual connection;
[0040] Output frequency domain attention weights through the FC layer;
[0041] Multiply the frequency domain attention weight by the input data and output the frequency domain data.
[0042] As a further solution of the present invention: the predicted action sequence and the real action sequence are compared through the mean square error to calculate the loss, and the step of back propagating and optimizing the weight parameters of the wavelet Transformer model includes:
[0043] Calculate the loss between the predicted action sequence obtained by forward propagation and the actual action sequence;
[0044] Perform back propagation and adjust the wavelet Transformer model parameters in reverse.
[0045] As a further solution of the present invention: the formula of low frequency component is,
[0046]
[0047] Among them, φ j,k [n] is the scaling function, which represents the low-frequency part of the signal, j represents the scale of decomposition, k represents the translation, and n represents the length of the sequence;
[0048] The formula for high frequency component is:
[0049]
[0050] Among them, ψ j,k [n] is the wavelet function, which represents the high-frequency part of the signal, j represents the scale of decomposition, k represents the translation, and n represents the length of the sequence.
[0051] Compared with the prior art, the beneficial effects of the present invention are: this humanoid robot learning method based on the frequency domain enhanced wavelet Transformer model introduces frequency domain enhancement in the FE-EMA module on the basis of EMA, uses wavelet transform to capture multi-scale information, achieves richer feature expression, and shows stronger robustness in frequency domain information extraction; it significantly improves the learning efficiency and task completion ability of the humanoid robot. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0053] Figure 1 This is a detailed algorithm architecture diagram of the frequency domain enhanced wavelet transformer model described in the embodiment provided by the present invention.
[0054] Figure 2 This is a frequency domain transformation diagram in the wavelet transformer described in the embodiment of the present invention.
[0055] Figure 3 The feature heat map of the output tensor of the Transformer encoder layer described in the embodiment provided by the present invention.
[0056] Figure 4 A feature heat map of the output tensor of the Wavelet-Transformer encoder layer described in the embodiment provided by the present invention.
[0057] Figure 5 This is an algorithm flow chart of the frequency domain enhanced multi-scale attention module (FE-EMA) described in the embodiment provided by the present invention.
[0058] Figure 6 This is a heat map of image features processed by the Resnet-18 module according to the embodiment provided by the present invention.
[0059] Figure 7 This is the image feature heat map processed by the FE-EMA module as described in the embodiment provided by the present invention.
[0060] Figure 8This is a training and verification curve diagram of the frequency domain enhanced wavelet transformer model described in the embodiment provided by the present invention. DETAILED DESCRIPTION
[0061] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings.
[0062] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0063] Secondly, the present invention is described in detail with reference to the schematic diagram. When describing the embodiments of the present invention in detail, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.
[0064] Furthermore, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selective embodiment that is mutually exclusive with other embodiments.
[0065] Example 1
[0066] like Figure 1 As shown, the present invention provides a technical solution: a humanoid robot learning method based on a frequency domain enhanced wavelet Transformer model, comprising:
[0067] S1: Collect data, specifically collect the RGB image dataset of the humanoid robot, the active arm joint information dataset, the slave arm joint information dataset, the chassis linear angular velocity dataset, and the IMU linear angular velocity dataset;
[0068] It should be noted that the RGB image dataset obtains image data through cameras respectively set on the head and left and right wrists of the humanoid robot, and the image data is obtained in real time through the cameras; the active arm joint information dataset, the slave arm joint information dataset, the chassis linear angular velocity dataset and the IMU linear angular velocity dataset obtain the real-time switching angle of the robotic arm through ROS (Robot Operating System); the active arm joint information dataset and the slave arm joint information dataset include a joint position information dataset with 14 degrees of freedom in total, the chassis linear angular velocity dataset includes the linear velocity and angular velocity dataset of the chassis, and the IMU linear angular velocity dataset includes the angular velocity and linear acceleration dataset of the IMU in the spatial coordinate system; the RGB image dataset, the joint position information dataset of the dual master / slave end robotic arms with a total of 14 degrees of freedom, the linear velocity and angular velocity dataset of the chassis, and the angular velocity and linear acceleration dataset of the IMU in the spatial coordinate system are a total of 6-dimensional datasets, and the overall degrees of freedom of the humanoid robot are 16 (14 degrees of freedom of both arms + 2 degrees of freedom of the mobile chassis).
[0069] S2: Infer z (latent space), specifically inputting the boom joint information dataset, chassis linear angular velocity dataset, IMU linear angular velocity dataset and CLS into the Transformer encoder module to obtain re-parameterized samples;
[0070] Specifically, the boom joint information dataset, chassis linear angular velocity dataset, IMU linear angular velocity dataset and CLS are uniformly mapped to 512 dimensions through a linear layer, and the action sequence is separately encoded with sinusoidal position encoding; the processed data are simultaneously placed in the Transformer encoder module, and the mean and variance of its latent space sample distribution are trained in combination with VAE, and the latent feature samples are re-parameterized and placed in the subsequent algorithm module as feature dimensions.
[0071] It should be noted that CLS is an input tag that represents the global representation of the entire input sequence (such as an image or a text).
[0072] S3: Action prediction: Specifically, the RGB image dataset processed by the FE-EMA module, the IMU linear angular velocity dataset, the joint information dataset, and the re-parameterized samples are input into the wavelet Transformer model, and the forward propagation calculation is performed to obtain the predicted action sequence;
[0073] Specifically, the image data collected by the camera is subjected to end-to-end object detection through an architecture that combines the FE-EMA module with the backbone network. The FE-EMA module can efficiently extract image features, retain channel information and enhance feature expression capabilities while reducing computational overhead. Subsequently, the processed image data is input into the Transformer encoder module together with the IMU data, slave end joint data, and re-parameterized feature samples. The output of the encoder passes through the DWT module in the wavelet Transformer model, and uses wavelet transform to fully fuse the frequency domain information in the sequence, thereby improving the deep expression capabilities of the features. Finally, the output result is calculated by the Transformer decoder to generate a predicted action sequence.
[0074] It should be noted that the model is trained in this step; the weight file of the trained model will be saved in the .ckpt file format and cannot be displayed intuitively. The specific training process is shown in the attached figure. Figure 8 Display; The trained model will save the trained weight file, which will be imported during verification. The effect of the previous training will be verified through the accuracy of completing some specific tasks, such as throwing garbage and opening drawers.
[0075] S4: Calculate the loss value of the predicted action sequence and the actual action sequence through the mean square error, update the gradient through back propagation, and optimize the weight parameters of the wavelet Transformer model;
[0076] S5: Obtain the real-time data set of the humanoid robot, and input the re-parameterized samples set to 0 into the trained model, and at the same time pass in the real-time observation data of the humanoid robot (the real-time observation data includes camera image data, joint information, and IMU linear angular velocity data set) to verify whether the actions of the humanoid robot can complete the previously trained tasks.
[0077] Further, the step of inputting the boom joint information dataset, chassis linear angular velocity dataset, IMU linear angular velocity dataset and CLS into the Transformer encoder module to obtain re-parameterized samples includes:
[0078] S21: The boom joint information dataset, chassis linear angular velocity dataset, IMU linear angular velocity dataset and CLS are uniformly mapped to 512 dimensions through a linear layer for processing, and the action sequence is separately added with sinusoidal position encoding;
[0079] Among them, in deep learning, the linear layer (also called the fully connected layer, Fully Connected Layer, FC layer) is a common neural network layer, which is used to learn the characteristics of data by linearly transforming the input; 512 dimensions are the dimensions of the data; for example, the original data is 1×14 in size, and it can be mapped to 1×512 in size through the linear layer; as for the principle, it is understood as a step of matrix operation; in deep learning, the dimension of the hidden layer is considered to be a hyperparameter, which can be set by yourself, usually 2 to the power of n. The setting of 512 dimensions here is self-defined. Through experiments, it can be found that the reasoning effect of humanoid robots is better when the data is unified to 512 dimensions for learning; unifying the data to 512 dimensions is essentially just to facilitate matrix splicing of the data; for example, if the image data becomes 300×512 data after passing through DETR and some network architectures, and the joint angles map the 1×14 data to 1×512 dimensional data through the linear layer, the data is spliced by column, so that the original data of different dimensions can be unified to 512 dimensions.
[0080] It should be noted that the action sequence includes the active arm joint information dataset and the chassis linear angular velocity dataset, specifically including the active end dual arms, a total of 14 joint angle data and chassis 2 degrees of freedom data (linear velocity, angular velocity), and the data dimension is k×16 (k is a hyperparameter for pre-action segmentation); the action sequence is no different from the above dataset collection method. This step is to put the collected data into VAE (variational autoencoder) to train the potential features of the data.
[0081] S22: Input the processed data into the Transformer encoder module;
[0082] S23: The data processed by the Transformer encoder module is trained by VAE to obtain the mean and variance of the latent space sample distribution and obtain the re-parameterized samples.
[0083] Furthermore, the RGB image data set processed by the FE-EMA module, the IMU linear angular velocity data set, the active arm joint information data set, the slave arm joint information data set and the re-parameterized samples are input into the wavelet Transformer model. The steps of calculating and predicting the action sequence include:
[0084] S31: The RGB image dataset is processed by the FE-EMA module to extract the frequency domain features of the image;
[0085] S32: The processed image data set, the IMU linear angular velocity data set, the active arm joint information data set, the slave arm joint information data set and the re-parameterized samples are input into the Transformer encoder module to obtain time domain data;
[0086] S33: input the time domain data into the DWT module for processing and then output the frequency domain data;
[0087] S34: Simultaneously fuse the time domain data and the frequency domain data into time-frequency data, input the data into the Transformer decoder module for calculation, and generate a predicted action sequence.
[0088] It should be noted that the wavelet Transformer model includes a Transformer encoder module, a DWT module and a Transformer decoder module, and the Transformer encoder module is connected to the Transformer decoder module through the DWT module.
[0089] Furthermore, the predicted action sequence and the actual action sequence are compared by calculating the loss through the mean square error. The steps of back-propagation optimization of the weight parameters of the wavelet Transformer model include:
[0090] S41: Calculate the loss between the predicted action sequence obtained by forward propagation and the actual action sequence;
[0091] S42: Perform back propagation to adjust the parameters of the wavelet Transformer model in reverse.
[0092] This solution adopts a deep learning model framework and uses an end-to-end learning strategy for training. The significant advantage of end-to-end learning is that it can uniformly model the entire process from raw input to target output, avoiding complex manual design of intermediate features or steps. In addition, end-to-end learning can automatically extract and fuse multi-dimensional features, more efficiently capture the deep correlation between input and output, and is particularly suitable for nonlinear mapping between high-dimensional input (such as visual information) and complex output (such as joint movements and trajectory planning) in humanoid robots. In this embodiment, the robot uses an end-to-end learning strategy to use environmental images, chassis speed, joint angles, chassis angular velocity and linear acceleration recorded by IMU, and other data to make real-time and accurate predictions of the motion state of future time steps. This strategy gives full play to the advantages of deep learning, enabling robots to efficiently complete complex tasks in dynamic environments.
[0093] Example 2
[0094] The difference from the previous embodiment is that: the specific steps of inputting time domain data into the DWT module for processing and then outputting frequency domain data are recorded. Unlike the traditional Fourier transform (FT), DWT does not need to assume the periodicity of the signal, and effectively avoids the introduction of high-frequency noise (such as Gibbs phenomenon) through multi-resolution analysis; it makes DWT more robust in extracting frequency domain information and can better preserve the original characteristics of the signal; in addition, the energy concentration of DWT further enhances the ability to capture low-frequency signals, providing important support for long-period time series modeling. Therefore, the discrete wavelet transform (DWT) is introduced into the time series prediction model (in Figure 1 In the second figure, the DWT algorithm module is introduced at the output of the Transformer encoder and the input of the decoder, which not only enhances its modeling ability for complex time series features, but also significantly improves the reliability and accuracy of the prediction results.
[0095] Its discrete wavelet transform (DWT) is a mathematical tool that decomposes signals into different scales and frequency bands. It is applicable to various signal forms such as one-dimensional time series and two-dimensional images.
[0096] In this embodiment, the step of inputting time domain data into the DWT module for processing and then outputting frequency domain data includes:
[0097] S331: input time domain data, and flatten the three-dimensional data into two-dimensional data;
[0098] Specifically, the time series data collected by the humanoid robot (time series data is the image, joint angle, chassis linear velocity and angular velocity and other data at each time) is input into the DWT module in the form of tensors (tensors are the basic data structure in deep learning and are widely used to represent data. It can be regarded as a multidimensional array with a specific shape (i.e., dimension). Tensors are generalizations of scalars, vectors, and matrices, and can represent higher-dimensional data). In order to facilitate subsequent processing, the input three-dimensional time series data is flattened (the three-dimensional data is directly flattened into two-dimensional data. The specific implementation process can be understood through the following example: if there is a cube with length × width × height, the height of the cube is divided into 1 cm at a time, and after the division, it is spliced layer by layer on the plane) into a two-dimensional form to adapt it to the needs of DWT decomposition. This step aims to reduce data complexity while ensuring the efficiency of decomposition and feature extraction in subsequent steps.
[0099] S332: Decomposing the input data into low frequency (cA) and high frequency (cD) components;
[0100] It should be noted that the low-frequency component cA mainly retains the global trend information in the time series, while the high-frequency component cD focuses on local details and short-term changes. Through this decomposition step, a more structured feature basis can be provided for the subsequent feature extraction module.
[0101] S333: Process low-frequency and high-frequency components separately through a specific feature extraction module.
[0102] Need to explain, such as Figure 2 As shown, the specific feature extraction module is part of the wavelet Transformer model; the feature extraction module firstly performs data dimension upgrading and feature extraction on the low-frequency and high-frequency data through the linear layer, and then reshapes the data into three-dimensional data.
[0103] S334: concatenate the low-frequency and high-frequency component data and perform feature fusion through 1D convolution, and reconstruct them back to the original shape through residual connection;
[0104] After completing the independent processing of low-frequency and high-frequency components, the two parts of data are spliced and feature fusion is achieved through one-dimensional convolution (1D convolution). The 1D convolution operation can integrate feature information of different frequencies and further compress and optimize the time series data. The fused feature data is also reconstructed back to the same original shape as the input tensor through residual connection to ensure the integrity of information transmission.
[0105] S335: Output frequency domain attention weights through the FC layer;
[0106] These weights represent the model's assessment of the importance of features of different frequencies and can dynamically adjust the contribution of features to more accurately capture the core patterns of time series.
[0107] Among them, the FC layer is used to process the fused frequency domain information data into frequency domain attention weights, which is part of the wavelet Transformer model.
[0108] S336: Multiply the frequency domain attention weight by the input data and output the frequency domain data.
[0109] The frequency domain attention weights are multiplied element by element with the input data to obtain the output data combined with the frequency domain information. This step effectively strengthens the key features while weakening noise interference and irrelevant information, providing a more reliable and efficient feature representation for subsequent task modeling and action prediction.
[0110] Specifically, the basic formula of DWT is to perform discrete wavelet transform decomposition on the data output by the Transformer encoder. The discrete wavelet transform projects the signal x(t) onto a set of orthogonal wavelet basis functions to obtain low-frequency (approximate) components and high-frequency (detail) components. For a discrete signal x[n], its discrete wavelet transform can be expressed as:
[0111] The formula for low-frequency component (approximate coefficient) is,
[0112]
[0113] Among them, φ j,k [n] is the scaling function, which represents the low-frequency part of the signal, j represents the scale of decomposition, k represents the translation, and n represents the length of the sequence;
[0114] The formula for high frequency component (detail coefficient) is:
[0115]
[0116] Among them, ψ j,k [n] is the wavelet function, which represents the high-frequency part of the signal, j represents the scale of decomposition, k represents the translation, and n represents the length of the sequence.
[0117] Among them, the definition of wavelet function and scaling function: wavelet function and scaling function are constructed through scaling and translation relationship.
[0118] Scaling function:
[0119] φ j,k [n] = 2 j / 2 ·φ(2 j nk)
[0120] Among them, φ is the basic scaling function, j controls the resolution (scale), and k controls the time position (translation).
[0121] Wavelet function:
[0122] ψ j,k [n] = 2 j / 2 ·ψ(2 j nk)
[0123] ψ is the basic wavelet function that can capture the local changes of the signal in time and frequency.
[0124] Among them, the recursive decomposition process: through DWT, the signal is decomposed into low-frequency and high-frequency parts. The specific steps are as follows:
[0125] Convolution operation: Use low-pass filter h[n] and high-pass filter g[n] to convolve the signal respectively:
[0126]
[0127] The above is the basic formula and process for discrete wavelet transform of time series. The specific algorithm process is as follows: Figure 2 shown.
[0128] The DWT module provided in the present invention plays an important complementary role in time domain analysis, in which the low-frequency component cA retains most of the energy in the time series and can clearly reflect the global trend of the data. This feature is particularly critical for long-time span prediction tasks, and the global trend information can effectively help the model understand the macroscopic change law of the data, thereby improving the accuracy of the prediction.
[0129] At the same time, the high-frequency component cD of the module provides rich local fluctuation information, which can accurately capture the short-term drastic changes in the movement of humanoid robots, which is crucial for realizing short-term motion prediction and dynamic adjustment. By processing the high-frequency components separately, the DWT module can effectively separate noise and useful signals, significantly improving the system's anti-noise performance. By modeling and fusing low-frequency and high-frequency features separately, the present invention can more deeply explore the potential feature expression capabilities of time series, and compared with directly processing time domain data, it improves the accuracy of humanoid robot motion sequence prediction and the success rate of task completion.
[0130] also, Figure 3 and Figure 4 The characteristic heat maps of the output tensors of the traditional Transformer encoder layer and the wavelet Transformer encoder layer are shown respectively. It can be seen intuitively from the figure that the wavelet Transformer can more effectively capture the feature information in different channels and different frequency ranges. Compared with traditional methods, it has obvious advantages in both the depth and breadth of feature extraction. This ability enables the model to not only accurately model the global trend of the time series, but also capture subtle fluctuations and local changes in the short term. By combining frequency domain information, the sequence features generated by the wavelet Transformer are richer in structure and expression ability, and are more in line with the high requirements of humanoid robot complex tasks for time series data analysis. The experimental results further prove that the model has achieved significant improvements in the accuracy, robustness and task adaptability of action sequence prediction, and can provide more powerful technical support for the behavior planning of humanoid robots in complex scenarios.
[0131] Example 3
[0132] The difference between this embodiment and the previous embodiment is that the principle and optimization points of the frequency domain enhanced multi-scale attention module (FE-EMA) are further disclosed, so as to improve the adaptability and expressiveness of the visual detection of humanoid robots in complex tasks; compared with the traditional multi-scale attention module (EMA), the frequency domain enhanced multi-scale attention module proposed in this scheme achieves significant improvement in feature extraction and expression capabilities.
[0133] Specifically, FE-EMA introduces a frequency domain enhancement mechanism based on EMA, effectively captures multi-scale information through wavelet transform, makes up for the shortcomings of the EMA module in frequency domain feature extraction, and significantly enriches the feature expression dimension. In addition, FE-EMA sets an α weight parameter, which can adaptively adjust the fusion ratio of time domain and frequency domain features, thereby dynamically optimizing the weight distribution of feature expression in complex data scenarios, enhancing the adaptability and robustness of the model. FE-EMA comprehensively surpasses the traditional EMA module in capturing global and local information, detail retention and noise resistance, providing more powerful technical support for humanoid robot image detection and classification tasks.
[0134] Efficient Multi-Scale Attention Module with Cross-Spatial Learning proposes a new multi-scale attention module (EMA) to improve the efficiency of feature extraction in computer vision tasks. Traditional channel or spatial attention mechanisms may bring large computational overhead in the process of deep visual representation extraction, while the EMA module effectively reduces this overhead through innovative design, while retaining channel information and enhancing feature representation capabilities.
[0135] The EMA module introduces a cross-space learning method and designs a multi-scale parallel sub-network that can establish short-term and long-term dependencies. By reshaping part of the channel dimension into the batch dimension, EMA effectively avoids the information loss caused by the reduction of the channel dimension. In addition, EMA achieves local cross-channel interaction through parallel 1×1 and 3×3 convolution branches while preserving spatial structure information.
[0136] The model also proposes an innovative cross-spatial information aggregation method to achieve richer feature fusion by aggregating in different spatial dimensions; specifically, this method combines 1×1 convolution, 3×3 convolution and global average pooling to more efficiently aggregate and encode global spatial information; the 1×1 branch mainly captures local cross-channel interactions and encodes global spatial information into the output through 2D global average pooling; the 3×3 branch expands the feature space, captures contextual information, and performs shape matching in the channel dimension to participate in fusion and activation with other branches; finally, the EMA module generates a cross-spatial attention map through matrix dot product, thereby effectively capturing pixel-level relationships.
[0137] This innovative EMA module can significantly improve feature extraction capabilities and preserve spatial information while reducing computational overhead. It has broad application potential and demonstrates strong advantages in computer vision tasks.
[0138] Combining the FE-EMA module in the Backbone is a reasonable and practical choice, especially in computer vision tasks; the advantages of using FE-EMA combined with ResNet-18 are:
[0139] 1. Enhance high-level feature expression: Using FE-EMA after the output of layer4 can enhance the fusion ability of global and local information, making the model more expressive in classification or detection tasks.
[0140] 2. Improve model robustness: Perform multi-scale processing on features to make the model perform better when facing complex backgrounds or targets of different scales.
[0141] 3. Improve model performance: Features enhanced by FE-EMA can be better utilized by the classifier or detection head, improving the overall performance of the model.
[0142] Integrating the FE-EMA module into ResNet-18 can enhance the model's multi-scale feature expression capability, more accurately capture local and global information through spatial and channel attention mechanisms, and improve the performance and robustness of the model in tasks such as object detection and classification.
[0143] From the perspective of multiresolution analysis (MRA), taking the application of wavelet on two-dimensional feature maps as the starting point, combined with tensor transformation and convolution operations, the mathematical expression of two-dimensional wavelet transform on feature maps is elaborated in detail below:
[0144] Wavelet transform is a localized multi-scale decomposition of a signal. When applied to a feature map, the purpose of wavelet transform is to decompose the input feature map X into coefficient matrices of different frequency bands (approximate components and detail components) and capture the local information of the feature map at multiple scales.
[0145] It should be noted that the wavelet transform here is not the wavelet transform of the Transformer module, but the wavelet transform of the image data. The present invention highlights the influence of frequency domain information on the learning method of humanoid robots, so the present invention is actually composed of two modules, FE-EMA and wavelet Transformer. FE-EMA is responsible for processing two-dimensional image data, while the wavelet transformer model is responsible for processing one-dimensional time series data, both of which are to capture the rich frequency domain information in the data to improve the learning efficiency and reasoning accuracy of humanoid robots. For details on how to integrate the two modules, please refer to Figure 1 The second picture.
[0146] (1) Core principle of wavelet transform: Given an input feature map X∈R h×w, whose wavelet transform is constructed by constructing a set of scaled and translated wavelet basis functions ψ j,k (x) and the scaling function φ j,k (x), the signal is expressed as:
[0147]
[0148] Among them, j is the scale parameter (controlling frequency resolution); k is the translation parameter (controlling time and space positioning); cA j,k is the low-frequency approximation coefficient; cD i,j,k are high-frequency detail coefficients (capturing details in the horizontal, vertical, and diagonal directions, respectively); Z represents the set of integers, that is, all integers (including positive integers, negative integers, and 0), used to indicate that variables j and k can take any integer values.
[0149] (2) Two-dimensional discrete wavelet transform (DWT) on the feature map: On the two-dimensional feature map X, the wavelet transform is implemented using a two-dimensional filter bank. The feature map is decomposed by rows and columns, and the following coefficients are obtained:
[0150] DWT(X)=(cA,cH,cV,cD)
[0151] Among them, cA is the low-frequency approximation component, which captures the overall structural information; cH is the horizontal detail component, which captures the edge features in the horizontal direction; cV is the vertical detail component, which captures the edge features in the vertical direction; cD is the diagonal detail component, which captures the edge changes in the diagonal direction.
[0152] (3) Convolution calculation formula of wavelet filter: The convolution kernel of db1 (Haar wavelet) is defined as follows:
[0153]
[0154] For convolution calculations on two-dimensional feature maps, these filters can be used on rows and columns respectively:
[0155] Horizontal convolution (along the rows):
[0156]
[0157] Vertical convolution (along the columns):
[0158]
[0159] The final coefficients cA, cH, cV, cD are obtained by combining the horizontal and vertical convolution results.
[0160] The above is the basic formula and process for discrete wavelet transform of images. The specific algorithm process is as follows: Figure 5 As shown, the figure gives Figure 1The detailed description of FE-EMA module includes:
[0161] 1. Multi-branch convolutional structure: 1×1 convolution is used for dimensionality reduction and global information extraction to reduce computational complexity. 3x3 convolution captures local features through ResBlock to enhance the fine-grained information expression of the model.
[0162] 2. Introduction of wavelet transform (DWT): Decompose the input features into low-frequency (cA) and high-frequency (cH, cV, cD) parts. The low-frequency component (cA) captures the global structure and main features. The high-frequency components (cH, cV, cD) capture details, textures, and edge information. The frequency domain features are enhanced through ResBlock.
[0163] 3. Adaptive weight fusion (α weight): According to the global pooling result, the adaptive fusion weight α is dynamically generated to dynamically balance the contribution of time domain and frequency domain features.
[0164] 4. Bilinear interpolation: Interpolate the features after frequency domain enhancement to ensure that the feature size is consistent with the original input.
[0165] In the FE-EMA module, the high-frequency part represents the detailed features in the image, such as edges and textures, which are separated by wavelet transform. The high-frequency part provides details and edge information, while the low-frequency part retains the global structural information. When the two are combined, the model can express the input features more comprehensively. The 3×3 convolution enhances these high-frequency details in ResBlock and suppresses noise through nonlinear activation, allowing the model to capture fine-grained information more accurately. The high-frequency part is highly consistent with the locality of the convolution operation, ensuring that the model performs better in different tasks. This design enables the FE-EMA module to capture global information while fully retaining key details when processing complex data, providing strong support for the feature expression of the model.
[0166] The optimization effect of wavelet transform in FE-EMA is reflected in:
[0167] (1) Wavelet transform divides the features into low-frequency and high-frequency parts, and enhances the features in the frequency domain through the coordination of residual blocks.
[0168] (2) In the FE-EMA module, wavelet transform splits the original feature map into multiple frequency components (reduced in size). Convolution only acts on these reduced features, so the overall FLOPs are reduced.
[0169] (3) Like EMA, grouped convolution is also used in FE-EMA. Since the convolution kernel acts on the downsampled features and processes them in groups, the amount of computation is further reduced.
[0170] In deep learning, given an input feature map X∈RB×C×H×W (where B is the batch size and C is the number of channels), the two-dimensional wavelet transform can be implemented as follows:
[0171] Step 1: Expand and reconstruct. Expand the feature map by group along the batch and channel dimensions:
[0172]
[0173] B is the batch size, C is the number of channels, and H and W are the height and width of the image respectively.
[0174] Step 2: Convolution operation realizes wavelet decomposition, and calculates cA, cH, cV, cD for each set of feature maps (here represents the Kronecker product of the two-dimensional convolution kernel):
[0175]
[0176] Step 3: Fusion of high-frequency information. In the FE-EMA module, the high-frequency information HighFreq = cH + cV + cD is added to the low-frequency component cA to obtain enhanced features:
[0177] Enhanced_Wavelet=ResBlock(cA+HighFreq)
[0178] Step 4: Fusion of time domain and frequency domain features. A dynamic weight α is used in this module to control the fusion of time domain and frequency domain features:
[0179] Fused_X=α·X+(1-α)·Enhanced_Wavelet
[0180] The weight α is generated by global average pooling and a linear layer:
[0181] α=σ(W·GAP(X))
[0182] where σ(·) is the Sigmoid function.
[0183] The FE-EMA module combines multi-scale analysis of wavelet decomposition and residual learning of convolutional neural networks. By dynamically fusing features between the time domain and the frequency domain, this method can efficiently capture information at different scales and directions: the low-frequency part captures global structural features, and the high-frequency part provides local edge and detail information. Residual learning ensures that while enhancing frequency domain features, key information of the original feature map is not lost.
[0184] Example 4
[0185] The difference from the previous embodiment is that in order to verify the effectiveness of the FE-EMA module in the present invention in the humanoid robot visual detection task, we visualized the image data processed by the ResNet-18 and FE-EMA modules using heat maps, as shown in Figure 2. Figure 6 and Figure 7 As shown in the figure. It can be clearly seen from the heat map that compared with the images processed by the traditional ResNet-18 model, the FE-EMA module can focus on the key areas in the image more accurately. Especially when the humanoid robot performs the garbage disposal task, the FE-EMA module pays more attention to the gripper and the target object, and can effectively capture the spatial relationship between the gripper and the target object. This result shows that the FE-EMA module has shown stronger accuracy and local feature extraction capabilities in visual information processing, providing strong support for the precise operation of humanoid robots in complex tasks.
[0186] The humanoid robot learning method of the present invention combines discrete wavelet transform (DWT) with the Transformer module, and proposes a new network architecture wavelet transformer (Wavelet-Tranformer). Compared with the traditional Fourier transform (FT), DWT does not need to assume the periodicity of the signal, and effectively avoids the introduction of high-frequency noise (such as Gibbs phenomenon) through multi-resolution analysis, showing stronger robustness in frequency domain information extraction. At the same time, the energy concentration of DWT further enhances the ability to capture low-frequency signals, providing key support for long-period time series modeling. Integrating discrete wavelet transform into the time series prediction model not only enhances its modeling ability for complex time series features, but also significantly improves the reliability and accuracy of the prediction results. This method provides a new technical idea and solution for time series prediction in diversified scenarios in the real world, and verifies its feasibility and superiority in humanoid robot learning tasks through a large number of experiments.
[0187] Furthermore, the present invention combines the residual network (ResNet) with the multi-scale attention module (EMA) in image detection, improves the EMA module, and proposes a frequency domain enhanced multi-scale attention module (FE-EMA). By integrating the improved EMA module into ResNet, the multi-scale feature expression ability of the model is significantly improved. Combined with the spatial and channel attention mechanisms, the model can capture local and global information more accurately, thereby showing higher performance and robustness in tasks such as target detection and classification. The improved FE-EMA module introduces frequency domain enhancement on the basis of EMA, uses wavelet transform to capture multi-scale information, and achieves richer feature expression. At the same time, by setting the α weight, the ratio of time domain and frequency domain features is adaptively adjusted, which further enhances the robustness of the model. In addition, the nonlinear activation mechanism effectively suppresses noise, allowing the model to extract fine-grained information more accurately. When processing complex data, this method can not only capture global information, but also fully retain key details, providing strong technical support for feature expression, and significantly improving the adaptability and expressiveness of the model in complex tasks.
[0188] The present invention builds a rich expert data set based on multimodal data integration, covering the action sequence, joint position, speed collected by the humanoid robot, as well as the angular velocity and linear acceleration recorded by the IMU (inertial measurement unit), monocular camera data, chassis linear velocity and angular velocity and other information. These multidimensional data are jointly input into the encoder and Transformer modules of the algorithm network to train an efficient environmental perception model. The model can accurately capture the key details of the operating environment. At the same time, the expert data set is highly scalable, providing strong environmental perception support for the robot to perform complex tasks, thereby significantly improving the accuracy and quality of task execution.
[0189] The present invention constructs a frequency domain enhanced wavelet transformer model by combining the Wavelet-Transformer network architecture with the characteristics of the FE-EMA module. The model realizes the deep fusion of multimodal and multidimensional data, and provides an innovative solution for the learning and operation of humanoid robots in complex task environments. In the processing of one-dimensional time series data, Wavelet-Transformer improves the ability to capture long-period signals and complex dynamic changes by collaboratively modeling the time domain and frequency domain features. Combined with the enhancement of frequency domain information, this architecture can effectively suppress noise interference and ensure the high accuracy and stability of the prediction results. In the processing of two-dimensional image data, the FE-EMA module extracts multi-scale features of the image through wavelet transform, and uses frequency domain enhancement technology to achieve a dynamic trade-off between global and local features. Combined with the nonlinear activation mechanism, the performance of the model in processing high-frequency details and low-frequency background information is further enhanced, so that the robot can show stronger robustness and flexibility in tasks such as target detection, image classification, and object recognition.
[0190] In addition, the expert dataset of the present invention not only covers multimodal perception data, but also provides high-quality annotations for network training, ensuring that the model can achieve efficient generalization in different complex scenarios. This integrated approach is not only applicable to traditional robot task scenarios, such as path planning and navigation, but can also be extended to more challenging scenarios, such as dynamic target tracking, fine grasping operations, and collaborative tasks.
[0191] Through a large number of experimental verifications, this method has achieved significant improvements in the learning efficiency, task adaptability and operation accuracy of humanoid robots. Compared with existing methods, this invention not only achieves innovative breakthroughs in model architecture design, but also provides higher reliability and practicality for the application of humanoid robots in actual scenarios. In the future, this technology can be further expanded to other robot platforms, promoting the widespread application of humanoid robots in industry, services, medical care and other fields, and opening up new paths for the development of intelligent robot technology.
[0192] Importantly, it should be noted that the construction and arrangement of the present application shown in a number of different exemplary embodiments are only exemplary. Although only a few embodiments are described in detail in this disclosure, it should be readily understood by those who refer to this disclosure that many modifications are possible, for example, the size, scale, structure, shape and proportion of various elements, and parameter values such as temperature, pressure, etc., mounting arrangements, use of materials, color, directional changes, etc., without substantially departing from the novel teachings and advantages of the subject matter described in the application. For example, the element shown as integrally formed can be composed of multiple parts or elements, the position of the element can be inverted or otherwise changed, and the nature or number or position of the discrete element can be changed or changed. Therefore, all such modifications are intended to be included in the scope of the present invention. The order or sequence of any process or method steps can be changed or reordered according to alternative embodiments. In the claims, any "device plus function" clause is intended to cover the structure of the execution function described herein, and is not only structurally equivalent but also equivalent structure. Without departing from the scope of the present invention, other substitutions, modifications, changes and omissions can be made in the design, operating conditions and arrangement of the exemplary embodiments. Therefore, the invention is not limited to a specific embodiment, but extends to numerous modifications still falling within the scope of the appended claims.
[0193] Furthermore, in order to provide a concise description of exemplary embodiments, all features of an actual embodiment may not be described, i.e., those features that are not relevant to the best mode presently contemplated for carrying out the invention or those features that are not relevant to implementing the invention.
[0194] It should be understood that in the development of any actual implementation, as in any engineering or design project, numerous implementation-specific decisions may be made. Such a development effort may be complex and time-consuming, but for those of ordinary skill having the benefit of this disclosure, the development effort will be a routine task of design, fabrication, and production without undue experimentation.
[0195] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A humanoid robot learning method based on frequency domain enhanced wavelet transformer model, characterized by: include, Collect the RGB image dataset, active arm joint information dataset, slave arm joint information dataset, chassis linear angular velocity dataset and IMU linear angular velocity dataset of the humanoid robot; Input the boom joint information dataset, chassis linear angular velocity dataset, IMU linear angular velocity dataset, and CLS classification head into the Transformer encoder module to obtain re-parameterized samples; The RGB image dataset processed by the FE-EMA module, the IMU linear angular velocity dataset, the joint information dataset, and the re-parameterized samples are input into the wavelet Transformer model, and the forward propagation calculation is performed to obtain the predicted action sequence; The loss value is calculated by the mean square error between the predicted action sequence and the actual action sequence, and the gradient is updated through back propagation to optimize the weight parameters of the wavelet Transformer model; Obtain a real-time data set of a humanoid robot, and input the re-parameterized samples set to 0 into the trained model. At the same time, pass in the real-time observation data of the humanoid robot to verify whether the actions of the humanoid robot can complete the previously trained tasks.
2. A humanoid robot learning method based on frequency domain enhanced wavelet transformer model as claimed in claim 1, characterized in that: The RGB image data set obtains image data through cameras respectively arranged on the head and left and right wrists of the humanoid robot.
3. A humanoid robot learning method based on frequency domain enhanced wavelet transformer model as claimed in claim 1 or 2, characterized in that: The active arm joint information data set, the passive arm joint information data set, the chassis linear angular velocity data set and the IMU linear angular velocity data set are acquired through ROS.
4. A humanoid robot learning method based on frequency domain enhanced wavelet transformer model as claimed in claim 3, characterized in that: The steps of inputting the slave arm joint information dataset, chassis linear angular velocity dataset, IMU linear angular velocity dataset and CLS into the Transformer encoder module to obtain the re-parameterized samples include: The slave arm joint information dataset, chassis linear angular velocity dataset, IMU linear angular velocity dataset and CLS are uniformly mapped to 512 dimensions through a linear layer for processing, and the action sequence is separately added with sinusoidal position encoding; Input the processed data into the Transformer encoder module; The data processed by the Transformer encoder module is trained by VAE to obtain the mean and variance of the sample distribution in the latent space and obtain the reparameterized samples.
5. A humanoid robot learning method based on frequency domain enhanced wavelet transformer model as claimed in claim 4, characterized in that: The action sequence includes the active arm joint information dataset and the chassis linear angular velocity dataset.
6. A humanoid robot learning method based on frequency domain enhanced wavelet transformer model as claimed in claim 4 or 5, characterized in that: The RGB image data set processed by the FE-EMA module, the IMU linear angular velocity data set, the active arm joint information data set, the slave arm joint information data set and the re-parameterized samples are input into the wavelet Transformer model. The steps of calculating and predicting the action sequence include: The RGB image dataset is passed through the FE-EMA module to extract the frequency domain features of the image; The processed image data set is input into the Transformer encoder module together with the IMU linear angular velocity data set, the active arm joint information data set, the slave arm joint information data set and the re-parameterized samples to obtain the time domain data; Input the time domain data into the DWT module for processing and then output the frequency domain data; At the same time, the time domain data and frequency domain data are fused into time-frequency data and input into the Transformer decoder module for calculation to generate a predicted action sequence.
7. A humanoid robot learning method based on frequency domain enhanced wavelet transformer model as claimed in claim 6, characterized in that: The steps of extracting the frequency domain features of an image using wavelet transform in the FE-EMA module are as follows: Step 1: Expand and reconstruct; expand the feature map by group along the batch and channel dimensions: B is the batch size, C is the number of channels, H and W are the height and width of the image respectively; Step 2: Convolution operation realizes wavelet decomposition, and calculates cA, cH, cV, cD for each set of feature maps: Step 3: Fusion of high-frequency information; in the FE-EMA module, the high-frequency information HighFreq = cH + cV + cD is added to the low-frequency component cA to obtain enhanced features: Enhanced_Wavelet=ResBlock(cA+HighFreq) Step 4: Fusion of time domain and frequency domain features; a dynamic weight α is used to control the fusion of time domain and frequency domain features: Fused_X=α·X+(1-α)·Enhanced_Wavelet The weight α is generated by global average pooling and a linear layer: α=σ(W·GAP(X)) where σ(·) is the Sigmoid function.
8. A humanoid robot learning method based on frequency domain enhanced wavelet transformer model as claimed in claim 7, characterized in that: The steps of inputting time domain data into the DWT module for processing and outputting frequency domain data include: Input time domain data and flatten the three-dimensional data into two-dimensional data; Perform discrete wavelet transform on the input data to decompose it into low-frequency and high-frequency components; Through specific feature extraction modules, low-frequency and high-frequency components are processed separately; The low-frequency and high-frequency component data are concatenated and feature fused through 1D convolution, and reconstructed back to the original shape through residual connection; Output frequency domain attention weights through the FC layer; Multiply the frequency domain attention weight by the input data and output the frequency domain data.
9. A humanoid robot learning method based on frequency domain enhanced wavelet transformer model as claimed in claim 7 or 8, characterized in that: The predicted action sequence and the actual action sequence are compared through the mean square error to calculate the loss. The steps of back-propagation optimization of the weight parameters of the wavelet Transformer model include: Calculate the loss between the predicted action sequence obtained by forward propagation and the actual action sequence; Perform back propagation and adjust the wavelet Transformer model parameters in reverse.
10. The humanoid robot learning method based on frequency domain enhanced wavelet transformer model as claimed in claim 8, characterized in that: The formula for the low-frequency component is: Among them, φ j,k [n] is the scaling function, which represents the low-frequency part of the signal, j represents the scale of decomposition, k represents the translation, and n represents the length of the sequence; The formula for high frequency component is: Among them, ψ j,k [n] is the wavelet function, which represents the high-frequency part of the signal, j represents the scale of decomposition, k represents the translation, and n represents the length of the sequence.