Humanoid robot welding pose adjustment method and system

By constructing the initial state vector and performing feature compression processing, combining the strategy optimization model and the multi-head attention mechanism, the welding posture of the humanoid robot is adjusted, which solves the problem of low accuracy of welding posture adjustment in the existing technology and achieves the adaptability and stability improvement of high-quality welding.

CN120326636BActive Publication Date: 2025-10-10INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510806676.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-10-10
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Existing technologies in humanoid robot welding lack the ability to perceive real-time changes in the environment, resulting in low accuracy in welding posture adjustment and difficulty in meeting high-quality welding requirements.

Method used

By constructing the initial state vector, performing temporal feature compression processing, and using the policy optimization model to generate an optimized action strategy, the welding posture of the humanoid robot is adjusted by combining the multi-head attention mechanism and the proximal policy optimization algorithm.

Benefits of technology

It significantly improves the adaptability and accuracy of the robot's welding movements, reduces the computational complexity, and improves the accuracy and stability of welding posture adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120326636B_ABST
    Figure CN120326636B_ABST
Patent Text Reader

Abstract

The application provides a humanoid robot welding pose adjustment method and system, and belongs to the technical field of artificial intelligence, comprising: constructing an initial state vector according to state information of the humanoid robot in a welding process; performing time sequence feature compression processing on the initial state vector to obtain a compressed state vector; inputting the compressed state vector into a strategy optimization model to obtain an optimized action strategy output by the strategy optimization model; wherein the strategy optimization model is trained according to state information samples; generating discrete joint rotation angle values of each joint of the humanoid robot according to the optimized action strategy; and adjusting the welding pose of the humanoid robot according to each discrete joint rotation angle value. The application realizes self-adaptive adjustment and accurate control of the welding pose of the humanoid robot in a dynamic welding environment by introducing a strategy optimization model based on an optimized long short-term memory network and a multi-head attention mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for adjusting the welding posture of a humanoid robot. Background Art

[0002] In the current technological context, with the increasing application of humanoid robots in high-precision welding scenarios, the environmental complexity faced by welding operations has increased significantly, including factors such as nonlinear changes in weld trajectories, dynamic interference from obstacles, and posture constraints of the robot itself.

[0003] Existing technologies rely heavily on rule-based or traditional sensor feedback control methods, typically employing preset path planning or static obstacle avoidance strategies. These methods lack the ability to perceive real-time environmental changes, preventing the robot from flexibly adjusting its welding posture based on its current state. Furthermore, existing control systems typically process state information using fixed-length static inputs, making them unsuitable for scenarios where the number of obstacles changes or the state dimensions fluctuate. This results in low accuracy in the welding posture adjustment of humanoid robots, making it difficult to meet the demands of high-quality welding operations. Summary of the Invention

[0004] The present invention provides a method, system, electronic device and storage medium for adjusting the welding posture of a humanoid robot, which are used to solve the defects in the existing technology and improve the accuracy of adjusting the welding posture of the humanoid robot.

[0005] The present invention provides a method for adjusting the welding posture of a humanoid robot, comprising the following steps:

[0006] Constructing an initial state vector according to the state information of the humanoid robot during the welding process;

[0007] Performing time series feature compression processing on the initial state vector to obtain a compressed state vector;

[0008] Inputting the compressed state vector into a policy optimization model to obtain an optimized action strategy output by the policy optimization model; wherein the policy optimization model is trained based on state information samples;

[0009] generating, according to the optimized motion strategy, a discretized joint rotation angle value of each joint of the humanoid robot;

[0010] The welding posture of the humanoid robot is adjusted according to each of the discretized joint rotation angle values.

[0011] According to a method for adjusting the welding posture of a humanoid robot provided by the present invention, the strategy optimization model includes a strategy network and a value network, and the step of inputting the compressed state vector into the strategy optimization model to obtain an optimized action strategy output by the strategy optimization model includes:

[0012] input the compressed state vector into the policy network and the value network respectively to obtain an initial action policy output by the policy network and an initial value estimate output by the value network;

[0013] According to the initial action policy and the initial value estimate, an advantage function is calculated.

[0014] The different dimension features of the compressed state vector are differentially weighted through a multi-head attention mechanism to obtain a differential weight result.

[0015] According to the advantage function and the differential weight result and using a proximal policy optimization algorithm, the initial action policy is pruned and updated to obtain the optimized action policy.

[0016] According to the human-shaped robot welding pose adjustment method provided by the application, the training method of the policy network and the value network is also provided.

[0017] In the simulation environment, state information samples of the human-shaped robot welding process are collected, initial state vector samples are constructed, and the initial state vector samples are subjected to time series feature compression processing to obtain compressed state vector samples.

[0018] The compressed state vector samples are input into the to-be-trained policy network and the to-be-trained value network respectively to obtain an action policy of the current iteration and a corresponding state value estimate.

[0019] According to the action policy, an action is performed on the simulation environment to obtain a reward value fed back by the simulation environment, and an actual return value is calculated according to the reward value.

[0020] According to the action policy of the current iteration, the action policy of the previous iteration, the state value estimate and the actual return value, a policy loss function and a value loss function are constructed; wherein the policy loss function is used to measure the difference between the action policy of the current iteration and the action policy of the previous iteration, and the value loss function is used to measure the difference between the state value estimate and the actual return value.

[0021] According to the policy loss function and the value loss function, gradient update is alternately performed on the to-be-trained policy network and the to-be-trained value network, and iteration is repeated in the training process until the action policy and the state value estimate converge, thereby obtaining the trained policy network and the value network.

[0022] According to the human-shaped robot welding pose adjustment method provided by the application, the initial state vector is constructed according to the state information of the human-shaped robot in the welding process, which comprises:

[0023] Using the real-time joint angle value of each joint of the humanoid robot as a joint angle state component;

[0024] Taking the three-dimensional coordinates of the base of the humanoid robot in the world coordinate system as the base posture state component;

[0025] Taking the three-dimensional coordinates of the weld target trajectory in the world coordinate system as the weld trajectory state component;

[0026] The distance between the tool center point of the humanoid robot and the closest point of the weld target trajectory is used as a weld tracking distance state component;

[0027] The minimum distance from any key point of the humanoid robot to the surface of each obstacle is taken as the obstacle distance state component;

[0028] The joint angle state component, the base posture state component, the weld trajectory state component, the weld tracking distance state component and the obstacle distance state component are spliced ​​in a predetermined order to obtain the initial state vector.

[0029] According to a method for adjusting the welding posture of a humanoid robot provided by the present invention, the method of performing time series feature compression processing on the initial state vector to obtain a compressed state vector includes:

[0030] The initial state vector is input into a feature compression model to obtain a compressed state vector output by the feature compression model; wherein the feature compression model is obtained after optimizing the network structure of the long short-term memory network.

[0031] According to a method for adjusting the welding posture of a humanoid robot provided by the present invention, the feature compression model includes multiple layers of long short-term memory network units connected in sequence, the channel dimensions of each layer of long short-term memory network units are scaled by a proportional factor, and the operations of the input gate, forgetting gate and output gate in each long short-term memory network unit are fused into a single convolution operation to generate a bottleneck feature map with N channels; the feature compression model also includes a memory unit, which is used to screen out obstacle distance state components that do not meet preset standards from all the obstacle distance state components contained in the initial state vector according to the bottleneck feature map, and output the compressed state vector.

[0032] According to a method for adjusting the welding posture of a humanoid robot provided by the present invention, adjusting the welding posture of the humanoid robot according to each of the discretized joint rotation angle values ​​includes:

[0033] Performing real-time trajectory interpolation processing on each of the discretized joint rotation angle values ​​to obtain an interpolated trajectory of the corresponding joint;

[0034] Real-time control instructions are generated according to all the interpolation trajectories, and the welding pose of the humanoid robot is adjusted according to the real-time control instructions.

[0035] The application further provides a humanoid robot welding pose adjustment system, comprising the following modules:

[0036] The first processing module is configured to construct an initial state vector according to state information of the humanoid robot in a welding process;

[0037] The second processing module is configured to perform time sequence feature compression processing on the initial state vector to obtain a compressed state vector;

[0038] The third processing module is configured to input the compressed state vector into a policy optimization model to obtain an optimized action policy output by the policy optimization model; wherein the policy optimization model is trained according to state information samples;

[0039] The fourth processing module is configured to generate discrete joint rotation angle values of each joint of the humanoid robot according to the optimized action policy;

[0040] The fifth processing module is configured to adjust the welding pose of the humanoid robot according to each discrete joint rotation angle value.

[0041] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned humanoid robot welding pose adjustment method when executing the program.

[0042] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the above-mentioned humanoid robot welding pose adjustment method.

[0043] In summary, the one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0044] By constructing an initial state vector based on the humanoid robot's state information during the welding process and performing time-series feature compression on this initial state vector, a compressed state vector is obtained that represents key factors such as changes in the welding environment, weld trajectory location, and the robot's own posture. This effectively addresses the issues of non-fixed state dimensions and information redundancy, making the model input structure more compact, significantly reducing computational complexity, and improving the generalization and stability of the subsequent policy model. By inputting the compressed state vector into a policy optimization model, the optimized action strategy is output by the policy optimization model. This enables the system to optimize action decisions in real time based on current state characteristics, improving the robot's welding action's adaptability to environmental changes, and effectively enhancing decision-making efficiency and action accuracy. By generating discrete joint rotation angle values ​​for each joint of the humanoid robot based on the optimized action strategy, the continuous action space is converted into a finite, standardized set of discrete action instructions, effectively reducing the complexity of action control and improving the real-time and stability of the robot's action execution. By adjusting the humanoid robot's welding posture based on each discrete joint rotation angle value, the robot can precisely follow the target weld path with a smooth, continuous motion trajectory, significantly improving the accuracy of the humanoid robot's welding posture adjustment. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0046] Figure 1 This is one of the flow charts of the method for adjusting the welding posture of a humanoid robot provided by the present invention.

[0047] Figure 2 This is the second flow chart of the humanoid robot welding posture adjustment method provided by the present invention.

[0048] Figure 3 This is the third flow chart of the method for adjusting the welding posture of a humanoid robot provided by the present invention.

[0049] Figure 4 This is the fourth flow chart of the method for adjusting the welding posture of a humanoid robot provided by the present invention.

[0050] Figure 5 This is the fifth flow chart of the method for adjusting the welding posture of a humanoid robot provided by the present invention.

[0051] Figure 6 It is a structural schematic diagram of the humanoid robot welding posture adjustment system provided by the present invention.

[0052] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0053] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field without making creative efforts based on the embodiments of the present invention are within the scope of protection of the present invention.

[0054] It should be noted that, in the description of the present invention, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. The orientation or positional relationship indicated by the terms "upper" and "lower" is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the system or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.

[0055] The terms "first," "second," and so forth, used herein are used to distinguish similar objects, not to describe a specific order or precedence. It should be understood that such terms are interchangeable where appropriate, allowing embodiments of the present invention to be implemented in an order other than that illustrated or described herein. Furthermore, the terms "first," "second," and so forth generally distinguish objects of a single type, and do not limit the number of objects. For example, the first object may be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the connected objects.

[0056] The following combination Figure 1-Figure 7 The present invention describes a method, system, electronic device and storage medium for adjusting the welding posture of a humanoid robot.

[0057] Figure 1 This is one of the flow charts of the method for adjusting the welding posture of a humanoid robot provided by the present invention, such as Figure 1As shown, including but not limited to the following steps:

[0058] Step 101: Construct an initial state vector based on the state information of the humanoid robot during the welding process.

[0059] In this embodiment, the purpose of step 101 is to comprehensively and accurately acquire multiple key parameters that characterize the humanoid robot's welding process state, thereby constructing an initial state vector for deep reinforcement learning input. Because the robot faces multiple uncertainties during welding, such as complex spatial trajectories, dynamic obstacles, and posture disturbances, it is necessary to systematically integrate these core information that influence the robot's welding posture adjustment to provide a high-quality input data foundation for subsequent strategy learning and decision-making.

[0060] Specifically, the construction of the initial state vector includes the following key state components: joint angle state component, base posture state component, weld trajectory state component, weld tracking distance state component, and obstacle distance state component. Therefore, the initial state vector can be expressed as:

[0061]

[0062] in, That is, the joint angle state component, m is the number of joints of the humanoid robot; That is the base seat posture state component; That is the weld trajectory state component; That is the weld tracking distance state component, That is the obstacle distance state component, and n is the number of obstacles.

[0063] In the following embodiments, a method for obtaining each key state component according to the state information of the humanoid robot during the welding process and a method for constructing an initial state vector according to each key state component will be described in detail.

[0064] In a possible implementation, step 101 specifically includes the following steps:

[0065] Step 201: The real-time joint angle value of each joint of the humanoid robot is used as a joint angle state component.

[0066] Step 202: The three-dimensional coordinates of the base of the humanoid robot in the world coordinate system are used as the base posture state components.

[0067] Step 203: The three-dimensional coordinates of the weld target trajectory in the world coordinate system are used as the weld trajectory state component.

[0068] Step 204: The distance between the tool center point of the humanoid robot and the nearest point of the weld target trajectory is used as the weld tracking distance state component.

[0069] Step 205: The minimum distance from any key point of the humanoid robot to the surface of each obstacle is used as the obstacle distance state component.

[0070] Step 206: The joint angle state component, the base posture state component, the weld trajectory state component, the weld tracking distance state component, and the obstacle distance state component are concatenated in a predetermined order to obtain an initial state vector.

[0071] In this embodiment, steps 201 to 206 are used to specifically implement the initial state vector construction process of step 101, which aims to systematically collect and encode multi-source state information of the humanoid robot body state, welding task objectives and welding environment.

[0072] First, during the robot welding process, the real-time joint angle values ​​of each joint of the humanoid robot are obtained as the joint angle state components , which is the basis for achieving motion control and posture reasoning. Since humanoid robots usually have multiple degrees of freedom, the joint angles directly determine the position and posture of the end welding gun. The accurate expression of joint information is of great significance for establishing the mapping relationship between motion and posture.

[0073] Secondly, the three-dimensional coordinates of the humanoid robot base in the world coordinate system are used as the base posture state components This component is introduced because welding trajectories are often distributed in three-dimensional space. The robot's overall position affects its workspace, path planning, and obstacle avoidance. By acquiring the robot's base position in real time, we can achieve global control of the robot's overall motion and enhance the accuracy and consistency of path planning.

[0074] Furthermore, in order to ensure that the robot can perform actions closely following the weld trajectory during the welding process, the three-dimensional coordinates of the weld target trajectory in the world coordinate system need to be used as the weld trajectory state component The introduction of this component enables the model to have the ability to perceive the geometric shape and spatial distribution of the target trajectory, thereby guiding the robot welding gun end to perform high-precision tracking operations.

[0075] In addition, the key control quantity in the welding process also includes the distance between the robot tool center point (TCP) and the closest point of the weld target trajectory, which is constructed as the weld tracking distance state component This state component can be used to evaluate the degree of fit between the welding gun and the weld in real time. It is an important indicator for measuring whether the current welding effect meets the process requirements and also provides feedback for fine-tuning the control system.

[0076] In the dynamic human-machine collaboration scenario, the obstacle information in the welding environment has a direct impact on the feasibility of the welding path. Therefore, it is necessary to further extract the minimum distance from any key point of the humanoid robot to the surface of each obstacle and use it as the obstacle distance state component. This component can quantify the spatial relationship between the robot and the environmental interference, providing a numerical basis for the safety constraints of subsequent action strategies.

[0077] After extracting the five state components, the joint angle state components, base posture state components, weld trajectory state components, weld tracking distance state components, and obstacle distance state components are concatenated in a predetermined order to generate a standardized, dimensionally fixed initial state vector. This initial state vector maintains information integrity and provides standardized, high-quality input for subsequent temporal feature compression and policy learning modules, effectively improving the deep reinforcement learning model's perceptual expression capabilities and control decision-making efficiency in complex welding scenarios.

[0078] Step 102: Perform time series feature compression processing on the initial state vector to obtain a compressed state vector.

[0079] In this embodiment, step 102 is intended to perform time-series feature compression processing on the initial state vector constructed in step 101 to generate a compressed state vector with a more compact structure and more effective expression. The reason for implementing this step is that the environment in which the robot is located during welding is dynamic and uncertain, especially the number and position of obstacles may change at any time, resulting in the dimension of the initial state vector being inconsistent at different times. If such a high-dimensional and non-fixed-structure vector is directly input into the strategy model, it may cause unstable model learning, weak generalization ability, and even the training process to fail to converge. To this end, this embodiment introduces a feature compression model to perform unified structural compression on the initial state vector to ensure that the input of the subsequent strategy optimization model has a fixed length and trainability.

[0080] In a possible implementation, step 102 specifically includes the following steps:

[0081] The initial state vector is input into a feature compression model to obtain a compressed state vector output by the feature compression model. The feature compression model is obtained by optimizing the network structure of the long short-term memory network. The feature compression model includes multiple layers of long short-term memory network units connected in sequence. The channel dimension of each layer of long short-term memory network units is scaled by a scaling factor. The input gate, forget gate, and output gate operations in each long short-term memory network unit are fused into a single convolution operation to generate a bottleneck feature map with N channels. The feature compression model also includes a memory unit, which is used to filter out obstacle distance state components that do not meet preset standards from all obstacle distance state components contained in the initial state vector based on the bottleneck feature map, and output a compressed state vector.

[0082] During robotic welding, the initial state vector often contains numerous components affected by dynamic environmental changes. In particular, the number of obstacle distance state components fluctuates with the number of obstacles in the scene. Directly inputting this variable-length data into the policy model will result in inconsistent input dimensions, severely impacting the convergence and generalization capabilities of the policy and value networks. Therefore, it is necessary to introduce a method that can model time series at the input stage and output fixed-length vectors. This involves using a feature compression model to uniformly compress and filter information in the initial state vector that is valuable to the policy.

[0083] During specific implementation, the initial state vector generated in step 101 is first input into the feature compression model. The feature compression model is obtained by optimizing the network structure of the long short-term memory network. By introducing a scaling factor into the network structure of the long short-term memory network, the channel dimension is scaled according to a preset ratio in each layer, thereby reducing the number of model parameters while retaining key channel information and improving operational efficiency. In addition, the feature compression model optimizes the original long short-term memory network structure through computational fusion, that is, integrating the independent matrix calculations of the input gate, forget gate, and output gate into a unified convolution operation to generate a bottleneck feature map with N channels, which not only reduces computational overhead but also improves network depth and expression capabilities.

[0084] After generating the bottleneck feature map, the feature compression model further uses its internal memory unit to filter the obstacle distance state components in the initial state vector. Based on the importance of each channel feature in the bottleneck feature map, the feature compression model selectively processes all obstacle-related components, retaining those that meet preset criteria and have a substantial impact on welding pose adjustment, while discarding the remaining redundant information. Ultimately, the feature compression model outputs a fixed-length, semantically refined compressed state vector, which serves as the standard input for the policy optimization model.

[0085] Step 103: input the compressed state vector into a policy optimization model to obtain an optimized action policy output by the policy optimization model; wherein the policy optimization model is trained according to the state information sample.

[0086] In this embodiment, the core purpose of step 103 is to generate an optimized action policy that can be used to control the welding pose of the humanoid robot from the optimized compressed state vector, so as to realize intelligent decision and adaptive adjustment of the robot action in the dynamic welding environment.

[0087] During the welding operation, the state of the robot is disturbed by various factors, including changes in the weld seam trajectory, obstacle interference, and attitude deviation. These factors represent nonlinear dynamic changes in the state space. In order to adapt to such changes and achieve the extraction of a high-quality control policy, the embodiment uses a policy optimization model based on deep reinforcement learning. The policy optimization model jointly decides according to the pre-trained policy network and value network, can extract the most valuable information for the task from the compressed state representation, and output the optimal action scheme accordingly.

[0088] In one possible implementation, the policy optimization model includes a policy network and a value network; and step 103 specifically includes the following steps:

[0089] Step 301: input the compressed state vector into the policy network and the value network respectively to obtain an initial action policy output by the policy network and an initial value estimate output by the value network.

[0090] Under the reinforcement learning framework, the policy network is responsible for outputting the action to be taken for the current state, while the value network is used to evaluate the long-term return of the state. By inputting the same compressed state vector into the two network models, parallel interpretation and different angle modeling of the same state can be achieved, making the policy optimization more comprehensive and having feedback constraint capability. Especially in a dynamic welding environment, the state changes are complex, and if only the policy network is used to output the action without the supervision of the value evaluation, it is easy to cause the policy to diverge or converge slowly.

[0091] In the specific implementation process, first, the compressed state vector output by the feature compression model is input into the policy network, and the policy network processes the state features through a multi-layer perceptron structure to output an initial action policy under the current state, representing the joint rotation direction and amplitude that the robot should take in that state. Subsequently, the same compressed state vector is synchronously input into the value network, and the value network performs regression calculation on the state through a similar deep structure to output an initial value estimate of the state, which reflects the expected cumulative return that the state can bring under the current policy.

[0092] Step 302: calculate the advantage function according to the initial action policy and the initial value estimate.

[0093] Robotic welding tasks often occur in environments with frequent dynamic interference and posture disturbances. Directly using the original state value or action value for policy updates can be susceptible to policy estimation bias or single-step reward fluctuations, resulting in model instability or slow convergence. Therefore, the advantage function is introduced in step 302. In deep reinforcement learning, the advantage function is an important metric for measuring whether a specific action performs better than the average policy. It not only improves the directionality and stability of policy updates but also effectively suppresses fluctuations during the policy update process. It is one of the core computational units in the proximal policy optimization (PPO) algorithm. By combining the outputs of the policy network and the value network and introducing a baseline value as a reference, the advantage function can more accurately measure the relative merits of an action, thereby enhancing the ability to identify the direction of policy optimization.

[0094] In specific implementation, the system first executes the behavior corresponding to the initial action strategy output by the policy network in step 301, combined with the current state information, and obtains an immediate reward value in a simulation environment or real-world execution. Simultaneously, based on a certain discount factor, these immediate rewards are accumulated to form an actual return value, which serves as an estimate of future returns. This actual return value is then compared with the initial value estimate output by the value network. The difference between the two constitutes the advantage function value of the current state-action pair.

[0095] Step 303: Differentiated weighting is performed on the different dimensional features of the compressed state vector through a multi-head attention mechanism to obtain differentiated weight results.

[0096] The state vector in the robot welding process is not only of high dimension, but also contains different types of data such as joint angles, base posture, weld trajectory, weld tracking distance, and distances to multiple obstacles. In the compressed state vector, the importance of each dimensional feature to the final action decision varies. For example, weld tracking distance and obstacle distance have a more direct impact on the triggering of obstacle avoidance actions, while some dimensions may be redundant or noisy. Traditional neural networks often treat all dimensions equally when processing such state inputs, which can easily lead to unclear focus targets during strategy optimization, affecting training efficiency and strategy performance. Therefore, the introduction of a multi-head attention mechanism in step 303 can simulate the attention mechanism of humans in the decision-making process, focusing model resources on more critical feature dimensions, thereby improving the accuracy and convergence speed of the strategy output, and improving the perception and utilization efficiency of the strategy network for key state information.

[0097] In practice, the system first performs a linear transformation on the compressed state vector to generate multiple feature subspace representations, each corresponding to an attention head. Each attention head then calculates the weight coefficients for the relevance of different feature dimensions in the state vector to the current task. These weight coefficients are then used to perform a weighted combination of these feature dimensions, resulting in multiple differentiated attention outputs. Finally, the system concatenates or sums the outputs of all attention heads to generate a differentiated, weighted state representation, which is used to guide the policy network to output a more targeted, optimized action strategy.

[0098] Step 304: The initial action strategy is trimmed and updated according to the advantage function and the differentiated weight results and the proximal strategy optimization algorithm is used to obtain an optimized action strategy.

[0099] In humanoid robot welding tasks, environmental disturbances and state changes are complex, and directly using the initial action strategy can lead to unstable behavior or decision-making errors. Traditional policy update methods can easily cause policy divergence if the update amplitude is too large; if the update amplitude is too small, it may fall into a local optimum. To this end, the clipping mechanism in the PPO algorithm can effectively limit the amplitude of each policy update, preventing the strategy from deviating too far from the original direction. In combination with the differentiated weight results output by the multi-head attention mechanism, the responsiveness to key features is further enhanced, thereby obtaining an optimized action strategy with higher execution effectiveness and policy stability, ensuring that the generated optimized action strategy strikes a good balance between performance and stability.

[0100] In the specific implementation process, first, based on the advantage function constructed in step 302 , combined with the attention mechanism output in step 303, the initial action strategy in the current state With the old strategy The ratio between Perform cropping and construct the loss function as follows:

[0101]

[0102] Among them, the clipping factor is a preset threshold that controls the maximum magnitude of policy updates. The above loss function amplifies responses to key state features through differentiated attention weighting, allowing policy updates to focus more on high-weight dimensions such as seam tracking error and obstacle proximity.

[0103] On this basis, the system performs gradient backpropagation and updates on the policy network parameters, thereby converging the policy toward a more optimal direction while preserving the stability of the original policy structure. Once the update is complete, an optimized motion policy for the humanoid robot is obtained, which can be used for real-time control.

[0104] In one possible implementation, the method also includes a policy network and a value network training process. The purpose is to continuously optimize the policy model's behavioral decision-making and state assessment capabilities by collecting and utilizing training data in a simulation environment, thereby obtaining a control strategy with good generalization performance and stability in actual welding tasks. This training process relies on a proximal policy optimization algorithm and combines previously designed key structures such as compressed state vectors, advantage functions, and multi-head attention mechanisms to effectively support the humanoid robot's adaptive welding posture adjustment in complex environments. Specifically, the policy network and value network training method includes the following steps:

[0105] Step 401: collecting state information samples of the humanoid robot welding process in a simulation environment, constructing initial state vector samples, and performing time series feature compression processing on the initial state vector samples to obtain compressed state vector samples.

[0106] To train control strategies capable of practical deployment, a large amount of sample data covering diverse scenarios and state changes is required. Because welding processes in real environments involve high temperatures, strong light, and other factors, directly collecting training data poses security and cost issues. Therefore, a high-fidelity simulation environment is used instead of real-world scenarios for data collection. This not only generates samples on a large scale but also allows for modeling and recording the humanoid robot's state under various interference conditions, significantly improving training efficiency and coverage.

[0107] During the implementation process, a simulated welding environment for reinforcement learning training was first constructed. This environment had the same geometry, dynamic parameters, weld trajectory model, and obstacle layout as the actual robot operating conditions. Within this environment, the humanoid robot was controlled to perform welding tasks, and state information was collected in real time at every moment. This included five state components: joint angles, base posture, weld trajectory position, weld tracking distance, and the minimum distance between the robot's key points and obstacles. These state components were then spliced ​​into a standardized initial state vector sample according to a predetermined format.

[0108] To address the issue of inconsistent state vector dimensions caused by the changing number of obstacles, the system further feeds the initial state vector samples into a feature compression model for processing. The optimized LSTM network significantly improves computational efficiency while maintaining time series modeling capabilities through structural improvements such as scaling channel dimensions, gated computation fusion, and bottleneck feature map generation. This internal memory mechanism eliminates redundant obstacle distance state components, ultimately outputting fixed-length compressed state vector samples.

[0109] Step 402: Input the compressed state vector samples into the policy network to be trained and the value network to be trained respectively to obtain the action strategy of the current iteration and the corresponding state value estimation.

[0110] In the implementation process, first, the compressed state vector sample obtained in step 401 is input into the policy network to be trained in sequence. The policy network generally adopts a multi-layer perceptron structure, and after receiving the compressed state vector, outputs an action policy distribution corresponding to the current state, representing the selection probability of each candidate action of the robot in the state. In this embodiment, the form of the action policy output is a probability prediction of the discretized rotation angle of each joint of the humanoid robot, which is used to guide the subsequent action execution in the simulation environment.

[0111] At the same time, the same compressed state vector sample is input into the value network to be trained. The value network is also constructed based on a deep neural structure, and outputs a scalar value as the state value estimation in the current state after encoding the state features, which is used to predict the cumulative expected return that can be obtained in the existing policy. This predicted value will be compared with the actual return obtained through interaction in the next step to measure the prediction accuracy of the value network.

[0112] Step 403: According to the action policy, perform the action in the simulation environment to obtain the reward value fed back by the simulation environment, and calculate the actual return value according to the reward value.

[0113] In the implementation, the system first selects the corresponding control action in a sampling manner according to the action policy output by the policy network in step 402, i.e., determines the discretized rotation angle value of each joint of the humanoid robot. Then, the humanoid robot is controlled to perform the control action in the simulation environment, and real-time feedback of various index data of the welding process is obtained from the simulation model, such as the weld trajectory tracking error, the minimum distance between the robot and the obstacle, the action smoothness, etc.

[0114] According to the above feedback information, the system calculates the immediate reward value according to the preset reward function construction rule. The reward function generally includes multiple weighted terms for measuring the positive promotion degree of the current action to the task goal and the inhibition effect of the potential risk. For example, a positive reward is given when the weld tracking error is small, and a penalty is given otherwise; when the robot approaches the obstacle, an additional negative reward is applied to guide the policy to generate a safe action.

[0115] After obtaining the immediate reward value, the system weights and accumulates the current reward and the reward of the subsequent state by introducing a discount factor to obtain the actual return value. The actual return value is used to measure the comprehensive income that can be obtained by executing the action from the current state under the current policy in the entire future time period, and is the basis for subsequent value estimation error calculation.

[0116] Step 404: According to the action policy of the current iteration, the action policy of the last iteration, the state value estimation and the actual return value, construct the policy loss function and the value loss function; wherein the policy loss function is used to measure the difference between the action policy of the current iteration and the action policy of the last iteration, and the value loss function is used to measure the difference between the state value estimation and the actual return value.

[0117] Specifically, in the embodiment of step 404, the method of constructing the policy loss function is consistent with the method of constructing the loss function in step 404, that is, first, according to the state value estimation and the actual return value , the advantage function is constructed, and then the ratio between the action policy of the current iteration and the action policy of the last iteration is combined to construct the policy loss function:

[0118]

[0119] Wherein, the clipping coefficient is a preset threshold value for controlling the maximum amplitude of policy update. The policy loss function controls the range of change of policy probability by clipping, ensuring that each policy adjustment is within a controlled range, thereby improving training stability.

[0120] At the same time, the value loss function is constructed to optimize the prediction performance of the value network, which usually adopts the form of mean square error:

[0121]

[0122] Wherein, represents the state value estimated by the current value network parameter . The value loss function can quantify the estimation error of the current network for the true long-term return, and is minimized as the objective function in the subsequent training process.

[0123] Step 405: According to the policy loss function and the value loss function, perform gradient update on the trained policy network and the trained value network alternately, and repeat the iteration in the training process until the action policy and the state value estimation converge, to obtain the trained policy network and the trained value network.

[0124] In the reinforcement learning framework, both the policy network and the value network are neural network structures with trainable parameters. These parameters must be continuously adjusted through backpropagation and gradient optimization algorithms to continuously approach optimal performance for a given task objective. Step 404 clearly defines the optimization objectives for the policy and value networks: minimizing policy loss and value loss. Parameter updates in this step effectively apply these optimization objectives to the model structure itself, thereby effectively improving policy capabilities. Furthermore, multiple rounds of iterative training enable the network to achieve more robust policy generalization and accurate decision-making in the face of diverse input conditions.

[0125] During implementation, the policy network is first backpropagated and its weights updated using a gradient descent optimization algorithm (such as Adam or SGD) based on the constructed policy loss function. This allows the action policy it outputs to maintain close proximity to the old policy while maximizing the advantage function value in the current state. Simultaneously, based on the value loss function, the value network parameters are optimized through backpropagation, so that its output state value estimate more closely matches the reward value obtained from actual environmental interactions. This optimization process typically employs a small batch iterative update approach, selecting batches of compressed state vector samples, action policies, and actual reward data from the training set for training. This improves training efficiency and reduces the impact of gradient fluctuations on the model.

[0126] After each round of gradient update is completed, the system will re-execute steps 401 to 404, that is, collect new state samples in the simulation environment, calculate new strategies and value outputs, obtain reward feedback and reconstruct the loss function, and continue to iterate until the strategy loss and value loss during training are both stable, indicating that the action strategy output by the strategy network is consistent with the actual reward, and the value network has the ability to accurately evaluate rewards.

[0127] Through this step, the system not only efficiently trains the deep policy structure but also, through multiple rounds of closed-loop iteration, improves the policy's decision-making ability and stability in dynamic welding tasks. The resulting policy network and value network can be deployed in a practical control system to guide the humanoid robot in making dynamic, stable, and highly precise posture adjustments based on real-time conditions during the welding process. This improves welding quality, obstacle avoidance, and operational safety, meeting the stringent requirements placed on autonomous welding systems in complex industrial environments.

[0128] Step 104: Generate discretized joint rotation angle values ​​of each joint of the humanoid robot according to the optimized motion strategy.

[0129] Optimized motion strategies are typically expressed as continuous motion vectors, encompassing the rotation angle trends and amplitudes for each joint. However, given the requirements for control accuracy, response speed, and safety in practical robotic control systems, directly employing high-dimensional continuous motion increases system computational overhead and can lead to high-frequency fine-tuning, resulting in unstable execution. To address this, this embodiment discretizes the strategy output, mapping continuous joint motions into a finite set of discrete rotation angle values ​​with a preset step size. This reduces the complexity of the motion space and improves the real-time and reliability of system control.

[0130] During the specific implementation process, the system first reads the rotation angle change instruction corresponding to each joint in the optimized action strategy. Then, based on the limited action change range, the action value of each joint is discretized. The discrete interval is usually set to , and is divided into equally spaced intervals with a fixed step size (for example, 0.2° or 0.5°). In this way, the continuous action space is quantized into a limited set of action candidates, forming a standard set of discretized joint actions. Based on the action trend output by the strategy, the system selects the closest rotation angle in this discrete set as the actual control value, ultimately forming the discrete rotation angle command for each joint of the robot.

[0131] By implementing this step, on the one hand, the computational complexity of strategy execution can be significantly reduced, the robot's response speed can be accelerated, and the real-time control requirements of the welding process can be met; on the other hand, the introduction of discrete actions can also help enhance the robustness and stability of the strategy, avoiding posture jumps or redundant adjustments caused by subtle movement fluctuations.

[0132] Step 105: Adjust the welding posture of the humanoid robot according to each discretized joint rotation angle value.

[0133] Although the strategy model has output the optimal action plan adapted to the current state and discretized it into controllable angle commands, the robot's execution system still needs to dynamically interpolate and smooth these discrete angle values ​​to generate a continuous, executable joint motion trajectory, thereby ensuring the physical feasibility and high-precision continuity of the welding action. Especially in the control of multi-degree-of-freedom humanoid robots, directly using discrete angle commands as joint target pose input can easily cause oscillations or sudden jumps, affecting the stability of the welding process. Therefore, the purpose of step 105 is to adjust the posture of each joint of the humanoid robot based on the discrete joint rotation angle values ​​generated in step 104 to achieve precise control of the welding posture.

[0134] In a possible implementation, step 105 specifically includes the following steps:

[0135] Step 501: Perform real-time trajectory interpolation processing on each discretized joint rotation angle value to obtain an interpolated trajectory of the corresponding joint.

[0136] Step 502: Generate real-time control instructions according to all interpolation trajectories, and adjust the welding posture of the humanoid robot according to the real-time control instructions.

[0137] The purpose of implementing step 501 is to convert the discretized joint rotation angle values ​​generated in step 104 into a continuous and smooth motion trajectory to meet the requirements of the robot execution system for the continuity of control input. Specifically, the system uses the current joint state and the target discrete angle value as boundary conditions, and adopts methods such as spline interpolation, cubic polynomial interpolation or quintic polynomial interpolation to construct a continuous interpolation trajectory curve for each joint. The trajectory generation process takes into account execution conditions such as the robot's kinematic constraints, maximum speed and acceleration limits to ensure that the trajectory is mathematically continuous and physically controllable. After the interpolation is completed, the system can obtain a continuous position sequence of each joint over a period of time, which serves as the basis for subsequent generation of control instructions.

[0138] In step 502, based on the interpolated trajectory generated in step 501, the system generates real-time control instructions for the humanoid robot and drives the execution system to implement posture adjustments. Specifically, the system interprets the interpolated position of each joint at that moment as a control target within a certain time period (e.g., milliseconds). This system then sends position or speed control signals to the robot's underlying servo drive module, achieving refined control of each joint. This ensures the continuity of the humanoid robot's posture changes and the kinematic stability of its mechanical structure, enabling high-precision adjustment of the humanoid robot's welding posture.

[0139] Reference Figure 6 , Figure 6 This is a structural diagram of the humanoid robot welding posture adjustment system provided by the present invention, the system includes:

[0140] A first processing module is used to construct an initial state vector according to state information of the humanoid robot during the welding process;

[0141] The second processing module is used to perform time series feature compression processing on the initial state vector to obtain a compressed state vector;

[0142] a third processing module, configured to input the compressed state vector into a policy optimization model to obtain an optimized action policy output by the policy optimization model; wherein the policy optimization model is trained based on the state information sample;

[0143] a fourth processing module, configured to generate a discretized joint rotation angle value of each joint of the humanoid robot according to the optimized motion strategy;

[0144] The fifth processing module is used to adjust the welding posture of the humanoid robot according to each discretized joint rotation angle value.

[0145] In a possible implementation, the third processing module is further configured to:

[0146] Input the compressed state vector into the policy network and the value network respectively to obtain the initial action strategy output by the policy network and the initial value estimate output by the value network;

[0147] Calculate the advantage function based on the initial action strategy and initial value estimate;

[0148] Through the multi-head attention mechanism, the different dimensional features of the compressed state vector are differentially weighted to obtain the differentiated weight results;

[0149] According to the advantage function and the differentiation weight results, the proximal strategy optimization algorithm is used to trim and update the initial action strategy to obtain the optimized action strategy.

[0150] In one possible implementation, the system further includes a training module for:

[0151] In a simulation environment, state information samples of the humanoid robot welding process are collected to construct initial state vector samples, and time series feature compression processing is performed on the initial state vector samples to obtain compressed state vector samples.

[0152] Input the compressed state vector samples into the policy network to be trained and the value network to be trained respectively to obtain the action strategy of the current iteration and the corresponding state value estimation;

[0153] Execute actions on the simulation environment according to the action strategy, obtain the reward value fed back by the simulation environment, and calculate the actual return value based on the reward value;

[0154] Based on the action strategy of the current iteration, the action strategy of the previous iteration, the state value estimate, and the actual return value, a policy loss function and a value loss function are constructed. The policy loss function is used to measure the difference between the action strategy of the current iteration and the action strategy of the previous iteration, and the value loss function is used to measure the difference between the state value estimate and the actual return value.

[0155] According to the policy loss function and the value loss function, gradient updates are performed alternately on the policy network to be trained and the value network to be trained, and iterations are repeated during the training process until the action policy and state value estimation converge, obtaining the trained policy network and value network.

[0156] In a possible implementation, the first processing module is further configured to:

[0157] The real-time joint angle value of each joint of the humanoid robot is used as the joint angle state component;

[0158] The three-dimensional coordinates of the base of the humanoid robot in the world coordinate system are used as the base posture state components;

[0159] The three-dimensional coordinates of the weld target trajectory in the world coordinate system are used as the weld trajectory state component;

[0160] The distance between the tool center point of the humanoid robot and the closest point of the weld target trajectory is used as the weld tracking distance state component;

[0161] The minimum distance from any key point of the humanoid robot to the surface of each obstacle is taken as the obstacle distance state component;

[0162] The joint angle state component, the base posture state component, the weld trajectory state component, the weld tracking distance state component and the obstacle distance state component are spliced ​​in a predetermined order to obtain an initial state vector.

[0163] In a possible implementation, the second processing module is further configured to:

[0164] The initial state vector is input into the feature compression model to obtain a compressed state vector output by the feature compression model; wherein the feature compression model is obtained after optimizing the network structure of the long short-term memory network.

[0165] In a possible implementation, the fifth processing module is further configured to:

[0166] Perform real-time trajectory interpolation processing on each discretized joint rotation angle value to obtain the interpolation trajectory of the corresponding joint;

[0167] Real-time control instructions are generated according to all interpolation trajectories, and the welding posture of the humanoid robot is adjusted according to the real-time control instructions.

[0168] It should be noted that the humanoid robot welding posture adjustment system provided by the present invention can execute the humanoid robot welding posture adjustment method of any of the above embodiments during specific operation, which will not be described in detail in this embodiment.

[0169] Figure 7 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 7As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 may call logic instructions in the memory 730 to execute a method for adjusting the welding posture of a humanoid robot. The method includes: constructing an initial state vector based on state information of the humanoid robot during the welding process; performing time series feature compression processing on the initial state vector to obtain a compressed state vector; inputting the compressed state vector into a policy optimization model to obtain an optimized action strategy output by the policy optimization model; wherein the policy optimization model is trained based on state information samples; generating discretized joint rotation angle values ​​for each joint of the humanoid robot based on the optimized action strategy; and adjusting the welding posture of the humanoid robot based on each discretized joint rotation angle value.

[0170] Furthermore, the logic instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0171] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the humanoid robot welding posture adjustment method provided in the above embodiments.

[0172] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the humanoid robot welding posture adjustment method provided in the above embodiments.

[0173] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. That is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0174] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods of each embodiment or certain portions of the embodiments.

[0175] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for adjusting the welding posture of a humanoid robot, characterized in that: include: Constructing an initial state vector according to the state information of the humanoid robot during the welding process; Performing time series feature compression processing on the initial state vector to obtain a compressed state vector; Inputting the compressed state vector into a policy optimization model to obtain an optimized action strategy output by the policy optimization model; wherein the policy optimization model is trained based on state information samples; the policy optimization model includes a policy network and a value network, and inputting the compressed state vector into the policy optimization model to obtain an optimized action strategy output by the policy optimization model includes: Inputting the compressed state vector into the policy network and the value network respectively to obtain an initial action strategy output by the policy network and an initial value estimate output by the value network; calculating an advantage function based on the initial action strategy and the initial value estimate; Differentiated weighting is performed on the different dimensional features of the compressed state vector through a multi-head attention mechanism to obtain differentiated weight results; According to the advantage function and the differentiation weight result, the initial action strategy is trimmed and updated by using a proximal strategy optimization algorithm to obtain the optimized action strategy; generating, according to the optimized motion strategy, a discretized joint rotation angle value of each joint of the humanoid robot; The welding posture of the humanoid robot is adjusted according to each of the discretized joint rotation angle values.

2. The method for adjusting the welding posture of a humanoid robot according to claim 1, wherein: Also included are training methods for the policy network and the value network: Collecting state information samples of the humanoid robot welding process in a simulation environment, constructing initial state vector samples, and performing time series feature compression processing on the initial state vector samples to obtain compressed state vector samples; Inputting the compressed state vector sample into the strategy network to be trained and the value network to be trained respectively to obtain the action strategy of the current iteration and the corresponding state value estimation; Executing an action on the simulation environment according to the action strategy, obtaining a reward value fed back by the simulation environment, and calculating an actual return value based on the reward value; Constructing a policy loss function and a value loss function based on the action strategy of the current iteration, the action strategy of the previous iteration, the state value estimate, and the actual reward value; wherein the policy loss function is used to measure the difference between the action strategy of the current iteration and the action strategy of the previous iteration, and the value loss function is used to measure the difference between the state value estimate and the actual reward value; According to the policy loss function and the value loss function, gradient updates are performed alternately on the policy network to be trained and the value network to be trained, and iterations are repeated during the training process until the action strategy and the state value estimation converge, thereby obtaining a trained policy network and value network.

3. The method for adjusting the welding posture of a humanoid robot according to claim 1, wherein: The initial state vector is constructed according to the state information of the humanoid robot during the welding process, including: Using the real-time joint angle value of each joint of the humanoid robot as a joint angle state component; Taking the three-dimensional coordinates of the base of the humanoid robot in the world coordinate system as the base posture state component; Taking the three-dimensional coordinates of the weld target trajectory in the world coordinate system as the weld trajectory state component; The distance between the tool center point of the humanoid robot and the closest point of the weld target trajectory is used as a weld tracking distance state component; The minimum distance from any key point of the humanoid robot to the surface of each obstacle is taken as the obstacle distance state component; The joint angle state component, the base posture state component, the weld trajectory state component, the weld tracking distance state component and the obstacle distance state component are spliced ​​in a predetermined order to obtain the initial state vector.

4. The method for adjusting the welding posture of a humanoid robot according to claim 3, wherein: The performing time series feature compression processing on the initial state vector to obtain a compressed state vector includes: The initial state vector is input into a feature compression model to obtain a compressed state vector output by the feature compression model; wherein the feature compression model is obtained after optimizing the network structure of the long short-term memory network.

5. The method for adjusting the welding posture of a humanoid robot according to claim 4, wherein: The feature compression model includes multiple layers of long short-term memory network units connected in sequence, the channel dimensions of each layer of long short-term memory network units are scaled by a scaling factor, and the operations of the input gate, forget gate, and output gate in each long short-term memory network unit are fused into a single convolution operation to generate a bottleneck feature map with N channels; the feature compression model also includes a memory unit, which is used to screen out obstacle distance state components that do not meet preset standards from all the obstacle distance state components contained in the initial state vector according to the bottleneck feature map, and output the compressed state vector.

6. The method for adjusting the welding posture of a humanoid robot according to claim 1, wherein: The step of adjusting the welding posture of the humanoid robot according to each of the discretized joint rotation angle values ​​includes: Performing real-time trajectory interpolation processing on each of the discretized joint rotation angle values ​​to obtain an interpolated trajectory of the corresponding joint; A real-time control instruction is generated according to all the interpolation trajectories, and the welding posture of the humanoid robot is adjusted according to the real-time control instruction.

7. A humanoid robot welding posture adjustment system, characterized in that: include: A first processing module is used to construct an initial state vector according to state information of the humanoid robot during the welding process; A second processing module is used to perform time series feature compression processing on the initial state vector to obtain a compressed state vector; The third processing module is used to input the compressed state vector into the policy optimization model to obtain the optimized action strategy output by the policy optimization model; wherein, the policy optimization model is obtained by training based on the state information sample; the policy optimization model includes a policy network and a value network, and inputting the compressed state vector into the policy optimization model to obtain the optimized action strategy output by the policy optimization model includes: inputting the compressed state vector into the policy network and the value network respectively to obtain the initial action strategy output by the policy network and the initial value estimate output by the value network; calculating the advantage function based on the initial action strategy and the initial value estimate; differentially weighting the different dimensional features of the compressed state vector through a multi-head attention mechanism to obtain a differentiated weight result; pruning and updating the initial action strategy based on the advantage function and the differentiated weight result and using a proximal policy optimization algorithm to obtain the optimized action strategy; a fourth processing module, configured to generate a discretized joint rotation angle value of each joint of the humanoid robot according to the optimized motion strategy; The fifth processing module is used to adjust the welding posture of the humanoid robot according to each of the discretized joint rotation angle values.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the humanoid robot welding posture adjustment method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the humanoid robot welding posture adjustment method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Mechanical arm 6D pose grabbing method based on deep reinforcement learning

    CN118544353A

  • Posture recognition method based on neural network

    CN118968632A

  • Multi-target collaborative optimization control method and system for igniter robot production line

    CN119217389A

  • Multi-machine collaborative industrial robot intelligent scheduling system and application method

    CN119974019A