Humanoid robot badminton return control system based on multi-stage training model
By incorporating ballistic locking, thermal avoidance planning, impedance shaping, and impact redirection modules in a multi-stage training model, the problems of ballistic prediction lag, joint thermal saturation, and electromechanical response delay in existing humanoid robot badminton control systems have been solved. This has enabled high-precision, thermally safe badminton return control, improving competitive stability and motion smoothness.
Patent Information
- Application Number
- CN202610073284.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-02-27
AI Technical Summary
Existing humanoid robot badminton control systems suffer from issues such as trajectory prediction lag, joint thermal saturation, electromechanical response delay, and oscillation during ball strikes, making it difficult to meet the demands of high-frequency, high-dynamic continuous competition.
The system adopts a multi-stage training model-based control system. It identifies the badminton shuttlecock's flipping phase through morphological gradient recognition, dynamically adjusts the trajectory prediction accuracy, and combines zero-space projection to achieve joint thermal avoidance and impedance shaping. It also utilizes anisotropic admittance control to guide the impact retreat and return, achieving high-precision and thermally safe return control.
It improves the robot's success rate and stability in high-speed and irregular ball scenarios, reduces the risk of joint overheating, enhances continuous operation capability and competitive stability, and ensures the natural connection of the hitting-returning action and structural safety.
Smart Images

Figure CN121572328A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robot control, in particular to a humanoid robot badminton return control system based on a multi-stage training model. BACKGROUND
[0002] Existing humanoid robot badminton control systems usually adopt a serial architecture of "visual perception-trajectory prediction-full body planning", or a control strategy based on end-to-end reinforcement learning. In terms of perception and prediction, existing technologies mostly use extended Kalman filtering or neural networks to smooth fit the flight trajectory, assuming that the badminton ball follows a continuous aerodynamic model, ignoring the stepwise mutation of aerodynamic resistance caused by skirt overturning during flight, resulting in significant lag in the prediction of landing point and arrival time at the critical moment of hitting the ball.
[0003] In terms of motion planning and execution, existing technologies mainly focus on the reachability of the kinematics level, without considering the thermodynamic characteristics of the motor into the planning constraints, resulting in the robot being prone to thermal saturation of the main joints and shutdown in high-intensity continuous swings. At the same time, the underlying control mostly adopts fixed gain PD or simple variable stiffness control, which cannot effectively compensate for the mechatronic response delay of the servo system within the millisecond-level hitting window, and lacks damping matching when outputting high stiffness, which is extremely prone to causing oscillation at the end of the robot arm, resulting in "weak" or unstable hitting. In addition, in terms of balance recovery, existing systems usually treat the hitting reaction force as disturbance based on the zero moment point theory and passively resist it, which not only limits the motion amplitude of the robot, but also causes additional energy loss, making it difficult to meet the high-frequency and high-dynamic continuous competition requirements.
[0004] Therefore, a humanoid robot badminton return control system based on a multi-stage training model is proposed. SUMMARY
[0005] The purpose of the present application is to provide a humanoid robot badminton return control system based on a multi-stage training model, which identifies the badminton overturning stage through morphological gradient and adaptively adjusts the trajectory prediction accuracy, realizes joint thermal avoidance control through zero space projection, dynamically reshapes impedance parameters before and after hitting, and guides the impact retreat and return using anisotropic admittance, achieving high-precision, thermal-safe and impact-controllable return control.
[0006] To achieve the above purpose, the present application provides the following technical solutions:
[0007] The humanoid robot badminton return control system based on a multi-stage training model comprises:
[0008] A trajectory locking module is configured to collect images of the shuttlecock in flight, extract a skirt windward area change rate as a shape gradient, determine that the shuttlecock enters a turning stage when the shape gradient presents a bell-shaped curve feature, dynamically expand a noise covariance matrix according to a turning completion degree, smooth adjust a drag coefficient, and output a hitting point coordinate and a reaching time;
[0009] A thermal avoidance planning module is configured to receive the hitting point coordinate, acquire temperatures of joint motors, and calculate a thermal urgency index, construct a null space projection matrix based on a Jacobian matrix of the robot arm, generate a thermal repulsion velocity vector according to the thermal urgency index and project the thermal repulsion velocity vector to the null space, reconstruct joint angles, and output a joint instruction sequence;
[0010] An impedance shaping module is configured to execute the joint instruction sequence, take the reaching time as a trigger source, adjust joint position loop stiffness in a trapezoidal ramp manner and adjust velocity loop damping coefficients in a critical damping condition within a time window before hitting the shuttlecock.
[0011] An impact redirection module is configured to estimate a hitting reaction force vector according to a ball incoming speed and a racket swinging speed, guide a backward movement caused by the reaction force by using an anisotropic admittance control, and link a return action.
[0012] Preferably, the shape gradient acquisition process comprises:
[0013] An image sequence of the shuttlecock in flight is continuously collected, a contour region of the shuttlecock is extracted by performing background segmentation processing on each frame of the image sequence, a skirt windward region in the contour region is identified, a pixel area of the skirt windward region is calculated, a difference operation is performed on skirt pixel areas between adjacent frames to obtain an area change rate, and the area change rate is output as the shape gradient.
[0014] Preferably, the noise covariance matrix expansion process comprises:
[0015] A change trend of the shape gradient is continuously monitored within a preset sliding time window, it is determined that the shuttlecock enters the turning stage when the shape gradient exceeds a preset starting threshold and presents a bell-shaped change feature of first increasing and then decreasing within the sliding time window, an integral value of an absolute value of the shape gradient within the turning stage is time-integrated, the integral value is normalized with a preset total turning amount to obtain a turning completion degree, an expansion factor is calculated according to the turning completion degree, the expansion factor monotonically increases with the turning completion degree, a basic noise covariance matrix is multiplied by the expansion factor to obtain the expanded noise covariance matrix, and it is determined that the turning stage ends when the shape gradient falls below a preset ending threshold and remains stable, and the noise covariance matrix is restored to a basic value.
[0016] Preferably, the hitting point coordinate and the reaching time acquisition process comprises:
[0017] The state of the shuttlecock is filtered and estimated by using the inflated noise covariance matrix to obtain a position estimation value and a speed estimation value at the current time; a trajectory prediction model is established based on the adjusted resistance coefficient; the position estimation value and the speed estimation value are substituted into the trajectory prediction model to extrapolate the flight trajectory of the shuttlecock; the intersection of the flight trajectory and a preset hitting plane is calculated as a hitting point coordinate; the flight time of the shuttlecock from the current position to the hitting point is calculated according to the flight trajectory, and the arrival time is obtained in combination with the current time.
[0018] Preferably, the construction process of the null space projection matrix comprises:
[0019] The current joint angle configuration of the redundant degree of freedom manipulator is obtained; the mapping relationship between the end effector speed and the joint speed under the current configuration is calculated according to the kinematics model of the manipulator to obtain a Jacobian matrix; pseudo-inverse operation is performed on the Jacobian matrix to obtain a Jacobian pseudo-inverse matrix; the product of the Jacobian pseudo-inverse matrix and the Jacobian matrix is subtracted from the unit matrix to obtain a null space projection matrix.
[0020] Preferably, the obtaining process of the joint instruction sequence comprises:
[0021] The temperature sensor data of each joint motor is read, and the residual heat capacity time under the current torque and speed working condition is queried; the residual heat capacity time is compared with a preset safety threshold time to obtain a thermal urgency index of each joint by normalization; a thermal repulsion potential field is constructed according to the thermal urgency index of each joint, and a negative gradient of the thermal repulsion potential field is calculated to generate a thermal repulsion speed vector that promotes each joint to move away from the overheated state.
[0022] The expected speed of the end effector is calculated according to the hitting point coordinate and the motion time planning; the Jacobian pseudo-inverse matrix is used to map the end effector expected speed to the main task joint speed; the thermal avoidance speed in the null space is obtained by projecting the thermal repulsion speed vector by the null space projection matrix; the main task joint speed and the null space thermal avoidance speed are superimposed to obtain a composite joint speed; the composite joint speed is time-integrated to obtain a joint angle sequence as the joint instruction sequence output.
[0023] Preferably, the adjustment process of the joint position loop stiffness and the speed loop damping coefficient comprises:
[0024] Obtain the electromechanical delay time and the preset stiffness adjustment duration. Determine the adjustment start time by subtracting the electromechanical delay time and stiffness adjustment duration from the arrival time. Obtain the equivalent inertia of each joint in the current configuration. Starting from the adjustment start time, increase the joint position ring stiffness from the initial value to the target peak value according to the trapezoidal ramp law. Calculate the corresponding damping coefficient according to the critical damping condition based on the equivalent inertia and the current stiffness value, so that the damping ratio is maintained in the critical damping and / or overdamped state. After the shot is completed, restore the stiffness and damping coefficient to the initial values.
[0025] Preferably, the control process for anisotropic admittance control to guide the reaction force to generate backward motion includes:
[0026] Based on the ball velocity, swing velocity, and estimated contact time, calculate the direction and peak value of the reaction force vector; establish a local coordinate system related to the reaction force direction to distinguish between the reaction force direction and the anti-collapse direction; set low stiffness parameters and moderate damping parameters in the reaction force direction, and high stiffness parameters in the anti-collapse direction; within the hitting window, read the force and / or torque signals of the lower limb joints; calculate the allowable displacement response in each direction based on the stiffness and damping parameters; synthesize the displacement responses in each direction and convert them into lower limb joint angle correction values, which are then superimposed on the original joint commands; use the controlled backward movement generated by the reaction force direction to connect the subsequent return motion.
[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0028] 1. This invention introduces a morphological gradient modeling method based on the rate of change of the windward area of the badminton shuttlecock's skirt, elevating traditional single-vision tracking that relies on position and velocity to the identification of the shuttlecock's flipping dynamics. When a bell-shaped curve characteristic is detected in the morphological gradient, the system can accurately determine that the shuttlecock has entered the flipping phase. Furthermore, based on the dynamic expansion of the noise covariance matrix of the flipping completion degree, it achieves adaptive modeling of uncertain aerodynamic disturbances. This mechanism avoids filter divergence or prediction offset caused by abrupt changes in aerodynamic parameters during the flipping phase, ensuring the ballistic prediction model maintains continuity and smoothness even under complex flight attitudes. This allows for the output of the hitting point coordinates and arrival time, improving the robot's return success rate and stability in high-speed, irregularly shaped ball scenarios.
[0029] 2、The joint motor thermal state is introduced into the ball return control decision process, a thermal urgency index is constructed, and the thermal avoidance and ball hitting task decoupling control is realized through zero space projection by combining the mechanical arm redundant freedom characteristics. While ensuring that the end effector strictly meets the ball hitting point and timing requirements, the thermal repulsion velocity vector is projected to the zero space, and only the degree of freedom direction that does not affect the main task is adjusted, thereby effectively reducing the continuous heating risk of high-load joints. Fine-grained and real-time thermal safety adjustment can be realized in high-speed confrontation movements, prolonging the life of the motor and reducer, and improving the continuous operation ability and competitive stability.
[0030] 3、The application introduces a multi-stage dynamic adjustment mechanism combining impedance shaping and anisotropic admittance control before and after hitting the ball, so that the robot can be directed to adapt to the impact characteristics at the moment of hitting the ball. By increasing the joint stiffness according to the trapezoidal slope before hitting the ball and matching the critical damping, the swing precision and energy transmission efficiency are ensured; during the ball contact stage, the direction-related stiffness and damping parameters are set according to the estimated reaction force direction to guide the controlled backward movement and suppress the structural deformation in the anti-collapse direction. This method avoids impact concentration and vibration accumulation under rigid control, realizes natural connection of the hitting and returning actions, reduces the impact load of the joints and the body structure, and improves the safety and smoothness of the humanoid robot in high-frequency sparring scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 A humanoid robot badminton return ball control system structure diagram based on a multi-stage training model is provided for the application.
[0032] Figure 2 A humanoid robot badminton return ball control flowchart is provided for the application.
[0033] Figure 3 A joint instruction sequence acquisition process flowchart is provided for the embodiment of the application. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical scheme and advantages of the application clearer and more understandable, the application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the application and do not limit the application.
[0035] Embodiment one:
[0036] As Figure 1 and Figure 2As shown, this invention provides a humanoid robot badminton return control system based on a multi-stage training model, including a ballistic locking module, a thermal avoidance planning module, an impedance shaping module, and an impact redirection module. There are clear data flow relationships between the four modules: the ballistic locking module outputs the coordinates of the hitting point and the arrival time; the thermal avoidance planning module receives the hitting point coordinates and outputs a joint command sequence; the impedance shaping module executes the joint command sequence using the arrival time as the trigger source; and the impact redirection module guides the backward movement within the hitting window and connects it to the return motion. The specific technical solutions are as follows: A ballistic locking module is used to acquire badminton shuttlecock flight images and extract the rate of change of the skirt's windward area as a morphological gradient. When the morphological gradient exhibits a bell-shaped curve characteristic, it determines that the flipping phase has begun. Based on the dynamic expansion noise covariance matrix of the flipping completion degree, it smoothly adjusts the drag coefficient and outputs the hitting point coordinates and arrival time. A thermal avoidance planning module is used to receive the hitting point coordinates, obtain the temperature of each joint motor, and calculate the thermal stress index. Based on the Jacobian matrix of the robotic arm, a null space projection matrix is constructed. Based on the thermal stress index, a thermal repulsion velocity vector is generated and projected onto the null space to reconstruct the joint angles and output the joint command sequence. An impedance shaping module is used to execute the joint command sequence. Using the arrival time as the trigger source, within the time window before the hit, it adjusts the joint position ring stiffness according to a trapezoidal ramp and adjusts the velocity ring damping coefficient according to the critical damping condition. An impact redirection module is used to estimate the hitting reaction force vector based on the incoming ball speed and swing speed, and uses anisotropic admittance control to guide the backward motion generated by the reaction force, connecting the return motion.
[0037] Furthermore, the process of obtaining the morphological gradient includes:
[0038] A series of images of the badminton shuttlecock in flight are continuously acquired. Background segmentation is performed on each frame to extract the outline region of the shuttlecock. The windward area of the skirt in the outline region is identified, and the pixel area of the windward area of the skirt is calculated. The difference operation is performed on the skirt pixel area between adjacent frames to obtain the area change rate. The area change rate is output as the morphological gradient.
[0039] Specifically, in this embodiment, the process of obtaining the morphological gradient is as follows:
[0040] A global shutter camera was used as the visual acquisition device, with a frame rate of 500 frames per second and a resolution of 640 x 480 pixels. The camera was installed above or to the side of the badminton court, covering the main flight area of the shuttlecock. A ring LED fill light was provided to ensure the brightness of the image under short exposure time.
[0041] Within each acquisition cycle, the following processing steps are performed:
[0042] The first step is to perform background subtraction or color threshold-based segmentation on the current frame image to separate the badminton shuttlecock from the background and obtain a binarized mask image of the badminton shuttlecock.
[0043] The background segmentation process uses a Gaussian mixture model algorithm to model the static and quasi-static parts of the background. The specific process is as follows: (1) In the initialization stage, 100 frames of images are collected to establish an initial background model. Each pixel is modeled using 3 Gaussian distribution components; (2) In the running stage, the background model is updated once every control cycle (16.67ms). The learning rate is set to 0.005 to ensure that the background model can slowly adapt to changes in illumination but will not be contaminated by moving targets; (3) The foreground / background classification uses Mahalanobis distance for decision-making. The distance threshold is set to 2.5σ (that is, pixels whose Mahalanobis distance does not exceed 2.5 times the standard deviation are judged as background).
[0044] The second step is to extract the contour of the binarized mask image to obtain the set of boundary points of the outer contour of the badminton shuttlecock. Then, fit the minimum bounding ellipse based on the set of boundary points to obtain the major axis direction vector and the area of the ellipse.
[0045] The third step is to divide the ellipse into a front half and a back half along the major axis direction. The back half corresponds to the skirt area of the badminton shuttlecock. The number of pixels in the skirt area is counted as the skirt windward pixel area of the current frame.
[0046] The fourth step is to subtract the area of the skirt facing the wind from the area of the skirt facing the wind in the previous frame, and divide by the time interval between the two adjacent frames to obtain the area change rate; this area change rate is the morphological gradient.
[0047] Furthermore, the expansion process of the noise covariance matrix includes:
[0048] Within a preset sliding time window, the changing trend of the morphological gradient is continuously monitored. When the morphological gradient exceeds a preset starting threshold and exhibits a bell-shaped change characteristic of first increasing and then decreasing within the sliding time window, it is determined that the flipping phase has begun. The absolute value of the morphological gradient within the flipping phase is integrated over time, and the integrated value is normalized with a preset total flipping amount to obtain the flipping completion degree. An expansion factor is calculated based on the flipping completion degree, and the expansion factor monotonically increases with the flipping completion degree. The basic noise covariance matrix is multiplied by the expansion factor to obtain the expanded noise covariance matrix. When the morphological gradient falls back to below a preset ending threshold and remains stable, the flipping phase is determined to end, and the noise covariance matrix is restored to its basic value.
[0049] Specifically, in this embodiment, the expansion process of the noise covariance matrix is as follows:
[0050] Maintain a sliding time window with a window length of 20ms, corresponding to 10 frames of image data acquired by the camera at a capture rate of 500 frames per second. Within the sliding time window, continuously monitor the changes in the morphological gradient value. When the morphological gradient value first exceeds the preset starting threshold, start judging the trend of the morphological gradient change within the window.
[0051] The system detects whether the morphological gradient sequence exhibits a characteristic of first increasing and then decreasing. Specifically, it searches for the maximum point of the morphological gradient within a sliding window. If the maximum point is located near the middle of the window, and the data before the maximum point shows an upward trend while the data after the maximum point shows a downward trend, then the morphological gradient is determined to exhibit a bell-shaped change characteristic. When the morphological gradient simultaneously meets both the conditions of exceeding the initial threshold and exhibiting a bell-shaped change characteristic, the shuttlecock is determined to have entered the flip phase, the flip flag is set to a valid state, and the start time of the flip is recorded. The process for determining the flip phase includes:
[0052] Within a preset sliding time window (window length 15-30 milliseconds, corresponding to 8-15 frames of data at a sampling rate of 500 frames per second), the morphological gradient sequence is continuously monitored. When the morphological gradient first exceeds the initial threshold (determined through offline calibration to ensure a false positive rate ≤5%; or set as the mean of the background gradient plus twice the standard deviation), morphological feature detection is initiated. First-order differencing is performed on the morphological gradient sequence within the sliding window to determine if there exists a maximum point where the difference before that point is positive and the difference after is negative. Simultaneously, the skewness and kurtosis of the sequence are calculated: when the skewness is close to 0 and the kurtosis is >3, the sequence is determined to be bell-shaped. Furthermore, the rate of change from the maximum to the initial value is calculated; when the rate of change is >50%, the flipping phase is further confirmed. When the morphological gradient remains below the end threshold (set as 60% of the initial threshold) for 3-5 consecutive cycles and the rate of change remains low, the flipping phase is determined to have ended.
[0053] When the morphological gradient is below the end threshold (set as 60% of the start threshold) for 3-5 consecutive cycles and the rate of change remains low, the flipping phase is considered to have ended.
[0054] During the flip phase, the absolute value of the morphological gradient is integrated over time. Specifically, the absolute value of the morphological gradient in each acquisition cycle is multiplied by the acquisition cycle duration and accumulated to obtain the cumulative change during the flip phase. The total cumulative morphological gradient when the shuttlecock completes a full flip is pre-calibrated; the current cumulative change is divided by the total cumulative change to obtain the normalized flip completion rate, which ranges from 0 to 1.
[0055] The inflation factor is calculated based on the flip completion rate. The inflation factor increases monotonically with the flip completion rate; when the flip completion rate is 0, the inflation factor is 1, and when the flip completion rate is 1, the inflation factor reaches the preset maximum inflation factor. In this embodiment, the maximum inflation factor is set to a value between 5 and 10. Each element of the basic noise covariance matrix is multiplied by the inflation factor to obtain the inflated noise covariance matrix. The inflated noise covariance matrix is used for subsequent state filtering estimation.
[0056] When the morphological gradient falls below the preset end threshold and remains stable for multiple consecutive acquisition cycles, the flipping phase is determined to be over, the flipping flag is set to invalid, and the noise covariance matrix is restored to the base value.
[0057] Furthermore, the process of obtaining the coordinates of the striking point and the arrival time includes:
[0058] The state of the badminton shuttlecock is filtered and estimated using the expanded noise covariance matrix to obtain the position and velocity estimates at the current moment. A trajectory prediction model is established based on the adjusted drag coefficient. The position and velocity estimates are substituted into the trajectory prediction model to extrapolate the flight trajectory of the badminton shuttlecock. The intersection of the flight trajectory and the preset hitting plane is calculated, and the intersection is used as the coordinates of the hitting point. The flight time of the badminton shuttlecock from the current position to the hitting point is calculated based on the flight trajectory, and the arrival time is obtained by combining the current moment.
[0059] Specifically, in this embodiment, the process of obtaining the coordinates of the hitting point and the arrival time is as follows:
[0060] A state filter is used to estimate the motion state of the badminton shuttlecock. The state vector is a 6-dimensional vector, which includes three components: the position coordinates of the shuttlecock in three-dimensional space and three components: velocity. The observation vector is a 3-dimensional vector, corresponding to the three-dimensional position coordinates of the shuttlecock measured by the vision system. In each filtering cycle, the state filter performs prediction and update steps.
[0061] In the prediction step, the state filter predicts the state of the shuttlecock at the next moment based on its dynamics. The shuttlecock's dynamics consider the effects of gravity and air resistance. Air resistance is proportional to the square of the shuttlecock's velocity, with the proportionality coefficient determined by the drag coefficient. During the flip phase, the drag coefficient gradually changes from its value before the flip to its value after the flip, following a smooth transition pattern. The drag coefficient before the flip is 0.5, and the drag coefficient after the flip is 0.7. The transition pattern of the drag coefficient is related to the degree of flip completion and uses an exponential function to achieve a smooth change.
[0062] In the update step, the state filter uses the expanded noise covariance matrix to calculate the filter gain. Since the value of the noise covariance matrix increases after expansion, the filter gain increases accordingly, making the state estimate more reliant on current visual observations and reducing the trust in historical predictions based on old models.
[0063] After the state filter outputs the current position and velocity estimates, a ballistic prediction model is established based on the adjusted drag coefficient. The ballistic prediction model uses a numerical integration method, taking the current position and velocity estimates as initial conditions, and progressively extrapolating the badminton shuttlecock's flight trajectory. The position of the striking plane is preset; the striking plane is typically a vertical plane perpendicular to the longitudinal direction of the court, located in the leading edge region of the robot's reachable workspace. The coordinates of the intersection point between the extrapolated flight trajectory and the striking plane are calculated; these coordinates are the striking point coordinates. The flight time required for the shuttlecock to reach the striking point from the current position is calculated based on the extrapolated flight trajectory. This flight time is added to the current system time to obtain the arrival time; the accuracy of the arrival time reaches the microsecond level.
[0064] The coordinates of the hitting point and the arrival time are sent to the thermal avoidance planning module and the impedance shaping module via a real-time communication bus.
[0065] Furthermore, the construction process of the null space projection matrix includes:
[0066] Obtain the current joint angle configuration of the redundant degree-of-freedom robotic arm; calculate the mapping relationship between the end effector velocity and the joint velocity under the current configuration based on the kinematic model of the robotic arm, and obtain the Jacobian matrix; perform pseudo-inverse operation on the Jacobian matrix to obtain the Jacobian pseudo-inverse matrix; subtract the product of the Jacobian pseudo-inverse matrix and the Jacobian matrix from the identity matrix to obtain the null space projection matrix.
[0067] Furthermore, the process of obtaining the joint command sequence refers to... Figure 3 ,include:
[0068] Read the temperature sensor data of each joint motor; according to the pre-established motor thermal characteristic lookup table, query the remaining heat capacity time under the current torque and speed conditions; compare the remaining heat capacity time with the preset safety threshold time, and normalize to obtain the thermal urgency index of each joint; construct a thermal repulsion potential field based on the thermal urgency index of each joint, calculate the negative gradient of the thermal repulsion potential field, and generate a thermal repulsion velocity vector that promotes each joint to move away from the overheated state.
[0069] The process of establishing the thermal characteristic lookup table is as follows: Calibration experiments were conducted on each joint motor at room temperature (20°C to 25°C). For torque, the sampling range was 0% to 100% of the rated torque, with one sampling point for every 10%, for a total of 11 sampling points. For speed, the sampling range was 0% to 100% of the rated speed, with one sampling point for every 10%, for a total of 11 sampling points. For the initial motor temperature, the sampling range was 20°C to 60°C (the maximum allowable temperature), with one sampling point for every 10°C, for a total of 5 sampling points. Under each sampling condition, the motor was continuously run until the temperature reached the maximum allowable temperature. The temperature rise curve over time was recorded, and the time required to rise from the current temperature to the maximum allowable temperature under that condition (i.e., the remaining heat capacity time) was obtained through time constant decomposition or exponential fitting. The table is stored using a three-dimensional array structure, with indices of [torque index, speed index, current temperature index], and the output value is the remaining heat capacity time (unit: seconds).
[0070] Based on the ball-hitting point coordinates and motion time planning, the desired velocity of the end effector is calculated; the desired velocity of the end effector is mapped to the main task joint velocity using the Jacobian pseudo-inverse matrix; the thermal repulsion velocity vector is projected using the null space projection matrix to obtain the thermal avoidance velocity in the null space; the main task joint velocity is superimposed with the null space thermal avoidance velocity to obtain the composite joint velocity; the composite joint velocity is integrated over time to obtain the joint angle sequence as the joint command sequence output.
[0071] Specifically, in this embodiment, the calculation of the thermal stress index and the generation of the thermal repulsion velocity vector are as follows:
[0072] In the offline phase, a thermal characteristic lookup table for each joint motor is pre-established. For each joint motor, thermal characteristic calibration experiments are conducted under different torque and speed combinations, and the motor temperature rise curve over time is recorded. Based on the motor's heat capacity and heat dissipation characteristics, the time required for the motor to rise from the current temperature to the maximum allowable temperature is obtained by fitting, and a three-dimensional lookup table is established with torque, speed, and current temperature as inputs and remaining heat capacity time as output.
[0073] During the real-time operation phase, the system performs the following steps in each control cycle:
[0074] The first step is to read the current temperature values from the built-in temperature sensors of each joint motor;
[0075] The second step is to read the current torque command and current speed value fed back by the servo driver of each joint;
[0076] The third step is to use the current torque, current speed, and current temperature as indexes to query the thermal characteristic lookup table and obtain the remaining thermal capacity time of the joint motor.
[0077] The fourth step is to compare the remaining heat capacity time with a preset safety threshold time. In this embodiment, the safety threshold time is set to 10 seconds. If the remaining heat capacity time is greater than or equal to the safety threshold time, the thermal urgency index is 0. If the remaining heat capacity time is less than the safety threshold time, the thermal urgency index is equal to 1 minus the ratio of the remaining heat capacity time to the safety threshold time. The thermal urgency index ranges from 0 to 1; a larger value indicates that the joint motor is closer to an overheated state.
[0078] The fifth step is to construct a thermal repulsion potential field based on the thermal stress index of each joint. The function value of the thermal repulsion potential field is equal to the weighted sum of the squares of the thermal stress indices of each joint. The weighting coefficients are preset according to the importance and thermal sensitivity of each joint.
[0079] The sixth step is to obtain the negative gradient of the thermal repulsion potential field to generate a thermal repulsion velocity vector. Each component of the thermal repulsion velocity vector is proportional to the thermal stress index of the corresponding joint and is in the negative direction, indicating that the joint is prompted to move in the direction of reducing thermal stress. The thermal repulsion velocity vector is also multiplied by the thermal avoidance gain coefficient to adjust the intensity of the thermal avoidance movement.
[0080] In this embodiment, the process of constructing the null projection matrix and obtaining the joint command sequence is as follows:
[0081] In this embodiment, the humanoid robot's upper limb robotic arm adopts a 7-DOF redundant configuration. The number of joint degrees of freedom is greater than the 6 dimensions of the end effector's task space, thus resulting in a 1-dimensional null space. The angle encoder values of each joint are read during each control cycle to obtain the current joint angle configuration of the robotic arm.
[0082] Based on the kinematic model of the robotic arm (including forward and inverse kinematics), the Jacobian matrix for the current configuration is calculated. The Jacobian matrix describes the linear mapping between the end effector velocity and the joint velocities. The matrix has 6 rows, corresponding to the 6 components of the end effector's translational and rotational velocities; and 7 columns, corresponding to the angular velocities of the 7 joints. A pseudo-inverse operation is performed on the Jacobian matrix. Since the Jacobian matrix has fewer rows than columns, the right pseudo-inverse calculation method is used, which involves multiplying the transpose of the Jacobian matrix by the inverse of the product of the Jacobian matrix and its transpose. To improve numerical stability, damped least squares is used for the pseudo-inverse calculation.
[0083] Subtracting the product of the Jacobian pseudo-inverse and the Jacobian matrix from the 7th-order identity matrix yields the null space projection matrix; the null space projection matrix is a 7x7 square matrix with a rank of 1, representing the 1-dimensional null space of the robotic arm.
[0084] Based on the striking point coordinates and motion time planning, the desired velocity of the end effector is calculated. The motion time planning is determined based on the difference between the arrival time and the current time. The desired end effector velocity includes the translational velocity component of the end effector position moving towards the striking point, and the rotational velocity component of the end effector attitude adjusting towards the striking attitude. The desired end effector velocity is mapped to the main task joint velocities using the Jacobian pseudo-inverse matrix. The main task joint velocities are 7-dimensional vectors, representing the angular velocities of each joint required to achieve the desired end effector velocity.
[0085] The thermal repulsion velocity vector is projected using the null space projection matrix. The thermal avoidance velocity in the null space is obtained by multiplying the null space projection matrix with the thermal repulsion velocity vector. The null space thermal avoidance velocity does not change the speed of the end effector, but only causes a change in the internal configuration of the robotic arm.
[0086] The primary task joint velocity is added component by component to the zero-space thermal avoidance velocity to obtain the synthetic joint velocity. The synthetic joint velocity simultaneously satisfies the primary task requirement of reaching the impact point at the end of the ball and the secondary task requirement of reducing the thermal load on the joint. The synthetic joint velocity is then integrated over time, that is, the joint angle at the current moment is added to the product of the synthetic joint velocity and the control cycle duration to obtain the joint angle at the next moment. This process is repeated to obtain a sequence of joint angles, which is then output as the joint command sequence.
[0087] Before outputting the joint command sequence, a constraint check is performed on each joint angle, including joint angle limit check, joint angular velocity limit check, and joint angular acceleration limit check. If the check fails, the corresponding values are trimmed.
[0088] Furthermore, the adjustment process for the joint position ring stiffness and velocity ring damping coefficient includes:
[0089] Obtain the electromechanical delay time and the preset stiffness adjustment duration. Determine the adjustment start time by subtracting the electromechanical delay time and stiffness adjustment duration from the arrival time. Obtain the equivalent inertia of each joint in the current configuration. Starting from the adjustment start time, increase the joint position ring stiffness from the initial value to the target peak value according to the trapezoidal ramp law. Calculate the corresponding damping coefficient according to the critical damping condition based on the equivalent inertia and the current stiffness value, so that the damping ratio is maintained in the critical damping and / or overdamped state. After the shot is completed, restore the stiffness and damping coefficient to the initial values.
[0090] Specifically, in this embodiment, the adjustment process for the joint position ring stiffness and velocity ring damping coefficient is as follows:
[0091] The electromechanical delay time is obtained in advance through online identification or offline calibration. The electromechanical delay time includes the sum of communication delay, servo drive response delay and mechanical transmission delay, and is set to a value between 5ms and 15ms in this embodiment.
[0092] The preset stiffness adjustment time is set to 30ms in this embodiment. The adjustment start time is determined by subtracting the electromechanical delay time and stiffness adjustment time from the arrival time.
[0093] Obtain the equivalent inertia of each joint under the current configuration. The equivalent inertia can be obtained in the following ways: The first way is to calculate the inertia parameters of each link based on the CAD model of the robotic arm, and calculate the equivalent inertia in combination with the inertia transfer relationship under the current configuration; The second way is to conduct an inertia identification experiment on each joint in the offline stage, establish a lookup table of equivalent inertia and joint angle configuration, and obtain the equivalent inertia by looking up the table in the real-time operation stage.
[0094] The initial value and target peak value of the joint position ring stiffness are preset. The initial value corresponds to the stiffness in the compliant control mode, which is set to 100 N·m / rad in this embodiment; the target peak value corresponds to the stiffness in the hitting mode, which is set to 2000 N·m / rad in this embodiment.
[0095] Starting from the initial adjustment moment, the joint position ring stiffness is adjusted according to the trapezoidal slope law; within the stiffness adjustment time, the stiffness increases linearly from the initial value to the target peak value; specifically, in each control cycle of the adjustment process, the time difference between the current moment and the initial adjustment moment is calculated, and divided by the stiffness adjustment time to obtain the normalized time; the initial stiffness value and the stiffness change are multiplied by the normalized time and added to obtain the stiffness value at the current moment.
[0096] While adjusting the stiffness, the system calculates the corresponding damping coefficient based on the equivalent inertia and the current stiffness value, so that the system damping ratio is kept in the critical damping or overdamped state. The critical damping condition requires that the damping coefficient be equal to twice the square root of the product of the equivalent inertia and the stiffness. The calculated damping coefficient is used as the velocity loop damping coefficient.
[0097] The current stiffness and damping coefficients are sent to the servo drivers of each joint via the servo communication bus. The servo drivers then update the control gains of the position and velocity loops accordingly. After the shot is completed, the stiffness and damping coefficients are restored to their initial values. The restoration process uses an exponential decay law to ensure that the stiffness quickly returns from the target peak value to the initial value.
[0098] Furthermore, the control process by which anisotropic admittance control guides the reaction force to generate recoil motion includes:
[0099] Based on the ball velocity, swing velocity, and estimated contact time, calculate the direction and peak value of the reaction force vector; establish a local coordinate system related to the reaction force direction to distinguish between the reaction force direction and the anti-collapse direction; set low stiffness parameters and moderate damping parameters in the reaction force direction, and high stiffness parameters in the anti-collapse direction; within the hitting window, read the force and / or torque signals of the lower limb joints; calculate the allowable displacement response in each direction based on the stiffness and damping parameters; synthesize the displacement responses in each direction and convert them into lower limb joint angle correction values, which are then superimposed on the original joint commands; use the controlled backward movement generated by the reaction force direction to connect the subsequent return motion.
[0100] Specifically, in this embodiment, the control process by which anisotropic admittance control guides the reaction force to generate backward motion is as follows:
[0101] The system acquires the incoming ball speed output by the ballistic locking module and the swing speed output by the motion planning module; it also estimates the contact time during the hitting process, which is set to 2ms in this embodiment.
[0102] The direction and peak value of the reaction force vector are calculated based on the momentum theorem. According to the momentum theorem, impulse equals change in momentum. The mass of the badminton shuttlecock is 5 grams. The difference between the incoming speed and the outgoing speed is the change in velocity. Multiplying the change in velocity by the mass gives the change in momentum. Dividing the change in momentum by the contact time gives the average reaction force. Multiplying the average reaction force by the peak factor gives the peak value of the reaction force. For an impact with an approximate half-sine waveform, the peak factor is approximately 1.57. The direction of the reaction force is opposite to the direction of the swing.
[0103] A local coordinate system related to the direction of the reaction force is established. The first axis of the local coordinate system is along the direction of the reaction force, the second axis is along the direction perpendicular to the ground as the anti-collapse direction, and the third axis is obtained by the cross product of the first and second axes. If the angle between the reaction force direction and the vertical direction is small, the lateral direction is used as the third axis, and the second axis is recalculated. The reaction force is measured using a force estimator from hybrid force / position control: based on joint torque sensor feedback or joint pressure derived from inverse kinematics, the force is transformed to Cartesian space through the Jacobian matrix to obtain the reaction force vector acting on the end effector.
[0104] Admittance parameters are set for each direction. In the reaction force direction, low stiffness and moderate damping parameters are set to allow the robot to generate controlled backward motion in this direction. In this embodiment, the low stiffness parameter is set to 50 N / m and the damping parameter is set to 200 N·s / m. In the anti-collapse direction, high stiffness and high damping parameters are set to prevent the robot from sinking or collapsing in this direction. In this embodiment, the high stiffness parameter is set to 5000 N / m and the high damping parameter is set to 500 N·s / m. In the third axis direction, medium stiffness and medium damping parameters are set.
[0105] Based on the stiffness and damping parameters in each direction, assemble the admittance parameter matrix; transform the admittance parameter matrix to the world coordinate system through coordinate transformation.
[0106] Within the striking window, the system reads the force or torque signal of the lower limb joints using torque sensors or current estimation methods; the striking window begins 5ms before the arrival time and ends 20ms after the arrival time.
[0107] The allowable displacement response in each direction is calculated based on the admittance control equation. The admittance control equation describes the dynamic relationship between external force and displacement response, and includes inertial, damping and stiffness terms. The displacement and velocity responses in each direction are obtained by numerical integration of the admittance control equation.
[0108] After synthesizing the displacement responses in each direction, they are converted into lower limb joint angle correction values through inverse kinematics calculations. The joint angle correction values are then superimposed on the original joint commands and sent to the lower limb joint servo driver for execution.
[0109] After the shot window ends, the system detects the backward movement speed generated by the reaction force. If the backward speed exceeds a preset threshold, the system activates the return motion planner, uses the current backward speed as the initial condition, plans the return trajectory from the current position to the next defensive position, and hands it over to the lower limb control module for execution.
[0110] This invention achieves high-precision and robust ball return control for a humanoid robot in high-speed badminton scenarios by constructing four tightly coupled functional modules: ballistic locking, thermal avoidance planning, impedance shaping, and impact redirection. The system uses the rate of change of the windward area of the shuttlecock's skirt as a morphological gradient to accurately identify the flip phase. Through dynamic expansion of the noise covariance matrix and smooth adjustment of the drag coefficient, the stability of ballistic estimation and shot timing prediction during the flip phase is significantly improved. Based on this, the system utilizes the zero-space characteristics of the redundant robotic arm to incorporate joint thermal stress into motion planning, achieving real-time avoidance of joint thermal loads without affecting the primary task of hitting the shuttlecock, ensuring the system's continuous high-intensity operation capability. Before and after the shot, an impedance shaping strategy triggered at the arrival time adaptively matches the joint stiffness and damping to the hitting conditions, balancing hitting accuracy and system compliance. At the moment of impact, anisotropic admittance control guides the reaction force to generate controlled backward motion, naturally connecting to the return motion, effectively dispersing the impact load and improving overall stability and motion continuity. Overall, this invention achieves multi-stage collaborative optimization of perception, planning, and execution, significantly enhancing the reliability, safety, and competitive performance of humanoid robots in badminton return.
[0111] Example 2:
[0112] This embodiment, based on Embodiment 1, further describes the enhanced functions of the system, including the specific implementation methods of the zero torque point pre-bias unit, the online identification unit of equivalent inertia, and the online correction unit of drag coefficient.
[0113] like Figure 1 As shown, the humanoid robot badminton return control system of this embodiment includes a ballistic locking module, a thermal avoidance planning module, an impedance shaping module, and an impact redirection module. The ballistic locking module acquires badminton flight images using a high frame rate camera, extracts the morphological gradient, determines the flip phase, and outputs the hitting point coordinates and arrival time after dynamically expanding the noise covariance matrix. The thermal avoidance planning module, based on null space projection technology, achieves joint thermal load redistribution while maintaining the end effector pose. The impedance shaping module adjusts the stiffness according to a trapezoidal ramp and synchronously adjusts the damping coefficient according to the critical damping condition. The impact redirection module uses anisotropic admittance control to guide the backward motion.
[0114] The focus of this embodiment is on the specific implementation of the following three enhancement units.
[0115] In this embodiment, the impact redirection module includes a zero-torque point pre-bias unit, which is used to actively adjust the robot's posture before the ball is hit, creating a stability margin for the impact response.
[0116] The impact redirection module further includes a zero-torque point pre-bias unit, used for:
[0117] Within a preset time window before the ball is struck, the zero-moment point offset required to maintain stability is calculated based on the estimated reaction force vector, the robot's center of mass height, and the robot's total mass. By adjusting the angles of the hip and ankle joints of the lower limbs, the robot's center of mass position is controlled to move in the opposite direction of the reaction force vector, causing the zero-moment point to shift accordingly. After the ball is struck, the position of the zero-moment point is continuously monitored to determine whether it is located within the support polygon formed by the support surfaces of the two feet.
[0118] Specifically, in this embodiment, the impact redirection module includes a zero-torque point pre-bias unit, which is used to actively adjust the robot's posture before hitting the ball, creating a stability margin for the impact response.
[0119] The zero-torque point pre-biasing unit starts working within a preset time window before impact; in this embodiment, the preset time window is set to 50ms before impact. The calculation process of the zero-torque point offset is as follows: First, the reaction force vector estimated by the impact redirection module is obtained, which includes two elements: direction and peak value. Simultaneously, the robot's current center of mass height and total mass are obtained. According to the simplified inverted pendulum model, when a horizontal external force acts on the center of mass position, the zero-torque point will shift. The zero-torque point offset is equal to the peak value of the external force multiplied by the center of mass height, and then divided by the product of the robot's total mass and gravitational acceleration.
[0120] In this embodiment, the total mass of the robot is set to 50 kg and the height of the center of mass is set to 0.8 m. When the estimated peak reaction force is 400 N, the calculated offset of the zero moment point is about 0.65 m. Considering the safety margin, the actual offset is taken as half of the calculated value, that is, about 0.32 m.
[0121] The process of adjusting the center of mass position is as follows: By adjusting the angles of the hip and ankle joints of the lower limbs, the robot's center of mass is controlled to move in the opposite direction of the reaction force vector; adjusting the hip joint angle causes the torso to tilt forward or backward, and adjusting the ankle joint angle causes the overall center of gravity to move forward or backward; the two joints coordinate their movements to achieve directional movement of the center of mass while maintaining stable support on both feet. Within a preset time window, joint angle adjustments are executed according to a smooth trajectory plan to avoid posture instability caused by sudden changes; after the adjustment is completed, the robot's center of mass has shifted in the opposite direction of the reaction force, and the zero torque point is pre-biased accordingly.
[0122] The stability monitoring process after the ball is hit is as follows: After the ball-hitting window ends, the real-time position of the zero-moment point is continuously monitored; the position of the zero-moment point is calculated based on the robot's current center of mass position, center of mass acceleration, and ground reaction force. It is then determined whether the zero-moment point is located within the support polygon formed by the two foot support surfaces; the support polygon is determined by the convex hull formed by the outer boundary points of the left and right foot contact areas with the ground.
[0123] If the zero-moment point is located inside the supporting polygon, the robot is determined to be in a stable state; if the zero-moment point is close to or exceeds the boundary of the supporting polygon, an emergency balance recovery strategy is triggered.
[0124] In this embodiment, the impedance shaping module includes an online equivalent inertia identification unit, used to estimate the equivalent inertia parameters of each joint in real time. The impedance shaping module further includes an online equivalent inertia identification unit, used for:
[0125] During the non-hitting phase of the movement, the torque signals output by the servo drives of each joint and the angular acceleration signals obtained by differential data from the angle encoder are collected. Based on the ratio of the torque signals to the angular acceleration signals, the equivalent inertia of each joint under the current configuration is estimated in real time using a recursive estimation method. The equivalent inertia obtained in real time is used to replace the offline calibration values for the calculation of the damping coefficient under critical damping conditions, so that the damping coefficient is adaptively adjusted with the configuration.
[0126] Specifically, the data acquisition process is as follows: During the non-hitting phase of the movement, motion data of each joint is continuously collected; torque signals are obtained from the servo drivers of each joint, and the servo drivers calculate the output torque based on current feedback and motor constants; angular acceleration signals are obtained from the angle encoder data through second-order differential operations.
[0127] To reduce noise introduced by differential operations, the angle data is first subjected to low-pass filtering. In this embodiment, the cutoff frequency of the low-pass filter is set to 50 Hz and the sampling frequency is 1000 Hz.
[0128] The equivalent inertia estimation process is as follows: According to Newton's second law of rotational motion, the inertial torque acting on the joint is equal to the product of the equivalent inertia and the angular acceleration; first, calculate the gravitational torque and Coriolis torque under the current configuration based on the robotic arm dynamics model, and subtract these components from the total torque to obtain the inertial torque components.
[0129] The equivalent inertia is estimated in real time using the recursive least squares method. In each control cycle, the current inertial torque components and angular acceleration data are input into the recursive estimator to update the estimated value of the equivalent inertia. In this embodiment, the forgetting factor of the recursive estimator is set to 0.98.
[0130] The adaptive damping adjustment process is as follows: The equivalent inertia obtained by real-time estimation is used to replace the offline calibration value for the calculation of the damping coefficient under critical damping conditions, so that the damping coefficient is adaptively adjusted with the configuration change.
[0131] In this embodiment, the ballistic locking module includes an online drag coefficient correction unit, used to adaptively adjust the drag coefficient parameters in the ballistic prediction model based on historical shot data. The ballistic locking module also includes an online drag coefficient correction unit, used for:
[0132] After each shot, the positional deviation between the predicted shot coordinates and the actual shot coordinates measured by the vision system is recorded. Based on the positional deviation data from multiple consecutive shots, the drag coefficient parameter in the ballistic prediction model is adjusted using an incremental correction method to gradually reduce the prediction deviation. The corrected drag coefficient parameter is stored and used for subsequent ballistic predictions, achieving self-learning update of the drag coefficient.
[0133] Specifically, the position deviation recording process is as follows: After each shot, the position deviation between the predicted shot point coordinates and the actual shot point coordinates is recorded; the position deviation is a three-dimensional vector, including three components along the longitudinal, lateral and vertical directions of the court; the position deviation data is stored in the historical record buffer, and in this embodiment the buffer capacity is set to store the data of the most recent 20 shots.
[0134] The drag coefficient correction process is as follows: Based on the position deviation data of multiple consecutive shots, analyze the statistical characteristics of the deviation; if the average deviation along the longitudinal direction of the field is positive, it indicates that the predicted landing point is far away, indicating that the current drag coefficient is too small; if the average deviation is negative, it indicates that the predicted landing point is too close, indicating that the current drag coefficient is too large.
[0135] An incremental correction method is used to adjust the drag coefficient parameter, with the correction increment proportional to the mean deviation. To avoid drastic parameter fluctuations caused by single outlier data, outlier removal is performed on the deviation data, eliminating data points whose absolute deviation exceeds three standard deviations. The parameter storage and update process is as follows: the corrected drag coefficient parameter is stored in non-volatile memory for subsequent trajectory prediction; as the number of hits accumulates, the drag coefficient parameter gradually converges to a value that matches the current individual badminton characteristics and environmental conditions, and the prediction deviation gradually decreases.
[0136] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A humanoid robot badminton shuttlecock return control system based on a multi-stage training model, characterized in that, include: The ballistic locking module is used to acquire images of the badminton shuttlecock in flight and extract the rate of change of the windward area of the skirt as the morphological gradient. When the morphological gradient exhibits a bell-shaped curve characteristic, it is determined that the flipping phase has begun. Based on the dynamic expansion noise covariance matrix of the flipping completion degree, the drag coefficient is smoothly adjusted, and the coordinates of the hitting point and the arrival time are output. The thermal avoidance planning module is used to receive the coordinates of the hitting point, obtain the temperature of each joint motor, and calculate the thermal stress index. The zero-space projection matrix is constructed based on the Jacobian matrix of the robotic arm. The thermal repulsion velocity vector is generated according to the thermal stress index and projected onto the zero space to reconstruct the joint angles and output the joint command sequence. The impedance shaping module is used to execute the joint command sequence, with the arrival time as the trigger source. In the time window before the ball is hit, the joint position ring stiffness is adjusted according to the trapezoidal slope, and the velocity ring damping coefficient is adjusted according to the critical damping condition. The impact redirection module is used to estimate the reaction force vector based on the incoming ball speed and swing speed, and uses anisotropic admittance control to guide the backward motion generated by the reaction force, connecting it to the return motion.
2. The humanoid robot badminton return control system based on a multi-stage training model according to claim 1, characterized in that, The process of obtaining the morphological gradient includes: A series of images of the badminton shuttlecock in flight are continuously acquired. Background segmentation is performed on each frame to extract the outline region of the shuttlecock. The windward area of the skirt in the outline region is identified, and the pixel area of the windward area of the skirt is calculated. The difference operation is performed on the skirt pixel area between adjacent frames to obtain the area change rate. The area change rate is output as the morphological gradient.
3. The humanoid robot badminton return control system based on a multi-stage training model according to claim 1, characterized in that, The expansion process of the noise covariance matrix includes: Within a preset sliding time window, the changing trend of the morphological gradient is continuously monitored. When the morphological gradient exceeds a preset starting threshold and exhibits a bell-shaped change characteristic of first increasing and then decreasing within the sliding time window, it is determined that the flipping phase has begun. The absolute value of the morphological gradient within the flipping phase is integrated over time, and the integrated value is normalized with a preset total flipping amount to obtain the flipping completion degree. An expansion factor is calculated based on the flipping completion degree, and the expansion factor monotonically increases with the flipping completion degree. The basic noise covariance matrix is multiplied by the expansion factor to obtain the expanded noise covariance matrix. When the morphological gradient falls back to below a preset ending threshold and remains stable, the flipping phase is determined to end, and the noise covariance matrix is restored to its basic value.
4. The humanoid robot badminton return control system based on a multi-stage training model according to claim 1, characterized in that, The process of obtaining the coordinates of the striking point and the time of arrival includes: The state of the badminton shuttlecock is filtered and estimated using the expanded noise covariance matrix to obtain the position and velocity estimates at the current moment. A trajectory prediction model is established based on the adjusted drag coefficient. The position and velocity estimates are substituted into the trajectory prediction model to extrapolate the flight trajectory of the badminton shuttlecock. The intersection of the flight trajectory and the preset hitting plane is calculated, and the intersection is used as the coordinates of the hitting point. The flight time of the badminton shuttlecock from the current position to the hitting point is calculated based on the flight trajectory, and the arrival time is obtained by combining the current moment.
5. The humanoid robot badminton return control system based on a multi-stage training model according to claim 1, characterized in that, The process of constructing the null projection matrix includes: Obtain the current joint angle configuration of the redundant degree-of-freedom robotic arm; calculate the mapping relationship between the end effector velocity and the joint velocity under the current configuration based on the kinematic model of the robotic arm, and obtain the Jacobian matrix; perform pseudo-inverse operation on the Jacobian matrix to obtain the Jacobian pseudo-inverse matrix; subtract the product of the Jacobian pseudo-inverse matrix and the Jacobian matrix from the identity matrix to obtain the null space projection matrix.
6. The humanoid robot badminton return control system based on a multi-stage training model according to claim 1, characterized in that, The process of obtaining the joint command sequence includes: Read the temperature sensor data of each joint motor; according to the pre-established motor thermal characteristic lookup table, query the remaining heat capacity time under the current torque and speed conditions; compare the remaining heat capacity time with the preset safety threshold time, and normalize to obtain the thermal urgency index of each joint; construct a thermal repulsion potential field based on the thermal urgency index of each joint, calculate the negative gradient of the thermal repulsion potential field, and generate a thermal repulsion velocity vector that promotes each joint to move away from the overheated state. Based on the ball-hitting point coordinates and motion time planning, the desired velocity of the end effector is calculated; the desired velocity of the end effector is mapped to the main task joint velocity using the Jacobian pseudo-inverse matrix; the thermal repulsion velocity vector is projected using the null space projection matrix to obtain the thermal avoidance velocity in the null space; the main task joint velocity is superimposed with the null space thermal avoidance velocity to obtain the composite joint velocity; the composite joint velocity is integrated over time to obtain the joint angle sequence as the joint command sequence output.
7. The humanoid robot badminton return control system based on a multi-stage training model according to claim 1, characterized in that, The adjustment process for the joint position ring stiffness and velocity ring damping coefficient includes: Obtain the electromechanical delay time and the preset stiffness adjustment duration. Determine the adjustment start time by subtracting the electromechanical delay time and stiffness adjustment duration from the arrival time. Obtain the equivalent inertia of each joint in the current configuration. Starting from the adjustment start time, increase the joint position ring stiffness from the initial value to the target peak value according to the trapezoidal ramp law. Calculate the corresponding damping coefficient according to the critical damping condition based on the equivalent inertia and the current stiffness value, so that the damping ratio is maintained in the critical damping and / or overdamped state. After the shot is completed, restore the stiffness and damping coefficient to the initial values.
8. The humanoid robot badminton return control system based on a multi-stage training model according to claim 1, characterized in that, The control process by which anisotropic admittance control guides the reaction force to generate backward motion includes: Based on the ball velocity, swing velocity, and estimated contact time, calculate the direction and peak value of the reaction force vector; establish a local coordinate system related to the reaction force direction to distinguish between the reaction force direction and the anti-collapse direction; set low stiffness parameters and moderate damping parameters in the reaction force direction, and high stiffness parameters in the anti-collapse direction; within the hitting window, read the force and / or torque signals of the lower limb joints; calculate the allowable displacement response in each direction based on the stiffness and damping parameters; synthesize the displacement responses in each direction and convert them into lower limb joint angle correction values, which are then superimposed on the original joint commands; use the controlled backward movement generated by the reaction force direction to connect the subsequent return motion.