A robot noise reduction directional sound pickup method and system, device, medium

By predicting the time and location of the quadruped robot's ground contact and assessing its reliability in real time, and actively suppressing noise interference, the problem of the quadruped robot's landing noise weakening speech recognition was solved. This achieved noise reduction within milliseconds before the noise occurred, thus improving the robustness of speech acquisition.

CN121340373BActive Publication Date: 2026-07-14SICHUAN EMBODIED HUMANOID ROBOT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN EMBODIED HUMANOID ROBOT TECHNOLOGY CO LTD
Filing Date
2025-12-17
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

The strong near-field, high-energy, transient impact noise generated by quadruped robots during walking, jumping, or landing weakens the effectiveness of conventional speech denoising and recognition algorithms. Existing methods lack proactive prediction and systematic motion-acoustic coordination schemes.

Method used

By acquiring the robot's global gait parameters, predicting the time and location of ground contact, establishing a ground contact model, evaluating its reliability in real time, sending control signals to suppress noise source interference, and using a microphone for active noise suppression.

Benefits of technology

Active adjustments can be made milliseconds before noise occurs, reducing ground noise interference and improving the robustness and integrity of voice acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121340373B_ABST
    Figure CN121340373B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of robot control, in particular to a robot noise reduction directional sound pickup method and system, equipment and medium, the present application mainly includes real-time acquisition of the motion parameters of the current robot, based on the current motion parameters, the calibration ground contact time and the calibration ground contact point to obtain the predicted ground contact time, comprehensive evaluation of the credibility of each foot end ground contact prediction based on real-time prediction of the ground contact time and confidence interval to determine the noise expected time window, based on the noise expected time window, send the control signal to the microphone to suppress the ground contact point direction sound source interference. Through the above method, the landing time is predicted by using the multi-modal sensing information of the robot body, and the suppression measures are actively taken within a short time before the impact occurs, the microphone gain, directional sound pickup and adaptive beam control are dynamically adjusted, so that the landing noise interference is effectively reduced while the integrity of the voice interaction is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot control technology, and more specifically, to a method, system, device, and medium for robot noise reduction and directional sound pickup. Background Technology

[0002] During walking, jumping, or landing, quadruped robots generate strong, near-field, high-energy, transient impact noise at their feet. This type of noise is concentrated in energy, short in duration, and often accompanied by structural vibration propagation, making it completely different from traditional far-field, stationary noise. This significantly weakens the effectiveness of conventional speech denoising and recognition algorithms. Existing methods are mostly passive, filtering or suppressing noise only after it has occurred, failing to predict the landing time and noise characteristics in advance, and lacking a systematic motion-acoustic coordination scheme. Therefore, a technical method is needed that can actively predict the landing time, impact intensity, and acoustic propagation characteristics of the feet to improve the robustness of robot speech acquisition in dynamic scenarios. Summary of the Invention

[0003] The purpose of this invention is to provide a method, system, device, and medium for directional sound pickup in robot noise reduction to solve the problems existing in the prior art.

[0004] This invention is achieved through the following technical solution:

[0005] In a first aspect, the present invention also provides a method for directional sound pickup in robot noise reduction, comprising:

[0006] Obtain the current robot's global gait parameters and the phase of each leg, determine the swing range of each leg, obtain the spherical coordinate trajectory of the foot based on the swing range, construct the ground contact model based on the foot height and the ground model, and obtain the calibrated ground contact time and calibrated ground contact point through the ground contact model;

[0007] The robot's motion parameters are acquired in real time, and the predicted contact time is obtained based on the current motion parameters, the calibrated contact time, and the calibrated contact point.

[0008] The reliability of each foot-to-ground-contact prediction was comprehensively evaluated, and confidence intervals were set.

[0009] Based on the real-time predicted ground contact time and confidence interval, a noise prediction time window is determined. Based on the noise prediction time window, a control signal is sent to the microphone to suppress interference from sound sources in the direction of the ground contact point.

[0010] Preferably, the method further includes establishing a spherical coordinate system with the microphone mounted on the robot as the origin, using the coordinates of the robot's foot as the coordinates of the noise source, and obtaining the specific location of the i-th noise source at time t based on the coordinates of the noise source. ,in, Let be the distance from the i-th noise source to the origin. The direction angle between the noise source and the origin of the coordinate system. Let be the pitch angle between the i-th noise source and the origin.

[0011] Preferably, obtaining the spherical coordinate trajectory of the foot based on the swing interval includes:

[0012] The swing range is , At the moment the swing begins, The moment when the swing ends, The current moment;

[0013] The spherical coordinate trajectory is , Let be the position of the foot at time t. Let be the distance from the foot to the origin at time t. Let be the angle between the foot and the origin at time t. Let be the pitch angle between the foot and the origin at time t.

[0014] Preferably, the construction of the ground contact model based on foot height and ground model includes:

[0015] The height of the foot is obtained by interpolating and smoothing the spherical coordinate trajectory: ,in, Let be the height of the foot at time t;

[0016] The ground model mentioned includes The construction of the ground contact model includes:

[0017]

[0018] Let x be the x-coordinate of the contact point in the Cartesian coordinate system of the fuselage. Let y be the vertical coordinate of the contact point in the Cartesian coordinate system of the fuselage. This is a function describing the ground height in the robot's body coordinate system.

[0019] Preferably, the predicted ground contact time based on the current motion parameters, the calibrated ground contact time, and the calibrated ground contact point includes:

[0020] Based on forward kinematics and Jacobi, the current position and velocity of the foot in the spherical coordinate system are obtained, and the relative height, relative velocity, and normal acceleration are obtained by comparing with the ground model.

[0021] Extended Kalman Filter (EKF) The core state is adaptively distinguished between the swinging or support phase. For height, To improve speed, anomalies are suppressed using observation consistency, and an approximate landing time model is solved. The landing time model includes:

[0022]

[0023] In the formula, Relative height, For relative velocity, Normal acceleration, To determine the solution time;

[0024] The predicted ground contact time is obtained based on the solution time. , For the predicted time of ground contact, To determine the time of ground contact.

[0025] Preferably, the setting of the confidence interval includes:

[0026] The time uncertainty is determined based on fitting error, gait cycle fluctuation, ground altitude error, and state estimation covariance. Confidence intervals are constructed based on time uncertainty. , represents the confidence interval coefficient.

[0027] Preferably, it also includes obtaining the impact energy based on the normal acceleration, etc.

[0028]

[0029] Based on the impact energy level and ground contact velocity, the energy characteristics of the foot contact with the ground are predicted, the energy characteristics are output to a microphone, and the predicted ground contact time is corrected by acoustic propagation delay.

[0030]

[0031]

[0032] In the formula, In order to reach the energy level, for Normal velocity at time t, As compensation for the delay, For the speed of sound, for The distance from the foot to the origin of the coordinate system at any given time. This is the revised predicted time of ground contact.

[0033] Secondly, the present invention provides a robot noise reduction directional sound pickup system for performing the above-described robot noise reduction directional sound pickup method, comprising:

[0034] The motion planning layer is configured to acquire the current global gait parameters of the robot and the phase of each leg, determine the swing range of each leg, obtain the spherical coordinate trajectory of the foot based on the swing range, construct a ground contact model based on the foot height and the ground model, and obtain the calibration ground contact time and calibration ground contact point through the ground contact model.

[0035] The state estimation layer is configured to acquire the robot's motion parameters in real time, and based on the current motion parameters, the calibrated ground contact time, and the calibrated ground contact point, to obtain the predicted ground contact time; it comprehensively evaluates the confidence of each foot's ground contact prediction and sets a confidence interval;

[0036] The execution module is configured to determine the noise prediction time window based on the real-time predicted ground contact time and confidence interval, and based on the noise prediction time window, send a control signal to the microphone to suppress the interference of the sound source in the direction of the ground contact point.

[0037] Thirdly, an electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the aforementioned robot noise reduction and directional sound pickup method.

[0038] Fourthly, a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for directional sound pickup in a robot with noise reduction capabilities.

[0039] The technical solution of the present invention has at least the following advantages and beneficial effects:

[0040] This invention mainly includes real-time acquisition of the robot's motion parameters, prediction of the landing time based on the current motion parameters, the calibrated landing time, and the calibrated landing point, comprehensive evaluation of the reliability of each foot's landing prediction, determination of a noise prediction time window based on the real-time predicted landing time and confidence interval, and sending a control signal to the microphone to suppress noise interference from the direction of the landing point based on the noise prediction time window. Through this method, multimodal sensing information from the robot body is used to predict the landing time, and proactive suppression measures are taken shortly before the impact occurs. Microphone gain, directional sound pickup, and adaptive beam control are dynamically adjusted, thereby effectively reducing landing noise interference while ensuring the integrity of voice interaction. Attached Figure Description

[0041] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a schematic diagram showing the location of the transient high-frequency noise source in this invention;

[0043] Figure 2 This is a schematic diagram illustrating the multimodal landing event prediction of the present invention;

[0044] Figure 3 This is a schematic diagram of the noise source direction angle of the present invention;

[0045] Figure 4 This is a schematic diagram of the pitch angle of the noise source in this invention;

[0046] Figure 5 This is a diagram showing the correspondence between the noise sources of this invention and the robot's gait.

[0047] Figure 6 This is a schematic diagram of the walking gait of the present invention;

[0048] Figure 7 This is a schematic diagram of the running gait of the present invention;

[0049] Figure 8 This is a schematic diagram of the intelligent directional sound pickup of the present invention. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0051] The module division in this application is a logical division. In actual application, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be ignored or not executed.

[0052] The independently described modules or sub-modules may or may not be physically separated; they may be implemented in software or hardware, and some modules or sub-modules may be implemented in software, with the processor calling the software to implement the function of these modules or sub-modules, while other modules or sub-modules may be implemented in hardware, such as through hardware circuits. Furthermore, some or all of the modules can be selected to achieve the purpose of this application's solution according to actual needs.

[0053] Please refer to Figures 1-4 This invention provides a method for directional sound pickup in robot noise reduction, comprising:

[0054] S101: Obtain the current robot's global gait parameters and the phase of each leg, determine the swing range of each leg, obtain the spherical coordinate trajectory of the foot based on the swing range, construct the ground contact model based on the foot height and the ground model, and obtain the calibration ground contact time and calibration ground contact point through the ground contact model;

[0055] The contact time and contact point obtained through the contact model can be understood as calibrated data, which are used to compare with the subsequent real-time prediction of the contact time.

[0056] S102: Real-time acquisition of the robot's motion parameters, and prediction of the ground contact time based on the current motion parameters, the calibrated ground contact time, and the calibrated ground contact point;

[0057] S103: Comprehensively evaluate the reliability of each foot contact prediction and set confidence intervals;

[0058] The tactile sensing layer combines plantar force triggering and observation quality to form a confidence score: based on the consistency between the triggering time of the plantar force sensor and the predicted ground contact time, the magnitude of the observation residual, the trace of the covariance matrix, and the quantity and quality of observations, the confidence of the ground contact prediction for each leg is comprehensively evaluated.

[0059] Based on the trigger time of the foot force sensor With predicted ground contact time Consistency, Observation Residual Size The trace of the covariance matrix P In addition to the number of observations N (the number of valid data points within a time window) and the observation quality Q (signal-to-noise ratio SNR), calculate the following 5 confidence functions:

[0060] Time consistency score: , It is the maximum allowable time deviation, on the order of milliseconds.

[0061] State estimation confidence level: residual score , This is the residual attenuation coefficient.

[0062] Covariance score , This is an uncertainty penalty factor used to adjust the weight of the covariance on the score.

[0063] Observation reliability score: Quantitative score , The rate at which trust is established characterizes how quickly confidence levels change. This is the threshold for the amount of data.

[0064] Quality rating , This is the quality normalization factor, used to standardize the original sensor quality indicators to the [0, 1] range.

[0065] Finally, the reliability of the predictions for each leg's ground contact was comprehensively evaluated. Constraints: , , , , To calculate the weights.

[0066] S104: Determine the noise prediction time window based on the real-time predicted ground contact time and confidence interval, and send a control signal to the microphone to suppress the interference of the sound source in the direction of the ground contact point based on the noise prediction time window.

[0067] Preferably, the method further includes establishing a spherical coordinate system with the microphone mounted on the robot as the origin, using the coordinates of the robot's foot as the coordinates of the noise source, and obtaining the specific location of the i-th noise source at time t based on the coordinates of the noise source. ,in, Let be the distance from the i-th noise source to the origin. The direction angle between the noise source and the origin of the coordinate system. Let be the pitch angle between the i-th noise source and the origin.

[0068] In this approach, the robot no longer passively waits for noise to occur before processing it. Instead, it uses its own sensing and control information to predict the moment of impact, thus proactively avoiding impact noise. The robot predicts impact events in a closed loop—from planning to estimation to verification—through a three-tiered information chain, driving the decision-making system with confidence levels. The end result is that the voice acquisition module can proactively adjust itself milliseconds before impact noise occurs, achieving an intelligent leap from "noise post-processing" to "noise pre-prevention."

[0069] In one exemplary embodiment of the present invention, obtaining the spherical coordinate trajectory of the foot based on the swing interval includes:

[0070] The swing range is , At the moment the swing begins, The moment when the swing ends, The current moment;

[0071] Obtain the spherical coordinate trajectory within the future window from the foot trajectory planner. , Let be the position of the foot at time t. Let be the distance from the foot to the origin at time t. Let be the angle between the foot and the origin at time t. Let be the pitch angle between the foot and the origin at time t.

[0072] Secondly, the ground contact model constructed based on foot height and ground model includes:

[0073] After aligning the machine trajectory with the microphone origin in the world frame, interpolate and smooth it in spherical coordinates to obtain the foot height: ,in, Let be the height of the foot at time t;

[0074] The ground model mentioned includes The construction of the ground contact model includes:

[0075]

[0076] in, Let x be the x-coordinate of the contact point in the Cartesian coordinate system of the fuselage. Let y be the vertical coordinate of the contact point in the Cartesian coordinate system of the fuselage. A function describing the ground height in the robot's body coordinate system;

[0077] Calculate the calibration time of contact and calibrated contact points , Let x be the x-coordinate of the contact point in the Cartesian coordinate system of the fuselage. Let y be the vertical coordinate of the contact point in the Cartesian coordinate system of the fuselage. It is a function describing the ground height in the robot's body coordinate system, usually a fixed constant determined by the four sets of robot hardware.

[0078] Rectangular coordinates and spherical coordinates can be converted to each other using coordinate transformation formulas:

[0079]

[0080]

[0081] In the formula, The radial distance at ground contact time. The polar angle at the time of ground contact. The azimuth angle represents the time of ground contact.

[0082] In one exemplary embodiment of the present invention, the predicted ground contact time based on current motion parameters, calibrated ground contact time, and calibrated ground contact point includes:

[0083] The state estimation layer performs real-time correction of the above trajectory through multi-sensor fusion: it completes time synchronization and extrinsic parameter calibration of IMU, joint encoder and foot force sensor, estimates body attitude, velocity and position in inertial navigation extended Kalman filter, and uses support phase zero velocity constraint to suppress drift.

[0084] Based on forward kinematics and Jacobi, the current position and velocity of the foot in the spherical coordinate system are obtained, and the relative height, relative velocity, and normal acceleration are obtained by comparing with the ground model.

[0085] Among them, relative height relative speed , normal acceleration .

[0086] in, This represents the normal velocity.

[0087] Extended Kalman Filter (EKF) The core state is adaptively distinguished between the swinging or support phase. For height, To address the velocity issue, anomalies are suppressed using observation consistency, thereby approximating the minimum positive root of the landing time model with constant acceleration or constant velocity within a few milliseconds to tens of milliseconds before impact. The landing time model includes:

[0088]

[0089] In the formula, Relative height, For relative velocity, Normal acceleration, To determine the solution time;

[0090] The predicted ground contact time is obtained based on the solution time. , For the predicted time of ground contact, To determine the time of ground contact.

[0091] In one exemplary embodiment of the present invention, setting the confidence interval includes:

[0092] The time uncertainty is determined based on fitting error, gait cycle fluctuation, ground altitude error, and state estimation covariance. Confidence intervals are constructed based on time uncertainty. , represents the confidence interval coefficient.

[0093] Based on fitting error gait cycle fluctuations Ground height error Determining the time uncertainty of state estimation covariance .

[0094] This confidence interval is the only reliable time window for the system to determine "when the impulse noise reaches the microphone." All subsequent active noise cancellation actions must strictly adhere to this interval for triggering and duration, achieving true "noise prevention." A higher confidence level S indicates... The smaller the value of C, the narrower the interval, indicating a more aggressive system; the lower the confidence level S, the more aggressive the system. The larger the value, the wider the interval C, and the more conservative the system.

[0095] The aforementioned time uncertainty confidence interval Further refined into confidence intervals for ground noise ,in , This interval represents the range of uncertainty at the moment of ground contact, providing a time window for subsequent acoustic signal processing.

[0096] An exemplary embodiment of the present invention further includes obtaining the impact energy level based on the normal acceleration.

[0097] Based on impact energy level, ground contact velocity, and contact material properties, predict the energy characteristics of foot contact. .

[0098] Specifically, the prediction of foot contact energy characteristics is based on Hertz contact theory, combined with dynamic models and damping extensions, to achieve quantitative calculation of impact energy peak, impact force peak, spectral distribution, main frequency components and energy attenuation characteristics.

[0099] First, calculate the impact energy level. This reflects the kinetic energy ratio. Based on E, With regard to material properties (Young's modulus) Poisson's ratio radius of curvature coefficient of friction coefficient of recovery The Hertzian model is used to calculate the nonlinear contact force. Where d is the penetration depth. , This represents the elastic flexibility parameter of a material. It is used to predict the peak energy of an impact. Maximum Hertz impulse ,in For equivalent Young's modulus, , i represents the robot's foot and j represents the ground.

[0100] The spectral distribution is modeled as a transfer function through modal analysis. ,in , Here is the modal stiffness matrix. Here is the modal damping matrix. Let J be the identity matrix and J be the impulse pulse. The main frequency components are natural frequencies. , For effective inertial mass, Derived from the contact stiffness matrix. Energy decay characteristics are incorporated into the Lankarani-Nikravesh model. Energy dissipation This is the proportional maximum normal strain energy, considering tangential attenuation. , The sliding velocity is used. The final energy decay characteristic over time is modeled as an exponentially decaying oscillation. ,in , .

[0101] The peak impact energy, peak impact force, spectral distribution, main frequency components, and energy attenuation characteristics are measured. The energy characteristics are output to a microphone, and the predicted ground contact time is corrected by acoustic propagation delay.

[0102]

[0103]

[0104] In the formula, In order to reach the energy level, for Normal velocity at time t, As compensation for the delay, For the speed of sound, for The distance from the foot to the origin of the coordinate system at any given time. This is the revised predicted time of ground contact.

[0105] Four-legged sequence Perform gait consistency and support polygon stability checks: Check whether the sequence of leg ground contact times conforms to the temporal constraints of the predetermined gait pattern, and verify whether the interval between adjacent ground contact times is within a reasonable range; simultaneously evaluate the geometric stability of the support polygon (the convex hull formed by the current supporting leg), including support area, centroid projection position, stability margin, etc., to ensure the robot is in a stable support state. Finally, output a complete set of information for each leg:

[0106] This set serves as prior information for subsequent multi-sensor fusion, acoustic signal processing, and noise modeling.

[0107] Based on the prior information provided above, the voice acquisition module can achieve an intelligent leap from "noise post-processing" to "noise pre-prevention": within milliseconds to tens of milliseconds before ground contact, the system predicts the ground contact time. confidence interval Determine the expected noise time window and combine it with energy / spectrum prediction. Impact energy level By acquiring prior information on noise amplitude and frequency characteristics, microphone array parameters can be adjusted in advance before the expected impact noise arrives. Directional pickup technology is used to point the main lobe toward the desired sound source and suppress interference from the point of impact. At the same time, an adaptive beamforming algorithm is used to optimize beam weights in real time based on predicted noise characteristics (including impact energy, spectral distribution, and ground impact position). The optimal directivity configuration is completed before the actual impact noise occurs. This transforms traditional passive noise reduction based on signal post-processing into active noise prevention based on motion state prediction, significantly improving the quality of speech acquisition and system robustness.

[0108] like Figure 5-8 As shown, secondly, based on the existing foot landing point and trajectory prediction, gait detection can be easily achieved, and different gait information corresponding to different motion states of the noise source direction at the current moment can be obtained. Different gait information corresponds to different noise source numbers and directions.

[0109] The ground state corresponds to the direction of the high-frequency noise source, based on the landing time of the robot's foot. and location From the prediction information within the confidence interval, a set of gait information can be obtained. Each set of gait information can yield the accurate set of directions of high-frequency noise sources at the current moment. Define the direction of the target sound source as... The VAD module is responsible for detecting whether a person is speaking and determining the direction of the target sound source. The directional sound pickup system is based on a set of prior information obtained through prediction. Intelligent sound pickup is performed. The access layer predicts the moment the foot touches the ground. Spatial location and the confidence interval of landing noise The parameters are mapped to a spherical coordinate system to generate predictions of noise direction and time window; at the same time, the spatial angle and time range are expanded according to the confidence interval to form a set of multi-directional steering vectors and complete the corresponding delay compensation.

[0110] The timing module sets a lead time based on the predicted ground contact time and applies a window function to ensure that the microphone array completes a smooth switch of beam weights before the impact sound wave arrives, thus achieving forward-looking control in time.

[0111] In terms of spatial processing, the spatial filter module is constrained by the geometry of the microphone array and combined with the target speech direction. With the predicted noise direction set An adaptive beamformer for MVDR is constructed. This beamformer effectively minimizes noise power from the prediction direction while maintaining the target speech gain. For cases with high prediction uncertainty, the system constructs a soft-constrained LCMV by expanding confidence intervals, balancing noise suppression depth and speech signal fidelity through angle and energy weighting.

[0112] The spectrum suppression layer utilizes the predicted impact energy With spectral morphology The system adaptively adjusts the frequency band weights, implementing gain attenuation in advance in the frequency bands where impact noise is expected to occur, achieving energy-aware spectral masking. During operation, the beam weights and frequency band masking coefficients are updated frame by frame to ensure that the array directivity and spectral response are adjusted to the optimal state at the moment the foot touches the ground, thus completing dual pre-suppression of spatial and spectral noise before it actually reaches the microphone. Simultaneously, the system continuously monitors residual noise levels, speech energy, and false suppression indicators, dynamically adjusting parameters through a closed-loop adaptive mechanism to maintain the stability of overall performance.

[0113] Ultimately, this directional sound pickup system, driven by predictive information, organically integrates three major mechanisms: spatial filtering, spectral masking, and confidence scheduling. While ensuring the integrity of the speech signal, it significantly suppresses high-frequency noise generated by landing impacts during dynamic gait, achieving intelligent forward-looking sound pickup that moves from passive post-processing to active prevention.

[0114] Specifically, this includes real-time signal stream acquisition, performance detection feature extraction, decision-making and parameter mapping, and simultaneous adjustment of spatial filtering parameters, spectrum suppression parameters, and confidence scheduling parameters. These parameters are then transmitted to the spatial filtering module, spectrum suppression layer, and prediction access layer, respectively, before outputting the audio stream and providing closed-loop feedback to the real-time signal stream.

[0115] Secondly, the present invention provides a robot noise reduction directional sound pickup system for performing the above-described robot noise reduction directional sound pickup method, comprising:

[0116] The motion planning layer is configured to acquire the current global gait parameters of the robot and the phase of each leg, determine the swing range of each leg, obtain the spherical coordinate trajectory of the foot based on the swing range, construct a ground contact model based on the foot height and the ground model, and obtain the calibration ground contact time and calibration ground contact point through the ground contact model.

[0117] The state estimation layer is configured to acquire the robot's motion parameters in real time, and based on the current motion parameters, the calibrated ground contact time, and the calibrated ground contact point, to obtain the predicted ground contact time; it comprehensively evaluates the confidence of each foot's ground contact prediction and sets a confidence interval;

[0118] The tactile perception layer is used to trigger verification.

[0119] The execution module is configured to determine the noise prediction time window based on the real-time predicted ground contact time and confidence interval, and based on the noise prediction time window, send a control signal to the microphone to suppress the interference of the sound source in the direction of the ground contact point.

[0120] It also includes an intelligent strategy decision-making and control module for predictive information. This module acts as the brain of the system, receiving a complete information flow from the prediction module—including but not limited to: which foot, when, with what confidence level, and the expected noise intensity and spectrum. Based on these inputs, a dynamic strategy decision-maker with embedded expert knowledge and learning capabilities begins to operate. Instead of employing fixed coping strategies, it runs a complex strategy decision tree. The reasoning process of this decision tree simultaneously considers the attributes of the predicted event (e.g., the left hind foot is about to land heavily), the prediction confidence level (high / medium / low), the current state of Voice Activity Detection (VAD) (whether the user is speaking), and the effectiveness of historical strategies. Based on these multi-dimensional inputs, the decision-maker outputs an optimal, or even combined, coping strategy.

[0121] Based on the above, here is a specific example:

[0122] The decision-maker predicts that the right forefoot will land with high confidence in 15 milliseconds, and the VAD is currently detecting that the user is issuing a voice command. At this point, the decision-maker may immediately generate a time-precise command sequence: First, at T-10ms, the command beamforming algorithm dynamically adjusts the weights to generate a deep "null" in the direction of the right forefoot, while ensuring that the main lobe beam is stably pointing towards the user's face; then, at T-2ms, a notch filter parameter set for hard ground impact noise is preloaded; finally, at T-0ms, a hardware command is issued to instantaneously reduce the global acquisition gain of the microphone array by 20dB for 30 milliseconds to prevent preamplifier saturation. In another scenario: with moderate prediction confidence and no user voice, the decision-maker may choose a more conservative strategy, such as only slightly adjusting the beamforming null and initiating a short-term high-speed environmental noise sampling process to learn and update the noise model using this impact event, providing more accurate parameters for future filtering. Through this real-time decision-making based on multi-factor perception and the temporal coordination of multiple strategies, the system can achieve maximum suppression of high-energy impact noise with minimal voice interruption cost.

[0123] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0124] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. This computer software product, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0125] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for directional sound pickup in robot noise reduction, characterized in that, include: Obtain the current robot's global gait parameters and the phase of each leg, determine the swing range of each leg, obtain the spherical coordinate trajectory of the foot based on the swing range, construct the ground contact model based on the foot height and the ground model, and obtain the calibrated ground contact time and calibrated ground contact point through the ground contact model; The robot's motion parameters are acquired in real time, and the predicted contact time is obtained based on the current motion parameters, the calibrated contact time, and the calibrated contact point. The reliability of each foot-to-ground-contact prediction was comprehensively evaluated, and confidence intervals were set. Based on the real-time predicted ground contact time and confidence interval, the noise prediction time window is determined, and based on the noise prediction time window, a control signal is sent to the microphone to suppress the interference of the sound source in the direction of the ground contact point. The predicted ground contact time, based on the current motion parameters, calibrated ground contact time, and calibrated ground contact point, includes: Based on forward kinematics and Jacobi, the current position and velocity of the foot in the spherical coordinate system are obtained, and the relative height, relative velocity, and normal acceleration are obtained by comparing with the ground model. Extended Kalman Filter (EKF) The core state is adaptively distinguished between the swinging or support phase. For height, To improve speed, anomalies are suppressed using observation consistency, and an approximate landing time model is solved. The landing time model includes: In the formula, Relative height, For relative velocity, Normal acceleration, To determine the solution time; The predicted ground contact time is obtained based on the solution time. , This is the predicted time of ground contact after initial correction. To determine the time of ground contact; It also includes obtaining the impact energy level based on the normal acceleration: Based on the impact energy level and ground contact velocity, the energy characteristics of the foot contact with the ground are predicted, the energy characteristics are output to a microphone, and the predicted ground contact time is corrected by acoustic propagation delay. In the formula, In order to reach the energy level, for Normal velocity at time t, As compensation for the delay, For the speed of sound, for The distance from the foot to the origin of the coordinate system at any given time. This is the revised predicted time of ground contact. To determine the time of ground contact, This is the predicted time of ground contact after the initial correction.

2. The robot noise reduction and directional sound pickup method according to claim 1, characterized in that, It also includes establishing a spherical coordinate system with the microphone mounted on the robot as the origin, using the coordinates of the robot's foot as the coordinates of the noise source, and obtaining the specific location of the i-th noise source at time t based on the coordinates of the noise source. ,in, Let be the distance from the i-th noise source to the origin. The direction angle between the noise source and the origin of the coordinate system. Let be the pitch angle between the i-th noise source and the origin.

3. The robot noise reduction and directional sound pickup method according to claim 1, characterized in that, The ball coordinate trajectory of the foot obtained based on the swing interval includes: The swing range is , At the moment the swing begins, The moment when the swing ends, The current moment; The spherical coordinate trajectory is , Let be the position of the foot at time t. Let be the distance from the foot to the origin at time t. Let be the angle between the foot and the origin at time t. Let be the pitch angle between the foot and the origin at time t.

4. The robot noise reduction and directional sound pickup method according to claim 3, characterized in that, The ground contact model constructed based on foot height and ground model includes: The height of the foot is obtained by interpolating and smoothing the spherical coordinate trajectory: ,in, Let be the height of the foot at time t; The ground model mentioned includes The construction of the ground contact model includes: in, Let x be the x-coordinate of the contact point in the Cartesian coordinate system of the fuselage. Let y be the vertical coordinate of the contact point in the Cartesian coordinate system of the fuselage. This is a function describing the ground height in the robot's body coordinate system.

5. A robot noise reduction and directional sound pickup method according to claim 4, characterized in that, The setting of the confidence interval includes: The time uncertainty is determined based on fitting error, gait cycle fluctuation, ground altitude error, and state estimation covariance. Confidence intervals are constructed based on time uncertainty. , represents the confidence interval coefficient.

6. A robot noise reduction directional sound pickup system, characterized in that, A robot noise reduction and directional sound pickup method for performing any one of claims 1-5 includes: The motion planning layer is configured to acquire the current global gait parameters of the robot and the phase of each leg, determine the swing range of each leg, obtain the spherical coordinate trajectory of the foot based on the swing range, construct a ground contact model based on the foot height and the ground model, and obtain the calibration ground contact time and calibration ground contact point through the ground contact model. The state estimation layer is configured to acquire the robot's motion parameters in real time, and based on the current motion parameters, the calibrated ground contact time, and the calibrated ground contact point, to obtain the predicted ground contact time; it comprehensively evaluates the confidence of each foot's ground contact prediction and sets a confidence interval; The execution module is configured to determine the noise prediction time window based on the real-time predicted ground contact time and confidence interval, and based on the noise prediction time window, send a control signal to the microphone to suppress the interference of the sound source in the direction of the ground contact point.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a robot noise reduction and directional sound pickup method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements a robot noise reduction and directional sound pickup method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • CN118280391A

  • US9586316B1