A desktop assistant interaction system and method based on intelligent voice and multi-axis servo linkage

CN122818063APending Publication Date: 2026-09-25HEFEI LINGGUANG YUANQI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610934028.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明要解决的技术问题是:针对现有桌面助理在嵌入式系统上运行时,语音流、TTS流与舵机控制指令之间时间同步精度差,以及多轴舵机运动不平滑、产生的机械噪声干扰语音识别的问题,提供一种能够实现高精度时空同步、动态平滑控制且具备主动消噪特性的桌面助理交互系统及方法

Benefits of technology

(1)微秒级时空同步:通过时空统一同步控制器,动态预测并补偿了嵌入式系统底层的音频驱动延迟、通信总线延迟以及舵机物理响应延迟,实现了TTS语音流(音素级)与舵机动作在时间上的精确对齐,彻底消除了动作与声音脱节的现象。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_137
    Figure SMS_137
Patent Text Reader

Abstract

The application discloses a kind of desktop assistant interaction systems and methods based on intelligent voice and multi-axis steering engine linkage, belong to intelligent voice interaction and robot motion control technical field.The present application solves the problem that the time synchronization precision of voice stream and steering engine control instruction is poor, multi-axis steering engine motion is not smooth and mechanical noise interferes with speech recognition in the embedded system of existing desktop assistant.System includes multi-channel voice acquisition and enhancement module, speech recognition and navigation control module, speech synthesis and phoneme analysis module, space-time unified synchronous controller, multi-axis steering engine smooth trajectory generator and driving module.The present application dynamically compensates audio and steering engine delay through Kalman filtering, realizes microsecond-level phoneme-action synchronization;Adopt five times spline curve to generate smooth trajectory with phoneme energy and acceleration, and actively offset mechanical noise using steering engine state feedback, significantly improve speech recognition rate and action naturalness in dynamic interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent voice interaction and robot motion control technology, specifically relating to a desktop assistant interaction system and method based on intelligent voice and multi-axis servo motor linkage. Background Technology

[0002] With the rapid development of artificial intelligence and embedded systems, desktop assistant robots have been widely used in smart home control, intelligent office assistance, voice navigation, and emotional companionship. These desktop assistants typically integrate microphone arrays, speakers, and multi-axis motion mechanisms (such as multi-axis servos). They interact with users bidirectionally through automatic speech recognition (ASR) and text-to-speech (TTS) technologies, and utilize the physical movements of multi-axis servos (such as head rotation and body tilting) to enhance the vividness and human-like experience of the interaction.

[0003] However, in existing desktop assistant interaction systems, the limited computing resources of embedded hardware platforms lead to severe asynchrony issues between voice stream processing, TTS speech synthesis output, and the issuance of multi-axis servo control commands. Because the computational time for voice enhancement algorithms, speech recognition decoding, semantic understanding, and TTS synthesis on embedded systems is dynamically fluctuating, while servo control commands are typically issued periodically, the robot's limb movements and emitted sounds often cannot be precisely synchronized, resulting in a sense of disharmony where "the movement stops before the voice finishes" or "the movement lags behind the sound." Furthermore, traditional servo control often employs piecewise linear interpolation or simple PID control, which is highly susceptible to transient shocks and joint jitter during multi-axis linkage. This not only reduces the smoothness and anthropomorphism of the movements but also generates high-frequency mechanical vibration noise that can be back-coupled into the microphone array, severely interfering with the performance of the voice enhancement algorithm and causing a sharp drop in speech recognition rate.

[0004] Therefore, solving the problem of precise time synchronization between voice streams, TTS streams, and motor control commands in embedded systems, and achieving smooth, jitter-free trajectory control of multi-axis servos to avoid interference from mechanical noise on voice enhancement and speech recognition, is a bottleneck problem that urgently needs to be solved in the field of desktop assistant interaction. Solving this problem is of paramount necessity and urgency for improving the naturalness of desktop assistant interaction and increasing the accuracy of speech recognition in noisy and dynamic conditions. Summary of the Invention

[0005] The technical problem to be solved by this invention is: to address the problems of poor time synchronization accuracy between voice stream, TTS stream and servo control commands when existing desktop assistants are running on embedded systems, as well as the problem of unsmooth multi-axis servo movement and mechanical noise interference with voice recognition, and to provide a desktop assistant interaction system and method that can achieve high-precision spatiotemporal synchronization, dynamic smooth control and active noise cancellation characteristics.

[0006] A desktop assistant interaction system based on intelligent voice and multi-axis servo motor linkage, the system comprising: The multi-channel voice acquisition and enhancement module is used to acquire user voice in real time through a microphone array, and to suppress environmental noise using an adaptive spatial beamforming algorithm. At the same time, combined with the real-time motion status feedback of the multi-axis servo, the mechanical noise generated by the servo motion is actively filtered out through an adaptive cancellation filter, and the enhanced voice stream is output. The voice recognition and navigation control module is used to extract and decode features from the enhanced voice stream, recognize the user's voice commands, and generate voice navigation data or control decisions based on the commands. The speech synthesis and phoneme analysis module is used to convert the text to be output into a TTS speech stream and to analyze the start time, duration and energy characteristics of each phoneme in real time. A unified time-space synchronization controller, deployed on an embedded system, is used to dynamically estimate the rendering latency of the TTS stream and the mechanical response latency of the servo motor, establish a unified time reference, and align the TTS voice stream and servo motor action commands at the microsecond level on the time axis. The multi-axis servo smooth trajectory generator is used to generate a jitter-free and highly smooth multi-axis servo linkage control trajectory based on the alignment command issued by the spatiotemporal unified synchronization controller, combined with the phoneme energy characteristics, and using a dynamic tension-limited quintic spline curve algorithm. The multi-axis servo drive module is used to convert smooth trajectories into high-frequency PWM control signals to drive multi-axis servos to perform corresponding linkage actions.

[0007] Furthermore, in the multi-channel voice acquisition and enhancement module, a delayed summation beamformer is used to perform spatial filtering on the target user's direction. Simultaneously, a reference noise source vector composed of the current angular velocity and angular acceleration of the multi-axis servo motor is constructed. Mechanical noise is estimated through an adaptive filter, and the estimated mechanical noise is subtracted from the beamformed signal. The normalized least mean square (NLMS) algorithm is used to update the weight coefficients of the adaptive filter in real time to achieve active cancellation of mechanical noise.

[0008] Furthermore, in the speech synthesis and phoneme parsing module, while synthesizing audio, a forced alignment algorithm is used to parse the identifier, start time, end time and phoneme duration of each phoneme in the synthesized speech, and the average short-time energy within each phoneme interval is calculated.

[0009] Furthermore, in the spatiotemporal unified synchronization controller, a state transition equation and an observation equation are established with audio rendering delay and servo physical response delay as state vectors. The optimal delay estimate is calculated dynamically and iteratively using a Kalman filter. The compensation offset of the action command relative to the audio stream is calculated, and the audio data frame and servo control frame are delayed or muted and aligned through a dual-path ring synchronization buffer.

[0010] Furthermore, in the multi-axis servo smooth trajectory generator, a fifth-order polynomial trajectory equation with zero velocity and acceleration at the start and end points is constructed using the phoneme duration as the boundary condition. At the same time, a nonlinear time scaling function based on phoneme energy is introduced to integrate and resample the physical time to obtain a virtual time axis, and the virtual time axis is substituted into the fifth-order polynomial trajectory equation to generate a control trajectory with continuous jerk.

[0011] As a further aspect of the present invention, an interaction method for a desktop assistant interaction system based on intelligent voice and multi-axis servo motor linkage is provided, the method comprising the following steps: Multi-channel adaptive voice enhancement steps: acquire multiple raw voice signals through a microphone array, use the current motion state feedback of the multi-axis servo to dynamically adjust the tap coefficients of the adaptive cancellation filter, actively filter out servo operation noise from the raw voice signal, and complete the voice enhancement. Speech recognition and semantic navigation steps: Extract acoustic features from the enhanced speech signal, perform speech recognition through a deep neural network, and combine the context for semantic understanding to generate navigation or interactive control commands; TTS phoneme-level feature extraction steps: Input the interactive response text into the TTS engine, and extract the phoneme sequence and its corresponding timestamp interval and instantaneous energy value while synthesizing the audio stream; Spatiotemporal unified synchronization alignment steps: Dynamically predict the current audio output delay and servo communication and physical response delay of the embedded system using the Kalman filter algorithm, calculate the compensation offset of the action command relative to the audio stream, and align the voice stream and action command in the circular synchronization buffer. Multi-axis servo motor dynamic smooth trajectory planning steps: Based on the aligned time nodes, with phoneme duration as the boundary condition and phoneme energy value as the speed scaling factor, construct a fifth-order polynomial trajectory equation, and solve for the multi-axis servo motor joint motion trajectory with continuous acceleration and minimum jerk. Linked execution steps: The planned smooth trajectory is sent to the multi-axis servo drive module in the form of high-frequency control frames, driving the desktop assistant to make smooth and natural linked body movements while making a sound.

[0012] Furthermore, the multi-channel speech adaptive enhancement step specifically includes: Step 7.1: Use a delay summation beamformer to perform spatial filtering on the target user direction to obtain a spatially filtered signal; Step 7.2: Collect the current angular velocity and angular acceleration of the multi-axis servo motor and construct a reference noise source vector; Step 7.3: Input the reference noise source vector into the adaptive filter to calculate the estimated mechanical noise; Step 7.4: Subtract the estimated mechanical noise from the spatially filtered signal to obtain the enhanced speech signal; Step 7.5: Using the Normalized Least Mean Square (NLMS) algorithm, update the weight coefficients of the adaptive filter in real time based on the enhanced speech signal and the reference noise source vector.

[0013] Furthermore, the spatiotemporal unified synchronization alignment step specifically includes: Step 8.1: In the audio DMA transfer interrupt service routine, read the hardware timer timestamp, subtract it from the timestamp written to the audio buffer by the upper layer, and obtain the audio rendering delay observation value; Step 8.2: Send a position query command to the servo motor via the communication bus, read the actual motion start timestamp fed back by the servo motor's built-in encoder, subtract the command issuance timestamp from the timestamp, and obtain the servo motor response delay observation value; Step 8.3: Use a Kalman filter to perform state estimation on audio rendering latency and servo response latency, and output the optimal latency estimate; Step 8.4: Calculate the difference between the two as the compensation offset. If the audio path delay is greater than the servo path delay, the servo control command will be delayed in the action ring buffer. If the servo response delay is greater than the audio path delay, a mute data frame will be inserted into the audio ring buffer.

[0014] Furthermore, the specific steps for the dynamic smooth trajectory planning of the multi-axis servo motor are as follows: Step 9.1: Set the starting angle and target angle of the single-axis servo motor in the current phoneme cycle, as well as the velocity and acceleration boundary conditions at the starting and ending points. Solve the coefficients of the fifth-degree polynomial equation and establish the initial trajectory equation with continuous acceleration. Step 9.2: Obtain the normalized energy value of the current phoneme and construct a nonlinear time scaling function using the hyperbolic tangent function; Step 9.3: Integrate the nonlinear time scaling function over time to generate a resampled virtual time axis; Step 9.4: Map the virtual time axis to the initial trajectory equation to generate a smooth control trajectory with continuous jerk and deep coupling with phoneme energy.

[0015] Furthermore, the speech recognition and semantic navigation steps are specifically as follows: Step 10.1: Perform frame segmentation and windowing on the enhanced speech signal, calculate the Mel-frequency cepstral coefficients (MFCC) of each frame, and obtain the feature sequence; Step 10.2: Input the feature sequence into the pre-trained acoustic model to calculate the posterior probability of the state, and combine it with the language model to decode using the Viterbi search algorithm, and output the recognized text sequence; Step 10.3: Perform intent recognition and slot filling on the text sequence to extract the action intent and navigation intent, and generate the corresponding servo control target angle and voice navigation path.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) Microsecond-level time and space synchronization: Through the unified time and space synchronization controller, the audio drive delay, communication bus delay and servo physical response delay of the embedded system are dynamically predicted and compensated, realizing the precise alignment of TTS voice stream (phoneme level) and servo action in time, and completely eliminating the phenomenon of action and sound being out of sync.

[0017] (2) Extremely smooth motion control: The multi-axis servo smooth trajectory generator introduces the minimum jerk optimization algorithm and dynamically adjusts the motion speed with phoneme energy, making the servo rotation extremely smooth and natural, eliminating the transient shock and joint vibration caused by traditional control methods.

[0018] (3) Active noise cancellation and improved speech recognition rate: Due to the smoothing and optimization of the servo motor's motion trajectory, the mechanical vibration and high-frequency electromagnetic noise generated are significantly reduced; at the same time, the system feeds back the servo motor's motion status to the speech enhancement module in real time, realizing active noise cancellation. The synergistic effect of these two factors improves the recognition accuracy of the user's secondary voice commands by more than 30% during the dynamic interaction process of the desktop assistant while speaking and moving, effectively solving the industry problem of motion noise interfering with voice interaction. Detailed Implementation

[0019] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] This invention provides a desktop assistant interaction system and method based on intelligent voice and multi-axis servo linkage, aiming to solve the problem of microsecond-level precise time synchronization between voice stream, TTS stream and multi-axis servo control commands in embedded systems, and to eliminate the interference of mechanical noise generated by multi-axis servo movement on voice enhancement and voice recognition.

[0021] I. Overall System Hardware and Software Architecture Design: In this embodiment, the desktop assistant interaction system is deployed on a highly integrated embedded hardware platform. This platform uses a dual-core heterogeneous processor (e.g., an ARM Cortex-A7 @ 1.2GHz for running an embedded Linux system to handle high-computing tasks such as speech recognition, semantic understanding, and TTS synthesis; and an ARM Cortex-M7 @ 400MHz real-time kernel for running an RTOS to handle precise trajectory control of multi-axis servos, Kalman filter delay estimation, and high-frequency sensor data acquisition).

[0022] The system hardware peripherals include: 1. Microphone array: A 4-channel MEMS microphone array with uniform ring distribution and a spacing of d=40mm is used to collect spatial multi-channel voice signals.

[0023] 2. Audio output unit: integrates a high-performance audio codec chip and a Class D audio amplifier, connecting to the desktop assistant's built-in speakers.

[0024] 3. Multi-axis motion mechanism: Includes three high-precision serial bus intelligent servos deployed in the desktop assistant's head (yaw axis, pitch axis) and body (roll axis). The servos have built-in 12-bit absolute magnetic encoders, supporting real-time position, speed, current, and temperature feedback, and are connected to the processor via RS-485 bus. II. Specific functions and underlying algorithm implementation of each functional module: 1. Specific implementation of the multi-channel voice acquisition and enhancement module: This module is deployed in a real-time kernel, and its core task is to filter out environmental noise and high-frequency mechanical vibration noise generated by the servo gearbox in real time during the dynamic interaction of the multi-axis servo motor's wide range of movements.

[0025] (1) Spatial Beamforming: Suppose that the four raw speech signals acquired by the microphone array are represented in the frequency domain as follows: The system utilizes a delay-and-sum beamformer to target the user's direction. Spatial filtering is performed to suppress ambient noise from non-target directions: ,in The steering vector, whose m-th component is defined as: ,in, This is the audio sampling rate (16kHz in this example). Let be the physical propagation delay of the sound wave reaching the m-th microphone relative to the reference center point.

[0026] (2) Adaptive noise cancellation (ANC) based on servo state feedback: Since multi-axis servos generate mechanical noise that is highly correlated with their speed and acceleration when rotating, this invention introduces an adaptive noise cancellation algorithm based on servo state feedback.

[0027] Let N be the number of axes of the servo motor (N=3 in this embodiment). The system reads the current angular velocity of each axis of the servo motor in real time at a period of 1ms via RS-485 bus. and angular acceleration Considering the physical propagation delay of mechanical vibration transmitted through structural components to the microphone array. The system constructs a reference noise source vector that includes delay compensation. : Among them, the delay in propagation. Based on the physical distance between the servo motor and the microphone array and solid sound velocity Preliminary determination, the calculation formula is: .

[0028] The output of the adaptive filter (i.e., the estimated mechanical noise) is expressed in the time domain as: ,in This is the weight coefficient vector of the adaptive filter.

[0029] From the time-domain signal after beamforming Subtract estimated noise The enhanced speech signal is obtained. : .

[0030] The weight coefficients are updated in real time using the Normalized Least Mean Square (NLMS) algorithm: ,in, This is the step size factor, with a value range of [0.01, 0.1], used to balance convergence speed and steady-state error; This is a regularization parameter to prevent the denominator from being zero when the servo is stationary (reference vector is zero). Through this closed-loop feedback, mechanical noise during motion can be dynamically and accurately filtered out.

[0031] 2. Specific implementation of the voice recognition and navigation control module: This module receives the enhanced voice signal. It also completes feature extraction, text decoding, and navigation decision-making.

[0032] (1) Acoustic feature extraction: The system The process involves framing (25ms frame length, 10ms frame shift) and applying a Hanning window. For each frame, a Fast Fourier Transform (FFT) is performed, and the signal is passed through a Mel-Filter Bank of 40 filters to calculate the logarithmic energy. Finally, a Discrete Cosine Transform (DCT) is performed to extract the 13-dimensional Mel-frequency cepstral coefficients (MFCCs).

[0033] To capture the dynamic features of speech, its first-order difference (Delta) and second-order difference (Delta-Delta) are calculated and finally concatenated into a 39-dimensional acoustic feature vector. .

[0034] (2) Acoustic model decoding and semantic understanding: feature sequence The input is fed into a pre-trained deep neural network-hidden Markov model (DNN-HMM) acoustic model to calculate the posterior probability of the state. Combined with a language model, the Viterbi search algorithm is used for decoding, outputting the recognized text sequence. .

[0035] The navigation control unit performs intent parsing on the text sequence W. For example, when it recognizes "Please turn left and navigate to the kitchen," the semantic parsing engine extracts: Action intent: TURN_ANGLE, slot: direction=left, angle=45°; Navigation intent: NAVIGATE_TO, slot: destination=kitchen.

[0036] 3. Specific implementation of the speech synthesis and phoneme parsing module: When the system needs to provide voice broadcasts or interactive responses, this module converts text into speech and extracts phoneme features to guide servo movements.

[0037] (1) TTS speech synthesis: Using an end-to-end speech synthesis model (such as Tacotron2 combined with WaveGlow synthesizer), the response text is synthesized. Convert to time-domain audio waveform .

[0038] (2) Phoneme alignment and analysis: While synthesizing audio, a forced alignment algorithm is used to extract each phoneme from the synthesized speech. The boundary information. Output a set of phoneme sequences: .in, Let i be the identifier for the i-th phoneme. This is the start time of the phoneme. The end time is the duration of the phoneme. .

[0039] The average short-time energy within this phoneme range is calculated using the following formula: .in, These are discretized TTS audio sample values. , , This represents the total number of sampling points within the phoneme interval.

[0040] 4. Specific implementation of the spatiotemporal unified synchronization controller: In embedded systems, the audio playback path (from user-space memory to the ALSA driver, DMA buffer, and then to the DAC chip for sound output) experiences dynamically fluctuating rendering latency. Meanwhile, there is a response delay in the servo control path (from sending commands via serial port and bus transmission to the servo controller decoding and driving the motor to rotate). This module dynamically estimates these two delays using a Kalman filter, achieving microsecond-level spatiotemporal alignment.

[0041] (1) Delayed state estimation: Define the system state vector .

[0042] The state transition equation is: .

[0043] in, (Identity matrix) This represents process noise. Its covariance matrix... Based on the current CPU utilization of the embedded system Dynamically adjust to reflect the impact of system load on latency fluctuations: Where q0 and q1 are the variances of the basic process noise. This is the load sensitivity coefficient.

[0044] The observation equation is: Among them, the observation vector .

[0045] A. The acquisition method is as follows: Read the hardware timer timestamp in the audio DMA transfer interrupt service routine. The timestamp of the audio buffer written by the upper layer Subtraction, that is .

[0046] B. Acquisition method: Send a position query command to the servo via RS-485 bus to read the actual motion start timestamp fed back by the servo's built-in encoder. With the command issuance timestamp Subtraction, that is .

[0047] Observation matrix , To observe the noise covariance.

[0048] Using the standard Kalman filter recursive formula (prediction and update), the optimal delay estimate is output through iterative calculation within each control cycle. and .

[0049] (2) Circular buffer alignment mechanism: Calculate the compensation offset ΔT of the action command relative to the audio stream: .

[0050] The system constructs a dual-path ring buffer to store audio data frames and servo control frames respectively.

[0051] A. If ΔT>0, it means the audio path delay is greater than the servo path delay. In this case, the spatiotemporal unified synchronization controller will delay the servo control command generated at the current moment in the action ring buffer by ΔT time before issuing it.

[0052] B. If ΔT < 0, it indicates that the servo response is slow. In this case, the controller inserts mute frames of length |ΔT| into the audio ring buffer through the audio playback interface (such as ALSA API), artificially delaying the audio output to ensure that the synchronization error between the acoustic signal and the physical action is within ±5ms.

[0053] 5. Specific implementation of the multi-axis servo smooth trajectory generator: To completely eliminate servo jitter during startup and shutdown, and to integrate the rhythm of the motion with the energy fluctuations of the voice phonemes, this module employs fifth-order polynomial trajectory planning based on minimum jerk and introduces phoneme energy for dynamic time resampling.

[0054] (1) Establishment and rigorous solution of the fifth-order polynomial trajectory equation: Suppose that a certain axis servo motor has an initial angle of θ0 and a target angle of θ within the current phoneme period t∈[0,T]. fTo ensure continuous acceleration during joint movement, and that the joints are stationary at both the start and end points (i.e., both velocity and acceleration are zero), we construct the following fifth-order polynomial trajectory equation: Find the first derivative of the velocity with respect to time t. ) and second derivative (acceleration) ): ; ; Its boundary condition constraints are: At the starting point t=0: ; At the endpoint t=T: ; The following is a rigorous algebraic solution: Substituting the boundary conditions at t=0 into the equation: ,have to ; ,have to ; ,have to ; This determined the first three coefficients: , , .

[0055] Substituting the calculated coefficients and the boundary conditions of t=T, we obtain the following about , , A system of three linear equations in three variables: ; Solve this system of equations: First, divide both sides of equation 2 by Divide both sides of equation 3 by : ; By transforming equation 2', we can obtain: ; Multiplying both sides by 2 gives Substituting this expression into equation 3': , simplify: ; Bundle Substitute into equation 2': , ; Now , Substituting the expression into equation 1: ; Extract common factors : ; Find a common denominator for the value inside the parentheses: ; then: ; Will Substitute return, seek ; .

[0056] All coefficients , , , , , Substituting back into the original equation, we obtain the unique equation for a smooth trajectory with continuous acceleration and minimal jerk: .

[0057] (2) Dynamic temporal resampling based on phoneme energy: To link the movement speed of a multi-axis servo motor with the intensity of speech (phoneme energy), this invention introduces a dynamic time scaling factor. .

[0058] Let the normalized energy of the currently emitted speech phoneme at time t be E(t)∈[0,1]. Define a nonlinear time scaling function: .

[0059] in: This is the energy sensitivity coefficient, used to control the magnitude of the effect of energy changes on velocity (in this embodiment, it is taken as...). =0.5); For nonlinear scaling factors, the hyperbolic tangent function is used. The saturation characteristics prevent speed loss due to excessive phoneme energy (in this embodiment, it is taken as...). =2.0); Based on the base speed offset, this ensures that the servo maintains a basic, smooth transition speed during silent or low-energy tones (in this embodiment, we take...). =0.8).

[0060] pass Integrating over physical time t yields the resampled virtual time axis. : Virtual time Substituting these equations into the fifth-order polynomial trajectory equation obtained above, we obtain the final multi-axis servo control trajectory, which is deeply coupled with voice energy: ; in, This is the virtual total time at the end of the entire phoneme cycle.

[0061] The trajectory generated by this method has a completely continuous jerk (Jerk, i.e., the third derivative of position) throughout the entire motion cycle, eliminating the step impact in the start and end phases of traditional control algorithms and suppressing the generation of mechanical noise from the physical source.

[0062] 6. Specific implementation of the multi-axis servo drive module: This module is responsible for processing the continuous angle values ​​output by the smooth trajectory generator. It is converted into a control signal that the actuator can recognize.

[0063] (1) High-frequency interpolation discretization: Because the transmission cycle of the embedded control bus is usually a fixed value The drive module controls the continuous trajectory Perform equal-interval sampling to generate a discrete sequence of target angles: .

[0064] (2) Protocol encapsulation and distribution: Target angle for each axis The data is encapsulated as a serial bus servo protocol data frame. The data frame format is defined in the table below: (3) Hardware driver execution: The packaged data frame is sent to the multi-axis servo motor at a frame rate of 100Hz via the UART interface of the embedded chip. After receiving the data frame, the microcontroller inside the servo motor drives the brushless DC motor through the built-in 200Hz high-frequency PWM generator, so that the servo motor rotates precisely and smoothly to the specified angle. III. Example of a closed-loop workflow for a complete interactive method: To more clearly demonstrate the actual operation of this invention, the following uses a specific interactive scenario as an example to explain in detail the complete execution steps of the method of this invention: Scene setting: The desktop assistant is in a static state.

[0065] The user issues a voice command from a position 45° to the right of the desktop assistant: "Xiao Zhi, turn right and check the kitchen for me." Step 1: Multi-channel adaptive speech enhancement: 1.4-channel microphone array for real-time acquisition of noisy speech signals .

[0066] 2. The space beamformer is based on the estimated user direction. Spatial filtering is performed to enhance the acoustic signal in that direction.

[0067] 3. Since the servo is stationary at this time, refer to the noise source vector. Adaptive filter output The system outputs enhanced voice that is unaffected by servo noise. .

[0068] Step 2: Speech Recognition and Semantic Navigation 1. The system has the following characteristics: Frame segmentation and windowing are performed to extract 39-dimensional MFCC feature sequences. .

[0069] 2. The acoustic model and decoder decode the feature sequence into text: “Xiao Zhi, turn right and check the kitchen for me.” 3. The semantic understanding module identifies the user's dual intent: Action intention: Turn 45° to the right (corresponding to the target angle of the yaw axis servo) Starting angle ).

[0070] Navigation intent: Query kitchen status (triggers camera activation and image recognition).

[0071] Step 3: TTS phoneme-level feature extraction: 1. The system generates the response text: "Okay, redirecting you to the kitchen, please wait." 2. The TTS engine synthesizes text into an audio stream. Simultaneously, a forced alignment algorithm is used to parse the phoneme sequence. For example, for the phoneme "zhuan," its time interval is parsed as follows: Duration And calculate the normalized short-time energy within this interval. .

[0072] Step 4: Spatiotemporal synchronization and alignment: 1. The Kalman filter runs in real time. The optimal estimate of the current audio rendering latency is measured during the audio DMA interrupt. The optimal estimate of the servo motor's physical response delay was obtained by measuring the RS-485 bus. .

[0073] 2. Calculate the compensation offset .

[0074] 3. Due to The time-space unified synchronization controller delays the rotation command of the yaw axis servo by 85ms in the action ring buffer, ensuring that the servo begins to physically rotate just as the speaker emits the word "turn".

[0075] Step 5: Dynamic smooth trajectory planning for multi-axis servos: 1. The trajectory generator is designed for yaw axis servos, with... For a period of time, , Using the boundary conditions, solve the fifth-degree polynomial to obtain the initial smooth trajectory.

[0076] 2. Utilizing the real-time energy of phoneme "turns" Calculate the time scaling factor When the sound energy is high (such as when producing a plosive sound), As the sound energy increases, virtual time passes faster, and the servo speed increases; when the sound energy is low, the servo speed decreases.

[0077] 3. Through the analysis of Integral generation virtual timeline The initial trajectory is resampled to generate the final control trajectory. The jerk of this trajectory is continuous and without abrupt changes.

[0078] Step 6: Linked Execution and Active Noise Cancellation Closed Loop: 1. The driver module will The control commands, discretized into 10ms frames, are encapsulated into protocol frames and sent to the yaw axis servo to drive the desktop assistant to smoothly turn to the right.

[0079] 2. As the servo motor rotates, its angular velocity... and angular acceleration Real-time feedback is provided to the multi-channel voice acquisition and enhancement module.

[0080] 3. The Adaptive Noise Cancellation (ANC) algorithm utilizes this feedback to dynamically adjust the filter weights. It actively filters out the weak gear noise generated by rotation from the microphone input.

[0081] 4. During this period, if the user suddenly interrupts and issues a second command (such as "stop"), the voice recognition module can still accurately recognize the interruption command with an accuracy rate of over 95% because the mechanical noise has been actively filtered out, and the movement will stop immediately.

[0082] Those skilled in the art should understand that the above embodiments are merely preferred embodiments of the present invention. The specific parameters in the formulas (such as Kalman filter parameters, time scaling factors, etc.) can be fine-tuned according to specific hardware selections, and these fine-tunings do not depart from the protection scope of the present invention. Those skilled in the art will understand that the above content is merely a preferred embodiment of the present invention; all parameters in the formulas (Kalman filter parameters, time scaling factors, etc.) can be fine-tuned according to the hardware model, and such modifications fall within the protection scope of the present invention.

[0083] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.

[0084] It should be understood that the present invention is not limited to the precise structure described above, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A desktop assistant interaction system based on intelligent voice and multi-axis servo motor linkage, characterized in that, The system includes: The multi-channel voice acquisition and enhancement module is used to acquire user voice in real time through a microphone array, and to suppress environmental noise using an adaptive spatial beamforming algorithm. At the same time, combined with the real-time motion status feedback of the multi-axis servo, the mechanical noise generated by the servo motion is actively filtered out through an adaptive cancellation filter, and the enhanced voice stream is output. The voice recognition and navigation control module is used to extract and decode features from the enhanced voice stream, recognize the user's voice commands, and generate voice navigation data or control decisions based on the commands. The speech synthesis and phoneme analysis module is used to convert the text to be output into a TTS speech stream and to analyze the start time, duration and energy characteristics of each phoneme in real time. A unified time-space synchronization controller, deployed on an embedded system, is used to dynamically estimate the rendering latency of the TTS stream and the mechanical response latency of the servo motor, establish a unified time reference, and align the TTS voice stream and servo motor action commands at the microsecond level on the time axis. The multi-axis servo smooth trajectory generator is used to generate a jitter-free and highly smooth multi-axis servo linkage control trajectory based on the alignment command issued by the spatiotemporal unified synchronization controller, combined with the phoneme energy characteristics, and using a dynamic tension-limited quintic spline curve algorithm. The multi-axis servo drive module is used to convert smooth trajectories into high-frequency PWM control signals to drive multi-axis servos to perform corresponding linkage actions.

2. The desktop assistant interaction system based on intelligent voice and multi-axis servo motor linkage according to claim 1, characterized in that, In the multi-channel voice acquisition and enhancement module, a delay-summing beamformer is used to perform spatial filtering on the target user's direction. Simultaneously, a reference noise source vector composed of the current angular velocity and angular acceleration of the multi-axis servo motor is constructed. Mechanical noise is estimated through an adaptive filter, and the estimated mechanical noise is subtracted from the beamformed signal. The normalized least mean square (NLMS) algorithm is used to update the weight coefficients of the adaptive filter in real time to achieve active cancellation of mechanical noise.

3. The desktop assistant interaction system based on intelligent voice and multi-axis servo motor linkage according to claim 1, characterized in that, In the speech synthesis and phoneme analysis module, while synthesizing audio, a forced alignment algorithm is used to parse the identifier, start time, end time and duration of each phoneme in the synthesized speech, and the average short-time energy within each phoneme interval is calculated.

4. A desktop assistant interaction system based on intelligent voice and multi-axis servo motor linkage according to claim 1, characterized in that, In the spatiotemporal unified synchronization controller, a state transition equation and an observation equation are established with audio rendering delay and servo physical response delay as state vectors. The optimal delay estimate is calculated dynamically and iteratively using a Kalman filter. The compensation offset of the action command relative to the audio stream is calculated, and the audio data frame and servo control frame are delayed or muted and aligned by a dual-path ring synchronization buffer.

5. A desktop assistant interaction system based on intelligent voice and multi-axis servo motor linkage according to claim 1, characterized in that, In the multi-axis servo smooth trajectory generator, a fifth-order polynomial trajectory equation with zero velocity and acceleration at the start and end points is constructed using the phoneme duration as the boundary condition. At the same time, a nonlinear time scaling function based on phoneme energy is introduced to integrate and resample the physical time to obtain a virtual time axis. The virtual time axis is then substituted into the fifth-order polynomial trajectory equation to generate a control trajectory with continuous jerk.

6. A desktop assistant interaction method based on intelligent voice and multi-axis servo motor linkage, based on the interactive system described in any one of claims 1-5, characterized in that, The method includes the following steps: Multi-channel adaptive voice enhancement steps: acquire multiple raw voice signals through a microphone array, use the current motion state feedback of the multi-axis servo to dynamically adjust the tap coefficients of the adaptive cancellation filter, actively filter out servo operation noise from the raw voice signal, and complete the voice enhancement. Speech recognition and semantic navigation steps: Extract acoustic features from the enhanced speech signal, perform speech recognition through a deep neural network, and combine the context for semantic understanding to generate navigation or interactive control commands; TTS phoneme-level feature extraction steps: Input the interactive response text into the TTS engine, and extract the phoneme sequence and its corresponding timestamp interval and instantaneous energy value while synthesizing the audio stream; Spatiotemporal unified synchronization alignment steps: Dynamically predict the current audio output delay and servo communication and physical response delay of the embedded system using the Kalman filter algorithm, calculate the compensation offset of the action command relative to the audio stream, and align the voice stream and action command in the circular synchronization buffer. Multi-axis servo motor dynamic smooth trajectory planning steps: Based on the aligned time nodes, with phoneme duration as the boundary condition and phoneme energy value as the speed scaling factor, construct a fifth-order polynomial trajectory equation, and solve for the multi-axis servo motor joint motion trajectory with continuous acceleration and minimum jerk. Linked execution steps: The planned smooth trajectory is sent to the multi-axis servo drive module in the form of high-frequency control frames, driving the desktop assistant to make smooth and natural linked body movements while making a sound.

7. A desktop assistant interaction method based on intelligent voice and multi-axis servo motor linkage according to claim 6, characterized in that, The multi-channel speech adaptive enhancement steps are as follows: Step 7.1: Use a delay summation beamformer to perform spatial filtering on the target user direction to obtain a spatially filtered signal; Step 7.2: Collect the current angular velocity and angular acceleration of the multi-axis servo motor and construct a reference noise source vector; Step 7.3: Input the reference noise source vector into the adaptive filter to calculate the estimated mechanical noise; Step 7.4: Subtract the estimated mechanical noise from the spatially filtered signal to obtain the enhanced speech signal; Step 7.5: Using the Normalized Least Mean Square (NLMS) algorithm, update the weight coefficients of the adaptive filter in real time based on the enhanced speech signal and the reference noise source vector.

8. A desktop assistant interaction method based on intelligent voice and multi-axis servo motor linkage according to claim 6, characterized in that, The specific steps for spatiotemporal unified synchronization alignment are as follows: Step 8.1: In the audio DMA transfer interrupt service routine, read the hardware timer timestamp, subtract it from the timestamp written to the audio buffer by the upper layer, and obtain the audio rendering delay observation value; Step 8.2: Send a position query command to the servo motor via the communication bus, read the actual motion start timestamp fed back by the servo motor's built-in encoder, subtract the command issuance timestamp from the timestamp, and obtain the servo motor response delay observation value; Step 8.3: Use a Kalman filter to perform state estimation on audio rendering latency and servo response latency, and output the optimal latency estimate; Step 8.4: Calculate the difference between the two as the compensation offset. If the audio path delay is greater than the servo path delay, the servo control command will be delayed in the action ring buffer. If the servo response delay is greater than the audio path delay, a mute data frame will be inserted into the audio ring buffer.

9. A desktop assistant interaction method based on intelligent voice and multi-axis servo motor linkage according to claim 6, characterized in that, The specific steps for the dynamic smooth trajectory planning of the multi-axis servo motor are as follows: Step 9.1: Set the starting angle and target angle of the single-axis servo motor in the current phoneme cycle, as well as the velocity and acceleration boundary conditions at the starting and ending points. Solve the coefficients of the fifth-degree polynomial equation and establish the initial trajectory equation with continuous acceleration. Step 9.2: Obtain the normalized energy value of the current phoneme and construct a nonlinear time scaling function using the hyperbolic tangent function; Step 9.3: Integrate the nonlinear time scaling function over time to generate a resampled virtual time axis; Step 9.4: Map the virtual time axis to the initial trajectory equation to generate a smooth control trajectory with continuous jerk and deep coupling with phoneme energy.

10. A desktop assistant interaction method based on intelligent voice and multi-axis servo motor linkage according to claim 6, characterized in that, The specific steps of speech recognition and semantic navigation are as follows: Step 10.1: Perform frame segmentation and windowing on the enhanced speech signal, calculate the Mel-frequency cepstral coefficients (MFCC) of each frame, and obtain the feature sequence; Step 10.2: Input the feature sequence into the pre-trained acoustic model to calculate the posterior probability of the state, and combine it with the language model to decode using the Viterbi search algorithm, and output the recognized text sequence; Step 10.3: Perform intent recognition and slot filling on the text sequence to extract the action intent and navigation intent, and generate the corresponding servo control target angle and voice navigation path.