Voice robot interaction effect optimization method based on intelligent sensor

By constructing a lip-shape database and introducing asynchronous lip-shape switching, trajectory adaptive optimization, and a flexible dynamic model, the lip-shape driven control of the voice robot was optimized, solving the problems of abrupt lip-shape changes and discontinuous motion, and achieving highly natural and stable voice robot interaction.

CN122090843APending Publication Date: 2026-05-26WUXI INSTITUTE OF TECHNOLOGY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUXI INSTITUTE OF TECHNOLOGY
Filing Date
2026-03-09
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing voice robots suffer from problems such as abrupt lip movements, discontinuous motion, jitter, and dynamic vibration during lip-syncing and execution control. Furthermore, they lack real-time feedback and error correction mechanisms, which affect the interaction effect and stability.

Method used

By constructing a lip shape database, extracting phoneme time series, introducing asynchronous lip shape switching algorithm and lip shape trajectory adaptive optimization algorithm, combining flexible dynamic model to calculate robot mouth driving torque, and using intelligent sensors for real-time status acquisition and error correction, the servo motor drive control parameters are optimized.

Benefits of technology

It improves the continuity of lip-sync transitions and motion stability, enhances the naturalness and stability of voice robot interaction, and ensures that lip-sync changes are consistent with speech rhythm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090843A_ABST
    Figure CN122090843A_ABST
Patent Text Reader

Abstract

The invention discloses a voice robot interaction effect optimization method based on an intelligent sensor, and the method comprises the following steps: S1, obtaining a user voice signal, and converting the user voice signal into text information; s2, extracting a phoneme sequence; s3, constructing a mouth shape database, and executing phoneme matching; s4, proposing an asynchronous mouth shape switching algorithm, and generating a prediction dynamic mouth shape sequence by adopting a mouth shape track adaptive optimization algorithm; s5, a flexible rigidity adjusting item and a damping adjusting item are introduced into the flexible dynamic model, an improved flexible dynamic model is constructed, the driving torque of the mouth of the robot is calculated, and steering engine driving control parameters are obtained; s6, a steering engine control instruction is generated, and a mouth shape action is executed; and S7, updating a steering engine control instruction. According to the method, the synchronization degree between the mouth shape motion of the robot mouth and the voice rhythm is improved, the mouth shape angle deviation and the mouth shape response time delay are reduced, and therefore the naturalness and stability in the voice robot interaction process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice interaction technology, and in particular to a method for optimizing the interaction effect of a voice robot based on intelligent sensors. Background Technology

[0002] With the continuous development of artificial intelligence, intelligent sensor, and service robot technologies, voice interaction has gradually become one of the important ways for intelligent robots to achieve human-computer interaction. To improve the naturalness and expressiveness of voice robots in practical applications, researchers typically drive the robot's mouth structure to perform lip-syncing movements, thereby achieving a synchronized voice and lip-syncing interaction. However, existing voice robots still face many problems in lip-syncing and execution control.

[0003] Existing voice robots typically directly match fixed lip-shape templates to text or phoneme sequences during speech-driven lip-shape generation. This lacks a continuous transition in lip-shape changes, easily leading to abrupt lip-shape changes during phoneme switching. This results in disjointed robot mouth movements and reduces the naturalness of voice interaction. Furthermore, in the process of generating lip-shape trajectories, traditional methods often employ simple interpolation or fixed trajectory generation, making it difficult to dynamically optimize lip-shape movement trajectories based on speech rhythm and lip-shape change characteristics. This leads to tremors or instability in the robot's mouth movements, affecting the accuracy of lip-shape expression.

[0004] Furthermore, in the process of driving and controlling the robot's mouth, existing technologies mostly use rigid body dynamics models or simple position control methods to calculate the driving torque. However, the robot's mouth structure usually has certain flexible characteristics, and it is prone to structural deformation and dynamic vibration during rapid mouth shape changes. Traditional dynamic models are difficult to accurately describe this flexible motion characteristic, resulting in insufficient accuracy in the calculation of servo motor drive control parameters, which in turn leads to problems such as mouth shape angle deviation or mouth shape response delay. At the same time, existing systems lack real-time feedback and error correction mechanisms for the mouth movement state during mouth shape execution, making it difficult to dynamically adjust the control parameters according to the actual execution state, thus affecting the overall interactive effect and stability of the voice robot.

[0005] Therefore, how to provide a method for optimizing the interaction effect of voice robots based on intelligent sensors is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a method for optimizing the interactive effect of a voice robot based on intelligent sensors. This invention utilizes speech denoising and speech feature extraction techniques to obtain text information, combines pinyin conversion, tone elimination, and pinyin error correction to generate a standardized pinyin sequence and extract phoneme time series, constructs a lip-shape database to generate target lip-shape parameters, uses an asynchronous lip-shape switching algorithm to generate an initial dynamic lip-shape sequence, and combines this with a lip-shape trajectory adaptive optimization algorithm to generate a predicted dynamic lip-shape sequence. An improved flexible dynamics model, including flexible stiffness adjustment and damping adjustment terms, is introduced to calculate the robot's mouth driving torque and obtain servo motor drive control parameters. Based on servo motor control commands, the robot's mouth servo motor is driven to perform lip-shape actions, and intelligent sensors are used to collect mouth execution state information. The servo motor drive control parameters are corrected based on lip-shape execution errors. This invention has the advantages of high lip-shape transition continuity, high lip-shape motion stability, and high naturalness of voice robot interaction.

[0007] A method for optimizing the interaction effect of a voice robot based on a smart sensor, according to an embodiment of the present invention, includes the following steps: S1. Acquire user voice signal, perform voice noise reduction and voice feature extraction on user voice signal, and convert user voice signal into text information; S2. Perform pinyin conversion, tone elimination and pinyin error correction on the text information to generate a standardized pinyin sequence. Then, perform syllable parsing on the standardized pinyin sequence to obtain the initial and final combinations, extract the phoneme sequence and generate a phoneme time series. S3. Extract lip shape feature parameters based on lip shape training sample data, construct a lip shape database, and perform phoneme matching in the lip shape database according to the phoneme time series to generate target lip shape parameters; S4. Based on the phoneme time series, a lip shape time series is constructed. An asynchronous lip shape switching algorithm is proposed to calculate the transition trajectory between adjacent phonemes to obtain the initial dynamic lip shape sequence. The lip shape trajectory adaptive optimization algorithm is then used to optimize the lip shape trajectory of the initial dynamic lip shape sequence to generate the predicted dynamic lip shape sequence. S5. Calculate the opening and closing displacement of the robot's mouth based on the mouth feature parameters corresponding to each time position in the predicted dynamic mouth shape sequence, establish the servo motor rotation angle sequence, angular velocity sequence and angular acceleration sequence, introduce flexible stiffness adjustment term and damping adjustment term into the flexible dynamic model, construct an improved flexible dynamic model, calculate the robot's mouth driving torque, and obtain the servo motor driving control parameters. S6. Generate servo control commands based on servo drive control parameters, and drive the robot's mouth servo to perform mouth movements according to the servo control commands, and use intelligent sensors to collect mouth execution status information in real time. S7. If there is a mouth shape angle deviation or mouth shape response delay between the mouth shape execution status information and the target mouth shape parameter, calculate the mouth shape execution error, adjust the servo drive control parameters according to the mouth shape execution error, and update the servo control command.

[0008] Optionally, S1 specifically includes: The system collects user voice signals and generates a continuous voice signal sequence. It then performs frame segmentation on the continuous voice signal sequence according to a preset frame length and frame shift, and performs windowing on the segmented voice frame sequence to obtain a voice frame data sequence. Speech denoising processing is performed on the speech frame data sequence. The background noise spectrum is calculated by noise spectrum estimation, and the speech frame data sequence is subjected to spectrum subtraction operation based on the background noise spectrum to obtain the denoised speech frame sequence. Speech feature extraction is performed on the denoised speech frame sequence, and a fast Fourier transform is performed on the denoised speech frame sequence to obtain speech spectrum data. Speech feature parameters are calculated based on the speech spectrum data to generate a speech feature sequence. Input the speech feature sequence into the speech recognition model to perform speech recognition calculations and output the corresponding text information.

[0009] Optionally, S2 specifically includes: The text information is segmented to obtain a text character sequence, and then the text character sequence is converted into a pinyin string sequence according to the Chinese pinyin mapping rules. Perform tone removal processing on the pinyin string sequence to convert pinyin with tone into pinyin without tone, resulting in a pinyin sequence without tone. Pinyin error correction is performed on the pinyin sequence without tone marks, and the pinyin character combination is corrected according to the rules of legal pinyin combination to obtain a standardized pinyin sequence; The standardized pinyin sequence is divided into syllables. Based on the rules of initial consonant and final vowel combination, the standardized pinyin sequence is parsed to obtain the initial consonant sequence and the final vowel sequence. Phoneme mapping is performed based on the initial consonant sequence and the final vowel sequence to obtain the phoneme sequence. Then, the phoneme sequence is time-arranged according to the speech time order to generate the phoneme time sequence.

[0010] Optionally, S3 specifically includes: Collect lip shape training sample data and generate lip shape sample sequences. Perform image preprocessing on the lip shape sample sequences to obtain lip shape image sequences. Key point detection processing is performed on the lip shape image sequence to extract the coordinates of key points of the lip contour, and the lip shape geometric feature parameters are calculated based on the coordinates of the key points of the lip contour to obtain the lip shape feature parameter sequence; Establish the correspondence between phonemes and lip shape feature parameters based on the sequence of lip shape feature parameters and the corresponding phoneme annotation information, and construct a lip shape database based on the correspondence. Phoneme matching is performed on the lip shape database based on the phoneme time series to obtain the lip shape feature parameter sequence corresponding to the phoneme time series, and the target lip shape parameters are generated based on the lip shape feature parameter sequence.

[0011] Optionally, the step of constructing a lip-sync time series based on phoneme time series and proposing an asynchronous lip-sync switching algorithm to calculate the transition trajectory between adjacent phonemes to obtain an initial dynamic lip-sync sequence is as follows: The phoneme time sequence is obtained by acquiring the time position of the phonemes based on the phoneme time sequence, and the phoneme sequence is sorted by time according to the time position of the phonemes to obtain the phoneme time arrangement sequence. Based on the phoneme time arrangement sequence, the corresponding mouth shape feature parameters are obtained from the mouth shape database, and the mouth shape parameter sequence is generated according to the phoneme time arrangement sequence; Perform lip-shape time alignment processing on the lip-shape parameter sequence to map the lip-shape parameter sequence to the corresponding time position of the phoneme time arrangement sequence, and generate the lip-shape time sequence; Interpolation calculation is performed based on the difference between the corresponding lip shape feature parameters of adjacent phonemes in the lip shape time series. The time interval between the time positions of adjacent phonemes is divided into equally spaced time intervals. At each time position, the corresponding lip shape feature parameters are calculated based on the lip shape feature parameters of the starting phoneme and the ending phoneme, thus obtaining the lip shape transition parameter sequence. An initial dynamic lip shape sequence is generated by combining the lip shape time series and the lip shape transition parameter sequence in chronological order.

[0012] Optionally, the step of using a lip trajectory adaptive optimization algorithm to optimize the initial dynamic lip trajectory of the sequence and generate a predicted dynamic lip sequence specifically involves: The initial dynamic lip shape sequence is processed by trajectory segmentation according to a preset time window to obtain multiple lip shape trajectory segments; For each lip movement trajectory segment, calculate the variation amplitude between lip movement feature parameters at adjacent time positions, and generate trajectory smoothing parameters based on the variation amplitude; A weighted smoothing calculation is performed on the mouth feature parameters in the mouth trajectory segment based on the trajectory smoothing parameters to obtain a smoothed mouth trajectory segment. The smooth lip shape trajectory segments are spliced ​​in chronological order to generate a continuous lip shape trajectory sequence, and a predicted dynamic lip shape sequence is generated based on the continuous lip shape trajectory sequence.

[0013] Optionally, the improved flexible dynamics model is specifically as follows: Based on the mouth shape feature parameters corresponding to each time position in the predicted dynamic mouth shape sequence, and the mouth opening and closing displacement is calculated according to the lip width, lip height and lip opening and closing area in the mouth shape feature parameters, the mouth opening and closing displacement sequence is formed by arranging them in the time order of the predicted dynamic mouth shape sequence. A geometric relationship between mouth opening and closing displacement and servo motor rotation angle is established by combining the structural dimensions of the robot's mouth. The servo motor rotation angle sequence is calculated based on the mouth opening and closing displacement sequence and arranged in the time sequence of the predicted dynamic mouth shape sequence to form the servo motor rotation angle sequence. The angular velocity sequence and angular acceleration sequence are calculated based on the servo angle sequence. The angular velocity is obtained by performing differential operation on the servo angle at adjacent time positions, and the angular acceleration is obtained by performing differential operation on the angular velocity at adjacent time positions. The angular velocity sequence and angular acceleration sequence are formed according to the time sequence of the predicted dynamic lip shape sequence. A flexible dynamic model is established by combining the structural mass, rotation radius and structural damping characteristics of the robot mouth. Flexible stiffness adjustment term and damping adjustment term are introduced into the flexible dynamic model. The servo motor rotation angle sequence, angular velocity sequence and angular acceleration sequence are substituted into the flexible dynamic equation to calculate the driving torque of the robot mouth. The servo motor drive control parameters are obtained by combining the calculated robot mouth driving torque with the servo motor drive characteristics, and the servo motor drive control parameter sequence is generated according to the time sequence of the predicted dynamic mouth shape sequence.

[0014] Optionally, the step of generating servo control commands based on servo drive control parameters specifically includes: Obtain the servo drive control parameter sequence and establish a time arrangement of the servo drive control parameters according to the time order of the predicted dynamic lip shape sequence; By combining the correspondence between servo drive control parameters and servo control signals, the servo drive control parameters are converted into servo control signal amplitudes, and a control signal time arrangement is established based on the predicted dynamic lip sequence time interval. The servo control command sequence is generated by arranging the control signal amplitude and control signal time.

[0015] Optionally, the step of driving the robot's mouth servo motor to perform lip-shape movements according to servo motor control commands, and using intelligent sensors to collect mouth movement status information in real time, specifically includes: According to the servo control command, the robot mouth servo sends a control signal to the servo motor. The robot mouth servo motor drives the robot mouth to perform mouth shape movements according to the corresponding rotation angle of the servo control command, and forms the mouth shape movement trajectory according to the predicted dynamic mouth shape sequence time order. During the mouth movement, intelligent sensors are used to collect information on the robot's mouth movement status, and the mouth movement status data is recorded in the time sequence of the predicted dynamic mouth movement sequence. Based on the mouth movement data, the mouth opening and closing angle and mouth movement time information are extracted to form mouth execution status information.

[0016] Optionally, S7 specifically includes: Acquire mouth execution state information and target mouth shape parameters, and establish a time arrangement of mouth execution state information and target mouth shape parameters according to the time sequence of the predicted dynamic mouth shape sequence; The mouth shape angle deviation is calculated based on the mouth opening and closing angle in the mouth execution status information and the lip height, lip width and lip opening and closing area in the target mouth shape parameters, and a mouth shape angle deviation sequence is formed according to the time sequence of the predicted dynamic mouth shape sequence. The mouth movement time information in the mouth execution state information is combined with the predicted dynamic mouth shape sequence time position to calculate the mouth shape response delay, and a mouth shape response delay sequence is formed according to the predicted dynamic mouth shape sequence time order; The lip-shape execution error sequence is calculated based on the lip-shape angle deviation sequence and the lip-shape response delay sequence, and the lip-shape execution error time arrangement is established according to the time order of the predicted dynamic lip-shape sequence. The servo drive control parameter sequence is corrected and calculated based on the lip-sync error sequence to obtain the updated servo drive control parameter sequence, and the servo control command sequence is regenerated.

[0017] The beneficial effects of this invention are: By constructing a lip shape database and extracting phoneme time series, accurate matching between phonemes and lip shape feature parameters is achieved. An asynchronous lip shape switching algorithm is introduced in the lip shape generation stage. The lip shape transition trajectory between adjacent phonemes is calculated based on the phoneme time position to generate an initial dynamic lip shape sequence. This effectively avoids the lip shape abruptness phenomenon in traditional lip shape template matching methods and improves the continuity of lip shape transition and the consistency of speech expression. Based on the initial dynamic lip shape sequence, an adaptive optimization algorithm for lip shape trajectory is introduced. Through trajectory segmentation, change amplitude calculation and trajectory smoothing parameter generation, weighted smoothing processing is performed on the lip shape trajectory segment to generate a predicted dynamic lip shape sequence. This effectively reduces high-frequency fluctuations in the lip shape trajectory, improves the stability of lip shape movement, and keeps the lip shape changes consistent with the speech rhythm. An improved flexible dynamics model, including flexible stiffness adjustment term and damping adjustment term, is introduced in the dynamic modeling stage. The opening and closing displacement of the robot's mouth is calculated based on the predicted dynamic mouth shape sequence, and the servo motor rotation angle sequence, angular velocity sequence and angular acceleration sequence are established to achieve an accurate description of the motion characteristics of the robot's flexible mouth structure, improve the calculation accuracy of the robot's mouth driving torque, and enhance the stability of mouth shape drive control. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a method for optimizing the interaction effect of a voice robot based on intelligent sensors, as proposed in this invention. Figure 2 This is a schematic diagram of the structure of the improved flexible dynamics model proposed in this invention; Figure 3 This is a schematic diagram of the asynchronous lip-syncing algorithm proposed in this invention; Figure 4 This is a schematic diagram of the adaptive lip trajectory optimization algorithm proposed in this invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0020] refer to Figures 1-4 A method for optimizing the interaction effect of a voice robot based on intelligent sensors includes the following steps: S1. Acquire user voice signal, perform voice noise reduction and voice feature extraction on user voice signal, and convert user voice signal into text information; S2. Perform pinyin conversion, tone elimination and pinyin error correction on the text information to generate a standardized pinyin sequence. Then, perform syllable parsing on the standardized pinyin sequence to obtain the initial and final combinations, extract the phoneme sequence and generate a phoneme time series. S3. Extract lip shape feature parameters based on lip shape training sample data, construct a lip shape database, and perform phoneme matching in the lip shape database according to the phoneme time series to generate target lip shape parameters; S4. Based on the phoneme time series, a lip shape time series is constructed. An asynchronous lip shape switching algorithm is proposed to calculate the transition trajectory between adjacent phonemes to obtain the initial dynamic lip shape sequence. The lip shape trajectory adaptive optimization algorithm is then used to optimize the lip shape trajectory of the initial dynamic lip shape sequence to generate the predicted dynamic lip shape sequence. S5. Calculate the opening and closing displacement of the robot's mouth based on the mouth feature parameters corresponding to each time position in the predicted dynamic mouth shape sequence, establish the servo motor rotation angle sequence, angular velocity sequence and angular acceleration sequence, introduce flexible stiffness adjustment term and damping adjustment term into the flexible dynamic model, construct an improved flexible dynamic model, calculate the robot's mouth driving torque, and obtain the servo motor driving control parameters. S6. Generate servo control commands based on servo drive control parameters, and drive the robot's mouth servo to perform mouth movements according to the servo control commands, and use intelligent sensors to collect mouth execution status information in real time. S7. If there is a mouth shape angle deviation or mouth shape response delay between the mouth execution status information and the target mouth shape parameters, calculate the mouth shape execution error, adjust the servo drive control parameters and update the servo control command according to the mouth shape execution error.

[0021] In this embodiment, S1 specifically refers to: The system collects user voice signals and generates a continuous voice signal sequence. It then performs frame segmentation on the continuous voice signal sequence according to a preset frame length and frame shift, and performs windowing on the segmented voice frame sequence to obtain a voice frame data sequence. Speech denoising processing is performed on the speech frame data sequence. The background noise spectrum is calculated by noise spectrum estimation, and the speech frame data sequence is subjected to spectrum subtraction operation based on the background noise spectrum to obtain the denoised speech frame sequence. Speech feature extraction is performed on the denoised speech frame sequence, and a fast Fourier transform is performed on the denoised speech frame sequence to obtain speech spectrum data. Speech feature parameters are calculated based on the speech spectrum data to generate a speech feature sequence. Input the speech feature sequence into the speech recognition model to perform speech recognition calculations and output the corresponding text information.

[0022] In this embodiment, S2 specifically refers to: The text information is segmented to obtain a text character sequence, and then the text character sequence is converted into a pinyin string sequence according to the Chinese pinyin mapping rules. Perform tone removal processing on the pinyin string sequence to convert pinyin with tone into pinyin without tone, resulting in a pinyin sequence without tone. Pinyin error correction is performed on the pinyin sequence without tone marks, and the pinyin character combination is corrected according to the rules of legal pinyin combination to obtain a standardized pinyin sequence; The standardized pinyin sequence is divided into syllables. Based on the rules of initial consonant and final vowel combination, the standardized pinyin sequence is parsed to obtain the initial consonant sequence and the final vowel sequence. Phoneme mapping is performed based on the initial consonant sequence and the final vowel sequence to obtain the phoneme sequence. Then, the phoneme sequence is time-arranged according to the speech time order to generate the phoneme time sequence.

[0023] In this embodiment, S3 specifically refers to: Collect lip shape training sample data and generate lip shape sample sequence. Perform grayscale normalization and size unification processing on each frame of the lip shape sample sequence. Perform linear normalization on the image pixel values ​​in the range of 0 to 1, and resample according to the uniform resolution to form lip shape image sequence. For each frame of the lip shape image sequence, the pixels of the lip contour edge are extracted, and a fixed number of lip contour key point coordinates are determined on the lip contour edge. The lip contour key point coordinates are composed of two-dimensional plane coordinates, and each key point coordinate is represented as (x_i, y_i), where x_i represents the key point position coordinate in the horizontal direction of the image, and y_i represents the key point position coordinate in the vertical direction of the image. The lip contour closed curve is formed by connecting the key point coordinates in the order of the lip contour boundary. The lip shape geometric parameters are calculated based on the coordinates of the key points of the lip contour, and a sequence of lip shape feature parameters is generated. The lip shape geometric parameters include lip width, lip height, and lip opening area. The lip width is equal to the difference between the x-coordinate of the leftmost key point and the x-coordinate of the rightmost key point. The lip height is equal to the difference between the y-coordinate of the topmost key point and the y-coordinate of the bottommost key point. The lip opening area is calculated according to the area of ​​the polygon enclosed by the coordinates of the key points of the lip contour. The area of ​​the polygon is formed by connecting the coordinates of adjacent key points in sequence and then calculating it according to the area summation formula. The area is equal to the sum of the x-coordinates of each adjacent key point multiplied by the y-coordinate of the next key point minus half the absolute value of the sum of the y-coordinates of each adjacent key point multiplied by the x-coordinate of the next key point. The correspondence between phonemes and lip shape feature parameters is established based on the sequence of lip shape feature parameters and the corresponding phoneme annotation information. The phoneme annotation information is established in a one-to-one correspondence with the lip shape image frame according to the time sequence of speech pronunciation. The phoneme annotation information and the lip shape feature parameter sequence are aligned according to the time sequence and combined item by item to form a phoneme feature pair sequence. The phoneme feature pair sequence consists of a phoneme identifier and a corresponding lip shape feature parameter vector. The lip shape database is formed by classifying and storing the phoneme feature pair sequence according to the phoneme category. Phoneme matching is performed on the lip shape database based on the phoneme time series. The lip shape feature parameter set corresponding to the same phoneme category is searched in the lip shape database according to the phoneme order in the phoneme time series. The corresponding lip shape feature parameter vector is obtained in time order to form a lip shape feature parameter sequence. The lip shape feature parameter of the next frame is directly appended to the lip shape feature parameter of the previous frame in the phoneme time order to form the target lip shape parameter.

[0024] In this embodiment, a lip-sync time series is constructed based on a phoneme time series, and an asynchronous lip-sync switching algorithm is proposed to calculate the transition trajectory between adjacent phonemes, thereby obtaining an initial dynamic lip-sync sequence, specifically as follows: Based on the phoneme time series, the phoneme time position is obtained and a phoneme time arrangement sequence is formed. The phoneme time position represents the start time and end time of speech pronunciation. By sorting the phonemes according to the order of pronunciation time, an increasing time sequence is formed. Based on the phoneme time arrangement sequence, the corresponding mouth shape feature parameters are obtained from the mouth shape database and a mouth shape parameter sequence is generated. The mouth shape feature parameters consist of lip width, lip height and lip opening area. Lip width represents the horizontal distance between the farthest key points on the left and right sides of the lips, lip height represents the vertical distance between the farthest key points on the upper and lower sides of the lips, and lip opening area represents the area of ​​the region enclosed by the key points of the lip contour. The lip shape parameter sequence is subjected to lip shape time alignment processing, which maps the lip shape feature parameter corresponding to each phoneme to the corresponding phoneme time position. A lip shape time sequence is formed by establishing a one-to-one correspondence between the time position and the lip shape feature parameter. Interpolation is performed based on the difference between the lip shape feature parameters of adjacent phonemes in the lip shape time series. The time is divided into multiple time sampling positions according to a fixed time interval between adjacent phoneme time positions. The lip shape feature parameters are calculated at each time sampling position. The lip shape feature parameters are equal to the lip shape feature parameters of the starting phoneme plus the time ratio multiplied by the difference between the lip shape feature parameters of the ending phoneme and the lip shape feature parameters of the starting phoneme. The time ratio is equal to the current time sampling position minus the time position of the starting phoneme and then divided by the time position of the ending phoneme minus the time position of the starting phoneme, forming a lip shape transition parameter sequence. Based on the lip shape time series and the lip shape transition parameter series, the lip shape feature parameters of the later time position are directly appended to the lip shape feature parameters of the previous time position in chronological order to form the initial dynamic lip shape sequence.

[0025] In this embodiment, an adaptive lip trajectory optimization algorithm is used to optimize the lip trajectory of the initial dynamic lip sequence to generate a predicted dynamic lip sequence, specifically as follows: The initial dynamic lip shape sequence is processed by trajectory segmentation according to a preset time window. The preset time window represents a fixed time length. Multiple lip shape trajectory segments are formed by extracting continuous lip shape feature parameters in chronological order. For each mouth shape trajectory segment, calculate the variation amplitude between the mouth shape feature parameters of adjacent time positions. The variation amplitude is equal to the square root of the sum of the difference between the lip width at the next time position and the lip width at the previous time position, plus the square root of the difference between the lip height at the next time position and the lip height at the previous time position, plus the square root of the difference between the lip opening area at the next time position and the lip opening area at the previous time position. The trajectory smoothing parameter is generated based on the magnitude of change. The trajectory smoothing parameter is equal to 1 divided by 1 plus the smoothing adjustment coefficient multiplied by the magnitude of change. The larger the magnitude of change, the smaller the trajectory smoothing parameter, and the smaller the magnitude of change, the larger the trajectory smoothing parameter. The lip feature parameters in the lip trajectory segment are weighted and smoothed according to the trajectory smoothing parameter. The lip feature parameter at the current time position is equal to the trajectory smoothing parameter multiplied by the lip feature parameter at the current time position plus 1 minus the trajectory smoothing parameter multiplied by the average of the lip feature parameters at the previous time position and the lip feature parameters at the next time position, thus obtaining the smoothed lip trajectory segment. For smooth lip shape trajectory segments, the lip shape feature parameters of the first time position of the next lip shape trajectory segment are aligned with the lip shape feature parameters of the last time position of the previous lip shape trajectory segment in chronological order and then directly connected to form a continuous lip shape trajectory sequence. Based on the continuous lip shape trajectory sequence, the lip shape feature parameter sequence is reorganized in chronological order. The lip shape feature parameters at each time position in the continuous lip shape trajectory sequence are arranged sequentially to form a complete time series lip shape parameter set, generating a predicted dynamic lip shape sequence.

[0026] In this embodiment, the asynchronous lip-syncing algorithm establishes a lip-sync time sequence based on the time position corresponding to each phoneme in the phoneme time sequence, and constructs a lip-sync transition trajectory based on the difference between the corresponding lip-sync feature parameters of adjacent phonemes. By dividing the time intervals of adjacent phonemes into equal time intervals, the corresponding lip-sync feature parameters are calculated at each time position based on the lip-sync feature parameters of the starting phoneme and the ending phoneme, so that the lip-sync changes are continuously transitioned between phoneme boundaries, thereby avoiding abrupt lip-sync changes during phoneme switching. The lip-sync trajectory adaptive optimization algorithm segments the lip-sync trajectory according to a preset time window based on the initial dynamic lip-sync sequence, calculates the change amplitude between the lip-sync feature parameters of adjacent time positions for each lip-sync trajectory segment, generates trajectory smoothing parameters based on the change amplitude, performs weighted smoothing calculation on the lip-sync feature parameters in the lip-sync trajectory segment, and splices the trajectory according to the time sequence, so that the continuous lip-sync trajectory sequence reduces high-frequency fluctuations while maintaining the phoneme correspondence, improves the continuity and stability of lip-sync movement, and thus keeps the predicted dynamic lip-sync sequence consistent with the speech rhythm and improves the naturalness of the voice robot interaction.

[0027] In this embodiment, the improved flexible dynamics model is specifically as follows: The mouth opening and closing displacement is calculated based on the mouth shape feature parameters corresponding to each time position in the predicted dynamic mouth shape sequence. The mouth shape feature parameters include lip width, lip height, and lip opening and closing area. The mouth opening and closing displacement is obtained by linearly combining the lip height and lip width according to a preset proportional coefficient and then superimposing the change in lip opening and closing area. The mouth opening and closing displacement is equal to the first proportional coefficient multiplied by the lip height, the second proportional coefficient multiplied by the lip width, and the third proportional coefficient multiplied by the change in lip opening and closing area. The mouth opening and closing displacement sequence is formed by arranging them in the time order of the predicted dynamic mouth shape sequence. Preferably, the first proportional coefficient is 0.55, the second proportional coefficient is 0.15, and the third proportional coefficient is 0.30. A geometric relationship between mouth opening and closing displacement and servo motor rotation angle is established by combining the robot mouth structure dimensions. The robot mouth structure dimensions include the mouth link length and the servo motor rotation radius. The servo motor rotation angle is determined by the proportional relationship between the mouth opening and closing displacement and the mouth link length. The servo motor rotation angle is equal to the mouth opening and closing displacement divided by the mouth link length. The servo motor rotation angle sequence is calculated according to the predicted dynamic mouth shape sequence time sequence. The angular velocity sequence and angular acceleration sequence are calculated based on the servo angle sequence. The ratio of the difference between the servo angles at adjacent time positions to the adjacent time interval forms the angular velocity, which is equal to the servo angle at the next time position minus the servo angle at the previous time position, and then divided by the time interval. The ratio of the difference between the angular velocities at adjacent time positions to the adjacent time interval forms the angular acceleration, which is equal to the angular velocity at the next time position minus the angular velocity at the previous time position, and then divided by the time interval. The angular velocity sequence and angular acceleration sequence are formed by arranging them in the time order of the predicted dynamic lip shape sequence. A flexible dynamic model is established by combining the structural mass, rotation radius, and structural damping characteristics of the robot's mouth. The flexible stiffness adjustment term is composed of the change in mouth opening and closing displacement and the structural stiffness coefficient, and is equal to the structural stiffness coefficient multiplied by the change in mouth opening and closing displacement. The damping adjustment term is composed of the structural damping coefficient and angular velocity, and is equal to the structural damping coefficient multiplied by the angular velocity. Substituting the angular acceleration, angular velocity, flexible stiffness adjustment term, and inertia term formed by structural mass and rotation radius into the flexible dynamic equation, the driving torque of the robot's mouth is calculated. The driving torque of the robot's mouth is equal to the product of the inertia matrix and angular acceleration, plus the damping adjustment term, plus the flexible stiffness adjustment term. The inertia matrix is ​​determined by the structural mass and rotation radius, and is equal to the structural mass multiplied by the square of the rotation radius. Summing the product of the inertia matrix and angular acceleration, the damping adjustment term, and the flexible stiffness adjustment term yields the sequence of driving torques of the robot's mouth. The servo motor drive control parameters are obtained by combining the robot mouth drive torque sequence with the servo motor drive characteristics. The servo motor drive control parameters are determined by the proportional relationship between the robot mouth drive torque and the servo motor rotation efficiency. The servo motor drive control parameters are equal to the robot mouth drive torque divided by the servo motor rotation efficiency. The servo motor drive control parameter sequence is formed by arranging the parameters in the time sequence of the predicted dynamic mouth shape sequence.

[0028] In this embodiment, the improved flexible dynamics model introduces flexible stiffness and damping adjustment terms driven by lip feature parameters based on the flexible dynamics description. The mouth opening and closing displacement is calculated by predicting the lip feature parameters corresponding to each time position in the dynamic lip shape sequence, and this displacement is converted into a servo motor rotation angle sequence, angular velocity sequence, and angular acceleration sequence. In the dynamic equations, the inertia term, flexible stiffness adjustment term, and damping adjustment term are coupled and calculated. The flexible stiffness adjustment term is determined by the change in mouth opening and closing displacement and the structural stiffness coefficient, while the damping adjustment term is determined by the angular velocity... Together with the structural damping coefficient, the elastic deformation of the flexible structure of the mouth during the mouth shape change process is characterized by the flexible stiffness adjustment term, and the energy dissipation during the mouth movement process is characterized by the damping adjustment term. This allows the dynamic equation to simultaneously reflect the inertial motion characteristics, structural flexibility characteristics, and damping characteristics of the robot's mouth, thereby improving the calculation accuracy of the robot's mouth driving torque, ensuring consistency between the robot's mouth servo drive control parameters and the predicted dynamic mouth shape sequence, reducing the angle deviation and response delay during mouth shape execution, and improving the stability and naturalness of the voice robot's mouth movement.

[0029] In this embodiment, servo control commands are generated based on the servo drive control parameters, specifically as follows: The sequence of servo drive control parameters is obtained and arranged in the time order of the predicted dynamic lip sequence to form a time arrangement of servo drive control parameters. The proportional relationship between the servo drive control parameters and the servo rotation angle is converted to obtain the amplitude of the servo control signal. The amplitude of the servo control signal is equal to the servo drive control parameter divided by the servo drive gain coefficient. The servo drive gain coefficient represents the proportional constant between the servo input control signal and the output rotation angle. The control signal time arrangement is established according to the time interval of the predicted dynamic lip sequence to form a control signal amplitude sequence. The servo control command sequence is generated based on the control signal amplitude sequence and the control signal time arrangement. The servo control command consists of the control signal amplitude and the control signal time position. The control signal time position is equal to the cumulative time of adjacent time positions in the predicted dynamic lip sequence. The control signal amplitude sequence and the control signal time arrangement are sequentially connected in ascending order of time to form the servo control command sequence.

[0030] In this embodiment, the robot's mouth servo motor is driven to perform lip-gesture actions according to servo control commands, and intelligent sensors are used to collect mouth execution status information in real time, specifically: According to the servo control command sequence, control signals are sent to the servo motors in the robot's mouth. The servo motors in the robot's mouth adjust the rotation angle according to the amplitude of the control signal. The rotation angle of the servo motor is equal to the amplitude of the control signal multiplied by the servo motor rotation angle proportional coefficient. The servo motor rotation angle proportional coefficient represents the proportional relationship between the amplitude of the servo motor input signal and the output rotation angle. The servo motor rotation angle changes in the time sequence of the predicted dynamic mouth shape sequence to form the robot's mouth shape movement trajectory. During the robot's mouth movement, intelligent sensors collect information on the robot's mouth movement status. The intelligent sensors collect data on the position change of the mouth link and calculate the mouth opening angle based on the link length. The mouth opening angle is equal to the link displacement divided by the link length. The link displacement represents the displacement change of the mouth link between adjacent time positions. Mouth movement state data is generated based on the mouth opening and closing angle and the acquisition time. The mouth movement time information is equal to the current acquisition time minus the initial acquisition time. The mouth opening and closing angle is arranged in the time sequence of the predicted dynamic mouth shape sequence to form mouth execution state information.

[0031] In this embodiment, S7 specifically refers to: The mouth execution state information and target mouth shape parameters are acquired, and a time correspondence is established according to the time sequence of the predicted dynamic mouth shape sequence to form a time arrangement of mouth execution state information and target mouth shape parameters. The target mouth shape parameters consist of lip height, lip width, and lip opening and closing area. The target mouth opening and closing angle is calculated by combining lip height and lip width. The target mouth opening and closing angle is equal to the first proportional coefficient multiplied by the lip height, the second proportional coefficient multiplied by the lip width, and the third proportional coefficient multiplied by the change in lip opening and closing area. The mouth shape angle deviation is formed by the difference between the mouth opening and closing angle in the mouth execution state information and the target mouth shape opening and closing angle. The mouth shape angle deviation is equal to the mouth opening and closing angle minus the target mouth shape opening and closing angle, and is arranged according to the time sequence of the predicted dynamic mouth shape sequence to form a mouth shape angle deviation sequence. By combining the mouth movement time information in the mouth execution state information with the predicted dynamic mouth shape sequence time position, a time correspondence is established to form a mouth shape response delay sequence. The mouth shape response delay is equal to the mouth movement time information minus the predicted dynamic mouth shape sequence time position. The mouth shape response delay sequence is formed by arranging the time difference between adjacent time positions. The lip shape execution error sequence is formed by combining the lip shape angle deviation sequence and the lip shape response delay sequence. The lip shape execution error is equal to the sum of the absolute value of the angle deviation and the product of the response delay. The lip shape execution error is equal to the absolute value of the angle deviation plus the response delay multiplied by the time weighting coefficient and arranged in the time order of the predicted dynamic lip shape sequence. The servo drive control parameter sequence is corrected and calculated based on the lip-sync error sequence. The corrected servo drive control parameter is equal to the original servo drive control parameter minus the error adjustment coefficient multiplied by the lip-sync error. The servo drive control parameter is dynamically compensated by the error adjustment coefficient to form an updated servo drive control parameter sequence. Based on the updated servo drive control parameter sequence, the servo drive control parameter time arrangement is re-established and a new servo control command sequence is generated. The robot's mouth servo is then driven to perform the corrected mouth shape movements through the updated servo control command sequence.

[0032] Example 1: To verify the feasibility of this invention in practice, it was applied to an intelligent service robot interaction system. The robot is equipped with an array of voice acquisition sensors, a mouth-actuating servo structure, and angle and time information acquisition sensors. The robot's mouth structure weighs 0.18 kg, has a mouth rotation radius of 0.045 m, a servo's rated rotation angle range of 0°-120°, and a control refresh cycle of 20 ms. In practical applications, voice robots need to perform numerous voice interaction tasks in scenarios such as voice broadcasting, customer service consultation, and public service guidance. Traditional voice robots typically use a fixed lip-sync method for voice broadcasting, lacking an accurate correspondence between lip movements and voice content. This leads to a lack of synchronization between mouth movements and voice rhythm, easily producing noticeable lip-sync delays and opening / closing angle errors, thereby reducing the naturalness and realism of the voice robot's interaction.

[0033] To verify the feasibility of this invention in practice, it was applied to an intelligent voice service robot system. During the robot's voice broadcasting process, user voice signals were collected and subjected to voice noise reduction and voice feature extraction, converting the user voice signals into text information. The text information underwent pinyin conversion, tone elimination, and pinyin error correction to generate a standardized pinyin sequence, which was further parsed to obtain a phoneme sequence and a phoneme time series. Based on lip-shape training sample data, lip-shape feature parameters were extracted and a lip-shape database was constructed. Phoneme matching was performed in the lip-shape database according to the phoneme time series to obtain target lip-shape parameters. A lip-shape time series was constructed based on the phoneme time series. An asynchronous lip-shape switching algorithm was used to calculate the transition trajectory between adjacent phonemes to generate an initial dynamic lip-shape sequence, and adaptive optimization of the lip-shape trajectory was performed. The algorithm optimizes the lip trajectory to obtain a predicted dynamic lip sequence. Based on the lip feature parameters corresponding to each time position in the predicted dynamic lip sequence, the robot's mouth opening and closing displacement is calculated, and servo motor rotation angle, angular velocity, and angular acceleration sequences are established. An improved flexible dynamic model is constructed by introducing flexible stiffness and damping adjustment terms into the flexible dynamic model. The servo motor drive control parameters are obtained by calculating the robot's mouth driving torque through the flexible dynamic equations. Based on the servo motor drive control parameters, servo motor control commands are generated to drive the robot's mouth servo motors to perform lip movements. During the robot's mouth movement, intelligent sensors collect mouth execution state information, and the lip execution error is calculated based on the deviation between the mouth execution state information and the target lip parameters. The servo motor drive control parameters are then updated and adjusted accordingly.

[0034] To compare and verify the effectiveness of the method of the present invention in the lip-sync control of a voice robot, the traditional lip-sync control method and the method of the present invention were selected and tested on the same robot platform. The experimental environment was an indoor service robot interaction scenario. The number of test sentences was 120, with an average length of 6-10 Chinese characters per sentence and a speech playback duration of about 3-5 seconds. The mouth opening and closing angle error, lip-sync response delay, and lip-sync trajectory smoothness were recorded by intelligent sensors. The following experimental results were obtained.

[0035] Table 1. Comparison of Lip Execution Accuracy of Voice Robots

[0036] As shown in Table 1, traditional lip-sync control methods maintain an average mouth opening / closing angle error between 6.5° and 7.4° during robot voice playback, with a maximum angle error exceeding 13°. Simultaneously, the lip-sync response delay is generally higher than 170ms, easily leading to significant lip-sync delay during continuous voice playback. After adopting the method of this invention, the average mouth opening / closing angle error is reduced to 2.1°-2.6°, the maximum angle error is reduced to within 5.6°, and the lip-sync response delay is reduced to between 64ms and 75ms, significantly reducing the fluctuation amplitude of the lip trajectory. Experimental results demonstrate that constructing a lip-sync time series using phoneme time series and combining it with an asynchronous lip-sync switching algorithm and a lip trajectory adaptive optimization algorithm can significantly improve the synchronization between the robot's mouth movements and the speech rhythm. Furthermore, improving the flexible dynamics model to calculate servo drive control parameters can effectively reduce the mouth opening / closing angle error, improving the realism and stability of the voice robot's interactive effect.

[0037] To further verify the stability of the method of the present invention in continuous voice interaction scenarios, a long-term voice broadcast test was conducted on the same robot platform. The test lasted for 30 minutes, and a total of 420 sentences were broadcast. The mouth execution status information was continuously recorded by intelligent sensors, and the stability index of mouth shape control was statistically analyzed. The experimental results are as follows.

[0038] Table 2. Comparison of Lip-Shape Control Stability in Continuous Voice Interaction Scenarios

[0039] As shown in Table 2, during continuous voice interaction, the traditional method experiences a gradual increase in average angle error and response delay with increasing runtime, resulting in a drop in lip-sync accuracy to approximately 75%. However, the method of this invention maintains an average angle error between 2.4° and 2.7°, a lip-sync response delay of around 70ms, and a lip-sync accuracy consistently above 92%, with a significantly higher trajectory smoothness index than the traditional method. Experimental results demonstrate that the method of this invention maintains stable lip-sync control performance even in long-term voice interaction environments. By using phoneme time-series to drive lip-sync time-series generation and improving the flexible dynamics model to calculate servo drive control parameters, the dynamic response capability and control stability of the robot's mouth movements can be effectively improved, thereby significantly enhancing the naturalness and interactive experience of the voice robot in human-computer interaction scenarios.

[0040] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for optimizing voice robot interaction effects based on intelligent sensors, characterized in that, The method comprises the following steps: S1, acquiring a user voice signal, performing voice noise reduction processing and voice feature extraction on the user voice signal, and converting the user voice signal into text information; S2, performing pinyin conversion, tone elimination and pinyin error correction processing on the text information to generate a standardized pinyin sequence, and performing syllable analysis on the standardized pinyin sequence to obtain an initial consonant and vowel combination, extract a phoneme sequence and generate a phoneme time sequence; S3, extracting mouth shape feature parameters based on mouth shape training sample data, constructing a mouth shape database, and performing phoneme matching on the phoneme time sequence in the mouth shape database to generate target mouth shape parameters; S4, constructing a mouth shape time sequence based on the phoneme time sequence, proposing an asynchronous mouth shape switching algorithm to calculate the transition trajectory between adjacent phonemes, obtaining an initial dynamic mouth shape sequence, and using a mouth shape trajectory adaptive optimization algorithm to optimize the initial dynamic mouth shape sequence to generate a predicted dynamic mouth shape sequence; S5, calculating the opening and closing displacement of the robot mouth based on the mouth shape feature parameters corresponding to each time position in the predicted dynamic mouth shape sequence, establishing a rudder angle sequence, an angular velocity sequence and an angular acceleration sequence, introducing a flexible stiffness adjustment term and a damping adjustment term into a flexible dynamics model, constructing an improved flexible dynamics model, calculating the driving torque of the robot mouth, and obtaining the rudder driving control parameters; S6, generating a rudder control instruction according to the rudder driving control parameters, and driving the robot mouth rudder to perform mouth shape actions according to the rudder control instruction, and collecting the mouth execution state information in real time by using an intelligent sensor; S7, if there is a mouth shape angle deviation or a mouth shape response time delay between the mouth execution state information and the target mouth shape parameters, calculating a mouth shape execution error, adjusting the rudder driving control parameters according to the mouth shape execution error and updating the rudder control instruction. 2.The method of claim 1, wherein, The S1 is specifically: Collecting a user voice signal and generating a continuous voice signal sequence, performing frame processing on the continuous voice signal sequence according to a preset frame length and frame shift, and performing windowing processing on the frame-processed voice frame sequence to obtain a voice frame data sequence; Performing voice noise reduction processing on the voice frame data sequence, calculating the background noise spectrum through noise spectrum estimation, and performing spectral subtraction operation on the voice frame data sequence according to the background noise spectrum to obtain a noise reduction voice frame sequence; Performing voice feature extraction processing on the noise reduction voice frame sequence, performing fast Fourier transform on the noise reduction voice frame sequence to obtain voice spectrum data, and calculating voice feature parameters according to the voice spectrum data to generate a voice feature sequence; Inputting the voice feature sequence into a speech recognition model to perform speech recognition calculation, and outputting corresponding text information. 3.The method of claim 1, wherein the method further comprises: determining a voice robot interaction effect based on the voice robot interaction effect model; and providing the determined voice robot interaction effect to the user. The S2 is specifically: Performing character segmentation processing on the text information to obtain a text character sequence, and performing pinyin conversion processing on the text character sequence according to the Chinese pinyin mapping rule to obtain a pinyin string sequence; Performing tone elimination processing on the pinyin string sequence to convert the pinyin with tone into the pinyin without tone, and obtaining a toneless pinyin sequence; Performing pinyin error correction processing on the toneless pinyin sequence, correcting the pinyin character combination through the pinyin legal combination rule to obtain a standardized pinyin sequence; The syllable division processing is performed on the standardized pinyin sequence, syllable analysis is performed on the standardized pinyin sequence according to a combination rule of initials and finals, and an initial sequence and a final sequence are obtained; The phoneme mapping processing is performed according to the initial sequence and the final sequence, a phoneme sequence is obtained, and the time arrangement is performed on the phoneme sequence according to a voice time sequence, and a phoneme time sequence is generated.

4. The method of claim 1, wherein the method further comprises: The S3 is specifically: The mouth shape training sample data is collected and a mouth shape sample sequence is generated, image preprocessing is performed on the mouth shape sample sequence, and a mouth shape image sequence is obtained; The key point detection processing is performed on the mouth shape image sequence, the lip contour key point coordinates are extracted, the mouth shape geometric feature parameters are calculated according to the lip contour key point coordinates, and a mouth shape feature parameter sequence is obtained; The corresponding relationship between the phonemes and the mouth shape feature parameters is established according to the mouth shape feature parameter sequence and the corresponding phoneme annotation information, and the mouth shape database is constructed according to the corresponding relationship; The phoneme matching processing is performed on the mouth shape database according to the phoneme time sequence, the mouth shape feature parameter sequence corresponding to the phoneme time sequence is obtained, and the target mouth shape parameter is generated according to the mouth shape feature parameter sequence.

5. The method of claim 1, wherein the method further comprises: The mouth shape time sequence is constructed based on the phoneme time sequence, an asynchronous mouth shape switching algorithm is used to calculate the transition trajectory between adjacent phonemes, and an initial dynamic mouth shape sequence is obtained, specifically: The time position of the phoneme is obtained based on the phoneme time sequence, and the time sorting is performed on the phoneme sequence according to the time position of the phoneme, and a phoneme time arrangement sequence is obtained; The corresponding mouth shape feature parameters are obtained in the mouth shape database based on the phoneme time arrangement sequence, and the mouth shape parameter sequence is generated according to the phoneme time arrangement sequence; The mouth shape time alignment processing is performed on the mouth shape parameter sequence, the mouth shape parameter sequence is mapped to the corresponding time position of the phoneme time arrangement sequence, and the mouth shape time sequence is generated; The interpolation calculation is performed according to the difference between the adjacent phoneme corresponding mouth shape feature parameters in the mouth shape time sequence, the time interval between the adjacent phoneme time positions is equally divided, and the corresponding mouth shape feature parameters are calculated according to the initial phoneme mouth shape feature parameters and the terminal phoneme mouth shape feature parameters at each time position, and a mouth shape transition parameter sequence is obtained; The combination processing is performed according to the time sequence based on the mouth shape time sequence and the mouth shape transition parameter sequence, and an initial dynamic mouth shape sequence is generated.

6. The method of claim 1, wherein the method further comprises: The mouth shape trajectory optimization is performed on the initial dynamic mouth shape sequence by using the mouth shape trajectory adaptive optimization algorithm, and a predicted dynamic mouth shape sequence is generated, specifically: The trajectory segmentation processing is performed on the initial dynamic mouth shape sequence according to a preset time window, and a plurality of mouth shape trajectory segments are obtained; The change amplitude between the adjacent time position mouth shape feature parameters of each mouth shape trajectory segment is calculated, and the trajectory smoothing parameter is generated according to the change amplitude; The weighted smoothing calculation is performed on the mouth shape feature parameters in the mouth shape trajectory segment according to the trajectory smoothing parameter, and a smoothed mouth shape trajectory segment is obtained; The trajectory splicing processing is performed on the smoothed mouth shape trajectory segment according to the time sequence, a continuous mouth shape trajectory sequence is generated, and a predicted dynamic mouth shape sequence is generated according to the continuous mouth shape trajectory sequence.

7. The method of claim 1, wherein the method further comprises: The improved flexible dynamics model is specifically: The mouth opening and closing displacement sequence is formed according to the time sequence of the predicted dynamic mouth shape sequence, based on the mouth shape feature parameters corresponding to each time position in the predicted dynamic mouth shape sequence, and according to the mouth width, mouth height and mouth opening and closing area in the mouth shape feature parameters. The geometric relationship between the mouth opening and closing displacement and the steering angle is established in combination with the size of the mouth structure of the robot, the steering angle sequence is calculated according to the mouth opening and closing displacement sequence, and the steering angle sequence is formed according to the time sequence of the predicted dynamic mouth shape sequence. The angular velocity sequence and the angular acceleration sequence are calculated according to the steering angle sequence, the angular velocity is obtained by performing difference operation on the steering angles of adjacent time positions, the angular acceleration is obtained by performing difference operation on the angular velocities of adjacent time positions, and the angular velocity sequence and the angular acceleration sequence are formed according to the time sequence of the predicted dynamic mouth shape sequence. The flexible dynamics model is established in combination with the mass, rotating radius and structural damping characteristics of the mouth structure of the robot, the flexible stiffness adjustment term and the damping adjustment term are introduced into the flexible dynamics model, the steering angle sequence, the angular velocity sequence and the angular acceleration sequence are substituted into the flexible dynamics equation to calculate the driving torque of the mouth of the robot. The steering driving control parameters are obtained according to the calculated driving torque of the mouth of the robot in combination with the steering driving characteristics, and the steering driving control parameter sequence is generated according to the time sequence of the predicted dynamic mouth shape sequence. 8.The method of claim 1, wherein the method further comprises: receiving a voice input from the user; and determining whether the voice input is a voice command or a voice query based on the received voice input. The steering control instruction is generated according to the steering driving control parameters, specifically as follows: The steering driving control parameter sequence is obtained, and the steering driving control parameter time arrangement is established according to the time sequence of the predicted dynamic mouth shape sequence; The steering driving control parameters are converted into the steering control signal amplitude in combination with the corresponding relationship between the steering driving control parameters and the steering control signal, and the control signal time arrangement is established according to the time interval of the predicted dynamic mouth shape sequence; The steering control instruction sequence is generated according to the control signal amplitude and the control signal time arrangement.

9. The method for optimizing the interactive effect of a voice robot based on intelligent sensors according to claim 1, characterized in that, The mouth shape action is performed by the robot mouth steering according to the steering control instruction, and the mouth execution state information is collected in real time by using the intelligent sensor, specifically as follows: The control signal is sent to the robot mouth steering according to the steering control instruction, the robot mouth steering drives the robot mouth to perform the mouth shape action according to the corresponding steering angle of the steering control instruction, and the mouth shape movement trajectory is formed according to the time sequence of the predicted dynamic mouth shape sequence; The mouth movement state information is collected by using the intelligent sensor during the mouth shape movement, and the mouth movement state data is recorded according to the time sequence of the predicted dynamic mouth shape sequence; The mouth opening and closing angle and the mouth movement time information are extracted from the mouth movement state data to form the mouth execution state information.

10. The method of claim 1, wherein the method further comprises: The S7 is specifically as follows: The mouth execution state information and the target mouth shape parameters are obtained, and the mouth execution state information time arrangement and the target mouth shape parameter time arrangement are established according to the time sequence of the predicted dynamic mouth shape sequence; The mouth shape angle deviation is calculated according to the mouth opening and closing angle in the mouth execution state information and the mouth width, mouth height and mouth opening and closing area in the target mouth shape parameters, and the mouth shape angle deviation sequence is formed according to the time sequence of the predicted dynamic mouth shape sequence. The mouth movement time information in the mouth execution state information is combined with the predicted dynamic mouth shape sequence time position to calculate the mouth shape response delay, and a mouth shape response delay sequence is formed according to the predicted dynamic mouth shape sequence time order; The lip-shape execution error sequence is calculated based on the lip-shape angle deviation sequence and the lip-shape response delay sequence, and the lip-shape execution error time arrangement is established according to the time order of the predicted dynamic lip-shape sequence. The servo drive control parameter sequence is corrected and calculated based on the lip-sync error sequence to obtain the updated servo drive control parameter sequence, and the servo control command sequence is regenerated.