A method for synchronizing speech and motion

By segmenting and feature extraction of speech signals, combining the timing alignment and reorganization of action elements, the space-time continuity of action sequences is optimized, and the problem of speech and action synchronization in the animation intelligent synthesis system is solved, and action generation with high naturalness and fluency is achieved.

CN119672185BActive Publication Date: 2025-05-16XIAMEN YOYA NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510175788.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

In the animation intelligent synthesis system, how to achieve accurate synchronization of speech and actions, especially when the dialogue rhythm changes, maintain the consistency and fluency of actions, and take into account the personalized characteristics of the characters.

Method used

By obtaining the duration characteristics of the speech signal, performing segmentation processing and extracting short-time features, combining preset action elements, calculating the timing alignment relationship between action elements and continuous speech segments, and generating a preliminary aligned action sequence. For speech pause segments, the action element sequence is reorganized and the spatiotemporal continuity of the action sequence is optimized through secondary alignment and fluency evaluation.

Benefits of technology

It realizes precise synchronization of voice and actions, improves the natural fluency of action generation, and is suitable for fields such as virtual anchors and intelligent robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672185B_ABST
    Figure CN119672185B_ABST
Patent Text Reader

Abstract

The present application provides a method for synchronizing speech and action, comprising: acquiring a speech signal, extracting speech duration features of the speech signal, segmenting the speech signal in combination with a preset speech segmentation rule, dividing the speech into speech continuous segments and speech pause segments, and extracting short-term features corresponding to each segment to obtain speech timing distribution features; if a fluency score is lower than a preset score threshold, adjusting the timing and spatial position of an action element, and obtaining and storing an optimized action sequence by fine-tuning the action in time and space; inputting the optimized action sequence into an action controller, generating an action execution instruction in combination with the speech timing distribution features, and completing the synchronous output of the action and speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a method for synchronizing voice and action. Background Art

[0002] In the animation intelligent synthesis system, the synchronization of dialogue rhythm and action sequence is a complex technical problem. The coupling relationship between speech duration and body movement timing directly affects the naturalness and realism of the character's performance. When the dialogue rhythm pauses, the action controller needs to adjust the action element sequence in real time to maintain the continuity and fluency of the character's action. However, in the process of action reorganization, how to accurately grasp the time nodes and avoid stiff, unnatural or overly smooth movements is a key problem that needs to be solved urgently. At the same time, action reorganization also needs to consider the impact of factors such as character emotions and personality on action performance. Characters with different personalities often have different body language expressions when facing the same dialogue rhythm. How to take into account the personalized characteristics of the characters in action reorganization, make their actions more diverse and expressive, and automatically generate interactive actions that conform to the context according to the dialogue content and role relationships, and achieve smooth transition and natural connection of actions when the dialogue rhythm changes is a very challenging technical problem. Summary of the invention

[0003] The present invention provides a method for synchronizing speech and motion, which is used for intelligent animation synthesis and mainly includes:

[0004] Acquire a speech signal, extract the speech duration feature of the speech signal, perform segmentation processing on the speech signal in combination with a preset speech segmentation rule, divide the speech into speech continuous segments and speech pause segments, and extract the short-term features corresponding to each segment to obtain the speech timing distribution feature; detect the starting point and the ending point of the speech pause segment according to the speech timing distribution feature, determine the starting time, the ending time and the duration of the pause segment, combine the speech timing distribution feature with the preset action element, calculate the timing alignment relationship between the action element and the speech continuous segment, and generate a preliminary aligned action sequence; detect whether the preliminary aligned action sequence has a speech pause segment, and if so, determine whether the action element sequence needs to be reorganized according to the pause duration and the duration of the action element, and when the pause duration is greater than the threshold of the target action element duration, reorganize the action element sequence to generate a reorganized action sequence; reorganize the reorganized action sequence The action sequence is aligned with the timing distribution of the continuous speech segment twice. When the difference between the start time and the end time of the action element and the start time and the end time of the corresponding speech segment is less than the preset synchronization threshold, it is determined that the coupling relationship between the reorganized action sequence and the speech meets the synchronization requirement; the transition time and spatial continuity between the action elements of the reorganized action sequence are calculated, and the action fluency is evaluated by calculating the difference between adjacent action elements in time and space, and obtaining the speed, acceleration and angle changes during the action execution process to generate a fluency score result; if the fluency score is lower than the preset score threshold, the timing and spatial position of the action element are adjusted, and the optimized action sequence is obtained and stored by fine-tuning the action in time and space; the optimized action sequence is input into the action controller, and the action execution instruction is generated in combination with the timing distribution characteristics of the speech, so as to complete the synchronous output of the action and the speech.

[0005] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:

[0006] The present invention discloses an intelligent animation synthesis method that integrates production elements and script elements. The method solves the problem of precise synchronization between action and voice for a voice-driven action generation scenario. First, the present invention performs segmentation processing on the input voice signal and extracts the temporal distribution characteristics of the voice. Then, the voice features are aligned with preset action elements to generate a preliminary action sequence. For the pause segments in the voice, the present invention optimizes the action sequence by reorganizing the action elements. Through secondary alignment and fluency evaluation, the spatiotemporal continuity of the action sequence is further optimized. Finally, the present invention combines the optimized action sequence with the voice features to generate synchronized action execution instructions. This method achieves precise synchronization between voice and action, improves the natural fluency of action generation, and can be widely used in the fields of virtual anchors, intelligent robots, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1The present invention is a flow chart of a method for synchronizing speech and motion.

[0008] Figure 2 A schematic diagram of a method for synchronizing speech and motion according to the present invention.

[0009] Figure 3 It is another schematic diagram of a method for synchronizing speech and motion according to the present invention. DETAILED DESCRIPTION

[0010] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0011] like Figure 1-3 In this embodiment, a method for synchronizing voice and action may specifically include:

[0012] S101, acquiring a speech signal, extracting speech duration features of the speech signal, segmenting the speech signal in combination with a preset speech segmentation rule, dividing the speech into speech continuous segments and speech pause segments, and extracting short-term features corresponding to each segment to obtain speech temporal distribution features.

[0013] The speech signal is smoothed and denoised by using a wavelet transform method, the signal amplitude is normalized by setting high and low double thresholds, and a Butterworth low-pass filter is used to eliminate high-frequency noise interference to obtain a speech signal after interference is eliminated; a short-time energy function and a short-time zero-crossing rate function are calculated according to the speech signal after interference is eliminated to obtain a speech signal strength sequence, if the signal strength is greater than a preset volume threshold, the speech starting point is marked, and the speech duration sequence is obtained according to the interval between adjacent starting points; the speech signal after interference is eliminated is segmented according to the speech duration sequence, the speech frame is divided by a Hamming window function, the speech frame boundary is determined by calculating the energy ratio between frames, the fundamental pitch period and the resonance peak frequency are extracted from the speech frame to obtain a speech basic feature sequence; continuity analysis is performed according to the speech basic feature sequence, if the pitch period change rate of adjacent speech frames is less than a preset threshold, it is determined to be a continuous speech segment, the maximum entropy model is used to optimize the speech segment boundary, and the energy feature vector and the time-varying feature vector are extracted from the continuous speech segment and the pause segment to obtain the speech time series distribution feature.

[0014] Exemplarily, after sampling and acquiring the speech signal, the wavelet transform method is used to smooth and reduce the noise of the speech signal, and the signal amplitude is normalized by setting high and low double thresholds. The normalized signal is subjected to a Butterworth low-pass filter to eliminate high-frequency noise interference to obtain the first speech signal. The first speech signal is subjected to time feature extraction, and the speech signal strength sequence is obtained by calculating the short-time energy function and the short-time zero-crossing rate function. If the signal strength is greater than the preset volume threshold, it is marked as the speech starting point, and the time interval between adjacent starting points is calculated to obtain the speech duration sequence. The first speech signal is segmented according to the speech duration sequence, and the Hamming window function is used to divide the speech frame. The speech frame boundary is determined by calculating the energy ratio between frames, and the pitch period and the formant frequency are extracted for each speech frame to obtain the speech basic feature sequence. The speech continuity analysis is performed on the speech basic feature sequence. If the pitch period change rate of adjacent speech frames is less than the preset threshold, it is determined as a continuous speech segment, otherwise it is determined as a pause segment, and the maximum entropy model is used to optimize the positioning of the speech segment boundary. The time series distribution features are extracted for continuous speech segments and pause segments respectively. The energy feature vector of the speech segment is obtained by calculating the short-time energy mean, variance and skewness in the speech segment. The time-varying feature vector of the speech segment is obtained by calculating the pitch period change rate in the speech segment. The energy feature vector and the time-varying feature vector of the speech segment are combined by the sliding time window method, and the speech time series distribution feature matrix is ​​obtained by calculating the Euclidean distance between the feature vectors. The smooth noise reduction processing of the speech signal adopts the wavelet transform method. For the speech signal with a sampling frequency of 16000 Hz, the Debesey 4 wavelet basis function is selected for 5-layer decomposition to obtain the wavelet coefficients of different frequency bands. The wavelet coefficients are processed by the soft threshold method, and the threshold is set to 3 times the standard deviation of the wavelet coefficients of this layer to achieve effective noise suppression. For the normalization processing of the signal amplitude, the high threshold is set to 0.8 and the low threshold is set to 0.2, and the signal amplitude is linearly mapped by iteration. In the stage of speech duration feature extraction, the speech signal is framed, the frame length is set to 320 sampling points, and the frame shift is set to 160 sampling points. When calculating the short-time energy function, a rectangular window function is used to window each frame of the signal to obtain a speech signal strength sequence. For the determination of the speech starting point, the volume threshold is set to 0.3. When the signal strength of 5 consecutive frames is greater than the threshold, the starting point of the first frame is marked as the speech starting point. The boundary determination of the speech frame is processed by the Hamming window function, and the window length is set to 320 sampling points. The energy ratio between adjacent frames is calculated. If the energy ratio is greater than 2 or less than 0.5, the frame is marked as a boundary frame. For the extraction of the fundamental period, the autocorrelation method is used for calculation, and the fundamental period is searched in the frequency range of 50 Hz to 400 Hz. The linear prediction analysis method is used to extract the resonance peak frequency, and the prediction order is set to 12 to extract the first 3 resonance peak frequencies.In the speech continuity analysis, the threshold of the pitch period change rate is set to 0.15. When the pitch period change rate of adjacent frames is less than the threshold, it is determined to be a continuous speech segment. When the maximum entropy model optimizes the speech segment boundary, the feature selection includes parameters such as pitch period, short-time energy, and zero-crossing rate, and the number of model iterations is set to 100 times. In the extraction of speech segment time series features, a 64-ms sliding time window is used, and the window overlap rate is 50%. The mean, variance, and skewness of the short-time energy in the window are calculated to obtain a 3D energy feature vector. The difference method is used to calculate the pitch period change rate to obtain a 1D time-varying feature vector. After combining the feature vectors, the speech time series distribution feature matrix is ​​obtained by Euclidean distance calculation. The number of rows of the matrix corresponds to the number of time windows, and the number of columns is 4, which represent the energy mean, variance, skewness, and pitch period change rate, respectively. For the determination of continuous speech segments and pause segments, the zero crossing rate feature is introduced for auxiliary judgment. When the zero crossing rate is greater than 1500 times per second and the energy is low, it is determined to be a silent segment. The localization accuracy of speech segment boundaries can be further optimized by calculating the first-order derivative of the energy envelope within the speech segment. The minimum duration of a speech segment is set to 200 milliseconds, and speech segments shorter than this threshold will be merged into adjacent segments.

[0015] S102. According to the speech timing distribution characteristics, the starting point and the ending point of the speech pause segment are detected, the starting time, the ending time and the duration of the pause segment are determined, the speech timing distribution characteristics and the preset action elements are combined, the timing alignment relationship between the action elements and the speech continuous segment is calculated, and a preliminary aligned action sequence is generated.

[0016] According to the time series distribution characteristics of the speech signal, the short-time energy difference value and the zero-crossing rate difference value of adjacent speech frames are obtained. If the short-time energy difference value is greater than a first preset threshold value and the zero-crossing rate difference value is greater than a second preset threshold value, the speech frame is marked as the start frame of the pause segment; the recursive bisection method is used to locate the boundary of the pause segment, and the characteristic distance sequence is obtained by calculating the Mel-frequency cepstral coefficient distance between adjacent speech frames, and the second pause segment position sequence is determined according to the characteristic distance sequence; the speech continuous segment is divided according to the second pause segment position sequence, and the speech continuous segment feature vector is obtained by calculating the duration, average energy and fundamental frequency period parameters of the speech continuous segment; the dynamic time warping algorithm is used to calculate the matching degree between the speech continuous segment feature vector and the preset action element feature vector, and the corresponding action element is selected according to the matching degree, and the execution timing of the action element is nonlinearly scaled by the cubic spline interpolation algorithm to obtain a timing-aligned action sequence.

[0017] Exemplarily, the short-time energy difference value and the zero-crossing rate difference value between adjacent speech frames are extracted according to the speech time series distribution characteristics. If the short-time energy difference value of the speech frame is greater than the first preset threshold value and the zero-crossing rate difference value is greater than the second preset threshold value, the speech frame is marked as the start frame of the pause segment, and the first pause segment position sequence is obtained. The boundary of each pause segment in the first pause segment position sequence is accurately located, and the characteristic distance sequence is obtained by calculating the Mel frequency cepstral coefficient distance between adjacent speech frames. The characteristic distance sequence is segmented by recursive dichotomy to obtain the second pause segment position sequence. The speech signal is time segmented according to the second pause segment position sequence to obtain multiple speech continuous segments. The speech continuous segment feature vector is generated by calculating the duration, average energy and pitch period parameters of the speech continuous segment. The dynamic time warping algorithm is used to calculate the matching degree between the speech continuous segment feature vector and the preset action element feature vector, and the action element with the highest matching degree is selected as the corresponding item for each speech continuous segment to generate a corresponding relationship sequence between the speech segment and the action element. For the corresponding relationship sequence between speech segments and action elements, the ratio of the standard duration of each action element to the actual duration of the corresponding speech continuous segment is calculated to obtain a duration ratio sequence. The cubic spline interpolation algorithm is used to perform nonlinear scaling on the execution timing of the action elements to generate a time-aligned action sequence. The time-aligned action sequence is smoothed by calculating the feature similarity between adjacent action elements. If the similarity is greater than a preset threshold, the adjacent action elements are merged, and the timing parameters of the merged action elements are calculated using a weighted average method. In the process of extracting the speech timing distribution features, the speech pause segment is identified by calculating the short-time energy difference value and the zero-crossing rate difference value of adjacent speech frames. For a speech signal with a sampling frequency of 16000 Hz, the frame length is set to 320 sampling points and the frame shift is set to 160 sampling points. After framing by a rectangular window, the first preset threshold of the short-time energy difference value is set to 0.3, and the second preset threshold of the zero-crossing rate difference value is set to 0.25. When the two difference values ​​exceed the corresponding thresholds at the same time, it is marked as the start frame of the pause segment. When accurately locating the pause segment boundary, the 13-dimensional Mel-frequency cepstrum coefficients are used as feature parameters, and the Euclidean distance between adjacent frames is calculated to obtain the feature distance sequence. When the recursive binary method is used to segment the feature distance sequence, the initial segmentation point is selected at the local peak position of the feature distance. The segmentation point is optimized by calculating the variance ratio of the subsequences before and after the segmentation point, and the variance ratio threshold is set to 2.5. The feature extraction of speech continuous segments includes two aspects: time domain and frequency domain. In the time domain, the duration of each speech continuous segment is calculated, which is usually between 200 milliseconds and 2000 milliseconds. The average energy is obtained by calculating the arithmetic mean of the short-time energy of all frames in the speech continuous segment. The fundamental frequency period parameter is extracted using the autocorrelation method, and the search range is set between 50 Hz and 400 Hz.When calculating the matching degree between the continuous segment of speech and the preset action element, the dynamic time warping algorithm sets the local path constraint to symmetric type and the slope constraint range is 0.5 to 2. The preset action element feature vector contains three dimensions: the duration of the action, the rate of change of speed and the rate of change of acceleration. The matching degree threshold is set to 0.75, and only matches exceeding this threshold are considered valid matches. When performing time scaling of action elements, the duration ratio sequence reflects the difference in duration between the speech segment and the action element. The cubic spline interpolation algorithm is used to perform nonlinear scaling on the execution timing of the action element. The interpolation nodes are selected at the key time points of the action element, usually including the starting point, peak point and end point of the action. For the control of the action speed, the duration ratio of the acceleration segment and the deceleration segment is maintained at 3:2. The feature similarity calculation of adjacent action elements is based on the spatial position and movement direction of the action. The spatial position is represented by three-dimensional coordinates, and the movement direction is represented by a unit vector. The feature similarity threshold is set to 0.8, and adjacent action elements exceeding this threshold are merged. The timing parameters of the merged action elements are calculated using a weighted average method, with the weight being proportional to the duration of the original action elements.

[0018] S103, detecting whether there is a speech pause segment in the initially aligned action sequence, if so, judging whether the action element sequence needs to be reorganized according to the pause duration and the duration of the action element, and when the pause duration is greater than the threshold of the target action element duration, reorganizing the action element sequence to generate a reorganized action sequence.

[0019] A short-time energy calculation method is used to analyze the speech signal to obtain a first speech pause feature sequence, and a short-time zero-crossing rate detection method is used to perform secondary judgment on the first speech pause feature sequence to obtain a second speech pause feature sequence; the ratio of the duration of the speech pause segment to the target duration of the action element is calculated according to the second speech pause feature sequence, and if the duration ratio is greater than a preset duration threshold, the action element is marked as an element to be reorganized to obtain a sequence of action elements to be reorganized; for the sequence of action elements to be reorganized, a hierarchical clustering algorithm is used to calculate the similarity matrix between the action elements, and the similarity matrix is ​​segmented by the action feature threshold to obtain a sequence of action element reorganization schemes; according to the sequence of reorganization schemes, a minimum spanning tree algorithm is used to construct a connection relationship graph of the action elements, and the action element combination order is determined by calculating the weight values ​​of the connecting edges in the connection relationship graph to obtain a reorganization time sequence, and the weight values ​​of the edges are calculated by the following formula:

[0020] ,

[0021] w ij represents the edge weight from node i to node j, s ij represents the similarity between nodes i and j, N iRepresents the neighbor set of node i.

[0022] Exemplarily, according to the preliminary aligned action sequence, the time interval sequence between adjacent action elements is calculated, and the speech signal in each time interval is analyzed by using a short-time energy calculation method to obtain a first speech pause feature sequence. For the first speech pause feature sequence, a short-time zero-crossing rate detection method is used to perform a secondary judgment on the speech signal, and the precise boundary position of the speech pause segment is determined by calculating the zero-crossing rate difference between adjacent frames to obtain a second speech pause feature sequence. According to the second speech pause feature sequence, the ratio of the duration of the speech pause segment to the target duration of the corresponding action element is calculated. If the ratio is greater than a preset duration threshold, the action element is marked as an element to be reorganized, and a sequence of action elements to be reorganized is obtained. For the sequence of action elements to be reorganized, a hierarchical clustering algorithm is used to calculate the similarity matrix between the action elements, and the similarity matrix is ​​segmented by setting an action feature threshold to obtain a sequence of reorganization schemes of the action elements. According to the sequence of reorganization schemes, a minimum spanning tree algorithm is used to construct a connection relationship diagram between the action elements, and the combination order of the action elements is determined by calculating the weight value of the connecting edge to generate a reorganization timing sequence of the action elements. The action elements in the reorganized timing sequence are adjusted in duration. By calculating the ratio of the duration of the speech pause segment to the execution duration of the reorganized action sequence, the linear interpolation algorithm is used to scale the execution speed of the action elements to obtain the reorganized action sequence. In the analysis of speech pause features, the speech signal is processed by short-time energy calculation. For the speech signal with a sampling frequency of 16000 Hz, the frame length is set to 320 sampling points, the frame shift is set to 160 sampling points, and the speech signal is framed by a rectangular window. The short-time energy value of each frame signal is obtained by calculating the square sum of the sampling points. For normal speech segments, the short-time energy value is usually between 0.5 and 1, while the short-time energy value of the pause segment is less than 0.1. Zero-crossing rate detection is used to further confirm the speech pause boundary, which is measured by calculating the number of times the amplitude of the speech signal changes from positive to negative or from negative to positive. For a speech signal sampled at 16,000 Hz, the zero-crossing rate of the voiced segment is usually between 1,000 and 2,000 times per second, while the zero-crossing rate of the pause segment is higher than 2,500 times per second. When judging the boundary of the pause segment, if the zero-crossing rate of 5 consecutive frames exceeds the threshold, it is confirmed as the starting point of the pause segment. The target duration of the action element is determined based on the preset standard action library. The standard action library contains a variety of basic action elements, each of which has a predefined standard execution duration, such as 800 milliseconds for a waving action and 500 milliseconds for a nodding action. When the ratio of the duration of the speech pause segment to the standard duration of the action element exceeds 1.5, the action reorganization mechanism is triggered. When calculating the similarity of action elements, the hierarchical clustering algorithm uses multidimensional feature vectors to represent the action elements, including parameters such as spatial position, movement speed, and movement direction. The spatial position is represented by three-dimensional coordinates, the movement speed is represented by a scalar value, and the movement direction is represented by a unit vector.Cosine similarity is used to calculate the similarity between feature vectors, and the similarity threshold is set to 0.75. In the minimum spanning tree algorithm, the weight value of the connecting edge reflects the difficulty of transition between action elements. The calculation of the weight value takes into account factors such as the spatial distance of the starting position of the action, the angle of the movement direction, and the rate of change of the speed. A smaller weight value indicates a smoother action transition. The weight threshold is usually set to 0.4, and the connecting edges below this threshold are preferred. During the reorganization of the action sequence, the execution speed of the action element needs to be adjusted to match the duration of the speech pause segment. The action execution speed is scaled by a linear interpolation algorithm, and the scaling ratio range is limited to between 0.5 and 2 to avoid the action speed being too fast or too slow. The interpolation nodes are selected at the key time points of the action, usually including the starting point, peak point, and end point of the action.

[0023] S104, performing secondary alignment on the reorganized action sequence and the timing distribution of the continuous speech segment. When the difference between the start time and the end time of the action element and the start time and the end time of the corresponding speech segment is less than a preset synchronization threshold, it is determined that the coupling relationship between the reorganized action sequence and the speech meets the synchronization requirement.

[0024] According to the reorganized action sequence, a short-time window detection method is used to extract action element feature parameters, and an action timing feature sequence is obtained by calculating the amplitude change rate and speed change rate of the action element; for the action timing feature sequence, an adaptive threshold method is used to detect the local extreme value points of the action amplitude, and the start and end time sequence of the action element is obtained by calculating the time interval between the local extreme value points; according to the start and end time sequence of the action element, a speech energy envelope extraction method is used to obtain the boundary time sequence of the speech paragraph by calculating the change trend of the speech energy; the time difference between the start and end time sequence of the action element and the boundary time sequence of the speech paragraph is calculated, and if the time difference is less than a first preset threshold, the dynamic programming algorithm is used to calculate the alignment path; according to the alignment path, the execution speed of the action element is adjusted to obtain a corrected action timing sequence; for the corrected action timing sequence, the start time difference and the end time difference between each action element and the corresponding speech paragraph are calculated, and if the start time difference and the end time difference are both less than a second preset threshold, it is determined that the action element meets the synchronization requirement, wherein the start time difference and the end time difference are calculated by the following formula:

[0025] ,

[0026] ΔT start Indicates the starting time difference between the action element and the corresponding speech segment, T action,start Indicates the start time of the action element, T audio,start Indicates the start time of the corresponding speech segment.

[0027] ,

[0028] ΔT end Indicates the end time difference between the action element and the corresponding speech segment, T action,end Indicates the end time of the action element, T audio,end Indicates the end time of the corresponding voice segment.

[0029] Exemplarily, according to the reorganized action sequence, a short-time window detection method is used to extract the characteristic parameters of each action element, and the first action timing feature sequence is obtained by calculating the action amplitude change rate and speed change rate. The boundary of the first action timing feature sequence is located, and the local extreme point of the action amplitude is detected by the adaptive threshold method. The start and end time sequence of the action element is obtained by calculating the time interval between adjacent extreme points. For the start and end time sequence of the action element, the speech energy envelope extraction method is used to obtain the boundary features of the speech continuous segment, and the boundary time sequence of the speech paragraph is obtained by calculating the change trend of the speech energy. The start and end time sequence of the action element is aligned and matched with the boundary time sequence of the speech paragraph, and the time difference between the time sequences is calculated. If the time difference is less than the first preset threshold, it is marked as a timing matching point. According to the timing matching point, the dynamic programming algorithm is used to calculate the optimal alignment path between the action element and the speech paragraph, and the corrected action timing sequence is obtained by adjusting the execution speed of the action element. For the corrected action timing sequence, the start time difference and end time difference between each action element and the corresponding speech paragraph are calculated. If both differences are less than the second preset threshold, the action element is determined to meet the synchronization requirements. The action timing feature extraction adopts a short-time window detection method, the window length is set to 200 milliseconds, and the window overlap rate is 50%. The action amplitude change rate is obtained by calculating the displacement difference between adjacent sampling points. For a standard action sequence, the peak value of the amplitude change rate is usually between 0.5 meters and 2 meters per second. The velocity change rate reflects the acceleration characteristics of the action, which is obtained by calculating the first-order difference of the velocity sequence. The typical action acceleration peak is between 2 meters and 5 meters per square second. When detecting the local extreme point of the action amplitude, the adaptive threshold method uses the sliding average method to calculate the background noise level, and the threshold value is set to 3 times the background noise mean. For action elements with a duration of more than 500 milliseconds, there are usually 2 to 4 local extreme points in the action sequence. The time interval between adjacent extreme points reflects the rhythmic characteristics of the action, and the interval duration is usually between 200 milliseconds and 800 milliseconds. The extraction of speech energy envelope adopts short-time energy calculation method. For speech signal sampled at 16000 Hz, the frame length is set to 320 sampling points and the frame shift is 160 sampling points. The speech signal is windowed by Hamming window to calculate the energy value of each frame signal. The boundary characteristics of speech paragraphs are reflected by the energy change rate. At the start and end positions of speech, the absolute value of the energy change rate is usually greater than 0.3. The determination of timing matching points adopts double threshold criteria. The first preset threshold is set to 50 milliseconds to determine the degree of alignment between action elements and speech paragraphs on the time axis. When the time difference between the start and end time of the action and the speech boundary is less than the threshold, it is considered that the two are aligned at this time point. In practical applications, about 70% of the key time points in the action sequence can match the speech boundary.When calculating the optimal alignment path, the dynamic programming algorithm sets the time window width to 400 milliseconds and searches for the best motion speed scaling factor within this window. The scaling factor value range is limited to between 0.5 and 2 to avoid excessive changes in motion speed. For standard motion sequences, about 80% of the motion elements only require less than 30% speed adjustment to achieve good synchronization with speech. The synchronization requirement is determined using a strict dual-threshold standard, with the second preset threshold set to 30 milliseconds, which is used to limit the start time difference and the end time difference of the action, respectively. Practice has shown that when the time difference is controlled within this threshold range, it is difficult for the human eye to detect the asynchrony between the action and speech. For complex motion and speech sequences, usually more than 90% of the motion elements can meet this synchronization requirement.

[0030] S105, calculating the transition time and spatial continuity between the action elements of the reorganized action sequence, evaluating the action fluency by calculating the differences in time and space between adjacent action elements, and obtaining the speed, acceleration and angle changes during the action execution process to generate a fluency score result.

[0031] A three-dimensional coordinate sampling method is used to obtain spatial sampling points between adjacent action elements, and the Euclidean distance value and the direction vector are calculated based on the spatial sampling points to obtain a displacement sequence and an angle sequence; the displacement sequence and the angle sequence are smoothed, and a cubic spline interpolation algorithm is used to calculate a transition trajectory curve, and the curvature value and the torsion value at the node are calculated based on the transition trajectory curve to obtain a first transition feature sequence; based on the first transition feature sequence, the instantaneous velocity value and the instantaneous acceleration value of the sampling point in the first transition feature sequence are calculated using the central difference method, and the acceleration threshold is used for segmentation processing to obtain a second transition feature sequence; for the second transition feature sequence, the angular velocity parameters and the angular acceleration parameters of the action transition stage are calculated, and the least squares method is used to perform curve fitting on the parameters to obtain a third transition feature sequence; feature extraction is performed on the third transition feature sequence, and the velocity fluctuation coefficient, the acceleration peak ratio, the angle change rate and the trajectory smoothness are calculated to generate a fluency score result.

[0032] Exemplarily, according to the reorganized action sequence, a three-dimensional coordinate sampling method is used to extract the spatial position data between adjacent action elements, the displacement sequence is obtained by calculating the Euclidean distance between the sampling points, and the angle sequence is obtained by calculating the direction vector between the adjacent sampling points. The displacement sequence and the angle sequence are smoothed, and the transition trajectory curve is calculated by using the cubic spline interpolation algorithm. The first transition feature sequence is obtained by calculating the curvature value and the torsion value of the curve at each node. For the first transition feature sequence, the instantaneous velocity value and the instantaneous acceleration value of each sampling point are calculated by the central difference method, and the acceleration curve is segmented by setting the maximum acceleration threshold to obtain the second transition feature sequence. According to the second transition feature sequence, the angular velocity and angular acceleration parameters of the action transition stage are calculated, and the least squares method is used to fit the angle change curve to obtain the third transition feature sequence. Feature extraction is performed on the third transition feature sequence, and the velocity fluctuation coefficient, the acceleration peak ratio, the angle change rate and the trajectory smoothness are calculated to obtain the action fluency feature vector. For the action fluency feature vector, the support vector regression algorithm is used to construct the scoring function, and the action fluency scoring result is obtained by calculating the distance value between the feature vector and the preset standard feature template.

[0033] ,

[0034] S represents the result of action fluency scoring, n represents the number of feature vectors, v i represents the i-th eigenvector, t i represents the i-th preset standard feature template, d(v i ,t i) represents the distance value between the feature vector and the preset standard feature template, and max(d) represents the maximum value of all distance values. This formula calculates the average normalized distance between all feature vectors and the corresponding preset standard feature template, and then subtracts this value from 1 to obtain the final score. The closer the score result is to 1, the smoother the action. The spatial position sampling of the action sequence is recorded in three-dimensional coordinates with a fixed frequency. For the standard action sequence, the sampling frequency is set to 60 Hz, and each sampling point contains three coordinate components: X, Y, and Z. When calculating the Euclidean distance between adjacent sampling points, the displacement sequence reflects the spatial trajectory characteristics of the action. When the displacement value is less than 5 cm, it is considered to be in a state of action pause. The direction vector is obtained by calculating the displacement vectors of adjacent sampling points, and the change in the direction angle reflects the turning characteristics of the action. When processing the transition trajectory, the cubic spline interpolation algorithm sets the boundary condition to the natural boundary condition to ensure that the second-order derivative of the curve at the node is continuous. The curvature value reflects the curvature of the trajectory. Under normal circumstances, the maximum curvature value of the transition trajectory is controlled within 0.5. The torsion value describes the degree of twisting of the spatial curve. The torsion peak value of the standard action sequence usually does not exceed 0.3. The instantaneous velocity and acceleration are calculated using the 5-point central difference formula. For a 60 Hz sampled motion sequence, the time step is 16.7 milliseconds. The maximum acceleration threshold is set to twice the gravity acceleration, that is, 19.6 meters per square second. Motions exceeding this threshold are judged as non-smooth transitions. The volatility of the velocity curve is evaluated by calculating the standard deviation of adjacent velocity values. The angular velocity and angular acceleration parameters reflect the severity of the change in the direction of the motion. In a standard motion transition, the angular velocity peak value usually does not exceed 90 degrees per second, and the angular acceleration peak value is controlled within 180 degrees per square second. When fitting the angle change curve, the least squares method uses a third-order polynomial as the fitting function. The motion smoothness feature vector contains four dimensions: the velocity fluctuation coefficient reflects the stability of the velocity curve, with a value range between 0 and 1; the acceleration peak ratio represents the ratio of the maximum acceleration to the average acceleration, which is usually controlled within 3; the angle change rate describes the smoothness of the direction change, ideally not exceeding 0.5; the trajectory smoothness is calculated by the weighted average of the curvature and torsion. The support vector regression algorithm uses the radial basis kernel function, and the kernel parameter is set to 0.1. The preset standard feature template comes from a professional action database, which contains the smoothness features of multiple sets of typical action sequences. The distance value is obtained by calculating the Mahalanobis distance between the feature vector and the template. The final score result is expressed in percentage. A score above 80 indicates smooth action transition.

[0035] S106: If the fluency score is lower than a preset score threshold, the timing and spatial position of the action elements are adjusted, and an optimized action sequence is obtained and stored by fine-tuning the action in time and space.

[0036] Action elements are detected according to a preset fluency score threshold, and the deviation value of the action element on the time axis and the space trajectory is calculated using a dual threshold detection method to obtain a first space-time optimization vector; for the timing parameters in the first space-time optimization vector, a particle swarm optimization algorithm is used to adjust the start and end time of the action, and the second space-time optimization vector is obtained by calculating the minimum time interval constraint and the maximum delay constraint; according to the spatial parameters in the second space-time optimization vector, a three-dimensional spline interpolation algorithm is used to reconstruct the control point position coordinates and tangent direction of the action trajectory to obtain a third space-time optimization vector; for the trajectory nodes in the third space-time optimization vector, a Bezier curve fitting algorithm is used to calculate the continuity constraint conditions of the trajectory nodes to obtain a fourth space-time optimization vector; for the fourth space-time optimization vector, a kinematic constraint verification method is used to verify the speed continuity and acceleration continuity of the action, and after the verification is passed, an optimized action sequence is formed and stored.

[0037] Exemplarily, according to the fluency score result, a dual threshold detection method is used to identify action elements with scores lower than a preset threshold, and the first spatiotemporal optimization vector is obtained by calculating the overlap of the action elements on the time axis and the deviation value on the spatial trajectory. For the timing parameters in the first spatiotemporal optimization vector, a particle swarm optimization algorithm is used to adjust the start and end time of the action, and the second spatiotemporal optimization vector is obtained by setting the minimum time interval constraint and the maximum delay constraint. According to the spatial parameters in the second spatiotemporal optimization vector, a three-dimensional spline interpolation algorithm is used to reconstruct the action trajectory, and the third spatiotemporal optimization vector is obtained by adjusting the position coordinates and tangent direction of the control point. The third spatiotemporal optimization vector is optimized for curve smoothness, and the trajectory nodes are adjusted by using the Bezier curve fitting algorithm. The fourth spatiotemporal optimization vector is obtained by calculating the continuity constraint of the curve. For the fourth spatiotemporal optimization vector, the kinematic constraint verification method is used to verify the speed continuity and acceleration continuity of the action. If the constraint condition is not met, the corresponding parameters are corrected to obtain the fifth spatiotemporal optimization vector. The optimized action sequence is constructed according to the fifth spatiotemporal optimization vector, and the timing parameters, spatial parameters and constraints of the action sequence are classified and stored using a relational database. In the process of action optimization, the dual threshold detection uses two dimensions, fluency score and spatiotemporal consistency, for judgment. The fluency score threshold is set to 80 points, and the spatiotemporal consistency threshold is set to 0.85. For the standard action sequence, the overlap on the time axis is obtained by calculating the time interval between adjacent action elements. Under normal circumstances, the overlap should be controlled within 15%. The deviation value of the spatial trajectory is obtained by calculating the Euclidean distance between the action trajectory and the standard trajectory template. The deviation value is usually no more than 10 cm. When adjusting the action timing, the particle swarm optimization algorithm sets the population size to 50 and the maximum number of iterations to 100. The minimum time interval constraint is set to 200 milliseconds to ensure sufficient transition time between actions. The maximum delay constraint is set to 500 milliseconds to avoid the overall rhythm of the action sequence being too slow. During the optimization process, each particle represents a set of possible timing parameters, and the particle position is updated by continuous iteration to search for the optimal timing solution. The three-dimensional spline interpolation algorithm uses cubic Bezier curves as the basic curves, and each curve is defined by 4 control points. The position coordinates of the control points are obtained by minimizing the curve energy functional, and the tangent direction is determined by calculating the positional relationship between adjacent control points. For complex motion trajectories, they are usually decomposed into multiple Bezier curves, and the second-order derivative continuity is guaranteed between segments. During the Bezier curve fitting process, the trajectory nodes are resampled using a uniform parameterization method, and the number of sampling points is set to twice the number of original nodes. The continuity constraints include three levels: position continuity, velocity continuity, and acceleration continuity. The optimal position of the control point is solved by constructing a set of constraint equations. For standard motion trajectories, the fitting error is usually controlled within 5 mm.Kinematic constraint verification includes two aspects: velocity constraint and acceleration constraint. Velocity continuity requires that the velocity change rate between adjacent sampling points does not exceed 30%, while acceleration continuity requires that the acceleration change rate be controlled within 50%. For parameters that do not meet the constraints, linear interpolation method is used for correction, and the interpolation weight is inversely proportional to the degree of constraint violation. The relational database adopts a hierarchical storage structure. The first layer stores the basic information of the action sequence, including sequence identification, creation time and number of optimizations. The second layer stores timing parameters, including the start time, duration and transition time of each action element. The third layer stores spatial parameters, including control point coordinates, curve parameters and constraints. The efficiency of data retrieval is improved by establishing indexes, and the typical retrieval time is controlled within 10 milliseconds.

[0038] S107, inputting the optimized action sequence into the action controller, generating action execution instructions in combination with the speech time distribution characteristics, and completing the synchronous output of the action and the speech.

[0039] A high-precision clock source is used to obtain the timestamp information of the action elements and the timestamp information of the voice segments, and a timing association table is generated to record the start and end times of the voice segments corresponding to the action elements to obtain a first synchronization control sequence; according to the first synchronization control sequence, a joint interpolation algorithm is used to calculate the joint position parameters and the joint speed parameters, and the motion trajectory is planned by the kinematic forward equation to obtain a second synchronization control sequence; for the second synchronization control sequence, a circular queue structure is used to cache the action parameters, and the action data is obtained by setting the queue read and write pointers. If the amount of data in the queue reaches a preset cache threshold, the action control signal is output; according to the sampling phase difference between the action control signal and the voice playback signal, a proportional-integral controller is used to adjust the action execution speed, and the timing deviation is compensated by adjusting the execution cycle to obtain a corrected action control signal; the corrected action control signal and voice playback signal are output in parallel, and a synchronous clock signal is generated by a hardware timer to achieve the synchronous output of the action execution signal and the voice playback signal.

[0040] Exemplarily, according to the optimized action sequence, a high-precision clock source is used to extract the timestamp information of the action elements and the voice segments, and the start and end times of the voice segments corresponding to each action element are recorded by generating a timing association table to obtain a first synchronization control sequence. For the first synchronization control sequence, a joint interpolation algorithm is used to calculate the joint position parameters and joint velocity parameters during the execution of the action, and the action trajectory is planned in real time through the kinematic forward solution equation to obtain a second synchronization control sequence.

[0041] ,

[0042] θ(t) represents the joint angle at time t, θ0 represents the initial angle, and θ fRepresents the final angle, t represents the current time, and T represents the total execution time. This formula uses quintic polynomial interpolation to calculate the joint position.

[0043] ,

[0044] ω(t) represents the joint angular velocity at time t, θ0 represents the initial angle, and θ fRepresents the final angle, t represents the current time, and T represents the total execution time. This formula obtains the joint speed by taking the derivative of the joint position interpolation formula. According to the second synchronization control sequence, a circular queue structure is used to cache the action parameters, and the action data is synchronously read by setting the queue read and write pointer. If the amount of data in the queue reaches the preset cache threshold, the action control signal is output. The output action control signal is sampled and detected, and the real-time synchronization deviation sequence is obtained by calculating the phase difference between the action sampling signal and the voice sampling signal. According to the real-time synchronization deviation sequence, a proportional integral controller is used to adjust the action execution speed, and the timing deviation is compensated by adjusting the execution cycle to obtain the corrected action control signal. The corrected action control signal and voice playback signal are output in parallel, and a synchronous clock signal is generated by a hardware timer to achieve the synchronous output of the action execution signal and the voice playback signal. The high-precision clock source uses a 100MHZ crystal oscillator to provide a reference clock signal, and a 1000MHZ system clock is generated by a frequency multiplication circuit to achieve microsecond timestamp accuracy. The timing association table uses a key-value pair structure to record the corresponding relationship, in which the action element identifier is used as the key and the start and end times of the voice segment are used as the value. For standard action sequences, each action element corresponds to 2 to 3 voice clips on average. The joint interpolation algorithm uses a quintic polynomial interpolation method to ensure that the position, velocity, and acceleration are continuous at the end points of the trajectory. The joint position parameters contain 6 degrees of freedom, and the range of motion of each degree of freedom is subject to mechanical limit constraints. The kinematics forward solution uses the DH parameter method to obtain the spatial position of the end effector by calculating the link coordinate transformation. For complex action sequences, the forward solution calculation frequency is maintained at 1000 Hz. The circular queue uses a dual pointer structure to implement circular storage of data. The queue length is set to 1024, and the initial interval between read and write pointers is 512. The preset cache threshold is set to 75% of the queue length, that is, the output operation is triggered when the data volume reaches 768. This design provides sufficient buffer space to cope with data flow fluctuations while ensuring data continuity. During the sampling detection process, the sampling frequency of the action signal and the voice signal is set to 44.1KHZ, and each frame of sampled data contains 2048 sampling points. The phase difference is obtained by calculating the cross-correlation function of the two signals. When the phase difference exceeds plus or minus 5 degrees, the synchronization compensation mechanism is triggered. The parameters of the proportional-integral controller are obtained through experimental calibration, with the proportional coefficient set to 0.8 and the integral coefficient set to 0.2. The adjustment range of the execution cycle is limited to plus or minus 10% to avoid excessive changes in the action speed. For a typical action sequence, the response time of the controller is less than 10 milliseconds, and the steady-state error is controlled within 1 millisecond. The hardware timer is implemented using a 16-bit counter, and the clock division factor is set to 8 to generate a 125KHZ synchronous clock signal. The action execution signal and the voice playback signal are cached through a dual-port RAM and output simultaneously on the rising edge of the synchronous clock.In practical applications, the time deviation of the two signals is controlled within 40 microseconds, which has achieved complete synchronization for the perception of the human eye and ear.

[0045] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the concept of the present application. For example, the above features are replaced with the technical features with similar functions disclosed in the present application (but not limited to) to form a technical solution.

Claims

1. A method for synchronizing speech and motion, the method for synchronizing speech and motion being used for intelligent animation synthesis, characterized in that: The method comprises: Acquire a speech signal, extract the speech duration feature of the speech signal, perform segmentation processing on the speech signal in combination with a preset speech segmentation rule, divide the speech into speech continuous segments and speech pause segments, and extract the short-term features corresponding to each segment to obtain the speech timing distribution feature; detect the starting point and the ending point of the speech pause segment according to the speech timing distribution feature, determine the starting time, the ending time and the duration of the pause segment, combine the speech timing distribution feature with the preset action element, calculate the timing alignment relationship between the action element and the speech continuous segment, and generate a preliminary aligned action sequence; detect whether the preliminary aligned action sequence has a speech pause segment, and if so, determine whether the action element sequence needs to be reorganized according to the pause duration and the duration of the action element, and when the pause duration is greater than the threshold of the target action element duration, reorganize the action element sequence to generate a reorganized action sequence; reorganize the reorganized action sequence The action sequence is aligned with the timing distribution of the continuous speech segment twice. When the difference between the start time and the end time of the action element and the start time and the end time of the corresponding speech segment is less than the preset synchronization threshold, it is determined that the coupling relationship between the reorganized action sequence and the speech meets the synchronization requirement; the transition time and spatial continuity between the action elements of the reorganized action sequence are calculated, and the action fluency is evaluated by calculating the difference between adjacent action elements in time and space, and obtaining the speed, acceleration and angle changes during the action execution process to generate a fluency score result; if the fluency score is lower than the preset score threshold, the timing and spatial position of the action element are adjusted, and the optimized action sequence is obtained and stored by fine-tuning the action in time and space; the optimized action sequence is input into the action controller, and the action execution instruction is generated in combination with the timing distribution characteristics of the speech, so as to complete the synchronous output of the action and the speech.

2. The method according to claim 1, characterized in that The method of acquiring a speech signal, extracting the speech duration feature of the speech signal, segmenting the speech signal in combination with a preset speech segmentation rule, dividing the speech into speech continuous segments and speech pause segments, and extracting the short-term features corresponding to each segment to obtain the speech time series distribution feature includes: The speech signal is smoothed and denoised by using a wavelet transform method, the signal amplitude is normalized by setting a high and low double threshold, and a Butterworth low-pass filter is used to eliminate high-frequency noise interference to obtain a speech signal after interference is eliminated; The short-time energy function and the short-time zero-crossing rate function are calculated according to the speech signal after eliminating interference to obtain a speech signal strength sequence, and if the signal strength is greater than a preset volume threshold, the speech starting point is marked, and the speech duration sequence is obtained according to the interval between adjacent starting points; Segmenting the speech signal after eliminating interference according to the speech duration sequence, dividing the speech frames by using a Hamming window function, determining the speech frame boundary by calculating the energy ratio between frames, and extracting the fundamental pitch period and the formant frequency from the speech frame to obtain a speech basic feature sequence; A continuity analysis is performed based on the basic speech feature sequence. If the pitch period change rate of adjacent speech frames is less than a preset threshold, it is determined to be a continuous speech segment. The maximum entropy model is used to optimize the speech segment boundary. Energy feature vectors and time-varying feature vectors are extracted from the continuous speech segment and pause segment to obtain the speech timing distribution characteristics.

3. The method according to claim 1, characterized in that The method comprises: detecting the starting point and the ending point of the speech pause segment according to the speech timing distribution feature, determining the starting time, the ending time and the duration of the pause segment, combining the speech timing distribution feature with the preset action element, calculating the timing alignment relationship between the action element and the speech continuous segment, and generating a preliminary aligned action sequence, including: Acquire a short-time energy difference value and a zero-crossing rate difference value of adjacent speech frames according to the time series distribution characteristics of the speech signal, and if the short-time energy difference value is greater than a first preset threshold and the zero-crossing rate difference value is greater than a second preset threshold, mark the speech frame as a pause segment start frame; The pause segment boundary is located by using a recursive binary division method, a characteristic distance sequence is obtained by calculating the Mel frequency cepstral coefficient distance between adjacent speech frames, and a second pause segment position sequence is determined according to the characteristic distance sequence; Dividing the speech continuous segments according to the second pause segment position sequence, and obtaining the speech continuous segment feature vector by calculating the duration, average energy and pitch period parameters of the speech continuous segments; A dynamic time warping algorithm is used to calculate the matching degree between the feature vector of the continuous speech segment and the feature vector of the preset action element, and the corresponding action element is selected according to the matching degree. The execution timing of the action element is nonlinearly scaled by the cubic spline interpolation algorithm to obtain a time-aligned action sequence.

4. The method according to claim 1, characterized in that The detecting whether there is a speech pause segment in the preliminary aligned action sequence, and if so, judging whether it is necessary to reorganize the action element sequence according to the pause duration and the duration of the action element, and when the pause duration is greater than the threshold of the target action element duration, reorganizing the action element sequence to generate a reorganized action sequence, including: A short-time energy calculation method is used to analyze the speech signal to obtain a first speech pause feature sequence, and a short-time zero-crossing rate detection method is used to perform a secondary judgment on the first speech pause feature sequence to obtain a second speech pause feature sequence; Calculating the ratio of the duration of the speech pause segment to the target duration of the action element according to the second speech pause feature sequence, and if the duration ratio is greater than a preset duration threshold, marking the action element as an element to be reorganized, and obtaining a sequence of action elements to be reorganized; For the sequence of action elements to be reorganized, a hierarchical clustering algorithm is used to calculate a similarity matrix between the action elements, and the similarity matrix is ​​segmented by an action feature threshold to obtain a sequence of action element reorganization schemes; According to the reorganization scheme sequence, a minimum spanning tree algorithm is used to construct an action element connection relationship graph, and the action element combination order is determined by calculating the weight values ​​of the connecting edges in the connection relationship graph to obtain a reorganization time sequence. The weight values ​​of the edges are calculated by the following formula: , w ij represents the edge weight from node i to node j, s ij represents the similarity between nodes i and j, N i Represents the neighbor set of node i.

5. The method according to claim 1, characterized in that The reorganized action sequence is aligned with the time distribution of the continuous speech segment twice, and when the difference between the start time and the end time of the action element and the start time and the end time of the corresponding speech segment is less than a preset synchronization threshold, it is determined that the coupling relationship between the reorganized action sequence and the speech meets the synchronization requirement, including: According to the reorganized action sequence, the action element characteristic parameters are extracted by using the short-time window detection method, and the action time sequence characteristic sequence is obtained by calculating the amplitude change rate and speed change rate of the action element; Adopting an adaptive threshold method to detect the local extreme value points of the action amplitude for the action time sequence feature sequence, and obtaining the start and end time sequence of the action element by calculating the time intervals between the local extreme value points; A speech energy envelope extraction method is used according to the start and end time sequence of the action element, and a boundary time sequence of the speech paragraph is obtained by calculating the change trend of the speech energy; Calculating a time difference between the start and end time sequence of the action element and the boundary time sequence of the speech paragraph, and if the time difference is less than a first preset threshold, using a dynamic programming algorithm to calculate an alignment path; According to the alignment path, by adjusting the execution speed of the action elements, a corrected action timing sequence is obtained; For the corrected action timing sequence, the start time difference and the end time difference between each action element and the corresponding speech segment are calculated. If the start time difference and the end time difference are both less than the second preset threshold, it is determined that the action element meets the synchronization requirement, where the start time difference and the end time difference are calculated by the following formula: , ΔT start Indicates the starting time difference between the action element and the corresponding speech segment, T action,start Indicates the start time of the action element, T audio,start Indicates the start time of the corresponding speech segment. , ΔT end Indicates the end time difference between the action element and the corresponding speech segment, T action,end Indicates the end time of the action element, T audio,end Indicates the end time of the corresponding voice segment.

6. The method according to claim 1, characterized in that The calculation of the transition time and spatial continuity between the action elements of the reorganized action sequence, by calculating the difference in time and space between adjacent action elements, by obtaining the speed, acceleration and angle changes during the action execution to evaluate the action fluency, and generate a fluency score result, including: A three-dimensional coordinate sampling method is used to obtain spatial sampling points between adjacent action elements, and a Euclidean distance value and a direction vector are calculated according to the spatial sampling points to obtain a displacement sequence and an angle sequence; Smoothing the displacement sequence and the angle sequence, calculating the transition trajectory curve using a cubic spline interpolation algorithm, and calculating the curvature value and the torsion value at the node according to the transition trajectory curve to obtain a first transition characteristic sequence; According to the first transition characteristic sequence, the instantaneous velocity value and the instantaneous acceleration value of the sampling point in the first transition characteristic sequence are calculated by using the central difference method, and segmented processing is performed by using the acceleration threshold to obtain a second transition characteristic sequence; For the second transition characteristic sequence, the angular velocity parameter and the angular acceleration parameter of the motion transition phase are calculated, and the least square method is used to perform curve fitting on the parameters to obtain a third transition characteristic sequence; The third transition feature sequence is extracted, and the velocity fluctuation coefficient, acceleration peak ratio, angle change rate and trajectory smoothness are calculated to generate a fluency score result.

7. The method according to claim 1, characterized in that If the fluency score is lower than the preset score threshold, the timing and spatial position of the action elements are adjusted, and the optimized action sequence is obtained and stored by fine-tuning the action in time and space, including: Detecting action elements according to a preset fluency score threshold, calculating deviation values ​​of the action elements on the time axis and the spatial trajectory using a dual threshold detection method, and obtaining a first spatiotemporal optimization vector; According to the timing parameters in the first spatiotemporal optimization vector, the particle swarm optimization algorithm is used to adjust the start and end time of the action, and the second spatiotemporal optimization vector is obtained by calculating the minimum time interval constraint and the maximum delay constraint; According to the spatial parameters in the second spatiotemporal optimization vector, a three-dimensional spline interpolation algorithm is used to reconstruct the position coordinates and tangent direction of the control point of the motion trajectory to obtain a third spatiotemporal optimization vector; For the trajectory nodes in the third space-time optimization vector, a Bezier curve fitting algorithm is used to calculate the continuity constraint conditions of the trajectory nodes to obtain a fourth space-time optimization vector; For the fourth space-time optimization vector, a kinematic constraint verification method is used to verify the velocity continuity and acceleration continuity of the action. After the verification is passed, an optimized action sequence is formed and stored.

8. The method according to claim 1, characterized in that The optimized action sequence is input into the action controller, and the action execution instruction is generated in combination with the speech timing distribution characteristics to complete the synchronous output of the action and the speech, including: A high-precision clock source is used to obtain action element timestamp information and voice segment timestamp information, and a timing association table is generated to record the start and end times of the voice segment corresponding to the action element, thereby obtaining a first synchronization control sequence; According to the first synchronous control sequence, the joint position parameters and the joint velocity parameters are calculated by using the joint interpolation algorithm, and the motion trajectory is planned by using the kinematics forward solution equation to obtain the second synchronous control sequence; For the second synchronous control sequence, a circular queue structure is used to cache the action parameters, and the action data is obtained by setting the queue read and write pointers. If the amount of data in the queue reaches a preset cache threshold, an action control signal is output; According to the sampling phase difference between the action control signal and the voice playback signal, a proportional-integral controller is used to adjust the action execution speed, and the timing deviation is compensated by adjusting the execution cycle to obtain a corrected action control signal; The corrected action control signal and voice playback signal are output in parallel, and a synchronous clock signal is generated by a hardware timer to achieve synchronous output of the action execution signal and the voice playback signal.

Citation Information

Patent Citations

  • Mouth shape animation synthesis method based on comprehensive weighted algorithm

    CN104361620A

  • Voice sample collection method based on network dubbing game

    CN107293286A