Enhanced voice feedback method based on dynamic time warping

By evaluating the completion of user actions through a dynamic time warping algorithm and combining it with a language generation model and voice cloning technology, emotional voice feedback that matches the user's voice is generated. This solves the problem of feedback content being out of touch with user performance and lacking personalization in existing systems, and achieves highly consistent and personalized voice feedback effects.

CN120612920APending Publication Date: 2025-09-09SHENZHEN HULE TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510933030.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing emotional voice feedback systems cannot effectively combine users' real-time behaviors and action states, resulting in the feedback content being disconnected from user performance, lacking personalization and dynamism, and unable to generate emotional voice that matches the user's timbre characteristics.

Method used

By obtaining the skeleton point sequence corresponding to the user's action, the kinematic dynamic time warping algorithm is applied to evaluate the completion of the action, and the preset score-emotion mapping rules and language generation model are combined to generate personalized emotional feedback text. The user's voice samples are used to train the timbre cloning model to generate voice feedback that matches the user's timbre.

Benefits of technology

It achieves a deep correlation between feedback content and user performance, ensuring that the emotional expression of voice feedback is highly consistent with the user's status, providing the ultimate personalized and immersive listening experience, and eliminating the mechanical feel and sense of distance of traditional systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612920A_ABST
    Figure CN120612920A_ABST
Patent Text Reader

Abstract

The invention discloses an enhanced voice feedback method, device and equipment based on dynamic time warping, and a computer readable storage medium, and the method comprises the steps: obtaining a user skeleton point sequence corresponding to a user action, and comparing the user skeleton point sequence with a standard action template through a kinematics dynamic time warping algorithm, generating an action evaluation score representing the degree of completion of the action; determining a user emotion level based on the action evaluation score; inputting the user emotion level and the action evaluation score into a pre-trained language generation type model, and generating an emotional feedback text matched with the user emotion level; determining a speech synthesis parameter based on the user emotion level; training and generating a personalized timbre clone model; and generating enhanced voice feedback with user timbre and emotion expression matched with the user emotion level. The method has the advantage of providing emotional speech enhancement feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction technology, and in particular to a method, device, equipment and computer-readable storage medium for enhanced voice feedback based on dynamic time regularization. Background Art

[0002] With the advancement of artificial intelligence, speech synthesis, and emotion recognition technologies, human-computer interaction is shifting from traditional one-way command feedback to more intelligent and emotional multimodal interactions. In line with this trend, speech feedback, as a key means of interaction between smart devices and users, has been widely used in fields such as smart assistants, virtual reality, and health monitoring. Among these, emotional speech synthesis technology, which generates speech with specific emotional overtones, is a key research direction for enhancing the naturalness and affinity of interactions.

[0003] However, existing emotional voice feedback systems have two main inherent flaws. First, emotional feedback is disconnected from the user's real-time behavior. Most systems focus only on the generation of voice emotions, but ignore the direct impact of the user's action status or task completion on the feedback content and emotional color, and fail to effectively link the two. Second, the feedback is not personalized and dynamic enough. Most existing systems rely on general emotional models or fixed static voice libraries. They are unable to dynamically adjust the tone, speed and content of the voice according to the user's real-time emotions or performance changes, and are even unable to generate personalized voice that matches the user's own timbre characteristics, resulting in feedback that appears monotonous, mechanical and lacks emotional involvement.

[0004] In summary, existing technologies for providing voice feedback fail to leverage the user's movements as the core driver of feedback content and emotional tone. Furthermore, they suffer from serious deficiencies in personalized feedback and dynamic responsiveness. Therefore, a new method for enhancing voice feedback is urgently needed to overcome the limitations of existing technologies. Summary of the Invention

[0005] The embodiments of the present application provide a method for enhancing voice feedback based on dynamic time warping, aiming to provide emotionally rich and personalized voice feedback interaction.

[0006] To achieve the above objectives, the present invention provides a method for enhancing speech feedback based on dynamic time warping, comprising:

[0007] Obtaining a user skeleton point sequence corresponding to the user action, and applying a kinematic dynamic time warping algorithm to compare the user skeleton point sequence with a standard action template to generate an action evaluation score representing the degree of action completion;

[0008] Based on the action evaluation score, determining the user emotion level corresponding to the action evaluation score through a preset score-emotion mapping rule;

[0009] Inputting the user's emotion level and the action evaluation score into a pre-trained language generation model to generate emotional feedback text that matches the user's emotion level;

[0010] Based on the user's emotion level, querying a preset emotion-speech parameter mapping library to determine speech synthesis parameters including intonation, speaking speed and tone;

[0011] Based on pre-collected user voice samples, a personalized voice cloning model is trained and generated to match the user's voice characteristics.

[0012] Based on the emotional feedback text, the speech synthesis parameters, and the personalized timbre cloning model, speech synthesis processing is performed to generate enhanced speech feedback with the user's timbre and emotional expression matching the user's emotional level.

[0013] In one embodiment, applying a kinematic dynamic time warping algorithm to generate the motion assessment score includes:

[0014] Calculate the position difference, velocity difference, and acceleration difference between each pair of frames in the user's skeleton point sequence and the standard action template;

[0015] Based on the position difference value, the velocity difference value, and the acceleration difference value, a kinematic local cost for dynamic time warping algorithm comparison is obtained;

[0016] Constructing a cost matrix using the kinematic local costs as matrix elements;

[0017] Based on the cost matrix, constructing a cumulative cost matrix through a dynamic programming recursive formula;

[0018] The final value of the cumulative cost matrix is ​​normalized to obtain the action evaluation score.

[0019] In one embodiment, obtaining a kinematic local cost for dynamic time warping algorithm comparison based on the position difference value, the velocity difference value, and the acceleration difference value includes:

[0020] Establishing a kinematic chain model that characterizes the force generation mode of the standard movement;

[0021] According to the kinematic chain model, quantify the contribution of different joints to the completion of the movement;

[0022] Determining a weight coefficient of each joint point for calculating the local cost based on the contribution;

[0023] Performing inter-dimensional weighted summation on the position difference value, velocity difference value, and acceleration difference value to obtain a kinematic comprehensive difference value for each joint point;

[0024] Based on the weight coefficient, the kinematic comprehensive difference values ​​of all joint points are summed to obtain the local cost.

[0025] In one embodiment, based on the action evaluation score, determining the user emotion level corresponding to the action evaluation score by using a preset score-emotion mapping rule includes:

[0026] Access a predefined mapping table that stores the correspondence between rating intervals and emotion level identifiers;

[0027] Comparing the action evaluation score with a plurality of scoring intervals in the mapping table to find and determine the scoring interval to which the action evaluation score belongs;

[0028] The emotion level identifier corresponding to the interval is output as the user emotion level.

[0029] In one embodiment, the user's emotion level and the action evaluation score are input into a pre-trained language generation model to generate emotional feedback text that matches the user's emotion level, including:

[0030] constructing the user emotion level and the action evaluation score into a structured input prompt;

[0031] Submitting the structured input prompt to a language generation model based on the Transformer architecture;

[0032] A text sequence output by the language generation model is received and decoded as the emotional feedback text.

[0033] In one embodiment, based on pre-collected user voice samples, training and generating a personalized voice cloning model that matches the user's voice characteristics includes:

[0034] Extracting a timbre feature vector representing the user's timbre characteristics from the user's voice sample;

[0035] A pre-trained general speech synthesis model is fine-tuned using the timbre feature vector to generate the personalized timbre cloning model.

[0036] In one embodiment, performing speech synthesis processing includes:

[0037] Initializing the personalized timbre cloning model;

[0038] Configuring the speech synthesis parameters to the speech synthesis engine of the model;

[0039] The speech synthesis engine is driven to process the emotional feedback text to generate and output the enhanced speech feedback in the form of a waveform sound file.

[0040] To achieve the above objectives, the present application further proposes an enhanced speech feedback device based on dynamic time warping, comprising:

[0041] An action evaluation module is used to obtain a user skeleton point sequence corresponding to a user action and compare the user skeleton point sequence with a standard action template using a kinematic dynamic time warping algorithm to generate an action evaluation score representing the degree of action completion;

[0042] An emotion determination module is used to determine the user emotion level corresponding to the action evaluation score based on the action evaluation score and using a preset score-emotion mapping rule;

[0043] A text generation module, configured to input the user's emotion level and the action evaluation score into a pre-trained language generation model to generate an emotional feedback text that matches the user's emotion level;

[0044] A parameter mapping module is used to query a preset emotion-speech parameter mapping library based on the user's emotion level to determine speech synthesis parameters including intonation, speaking speed and tone;

[0045] The voice modeling module is used to train and generate a personalized voice cloning model that matches the user's voice characteristics based on pre-collected user voice samples;

[0046] The speech synthesis module is used to perform speech synthesis processing based on the emotional feedback text, the speech synthesis parameters, and the personalized timbre cloning model to generate enhanced speech feedback with the user's own timbre and emotional expression matching the user's emotional level.

[0047] To achieve the above-mentioned purpose, an embodiment of the present application also proposes an enhanced voice feedback device based on dynamic time regularization, including a memory, a processor, and an enhanced voice feedback program based on dynamic time regularization stored in the memory and runnable on the processor. When the processor executes the enhanced voice feedback program based on dynamic time regularization, it implements the enhanced voice feedback method based on dynamic time regularization as described in any one of the above items.

[0048] To achieve the above-mentioned purpose, an embodiment of the present application also proposes a computer-readable storage medium, on which an enhanced speech feedback program based on dynamic time regularization is stored. When the enhanced speech feedback program based on dynamic time regularization is executed by a processor, the enhanced speech feedback method based on dynamic time regularization as described in any one of the above items is implemented.

[0049] Based on the above embodiments, the method for enhancing speech feedback based on dynamic time warping of the present application has at least the following beneficial effects:

[0050] 1. Achieve a deep correlation between feedback content and user performance

[0051] This application first accurately evaluates the user's real-time action completion by applying a kinematic dynamic time warping algorithm, and uses this evaluation result as one of the core inputs driving the language generation model. This ensures that the final generated voice feedback content is no longer a general phrase unrelated to the user's performance, but is instead highly relevant and intelligent feedback that can be dynamically "created" based on the user's specific performance (high or low score).

[0052] 2. Achieve a high degree of consistency between feedback emotion and user status

[0053] This application not only generates content, but also finely controls the "emotional expression" of feedback. By mapping the action evaluation score to a clear user emotion level, and further mapping this level to specific speech synthesis parameters (such as intonation, speaking speed, and tone), this method ensures that the acoustic characteristics of the speech output are consistent with the emotional tone of the text content. For example, encouraging text will be spoken in a gentle and soothing tone, while praising text will be spoken in a cheerful and upbeat tone, achieving a unity of content and form.

[0054] 3. Provides the ultimate personalized and immersive listening experience

[0055] This application innovatively introduces user voice cloning technology, fine-tuning a powerful general-purpose speech synthesis model to generate voices that closely resemble the user's voice. This feedback mechanism, which "uses my own voice to speak to me in a way that best understands my performance and emotions," completely eliminates the mechanical and distant feel of traditional voice systems, creating an unprecedented, highly engaging, and personalized immersive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0057] Figure 1 This is a module structure diagram of an embodiment of an enhanced speech feedback device based on dynamic time warping according to the present invention;

[0058] Figure 2 2. It is a flow chart of an embodiment of a method for enhancing speech feedback based on dynamic time warping of the present invention.

[0059] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0060] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0061] To better understand the above technical solutions, exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0062] It should be noted that in the claims, any reference signs placed between brackets shall not be construed as limiting the claims. The presence of "comprising" in the text does not exclude the presence of components or steps not listed in the claims. The quantifier "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim that lists several means, several of these means may be embodied by the same hardware item. The use of "first", "second", and "third" etc. does not indicate any order and these words may be interpreted as names.

[0063] like Figure 1 As shown, Figure 1 It is a structural diagram of server 1 (also called enhanced voice feedback device based on dynamic time regularization) of the hardware operating environment involved in the embodiment of the present invention.

[0064] The server of the embodiment of the present invention is a device with display function such as "Internet of Things devices", smart air conditioners, smart lights, smart power supplies with networking functions, AR / VR devices with networking functions, smart speakers, self-driving cars, PCs, smart phones, tablet computers, e-book readers, portable computers, etc.

[0065] like Figure 1 As shown, the server 1 includes: a memory 11 , a processor 12 and a network interface 13 .

[0066] The memory 11 includes at least one type of readable storage medium, including a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the server 1, such as a hard disk of the server 1. In other embodiments, the memory 11 may also be an external storage device of the server 1, such as a plug-in hard disk equipped on the server 1, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0067] Furthermore, the memory 11 may include both an internal storage unit of the server 1 and an external storage device. The memory 11 can be used not only to store application software installed on the server 1 and various data, such as the code of the enhanced speech feedback program 10 based on dynamic time warping, but also to temporarily store data that has been output or is about to be output.

[0068] In some embodiments, the processor 12 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip, used to run the program code or process data stored in the memory 11, such as executing the enhanced speech feedback program 10 based on dynamic time warping.

[0069] The network interface 13 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface), and is generally used to establish a communication connection between the server 1 and other electronic devices.

[0070] The network may be the Internet, a cloud network, a wireless fidelity (Wi-Fi) network, a personal area network (PAN), a local area network (LAN), and / or a metropolitan area network (MAN). Various devices in the network environment may be configured to connect to the communication network according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of the following: Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Light Fidelity (Li-Fi), 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocol, and / or Bluetooth communication protocol, or a combination thereof.

[0071] Optionally, the server may further include a user interface, which may include a display and an input unit such as a keyboard. The optional user interface may also include a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display, which may also be referred to as a display screen or display unit, is used to display information processed in the server 1 and to display a visual user interface.

[0072] Figure 1 Only the server 1 having components 11-13 and the enhanced speech feedback program 10 based on dynamic time warping is shown. It can be understood by those skilled in the art that Figure 1 The structure shown does not constitute a limitation on the server 1 , and the server 1 may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0073] In this embodiment, the processor 12 may be configured to call the enhanced speech feedback program based on dynamic time warping stored in the memory 11 and perform the following operations:

[0074] Obtaining a user skeleton point sequence corresponding to the user action, and applying a kinematic dynamic time warping algorithm to compare the user skeleton point sequence with a standard action template to generate an action evaluation score representing the degree of action completion;

[0075] Based on the action evaluation score, determining the user emotion level corresponding to the action evaluation score through a preset score-emotion mapping rule;

[0076] Inputting the user's emotion level and the action evaluation score into a pre-trained language generation model to generate emotional feedback text that matches the user's emotion level;

[0077] Based on the user's emotion level, querying a preset emotion-speech parameter mapping library to determine speech synthesis parameters including intonation, speaking speed and tone;

[0078] Based on pre-collected user voice samples, a personalized voice cloning model is trained and generated to match the user's voice characteristics.

[0079] Based on the emotional feedback text, the speech synthesis parameters, and the personalized timbre cloning model, speech synthesis processing is performed to generate enhanced speech feedback with the user's timbre and emotional expression matching the user's emotional level.

[0080] Based on the hardware architecture of the aforementioned enhanced speech feedback device based on dynamic time warping, an embodiment of the present invention's enhanced speech feedback method based on dynamic time warping is proposed. The present invention's enhanced speech feedback method based on dynamic time warping aims to provide emotionally rich and personalized speech feedback interactions.

[0081] Reference Figure 2 , Figure 2 This is an embodiment of a method for enhancing speech feedback based on dynamic time warping of the present invention, and the method for enhancing speech feedback based on dynamic time warping comprises the following steps:

[0082] S10, obtaining a user skeleton point sequence corresponding to the user action, and applying a kinematic dynamic time warping algorithm to compare the user skeleton point sequence with a standard action template to generate an action evaluation score representing the completion degree of the action;

[0083] Specifically, a user's skeletal point sequence is a time series of data that records the three-dimensional spatial coordinates of multiple predefined joints on the user's body at consecutive time points while the user performs an action. A standard action template is a pre-stored reference action sequence associated with a specific emotion or intention. The action evaluation score is a quantitative value, typically normalized to a range of 0 to 100, that represents the degree of kinematic conformity between the user's real-time action and the standard action template.

[0084] In a specific embodiment, the process of acquiring a user skeletal point sequence includes the following steps: First, the system continuously captures a video stream containing user movements using a visual sensor (e.g., a standard webcam). Each frame of the video stream is then fed into a preset pose estimation model optimized for real-time interaction, such as the one in the MediaPipe library. This model rapidly identifies the human body in each frame and extracts the spatial coordinates of all predefined joints. To ensure precise temporal correspondence between subsequent tactile feedback and the action that triggered it, the system appends a high-precision timestamp to the extracted coordinate data for each frame. Finally, all consecutive, timestamped pose coordinate data is organized in chronological order to form the user skeletal point sequence. The second stage involves comparison using a kinematic dynamic time warping algorithm. The core of this embodiment lies in the use of a cost function that incorporates multi-dimensional kinematic information as the basis for the DTW algorithm. Specifically, when comparing any pair of frames in the user sequence with a standard template, the system calculates the differences in three different kinematic dimensions: position difference, velocity difference, and acceleration difference. Then, the system performs a weighted summation of the difference values ​​of these three dimensions to obtain a comprehensive kinematic local cost. The calculation formula for this local cost is:

[0085] In a specific embodiment, the process of applying the kinematic dynamic time warping algorithm to generate the motion assessment score is achieved through the following steps S11 to S15:

[0086] First, in step S11, when comparing the user's skeleton point sequence with any pair of frames in the standard template, the system calculates their differences in three different kinematic dimensions: position difference, velocity difference, and acceleration difference. The technical details of this calculation are as follows:

[0087] Position difference value (d_pos): The system calculates the Euclidean distance between the coordinate vectors of all K corresponding joint points in two frames and sums or averages them. The formula can be expressed as:

[0088]

[0089] Velocity difference (d_vel): The system first calculates the velocity vector of each joint point in the player and the standard template by taking the first-order difference of the position coordinates between consecutive frames (i.e., (pos_t - pos_{t-1}) / Δt). Then, the system calculates the sum of the Euclidean distances between the corresponding joint point velocity vectors between the two frames.

[0090] Acceleration difference (d_accel): Similarly, the system calculates the acceleration vector for each joint by taking the first-order difference of the velocity vector. The sum of the Euclidean distances between the corresponding joint acceleration vectors between the two frames is then calculated.

[0091] Next, in step S12, the system processes the difference values ​​of the three dimensions calculated in the previous step to obtain a comprehensive kinematic local cost. This local cost not only reflects the difference in static posture, but also reflects the difference in movement speed and force pattern.

[0092] In one embodiment, step S12 can be implemented by following the steps S121 to S125:

[0093] First, the system enters a biomechanical model-based weight determination phase. In step S121, the system builds a kinematic chain model that characterizes the main force generation methods and mechanical transmission paths of the standard movement currently being compared.

[0094] After establishing the model, in step S122, the system will quantify the contribution of different joints to the completion of the entire movement based on this kinematic chain model. A specific quantification method is to analyze the peak angular velocity or peak acceleration of each joint in the standard movement template during the entire movement process. Generally, the joints that play a key role in the mechanical transmission chain or have the fastest speed at the moment of force application are considered to have the highest contribution.

[0095] Then, in step S123, the system determines the weight coefficient w_j for each joint point used to calculate the local cost based on the contribution quantified in the previous step. For example, the peak angular velocity of each joint point can be normalized, and the resulting normalized value is the weight coefficient of the joint point. In this way, the higher the contribution of the joint point, the larger the weight coefficient is assigned.

[0096] After determining the weight of each joint point, the system begins weighted fusion. In step S124, the system will perform a weighted summation of the difference values ​​of the position, velocity, and acceleration calculated in step S11 for each joint point with a set of preset inter-dimensional weights (i.e., α, β, γ), thereby obtaining a kinematic comprehensive difference value that can reflect the comprehensive performance of the single joint point. The calculation formula of the kinematic comprehensive difference value d_kinematic is as follows:

[0097] d_kinematic=α·d_pos+β·d_vel+γ·d_accel, where α, β, and γ are preset weight coefficients whose sum is 1.

[0098] Finally, in step S125, the system multiplies the kinematic comprehensive difference values ​​of all relevant nodes obtained in the previous step by their respective weight coefficients w_j obtained in the weight determination phase for achieving weighted weighting between parts. The system then sums all the products to obtain the final kinematic local cost that can fully and accurately reflect the comprehensive difference between the two-frame postures. The calculation formula for this kinematic local cost is as follows:

[0099]

[0100] Among them, the part in the brackets is the process of calculating the kinematic comprehensive difference value of the j-th joint point, and the outer layer, multiplied by w_j and then summed up, is the process of finally obtaining the local cost.

[0101] It can be understood that by establishing a biomechanical kinematic chain model and using it to quantitatively determine the importance of each joint in the evaluation, the cost function calculation is no longer "uniform," but rather possesses expert-level analytical capabilities that capture the core essence of the movement. Furthermore, the calculation of local costs is deconstructed into a "two-layer" weighted summation process that first weights each joint "inter-dimensionally" and then weights all related nodes "contribution-wise." This significantly improves the accuracy and interpretability of the evaluation model, ensuring that the final evaluation results are not only accurate, but also that the underlying reasons can be clearly traced and analyzed.

[0102] Then, in steps S23 to S25, the system will obtain the final real-time action score based on this kinematic local cost through a complete dynamic time warping calculation process. The process specifically includes:

[0103] In step S23, a cost matrix is ​​constructed using the kinematic local costs as matrix elements. Specifically, the system initializes a matrix of size N x M, where N is the length of the user's skeletal point sequence and M is the length of the standard motion template. The system then populates the cell in row i and column j of the cost matrix with the local costs calculated between frame i of the user sequence and frame j of the template sequence.

[0104] In step S24, based on the cost matrix, the recursive formula is derived through dynamic programming:

[0105] γ(i,j)=D(i,j)+min{γ(i-1,j),γ(i-1,j-1),γ(i,j-1)}, construct the cumulative cost matrix.

[0106] Finally, in step S25, once the entire cumulative cost matrix is ​​constructed, the value γ(N,M) at the matrix endpoint (N,M) is the minimum cumulative total cost between the two complete sequences. The system processes this total cost value through a preset normalization function to obtain the final posture evaluation score. For example, the normalization function can be: posture evaluation score = 100 * max(0,1-(minimum cumulative cost / preset maximum possible cost)), thereby obtaining the action evaluation score.

[0107] For example, suppose there is an action template of a slight "head scratching" that represents the emotion of "confusion" in the standard action template library. A user unconsciously made a similar action while talking to an intelligent assistant. After the system captures the action sequence, it will compare it with the "head scratching" template. Because the action is slow and light in force, the system may give a higher weight to the position difference value when calculating the local kinematic cost. Through calculation, the system may eventually come up with a higher action evaluation score, such as 85 points. This score indicates that the user's body language is morphologically highly correlated with "confusion", and it will be used together with subsequent facial, physiological and other modal information to comprehensively judge the user's current emotional state.

[0108] As you can see, the kinematic dynamic time warping algorithm, which integrates position, velocity, and acceleration differences during motion comparison, provides a comprehensive and in-depth quantitative assessment of a user's body language. The results reflect not only the similarity of postures but also the dynamic characteristics of movements, more accurately than traditional methods. This provides a crucial and reliable feature input derived from user body language for subsequent multimodal fusion emotion assessment models.

[0109] S20 . Based on the action evaluation score, determine the user emotion level corresponding to the action evaluation score through a preset score-emotion mapping rule.

[0110] Specifically, the score-emotion mapping rules are a set of pre-defined logical conditions used to convert a quantitative action evaluation score into a discrete or continuous user emotion level that represents an emotional state. The user emotion level is an identifier that represents the user's current emotional state, such as discrete categories such as "positive," "negative," and "neutral."

[0111] In step S20, based on the action evaluation score generated in S10, this method uses a preset score-emotion mapping rule to determine the user's emotional level corresponding to that score. Its core purpose is to preliminarily and inferentially map an objective, quantified action completion score to a classification of emotional states. For example, if a user's imitation of a standard action template associated with "joy" scores highly, the system can reasonably infer that the user's current emotional level is biased towards "positive." This preliminary emotional level serves as the core basis for the subsequent generation of emotional text and voice parameters.

[0112] In a specific embodiment, the process of determining the user's emotion level is implemented through the following steps S21 to S23:

[0113] First, in step S21, the system accesses a predefined mapping table, which is a data structure that stores a preset correspondence between "rating intervals" and "emotional level identifiers".

[0114] Next, in step S22 , the system compares the action evaluation score generated in S10 with all the scoring intervals defined in the mapping table to find and determine which interval the score falls within.

[0115] Finally, in step S23, the system outputs the emotion level identifier corresponding to the scoring interval to which the score belongs, and uses it as the user emotion level for use in subsequent steps.

[0116] For example, a score-sentiment mapping table can be pre-configured as follows:

[0117] Score range [85,100] > Corresponding emotion level label: "positive"

[0118] Score range [60,84] > Corresponding emotion level label: "neutral"

[0119] Score range [0,59] > Corresponding emotion level label: "negative"

[0120] Suppose, in step S10, the system calculates the user's current action evaluation score as 92. In step S22, the system compares this value with the interval in the mapping table and determines that it falls within the interval [85, 100]. Therefore, in step S23, the system outputs the user's emotion level as "positive." If another user scores 45, the system will determine their emotion level as "negative."

[0121] As you can understand, a predefined mapping table converts continuous, quantified action assessment scores into discrete, emotionally explicit emotion-level labels. This provides a clear, simple, and standardized input signal for subsequent modules that require differentiated processing based on different emotional states, such as language generation models and speech synthesis engines. This conversion from numerical values ​​to labels greatly simplifies the logical complexity of subsequent processing and enhances the stability and interpretability of the entire emotional feedback system.

[0122] S30: Input the user emotion level and the action evaluation score into a pre-trained language generation model to generate an emotional feedback text that matches the user emotion level.

[0123] Specifically, in step S30, this method performs the "content generation" phase, leveraging a pre-trained generative language model to dynamically create feedback text tailored to the user's current state. This model receives the quantified user emotion rating and action evaluation scores from the upstream evaluation module as input. Based on these inputs, it aims to automatically "create" a humanized text that is highly contextually appropriate in both content and tone.

[0124] In a specific embodiment, the process of generating the emotional feedback text is implemented by the following steps S31 to S33:

[0125] First, in step S31, the system needs to construct a structured input prompt from the user's emotion level determined in S20 and the action evaluation score obtained in S10. The design of the input prompt is crucial. It needs to convert the numerical information of multiple dimensions into a task instruction that can be clearly understood by the language model. For example, this prompt can be a string in JSON format: {"emotion":"positive","score":92,"task":"Generate a sentence of praise and encouragement"}, or a more direct natural language instruction: "Generate a sentence of feedback for a learner with a positive emotion and an action score of 92."

[0126] After constructing the input prompt, in step S32, the system submits the prompt to a pre-trained, Transformer-based language generation model. This model, such as a fine-tuned GPT or BERT model, leverages its powerful contextual understanding to parse the sentiment, score, and task instructions contained in the prompt. Based on this understanding, the model leverages the linguistic patterns and world knowledge learned from massive amounts of text data to begin generating a text sequence, word by word (or token), that best fits the input context in terms of logic, grammar, and emotional tone.

[0127] Finally, in step S33, the system receives and decodes the text sequence output by the language generation model and converts it into a human-readable, final emotional feedback text.

[0128] For example, if the input prompt includes "Emotion: Negative" and "Score: 55," the language model might generate a piece of text designed to provide gentle encouragement and specific guidance, such as: "It's okay, your performance is above the passing line this time. We see your hard work. Please note that if you can relax your movements a little, I believe your score will be higher!" This text "created" by the model in real time is clearly more humane and instructive than a fixed template of "Low score, please keep working hard."

[0129] As can be understood, by constructing the evaluation results of multiple dimensions into a structured input prompt, it provides an information-rich and well-defined context for the subsequent language generation model, ensuring the pertinence and accuracy of the generated content. Furthermore, by utilizing a large language model based on the Transformer architecture to generate the final feedback text, this method can automatically and intelligently convert a set of quantitative evaluation data into natural language that aligns with human communication habits and whose content and tone are highly aligned with the context, greatly enhancing the intelligence and emotional resonance of the voice feedback.

[0130] S40 : Based on the user's emotion level, query a preset emotion-speech parameter mapping library to determine speech synthesis parameters including intonation, speech speed, and tone.

[0131] Specifically, the emotion-speech parameter mapping library is a data repository that stores preset correspondences between different emotion levels and a set of specific speech synthesis parameters. It can be implemented as a database, configuration file, or lookup table. These speech synthesis parameters are a set of quantitative parameters that can be directly read and applied by the speech synthesis engine to control the acoustic characteristics of the generated speech. These parameters include at least intonation (which defines the pitch of the voice), speech rate (which defines the speed of speech), and tone (which defines the emotional overtones).

[0132] In step S40, the method executes the "emotion generation" phase. The core purpose of this step is to convert the relatively abstract "user emotion level" determined in S20 into a set of concrete acoustic parameters that can be executed by the speech synthesis engine. This step is responsible for defining the vocal qualities of the resulting speech, ensuring that the emotional tone of the speech is consistent with the previously determined user emotional state.

[0133] In a specific embodiment, the above-mentioned process of determining speech synthesis parameters is implemented through a "library search" process. The system uses the user's emotional level (for example, "positive") determined in S20 as a query index to search in a preset "emotion-speech parameter mapping library" that stores the mapping relationship between emotions and parameters. The library pre-configures a set of optimal speech synthesis parameters for each possible emotional level. After the system finds an entry that matches the current emotional level, it extracts the parameter group corresponding to the entry, including information such as intonation, speaking speed and tone, as the final output of this step.

[0134] For example, the emotion-speech parameter mapping library may preset the following correspondence:

[0135] Emotion level: "Positive" > Speech synthesis parameters: {intonation: 'high', speaking speed: '1.2x', tone: 'joyful'}

[0136] Emotion level: "Neutral" > Speech synthesis parameters: {intonation: 'smooth', speech rate: '1.0x', tone: 'statement'}

[0137] Emotion level: "Negative" > Speech synthesis parameters: {Intonation: 'Low', Speech speed: '0.8x', Tone: 'Soothing'}

[0138] Assume that in step S20, the system determines that the user's current emotional level is "positive." In step S40, the system searches the mapping library using "positive" as the keyword and determines the speech synthesis parameters to be used: a high pitch, a 1.2x faster speaking rate, and a pleasant tone. This set of parameters is then fed into the subsequent speech synthesis step to control the acoustic performance of the final generated speech.

[0139] As you can see, a preset emotion-speech parameter mapping library explicitly associates discrete emotion level identifiers with a specific set of speech synthesis parameters, providing the subsequent speech synthesis engine with clear, quantified instructions for controlling the emotional tone of generated speech. This ensures that the final synthesized speech maintains a high degree of consistency with the previously determined user emotional state in core acoustic dimensions such as intonation, speech rate, and tone, making the emotional feedback more authentic and trustworthy.

[0140] S50: Based on the pre-collected user voice sample, train and generate a personalized voice cloning model that matches the user's voice characteristics.

[0141] Specifically, a user voice sample refers to a pre-collected raw audio data containing the user's actual speaking voice. A voice cloning model is a specially trained or adapted speech synthesis model that can generate voice that is highly similar to a specific user in timbre, prosody, and accent.

[0142] In step S50, this method performs a key personalized model training step. The core purpose of this step is to train and generate a personalized voice clone model unique to the user, capable of mimicking their timbre characteristics, by analyzing the user's pre-provided voice samples. This step ensures that subsequent voice feedback generated by the system is no longer a cold, generic system voice, but rather a voice familiar to the user, creating a unique, self-conversational interactive experience.

[0143] In a specific embodiment, the process of training and generating the personalized timbre cloning model is implemented by the following steps S51 and S52:

[0144] First, in step S51, the system processes the voice sample provided by the user to extract a timbre feature vector that can represent its timbre characteristics. A timbre feature vector is a high-dimensional mathematical vector that is used to digitally represent the unique characteristics of a specific speaker's voice. This process is usually completed by a deep learning model called a "speaker encoder". The encoder can map a raw audio waveform into a high-dimensional embedding vector that condenses the speaker's identity information. This vector is the timbre feature vector.

[0145] After obtaining this feature vector representing the user's "voice fingerprint," in step S52, the system uses this vector to fine-tune a pre-trained, powerful universal speech synthesis model. This universal model (for example, a VITS or Tacotron 2 model) has been trained using a large amount of speech data from different speakers and has the ability to generate high-quality, natural and fluent speech, but its timbre is universal. The fine-tuning process is to inject the timbre feature vector into this universal model as a "condition" or "style guide" and retrain some of its parameters with a small number of iterations. This process allows the universal model to maintain its powerful synthesis capabilities while its output timbre will "approach" and "adapt" to the user's own timbre characteristics. The new model obtained after fine-tuning is the final, personalized timbre cloning model that is exclusive to the user.

[0146] For example, when a user first uses this method, the system instructs them to read a few specified sentences to collect a short speech sample. The system then extracts a 256-dimensional timbre feature vector from this sample in S51. In S52, the system loads a general speech synthesis model and fine-tunes the general model using this 256-dimensional vector as an additional input. Once fine-tuning is complete, the resulting personalized timbre clone model is saved and bound to the user's ID for subsequent use in S60.

[0147] It can be understood that by extracting a unique timbre feature vector from a user's voice sample and using it to fine-tune a general speech synthesis model, a speech synthesis model with a timbre highly similar to the user's own and highly personalized characteristics can be generated. Furthermore, because this is based on fine-tuning a powerful pre-trained model rather than training from scratch, the timbre cloning process can be completed in a short time with only a small number of user voice samples, greatly improving the efficiency and practicality of model training and lowering the barrier to personalized voice feedback for users.

[0148] S60: Perform speech synthesis processing based on the emotional feedback text, the speech synthesis parameters, and the personalized timbre cloning model to generate enhanced speech feedback with the user's timbre and emotional expression matching the user's emotional level.

[0149] Specifically, the speech synthesis engine is the core computing component of the speech synthesis model responsible for converting text information into acoustic features and ultimately generating an audio waveform. The waveform sound file is the resulting digital audio file that can be directly played by audio playback devices. Common formats include WAV or MP3.

[0150] In step S60, the system integrates the core elements generated by all previous steps—the emotionally charged feedback text, specific speech synthesis parameters, and the personalized voice cloning model—to perform speech synthesis processing. The core purpose of this step is to transform a complete feedback instruction encompassing the three dimensions of "content," "emotion," and "identity" into a realistic, audible, enhanced voice feedback with the user's own voice and accurate emotional expression through a technical process.

[0151] In a specific embodiment, the above-mentioned process of performing speech synthesis processing is implemented through the following steps S61 to S63:

[0152] First, in step S61, the system initializes the personalized voice cloning model trained and generated for the current user in step S50. This process includes loading the model from storage into memory and preparing its internal speech synthesis engine to be in a standby state.

[0153] Next, in step S62, the system applies the set of speech synthesis parameters (including intonation, speech rate, and tone) determined in step S40 to the speech synthesis engine of the initialized model. These parameters act like "style instructions" for a broadcaster, adjusting the specific acoustic performance of the synthesis engine when it subsequently generates speech. For example, a "high" intonation parameter will increase the fundamental frequency of the synthesized speech, while a "1.2x" speech rate parameter will increase the speaking speed.

[0154] After the model is correctly configured, in step S63, the system drives the parameterized speech synthesis engine to process the emotional feedback text generated in S30. The synthesis engine converts the text content word by word into acoustic features and renders these acoustic features according to the emotional parameters configured in S62. Ultimately, it generates and outputs a digitized, playable waveform sound file. This audio file is the final enhanced speech feedback.

[0155] For example, continuing with the above example, the system has loaded a personalized voice cloning model for the user in S50. In S40, the system determines that the current voice parameters to be adopted are {intonation: 'high', speaking speed: '1.2 times', tone: 'joyful'}. In S30, the system generates the text: "Great! Your movements are very standard!". In step S61, the model is initialized. In S62, this set of "positive" parameters is configured to the synthesis engine. In S63, the engine receives the text "Great!..." and, based on the configuration, synthesizes this sentence into a WAV audio file and plays it in a high-pitched, slightly fast, joyful voice with the user's own timbre.

[0156] As you can see, the final speech synthesis is performed through a clear process that includes initialization, parameter configuration, and engine driving. This ensures that all upstream analyzed and generated parameters (text content, emotional style, and personal timbre) are accurately and organically integrated into the final audio output. Furthermore, because the final output is a standard waveform sound file, the enhanced speech feedback generated by this method has excellent compatibility and can be easily integrated and played on any device or platform that supports audio playback.

[0157] Based on the above embodiments, the method for enhancing speech feedback based on dynamic time warping of the present application has at least the following beneficial effects:

[0158] 1. Achieve a deep correlation between feedback content and user performance

[0159] This application first accurately evaluates the user's real-time action completion by applying a kinematic dynamic time warping algorithm, and uses this evaluation result as one of the core inputs driving the language generation model. This ensures that the final generated voice feedback content is no longer a general phrase unrelated to the user's performance, but is instead highly relevant and intelligent feedback that can be dynamically "created" based on the user's specific performance (high or low score).

[0160] 2. Achieve a high degree of consistency between feedback emotion and user status

[0161] This application not only generates content, but also finely controls the "emotional expression" of feedback. By mapping the action evaluation score to a clear user emotion level, and further mapping this level to specific speech synthesis parameters (such as intonation, speaking speed, and tone), this method ensures that the acoustic characteristics of the speech output are consistent with the emotional tone of the text content. For example, encouraging text will be spoken in a gentle and soothing tone, while praising text will be spoken in a cheerful and upbeat tone, achieving a unity of content and form.

[0162] 3. Provides the ultimate personalized and immersive listening experience

[0163] This application innovatively introduces user voice cloning technology, fine-tuning a powerful general-purpose speech synthesis model to generate voices that closely resemble the user's voice. This feedback mechanism, which "uses my own voice to speak to me in a way that best understands my performance and emotions," completely eliminates the mechanical and distant feel of traditional voice systems, creating an unprecedented, highly engaging, and personalized immersive experience.

[0164] The embodiment of the present invention further provides an enhanced speech feedback device based on dynamic time warping, wherein the enhanced speech feedback device based on dynamic time warping includes:

[0165] An action evaluation module is used to obtain a user skeleton point sequence corresponding to a user action and compare the user skeleton point sequence with a standard action template using a kinematic dynamic time warping algorithm to generate an action evaluation score representing the degree of action completion;

[0166] An emotion determination module is used to determine the user emotion level corresponding to the action evaluation score based on the action evaluation score and using a preset score-emotion mapping rule;

[0167] A text generation module, configured to input the user's emotion level and the action evaluation score into a pre-trained language generation model to generate an emotional feedback text that matches the user's emotion level;

[0168] A parameter mapping module is used to query a preset emotion-speech parameter mapping library based on the user's emotion level to determine speech synthesis parameters including intonation, speaking speed and tone;

[0169] The voice modeling module is used to train and generate a personalized voice cloning model that matches the user's voice characteristics based on pre-collected user voice samples;

[0170] The speech synthesis module is used to perform speech synthesis processing based on the emotional feedback text, the speech synthesis parameters, and the personalized timbre cloning model to generate enhanced speech feedback with the user's own timbre and emotional expression matching the user's emotional level.

[0171] Among them, the steps for implementing each functional module of the enhanced speech feedback device based on dynamic time warping can refer to the various embodiments of the enhanced speech feedback method based on dynamic time warping of the present invention, and will not be repeated here.

[0172] In addition, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium can be any one of a hard disk, a multimedia card, an SD card, a flash memory card, an SMC, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination thereof. The computer-readable storage medium includes an enhanced speech feedback program 10 based on dynamic time warping. The specific implementation of the computer-readable storage medium of the present invention is substantially the same as the specific implementation of the enhanced speech feedback method based on dynamic time warping and the server 1 described above, and will not be repeated here.

[0173] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0174] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0175] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0177] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0178] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A method for enhancing speech feedback based on dynamic time warping, characterized in that: include: Obtaining a user skeleton point sequence corresponding to the user action, and applying a kinematic dynamic time warping algorithm to compare the user skeleton point sequence with a standard action template to generate an action evaluation score representing the degree of action completion; Based on the action evaluation score, determining the user emotion level corresponding to the action evaluation score through a preset score-emotion mapping rule; Inputting the user's emotion level and the action evaluation score into a pre-trained language generation model to generate emotional feedback text that matches the user's emotion level; Based on the user's emotion level, querying a preset emotion-speech parameter mapping library to determine speech synthesis parameters including intonation, speaking speed and tone; Based on pre-collected user voice samples, a personalized voice cloning model is trained and generated to match the user's voice characteristics. Based on the emotional feedback text, the speech synthesis parameters, and the personalized timbre cloning model, speech synthesis processing is performed to generate enhanced speech feedback with the user's timbre and emotional expression matching the user's emotional level.

2. The method for enhancing speech feedback based on dynamic time warping according to claim 1, wherein: Applying a kinematic dynamic time warping algorithm to generate the motion assessment score includes: Calculate the position difference, velocity difference, and acceleration difference between each pair of frames in the user's skeleton point sequence and the standard action template; Based on the position difference value, the velocity difference value, and the acceleration difference value, a kinematic local cost for dynamic time warping algorithm comparison is obtained; Constructing a cost matrix using the kinematic local costs as matrix elements; Based on the cost matrix, constructing a cumulative cost matrix through a dynamic programming recursive formula; The final value of the cumulative cost matrix is ​​normalized to obtain the action evaluation score.

3. The method for enhancing speech feedback based on dynamic time warping according to claim 2, wherein: Based on the position difference value, velocity difference value, and acceleration difference value, a kinematic local cost for dynamic time warping algorithm comparison is obtained, including: Establishing a kinematic chain model that characterizes the force generation mode of the standard movement; According to the kinematic chain model, quantify the contribution of different joints to the completion of the movement; Determining a weight coefficient of each joint point for calculating the local cost based on the contribution; Performing inter-dimensional weighted summation on the position difference value, velocity difference value, and acceleration difference value to obtain a kinematic comprehensive difference value for each joint point; Based on the weight coefficient, the kinematic comprehensive difference values ​​of all joint points are summed to obtain the local cost.

4. The method for enhancing speech feedback based on dynamic time warping according to claim 1, wherein: Based on the action evaluation score, determining the user emotion level corresponding to the action evaluation score through a preset score-emotion mapping rule includes: Access a predefined mapping table that stores the correspondence between rating intervals and emotion level identifiers; Comparing the action evaluation score with a plurality of scoring intervals in the mapping table to find and determine the scoring interval to which the action evaluation score belongs; The emotion level identifier corresponding to the interval is output as the user emotion level.

5. The method for enhancing speech feedback based on dynamic time warping according to claim 1, wherein: Inputting the user's emotion level and the action evaluation score into a pre-trained language generation model to generate emotional feedback text that matches the user's emotion level, including: constructing the user emotion level and the action evaluation score into a structured input prompt; Submitting the structured input prompt to a language generation model based on the Transformer architecture; A text sequence output by the language generation model is received and decoded as the emotional feedback text.

6. The method for enhancing speech feedback based on dynamic time warping according to claim 1, wherein: Based on pre-collected user voice samples, a personalized voice cloning model is trained and generated to match the user's voice characteristics, including: Extracting a timbre feature vector representing the user's timbre characteristics from the user's voice sample; A pre-trained general speech synthesis model is fine-tuned using the timbre feature vector to generate the personalized timbre cloning model.

7. The method for enhancing speech feedback based on dynamic time warping according to claim 1, wherein: Perform speech synthesis processing, including: Initializing the personalized timbre cloning model; Configuring the speech synthesis parameters to the speech synthesis engine of the model; The speech synthesis engine is driven to process the emotional feedback text to generate and output the enhanced speech feedback in the form of a waveform sound file.

8. An enhanced speech feedback device based on dynamic time warping, characterized in that: include: An action evaluation module is used to obtain a user skeleton point sequence corresponding to a user action and compare the user skeleton point sequence with a standard action template using a kinematic dynamic time warping algorithm to generate an action evaluation score representing the degree of action completion; An emotion determination module is used to determine the user emotion level corresponding to the action evaluation score based on the action evaluation score and using a preset score-emotion mapping rule; A text generation module, configured to input the user's emotion level and the action evaluation score into a pre-trained language generation model to generate an emotional feedback text that matches the user's emotion level; A parameter mapping module is used to query a preset emotion-speech parameter mapping library based on the user's emotion level to determine speech synthesis parameters including intonation, speaking speed and tone; The voice modeling module is used to train and generate a personalized voice cloning model that matches the user's voice characteristics based on pre-collected user voice samples; The speech synthesis module is used to perform speech synthesis processing based on the emotional feedback text, the speech synthesis parameters, and the personalized timbre cloning model to generate enhanced speech feedback with the user's own timbre and emotional expression matching the user's emotional level.

9. An enhanced speech feedback device based on dynamic time warping, characterized in that: It includes a memory, a processor, and an enhanced speech feedback program based on dynamic time regularization stored in the memory and runnable on the processor. When the processor executes the enhanced speech feedback program based on dynamic time regularization, it implements the enhanced speech feedback method based on dynamic time regularization as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an enhanced speech feedback program based on dynamic time regularization, and when the enhanced speech feedback program based on dynamic time regularization is executed by the processor, the enhanced speech feedback method based on dynamic time regularization as described in any one of claims 1-7 is implemented.

Citation Information

Cited By

  • Pet accompanying voice generation method and device based on voiceprint modeling, equipment and storage medium

    CN122337180A