A method and system for real-time affective speech conversion

By collecting and processing user voice data in real time, and combining timbre file matching and baud rate adjustment, the problem of insufficient voice conversion quality in traditional voice conversion methods when the network is unstable is solved, and high-quality, personalized voice conversion effects are achieved.

CN116453529BActive Publication Date: 2026-02-06SHANGHAI GEZI INTERACTIVE INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310538032.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-12
Publication Date
2026-02-06
Estimated Expiration
2043-05-12

AI Technical Summary

Technical Problem

Traditional speech conversion methods cannot adapt to various speech conversion quality requirements when the network transmission quality is unstable, resulting in insufficient naturalness of speech conversion and monotonous tone, leading to a poor user experience.

Method used

By collecting user voice data in real time, matching user IDs with timbre files, selecting sampling domains and baud rates for synchronous switching, and combining background noise separation methods, the baud rate is adaptively adjusted to enhance timbre quality, and timbre parameters are customized according to user needs.

Benefits of technology

It improves the naturalness and user experience of speech conversion, meets the speech conversion quality requirements in multiple scenarios, enhances speech recognition accuracy and personalized tone selection, and overcomes the mechanical tone problem caused by differences in speech data quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116453529B_ABST
    Figure CN116453529B_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to the technical field of audio processing, in particular to a method and system for real-time emotional speech conversion. The application specifically comprises the following steps: step one, collecting user-entered speech data in real time; step two, transmitting the user-entered speech data to a model file for preprocessing; and step three, performing audio output after the preprocessing is completed. The real-time emotional speech conversion method matches the user's timbre file with the model file for preprocessing, different model files correspond to different timbre data to be matched, so as to help the user freely select the timbre and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the present application relates to the technical field of audio processing, in particular to a method and system for real-time emotional speech conversion. BACKGROUND

[0002] In the traditional speech conversion method, the user input speech data is usually collected, and the collected speech data is converted into binary data, and then a network request based on data transmission is established. After that, the user speech data on the server is converted and fed back to the client output. However, the speech conversion quality of this speech conversion method depends on the quality of network transmission. Under a single network transmission modulation parameter, it cannot adapt to the transmission requirements of various speech conversion qualities. Therefore, due to the different quality of speech data input by different users, the quality of speech data transmission is different, which ultimately leads to insufficient naturalness of speech conversion and the problem of single emotional experience of output speech tone.

[0003] Chinese Patent No. CN113689867B provides a preprocessing method and device for speech conversion model, electronic equipment and medium. This patent extracts hidden features in original acoustics to further improve the matching degree between original acoustics and predicted acoustics. Chinese Patent No. CN112116904B provides a speech conversion method, device, equipment and storage medium. In this patent, the original speech can be converted in terms of both speech and language. However, the above-mentioned patents do not explicitly explain how to further enhance the acoustic quality of poor quality information in the original acoustics after matching or speech conversion.

[0004] Therefore, in view of the problems existing in the prior art speech conversion technology, the present application provides a method and system for real-time emotional speech conversion SUMMARY

[0005] In view of the above problems, the first aspect of the present application provides a method for real-time emotional speech conversion, which specifically includes the following steps: Step 1, real-time collection of user input speech data; Step 2, transmission of user input speech data to a model file for preprocessing; Step 3, audio output after preprocessing is completed.

[0006] Preferably, in the step of transmitting the user input speech data to the model file for preprocessing, the user input speech data is numbered according to the user number, and the timbre file is issued according to the user number.

[0007] Preferably, the model file is checked for existence. If yes, the timbre file is transmitted to the model file for preprocessing. If no, a model file import error is fed back.

[0008] Preferably, in the preprocessing of transmitting the timbre file into the model file, the sampling domain is selected according to the timbre quality.

[0009] Preferably, according to the selection of the sampling domain, the synchronous switching of the data transmission baud rate is performed, and the synchronous switching of the baud rate is performed to switch the timbre quality.

[0010] Preferably, in the real-time emotional speech conversion, the user switches the real-time timbre quality by switching different baud rate values, and the switching rate is between 40ms and 60ms.

[0011] Preferably, in the real-time timbre quality switching, a background noise separation method is established to extract the user timbre file, the sound quality in the timbre file is judged, the baud rate value is adaptively adjusted for the timbre file not meeting the judgment standard, and the timbre quality is enhanced.

[0012] The second aspect of the present application provides a system for real-time emotional speech conversion, specifically including a resource module, a preprocessing module and a conversion module.

[0013] Preferably, the preprocessing module includes a sound library, and the sound library stores model files to be converted by the user.

[0014] Preferably, the conversion module adjusts the timbre parameters according to the user's demand, and customizes the timbre to be converted.

[0015] Compared with the prior art, the present application has the following advantages:

[0016] (1) The real-time emotional speech conversion method of the present application pre-processes the user timbre file by matching the model file, different model files correspond to different timbre data to be matched, which helps the user to freely select the timbre and improves the user experience.

[0017] (2) On the basis of (1), the present application switches the data transmission baud rate by selecting the sampling domain, thereby switching the timbre quality. For different user speech conversion quality requirements and different recording scenes, the speech conversion quality is dynamically optimized, thereby further improving the naturalness of user speech conversion.

[0018] (3) On the basis of (2), the present application extracts the user timbre file by establishing a background noise separation method, so as to meet the speech conversion and output quality in multiple scenes and improve the user speech recognition accuracy.

[0019] (4) Based on (3), the timbre file that does not meet the judgment standard is adaptively adjusted in the value of baud rate, and the timbre quality is enhanced, so as to further overcome the problems that the voice data transmission quality is different due to the different voice quality of the specified frequency band in the voice data input by different users, and the naturalness of voice conversion is insufficient, and the output voice tone is single and the emotional experience is poor.

[0020] (5) Based on (4), a system for real-time emotional voice conversion is established in the application, and the setting of the voice to be converted can be customized according to the user's demand, so as to meet the user's individual needs and improve the application range of the voice conversion system. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 A flowchart of a method for real-time emotional voice conversion. DETAILED DESCRIPTION

[0022] Embodiment:

[0023] The first aspect of the embodiment provides a method for real-time emotional voice conversion, as shown in Figure 1 Specifically, it includes:

[0024] Step one, real-time acquisition of user input voice data;

[0025] Step two, the user input voice data is transmitted to the model file for preprocessing; wherein the specific process of the preprocessing is:

[0026] S1, the user input voice data is numbered according to the user number, and the timbre file is issued according to the user number;

[0027] S2, check whether the model file exists, if yes, transmit the timbre file to the model file for preprocessing; if not, feedback model file import error;

[0028] S3, in the step of transmitting the timbre file to the model file for preprocessing, the sampling domain is selected according to the timbre quality;

[0029] S5, in the step of transmitting the timbre file to the model file for preprocessing, the sampling domain is selected according to the timbre quality;

[0030] S6, according to the selection of the sampling domain, the data transmission baud rate is synchronously switched, and the timbre quality is switched according to the synchronous switching of the baud rate;

[0031] S7, in the real-time emotional voice conversion, the user switches the real-time timbre quality by switching different baud rate values, and the switching rate is 50ms;

[0032] S8, in the real-time switching of the different timbre qualities, a background noise separation method is used to extract a user timbre file, the sound quality in the timbre file is judged, the baud rate value of the timbre file not meeting the judgment standard is adaptively adjusted, and the timbre quality is enhanced;

[0033] Further, in the background noise separation method, the collected user input voice data is converted into voice features, and the voice features are screened, in the voice feature screening, the amplitude data and phase data of the voice frequency domain information in the voice features are extracted, and according to the distribution characteristics of the amplitude data and the phase data, the voice features and the noise are distinguished and screened, at the same time, the distribution coefficient of the amplitude data and the phase data is established according to the distribution degree, the voice features are further amplified according to the distribution coefficient, and the noise is further reduced, so as to improve the clarity of the background noise separation.

[0034] Further, in the judgment of the sound quality in the timbre file, the specific judgment method is as follows: for the user input timbre file, the timbre file is divided into multiple timbre frequency bands, the sound quality of each timbre frequency band is judged, and each timbre frequency band in the timbre file is divided into a corresponding sampling domain according to the sound quality judgment result. Different sampling domains correspond to different baud rates in the process of sound signal propagation, and different baud rates determine the propagation quality of sound signals. Through real-time dynamic adjustment of the baud rate of the timbre frequency band under different timbre qualities in the sound propagation process, the converted timbre quality is kept stable output.

[0035] Further, the sound quality judgment of the present application is applied to the method of real-time emotional voice conversion, in order to overcome the fluctuation of the sound in each timbre frequency band of the timbre file due to the change of emotion when the user performs voice conversion, and the sound quality in different frequency bands changes nonlinearly due to the fluctuation, so that the preprocessed timbre data output finally tends to be mechanical sound, lacking of continuity and emotional expression.

[0036] Step three, after the preprocessing is completed, the audio output is performed.

[0037] The second aspect of the embodiment provides a system for real-time emotional voice conversion, specifically including a resource module, a preprocessing module and a conversion module, wherein:

[0038] The resource module is used for loading model files and timbre files; specifically, the model files are introduced as follows:

[0039] std::map<int,float*>speakerBins;

[0040] The preprocessing module is used for selecting a sampling domain and a corresponding baud rate, selecting a sound color parameter to be converted, and initializing a sound color conversion engine, specifically, initializing the sound color conversion engine and loading a model file.

[0041]

[0042]

[0043]

[0044] The conversion module pre-processes a sound color file in the sound color conversion engine through a corresponding model file, and outputs pre-processed data; wherein the preprocessing module comprises a sound library, and the sound library stores a model file to be converted by a user; the conversion module adjusts a sound color parameter according to a user demand, and customizes a sound color to be converted.

[0045] According to the technical scheme of the present application, through the selection of the sampling domain, the synchronous switching of the data transmission baud rate is performed, the sound quality is switched according to the synchronous switching of the baud rate, when the user inputs the voice data, the different sound data transmission baud rate is adjusted according to the emotional change in the user voice data, so that the real-time switching of different voice qualities is performed, thereby improving the stability of the sound quality of the output after the pre-processing, and improving the authenticity of the output sound color data.

Claims

1. A method for real-time emotional speech conversion, characterized in that, Specifically, it includes: Step 1: Collect user-inputted voice data in real time; Step 2: Transmit the user-entered voice data to the model file for preprocessing; Step 3: Output the audio after preprocessing is complete; In step two, the user-inputted voice data is transmitted to the model file for preprocessing. The user-inputted voice data is assigned a user number, and the next voice tone file is generated based on the user number. Verify if the model file exists. If it does, transfer the timbre file to the model file for preprocessing; otherwise, report a model file import error. In the process of transmitting the timbre file to the model file for preprocessing, the sampling domain is selected based on the timbre quality. Based on the selection of the sampling domain, the data transmission baud rate is switched synchronously, and the tone quality is switched synchronously based on the baud rate switching. In the real-time emotional voice conversion, the user can switch between different timbre qualities in real time by switching different baud rate values, with a switching rate between 40ms and 60ms. In the real-time switching of different timbre qualities, a background noise separation method is established to extract the user's timbre file, the sound quality in the timbre file is judged, the baud rate value is adaptively adjusted for timbre files that do not meet the judgment criteria, and the timbre quality is enhanced.

2. A system for real-time emotion-based speech conversion, characterized in that, The system is used to perform the method of claim 1, and the system specifically includes a resource module, a preprocessing module, and a conversion module.

3. The system for real-time emotional speech conversion according to claim 2, characterized in that, The preprocessing module includes a sound library, which stores model files to be converted by the user.

4. The system for real-time emotional speech conversion according to claim 2, characterized in that, The conversion module adjusts the timbre parameters according to user needs and allows for the customization of the timbre to be converted.

Citation Information

Patent Citations

  • Speech conversion methods, devices, equipment and storage media

    CN112116904B

  • A training method, apparatus, electronic device, and medium for a speech conversion model.

    CN113689867B

  • Voice activated data rate change in simultaneous voice and data transmission

    CN1117228A

  • Chinese speech cloning method for end-to-end tone and emotion migration

    CN115359775A