A speaker-independent pronunciation inverse inference method and system based on variational autoencoder

Through the method based on the variational autoencoder, the content embedding and identity embedding information of speech and pronunciation motion data were extracted separately, and the pronunciation reversal model was constructed, which solved the problems of insufficient speed, flexibility and accuracy in the existing technology, and achieved rapid and high-precision reversal.

CN119741927BActive Publication Date: 2025-07-04INSTITUTE OF LANGUAGES CHINESE ACADEMY OF SOCIAL SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411754283.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-07-04
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

The existing speaker-independent pronunciation reverse method has shortcomings in terms of speed, flexibility and accuracy, and it is difficult to meet actual needs.

Method used

Using a method based on a variational autoencoder, the content embedding and speaker identity embedding information of acoustic features and pronunciation motion characteristics are extracted respectively by training the pronunciation motion characteristics, and a pronunciation reversal model is constructed for mapping and decoding, and pronunciation motion characteristics are generated.

Benefits of technology

It realizes fast, flexible and high-precision pronunciation reverse rumination, which is suitable for improving the robustness of virtual human model drivers and speech recognition systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741927B_ABST
    Figure CN119741927B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of speech recognition technology, and discloses a speaker-independent pronunciation inverse inference method and system based on a variational autoencoder, including: collecting speech acoustic data and pronunciation movement data; training a speech variational autoencoder according to the speech acoustic data, and using the speech variational autoencoder to extract content embedding information based on acoustic features and speaker identity embedding information from the speech acoustic data; training a pronunciation variational autoencoder according to the pronunciation movement data, and using the pronunciation variational autoencoder to extract content embedding information based on pronunciation movement features and speaker identity embedding information from the pronunciation movement data; constructing a pronunciation inverse inference model, and training the pronunciation inverse inference model by using the content embedding information based on acoustic features and speaker identity embedding information and the content embedding information based on pronunciation movement features and speaker identity embedding information; inputting the speech acoustic data into the trained pronunciation inverse inference model, and outputting a pronunciation movement trajectory corresponding to the speech acoustic data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and in particular to a speaker-independent pronunciation inversion method and system based on a variational autoencoder. Background Art

[0002] Pronunciation inversion refers to inferring pronunciation posture or pronunciation movement characteristics from speech signals. Pronunciation inversion can mine the relationship between human pronunciation process and acoustic signals from data, helping to explain the causes of certain speech phenomena; compared with acoustic speech features, pronunciation movement characteristics change slowly and are insensitive to environmental noise, which is beneficial for low-bitrate speech coding, speech parameter trajectory generation and control in speech synthesis, and improving the robustness of speech recognition systems; the inverted pronunciation movement characteristics can be used to drive virtual human models to show the movement process of the pronunciation organs for foreign language learners or hearing-impaired children.

[0003] At present, the mainstream speaker-independent pronunciation inversion methods are mainly divided into model adaptive methods, interpolation methods based on typical data, and direct methods based on aggregated multi-person data. The model-based adaptive method requires a certain amount of data of the target speaker to be obtained in advance, but in actual environments it is difficult for us to obtain the target speaker's data in advance; the typical data interpolation method does not require the target speaker's data to be obtained in advance, but it is necessary to calculate the similarity between the target speaker's speech features and the speech features of a large amount of typical data and assume that the similarity distribution in the pronunciation space is the same as the similarity distribution in the acoustic space. This assumption is unreasonable and the speed of the inference stage cannot meet actual needs; the direct method based on aggregated multi-person data mixes speaker content features and personality features together to jointly model the model, and the model is not flexible enough and the accuracy does not meet the requirements.

[0004] Therefore, how to overcome the deficiencies of the existing technology and realize a speaker-independent pronunciation inversion method with high speed, strong flexibility and high accuracy is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the invention

[0005] To achieve the purpose of the present invention, the present application provides a speaker-independent pronunciation inversion method based on a variational autoencoder, comprising:

[0006] Step S1: collecting speech acoustic data and pronunciation movement data;

[0007] Step S2: training a speech variational autoencoder according to the speech acoustic data, and using the speech variational autoencoder to extract content embedding information and speaker identity embedding information based on acoustic features from the speech acoustic data;

[0008] Step S3: Train a pronunciation variational autoencoder according to the pronunciation motion data, and use the pronunciation variational autoencoder to extract content embedding information and speaker identity embedding information based on pronunciation motion features from the pronunciation motion data;

[0009] Step S4: Construct a pronunciation inverse model, and use the content embedding information and speaker identity embedding information based on acoustic features and the content embedding information and speaker identity embedding information based on pronunciation motion features to train the pronunciation inverse model;

[0010] Step S5: Input the speech acoustic data into the trained pronunciation inverse model, and output the pronunciation motion features corresponding to the speech acoustic data.

[0011] In some specific embodiments, Steps S2 and S3 respectively optimize the parameters of the speech variational autoencoder and the pronunciation variational autoencoder according to the cost function of the standard variational autoencoder:

[0012] L = L rec + L KL

[0013]

[0014] where L, L rec , L KL respectively represent the total cost, the reconstruction error, and the KL divergence; z c , z s respectively represent the outputs of the content feature encoder and the personal feature encoder; N represents the number of speech features or pronunciation features; represents the j-th component of the mean of the normal distribution that z c follows, ∧ ij represents the j-th component of the covariance matrix of this normal distribution; d is the dimension of z c .

[0015] In some specific embodiments, Step S2 includes:

[0016] Step S21: Based on the parameters of the trained speech variational autoencoder, construct a speech feature extractor, a local feature extractor, a content feature encoder, and a personal feature encoder, and use the speech feature extractor to extract acoustic features from the speech acoustic data;

[0017] Step S22: Based on the acoustic features, use the local feature extraction, the content feature encoder, and the personal feature encoder respectively to obtain content embedding information and speaker identity embedding information based on acoustic features.

[0018] In some of the specific embodiments, the pronunciation reverse inference model includes a content reverse inference network and a speaker reverse inference network, and step S4 includes:

[0019] Step S41: Construct the content reverse inference network by using a multi-layer bidirectional long short-term recurrent neural network stacked with two layers of fully connected neural networks to realize the mapping from the content embedding information based on acoustic features to the content embedding information based on pronunciation movement features;

[0020] Step S42: Construct the speaker reverse inference network by using a single-layer linear network to realize the mapping from the speaker identity embedding information based on acoustic features to the speaker identity embedding information based on pronunciation movement features.

[0021] In some of the specific embodiments, step S3 further includes: constructing a pronunciation decoder based on the parameters of the trained voice variational autoencoder.

[0022] In some of the specific embodiments, step S5 includes:

[0023] The trained pronunciation reverse inference model outputs the content embedding information and speaker identity embedding information based on pronunciation movement features according to the speech acoustic data, and decodes the corresponding pronunciation movement features by using the pronunciation decoder.

[0024] In some of the specific embodiments, step S4 further includes:

[0025] Step S43: Train the content reverse inference network and the speaker reverse inference network according to the following formula:

[0026] L = L c + L s + L rec

[0027]

[0028]

[0029] Among them, L, L c , L s and L rec respectively represent the overall cost, content embedding error, speaker embedding error, and pronunciation movement reconstruction error; is the content embedding information of the real pronunciation space corresponding to the i-th sample, is the content embedding information of the pronunciation space estimated by the content reverse inference network; is the speaker embedding information of the real pronunciation space corresponding to the i-th sample, is the speaker embedding information of the pronunciation space estimated by the speaker reverse inference network; is the true articulator position vector corresponding to the i-th sample, is the articulator position vector estimated by using the pronunciation decoder.

[0030] To achieve the same invention purpose, the present application also provides a speaker-independent pronunciation inverse inference system based on a variational autoencoder, including:

[0031] A voice acquisition module, configured to acquire voice acoustic data and pronunciation movement data;;

[0032] A feature extraction module, configured to extract acoustic features in the voice acoustic data;

[0033] An acoustic information extraction module, configured to obtain content embedding information and speaker identity embedding information based on acoustic features from the acoustic features;

[0034] A pronunciation inverse inference module, including a content inverse inference network and a speaker inverse inference network, configured to map the content embedding information and speaker identity embedding information based on acoustic features to content embedding information and speaker identity embedding information based on pronunciation movement features respectively;

[0035] A pronunciation decoding module, configured to decode pronunciation movement data from the content embedding information and speaker identity embedding information based on pronunciation movement features.

[0036] In some specific embodiments, the pronunciation inverse inference module is configured to:

[0037] Construct the content inverse inference network by using a multi-layer bidirectional long short-term recursive neural network stacked with two layers of fully connected neural networks, so as to realize the mapping from the content embedding information based on acoustic features to the content embedding information based on pronunciation movement features;

[0038] Construct the speaker inverse inference network by using a one-layer linear network, so as to realize the mapping from the speaker identity embedding information based on acoustic features to the speaker identity embedding information based on pronunciation movement features.

[0039] In some specific implementation processes, the pronunciation inverse inference module is further configured to train the content inverse inference network and the speaker inverse inference network according to the following formula:

[0040] L = L c + L s + L rec

[0041]

[0042] wherein, L, L c , L s and L recrespectively represent the overall cost, content embedding error, speaker embedding error, and pronunciation motion reconstruction error; is the content embedding information of the true pronunciation space corresponding to the i-th sample, is the content embedding information of the pronunciation space estimated by the content inverse inference network; is the speaker identity embedding information of the true pronunciation space corresponding to the i-th sample, is the speaker identity embedding information of the pronunciation space estimated by the speaker inverse inference network; is the true pronunciation motion feature corresponding to the i-th sample, is the pronunciation motion feature estimated by the pronunciation decoder.

[0043] Beneficial effects of the above technical solution:

[0044] The present application provides a speaker-independent pronunciation inverse inference method based on a variational autoencoder. Different from the traditional method based on aggregating data of multiple people, the proposed solution in the present invention decouples speech features to obtain content embedding and speaker identity embedding based on speech features, then maps the content embedding and speaker identity embedding based on speech features to the content embedding and speaker identity embedding based on pronunciation features respectively, and finally generates pronunciation motion features by comprehensively considering the content embedding and speaker identity embedding through a decoder. The method proposed in the present invention performs inverse inference on two different information channels respectively, and then synthesizes the inverse inference results of different channels through a decoder; while the traditional method concatenates the content embedding and speaker identity embedding decoupled from speech features into a new embedding in the feature dimension, and then directly estimates pronunciation features from the concatenated embedding using an inverse inference network. Compared with the traditional method, the present application can flexibly construct different mappings according to the distribution characteristics of the embeddings in different channels, achieving better results. Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0046] Figure 1 is a schematic flowchart of a speaker-independent pronunciation inverse inference method based on a variational autoencoder provided by an embodiment of the present invention;

[0047] Figure 2 is a schematic system framework diagram of a speaker-independent pronunciation inverse inference method based on a variational autoencoder provided by an embodiment of the present invention;

[0048] Figure 3 A running framework diagram of a voice or pronunciation variational autoencoder for a speaker-independent pronunciation inverse inference method provided by an embodiment of the present invention;

[0049] Figure 4 A schematic structural diagram of a speaker-independent pronunciation inverse inference system based on a variational autoencoder provided by an embodiment of the present invention. Detailed implementation manners

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0051] Examples of the embodiments are shown in the accompanying drawings, where the same or similar symbols represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.

[0052] Embodiment 1

[0053] An embodiment of the present invention provides a speaker-independent pronunciation inverse inference method based on a variational autoencoder. Referring to Figures 1 - 3 as shown, it includes:

[0054] Step S1: Collect voice acoustic data and pronunciation movement data.

[0055] Specifically, an electromagnetic articulograph is used to collect voice acoustic data and pronunciation movement data.

[0056] Step S2: Train a voice variational autoencoder according to the voice acoustic data, and use the voice variational autoencoder to extract content embedding information and speaker identity embedding information based on acoustic features from the voice acoustic data.

[0057] In a specific embodiment of the present invention, step S2 includes:

[0058] Step S21: Based on the parameters of the trained voice variational autoencoder, construct a voice feature extractor, a local feature extractor, a content feature encoder, and a personal feature encoder, and use the voice feature extractor to extract acoustic features in the voice acoustic data;

[0059] Step S22: Based on the acoustic features, use the local feature extractor, the content feature encoder, and the personal feature encoder respectively to obtain content embedding information and speaker identity embedding information based on acoustic features.

[0060] Step S3: Train a pronunciation variational autoencoder based on the pronunciation motion data, and use the pronunciation variational autoencoder to extract content embedding information and speaker identity embedding information based on pronunciation motion features from the pronunciation motion data.

[0061] In a specific embodiment of the present invention, in steps S2 and S3, the parameters of the speech variational autoencoder and the pronunciation variational autoencoder are optimized according to the cost function of the standard variational autoencoder respectively:

[0062] L = L rec + L KL

[0063]

[0064] where L, L rec , L KL represent the total cost, the reconstruction error, and the KL divergence respectively; z c , z s represent the outputs of the content feature encoder and the personal feature encoder respectively; N represents the number of speech features or pronunciation features; represents the j-th dimensional component of the mean of the normal distribution that z c obeys, and ∧ ij represents the j-th component of the covariance matrix of this normal distribution; d is the dimension of z c .

[0065] In a specific embodiment of the present invention, step S3 further includes: constructing a pronunciation decoder based on the parameters of the trained speech variational autoencoder.

[0066] Step S4: Construct a pronunciation inverse model, and use the content embedding information and speaker identity embedding information based on acoustic features and the content embedding information and speaker identity embedding information based on pronunciation motion features to train the pronunciation inverse model.

[0067] In a specific embodiment of the present invention, the pronunciation inverse model includes a content inverse network and a speaker inverse network, and step S4 includes:

[0068] Step S41: Construct the content inverse network by stacking a multi-layer bidirectional long short-term recursive neural network and two layers of fully connected neural networks to realize the mapping from the content embedding information based on acoustic features to the content embedding information based on pronunciation motion features;

[0069] Step S42: Construct the speaker inverse network by using a one-layer linear network to realize the mapping from the speaker identity embedding information based on acoustic features to the speaker identity embedding information based on pronunciation motion features.

[0070] Step S5: Input the speech acoustic data into the trained pronunciation inverse inference model, and output the pronunciation motion data corresponding to the speech acoustic data.

[0071] In a specific embodiment of the present invention, step S5 includes:

[0072] The trained pronunciation inverse inference model outputs content embedding information and speaker identity embedding information based on pronunciation motion features according to the speech acoustic data, and decodes the corresponding pronunciation motion data by using the pronunciation decoder.

[0073] In a specific embodiment of the present invention, step S4 further includes:

[0074] Step S43: Train the content inverse inference network and the speaker inverse inference network according to the following formula:

[0075] L = L c + L s + L rec

[0076]

[0077] wherein, L, L c , L s and L rec respectively represent the overall cost, content embedding error, speaker embedding error, and pronunciation motion reconstruction error; is the content embedding information of the real pronunciation space corresponding to the i-th sample, is the content embedding information of the pronunciation space estimated by the content inverse inference network; is the speaker identity embedding information of the real pronunciation space corresponding to the i-th sample, is the speaker identity embedding information of the pronunciation space estimated by the speaker inverse inference network; is the real pronunciation motion feature corresponding to the i-th sample, is the pronunciation motion feature estimated by using the pronunciation decoder.

[0078] Among them, the content embedding error is used to optimize the parameters of the content embedding inverse inference network;; the speaker embedding error is used to optimize the parameters of the speaker inverse inference network; the pronunciation motion reconstruction error is used to comprehensively optimize the parameters of the pronunciation decoder, content embedding inverse inference network, and speaker inverse inference network.

[0079] Specifically, initialize the corresponding components shown in Figure 2 with the parameters of the feature extractor, local feature extractor, content feature encoder, and personal feature encoder in the speech variational autoencoder; during the subsequent training process Figure 2 the parameters of these components remain unchanged;

[0080] Initialize with the parameters of the decoder in the pronunciation variational auto - encoder Figure 2 the decoder shown in; during the subsequent training process for Figure 2 fine - tune the parameters of the decoder shown in;

[0081] Random initialization Figure 2 the parameters of the personal feature inverse - inference network and the content feature inverse - inference network in; adjust the parameters of these two networks during the subsequent training process.

[0082] Specifically, analyze and synthesize the speech signal through a speech feature extractor, a local feature extractor, a content encoder, a personal feature encoder, a content inverse - inference network, a speaker inverse - inference network, and a decoder to obtain the movement trajectory of the corresponding speech organs. Among them, the speech feature extractor, the local feature extractor, the content encoder, and the personal feature encoder are pre - trained through speech acoustic data and pronunciation movement data, and the parameters of each component remain unchanged during the entire pronunciation inverse - inference training process; the initial parameters of the decoder come from the pronunciation variational auto - encoder and are fine - tuned during the training process of the entire pronunciation inverse - inference system; the parameters of the content inverse - inference network and the speaker inverse - inference network are randomly initialized and the parameters are continuously updated during the training process of the entire pronunciation inverse - inference system.

[0083] This application is different from the traditional method based on aggregating multi - person data. The proposed solution of the present invention extracts content embeddings and speaker identity embeddings based on acoustic features, then maps the content embeddings and speaker identity embeddings based on acoustic features to content embeddings and speaker identity embeddings based on pronunciation movement features respectively, and finally synthesizes the content embeddings and speaker identity embeddings through a pronunciation decoder to generate pronunciation movement features.

[0084] Embodiment 2

[0085] An embodiment of the present invention provides a speaker - independent pronunciation inverse - inference system based on a variational auto - encoder. Referring to Figure 4 shown, it includes:

[0086] A speech acquisition module 10 for acquiring speech acoustic data and pronunciation movement data;

[0087] A feature extraction module 20 for extracting acoustic features from the speech acoustic data;

[0088] An acoustic information extraction module 30 for obtaining content embedding information and speaker embedding information based on the acoustic features from the acoustic features;

[0089] The pronunciation reverse inference module 40, including a content reverse inference network and a speaker reverse inference network, is used to map the content embedding information and speaker identity embedding information based on acoustic features to the content embedding information and speaker identity embedding information based on pronunciation movement features respectively;

[0090] The pronunciation decoding module 50 is used to decode the pronunciation movement data from the content embedding information and speaker embedding information based on pronunciation movement features.

[0091] In a specific embodiment of the present invention, the pronunciation reverse inference module 40 is used for:

[0092] The content reverse inference network is constructed by stacking a multi-layer bidirectional long short-term recurrent neural network and two layers of fully connected neural networks to realize the mapping from the content embedding information based on acoustic features to the content embedding information based on pronunciation movement features;

[0093] The speaker reverse inference network is constructed by using a single-layer linear network to realize the mapping from the speaker identity embedding information based on acoustic features to the speaker identity embedding information based on pronunciation movement features.

[0094] In a specific embodiment of the present invention, the pronunciation reverse inference module 40 is also used to train the content reverse inference network and the speaker reverse inference network according to the following formula:

[0095] L = L c + L s + L rec

[0096]

[0097]

[0098] where L, L c , L s and L rec represent the overall cost, content embedding error, speaker embedding error, and pronunciation movement reconstruction error respectively; is the content embedding information of the real pronunciation space corresponding to the i-th sample, is the content embedding information of the pronunciation space estimated by the content reverse inference network; is the speaker identity embedding information of the real pronunciation space corresponding to the i-th sample, is the speaker identity embedding information of the pronunciation space estimated by the speaker reverse inference network; is the real pronunciation organ position vector corresponding to the i-th sample, is the pronunciation organ position vector estimated by the pronunciation decoder.

[0099] Different from traditional methods based on aggregating data from multiple people, the solution proposed by the present invention extracts content embeddings and speaker identity embeddings based on acoustic features, then maps the content embeddings and speaker identity embeddings based on acoustic features to content embeddings and speaker identity embeddings based on articulatory motion features respectively, and finally generates articulatory motion features by comprehensively considering the content embeddings and speaker embeddings through an articulatory decoder.

[0100] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

[0101] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general computer, a special computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1Steps of the functions specified in one or more boxes. Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention. Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.

[0102] The above has introduced the method and device provided by the present invention in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0103] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", "one specific embodiment" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0104] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A speaker-independent pronunciation inverse inference method based on variational autoencoder, characterized in that Including: Step S1: Collect speech acoustic data and articulatory movement data; Step S2: Train a speech variational autoencoder according to the speech acoustic data, and use the speech variational autoencoder to extract content embedding information and speaker identity embedding information based on acoustic features from the speech acoustic data; Step S3: Train an articulatory variational autoencoder according to the articulatory movement data, and use the articulatory variational autoencoder to extract content embedding information and speaker identity embedding information based on articulatory movement features from the articulatory movement data; Step S4: Construct an articulatory inverse inference model, and use the content embedding information and speaker identity embedding information based on acoustic features and the content embedding information and speaker embedding identity information based on articulatory movement features to train the articulatory inverse inference model; Step S5: Input the speech acoustic data into the trained articulatory inverse inference model, and output the articulatory movement trajectory corresponding to the speech acoustic data.

2. The speaker-independent pronunciation inverse inference method based on a variational autoencoder according to claim 1, wherein In steps S2 and S3, optimize the parameters of the speech variational autoencoder and the articulatory variational autoencoder according to the cost function of the standard variational autoencoder respectively: L = L rec + L KL Among them, L, L rec , L KL represent the overall cost, reconstruction error, and KL divergence respectively; z c , z s represent the outputs of the content feature encoder and the personal feature encoder respectively; N represents the number of speech features or pronunciation features; represents the j-th dimensional component of the mean of the normal distribution that z c obeys, and ∧ ij represents the j-th component of the covariance matrix of this normal distribution; d is the dimension of z c .

3. The speaker-independent pronunciation inverse inference method based on variational autoencoder according to claim 1, characterized in that Step S2 includes: Step S21: Based on the parameters of the trained speech variational autoencoder, construct a speech feature extractor, a local feature extractor, a content feature encoder, and a personal feature encoder, and use the speech feature extractor to extract acoustic features from the speech acoustic data; Step S22: Based on the acoustic features, use the local feature extractor, the content feature encoder, and the personal feature encoder to obtain content embedding information and speaker identity embedding information based on acoustic features respectively.

4. The speaker-independent pronunciation inverse inference method based on variational autoencoder according to claim 1, characterized in that, The articulatory inverse inference model includes a content inverse inference network and a speaker inverse inference network. Step S4 includes: Step S41: Construct the content inverse inference network by stacking two fully connected neural networks on a multi-layer bidirectional long short-term recurrent neural network to realize the mapping from the content embedding information based on acoustic features to the content embedding information based on articulatory movement features; Step S42: Construct the speaker inverse inference network by using a single-layer linear network to realize the mapping from the speaker identity embedding information based on acoustic features to the speaker identity embedding information based on articulatory movement features.

5. The method for speaker-independent pronunciation inverse inference based on variational autoencoder according to claim 4, wherein Step S3 further includes: Construct an articulatory decoder based on the parameters of the trained speech variational autoencoder.

6. The speaker-independent pronunciation inverse inference method based on variational autoencoder according to claim 5, wherein Step S5 includes: The trained articulatory inverse inference model outputs content embedding information and speaker embedding information based on articulatory movement features according to the speech acoustic data, and decodes the corresponding articulatory movement data by using the articulatory decoder.

7. The speaker-independent pronunciation inverse inference method based on variational autoencoder according to claim 5, wherein Step S4 further includes: Step S43: Train the content inverse inference network and the speaker inverse inference network according to the following formula: L = L c + L s + L rec Among them, L, L c , L s and L rec respectively represent the overall cost, content embedding error, speaker identity embedding error, and pronunciation motion reconstruction error; is the content embedding information of the real pronunciation space corresponding to the i-th sample, is the content embedding information of the pronunciation space estimated by the content inverse inference network; is the speaker identity embedding information of the real pronunciation space corresponding to the i-th sample, is the speaker identity embedding information of the pronunciation space estimated by the speaker inverse inference network; is the real pronunciation feature corresponding to the i-th sample, is the pronunciation feature estimated by using the pronunciation decoder.

8. A speaker-independent pronunciation inverse inference system based on a variational autoencoder, characterized in that, Including: A speech acquisition module for collecting speech acoustic data and articulatory movement data; A feature extraction module for extracting acoustic features from speech acoustic data; An acoustic information extraction module for obtaining content embedding information and speaker identity embedding information based on acoustic features from the acoustic features; The pronunciation reverse inference module, including a content reverse inference network and a speaker reverse inference network, is used to map the content embedding information and speaker embedding information based on acoustic features to the content embedding information and speaker identity embedding information based on pronunciation motion features respectively; The pronunciation decoding module is used to decode pronunciation motion data from the content embedding information and speaker identity embedding information based on pronunciation motion features.

9. The speaker-independent pronunciation inverse inference system based on a variational autoencoder according to claim 8, characterized in that, The pronunciation reverse inference module is used for: Constructing the content reverse inference network by using a multi-layer bidirectional long short-term recursive neural network stacked with two layers of fully connected neural networks to realize the mapping from the content embedding information based on acoustic features to the content embedding information based on pronunciation motion features; Constructing the speaker reverse inference network by using a single-layer linear network to realize the mapping from the speaker embedding information based on acoustic features to the speaker embedding information based on pronunciation motion features.

10. The speaker-independent pronunciation inverse inference system based on variational autoencoder according to claim 9, characterized in that, The pronunciation reverse inference module is also used to train the content reverse inference network and the speaker reverse inference network according to the following formula: L = L c + L s + L rec where L, L c , L s and L rec represent the overall cost, content embedding error, speaker embedding error, and articulatory motion reconstruction error, respectively; is the content embedding information of the true articulatory space corresponding to the i-th sample, is the content embedding information of the articulatory space estimated by the content inverse network; is the speaker embedding information of the true articulatory space corresponding to the i-th sample, is the speaker embedding information of the articulatory space estimated by the speaker inverse network; is the true articulatory feature corresponding to the i-th sample, is the articulatory feature estimated by the articulatory decoder.

Citation Information

Patent Citations

  • Pronunciation inversion method based on feature fusion and attention mechanism

    CN111680591A

  • Pronunciation inversion method and system based on pronunciation physiological modeling

    CN116524896A