Voice style reconstruction method and device, equipment and medium

Through the potential spatial model, the source speech is subjected to information decomposition and noise diffusion processing, and style reconstruction is carried out in combination with style feature information, which solves the problems of insufficient quality of speech style conversion and poor style decoupling in the existing technology, and achieves efficient and accurate speech style conversion.

CN120183418APending Publication Date: 2025-06-20PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510439823.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art has problems in the speech style conversion with insufficient conversion quality, poor style decoupling and insufficient zero sample capability, which is difficult to meet the complex needs of financial and medical scenarios.

Method used

The preset latent spatial model is used to decompose the source speech information, obtain content feature information and style feature information, and generate noise diffusion latent variable information through random noise information, and combine style feature information for style reconstruction to achieve efficient conversion of speech style.

Benefits of technology

It significantly improves the quality and flexibility of speech style conversion, enhances style decoupling and zero-sample capabilities, and can achieve efficient and accurate speech style conversion in scenarios such as finance and medical care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183418A_ABST
    Figure CN120183418A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of artificial intelligence and the field of financial science and technology, and relates to a voice style reconstruction method, which comprises the following steps: acquiring a source voice of a first style; performing information decomposition on the source voice by adopting a preset potential space model to obtain content feature information and style feature information of the source voice; generating noise diffusion latent variable information of the source voice based on preset random noise information and the content feature information; combining the noise diffusion latent variable information and the style feature information to obtain reconstructed source voice latent variable information; and adopting a preset style decoder to perform style reconstruction on the source voice latent variable information and the style feature information to obtain a source voice of a second style. The invention further provides a device, equipment and a medium. In addition, the invention also relates to a block chain technology, and the source voice of the first style and the source voice of the second style can be stored in a block chain. According to the invention, accurate voice style conversion can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and financial technology, and particularly to a method, apparatus, device, and medium for reconstructing speech style. Background Art

[0002] The speech style conversion technology aims to achieve the conversion of specific attributes of speech while maintaining the semantic content unchanged, which is of great significance in the field of speech processing and is widely used in many key scenarios such as finance and healthcare.

[0003] In the financial scenario, the speech style conversion technology can be used to enhance the customer service experience. For example, the automatic speech response system of a bank or financial institution can use this technology to convert the cold machine voice into a more cordial and natural style to enhance the interaction and trust with customers. However, the existing technical methods in the industry still face challenges. Specifically, the methods based on statistical modeling, such as using traditional models like Gaussian mixture models and hidden Markov models, although they can achieve basic speech style conversion, the conversion quality is limited, making it difficult to meet the complex style migration requirements and unable to perfectly integrate into the detailed requirements of financial services.

[0004] In the medical scenario, the speech style conversion technology can be used to assist patient treatment or rehabilitation. For example, through this technology, the speech of a doctor or therapist can be converted into a style that the patient is more familiar with or likes, thereby enhancing the treatment effect. However, similarly, the existing technologies also have deficiencies. With the development of deep learning, methods based on models such as autoencoders, recurrent neural networks, and generative adversarial networks have emerged. These methods achieve style conversion by learning the latent space representation of speech features. However, they still have problems of insufficient style decoupling and zero-shot ability, that is, it is difficult to clearly separate the content and style representations of speech, and a large amount of training data with different styles is required.

[0005] In summary, whether in the financial or medical scenario, the speech style conversion technology faces challenges in terms of conversion quality, style decoupling, and zero-shot ability, and further technological breakthroughs and optimizations are urgently needed. Summary of the Invention

[0006] The purpose of the embodiments of this application is to propose a method, apparatus, device, and medium for reconstructing speech style to solve the problems of challenges in aspects such as conversion quality, style decoupling, and zero-shot ability faced by the speech style conversion technology in the existing technology.

[0007] In a first aspect, a method for reconstructing speech style is provided, adopting the following technical solution:

[0008] Obtain the source speech of the first style; use a preset latent space model to decompose the information of the source speech to obtain the content feature information and style feature information of the source speech; generate the noise diffusion latent variable information of the source speech based on the preset random noise information and content feature information; combine the noise diffusion latent variable information and style feature information to obtain the reconstructed source speech latent variable information; use a preset style decoder to perform style reconstruction on the source speech latent variable information and style feature information to obtain the source speech of the second style.

[0009] In a second aspect, a speech style reconstruction device is provided, which adopts the following technical solution:

[0010] A speech acquisition module, configured to acquire the source speech of the first style;

[0011] An information decomposition module, configured to use a preset latent space model to decompose the information of the source speech to obtain the content feature information and style feature information of the source speech;

[0012] An information generation module, configured to generate the noise diffusion latent variable information of the source speech based on the preset random noise information and content feature information;

[0013] An information combination module, configured to combine the noise diffusion latent variable information and style feature information to obtain the reconstructed source speech latent variable information;

[0014] A style reconstruction module, configured to use a preset style decoder to perform style reconstruction on the source speech latent variable information and style feature information to obtain the source speech of the second style.

[0015] In a third aspect, a computer device is provided, including a memory and a processor. The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the speech style reconstruction method as described above are implemented.

[0016] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable at least one processor to execute the steps of the speech style reconstruction method as described above.

[0017] In the solutions implemented by the above-mentioned method, apparatus, device, and medium for reconstructing speech style, through a preset latent space model, the source speech is accurately decomposed, effectively distinguishing content feature information from style feature information. This feature greatly enhances style decoupling and solves the problem in the prior art that it is difficult to clearly separate style from content. Further, by cleverly combining preset random noise information with content feature information, the noise-diffused latent variable information of the source speech is generated. This not only enriches the diversity of speech representation but also significantly improves the conversion quality, enabling it to easily meet the subtle service requirements in financial scenarios. Particularly importantly, during the style reconstruction process, relying only on the source speech latent variable information and style feature information greatly reduces the dependence on a large amount of training data. This feature significantly enhances the zero-shot style conversion ability, enabling efficient and accurate speech style conversion even in sensitive fields such as healthcare in the face of scarce training resources. Brief Description of the Drawings

[0018] To more clearly illustrate the solutions in this application, the following will briefly introduce the drawings required for the description of the embodiments of this application. Obviously, the drawings below are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 is an exemplary system architecture diagram to which this application can be applied;

[0020] Figure 2 is a flowchart of a method for reconstructing speech style provided by this application;

[0021] Figure 3 is a structural diagram of an apparatus for reconstructing speech style provided by this application;

[0022] Figure 4 is a structural diagram of a computer device provided by this application. Detailed Description of the Embodiments

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the description of the embodiments of this application in this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the description and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the description and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0024] References to "embodiments" in this specification mean that particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0025] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0026] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0027] A user may use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.

[0028] The terminal device 101 may be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012, or the mobile phone 1013, the terminal device 101 may also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop portable computer, a desktop computer, etc.

[0029] The server 103 may be a server providing various services, such as a background server supporting the pages displayed on the terminal device 101.

[0030] It should be noted that the method for reconstructing the speech style provided in the embodiments of this application is generally executed by the server. Correspondingly, the device for reconstructing the speech style is generally provided in the server.

[0031] It should be understood, Figure 1The numbers of the terminal devices, networks, and servers therein are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0032] Continuing to refer to Figure 2 , a flowchart of an embodiment of a method for reconstructing a voice style according to the present application is shown. The method for reconstructing a voice style includes the following steps:

[0033] Step S201, obtain the source voice of the first style.

[0034] Wherein, the first style refers to the original voice style of the source voice. This style is specific and can be formal, casual, regional, or any other recognizable voice trait. For example, in a financial scenario, the first style may be the original cold and mechanical style of a bank's automated voice response system.

[0035] Wherein, the source voice refers to the original voice signal to be subjected to style conversion. It contains specific semantic content and the original voice style. For example, a recording of a customer consulting a bank is the source voice.

[0036] Step S202, use a preset latent space model to decompose the information of the source voice to obtain the content feature information and style feature information of the source voice.

[0037] Wherein, the latent space model is a model obtained through training, used to map the voice signal into a latent space, and this space can represent the content and style features of the voice. Through the latent space model, effective separation of the voice style and content can be achieved. For example, through the latent space model, the voice signal can be decomposed into content latent variables and style latent variables.

[0038] Wherein, information decomposition refers to the process of decomposing the voice signal of the source voice into content feature information and style feature information through the latent space model. The purpose of information decomposition is to decouple the voice content and style and provide a basis for subsequent style conversion. For example, decomposing a voice containing specific semantic content into content and style parts.

[0039] Wherein, the content feature information refers to the feature information representing the semantic content in the source voice. These information remain unchanged during the voice style conversion process to ensure that the converted voice still retains the original semantics. For example, in voice style conversion, the semantic content of the greeting is kept unchanged.

[0040] Wherein, the style feature information refers to the feature information representing the voice style in the source voice. These information will be replaced by the feature information of the second style during the style conversion process. For example, replacing the cold style of the original voice with a more cordial style.

[0041] Step S203: Generate the noise diffusion latent variable information of the source speech based on the preset random noise information and content feature information.

[0042] Among them, the random noise information refers to the random noise vector used to generate the noise diffusion latent variable information of the source speech. By adding random noise, the diversity of the speech signal can be increased, and the robustness of style conversion can be improved. For example, in a generative adversarial network, random noise is used to generate diverse speech samples.

[0043] Among them, the noise diffusion latent variable information refers to the latent variable information obtained by fusing the random noise information and the content feature information. This kind of information combines the content of the original speech and random noise, providing a basis for subsequent style conversion. For example, in speech style conversion, the noise diffusion latent variable information is used to generate speech signals with the second style.

[0044] Step S204: Combine the noise diffusion latent variable information and the style feature information to obtain the reconstructed latent variable information of the source speech.

[0045] Among them, the latent variable information of the source speech refers to the latent variable representation of the source speech obtained through information decomposition and noise diffusion processing. This kind of information contains both the content features of the original speech and the latent features of the second style. For example, in speech style conversion, the latent variable information of the source speech is used to generate speech signals with the second style.

[0046] Step S205: Use the preset style decoder to perform style reconstruction on the latent variable information of the source speech and the style feature information to obtain the source speech with the second style.

[0047] Among them, the style decoder is a neural network model used to decode the latent variable information generated by fusing the latent variable information of the source speech and the style feature information into a speech signal with the second style. The style decoder learns through training how to map the latent variable information to the speech signal space. For example, decoding the fused latent variable information into a speech signal with the second style.

[0048] Among them, the second style refers to the style that the source speech is expected to be converted into. This style can be any preset speech trait different from the first style. For example, in a financial scenario, the second style may be a more cordial and natural speech style.

[0049] Embodiments of the present application can precisely decompose the information of the source speech through a preset latent space model, effectively distinguishing the content feature information from the style feature information. This characteristic greatly enhances the style decoupling and solves the problem that it is difficult to clearly separate the style and content in the prior art. Further, by cleverly combining the preset random noise information with the content feature information, the noise diffusion latent variable information of the source speech is generated. This not only enriches the diversity of speech representation but also significantly improves the conversion quality, enabling it to easily meet the nuanced service requirements in the financial scenario. Particularly importantly, during the style reconstruction process, it only relies on the source speech latent variable information and the style feature information, greatly reducing the dependence on a large amount of training data. This characteristic significantly enhances the zero-shot style conversion ability, enabling efficient and accurate speech style conversion even in sensitive fields such as healthcare in the face of scarce training resources.

[0050] In some alternative implementation manners of this embodiment, before step 201 of using a preset latent space model to decompose the information of the source speech to obtain the content feature information and the style feature information of the source speech, the following steps are specifically further included:

[0051] Obtain multiple sample speech data of different styles; use the multiple sample speech data to train a preset encoder to obtain a preset latent space model.

[0052] Among them, the multiple sample speech data refers to a diverse speech data set used to train the latent space model and the style decoder. These sample data contain different speech styles and semantic contents and are used to learn the latent representation of speech and the style conversion rules. For example, a speech data set containing different speaking styles of different people and different emotional expressions.

[0053] Among them, the encoder is a neural network model used to map the input speech signal into the latent space to obtain its latent representation. In speech style conversion, the encoder is used to extract the content and style features of the speech signal. For example, encoding the original speech signal into a vector representation in the latent space.

[0054] In one example, speech data of different styles, including formal, friendly, professional and other styles, can be collected from the automatic speech response systems of multiple banks and financial institutions. Ensure that the sample data is widely covered and diverse in style to fully train the encoder. Preprocess the collected speech data, including denoising, normalizing the sampling rate, etc., to ensure data quality. Use a deep learning framework to construct a preset encoder network. The encoder is designed as a multi-layer convolutional neural network structure to effectively extract speech features. Use the collected multi-style sample speech data to perform unsupervised training on the encoder. The goal is to learn a latent space model that can map different styles of speech to the same latent space while retaining the content and style information of the speech. During the training process, adopt a contrastive learning strategy to enhance the style decoupling ability of the model by maximizing the distance between different styles of speech in the latent space and minimizing the distance between the same content but different styles of speech in the latent space. For example, taking a large bank as an example, its original automatic speech response system used a formal machine speech style, and customer feedback was relatively cold. Through the method of this embodiment, a large amount of customer service speech data containing a friendly style was collected and the encoder was trained. Based on the trained encoder, the formal-style speech was converted into a friendly style, significantly improving the customer's interaction experience.

[0055] In one example, speech data of different therapists and doctors are collected from multiple medical institutions, covering various styles such as gentle, encouraging, professional, etc. Ensure that the sample data is representative and can reflect the common needs in medical rehabilitation. Preprocess the collected speech data, including removing background noise, normalizing the volume, etc., to improve data quality. Construct a preset encoder network based on a recurrent neural network or a long short-term memory network to capture the temporal characteristics of the speech. Use the collected multi-style sample speech data to perform supervised training on the encoder. The goal is to learn a latent space model that can map different styles of speech to the same latent space while keeping the content and style information of the speech separated. During the training process, adopt an attention mechanism to enhance the model's ability to capture key features of the speech and improve the accuracy of style conversion. For example, taking a rehabilitation center as an example, its therapists originally used a unified formal style for speech guidance, and some patients reported a lack of affinity. Through the method of this embodiment, a large amount of speech data containing an encouraging style was collected and the encoder was trained. Based on the trained encoder, the formal-style speech was converted into an encouraging style, significantly improving the patients' treatment enthusiasm and rehabilitation effect.

[0056] Embodiments of the present application can ensure the comprehensiveness and diversity of training data by obtaining sample speech data covering multiple styles, laying a solid foundation for the learning of the encoder. These samples not only contain rich style features but also imply the diversity of speech content, helping the encoder to form a more refined representation in the latent space. Subsequently, using these multi-style samples to train a preset encoder can effectively prompt the encoder to learn the effective separation of style and content. In the latent space model, the content features are stably retained, while the style features show high variability, which provides great flexibility for subsequent style conversion. This training process not only enhances the model's ability to capture speech styles but also, through the construction of the latent space, enables the clear decoupling of style and content features, laying a key technical foundation for subsequent tasks of style transfer based on content features, thus significantly improving the accuracy and naturalness of speech style conversion.

[0057] In some optional implementation manners of this embodiment, in step S202, a preset latent space model is used to decompose the information of the source speech to obtain the content feature information and style feature information of the source speech, which specifically includes the following steps:

[0058] Preprocess the source speech to obtain the preprocessed source speech; input the preprocessed source speech into the latent space model to map the preprocessed source speech to the latent space through the latent space model, and output the content latent variable and style latent variable of the source speech; determine the content latent variable as the content feature information of the source speech, and determine the style latent variable as the style feature information of the source speech.

[0059] Among them, preprocessing refers to a series of preprocessing operations performed on the source speech signal, such as format standardization, normalization, etc., to improve the accuracy and efficiency of subsequent processing. Preprocessing is one of the important steps in the speech style conversion process.

[0060] Among them, the latent space refers to the space obtained through the learning of the encoder and used to represent the latent features of the speech signal. The vectors in the latent space can represent the content and style features of the speech and are the basis for realizing speech style conversion. For example, the vectors in the latent space can represent the speaking styles of different people or the speech signals of different emotional expressions.

[0061] Among them, the content latent variable refers to the vector representing the semantic content of the speech signal in the latent space. The content latent variable remains unchanged during the speech style conversion process to ensure that the converted speech still retains the original semantics. For example, in speech style conversion, the content latent variable that keeps the semantic content of the greeting unchanged.

[0062] Among them, the style latent variable refers to a vector that represents the style characteristics of a speech signal in the latent space. The style latent variable will be replaced by the feature vector of the target style during the style conversion process. For example, replacing the cold style of the original speech with a more amiable style's feature vector is the style latent variable.

[0063] In one example, the original customer service speech is collected from a financial customer service system. This speech style is rather mechanical and formal and serves as the source speech to be converted. The collected source speech is denoised, normalized, and undergoes necessary audio format conversion to ensure the purity and consistency of the input data. The preprocessed source speech is input into a pre-trained latent space model. This model is based on a deep learning framework and adopts a variational autoencoder structure, capable of learning the latent representation space of the speech. The model maps the source speech to the latent space, separating out the content latent variable and the style latent variable. The content latent variable captures the semantic information of the speech, while the style latent variable represents the style characteristics of the speech. From the model output, the content latent variable is directly used as the content feature information of the source speech, and the style latent variable is used as the style feature information. This process achieves the decoupling of content and style, providing the possibility for subsequent style conversion.

[0064] Embodiments of this application can ensure the purity and format consistency of the input data by implementing a series of preprocessing measures on the source speech, such as noise reduction, normalization, and format adjustment, laying a solid foundation for subsequent model processing. The preprocessed source speech is fed into a carefully designed latent space model. This model is based on an advanced deep learning architecture and can accurately map the speech to the latent space, achieving a deep decoupling of content and style. During this process, the content latent variable precisely captures the semantic core of the source speech, while the style latent variable meticulously depicts its style traits. This ingenious separation not only retains the core information of the source speech but also provides infinite possibilities for style reconstruction and conversion. By clearly defining the content feature information and the style feature information, it lays an efficient and accurate path for subsequent style conversion steps, greatly enhancing the flexibility and precision of speech style conversion and laying a solid foundation for subsequent steps such as the generation of noise diffusion latent variable information and style reconstruction.

[0065] In some optional implementation manners of this embodiment, in step S203, based on the preset random noise information and content feature information, generate the noise diffusion latent variable information of the source speech, which specifically includes the following steps:

[0066] Obtain the content dimension information of the content feature information; generate a random noise vector with dimensions matching the content dimension information according to a preset noise distribution and the content dimension information; perform normalization processing on the random noise vector to obtain the preset random noise information; use a preset diffusion model to fuse the random noise information and the content feature information to generate the noise diffusion latent variable information of the source speech.

[0067] Among them, the content dimension information refers to the dimension information of the content latent variable, that is, the dimension size of the content latent variable in the latent space. The content dimension information is used to generate a random noise vector matching the dimension of the content latent variable. For example, if the dimension of the content latent variable is 128 dimensions, the generated random noise vector should also be 128 dimensions.

[0068] Among them, the noise distribution refers to the type of noise distribution used to generate the random noise vector, such as Gaussian distribution, uniform distribution, etc. The choice of the noise distribution will affect the characteristics and diversity of the generated random noise vector. For example, in the Gaussian mixture model, the Gaussian distribution is selected as the noise distribution type.

[0069] Among them, the random noise vector refers to a random vector generated according to a preset noise distribution, which is used to increase the diversity of the speech signal. The random noise vector is fused with the content latent variable during the style conversion process to generate the noise diffusion latent variable information. For example, in the Gaussian mixture model, the generated random noise vector is used to increase the diversity of the speech signal.

[0070] Among them, the normalization processing refers to a series of processing operations performed on the random noise vector, such as normalization, mean removal, etc., to ensure that it meets the requirements of subsequent processing. The normalization processing can improve the stability and accuracy of the style conversion. For example, the random noise vector is normalized so that its mean is 0 and its variance is 1.

[0071] Among them, the diffusion model is a generative model used to fuse the random noise information and the content feature information to generate a speech signal with a second style. The diffusion model realizes the generation and style conversion of the speech signal by gradually adding noise and learning the inverse process. For example, in the generative adversarial network, the diffusion model is used to generate diverse speech samples.

[0072] In one example, the source speech is decomposed by a preset latent space model to obtain content feature information and style feature information. On this basis, the content dimension information of the content feature information is further analyzed, including but not limited to the number of dimensions of the feature, the value range and distribution characteristics of each dimension, etc. According to the preset noise distribution (such as normal distribution, uniform distribution, etc.) and the content dimension information, a random noise vector whose dimensions match the content dimension information is generated. To ensure the stability and controllability of the noise vector, it is normalized to obtain the preset random noise information. This step aims to introduce appropriate noise to enrich the latent representation of the source speech and provide more possibilities for subsequent style conversion. A preset diffusion model is used to fuse the normalized random noise information and the content feature information. The diffusion model can gradually introduce noise and control the diffusion process of the noise, enabling the latent representation to be more flexible and diverse while keeping the content features unchanged. This step generates the noise diffusion latent variable information of the source speech, providing a basis for subsequent style reconstruction.

[0073] The embodiments of the present application can accurately obtain the content dimension information of the content feature information, deeply understand the internal structure of the source speech, and lay a solid foundation for subsequent steps. The random noise vector generated based on the content dimension information and the preset noise distribution not only strictly matches the content features in terms of dimensions but also ensures the rationality and effectiveness of noise introduction. The normalization process further improves the stability and controllability of the noise vector, creating favorable conditions for the subsequent fusion process. By using an advanced diffusion model, the random noise information and the content feature information are skillfully fused. This process not only enriches the latent representation of the source speech but also enhances the generalization ability of the model through the diffusion characteristics of the noise. The generated noise diffusion latent variable information not only retains the core content of the source speech but also contains rich style conversion potential, providing a high-quality data basis for subsequent style reconstruction steps, thus significantly improving the flexibility and accuracy of speech style conversion.

[0074] In some optional implementation manners of this embodiment, in step S204, the noise diffusion latent variable information and the style feature information are combined to obtain the reconstructed source speech latent variable information, which specifically includes the following steps:

[0075] Obtain a preset fusion strategy; based on the fusion strategy, fuse the noise diffusion latent variable information and the style feature information to obtain the fused information after content-style fusion; input the fused information into the diffusion model for reverse diffusion processing to obtain the reconstructed source speech latent variable information.

[0076] Among them, the fusion strategy refers to the method or strategy of fusing the noise diffusion latent variable information and the style feature information. The choice of the fusion strategy will affect the quality of the finally generated speech signal and the effect of style conversion. For example, fusion is performed by means of weighted summation or splicing.

[0077] Among them, the fusion information refers to the information obtained after fusing the noise diffusion latent variable information and the style feature information. The fusion information contains both the content features of the original speech (i.e., the source speech) and the latent features of the second style, providing a basis for subsequent style decoding. For example, in speech style conversion, the fusion information is used to generate a speech signal with the second style.

[0078] Among them, the reverse diffusion process refers to the process of gradually restoring the fusion information into a speech signal with the second style through the inverse process of the diffusion model. The reverse diffusion process is one of the important steps in style decoding and is used to generate high-quality speech signals. For example, in a generative adversarial network, the fusion information is restored into a speech signal with the second style through the reverse diffusion process.

[0079] In one example, the speech style conversion of a financial customer service in the financial customer service field is used as an illustration. The preset fusion strategy is designed according to the actual needs of the financial customer service field. Considering that the financial customer service speech needs to convey a kind, natural and professional style, the fusion strategy is set to moderately enhance the weight of the style feature information on the basis of retaining the content feature information. This strategy is represented by algorithmic parameters, which is convenient for flexible application in subsequent steps. After obtaining the content feature information and the style feature information of the source speech, the preset fusion strategy is adopted to fuse the noise diffusion latent variable information and the style feature information. The fusion process follows the weight distribution set in the strategy to ensure that the content feature information is not overly interfered with while the style feature information is effectively enhanced. This step generates the fusion information after content-style fusion, providing key data for subsequent reverse diffusion processing. The fusion information is input into the diffusion model for reverse diffusion processing. The reverse diffusion process restores the clearer and more stable source speech latent variable information by gradually removing noise. This process not only retains the core content of the source speech but also effectively integrates the style features, laying a solid foundation for the final style reconstruction.

[0080] In the embodiments of the present application, through a fusion strategy, it is possible to intelligently guide the fusion process of noise diffusion latent variable information and style feature information. This fusion strategy not only ensures the integrity of the content feature information but also subtly incorporates the style feature information, enabling the fused information to retain the core semantics of the source speech while containing the unique charm of the target style. Under the guidance of the fusion strategy, the noise diffusion latent variable information and the style feature information are precisely aligned and fused in the latent space, laying a solid foundation for subsequent reverse diffusion processing. The reverse diffusion processing utilizes the powerful capabilities of the diffusion model to gradually remove the noise components in the fused information and recover the clear and stable source speech latent variable information. This process not only enhances the coherence and naturalness of the speech but also effectively embeds the style features, providing high-quality data support for the final style reconstruction and thus significantly improving the overall effect of speech style conversion.

[0081] In some optional implementation manners of this embodiment, in step S205, a preset style decoder is used to perform style reconstruction on the source speech latent variable information and the style feature information to obtain the source speech in the second style, which specifically includes the following steps:

[0082] Fuse the source speech latent variable information and the style feature information to obtain the fused latent variable information; input the latent variable information into the preset style decoder, and gradually generate a speech signal through the neural network layers of the style decoder to obtain the source speech in the second style.

[0083] Among them, the latent variable information refers to the latent variable representation obtained after information decomposition, noise diffusion, and fusion processing. The latent variable information contains both the content features of the original speech and the latent features of the target style, and is the input for style decoding. For example, in speech style conversion, the latent variable information is used to generate a speech signal with the target style.

[0084] In an example, a preset latent space model is used to decompose the source speech into content feature information and style feature information. Subsequently, the source speech latent variable information and the style feature information are fused. This fusion process not only retains the core semantic content of the source speech but also subtly incorporates the features of the second style. The fusion strategy is adopted to ensure the effective combination of the style features and the information latent variables while avoiding information loss or confusion. The fused latent variable information is input into the preset style decoder. The style decoder contains multiple layers of neural networks, and these network layers gradually process the input latent variable information and finally generate a speech signal with the second style.

[0085] In the embodiments of the present application, through the deep fusion of the source speech latent variable information and the style feature information, the key steps of speech style conversion are achieved. In this process, the core semantic content of the source speech is retained, and at the same time, the features of the second style are accurately incorporated, ensuring that the converted speech not only maintains the integrity of the original information but also presents a brand-new style appearance. The fused latent variable information, as the input of the style decoder, is gradually decoded through its carefully designed neural network layers to gradually generate a speech signal with the second style. With its powerful learning and generalization capabilities, the style decoder can delicately capture and reproduce the unique charm of the second style, thereby outputting a high-quality source speech with the second style. This process not only improves the accuracy and naturalness of speech style conversion but also further enhances the practicality and adaptability of the technology, bringing significant optimization and improvement to the speech interaction experience in fields such as finance and healthcare.

[0086] In some alternative implementation manners of this embodiment, in step S205, after using a preset style decoder to perform style reconstruction on the source speech latent variable information and the style feature information to obtain the source speech with the second style, the following steps are specifically further included:

[0087] Obtain a pre-trained discriminator, input the source speech with the second style into the discriminator for discriminant analysis to obtain discriminant information; based on the discriminant information, adjust the parameters of the latent space model and the parameters of the style decoder.

[0088] Among them, the discriminator is a neural network model used to make true / false judgments and style discrimination on the generated speech signal (the source speech with the second style). The discriminator learns through training how to distinguish between real speech signals and generated speech signals, as well as the differences between different styles.

[0089] Among them, discriminant analysis refers to the process of using the discriminator to analyze the generated speech signal (the source speech with the second style). The results of discriminant analysis are used to adjust the parameters of the latent space model and the parameters of the style decoder to improve the quality and accuracy of style conversion.

[0090] Among them, discriminant information refers to the information obtained after the discriminator makes true / false judgments and style discrimination on the generated speech signal (the source speech with the second style). The discriminant information includes the true / false judgment result and the style discrimination result, which are used to guide the parameter adjustment of the latent space model and the style decoder.

[0091] In one example, first, a financial customer service source speech with a first style, such as a cold machine voice, can be obtained. Then, a preset latent space model is used to decompose the information of the source speech to obtain content feature information and style feature information. Based on the preset random noise information and content feature information, noise diffusion latent variable information of the source speech is generated. Subsequently, the noise diffusion latent variable information is combined with the style feature information to obtain reconstructed source speech latent variable information. A preset style decoder is used to reconstruct the style of the source speech latent variable information and style feature information to obtain a source speech with a second style (such as a kind and natural style). Then, a pre-trained discriminator is obtained, and the source speech with the second style is input into the discriminator for discriminant analysis. The discriminant analysis includes true / false judgment and style discrimination to evaluate whether the converted speech is real and natural and whether it conforms to the second style. Based on the discriminant information output by the discriminator, the parameters of the latent space model and the parameters of the style decoder are adjusted. Through an iterative optimization process, the quality and accuracy of speech style conversion are continuously improved.

[0092] In the embodiments of the present application, by introducing a pre-trained discriminator, accurate evaluation of the speech after style conversion can be achieved. This discriminator can not only judge the authenticity of the converted speech to ensure the integrity of semantic content during the conversion process, but also accurately discriminate the style of the speech, thereby verifying the effect of style conversion. Based on the discriminant information output by the discriminator, the parameters of the latent space model and the parameters of the style decoder can be accurately adjusted. This process forms a closed-loop feedback mechanism, enabling the model to continuously self-optimize and gradually improve the accuracy and naturalness of style conversion. Through this iterative optimization method, the problem of insufficient style decoupling is effectively solved, so that the converted speech can be closer to the second style while maintaining the semantic content unchanged, meeting the high-quality requirements for speech style conversion in scenarios such as finance and healthcare.

[0093] It should be emphasized that to further ensure the above-mentioned source speech with the first style, content feature information, style feature information, noise diffusion latent variable information, source speech latent variable information, and source speech with the second style, the above-mentioned source speech with the first style, content feature information, style feature information, noise diffusion latent variable information, source speech latent variable information, and source speech with the second style can also be stored in a node of a blockchain.

[0094] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, which is a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of its information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.

[0095] The embodiments of this application can construct and optimize relevant models and networks based on artificial intelligence technologies, such as latent space models, style decoders, diffusion models, etc. Among them, the artificial intelligence (AI) model is the crystallization of theory and practice that simulates the decision-making process of human intelligence through algorithms and data analysis to solve complex problems, predict future trends, or achieve automated tasks. These models utilize a large amount of historical data and real-time information and are trained and optimized through specific algorithm frameworks to achieve efficient, accurate, and reliable performance.

[0096] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), etc., or a random access memory (RAM), etc.

[0097] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0098] Further reference Figure 3 to Figure 2 As an implementation of the method shown above, this application provides an embodiment of a voice style reconstruction device. This device embodiment is related to Figure 2The method embodiments shown correspond to various electronic devices in which the apparatus can be specifically applied.

[0099] As Figure 3 shown, the speech style reconstruction apparatus 400 of this embodiment includes: a speech acquisition module 401, an information decomposition module 402, an information generation module 403, an information combination module 404, and a style reconstruction module 405. Among them:

[0100] The speech acquisition module 401 is configured to acquire the source speech of the first style;

[0101] The information decomposition module 402 is configured to decompose the information of the source speech by using a preset latent space model to obtain the content feature information and style feature information of the source speech;

[0102] The information generation module 403 is configured to generate the noise diffusion latent variable information of the source speech based on the preset random noise information and content feature information;

[0103] The information combination module 404 is configured to combine the noise diffusion latent variable information and the style feature information to obtain the reconstructed source speech latent variable information;

[0104] The style reconstruction module 405 is configured to perform style reconstruction on the source speech latent variable information and the style feature information by using a preset style decoder to obtain the source speech of the second style.

[0105] The embodiments of the present application can accurately decompose the information of the source speech through a preset latent space model, effectively distinguishing the content feature information and the style feature information. This feature greatly enhances the style decoupling and solves the problem that it is difficult to clearly separate the style and content in the prior art. Further, the preset random noise information and the content feature information are cleverly combined to generate the noise diffusion latent variable information of the source speech. This not only enriches the diversity of speech representation but also significantly improves the conversion quality, enabling it to easily meet the delicate service requirements in the financial scenario. Particularly importantly, in the style reconstruction process, it only depends on the source speech latent variable information and the style feature information, greatly reducing the dependence on a large amount of training data. This feature significantly enhances the zero-shot style conversion ability, enabling efficient and accurate speech style conversion even in sensitive fields such as medical care in the face of scarce training resources.

[0106] In one embodiment, the information decomposition module 402 includes:

[0107] The preprocessing sub-module is configured to preprocess the source speech to obtain the preprocessed source speech;

[0108] A voice input sub-module, configured to input the preprocessed source voice into a latent space model, so as to map the preprocessed source voice to a latent space through the latent space model, and output the content latent variable and style latent variable of the source voice;

[0109] An information determination sub-module, configured to determine the content latent variable as the content feature information of the source voice, and determine the style latent variable as the style feature information of the source voice.

[0110] In the embodiment of the present application, a series of preprocessing measures can be implemented on the source voice, such as noise reduction, normalization, and format adjustment, which ensure the purity and format consistency of the input data and lay a solid foundation for subsequent model processing. The preprocessed source voice is fed into a carefully designed latent space model, which is based on an advanced deep learning architecture and can accurately map the voice to the latent space to achieve deep decoupling of content and style. In this process, the content latent variable accurately captures the semantic core of the source voice, while the style latent variable meticulously depicts its style characteristics. This ingenious separation not only retains the core information of the source voice but also provides infinite possibilities for style reconstruction and conversion. By clearly defining the content feature information and style feature information, an efficient and accurate path is paved for subsequent style conversion steps, greatly improving the flexibility and accuracy of voice style conversion and laying a solid foundation for subsequent steps such as the generation of noise diffusion latent variable information and style reconstruction.

[0111] In one embodiment, the information generation module 403 includes:

[0112] A first acquisition sub-module, configured to acquire the content dimension information of the content feature information;

[0113] A sub-module, configured to generate a random noise vector with a dimension matching the content dimension information according to a preset noise distribution and the content dimension information;

[0114] A normalization sub-module, configured to perform normalization processing on the random noise vector to obtain preset random noise information;

[0115] A first fusion sub-module, configured to fuse the random noise information and the content feature information by using a preset diffusion model to generate the noise diffusion latent variable information of the source voice.

[0116] Embodiments of the present application can accurately obtain the content dimension information of content feature information, deeply understand the internal structure of the source speech, and lay a solid foundation for subsequent steps. The random noise vector generated based on the content dimension information and the preset noise distribution not only strictly matches the content features in terms of dimensions, but also ensures the rationality and effectiveness of noise introduction. The normalization process further improves the stability and controllability of the noise vector, creating favorable conditions for the subsequent fusion process. By adopting an advanced diffusion model, the random noise information and the content feature information are skillfully fused. This process not only enriches the latent representation of the source speech, but also enhances the generalization ability of the model through the diffusion characteristics of the noise. The generated noise diffusion latent variable information not only retains the core content of the source speech, but also contains rich potential for style conversion, providing a high-quality data basis for subsequent style reconstruction steps, thus significantly improving the flexibility and accuracy of speech style conversion.

[0117] In one embodiment, the information combination module 404 includes:

[0118] The second acquisition sub-module is used to acquire a preset fusion strategy;

[0119] The second fusion sub-module is used to fuse the noise diffusion latent variable information and the style feature information based on the fusion strategy to obtain the fused information after content-style fusion;

[0120] The diffusion sub-module is used to input the fused information into the diffusion model for reverse diffusion processing to obtain the reconstructed source speech latent variable information.

[0121] Embodiments of the present application can, through the fusion strategy, intelligently guide the fusion process of the noise diffusion latent variable information and the style feature information. This fusion strategy not only ensures the integrity of the content feature information, but also skillfully incorporates the style feature information, making the fused information retain the core semantics of the source speech and contain the unique charm of the target style. Under the guidance of the fusion strategy, the noise diffusion latent variable information and the style feature information are accurately aligned and fused in the latent space, laying a solid foundation for subsequent reverse diffusion processing. The reverse diffusion processing utilizes the powerful ability of the diffusion model to gradually remove the noise components in the fused information and restore clear and stable source speech latent variable information. This process not only enhances the coherence and naturalness of the speech, but also effectively embeds the style features, providing high-quality data support for the final style reconstruction, thus significantly improving the overall effect of speech style conversion.

[0122] In one embodiment, the style reconstruction module 405 includes:

[0123] The third fusion sub-module is used to fuse the source speech latent variable information and the style feature information to obtain the fused latent variable information;

[0124] An information input sub-module, configured to input latent variable information into a preset style decoder, and gradually generate a speech signal through the neural network layer of the style decoder to obtain a source speech in a second style.

[0125] In the embodiments of the present application, through the deep fusion of the latent variable information and style feature information of the source speech, the key steps of speech style conversion are realized. In this process, the core semantic content of the source speech is retained, and at the same time, the features of the second style are accurately incorporated, ensuring that the converted speech not only maintains the integrity of the original information but also presents a brand-new style appearance. The fused latent variable information, as the input of the style decoder, is gradually decoded through its carefully designed neural network layer, and a speech signal with the second style is gradually generated. With its powerful learning and generalization capabilities, the style decoder can delicately capture and reproduce the unique charm of the second style, thereby outputting a high-quality source speech in the second style. This process not only improves the accuracy and naturalness of speech style conversion but also further enhances the practicality and adaptability of the technology, bringing significant optimization and improvement to the speech interaction experience in fields such as finance and healthcare.

[0126] In one embodiment, the speech style reconstruction device 400 further includes:

[0127] A data acquisition module, configured to acquire multiple sample speech data of different styles;

[0128] A training module, configured to use the multiple sample speech data to train a preset encoder to obtain a preset latent space model.

[0129] In the embodiments of the present application, by acquiring sample speech data covering multiple styles, the comprehensiveness and diversity of the training data are ensured, laying a solid foundation for the learning of the encoder. These samples not only contain rich style features but also imply the diversity of speech content, helping the encoder to form a more refined representation in the latent space. Subsequently, using these multi-style samples to train the preset encoder can effectively prompt the encoder to learn the effective separation of style and content. In the latent space model, the content features are stably retained, while the style features show high variability, which provides great flexibility for subsequent style conversion. This training process not only enhances the model's ability to capture speech styles but also, through the construction of the latent space, enables the clear decoupling of style and content features, laying a key technical foundation for subsequent tasks of style transfer based on content features, thereby significantly improving the accuracy and naturalness of speech style conversion.

[0130] In one embodiment, the speech style reconstruction device 400 further includes:

[0131] A discriminator acquisition module, configured to obtain a pre-trained discriminator, input the source speech of the second style into the discriminator for discriminant analysis, and obtain discriminant information;

[0132] An adjustment module, configured to adjust the parameters of the latent space model and the parameters of the style decoder based on the discriminant information.

[0133] In the embodiment of the present application, by introducing a pre-trained discriminator, accurate evaluation of the speech after style conversion can be realized. This discriminator can not only judge the authenticity of the converted speech, ensure the integrity of the semantic content during the conversion process, but also accurately discriminate the style of the speech, so as to verify the effect of style conversion. Based on the discriminant information output by the discriminator, the parameters of the latent space model and the parameters of the style decoder can be accurately adjusted. This process forms a closed-loop feedback mechanism, enabling the model to continuously optimize itself and gradually improve the accuracy and naturalness of style conversion. Through this iterative optimization method, the problem of insufficient style decoupling is effectively solved, so that the converted speech can be closer to the second style while maintaining the semantic content unchanged, and can meet the high-quality requirements for speech style conversion in scenarios such as finance and medical care.

[0134] To solve the above technical problems, the embodiment of the present application also provides a computer device. For details, please refer to Figure 4 , Figure 4 which is the basic structural block diagram of the computer device in this embodiment.

[0135] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are communicatively connected to each other through a system bus. It should be noted that only the computer device 6 with a memory 61, a processor 62, and a network interface 63 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0136] The computer device can be a desktop computer, a notebook, a palm computer, a cloud server and other computing devices. The computer device can interact with the user through a keyboard, a mouse, a remote control, a touchpad or a voice control device and other means.

[0137] The memory 61 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk equipped on the computer device 6, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the memory 61 may also include both the internal storage unit and the external storage device of the computer device 6. In this embodiment, the memory 61 is generally used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions of the voice style reconstruction method, etc. In addition, the memory 61 may also be used to temporarily store various types of data that have been output or will be output.

[0138] In some embodiments, the processor 62 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 62 is generally used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to run the computer-readable instructions stored in the memory 61 or process data, such as running the computer-readable instructions of the voice style reconstruction method.

[0139] The network interface 63 may include a wireless network interface or a wired network interface, and the network interface 63 is generally used to establish a communication connection between the computer device 6 and other electronic devices.

[0140] In the embodiments of the present application, through a preset latent space model, the source speech can be accurately decomposed in terms of information, effectively distinguishing content feature information from style feature information. This characteristic greatly enhances the style decoupling property and solves the problem in the prior art that it is difficult to clearly separate style and content. Further, by skillfully combining preset random noise information with content feature information, the noise diffusion latent variable information of the source speech is generated. This not only enriches the diversity of speech representation but also significantly improves the conversion quality, enabling it to easily meet the nuanced service requirements in the financial scenario. Particularly importantly, during the style reconstruction process, relying only on the source speech latent variable information and style feature information greatly reduces the dependence on a large amount of training data. This characteristic significantly enhances the zero-shot style conversion ability, enabling efficient and accurate speech style conversion even in sensitive fields such as medical care when facing scarce training resources.

[0141] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor, so that at least one processor executes the steps of the method for reconstructing the speech style as described above.

[0142] In the embodiments of the present application, through a preset latent space model, the source speech can be accurately decomposed in terms of information, effectively distinguishing content feature information from style feature information. This characteristic greatly enhances the style decoupling property and solves the problem in the prior art that it is difficult to clearly separate style and content. Further, by skillfully combining preset random noise information with content feature information, the noise diffusion latent variable information of the source speech is generated. This not only enriches the diversity of speech representation but also significantly improves the conversion quality, enabling it to easily meet the nuanced service requirements in the financial scenario. Particularly importantly, during the style reconstruction process, relying only on the source speech latent variable information and style feature information greatly reduces the dependence on a large amount of training data. This characteristic significantly enhances the zero-shot style conversion ability, enabling efficient and accurate speech style conversion even in sensitive fields such as medical care when facing scarce training resources.

[0143] From the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0144] Obviously, the embodiments described above are only a part of the embodiments of the present application, rather than all of them. The preferred embodiments of the present application are shown in the drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments or equivalently replace some of the technical features. Any equivalent structure made by using the content of the specification and drawings of the present application, directly or indirectly applied in other related technical fields, is equally within the scope of the patent protection of the present application. The non-company enterprise software tools or components that appear in the embodiments of the present application are only introduced by way of example and do not represent actual use.

Claims

1. A method for reconstructing speech style, characterized in that: The steps include: Obtaining a source voice of a first style; Using a preset latent space model, the source speech is decomposed to obtain content feature information and style feature information of the source speech; Based on the preset random noise information and the content feature information, generating noise diffusion latent variable information of the source speech; Combining the noise diffusion latent variable information with the style feature information to obtain reconstructed source speech latent variable information; A preset style decoder is used to perform style reconstruction on the source speech latent variable information and the style feature information to obtain the source speech in a second style.

2. The method according to claim 1, characterized in that Before the step of using a preset latent space model to decompose the source speech to obtain content feature information and style feature information of the source speech, the method further includes: Obtain multiple sample speech data of different styles; The plurality of sample speech data are used to train a preset encoder to obtain a preset latent space model.

3. The method according to any one of claims 1 to 2, characterized in that: The step of using a preset latent space model to decompose the source speech to obtain content feature information and style feature information of the source speech specifically includes: Preprocessing the source speech to obtain preprocessed source speech; Inputting the preprocessed source speech into the latent space model, so as to map the preprocessed source speech to a latent space through the latent space model, and outputting content latent variables and style latent variables of the source speech; The content latent variable is determined as the content feature information of the source speech, and the style latent variable is determined as the style feature information of the source speech.

4. The method according to claim 1, characterized in that: The step of generating the noise diffusion latent variable information of the source speech based on the preset random noise information and the content feature information specifically includes: Acquire content dimension information of the content feature information; Generate a random noise vector whose dimension matches the content dimension information according to a preset noise distribution and the content dimension information; Performing standardization processing on the random noise vector to obtain preset random noise information; The random noise information and the content feature information are fused using a preset diffusion model to generate noise diffusion latent variable information of the source speech.

5. The method according to claim 4, characterized in that The step of combining the noise diffusion latent variable information and the style feature information to obtain the reconstructed source speech latent variable information specifically includes: Get the preset fusion strategy; Based on the fusion strategy, the noise diffusion latent variable information and the style feature information are fused to obtain fusion information after content and style fusion; The fusion information is input into the diffusion model for reverse diffusion processing to obtain the reconstructed source speech latent variable information.

6. The method according to claim 1, characterized in that The step of using a preset style decoder to perform style reconstruction on the source speech latent variable information and the style feature information to obtain the source speech of the second style specifically includes: Fusing the source speech latent variable information and the style feature information to obtain fused latent variable information; The latent variable information is input into a preset style decoder, and a speech signal is gradually generated through a neural network layer of the style decoder to obtain the source speech of the second style.

7. The method according to claim 1, characterized in that After the step of using a preset style decoder to reconstruct the style of the source speech latent variable information and the style feature information to obtain the source speech of the second style, the method further includes: Obtaining a pre-trained discriminator, inputting the source speech of the second style into the discriminator for discriminant analysis, and obtaining discriminant information; Based on the discriminant information, parameters of the latent space model and parameters of the style decoder are adjusted.

8. A device for reconstructing speech style, characterized in that: include: A speech acquisition module, used for acquiring a source speech of a first style; An information decomposition module, used to use a preset latent space model to perform information decomposition on the source speech to obtain content feature information and style feature information of the source speech; An information generating module, configured to generate noise diffusion latent variable information of the source speech based on preset random noise information and the content feature information; An information combining module, used for combining the noise diffusion latent variable information and the style feature information to obtain reconstructed source speech latent variable information; The style reconstruction module is used to use a preset style decoder to perform style reconstruction on the source speech latent variable information and the style feature information to obtain the source speech of the second style.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the method for reconstructing speech style according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the method for reconstructing a speech style according to any one of claims 1 to 7 are implemented.