Attribute interpolation method and device for realizing speech synthesis, equipment and medium

By constructing acoustic and content speech models and performing weighted average fusion, combined with attribute interpolation technology, the problem of poor fusion generalization ability of speech synthesis model is solved, high-quality and natural speech output is achieved, and the immersion of speech interaction is enhanced.

CN119964549APending Publication Date: 2025-05-09PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510243647.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, the generalization ability of the fusion of speech synthesis model is poor, and the smooth transition of speech synthesis cannot be accurately realized, resulting in inaccurate emotional expression of synthetic speech or unclear speaker characteristics.

Method used

By obtaining the voice data of the speaker, an acoustic voice model and content voice model are constructed, and the target voice fusion model is obtained using the weighted average fusion method, attribute interpolation is performed to generate interpolated voice data, and finally, a natural and smooth voice output is generated through the speech sequence extraction.

Benefits of technology

It improves the robustness and accuracy of the speech model, realizes a smooth transition of speech synthesis, and the generated speech output is more natural and smooth, enhancing the immersion and realism of speech interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964549A_ABST
    Figure CN119964549A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice processing, and discloses an attribute interpolation method and device for realizing voice synthesis, equipment and a medium, and the method comprises the steps: obtaining the voice data of a speaker, and constructing an acoustic voice model according to the voice acoustic features in the voice data; constructing a content voice model according to voice content characteristics in the voice data; performing weighted average fusion according to the fusion weights corresponding to the acoustic voice model and the content voice model to obtain a target voice fusion model; performing attribute interpolation on the voice data according to the target voice fusion model to obtain interpolated voice data; and performing voice sequence extraction on the interpolated voice data to obtain a voice feature sequence, and generating voice output corresponding to the interpolated voice data according to the voice feature sequence. The method can be applied to business scenes such as financial science and technology, medical health, old-age care and the like, and can effectively improve the generalization ability of model fusion and realize smooth transition of speech synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular to an attribute interpolation method, device, equipment and medium for realizing speech synthesis. Background Art

[0002] With the rapid development of deep learning technology, text-to-speech (TTS) systems have made remarkable achievements in speech synthesis, especially in speaker generation and emotion intensity control, and are now able to generate speech that is highly close to human naturalness.

[0003] In the prior art, there are many attribute smooth transition methods for speech synthesis based on model fusion, including methods based on speaker / emotion embedding and attribute interpolation. However, most traditional methods for speech synthesis based on model fusion usually rely on specific modules or complex training processes. The model may focus too much on the details of the classification task and ignore the essential universality of attribute interpolation, which leads to a series of problems such as increased demand for computing resources, high model fitting risk, and poor generalization ability.

[0004] For example, in medical application scenarios, speech synthesis technology can be used to convert medical reports, health science articles, etc. into speech, making it easier for patients to obtain information; in financial technology application scenarios, speech synthesis technology can realize the risk warning function in financial transactions. When users perform high-risk operations, risk warnings can be issued to users through speech synthesis technology to improve transaction security. However, the problems with traditional attribute interpolation methods may cause inaccurate emotional expression of synthesized speech or unclear speaker characteristics, causing trouble to patients. For example, when reading a medical report to a visually impaired patient, if the emotional intensity is not properly controlled, the key information and urgency of the report may not be accurately conveyed; or when simulating the voice of a financial institution's customer service for remote financial transaction guidance, if the speaker generation is inaccurate, it may reduce the user's trust in the transaction.

[0005] Therefore, how to improve the generalization ability of model fusion and achieve smooth transition of speech synthesis has become an urgent problem to be solved. Summary of the invention

[0006] The present invention provides an attribute interpolation method, device, equipment and medium for realizing speech synthesis, the main purpose of which is to solve the technical problems of poor generalization ability of model fusion and inability to accurately realize smooth transition of speech synthesis.

[0007] In a first aspect, to achieve the above-mentioned purpose, the present invention provides an attribute interpolation method for implementing speech synthesis, comprising:

[0008] Acquiring speech data of a speaker, and constructing an acoustic speech model according to speech acoustic features in the speech data;

[0009] Constructing a content speech model according to speech content features in the speech data;

[0010] Performing weighted average fusion according to the fusion weights corresponding to the acoustic speech model and the content speech model to obtain a target speech fusion model;

[0011] Performing attribute interpolation on the speech data according to the target speech fusion model to obtain interpolated speech data;

[0012] A speech sequence is extracted from the interpolated speech data to obtain a speech feature sequence, and a speech output corresponding to the interpolated speech data is generated according to the speech feature sequence.

[0013] In a second aspect, the present invention further provides a speech synthesis attribute interpolation device, comprising:

[0014] An acoustic model building module, used to obtain speech data of a speaker and build an acoustic speech model according to speech acoustic features in the speech data;

[0015] A content model building module, used to build a content speech model according to speech content features in the speech data;

[0016] A model fusion module, used to perform weighted average fusion according to the fusion weights corresponding to the acoustic speech model and the content speech model to obtain a target speech fusion model;

[0017] An attribute interpolation module, used for performing attribute interpolation on the speech data according to the target speech fusion model to obtain interpolated speech data;

[0018] The speech output module is used to extract the speech sequence of the interpolated speech data to obtain a speech feature sequence, and generate a speech output corresponding to the interpolated speech data according to the speech feature sequence.

[0019] In a third aspect, the present invention further provides an electronic device, the electronic device comprising:

[0020] at least one processor; and,

[0021] a memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned attribute interpolation method for realizing speech synthesis.

[0023] In a fourth aspect, the present invention further provides a computer-readable storage medium, wherein at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned attribute interpolation method for realizing speech synthesis.

[0024] The present invention obtains the speaker's voice data through professional audio acquisition equipment and other methods, ensuring the high quality of the voice data. By constructing a first basic voice model and a second basic voice model, the acoustic characteristics of the speaker can be deeply analyzed and understood, and the language content in the voice data can be parsed, providing a solid foundation for subsequent voice synthesis and improving the robustness and accuracy of the voice model. According to the weighted average fusion method, the weights are accurately calculated to ensure that the contribution of each basic model in the fusion process is reasonably reflected, avoiding the limitations that may exist in a single model, improving the overall performance and adaptability of the fusion model, and providing a more intelligent and natural solution for subsequent application scenarios such as voice interaction and voice control. Through attribute interpolation, parameter changes between voice data can be effectively smoothed to avoid sudden changes in sound quality, thereby improving the sound quality of the synthesized voice. According to the interpolated voice data, the target voice fusion model can generate a more natural and fluent voice output, and combined with the position-coded voice feature extraction method, the model can more accurately capture the emotional color in the voice, enhance the expressiveness of the voice output, and enhance the immersion and reality of the voice interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0026] Figure 1 It is a schematic diagram of an application environment of an attribute interpolation method for implementing speech synthesis in one embodiment of the present invention;

[0027] Figure 2 A flowchart of a method for implementing attribute interpolation for speech synthesis provided by an embodiment of the present invention;

[0028] Figure 3 A schematic diagram of a module of a speech synthesis attribute interpolation device provided by an embodiment of the present invention;

[0029] Figure 4 A schematic diagram of the structure of an electronic device for implementing a method for attribute interpolation for speech synthesis provided by an embodiment of the present invention;

[0030] Figure 5 It is another structural schematic diagram of an electronic device for implementing an attribute interpolation method for speech synthesis provided by an embodiment of the present invention.

[0031] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0032] In order to enable those skilled in the art to better understand the technical solution of the present disclosure, and to fully understand and implement how the present disclosure applies technical means to solve technical problems and achieve the corresponding technical effects, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only embodiments of a part of the present disclosure, not all of the embodiments. The embodiments of the present disclosure and the various features in the embodiments can be combined with each other without conflict, and the technical solutions formed are all within the scope of protection of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present disclosure.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or equipment that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0034] The embodiment of the present application provides a method for implementing attribute interpolation of speech synthesis, and the execution subject of the method for implementing attribute interpolation of speech synthesis includes but is not limited to at least one of the electronic devices such as a server and a terminal that can be configured to execute the device provided by the embodiment of the present application. In other words, the method for implementing attribute interpolation of speech synthesis can be executed by software or hardware installed on a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms.

[0035] The present invention provides an attribute interpolation method for realizing speech synthesis, which can be applied in the following aspects: Figure 1 In the application environment. Among them, the client communicates with the server through the network. The server can obtain the speaker's voice data through the client, ensuring the high quality of the voice data. By constructing the first basic voice model and the second basic voice model, the acoustic characteristics of the speaker can be deeply analyzed and understood, and the language content in the voice data can be parsed, providing a solid foundation for subsequent voice synthesis and improving the robustness and accuracy of the voice model; according to the weighted average fusion method, the weight is accurately calculated to ensure that the contribution of each basic model in the fusion process is reasonably reflected, avoiding the limitations that may exist in a single model, improving the overall performance and adaptability of the fusion model, and providing a more intelligent and natural solution for subsequent application scenarios such as voice interaction and voice control; through attribute interpolation, the parameter changes between voice data can be effectively smoothed to avoid sudden changes in sound quality, thereby improving the sound quality of the synthesized voice; according to the interpolated voice data, the target voice fusion model can generate a more natural and fluent voice output, combined with the position-encoded voice feature extraction method, so that the model can more accurately capture the emotional color in the voice, enhance the expressiveness of the voice output, and enhance the immersion and realism of the voice interaction, and feed the final voice output back to the client. The client may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. The server may be implemented by an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0036] The following is an explanation of the specification of the present invention. The present invention adopts a method based on attribute interpolation as a baseline, and adjusts model parameters by linear interpolation of speech data so that the smooth transition effect between different speaker features can be evaluated, and proposes a simple and efficient method, namely weighted average model fusion. This method achieves model fusion by performing weighted averaging between two pre-trained basic speech models. Specifically, the attributes of the generated speech can be flexibly controlled by adjusting the fusion coefficient or fusion weight, such as creating a new speaker or adjusting the emotion intensity, so that a smooth transition of attributes can be achieved while maintaining the language content, thereby generating high-quality speech output.

[0037] Reference Figure 2 FIG. 1 is a flow chart of a method for implementing attribute interpolation for speech synthesis according to an embodiment of the present invention. In this embodiment, the method for implementing attribute interpolation for speech synthesis includes:

[0038] S1. Acquire speech data of a speaker, and construct an acoustic speech model according to speech acoustic features in the speech data.

[0039] In an embodiment of the present invention, professional audio acquisition equipment and other methods can be used to accurately capture the speaker's voice clips in various scenarios, intonations, and speaking speeds, and an acoustic speech model can be constructed based on these rich voice data. The acoustic speech model focuses on the basic features of speech, such as pitch, loudness, timbre, etc., to form a basic acoustic model framework.

[0040] In the embodiment of the present invention, the voice data of the speaker can be obtained by directly recording the speaker's voice using a microphone, a voice recorder or other recording equipment, or by using recording software (such as Audacity, Quick Video Converter, etc.) on a computer or mobile device to record the speaker's voice data.

[0041] Furthermore, programming languages ​​such as Python and Java can be used to connect to a preset voice database through the HTTP protocol, and an HTTP request can be sent to obtain the voice data of the speaker at any time within the reference time period, so as to obtain the voice data, thereby ensuring that the obtained voice data is clear and accurate.

[0042] For example, in the voice data acquisition scenario in the medical and health field, doctors and researchers often need to use voice data to analyze the health status of patients. However, the voice data of different patients may show large differences due to different recording equipment, environmental noise and speaking habits. By performing data denoising and data enhancement operations on the original voice data, a variety of voice samples with different acoustic characteristics can be created, such as voice data under different volumes, different speech speeds or different background noise conditions. These enhanced voice samples can assist the preset machine learning model to better capture voice features, improve the accuracy of diagnosis or analysis tasks, and reduce the burden on doctors when manually screening and sorting voice data. For example, in voice pathology analysis, the original voice data may be a patient's cough with respiratory sound interference. Through data enhancement operations, cough samples with reduced volume, faster speech speed or different background noise can be generated. These enhanced samples help the model to more accurately identify the pathological features in the cough sound, and even in the actual diagnosis, the key voice information can be effectively identified when the volume is weak, the speech speed changes or the complex noise environment is encountered. This approach reduces the dependence on large amounts of high-quality original speech data and significantly enhances the recognition robustness of the model under different recording conditions.

[0043] Similarly, in the voice data acquisition scenario in the financial technology field, financial institutions need to verify the authenticity of user voices to prevent fraudsters from impersonating or committing fraud by repeatedly recording or tampering with voice content. Due to the large differences in recording equipment and environmental conditions among different users, voice data may have different sound quality, volume, background noise, etc., and fraudsters may even modify voice features through voice changing or editing. For example, fraudsters may imitate the user's voice or record voice in different environments in an attempt to deceive the voice model through these voice differences.

[0044] By performing relevant data processing on the original speech data set, the ability of the speech model to distinguish highly similar speech can be effectively improved. For example, the model can perform enhancement processing such as speed change, pitch change, and adding background noise on the user's original speech to generate a variety of enhanced speech with different feature changes. These enhanced speech can help the model learn the changes in the user's speech features under different recording conditions, thereby improving the generalization ability of the verification system.

[0045] In a specific application, suppose that a user's original voice is enhanced to generate a variety of enhanced voices under different conditions after data enhancement. The system can mark these voices as similar voices through the comparative learning module, and automatically identify whether the voice submitted by the user is a repeated recording of the same person during the actual voice verification process. In this way, even if fraudsters try to imitate or record voices in different environments, the model can recognize the similarities of these voices and effectively prevent identity theft and fraud. Through data enhancement and voice feature extraction technology, financial institutions can greatly improve the accuracy and timeliness of voice verification without increasing labor costs, thereby realizing automated anti-fraud and reducing the risks of financial institutions.

[0046] In the embodiment of the present invention, the step of constructing an acoustic speech model according to the speech acoustic features in the speech data includes:

[0047] Performing noise reduction preprocessing on the speech data to obtain noise-reduced speech data;

[0048] Extracting acoustic features of the noise-reduced speech data to obtain speech acoustic features of the noise-reduced speech data;

[0049] Determining a first model node and a first connection relationship between nodes according to the speech acoustic feature;

[0050] The first model nodes are concatenated according to the first connection relationship to obtain an acoustic speech model.

[0051] Among them, the noise reduction preprocessing refers to the noise reduction of the basic acoustic feature data of the speech data, that is, removing the background noise in the speech data to improve the clarity of the speech, which can be achieved through Gaussian filtering; in detail, the acoustic speech features obtained by extracting the noise reduction speech data features include pitch (fundamental frequency), that is, the frequency of the periodic components in the speech data, which is related to the speaker's tone, and sound intensity (amplitude), that is, the intensity or loudness of the speech data, which is related to the speaker's volume.

[0052] Furthermore, the first connection relationship refers to a connection established between first model nodes, where the first model nodes represent first feature vectors in the first speech data and are established based on similarity and correlation between the first feature vectors.

[0053] In an actual application scenario of the present invention, the first connection relationship may reflect the intrinsic connection between the acoustic features in the speech data. For example, the feature of the first node may represent the volume of the speech, and the feature of the second node may represent the pitch of the speech. If the two features show a certain correlation in multiple speech samples (such as the pitch increases when the volume increases), then a first connection relationship can be established between the first node and the second node.

[0054] In the embodiment of the present invention, the voice data is subjected to noise reduction processing. In a medical environment, background noise may come from various medical devices, patient activities, personnel conversations, etc. The noise may interfere with the clarity of the voice data and affect subsequent analysis and processing. The noise removal methods such as Gaussian filtering are used to effectively perform directional noise reduction, thereby improving the clarity of the voice.

[0055] S2. Constructing a content speech model according to speech content features in the speech data.

[0056] In the embodiment of the present invention, the content speech model focuses on the language content level, including parsing vocabulary, sentence structure, etc., and converting speech into understandable text information.

[0057] In the embodiment of the present invention, the step of constructing a content speech model according to speech content features in the speech data includes:

[0058] Performing text preprocessing on the voice data to obtain text voice data;

[0059] Extracting content features from the text and speech data to obtain speech content features of the text and speech data;

[0060] Determine a second model node and a second connection relationship between nodes according to the speech content feature;

[0061] The second model nodes are concatenated according to the second connection relationship to obtain a content speech model.

[0062] Among them, the text preprocessing refers to the text preprocessing of the language content level of the voice data, which aims to convert the voice signal into understandable text information, including voice segmentation, etc., dividing the continuous voice data into smaller voice units, such as syllables, words, etc., which helps to perform more detailed analysis of the voice data later and provide support for subsequent model building, synthesis and other tasks.

[0063] In an actual application scenario of the present invention, segmenting continuous speech data into smaller speech units such as syllables, words, etc. helps to perform more detailed analysis of the speech data later. In medical applications, it can help doctors more accurately identify key information such as patients' symptom descriptions and medical advice.

[0064] In detail, the content speech features obtained by extracting features from text speech data include vocabulary features, etc., which represent acoustic or semantic features of vocabulary in speech data and can be used to identify different vocabulary to provide data support for subsequent speech synthesis.

[0065] The second connection relationship refers to the connection established between the second model nodes, the second model nodes represent the second feature vectors in the second speech data, and the second connection relationship is established based on the semantic relationship and contextual connection between the second feature vectors.

[0066] In an actual application scenario of the present invention, the second connection relationship may reflect the logical relationship between the language content in the voice data. For example, the first node may represent the patient's symptom description, and the second node may represent the doctor's diagnosis or treatment suggestions based on the symptoms. The first node and the second node show a certain causal relationship in multiple medical conversations, such as specific symptoms leading to specific diagnoses, and a second connection relationship may be established between the first node and the second node.

[0067] The first model node and the second model node may be an acoustic speech model node and a content speech model node constructed subsequently.

[0068] The present invention splices the second model nodes according to the first connection relationship to construct a complete basic speech model, namely, a content speech model.

[0069] In the embodiment of the present invention, the speaker's voice data is obtained by means of professional audio acquisition equipment and other methods, thereby ensuring the high quality of the voice data. By constructing an acoustic first basic voice model and a content-based second basic voice model, the speaker's acoustic characteristics can be deeply analyzed and understood, and the language content in the voice data can be parsed, providing a solid foundation for subsequent speech synthesis and improving the robustness and accuracy of the voice model.

[0070] S3. Perform weighted average fusion according to the fusion weights corresponding to the acoustic speech model and the content speech model to obtain a target speech fusion model.

[0071] In an embodiment of the present invention, the acoustic first basic speech model and the content-based second basic speech model are fused according to the weighted average fusion method, aiming to integrate the advantages of both, thereby obtaining a target speech fusion model with better performance and stronger adaptability, so as to ensure that the target speech fusion model can accurately reflect the various speech features of the speech data.

[0072] In the embodiment of the present invention, the step of performing weighted average fusion according to the fusion weights corresponding to the acoustic speech model and the content speech model to obtain a target speech fusion model includes:

[0073] Obtaining a fusion weight of the acoustic speech model and the content speech model;

[0074] Extracting key speech features of the acoustic speech model and the content speech model respectively;

[0075] Performing weighted addition on the key speech features according to the fusion weights to obtain fused speech features;

[0076] A target speech fusion model is constructed according to the fused speech features and preset fusion coefficients.

[0077] In detail, the fusion weight refers to the weight of the acoustic speech model and the content speech model; the present invention can express the contribution of the acoustic speech model and the content speech model in the target speech fusion model by controlling the fusion coefficient. The value range of the fusion coefficient can be [0,1]. When the fusion coefficient is 0, the relevant attribute characteristics of the target speech fusion model are equivalent to the content speech model. When the fusion coefficient is 1, the relevant attribute characteristics of the target speech fusion model are equivalent to the acoustic speech model. When the fusion coefficient takes the average value, the relevant attribute characteristics of the target speech fusion model are between the attribute characteristics of the acoustic speech model and the content speech model.

[0078] In detail, the target speech fusion model is as follows:

[0079] θ ab =α*θ a +(1-α)*θ b

[0080] Among them, θ ab represents the model construction parameters of the target speech fusion model, α represents the preset fusion coefficient, θ a represents the fusion weight of the acoustic first basic speech model, θ b Represents the fusion weight of the second basic speech model of the content learning.

[0081] The model building parameters are used to build specific parameter values ​​of the acoustic speech model and the content speech model.

[0082] In detail, key speech features related to the acoustic characteristics of the acoustic speech model and key speech features related to the language content analysis are extracted from the acoustic speech model and the content speech model respectively, and for each extracted key speech feature, point-by-point weighted addition is performed according to its corresponding fusion weight, that is, for each feature point, the new value of the feature point in the target speech fusion model is calculated according to the weight, so as to obtain the fused speech feature, and the target speech fusion model is constructed according to the fused speech feature, which can more accurately capture and process the acoustic characteristics and language content in the speech signal, thereby showing higher performance in tasks such as speech synthesis.

[0083] Example description: In the speech model fusion scenario in the field of medical health, especially in the process of medical consultation and diagnosis, doctors can use speech model fusion technology to convey key information such as treatment plans, drug dosages, and condition descriptions to patients in the form of voice, so that patients can have a clearer understanding of their condition and treatment plans. Instead of using a single language model, the models of related functions such as medical voice assistants and voice medical records can be fused, so that the target fusion model can not only recognize the patient's voice commands, and use speech synthesis technology to provide disease advice, appointment registration, query test results and other services to improve the convenience and efficiency of medical services, but also quickly record the patient's medical record information through speech recognition technology to improve the accuracy and efficiency of medical record records.

[0084] In the embodiment of the present invention, the weights are accurately calculated according to the weighted average fusion method to ensure that the contribution of each basic model in the fusion process is reasonably reflected, thereby avoiding the limitations that may exist in a single model and improving the overall performance and adaptability of the fusion model, so that it can accurately understand and respond to various voice commands, providing a more intelligent and natural solution for subsequent application scenarios such as voice interaction and voice control.

[0085] S4. Perform attribute interpolation on the speech data according to the target speech fusion model to obtain interpolated speech data.

[0086] In an embodiment of the present invention, the preset weight coefficients of the speech data are used as interpolation coefficients, and the interpolation coefficients and the target speech fusion model are used to perform linear interpolation processing on the original speech data. The target speech fusion model provides key guidance for the interpolation process, and the interpolated speech data is estimated more accurately by analyzing the acoustic characteristics and language content of the speech data.

[0087] Among them, the interpolated speech data not only retains the main features of the original speech data, but also enhances the integrity and continuity of the speech data to a certain extent, improves the processing efficiency and accuracy of the speech data, and provides more reliable data support for subsequent tasks such as speech synthesis.

[0088] In the embodiment of the present invention, the method of performing attribute interpolation on the speech data according to the target speech fusion model to obtain interpolated speech data includes:

[0089] Acquire first voice data and second voice data from the voice data;

[0090] The target speech fusion model is used to perform linear interpolation on the first speech data and the second speech data according to preset interpolation coefficients to obtain interpolated speech data.

[0091] In detail, linear interpolation can be performed using the following formula:

[0092] F AB =ρ*F A +(1-ρ)*F B

[0093] Among them, F AB represents the interpolated speech data, ρ represents the preset interpolation coefficient, F A represents the first voice data, F B Represents the second voice data.

[0094] Among them, the interpolation coefficient represents the weight of controlling the basic speaker voice data in generating the interpolated voice data. The value range of the interpolation coefficient can be [0,1]. When the interpolation coefficient takes the value of 0 and 1, the interpolated voice data is equivalent to the first voice data and the second voice data. When the interpolation coefficient takes the average value, the interpolated voice data is between the attribute characteristics of the first voice data and the second voice data.

[0095] Furthermore, in speech processing, the linear interpolation is generally used to smooth the attribute changes of speech data, such as pitch, volume, etc. The attributes of the speech data are analyzed according to the target speech fusion model, and the interpolation coefficient is determined. The interpolation coefficient is a preset weight coefficient of the speech data. Between adjacent known speech data, the value of the unknown data point is estimated according to the interpolation coefficient and the linear interpolation formula, and this is used as the attribute of the interpolated speech data to generate interpolated speech data, so that the speech data is smoother and more continuous in attributes, thereby improving the overall quality of the speech data.

[0096] In the embodiment of the present invention, through attribute interpolation, parameter changes between speech data can be effectively smoothed to avoid sudden changes in sound quality, thereby improving the sound quality of synthesized speech; interpolation technology can fill in gaps or sparse parts in speech data, making the data more continuous and complete, which is crucial for subsequent language output and helps to extract speech features more accurately.

[0097] S5. Perform voice sequence extraction on the interpolated voice data to obtain a voice feature sequence, and generate a voice output corresponding to the interpolated voice data according to the voice feature sequence.

[0098] In an embodiment of the present invention, speech sequence extraction is performed on the interpolated speech data to obtain relevant speech feature vectors, and speech data matching the input features is quickly and accurately generated based on the speech feature vectors, thereby realizing the conversion from the interpolated speech data to the corresponding speech output, and providing a new method for the development of speech processing technology.

[0099] In the embodiment of the present invention, the step of extracting a speech sequence from the interpolated speech data to obtain a speech feature sequence includes:

[0100] Segment-processing the interpolated voice data to obtain a plurality of voice segment data corresponding to the interpolated voice data;

[0101] Extracting intensity features from the plurality of speech segment data to obtain a speech intensity feature sequence;

[0102] Determine the speech emotion feature sequence corresponding to the interpolated speech data according to the speech intensity feature sequence and the preset position code,

[0103] The speech intensity feature sequence and the speech emotion feature sequence are used as speech feature sequences.

[0104] The position code is used to represent the position information of the speech segment data in the interpolated speech data.

[0105] Among them, the speech feature sequence includes a speech intensity feature sequence and a speech emotion feature sequence. The speech intensity feature sequence includes the pitch and volume features of the interpolated speech data, and the speech emotion feature sequence includes the emotional color of the interpolated speech data, such as joy, sadness, etc. According to the speech feature sequence, a corresponding speech waveform is generated through a preset synthesis algorithm such as waveform splicing, parameter synthesis, etc., thereby generating a speech output corresponding to the interpolated speech data.

[0106] In detail, the interpolated speech data is usually a continuous speech signal. In order to perform subsequent feature extraction and processing, it needs to be cut into multiple shorter speech segment data. The speech signal can be smoothed during the segmentation process through a window function such as a rectangular window to reduce the influence of spectrum leakage and fence effect. By extracting the intensity feature of each speech segment, it can be obtained by calculating the short-time energy, that is, calculating the sum of the squares of the speech signal in a short time. The size of the short-time energy is proportional to the volume of the speech, so it can be used to extract the intensity feature of the interpolated speech data.

[0107] Furthermore, the position code is used to characterize the position information of each speech segment data in the interpolated speech data. Different positions may correspond to different emotional expressions. Feature splicing and deep learning are performed based on the position code and the speech intensity feature sequence to obtain the speech emotion feature sequence of the corresponding position. The speech intensity feature sequence and the speech emotion feature sequence are combined to form a complete speech feature sequence.

[0108] In the embodiment of the present invention, the step of generating a speech output corresponding to the interpolated speech data according to the speech feature sequence includes:

[0109] Performing feature standardization on the speech feature sequence to obtain target speech features;

[0110] Performing waveform synthesis on the target speech feature to obtain a speech waveform;

[0111] The speech waveform is subjected to audio conversion to obtain a speech output corresponding to the interpolated speech data.

[0112] In detail, the feature standardization processing of the speech feature sequence is to convert the speech feature sequence into a target speech feature with a uniform scale and distribution, and the feature standardization includes data normalization, denoising, smoothing and other processing steps to ensure the quality of input data of the subsequent speech model. The speech waveform synthesis according to the target speech feature is to convert the feature sequence into waveform data with natural speech characteristics, and convert the generated speech waveform data into a common audio format (such as WAV, MP3, etc.) for storage and playback, so as to obtain the speech output corresponding to the interpolated speech data.

[0113] Example description: The speech feature sequence can be the intensity feature and emotional feature of the doctor and patient's speech data. The doctor can issue medical advice to the preset hospital management system through voice. After extracting the speech feature sequence, the hospital management system uses advanced speech recognition and natural language processing technology to convert the speech feature sequence containing intensity features and emotional features into accurate text medical advice, just like a professional translator, converting the doctor's verbal instructions into clear and accurate written text to ensure that every detail is accurately recorded. The text medical advice will be quickly conveyed to the patient or caregiver. For patients, they can understand their condition, treatment plan and matters that need attention more clearly by reading the text medical advice. For caregivers, they can provide patients with more accurate and thoughtful nursing services according to the requirements of the medical advice.

[0114] The present invention not only greatly improves the level of intelligence and efficiency of medical services, but also provides patients with more intimate and personalized medical consultation and condition explanation services. Patients no longer need to worry about missing important instructions from doctors due to hearing problems or language barriers, nor do they need to be anxious about waiting for written medical orders. On the contrary, they can receive treatment with greater peace of mind and enjoy more convenient and efficient medical services, providing strong support for the intelligence and efficiency of medical services.

[0115] Similarly, in the field of financial technology, with the acceleration of digital transformation, speech synthesis and speech output technology is gradually becoming a key tool to improve user experience and enhance the level of service intelligence. By simulating the natural fluency and emotional expression of human speech, financial services can transcend the traditional interface interaction limitations and achieve more intuitive, convenient and humanized information transmission. Whether it is the voice response of the intelligent customer service system, the voice introduction of financial products, or the voice broadcast of personalized financial advice, speech synthesis technology has greatly enriched the interactive mode of financial services, not only improving the user's interactive experience, but also bringing financial institutions a broader innovation space and market competitive advantage.

[0116] In an embodiment of the present invention, based on the interpolated speech data, the target speech fusion model can generate a more natural and fluent speech output. Combined with the position-coded speech feature extraction method, the model can more accurately capture the emotional color in the speech, enhance the expressiveness of the speech output, and improve the immersion and realism of the speech interaction.

[0117] The present invention obtains the speaker's voice data through professional audio acquisition equipment and other methods, ensuring the high quality of the voice data. By constructing an acoustic voice model and a content voice model, the acoustic characteristics of the speaker can be deeply analyzed and understood, and the language content in the voice data can be parsed, providing a solid foundation for subsequent voice synthesis and improving the robustness and accuracy of the voice model. According to the weighted average fusion method, the weights are accurately calculated to ensure that the contribution of each basic model in the fusion process is reasonably reflected, avoiding the limitations that may exist in a single model, improving the overall performance and adaptability of the fusion model, and providing a more intelligent and natural solution for subsequent application scenarios such as voice interaction and voice control. Through attribute interpolation, parameter changes between voice data can be effectively smoothed to avoid sudden changes in sound quality, thereby improving the sound quality of the synthesized voice. According to the interpolated voice data, the target voice fusion model can generate a more natural and fluent voice output, and combined with the position-coded voice feature extraction method, the model can more accurately capture the emotional color in the voice, enhance the expressiveness of the voice output, and enhance the immersion and reality of the voice interaction.

[0118] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0119] like Figure 3 , which is a functional module diagram of a speech synthesis attribute interpolation device provided by an embodiment of the present invention.

[0120] In an embodiment of the present disclosure, a speech synthesis attribute interpolation device is provided, and the speech synthesis attribute interpolation device corresponds one-to-one to the attribute interpolation method for implementing speech synthesis in the above embodiment. Figure 3 As shown, the speech synthesis attribute interpolation device 100 can be installed in an electronic device. According to the functions to be implemented, the speech synthesis attribute interpolation device 100 includes an acoustic model construction module 101, a content model construction module 102, a model fusion module 103, an attribute interpolation module 104, and a speech output module 105. The functional modules are described in detail as follows:

[0121] The acoustic model building module 101 is used to obtain speech data of a speaker and build an acoustic speech model according to speech acoustic features in the speech data;

[0122] A content model building module 102, configured to build a content speech model according to speech content features in the speech data;

[0123] A model fusion module 103, configured to perform weighted average fusion according to the fusion weights corresponding to the acoustic speech model and the content speech model to obtain a target speech fusion model;

[0124] An attribute interpolation module 104, configured to perform attribute interpolation on the speech data according to the target speech fusion model to obtain interpolated speech data;

[0125] The speech output module 105 is used to extract a speech sequence from the interpolated speech data to obtain a speech feature sequence, and generate a speech output corresponding to the interpolated speech data according to the speech feature sequence.

[0126] In one embodiment, when executing the construction of the acoustic speech model according to the speech acoustic features in the speech data, the acoustic model construction module 101 is used to:

[0127] Performing noise reduction preprocessing on the speech data to obtain noise-reduced speech data;

[0128] Extracting acoustic features of the noise-reduced speech data to obtain speech acoustic features of the noise-reduced speech data;

[0129] Determining a first model node and a first connection relationship between nodes according to the speech acoustic feature;

[0130] The first model nodes are concatenated according to the first connection relationship to obtain an acoustic speech model.

[0131] In one embodiment, when executing the step of constructing a content speech model according to speech content features in the speech data, the content model construction module 102 is configured to:

[0132] Performing text preprocessing on the voice data to obtain text voice data;

[0133] Extracting content features from the text and speech data to obtain speech content features of the text and speech data;

[0134] Determine a second model node and a second connection relationship between nodes according to the speech content feature;

[0135] The second model nodes are concatenated according to the second connection relationship to obtain a content speech model.

[0136] In one embodiment, when the model fusion module 103 performs weighted average fusion according to the fusion weights corresponding to the acoustic speech model and the content speech model to obtain the target speech fusion model, it is used to:

[0137] Obtaining a fusion weight of the acoustic speech model and the content speech model;

[0138] Extracting key speech features of the acoustic speech model and the content speech model respectively;

[0139] Performing weighted addition on the key speech features according to the fusion weights to obtain fused speech features;

[0140] A target speech fusion model is constructed according to the fused speech features and preset fusion coefficients.

[0141] In one embodiment, when the attribute interpolation module 104 performs attribute interpolation on the speech data according to the target speech fusion model to obtain interpolated speech data, it is used to:

[0142] Acquire first voice data and second voice data from the voice data;

[0143] The target speech fusion model is used to perform linear interpolation on the first speech data and the second speech data according to preset interpolation coefficients to obtain interpolated speech data.

[0144] In one embodiment, when the speech output module 105 extracts the speech sequence from the interpolated speech data to obtain the speech feature sequence, it is used to:

[0145] Segment-processing the interpolated voice data to obtain a plurality of voice segment data corresponding to the interpolated voice data;

[0146] Extracting intensity features from the plurality of speech segment data to obtain a speech intensity feature sequence;

[0147] Determine the speech emotion feature sequence corresponding to the interpolated speech data according to the speech intensity feature sequence and the preset position code,

[0148] The speech intensity feature sequence and the speech emotion feature sequence are used as speech feature sequences.

[0149] In one embodiment, when the speech output module 105 generates the speech output corresponding to the interpolated speech data according to the speech feature sequence, it is used to:

[0150] Performing feature standardization on the speech feature sequence to obtain target speech features;

[0151] Performing waveform synthesis on the target speech feature to obtain a speech waveform;

[0152] The speech waveform is subjected to audio conversion to obtain a speech output corresponding to the interpolated speech data.

[0153] In the present invention, for a speech synthesis attribute interpolation device, first, the present invention obtains the speaker's speech data through professional audio acquisition equipment and other methods, thereby ensuring the high quality of the speech data. By constructing a first basic speech model and a second basic speech model, the speaker's acoustic characteristics can be deeply analyzed and understood, and the language content in the speech data can be parsed, providing a solid foundation for subsequent speech synthesis and improving the robustness and accuracy of the speech model; according to the weighted average fusion method, the weight is accurately calculated to ensure that the contribution of each basic model in the fusion process is reasonably reflected, avoiding the limitations that may exist in a single model, improving the overall performance and adaptability of the fusion model, and providing a more intelligent and natural solution for subsequent application scenarios such as speech interaction and speech control; through attribute interpolation, the parameter changes between speech data can be effectively smoothed to avoid sudden changes in sound quality, thereby improving the sound quality of the synthesized speech; according to the interpolated speech data, the target speech fusion model can generate a more natural and fluent speech output, and combined with the position-coded speech feature extraction method, the model can more accurately capture the emotional color in the speech, enhance the expressiveness of the speech output, and enhance the immersion and reality of the speech interaction. For the specific definition of a speech synthesis attribute interpolation device, please refer to the definition of an attribute interpolation method for realizing speech synthesis above, which will not be repeated here. Each module in the above-mentioned data visualization configuration device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0154] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements a function or step on the server side of an attribute interpolation method for realizing speech synthesis.

[0155] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements a function or step on the client side of an attribute interpolation method for realizing speech synthesis.

[0156] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:

[0157] Acquiring speech data of a speaker, and constructing an acoustic speech model according to speech acoustic features in the speech data;

[0158] Constructing a content speech model according to speech content features in the speech data;

[0159] Performing weighted average fusion according to the fusion weights corresponding to the acoustic speech model and the content speech model to obtain a target speech fusion model;

[0160] Performing attribute interpolation on the speech data according to the target speech fusion model to obtain interpolated speech data;

[0161] A speech sequence is extracted from the interpolated speech data to obtain a speech feature sequence, and a speech output corresponding to the interpolated speech data is generated according to the speech feature sequence.

[0162] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and apparatuses can be implemented in other ways. For example, the system embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0163] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.

[0164] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.

[0165] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0166] In some implementations of this embodiment, a computer-readable storage medium is provided, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method described in the above embodiment are implemented.

[0167] The readable storage medium of the present invention stores a computer program, and when the computer program is executed by a processor of an electronic device, the computer program can achieve:

[0168] Acquiring speech data of a speaker, and constructing an acoustic speech model according to speech acoustic features in the speech data;

[0169] Constructing a content speech model according to speech content features in the speech data;

[0170] Performing weighted average fusion according to the fusion weights corresponding to the acoustic speech model and the content speech model to obtain a target speech fusion model;

[0171] Performing attribute interpolation on the speech data according to the target speech fusion model to obtain interpolated speech data;

[0172] A speech sequence is extracted from the interpolated speech data to obtain a speech feature sequence, and a speech output corresponding to the interpolated speech data is generated according to the speech feature sequence.

[0173] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0174] The computer-readable storage medium may also store at least one computer executable program / instruction, which may be, for example, a computer-readable instruction. The computer-readable storage medium includes, but is not limited to, for example, a volatile memory and / or a non-volatile memory. The volatile memory may include, for example, a random access memory (RAM) and / or a cache memory (cache), etc. The computer-readable storage medium may include, for example, a read-only memory (ROM), a hard disk, a flash memory, etc. For example, a non-transitory computer-readable storage medium may be connected to a computing device such as a computer, and then, when the computing device runs the computer-readable instructions stored on the computer-readable storage medium, the various methods described above may be performed.

[0175] In addition, the computer device may also include (but not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (eg, keyboard, mouse, speaker, etc.), etc.

[0176] The processor may communicate with external devices via an I / O bus via a wired or wireless network.

[0177] In one embodiment, the at least one computer executable instruction may also be compiled into or constitute a software product / computer program product, wherein one or more computer executable instructions are executed by a processor to perform the various functions and / or method steps in the embodiments described in the present technology.

[0178] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0179] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0180] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0181] It should be noted that in the present disclosure, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element limited by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0182] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

[0183] It should be noted that if software tools or components other than those of our company appear in the embodiments of the present application, they are only used for illustrative purposes and do not represent actual use.

Claims

1. A method for implementing attribute interpolation for speech synthesis, characterized in that: The method comprises: Acquiring speech data of a speaker, and constructing an acoustic speech model according to speech acoustic features in the speech data; Constructing a content speech model according to speech content features in the speech data; Performing weighted average fusion according to the fusion weights corresponding to the acoustic speech model and the content speech model to obtain a target speech fusion model; Performing attribute interpolation on the speech data according to the target speech fusion model to obtain interpolated speech data; A speech sequence is extracted from the interpolated speech data to obtain a speech feature sequence, and a speech output corresponding to the interpolated speech data is generated according to the speech feature sequence.

2. The attribute interpolation method for realizing speech synthesis according to claim 1, characterized in that: The step of constructing an acoustic speech model according to the speech acoustic features in the speech data comprises: Performing noise reduction preprocessing on the speech data to obtain noise-reduced speech data; Extracting acoustic features of the noise-reduced speech data to obtain speech acoustic features of the noise-reduced speech data; Determining a first model node and a first connection relationship between nodes according to the speech acoustic feature; The first model nodes are concatenated according to the first connection relationship to obtain an acoustic speech model.

3. The attribute interpolation method for realizing speech synthesis according to claim 1, characterized in that: The step of constructing a content speech model according to speech content features in the speech data comprises: Performing text preprocessing on the voice data to obtain text voice data; Extracting content features from the text and speech data to obtain speech content features of the text and speech data; Determine a second model node and a second connection relationship between nodes according to the speech content feature; The second model nodes are concatenated according to the second connection relationship to obtain a content speech model.

4. The attribute interpolation method for realizing speech synthesis according to claim 1, characterized in that: The step of performing weighted average fusion according to the fusion weights corresponding to the acoustic speech model and the content speech model to obtain a target speech fusion model includes: Obtaining a fusion weight of the acoustic speech model and the content speech model; Extracting key speech features of the acoustic speech model and the content speech model respectively; Performing weighted addition on the key speech features according to the fusion weights to obtain fused speech features; A target speech fusion model is constructed according to the fused speech features and preset fusion coefficients.

5. The attribute interpolation method for realizing speech synthesis according to claim 1, characterized in that: The performing attribute interpolation on the speech data according to the target speech fusion model to obtain interpolated speech data includes: Acquire first voice data and second voice data from the voice data; The target speech fusion model is used to perform linear interpolation on the first speech data and the second speech data according to preset interpolation coefficients to obtain interpolated speech data.

6. The attribute interpolation method for realizing speech synthesis according to claim 1, characterized in that: The step of extracting a speech sequence from the interpolated speech data to obtain a speech feature sequence comprises: Segment-processing the interpolated voice data to obtain a plurality of voice segment data corresponding to the interpolated voice data; Extracting intensity features from the plurality of speech segment data to obtain a speech intensity feature sequence; Determine the speech emotion feature sequence corresponding to the interpolated speech data according to the speech intensity feature sequence and the preset position code, The speech intensity feature sequence and the speech emotion feature sequence are used as speech feature sequences.

7. The attribute interpolation method for realizing speech synthesis according to claim 1, characterized in that: Generating a speech output corresponding to the interpolated speech data according to the speech feature sequence includes: Performing feature standardization on the speech feature sequence to obtain target speech features; Performing waveform synthesis on the target speech feature to obtain a speech waveform; The speech waveform is subjected to audio conversion to obtain a speech output corresponding to the interpolated speech data.

8. A speech synthesis attribute interpolation device, characterized in that: The device comprises: An acoustic model building module, used to obtain speech data of a speaker and build an acoustic speech model according to speech acoustic features in the speech data; A content model building module, used to build a content speech model according to speech content features in the speech data; A model fusion module, used to perform weighted average fusion according to the fusion weights corresponding to the acoustic speech model and the content speech model to obtain a target speech fusion model; An attribute interpolation module, used for performing attribute interpolation on the speech data according to the target speech fusion model to obtain interpolated speech data; The speech output module is used to extract the speech sequence of the interpolated speech data to obtain a speech feature sequence, and generate a speech output corresponding to the interpolated speech data according to the speech feature sequence.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute an attribute interpolation method for implementing speech synthesis as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for implementing attribute interpolation for speech synthesis as claimed in any one of claims 1 to 7 is implemented.