Artificial intelligence-based voice conversion method, device, computer equipment and medium

By extracting the boundary frames and rhythmic features of phoneme sequences and text sequences in speech conversion, constructing a feature value sequence and reconstructing the speech using the target timbre, the problem of low speech conversion accuracy in the existing technology is solved, and the accuracy of speech conversion and the naturalness of robot customer service are improved.

CN116612773BActive Publication Date: 2025-10-14PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310724428.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-16
Publication Date
2025-10-14
Estimated Expiration
2043-06-16

AI Technical Summary

Technical Problem

Existing speech conversion methods are unable to accurately decouple and extract the text semantic information, speech prosody information and original speaker information from the speech to be converted, resulting in low accuracy of the reconstructed speech.

Method used

By obtaining the phoneme sequence and text sequence of the speech to be converted, determining M boundary frames and the duration of each boundary frame, extracting the rhythmic features of the text sequence, constructing a text rhythmic feature sequence and aligning it with the phoneme sequence, constructing a feature value sequence and performing speech reconstruction, and using the target timbre for speech reconstruction.

Benefits of technology

It improves the accuracy of voice conversion, enhances the naturalness and expressiveness of robot customer service, and improves the service quality of financial services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612773B_ABST
    Figure CN116612773B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of speech conversion, and particularly relates to a speech conversion method and device based on artificial intelligence, a computer device and a medium. The application determines M boundary frames in a phoneme sequence and corresponding duration, extracts a first text prosody feature sequence of a text sequence, constructs a feature value sequence of the duration of the corresponding boundary frame according to the feature value corresponding to the target position of the boundary frame, and sequentially combines the feature value sequences corresponding to all boundary frames to form a second text prosody feature sequence. The target reconstructed speech is obtained according to the text sequence, the second text prosody feature and the target timbre. The feature value correction is performed by extracting the first text prosody feature sequence, the representation accuracy of semantic information and prosody information is improved, the influence of speaker information in the speech to be converted on the reconstructed speech is reduced, the accuracy of speech conversion is improved, and the naturalness, expressiveness and service quality of the robot customer service in the financial scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention is applicable to the field of voice conversion technology, and in particular relates to a voice conversion method, device, computer equipment and medium based on artificial intelligence. Background Art

[0002] Speech conversion is to make what one person says sound like what another person said without changing the content of the speech. It has great application value in many fields such as driving navigation and video production.

[0003] Existing speech conversion methods typically extract textual semantic information from the speech to be converted and target speaker information of the target speaker. These information is then fused and mapped to produce the reconstructed speech. However, speech prosody, a hallmark of natural human language, characterizes language and emotion through properties such as pitch, intensity, and timing. This plays a crucial role in guiding the naturalness and expressiveness of the reconstructed speech in speech conversion tasks. However, these speech conversion methods fail to accurately decouple and extract the textual semantic information, speech prosody information, and original speaker information from the converted speech. Consequently, the converted target speech is contaminated by the original speaker information and low-quality speech prosody information, reducing the accuracy of the reconstructed speech.

[0004] Therefore, in the field of speech conversion technology, how to improve the accuracy of speech conversion methods has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method, apparatus, computer device, and medium for speech conversion based on artificial intelligence to solve the problem of low accuracy of reconstructed speech in existing speech conversion methods.

[0006] In a first aspect, an embodiment of the present invention provides a voice conversion method based on artificial intelligence, the voice conversion method comprising:

[0007] Obtain a phoneme sequence and a text sequence of the speech to be converted, and determine M boundary frames and the duration corresponding to each boundary frame based on the phoneme sequence, where M is an integer greater than 0;

[0008] Extracting prosodic features from the text sequence to obtain a first text prosodic feature sequence, aligning the first text prosodic feature sequence with the phoneme sequence, and determining a target position corresponding to each boundary frame in the first text prosodic feature sequence;

[0009] constructing a feature value sequence corresponding to the duration of the boundary frame according to the feature value corresponding to the target position of the boundary frame in the first text prosodic feature sequence;

[0010] According to the order of all boundary frames in the phoneme sequence, the feature value sequences corresponding to all boundary frames are combined into a second text prosodic feature sequence;

[0011] A target timbre is obtained, and speech reconstruction is performed according to the text sequence, the second text prosodic feature sequence, and the target timbre to obtain a target reconstructed speech.

[0012] In a second aspect, an embodiment of the present invention provides an artificial intelligence-based speech conversion device, the speech conversion device comprising:

[0013] A boundary frame determination module is configured to obtain a phoneme sequence and a text sequence of the speech to be converted, and determine M boundary frames and a duration corresponding to each boundary frame based on the phoneme sequence, where M is an integer greater than 0;

[0014] a feature extraction module configured to extract prosodic features from the text sequence to obtain a first text prosodic feature sequence, align the first text prosodic feature sequence with the phoneme sequence, and determine a target position corresponding to each boundary frame in the first text prosodic feature sequence;

[0015] a sequence construction module, configured to construct a feature value sequence corresponding to the duration of the boundary frame according to the feature value corresponding to the target position of the boundary frame in the first text prosodic feature sequence;

[0016] a sequence composition module, configured to compose a second text prosodic feature sequence from feature value sequences corresponding to all boundary frames according to the order of all boundary frames in the phoneme sequence;

[0017] The speech reconstruction module is used to obtain the target timbre, and reconstruct the speech according to the text sequence, the second text prosodic feature sequence and the target timbre to obtain the target reconstructed speech.

[0018] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the speech conversion method as described in the first aspect when executing the computer program.

[0019] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the speech conversion method as described in the first aspect is implemented.

[0020] Compared with the prior art, the embodiments of the present invention have the following advantages: determining M boundary frames and the duration corresponding to each boundary frame based on the phoneme sequence of the speech to be converted, extracting prosodic features from the text sequence of the speech to be converted to obtain a first text prosodic feature sequence, aligning the first text prosodic feature sequence with the phoneme sequence, determining the target position corresponding to each boundary frame in the first text prosodic feature sequence, constructing a feature value sequence corresponding to the duration of the boundary frame based on the feature value corresponding to the target position of the boundary frame in the first text prosodic feature sequence, forming a second text prosodic feature sequence from the feature value sequences corresponding to all boundary frames in the order of all boundary frames in the phoneme sequence, reconstructing speech based on the text sequence, the second text prosodic features, and the obtained target timbre to obtain a target reconstructed speech, and obtaining a second text prosodic feature sequence by extracting the first text prosodic feature sequence of the speech to be converted and performing feature value correction. This improves the accuracy of representing semantic and prosodic information in the speech to be converted, reduces the influence of speaker information in the speech to be converted on the reconstructed speech, improves the accuracy of speech conversion, improves the naturalness, expressiveness, and richness of robot customer service in financial scenarios, and improves the service quality of financial services. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 This is a schematic diagram of an application environment of an artificial intelligence-based voice conversion method provided in Example 1 of the present invention;

[0023] Figure 2 This is a flow chart of a voice conversion method based on artificial intelligence provided by the first embodiment of the present invention;

[0024] Figure 3 This is a schematic diagram of the structure of a speech conversion device based on artificial intelligence provided in the second embodiment of the present invention;

[0025] Figure 4 This is a structural diagram of a computer device provided in Example 3 of the present invention. DETAILED DESCRIPTION

[0026] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration, not limitation, to facilitate a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present invention with unnecessary detail.

[0027] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0028] It will also be understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0029] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0030] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0031] References to "one embodiment" or "some embodiments" in the present specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present invention. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0032] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0033] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0034] It should be understood that the order of execution of the steps in the following embodiments does not necessarily mean the order in which they are executed. The order in which each process is executed should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0035] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.

[0036] The first embodiment of the present invention provides an artificial intelligence-based voice conversion method that can be applied to Figure 1 In an application environment, a client communicates with a server. The client includes but is not limited to a PDA, a desktop computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud computer device, a personal digital assistant (PDA), and other computer devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The voice conversion method can be applied to various fields such as animation production, game production, voice navigation, and voice customer service. For example, due to the complexity and diversity of financial services, a large number of simple tasks such as consulting services and after-sales services will seriously occupy the energy and time of business personnel, reducing their work efficiency and work quality. The use of robot customer service can save a lot of labor costs. Whether the voice service of the robot customer service is natural and smooth will directly affect the user experience. Therefore, the voice conversion method can convert the timbre of the robot customer service in the financial scenario, providing customers with more natural, expressive, and rich voice customer service services, thereby improving the customer experience and thus improving the service quality of the financial business.

[0037] See also Figure 2, is a flow chart of a voice conversion method based on artificial intelligence provided by the first embodiment of the present invention. The above voice conversion method can be applied to Figure 1 In the client, the voice conversion method may include the following steps:

[0038] Step S201 : obtaining a phoneme sequence and a text sequence of the speech to be converted, and determining M boundary frames and a duration corresponding to each boundary frame according to the phoneme sequence.

[0039] A speech segment contains semantic information representing the content of the speech, speaker information representing the speaker's timbre and pronunciation habits, and prosodic information representing emotion, rhythm, pauses, etc. The task of speech conversion is to make the speech of one person sound like that of another without changing the semantic and prosodic information. In other words, the original speaker information in the speech is replaced with the speaker information of the target speaker.

[0040] In the speech conversion task, in order to reduce the influence of the speaker information of the speech to be converted on the reconstructed speech, it is necessary to decouple the semantic information, speaker information and prosodic information in the speech to be converted, so as to obtain the semantic information and prosodic information in the speech to be converted, and combine the speaker information of the target speaker to reconstruct the speech to obtain the reconstructed speech, thus completing the speech conversion task from the speech to be converted to the reconstructed speech.

[0041] Since phonemes are the smallest speech units divided according to the natural properties of speech, one action can constitute a phoneme when analyzed based on the pronunciation action in the syllable. In this embodiment, when performing speech conversion, the phoneme sequence and text sequence of the speech to be converted are obtained as the basis for performing speech conversion on the speech to be converted, so as to improve the accuracy of decoupling of the speech to be converted by converting the speech to be converted into a more basic phoneme form and text form, thereby improving the accuracy of speech conversion. Specifically, the speech to be converted can be subjected to phoneme conversion and text conversion to obtain the corresponding phoneme sequence and text sequence. Correspondingly, when performing speech conversion for robot customer service in a financial scenario, the speech to be converted can be the original speech corresponding to the robot customer service's communication with the customer, the text sequence of the speech to be converted can be the communication content between the robot customer service and the customer, and the phoneme sequence of the speech to be converted can be the phoneme sequence of the original speech.

[0042] In a phoneme sequence, when the duration of a phoneme changes, the semantic information and prosodic information contained in the speech to be converted will change. The duration of a phoneme is of great significance for characterizing the text prosodic features of the speech to be converted. Therefore, this embodiment segments each type of phoneme in the phoneme sequence to obtain boundary frames between M phonemes and the duration of each phoneme corresponding to each boundary frame as the basis for analyzing the text prosodic features of the speech to be converted, where M is an integer greater than 0. Optionally, determining the M boundary frames based on the phoneme sequence includes:

[0043] For the j-th frame phoneme in the phoneme sequence, compare the j-th frame phoneme with the j+1-th frame phoneme to see if they are consistent, and obtain a comparison result, where j=1, 2, ..., N-1, where N is the total number of phonemes in the phoneme sequence and N is an integer greater than 1;

[0044] If the comparison result is inconsistent, the frame number corresponding to any frame phoneme is determined to be a boundary frame;

[0045] Traverse the 1st, 2nd, ..., N-1th frame phonemes in the phoneme sequence and determine the M-1 boundary frames corresponding to the phoneme sequence.

[0046] Here, for the j-th frame phoneme in the phoneme sequence, j = 1, 2, ..., N-1, when the j-th frame phoneme is inconsistent with the j+1-th frame phoneme, it can be determined that the j-th frame phoneme is a boundary frame separating two types of phonemes in the phoneme sequence. When the j-th frame phoneme is consistent with the j+1-th frame phoneme, it can be determined that the j-th frame phoneme is not a boundary frame separating two types of phonemes in the phoneme sequence. Therefore, the 1st, 2nd, ..., N-1-th frame phoneme is traversed to determine M-1 boundary frames in the phoneme sequence.

[0047] This embodiment determines whether the phonemes in the jth frame are consistent with the phonemes in the j+1th frame to determine whether the phonemes in the 1st, 2nd, ..., N-1th frames are boundary frames separating two types of phonemes in the phoneme sequence, thereby improving the rationality and accuracy of boundary frame calculation.

[0048] Optionally, determining M boundary frames according to the phoneme sequence further includes:

[0049] The number N of frames corresponding to the N-th frame phoneme in the phoneme sequence is determined as a boundary frame, and M boundary frames are obtained.

[0050] Among them, since the N-th frame phoneme does not have a corresponding next frame phoneme, although the N-th frame phoneme cannot be used to divide into two categories of phonemes, it can still be determined that the N-th frame phoneme is a boundary frame in the phoneme sequence, and combined with the M-1 boundary frames determined by traversing the 1st, 2nd, ..., N-1th frame phonemes, M boundary frames are obtained.

[0051] This embodiment considers that the Nth frame phoneme does not have a corresponding next frame phoneme, and determines that the Nth frame phoneme is a boundary frame in the phoneme sequence, thereby improving the rationality and accuracy of boundary frame calculation.

[0052] Optionally, determining the duration corresponding to each boundary frame according to the phoneme sequence includes:

[0053] Sort the M boundary frames according to the number of frames corresponding to the M boundary frames;

[0054] For the i-th boundary frame, the frame number difference between the i-th boundary frame and the (i-1)-th boundary frame is calculated, and the difference is determined as the duration corresponding to the i-th boundary frame, i=2, 3, . . . , M.

[0055] Among them, first, the M boundary frames are sorted according to the frame number of the M boundary frames. For the i-th boundary frame, the phoneme corresponding to the i-th boundary frame and the phoneme between the i-th boundary frame and the i-1-th boundary frame belong to the same type of phoneme. Then, the frame number difference between the i-th boundary frame and the i-1-th boundary frame can be calculated, and the difference is determined as the duration corresponding to the i-th boundary frame, where i = 2, 3, ..., M.

[0056] Optionally, determining the duration of each boundary frame according to the phoneme sequence further includes:

[0057] For the first boundary frame, the number of frames corresponding to the boundary frame is determined as the duration corresponding to the boundary frame.

[0058] Among them, since the first boundary frame does not have a previous boundary frame, and the phoneme corresponding to the first boundary frame and all the phonemes corresponding to the first boundary frame before the first boundary frame belong to the same type of phoneme, the frame number of the first boundary frame can be determined as the duration corresponding to the first boundary frame.

[0059] This embodiment takes into account the difference between the first boundary frame and other boundary frames, determines the frame number of the first boundary frame as the duration corresponding to the first boundary frame, and determines the difference in the number of frames between other boundary frames and their previous boundary frames as the duration corresponding to the other boundary frames, thereby improving the calculation accuracy of the duration.

[0060] For example, for the phoneme sequence [a,a,b,b,b,c,c] of the speech to be converted, N = 7. For the second frame of phoneme a, the comparison results of the second frame phoneme a and the third frame phoneme b are inconsistent. Therefore, the frame number 2 corresponding to the second frame phoneme a can be determined to be a boundary frame. For the third frame of phoneme b, the comparison results of the third frame phoneme b and the fourth frame phoneme b are consistent. Therefore, the frame number 3 corresponding to the third frame phoneme b can be determined to be not a boundary frame. For the last frame of phoneme, that is, the seventh frame phoneme c, the frame number 7 corresponding to the seventh frame phoneme c can be determined to be a boundary frame. Then, by traversing the seven frames of phonemes in the above phoneme sequence [a,a,b,b,b,c,c], the boundary frames can be determined to be 2, 5, and 7, correspondingly, M = 3.

[0061] According to the order of the boundary frames, for the first boundary frame, the frame number 2 corresponding to the first boundary frame is determined as the duration corresponding to the first boundary frame; for the second boundary frame, the frame number difference between the second boundary frame and the first boundary frame is calculated, that is, 5-2=3, and the difference 3 is determined as the duration corresponding to the second boundary frame; for the third boundary frame, the frame number difference between the third boundary frame and the second boundary frame is calculated, that is, 7-5=2, and the difference 2 is determined as the duration corresponding to the third boundary frame.

[0062] The above steps of obtaining the phoneme sequence and text sequence of the speech to be converted, and determining M boundary frames and the duration corresponding to each boundary frame based on the phoneme sequence, convert the speech to be converted into a more basic phoneme form as the basis for characterizing the text prosodic features of the speech to be converted, thereby improving the accuracy of feature extraction and feature decoupling of the speech to be converted.

[0063] Step S202 : extracting prosodic features from the text sequence to obtain a first text prosodic feature sequence, aligning the first text prosodic feature sequence with the phoneme sequence, and determining the target position corresponding to each boundary frame in the first text prosodic feature sequence.

[0064] Among them, since the text sequence can represent the text information and prosodic information in the speech to be converted, it cannot represent the timbre information of the speech to be converted. Therefore, in order to accurately decouple the semantic information, speaker information and prosodic information in the speech to be converted, this embodiment performs feature extraction on the text sequence. Compared with feature extraction on the speech to be converted, feature extraction on the text sequence can focus on extracting the semantic information and prosodic information in the speech to be converted corresponding to the text sequence, ignoring the speaker information in the speech to be converted, and obtaining a first text prosodic feature sequence. The first text prosodic feature sequence is used as the basis for reconstructing the semantic information and prosodic information of the speech to be converted, reducing the influence of the speaker information of the speech to be converted on the reconstructed speech, thereby improving the accuracy of speech conversion.

[0065] In this embodiment, after obtaining the first text prosody feature sequence, the first text prosody feature sequence is aligned with the phoneme sequence, the target position corresponding to each boundary frame in the first text prosody feature sequence is determined, the feature value corresponding to each boundary frame is determined, the first text prosody feature sequence is corrected according to the feature value, and the representation accuracy of the text prosody feature sequence in the speech to be converted is further improved.

[0066] The above step of extracting the prosody feature of the text sequence, obtaining the first text prosody feature sequence, aligning the first text prosody feature sequence with the phoneme sequence, and determining the target position corresponding to each boundary frame in the first text prosody feature sequence, focuses on extracting semantic information and prosody information in the speech to be converted corresponding to the text sequence, ignores the speaker information in the speech to be converted, obtains the first text prosody feature sequence, and takes the target position corresponding to the boundary frame as the basis for correcting the first text prosody feature sequence, to reconstruct the semantic information and prosody information of the speech, reduces the influence of the speaker information of the speech to be converted on the reconstructed speech, and improves the accuracy of speech conversion.

[0067] In step S203, a feature value sequence corresponding to the duration of the boundary frame is constructed according to the feature value corresponding to the target position of the boundary frame in the first text prosody feature sequence.

[0068] For any boundary frame, the phonemes corresponding to the boundary frame in the phoneme sequence belong to the same type of phonemes, and the number of phonemes is consistent with the corresponding duration.

[0069] In this embodiment, the feature value of the target position corresponding to the boundary frame in the first text prosody feature sequence is taken as the feature value of the boundary frame, so the feature value corresponding to the boundary frame can be used to represent the feature value of each phoneme corresponding to the boundary frame. Therefore, the feature value sequence corresponding to the duration of the boundary frame can be constructed based on the feature value corresponding to the target position of the boundary frame in the first text prosody feature sequence, and the feature value sequence can be taken as the basis for correcting the first text prosody feature sequence, so as to improve the representation accuracy of the text prosody feature in the speech to be converted.

[0070] Optionally, constructing the feature value sequence corresponding to the duration of the boundary frame according to the feature value corresponding to the target position of the boundary frame in the first text prosody feature sequence comprises:

[0071] For any boundary frame, the feature value corresponding to the target position of the boundary frame in the first text prosody feature sequence is copied until the number of feature values is consistent with the duration corresponding to the boundary frame, to obtain the feature value sequence corresponding to the duration of the boundary frame;

[0072] All boundary frames are traversed to obtain the feature value sequence corresponding to the duration of all boundary frames.

[0073] Among them, since the boundary frame can separate two types of phonemes in the phoneme sequence, and the phonemes corresponding to the i-th boundary frame and the phonemes between the i-th boundary frame and the i-1-th boundary frame belong to the same type of phonemes, the phonemes corresponding to the first boundary frame and all the phonemes corresponding to the first boundary frame belong to the same type of phonemes, therefore, by copying the feature value until the number of feature values ​​is consistent with the duration corresponding to the boundary frame, the feature value sequence corresponding to the boundary frame can be obtained to complete the correction of the feature values ​​of all phonemes corresponding to the boundary frame, thereby ensuring the consistency and accuracy of the feature values ​​of all phonemes corresponding to the boundary frame.

[0074] For example, for the first text prosodic feature sequence extracted [A,A,D,B,B,C,C], M=3, where the number of frames corresponding to the first boundary frame is 2 and the duration is 2, the number of frames corresponding to the second boundary frame is 5 and the duration is 3, and the number of frames corresponding to the third boundary frame is 7 and the duration is 2.

[0075] Then, it can be determined that the feature value corresponding to the target position of the first boundary frame in the first text rhythmic feature sequence is A, and the feature value A is copied until the number of feature values ​​A is consistent with the corresponding duration 2, and the feature value sequence corresponding to the first boundary frame is [A, A]; the feature value corresponding to the target position of the second boundary frame in the first text rhythmic feature sequence is determined to be B, and the feature value B is copied until the number of feature values ​​B is consistent with the corresponding duration 3, and the feature value sequence corresponding to the second boundary frame is [B, B, B]; the feature value corresponding to the target position of the third boundary frame in the first text rhythmic feature sequence is determined to be C, and the feature value C is copied until the number of feature values ​​C is consistent with the corresponding duration 2, and the feature value sequence corresponding to the third boundary frame is [C, C].

[0076] The above-mentioned step of constructing a feature value sequence corresponding to the duration of the boundary frame according to the feature value corresponding to the target position corresponding to the boundary frame in the first text prosodic feature sequence, obtains the feature value sequence corresponding to the boundary frame by copying the feature value corresponding to the boundary frame, completes the correction of the feature values ​​of all phonemes corresponding to the boundary frame, ensures the consistency and accuracy of the feature values ​​of all phonemes corresponding to the boundary frame, and improves the accuracy of the representation of the text prosodic features in the converted speech.

[0077] Step S204 : According to the order of all boundary frames in the phoneme sequence, the feature value sequences corresponding to all boundary frames are combined into a second text prosodic feature sequence.

[0078] Wherein, after the above feature value determination and feature value copying operations are performed on all boundary frames to obtain the feature value sequences corresponding to all boundary frames, the feature values of all phonemes in the phoneme sequence can be corrected. Then, according to the order of all boundary frames in the phoneme sequence, the feature value sequences corresponding to all boundary frames are sorted and combined to obtain a second text prosody feature sequence as the basis of semantic information and prosody information during speech reconstruction.

[0079] For example, according to the order of all boundary frames in the phoneme sequence, the feature value sequence corresponding to the first boundary frame

A,A

B,B,B

C,C

A,A,B,B,B,C,C

[0080] The above step of combining the feature value sequences corresponding to all boundary frames to form a second text prosody feature sequence according to the order of all boundary frames in the phoneme sequence, the feature value determination and feature value copying operations performed on all boundary frames to obtain the feature value sequences corresponding to all boundary frames, and the sorting and combination to obtain the second text prosody feature sequence, improve the accuracy of representing the text prosody features in the speech to be converted.

[0081] In step S205, the target voice timbre is obtained, and speech reconstruction is performed according to the text sequence, the second text prosody feature sequence, and the target voice timbre to obtain a target reconstructed speech.

[0082] Wherein, the target voice timbre can represent the speaker information of the target speaker, the text sequence can represent the text information in the speech to be converted, and the second text prosody feature can represent the semantic information and prosody information in the speech to be converted. Then, speech reconstruction can be performed according to the text sequence, the second text prosody feature sequence, and the target voice timbre, the speaker information in the speech to be converted is replaced with the speaker information of the target speaker without changing the semantic information and prosody information, to obtain a target reconstructed speech and complete the speech conversion task. Correspondingly, in the financial scenario, the voice timbre of the artificial customer service can be collected as the target voice timbre for speech conversion of the robot customer service, and the target voice timbre can be stored in the corresponding database as the data basis for speech conversion.

[0083] In one embodiment, a trained decoder can be used to reconstruct speech from a text sequence, a second text prosodic feature, and a target timbre, outputting the target reconstructed speech. To improve the accuracy of the target reconstructed speech, the decoder can be trained by modifying its parameters. Therefore, this embodiment obtains a speech sample to be converted, a text sequence sample of the speech sample to be converted, a second text prosodic feature sequence sample, and a target timbre sample as training samples, and uses the speech sample to be converted as a training label to train the decoder.

[0084] Specifically, a decoder is used to reconstruct speech from a text sequence sample, a second text prosodic feature sequence sample, and a target timbre sample, yielding a reconstructed speech. This reconstructed speech represents the same semantic, speaker, and prosodic information as the speech sample to be converted. Therefore, the higher the similarity between the reconstructed speech and the speech sample to be converted, the higher the accuracy of the decoder. The model loss is then calculated by calculating the similarity between the speech sample to be converted and the reconstructed speech. The decoder parameters are then reversely modified using gradient descent until the model loss converges, resulting in a trained decoder.

[0085] This embodiment uses a trained decoder to reconstruct the text sequence, the second text prosodic features and the target timbre, outputs the target reconstructed speech, and obtains the speech sample to be converted, the text sequence sample of the speech sample to be converted, the second text prosodic feature sequence sample and the target timbre sample as training samples, uses the speech sample to be converted as the training label, trains the decoder, improves the accuracy of the decoder, and thus improves the accuracy of the target reconstructed speech.

[0086] In one embodiment, calculating the model loss based on the speech sample to be converted and the reconstructed speech includes:

[0087] Calculating the first Mel-cepstral coefficient of the speech sample to be converted and the second Mel-cepstral coefficient of the reconstructed speech;

[0088] The similarity between the first mel-cepstral coefficient and the second mel-cepstral coefficient is calculated, and the difference between the similarity and a preset value is determined as the model loss.

[0089] To accurately measure the similarity between speech sounds, this embodiment calculates the first mel-cepstral coefficient of the speech sample to be converted to represent the information of the speech sample to be converted, and the second mel-cepstral coefficient of the reconstructed speech to represent the information of the reconstructed speech. Therefore, the similarity between the first mel-cepstral coefficient and the second mel-cepstral coefficient can be calculated. The higher the similarity, the more similar the speech sample to be converted and the reconstructed speech are, and thus the higher the accuracy of the decoder. Therefore, the difference between the similarity and a preset value is calculated and determined as the model loss. Correspondingly, the smaller the model loss, the higher the accuracy of the decoder.

[0090] The preset value can be set according to actual conditions. For example, when the calculated similarity is in the range of [0, 1], the preset value can be set to 1.

[0091] This embodiment determines the difference between the similarity between the first Mel-cepstral coefficient and the second Mel-cepstral coefficient and a preset value as the model loss, thereby improving the calculation accuracy of the model loss.

[0092] The above steps of obtaining target features, reconstructing speech based on the text sequence, the second text prosodic feature sequence and the target features, and obtaining the target reconstructed speech, replace the speaker information in the speech to be converted with the speaker information of the target speaker without changing the semantic information and prosodic information, obtain the target reconstructed speech, complete the speech conversion task, and improve the accuracy of the target reconstructed speech.

[0093] An embodiment of the present invention determines, based on a phoneme sequence of a speech to be converted, M boundary frames and the duration corresponding to each boundary frame, extracts prosodic features from a text sequence of the speech to be converted to obtain a first text prosodic feature sequence, aligns the first text prosodic feature sequence with the phoneme sequence, determines the target position corresponding to each boundary frame in the first text prosodic feature sequence, constructs a feature value sequence corresponding to the duration of the boundary frame based on the feature value corresponding to the target position of the boundary frame in the first text prosodic feature sequence, and, in accordance with the order of all boundary frames in the phoneme sequence, assembles the feature value sequences corresponding to all boundary frames into a second text prosodic feature sequence. Speech is reconstructed based on the text sequence, the second text prosodic features, and the obtained target timbre to obtain a target reconstructed speech. By extracting the first text prosodic feature sequence of the speech to be converted and performing feature value correction to obtain a second text prosodic feature sequence, the accuracy of representing semantic and prosodic information in the speech to be converted is improved, the influence of speaker information in the speech to be converted on the reconstructed speech is reduced, the accuracy of speech conversion is improved, the naturalness, expressiveness, and richness of robot customer service in financial scenarios are improved, and the service quality of financial services is improved.

[0094] Corresponding to the voice conversion method of the above embodiment, Figure 3 A structural block diagram of an artificial intelligence-based speech conversion device provided in the second embodiment of the present invention is given. For ease of explanation, only the parts related to the embodiment of the present invention are shown.

[0095] See also Figure 3 , the voice conversion device comprises:

[0096] The boundary frame determination module 31 is used to obtain a phoneme sequence and a text sequence of the speech to be converted, and determine M boundary frames and the duration corresponding to each boundary frame according to the phoneme sequence, where M is an integer greater than 0;

[0097] A feature extraction module 32 is configured to extract prosodic features from the text sequence to obtain a first text prosodic feature sequence, align the first text prosodic feature sequence with the phoneme sequence, and determine a target position corresponding to each boundary frame in the first text prosodic feature sequence;

[0098] A sequence construction module 33 is configured to construct a feature value sequence corresponding to the duration of a boundary frame according to the feature value corresponding to the target position of the boundary frame in the first text prosodic feature sequence;

[0099] A sequence composition module 34 is configured to compose a second text prosodic feature sequence from feature value sequences corresponding to all boundary frames according to the order of all boundary frames in the phoneme sequence;

[0100] The speech reconstruction module 35 is used to obtain the target timbre, and reconstruct the speech according to the text sequence, the second text prosodic feature sequence and the target timbre to obtain the target reconstructed speech.

[0101] Optionally, the boundary frame determination module 31 includes:

[0102] A first phoneme comparison submodule is configured to compare the j-th frame phoneme in the phoneme sequence with the j+1-th frame phoneme to determine whether the phoneme is consistent, and obtain a comparison result, where j=1, 2, ..., N-1, where N is the total number of phonemes in the phoneme sequence and N is an integer greater than 1;

[0103] A first boundary frame determination submodule, configured to determine the frame number corresponding to any frame phoneme as a boundary frame if the comparison result is inconsistent;

[0104] The second boundary frame determination submodule is used to traverse the 1st, 2nd, ..., N-1th frame phonemes in the phoneme sequence and determine M-1 boundary frames corresponding to the phoneme sequence.

[0105] Optionally, the boundary frame determination module 31 further includes:

[0106] The third boundary frame determination submodule is used to determine the frame number N corresponding to the N-th frame phoneme in the phoneme sequence as a boundary frame, and obtain M boundary frames.

[0107] Optionally, the boundary frame determination module 31 includes:

[0108] A sorting submodule, for sorting the M boundary frames according to the number of frames corresponding to the M boundary frames;

[0109] The first duration calculation submodule is used to calculate the frame number difference between the i-th boundary frame and the i-1-th boundary frame for the i-th boundary frame, and determine the difference as the duration corresponding to the i-th boundary frame, i=2, 3, ..., M.

[0110] Optionally, the boundary frame determination module 31 further includes:

[0111] The second duration calculation submodule is configured to determine, for the first boundary frame, the number of frames corresponding to the boundary frame as the duration corresponding to the boundary frame.

[0112] Optionally, the sequence construction module 33 includes:

[0113] The first sequence construction submodule is used to copy the feature values ​​corresponding to the target position of any boundary frame in the first text prosodic feature sequence until the number of feature values ​​is consistent with the duration corresponding to the boundary frame, thereby obtaining a feature value sequence corresponding to the duration of the boundary frame;

[0114] The second sequence construction submodule is used to traverse all boundary frames and obtain a feature value sequence corresponding to the duration of all boundary frames.

[0115] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0116] Figure 4 This is a schematic diagram of the structure of a computer device provided in the third embodiment of the present invention. Figure 4 As shown, the computer device of this embodiment includes: at least one processor ( Figure 4 Only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, the steps of any of the above-mentioned speech conversion method embodiments are implemented.

[0117] The computer device may include, but is not limited to, a processor and a memory. It will be understood by those skilled in the art that Figure 4The above is merely an example of a computer device and does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include a network interface, a display screen, and an input device.

[0118] The processor may be a CPU, or other general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor, or any conventional processor.

[0119] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of the computer device, and in other embodiments, it can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Furthermore, the memory can also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders (BootLoader), data, and other programs, such as the program code of the computer program. The memory can also be used to temporarily store data that has been output or is about to be output.

[0120] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned method embodiment. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include at least: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0121] The present invention may implement all or part of the processes in the above-mentioned method embodiments, and may also be completed through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0122] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0123] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0124] In the embodiments provided by the present invention, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0125] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0126] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A voice conversion method based on artificial intelligence, characterized in that: The voice conversion method comprises: Obtain a phoneme sequence and a text sequence of the speech to be converted, and determine M boundary frames and a duration corresponding to each boundary frame based on the phoneme sequence, where M is an integer greater than 0; the duration is the number of frames between adjacent boundary frames; Extracting prosodic features from the text sequence to obtain a first text prosodic feature sequence, aligning the first text prosodic feature sequence with the phoneme sequence, and determining a target position corresponding to each boundary frame in the first text prosodic feature sequence; constructing a feature value sequence corresponding to the duration of the boundary frame according to the feature value corresponding to the target position of the boundary frame in the first text prosodic feature sequence; According to the order of all boundary frames in the phoneme sequence, the feature value sequences corresponding to all boundary frames are combined into a second text prosodic feature sequence; Acquire a target timbre, and reconstruct speech based on the text sequence, the second text prosodic feature sequence, and the target timbre to obtain a target reconstructed speech; Determining M boundary frames according to the phoneme sequence includes: For the j-th frame phoneme in the phoneme sequence, compare the j-th frame phoneme with the j+1-th frame phoneme to determine whether they are consistent, obtaining a comparison result, where j=1, 2, ..., N-1, where N is the total number of phonemes in the phoneme sequence and N is an integer greater than 1; If the comparison result is inconsistent, determining the frame number corresponding to any frame phoneme as a boundary frame; The 1st, 2nd, ..., N-1th frame phonemes in the phoneme sequence are traversed to determine M-1 boundary frames corresponding to the phoneme sequence.

2. The voice conversion method according to claim 1, wherein: The determining of M boundary frames according to the phoneme sequence further includes: The number N of frames corresponding to the N-th frame phoneme in the phoneme sequence is determined as a boundary frame, and M boundary frames are obtained.

3. The voice conversion method according to claim 2, wherein: Determining the duration corresponding to each boundary frame according to the phoneme sequence includes: sorting the M boundary frames according to the frame numbers corresponding to the M boundary frames; For the i-th boundary frame, the frame number difference between the i-th boundary frame and the (i-1)-th boundary frame is calculated, and the difference is determined as the duration corresponding to the i-th boundary frame, i=2, 3, ..., M.

4. The voice conversion method according to claim 3, wherein: Determining the duration of each boundary frame according to the phoneme sequence further includes: For the first boundary frame, the number of frames corresponding to the boundary frame is determined as the duration corresponding to the boundary frame.

5. The voice conversion method according to claim 1, wherein: The step of constructing a feature value sequence corresponding to the duration of the boundary frame according to the feature value corresponding to the target position of the boundary frame in the first text prosodic feature sequence includes: For any boundary frame, copy the feature value corresponding to the target position corresponding to the boundary frame in the first text prosodic feature sequence until the number of the feature values ​​is consistent with the duration corresponding to the boundary frame, thereby obtaining a feature value sequence corresponding to the duration of the boundary frame; Traverse all boundary frames and obtain the feature value sequence corresponding to the duration of all boundary frames.

6. A voice conversion device based on artificial intelligence, characterized in that: The voice conversion device comprises: A boundary frame determination module is configured to obtain a phoneme sequence and a text sequence of the speech to be converted, and determine M boundary frames and a duration corresponding to each boundary frame based on the phoneme sequence, where M is an integer greater than 0; the duration is the number of frames between adjacent boundary frames; a feature extraction module configured to extract prosodic features from the text sequence to obtain a first text prosodic feature sequence, align the first text prosodic feature sequence with the phoneme sequence, and determine a target position corresponding to each boundary frame in the first text prosodic feature sequence; a sequence construction module, configured to construct a feature value sequence corresponding to the duration of the boundary frame according to the feature value corresponding to the target position of the boundary frame in the first text prosodic feature sequence; a sequence composition module, configured to compose a second text prosodic feature sequence from feature value sequences corresponding to all boundary frames according to the order of all boundary frames in the phoneme sequence; A speech reconstruction module is used to obtain a target timbre, and reconstruct the speech according to the text sequence, the second text prosodic feature sequence and the target timbre to obtain a target reconstructed speech; The boundary frame determination module includes: A first phoneme comparison submodule is configured to compare the j-th frame phoneme in the phoneme sequence with the j+1-th frame phoneme to determine whether the j-th frame phoneme is consistent with the j+1-th frame phoneme, to obtain a comparison result, where j=1, 2, ..., N-1, where N is the total number of phonemes in the phoneme sequence and N is an integer greater than 1; A first boundary frame determination submodule, configured to determine, if the comparison result is inconsistent, the frame number corresponding to any frame phoneme as a boundary frame; The second boundary frame determination submodule is used to traverse the 1st, 2nd, ..., N-1th frame phonemes in the phoneme sequence and determine M-1 boundary frames corresponding to the phoneme sequence.

7. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the speech conversion method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the speech conversion method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Speech synthesis method and device, computer equipment and storage medium

    CN114360490A

  • Voice processing method, device and equipment and computer readable storage medium

    CN114582315A