A method, apparatus, device, and storage medium for generating accent data.

By slicing the original speech data and the target accent data, and using an accent recognition model to generate accent labels for the spliced ​​speech slices, the problem of inaccurate speech data generation in existing technologies is solved, and more refined and accurate accent data generation is achieved.

CN119600989BActive Publication Date: 2025-11-14GUANGZHOU QUWAN NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411784363.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-11-14
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Existing speech generation technologies directly generate speech data with specific accent requirements from raw speech data, resulting in inaccurate speech data.

Method used

By slicing the original speech data and the target accent data, spliced ​​speech slices are generated. Then, a pre-trained accent recognition model is used to identify the accent labels of the spliced ​​speech slices, and finally, the target accent data is generated.

Benefits of technology

It improves the precision and accuracy of voice data generation, avoiding the negative impact of the complex and variable features of the original voice data on the generation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600989B_ABST
    Figure CN119600989B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and storage medium for generating accent data. The method involves: acquiring raw speech data; generating first speech data according to a preset target accent; slicing the raw speech data to obtain individual raw speech slices, and slicing the first speech data to obtain individual first speech slices; determining individual concatenated speech slices from the raw speech slices and the first speech slices; inputting each concatenated speech slice into a pre-trained accent recognition model to obtain accent labels for each concatenated speech slice; and generating target accent data based on each raw speech slice, the accent labels of each concatenated speech slice, and the target accent. Compared to existing generation methods, this application offers greater precision. Directly generating accent data from raw speech data may be affected by complex and variable global features, leading to poor generation results. The slicing method used in this application effectively avoids this problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of accent data generation technology, specifically to a method, apparatus, device, and storage medium for generating accent data. Background Technology

[0002] With the acceleration of globalization and the continuous expansion of voice interaction applications, such as voice assistants, intelligent customer service, and language learning software, the demand for diverse accented voice data has increased dramatically. People from different regions and cultural backgrounds have their own unique accents. In order to make the use of voice technology more widespread, the industry is now focusing on generating multiple voice data with different accents from the same voice data.

[0003] However, existing speech generation technologies often simply generate speech data with specific accent requirements directly from the original speech data. This approach is rather crude, and the generated speech data is not accurate. Summary of the Invention

[0004] In view of this, this application provides a method, apparatus, device and storage medium for generating accent data, which solves the problem that existing speech generation technologies often simply generate speech data with specific accent requirements directly from the original speech data. This method is relatively crude and the generated speech data is not accurate.

[0005] To achieve the above objectives, the following solution is proposed:

[0006] Firstly, a method for generating accent data includes:

[0007] Acquire raw speech data;

[0008] Generate first voice data corresponding to the original voice data according to the preset target accent;

[0009] The original speech data is sliced ​​to obtain various original speech slices, and the first speech data is sliced ​​to obtain various first speech slices.

[0010] Each spliced ​​speech slice is determined by each of the original speech slices and each of the first speech slices;

[0011] Each of the spliced ​​speech segments is input into a pre-trained accent recognition model to obtain accent labels for each of the spliced ​​speech segments; the accent recognition model is trained using a speech sample set as training samples and the accent labels of each spliced ​​speech segment in the speech sample set as sample labels.

[0012] Based on each of the original speech segments, and based on the accent tags of each of the spliced ​​speech segments and the target accent, target accent data is generated.

[0013] Preferably, determining each concatenated speech slice from each of the original speech slices and each of the first speech slices includes:

[0014] The original speech slices are combined to form an original speech sequence, and the sequence position of each original speech slice is determined as the original sequence position.

[0015] Each of the first speech slices is combined to form a first speech sequence, and the sequence position of each of the first speech slices is determined as the first sequence position;

[0016] Based on the original sequence position and the first sequence position, each original speech slice and each first speech slice are mapped to another;

[0017] Align and concatenate the corresponding original speech slices with the first speech slice to obtain each concatenated speech slice.

[0018] Preferably, the step of mapping each original speech slice and each first speech slice to another based on the original sequence position and the first sequence position includes:

[0019] For each of the original speech segments, a first speech segment whose first sequence position is the same as the original sequence position of the original speech segment is determined;

[0020] Obtain the text data of the first speech slice, and at the same time obtain the text data of the original speech slice;

[0021] Determine whether the text data of the first speech slice is completely consistent with the text data of the original speech slice;

[0022] If so, then the first speech slice is taken as the first speech slice corresponding to the original speech slice.

[0023] Preferably, the process of establishing the speech sample set includes:

[0024] Multiple voice sample data are extracted from a pre-defined target accent database and used as the first sample data;

[0025] Multiple voice sample data are extracted from a pre-determined database of other accents and used as second sample data.

[0026] Each of the first sample data is sliced ​​to obtain a first sample slice, and each of the second sample data is sliced ​​to obtain a second sample slice.

[0027] For each of the first sample slices, a first target accent sample corresponding to the first sample slice is generated according to the target accent.

[0028] Each of the first sample slices is aligned and spliced ​​with its corresponding first target accent sample to obtain each first spliced ​​sample.

[0029] For each second sample slice, generate a second target accent sample corresponding to the second sample slice according to the target accent;

[0030] Each second sample slice is aligned and spliced ​​with its corresponding second target accent sample to obtain each spliced ​​second sample.

[0031] Each of the first spliced ​​samples and each of the second spliced ​​samples are used as spliced ​​sample slices and summarized to obtain a speech sample set.

[0032] Preferably, generating target accent data based on each of the original speech segments and the accent tags of each of the concatenated speech segments and the target accent includes:

[0033] Each concatenated speech slice with the accent label set to the preset first label is used as a first accent concatenated slice;

[0034] Determine the original speech slice in each of the first accent splicing slices;

[0035] Generate a target accent slice corresponding to the original speech slice in each of the first accent splicing slices according to the target accent;

[0036] Each concatenated speech slice with the accent label set to the preset second label is used as a second accent concatenation slice;

[0037] Identify the original speech slice in each of the second accent splicing slices;

[0038] The target accent slice corresponding to the original speech slice in each of the first accent splicing slices is combined with the original speech slice in each of the second accent splicing slices to generate target accent data.

[0039] Preferably, the step of slicing the original speech data to obtain various original speech slices includes:

[0040] Identify the pause locations in the raw speech data;

[0041] The original speech data is divided into first original slices according to each of the pause positions;

[0042] Determine the text data corresponding to the original speech data;

[0043] Identify each punctuation mark in the text data;

[0044] Each punctuation mark is matched with the original speech data to determine the position of each punctuation mark in the original speech data, which is then used as the position of each punctuation mark.

[0045] Determine whether the positions of each punctuation mark in the original speech data coincide with the pause positions;

[0046] Each non-overlapping punctuation mark is used as a segmentation point.

[0047] Each of the positions to be segmented in each of the first original slices is segmented to obtain each original speech slice.

[0048] Preferably, generating the first speech data corresponding to the original speech data according to a preset target accent includes:

[0049] Extract the timbre, prosodic, and emotional features from the original speech data;

[0050] Determine the text data corresponding to the original speech data;

[0051] Based on the text data, and using the target accent, timbre features, prosodic features, and emotional features as standards, first speech data corresponding to the original speech data is generated.

[0052] Secondly, an apparatus for generating accent data includes:

[0053] The raw speech data acquisition module is used to acquire raw speech data;

[0054] The first speech data generation module is used to generate first speech data corresponding to the original speech data according to a preset target accent.

[0055] The slicing module is used to slice the original speech data to obtain various original speech slices, and at the same time slice the first speech data to obtain various first speech slices.

[0056] A spliced ​​speech slice determination module is used to determine each spliced ​​speech slice from each of the original speech slices and each of the first speech slices;

[0057] The accent labeling module is used to input each of the spliced ​​speech segments into a pre-trained accent recognition model to obtain accent labels for each of the spliced ​​speech segments; the accent recognition model is trained using a speech sample set as training samples and the accent labels of each spliced ​​sample segment in the speech sample set as sample labels.

[0058] The target accent data generation module is used to generate target accent data based on each of the original speech slices and the accent tags of each of the spliced ​​speech slices and the target accent.

[0059] Thirdly, an apparatus for generating accent data, including a memory and a processor;

[0060] The memory is used to store programs;

[0061] The processor is configured to execute the program to implement the steps of the method for generating accent data as described in any of the first aspects.

[0062] Fourthly, a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for generating accent data as described in any of the first aspects.

[0063] As can be seen from the above technical solution, this application obtains original speech data; generates first speech data corresponding to the original speech data according to a preset target accent; slices the original speech data to obtain various original speech slices, and simultaneously slices the first speech data to obtain various first speech slices; determines various concatenated speech slices from the various original speech slices and the various first speech slices; inputs each concatenated speech slice into a pre-trained accent recognition model to obtain accent labels for each concatenated speech slice; the accent recognition model is trained using a speech sample set as training samples and the accent labels of each concatenated sample slice in the speech sample set as sample labels; and generates target accent data based on each original speech slice and the accent labels of each concatenated speech slice and the target accent. This application first generates first speech data corresponding to the original speech data according to a preset target accent. This can be regarded as standard speech data under the target accent, which is directly generated speech data. Then, the original speech data and the first speech data are sliced ​​to divide each into multiple small parts. This is to reduce the size of the original speech data. Generating accent data from small parts will be more detailed and accurate. The slices are spliced ​​to form various spliced ​​speech slices. This application pre-trains an accent recognition model. The spliced ​​speech slices are input into this accent recognition model, which can process the spliced ​​speech slices and predict the accent labels of the spliced ​​speech slices. This allows the accent information corresponding to the spliced ​​speech slices to be determined. Thus, the accent situation of each original speech slice can be determined. Then, based on each original speech slice, accent label, and target accent, the target accent data can be generated. Compared with the generation method of the prior art, the generation method of this application is more refined. In the process of directly generating accent data according to the original speech data, it may be affected by the complex and variable global features in the original speech data, resulting in poor generation effect. The slicing method of this application can effectively avoid this problem. Attached Figure Description

[0064] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0065] Figure 1 An optional flowchart of a method for generating accent data provided in an embodiment of this application;

[0066] Figure 2 This is a schematic diagram of the structure of a spliced ​​speech slice provided in an embodiment of this application;

[0067] Figure 3 A schematic diagram of the structure of an accent data generation device provided in an embodiment of this application;

[0068] Figure 4 This is a schematic diagram of the structure of an accent data generation device provided in an embodiment of this application. Detailed Implementation

[0069] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0070] In recent years, on the one hand, large-scale model technology has made great progress. Large oracle models, represented by ChatGPT, have continuously broken through the AI ​​ceiling in areas such as dialogue and text generation. At the same time, some large speech models have achieved outstanding results in areas such as speech synthesis and speech translation, and can realize realistic acoustic features in speech. On the other hand, more and more companies now have the need to go global, and cultural export is also booming. Non-English speaking countries are now more favored globally, but this has also increased the difficulty of communication. Therefore, mastering foreign languages ​​is not limited to daily communication and personal learning. Good speech technology is needed in film and television, e-commerce, cultural media and other fields.

[0071] In particular, there is a focus on generating speech data in other accents from a piece of speech data. However, existing speech generation technologies often simply generate speech data with specific accent requirements directly from the original speech data. This approach is rather crude, and the generated speech data is not accurate.

[0072] To address the problems of the prior art, embodiments of the present invention provide a method for generating accent data. This invention can be used in numerous general-purpose or specialized computing device environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor devices, distributed computing environments including any of the above devices, etc. It can also be applied to various computer terminals or smart terminals. The executing entity can be the processor or server of the computer terminal or smart terminal. The method flowchart is shown below. Figure 1 As shown, it specifically includes:

[0073] S1: Obtain raw voice data.

[0074] The raw speech data obtained in this application may be voice data directly emitted by a person, speech data emitted by speech synthesis software, or speech data emitted from a broadcaster.

[0075] This application does not limit the accent, rhythm, timbre, or length of the original speech data; it can be any speech data that can generate corresponding text data content.

[0076] S2: Generate first voice data corresponding to the original voice data according to the preset target accent.

[0077] The target accent preset in this application can be any dialect of Chinese, such as Sichuan accent, Cantonese accent, or Mandarin accent (this application considers Mandarin as an accent); it can also be any accent from a foreign language, because not only are there many different accents in Chinese, but foreign languages ​​also have various accents due to the influence of regional, cultural, and historical factors. For example, English has British accent, American accent, Australian accent, Indian accent, etc. This embodiment does not limit this.

[0078] The purpose of this solution is to generate speech data belonging to another accent from original speech data belonging to one accent. This is referred to as target accent data in this application, and the other accent is the target accent. For example, this can be achieved using a large speech model. The original speech data is input into a pre-acquired large speech model, and the target accent is set for the large speech model. The large speech model will output the first speech data under the target accent, corresponding to the original speech data. A large speech model with functions such as zero-shot speech cloning, speech editing, generation of standard pronunciation of the target language, and control of the duration of generated speech can be selected. Furthermore, during speech cloning, features such as timbre, emotion, and rhythm in the original speech data can be accurately transferred.

[0079] However, before generating the final desired speech data, this step first generates speech data corresponding to the original speech according to the target accent, which serves as the first speech data. Therefore, generating the first speech data in this step is to assist in the subsequent generation process of the target accent data.

[0080] S3: Slice the original speech data to obtain various original speech slices, and simultaneously slice the first speech data to obtain various first speech slices.

[0081] The methods for slicing the original speech data and the first speech data can be the same or different; slicing can be performed simultaneously, or the original speech data can be sliced ​​first and then the first speech data can be sliced, or vice versa. Slicing can be performed using a voice activity detection (VAD) method combined with the semantic information corresponding to the speech.

[0082] Taking raw speech data as an example, other methods of slicing include:

[0083] 1) Based on the voice endpoint.

[0084] Understandably, the energy of speech data differs significantly between parts with and without speech. Therefore, the original speech signal can be segmented into frames, the short-time energy of each frame can be calculated, and an energy threshold can be set. The short-time energy of each frame can be compared with the energy threshold. When the short-time energy of a frame exceeds the energy threshold, speech is considered to have started. When the short-time energy of a frame is below the energy threshold and lasts for a period of time (e.g., a few seconds), speech is considered to have ended. This method can be used to determine one or more speech endpoints in the original speech data. By segmenting from each speech endpoint, the original speech data can be divided into sentences, and each sentence can be regarded as a slice.

[0085] 2) Based on the prosodic features of speech.

[0086] Generally, the prosodic features of speech include intonation and stress. The prosodic features usually change between sentences or phrases. For example, the intonation may drop at the end of a declarative sentence and rise at the end of an interrogative sentence. Therefore, by analyzing the intonation and stress of the raw speech data, the prosodic boundaries can be determined, and this can be used for segmentation.

[0087] 3) Based on voice and text data.

[0088] In addition to directly slicing the raw speech data, it is also possible to slice the corresponding text data. First, the raw speech data is converted into text data, and then the words and sentences of the text data are identified and segmented based on the word and sentence division points.

[0089] This application abandons the existing technology of directly generating other accent data from a long speech segment. Instead, it segments the original speech data and the first speech data separately. This reduces the size of the original speech data. Starting from the smaller part, other accent segments can be generated for each segment, resulting in more accurate and detailed data.

[0090] S4: Each concatenated speech slice is determined by each of the original speech slices and each of the first speech slices.

[0091] The purpose of concatenating the original speech slice and the first speech slice is to enable the accent recognition model to better identify accents. The first speech slice provides a comparison, allowing the model to obtain richer accent comparison information. Since the first speech slice is generated according to the target accent and is identical in all aspects except for the accent, the original speech slice and the first speech slice will only form a sharp contrast in terms of accent, highlighting the characteristics of the original speech slice. In subsequent processing, the model can more accurately identify the accent features of the original speech slice.

[0092] There are various ways to splice the audio, such as horizontal splicing, where the original audio slice is spliced ​​after the first audio slice, or the first audio slice is spliced ​​after the original audio slice. Vertical splicing is also possible, such as using the Connectionist Temporal Classification (CTC) method for vertical alignment splicing.

[0093] S5: Input each of the spliced ​​speech segments into a pre-trained accent recognition model to obtain accent labels for each of the spliced ​​speech segments; the accent recognition model is trained using a speech sample set as training samples and the accent labels of each spliced ​​sample segment in the speech sample set as sample labels.

[0094] The model can employ a binary classification model, a supervised learning model primarily used to divide data into two categories. This is because the original speech slice and the first speech slice may be speech slices of the same accent or slices of two different accents. Using a binary classification model can quickly and accurately identify spliced ​​speech slices. However, although the accent recognition model pre-trained in this application processes the input spliced ​​speech slices and outputs the accent labels corresponding to the spliced ​​speech slices, in reality, the accent labels output by the model can be regarded as the accent labels of the original speech slices, that is, whether the original speech slice is a different accent or the same accent as the first accent slice. Specifically, the binary classification model can be a logistic regression model, a support vector machine model, a decision tree model, etc., and this application does not impose any restrictions on this.

[0095] Labels can be in the form of numbers or text:

[0096] In numerical terms, if the output of this accent recognition model is set to be in digital form, then the accent labels include two types: 1 and 0. "1" indicates that the two segments in the concatenated speech slice recognized by the model belong to two different accents, and "0" indicates that the two segments belong to the same accent. In one example, taking the first speech slice as the target accent, if the accent label of one of the concatenated speech slices output by the accent recognition model is 0, it means that the accent of the original speech slice in the concatenated speech slice belongs to the target accent. If the accent label of the concatenated speech slice is 1, it means that the accent of the original speech slice in the concatenated speech slice does not belong to the target accent, but belongs to another accent.

[0097] In terms of text, if the output of this accent recognition model is set to text, then the accent labels also include two types: yes and no. "Yes" means that the two segments in the spliced ​​speech segment recognized by the model belong to two different accents, and "No" means that the two segments belong to the same accent.

[0098] The speech sample set contains countless spliced ​​sample slices. A high-performing accent recognition model can be trained using this speech sample set. In this case, an accent recognition model can be trained for a specific accent. For example, the accent recognition model in this application is trained for the target accent, so its output is also related to the target accent.

[0099] S6: Generate target accent data based on each of the original speech segments and the accent tags of each of the spliced ​​speech segments and the target accent.

[0100] The accent labels of the concatenated speech segments have been clearly defined in the above steps, that is, it is clear whether each original speech segment belongs to the target accent. Then, operations such as conversion can be performed on the original speech segments that belong to the target accent or not. For example, for one or more original speech segments that do not belong to the target accent, corresponding segments belonging to the target accent can be generated and then replaced, so as to achieve the purpose of replacing and converting the original speech data. All the resulting segments belong to the speech data of the target accent and are combined to form the target accent data.

[0101] As can be seen from the above technical solution, this application first generates first speech data corresponding to the original speech data according to the preset target accent. This can be regarded as standard speech data under the target accent, which is directly generated speech data. Then, the original speech data and the first speech data are sliced ​​to divide each of them into multiple small parts. This is to reduce the size of the original speech data. Generating accent data from small parts will be more detailed and accurate. Among them, the slices are spliced ​​to form various spliced ​​speech slices. Furthermore, this application pre-trains an accent recognition model to slice the spliced ​​speech. When a speech slice is input into this accent recognition model, it can process the spliced ​​speech slice, predict the accent label of the spliced ​​speech slice, and thus determine the accent information corresponding to the spliced ​​speech slice. This allows the determination of the accent of each original speech slice. Then, based on each original speech slice, accent label, and target accent, the target accent data can be generated. Compared with the existing technology, the generation method of this application is more refined. In the process of directly generating accent data according to the original speech data, it may be affected by the complex and variable global features in the original speech data, resulting in poor generation effect. The slice format of this application can effectively avoid this problem.

[0102] On the other hand, this application can also be applied to the field of speech correction, such as correcting raw speech data into speech data with a target accent.

[0103] Optionally, the step of generating the first voice data corresponding to the original voice data according to a preset target accent in this application is as follows:

[0104] Extract the timbre, prosodic, and emotional features from the original speech data;

[0105] Determine the text data corresponding to the original speech data;

[0106] Based on the text data, and using the target accent, timbre features, prosodic features, and emotional features as standards, first speech data corresponding to the original speech data is generated.

[0107] Specifically, to ensure a more efficient generation process for the subsequent target accent data, when initially generating the first speech data corresponding to the original speech data according to the target accent, only the accent is considered, while other speech features remain unchanged. Therefore, the speech features of the original speech data can be extracted first, including but not limited to timbre features, prosodic features, and emotional features. In addition, the corresponding text data needs to be determined. A large language model can be used to recognize the text, or automatic speech recognition technology (ASR) can be used to recognize the text. Then, according to the text data, the first speech data corresponding to the original speech data can be generated based on the target accent, timbre features, prosodic features, and emotional features. For example, an existing large speech model can be called, and the text data, target accent, timbre features, prosodic features, and emotional features can all be input to output the first speech data.

[0108] Optionally, the above process ensures that the original speech data and the target accent data are identical except for the accent. Therefore, the acoustic features of the original speech data, such as timbre, rhythm, and emotion, are preserved, and only the accent is changed. For example, a user records their original speech data with a Sichuan accent and wants to generate target accent data with a Mandarin accent based on this original speech data, keeping everything else unchanged except for the accent. They also want to generate target accent data with a Northeastern accent based on this original speech data. This can meet the multiple accent generation needs of the same user.

[0109] In addition, acoustic features such as timbre, rhythm, and emotion can be arbitrarily changed or set to match different voice application scenarios. This can meet the needs of multiple users in multiple different scenarios, making the usability / application range of the original voice data wider and improving the user experience.

[0110] After obtaining the first speech data, it is necessary to slice the original speech data to obtain various original speech slices. At the same time, the first speech data is also sliced ​​to obtain various first speech slices. The same method can be used to slice both the original speech data and the first speech data. The following is a detailed explanation using the original speech data as an example.

[0111] Identify the pause locations in the raw speech data;

[0112] The original speech data is divided into first original slices according to each of the pause positions;

[0113] Determine the text data corresponding to the original speech data;

[0114] Identify each punctuation mark in the text data;

[0115] Each punctuation mark is matched with the original speech data to determine the position of each punctuation mark in the original speech data, which is then used as the position of each punctuation mark.

[0116] Determine whether the positions of each punctuation mark in the original speech data coincide with the pause positions;

[0117] Each non-overlapping punctuation mark is used as a segmentation point.

[0118] Each of the positions to be segmented in each of the first original slices is segmented to obtain each original speech slice.

[0119] Specifically, in order to ensure sufficient segmentation and that the resulting original speech slices are not abrupt, thus maintaining the authenticity of the original speech data, segmentation can be performed from two reference points: pauses at the speech level and punctuation marks at the text data level, to prevent omission of points that could have been segmented.

[0120] First, pauses at the speech level are used as reference points. It's understood that pauses exist in real speech data. Therefore, starting from these pauses, the original speech data is identified, and then segmented into first-level original slices based on these pauses. Next, the corresponding text data is determined, and punctuation marks are identified. Punctuation marks are used to separate sentences in the text and also serve as transitions, so using punctuation marks to assist segmentation can effectively and accurately produce slices. It's important to note that pauses and punctuation marks determined from the two reference points may overlap. Segmenting from each reference point separately could result in duplicate segments, making the process cumbersome and time-consuming. Therefore, after segmenting by pauses, it's determined whether there are overlapping pauses and punctuation marks. Only the non-overlapping punctuation marks are then segmented, thus improving efficiency.

[0121] To ensure the quality and accuracy of the obtained raw speech segments, verification can be performed after segmentation, which further guarantees the quality and accuracy of the final generated target accent data.

[0122] The method provided in this embodiment of the invention, which involves determining each concatenated speech slice from each of the original speech slices and each of the first speech slices, is described in detail below:

[0123] The original speech slices are combined to form an original speech sequence, and the sequence position of each original speech slice is determined as the original sequence position.

[0124] Each of the first speech slices is combined to form a first speech sequence, and the sequence position of each of the first speech slices is determined as the first sequence position;

[0125] Based on the original sequence position and the first sequence position, each original speech slice and each first speech slice are mapped to another;

[0126] Align and concatenate the corresponding original speech slices with the first speech slice to obtain each concatenated speech slice.

[0127] Specifically, since each first speech slice differs from each original speech slice only in that the accent is replaced, while everything else remains unchanged, for each original speech slice, there must exist a corresponding first speech slice, and their positions in the original and first speech data are identical. Furthermore, because slicing has already been performed, both the original and first speech slices are independent, allowing them to be recombine into a sequence. Their original positions in the speech data become their current sequence positions. Original and first speech slices with the same sequence position are considered corresponding, and the number of corresponding text data items is also the same. Then, alignment and concatenation are performed. Figure 2 As shown, Figure 2 This refers to the presentation of the original speech slice and the first speech slice after being spliced ​​together. They are spliced ​​one above the other, or the first speech slice can be placed on top and the original speech slice on the bottom. This embodiment does not restrict this. This splicing makes the spliced ​​speech slices look neater and is easier for the accent recognition model to process.

[0128] The process of assigning correspondences between each original speech slice and each first speech slice based on the original sequence position and the first sequence position may include:

[0129] For each of the original speech segments, a first speech segment whose first sequence position is the same as the original sequence position of the original speech segment is determined;

[0130] Obtain the text data of the first speech slice, and at the same time obtain the text data of the original speech slice;

[0131] Determine whether the text data of the first speech slice is completely consistent with the text data of the original speech slice;

[0132] If so, then the first speech slice is taken as the first speech slice corresponding to the original speech slice.

[0133] Specifically, please refer to Figure 2The original sequence position of the original speech slice "A1B1C1D1F1" in the original speech data is determined to be the same as the first sequence position of the first speech slice "A2B2C2D2F2" in the first speech data. In this example, A1 and A2 are assumed to be the same text, only to distinguish between the original speech slice and the first speech slice. It can be seen that the two corresponding text characters within a dotted line, such as A1 and A2, are identical. Therefore, the text data of the first speech slice is completely consistent with the text data of the original speech slice, and thus they correspond and can be concatenated. Furthermore, to enhance accuracy, the text data comparison process can also simultaneously determine whether the number of characters is the same, for example... Figure 2 The original speech slice has 5 characters, and the first speech slice also has 5 characters. Since each character is identical, these two slices are corresponding.

[0134] To enable a pre-trained accent recognition model to process spliced ​​speech data, a large number of slice-type samples need to be input into the model during training. This allows the model to learn the similarities and differences between two types of slices within the sample. Below is a preferred method for constructing a speech sample set:

[0135] Multiple voice sample data are extracted from a pre-defined target accent database and used as the first sample data;

[0136] Multiple voice sample data are extracted from a pre-determined database of other accents and used as second sample data.

[0137] Each of the first sample data is sliced ​​to obtain a first sample slice, and each of the second sample data is sliced ​​to obtain a second sample slice.

[0138] For each of the first sample slices, a first target accent sample corresponding to the first sample slice is generated according to the target accent.

[0139] Each of the first sample slices is aligned and spliced ​​with its corresponding first target accent sample to obtain each first spliced ​​sample.

[0140] For each second sample slice, generate a second target accent sample corresponding to the second sample slice according to the target accent;

[0141] Each second sample slice is aligned and spliced ​​with its corresponding second target accent sample to obtain each spliced ​​second sample.

[0142] Each of the first spliced ​​samples and each of the second spliced ​​samples are used as spliced ​​sample slices and summarized to obtain a speech sample set.

[0143] Specifically, the accent recognition model of this application enables the accent recognition model to identify whether the accent of the original speech slice in the input spliced ​​speech slice is the same as the accent of the first speech slice. Therefore, the sample data used in training the accent recognition model needs to include sample data with the same accent as the first speech slice, and also sample data with different accents. So, sample data are collected for the accent of the first speech slice (i.e., the target accent) and other accents respectively to determine the target accent database, which contains all speech data under the target accent. At the same time, other accent databases are determined, which contain speech data under various accents other than the target accent. Then, multiple sample data are extracted from these two databases as the first sample data and the second sample data, and then the slice is performed. In order to enrich the sample data, the sample data can be spliced ​​separately.

[0144] First, there is the first sample slice. Although the accent of the first sample slice is the target accent, it is real data obtained from the target accent database and may not be the standard target accent. Therefore, it is necessary to use the data of the corresponding standard target accent to splice it. Thus, the first target accent sample corresponding to the first sample slice is generated according to the target accent, and then the first sample slice and the first target accent sample are spliced ​​together.

[0145] Next is the second sample slice, which can be any accent. Therefore, it is necessary to generate a second target accent sample corresponding to the second sample slice according to the target accent, and then splice them together.

[0146] In summary, if the target accent is represented as 'a', other accents are represented as 'b', and the accent of the first sample slice is represented as 'a'', then according to the above method for constructing the speech sample set, the spliced ​​sample slices in the constructed speech sample set include two types, represented by accents, namely "a'-a" and "ba".

[0147] After the accent recognition model identifies the accent tags of the original speech segments in the concatenated speech segments, the target accent data can be generated based on each original speech segment, the accent tags of each concatenated speech segment, and the target accent. The specific process is as follows:

[0148] Each concatenated speech slice with the accent label set to the preset first label is used as a first accent concatenated slice;

[0149] Determine the original speech slice in each of the first accent splicing slices;

[0150] Generate a target accent slice corresponding to the original speech slice in each of the first accent splicing slices according to the target accent;

[0151] Each concatenated speech slice with the accent label set to the preset second label is used as a second accent concatenation slice;

[0152] Identify the original speech slice in each of the second accent splicing slices;

[0153] The target accent slice corresponding to the original speech slice in each of the first accent splicing slices is combined with the original speech slice in each of the second accent splicing slices to generate target accent data.

[0154] Specifically, following the example above, the first and second labels can be set to "1, 0" or "Yes, No". In one example, the first label is "1" and the second label is "0", meaning that the original speech slice in the concatenated speech slice with the accent label "1" does not belong to the same accent as the first speech slice, while the original speech slice in the concatenated speech slice with the accent label "0" belongs to the same accent as the first speech slice. If this example wants to generate speech data under the target accent, then the original speech slice in the concatenated speech slice with the accent label "1" needs to be converted into a slice under the target accent. That is, generate a target accent slice corresponding to the original speech slice according to the target accent. This ensures that all speech slices corresponding to the original speech data are under the target accent. The concatenated speech slice with the label "0" is used as the second accent concatenated slice. The original speech slice in the second accent concatenated slice is combined with the target accent slice that has been converted to the target accent to generate the target accent data.

[0155] In one example, the process of generating a target accent slice corresponding to the original speech slice according to the target accent may include: establishing a conversion mechanism to determine the accent of the original speech data, analyzing the differences between the original accent and the target accent, such as focusing on phonemes and intonation; in terms of phonemes, performing phoneme conversion and creating a phoneme mapping table to map the phonemes of the original accent to the corresponding phonemes of the target accent; in terms of intonation, determining the intonation patterns of multiple different sentence structures such as declarative sentences, interrogative sentences, and rhetorical questions in the target accent, such as the variation rules between rising, falling, or inflectional tones; and finally generating a target accent slice corresponding to the original speech slice according to the factor mapping table and intonation patterns.

[0156] Optionally, this application can also achieve accent mixing. Since it is clear whether the accent of each original speech slice is the same as the target accent, the original speech slice can be arbitrarily converted according to the requirements. For example, some original speech slices with accents other than the target accent can be converted to the target accent, or all or some original speech slices with the target accent can be converted to other accents. Alternatively, a gradual accent conversion can be set to create unique speech effects. For example, in the fields of film and television dubbing and audiobooks, this method can produce speech data with special styles, adding artistic color to the work.

[0157] Since each original speech slice is relatively independent, when the target accent slice generated from a certain original speech slice does not meet expectations, the original speech slice can be directly reprocessed without reprocessing the entire target accent data. In addition, generating target accent data in the form of slices can realize error location and correction. For example, if an error occurs in the final target accent data generation process, the slice form adopted in this application can more easily locate the specific location where the error occurred.

[0158] Preferably, the target accent slices generated from each original speech slice can be evaluated separately to more meticulously control the quality of accent conversion. For example, the authenticity and naturalness of each original speech slice before and after conversion can be compared separately, thereby optimizing one or more target accent slices in a targeted manner. This helps to improve the overall quality of accent conversion and make the generated target accent data more detailed, natural and accurate.

[0159] and Figure 1 Corresponding to the method described above, embodiments of the present invention also provide an accent data generation apparatus for processing... Figure 1 In a specific implementation of the method, the accent data generation device provided in this embodiment of the invention can be used in a computer terminal or various mobile devices, combined with Figure 3 The device for generating accent data is introduced, such as... Figure 3 As shown, the device may include:

[0160] Raw speech data acquisition module 10 is used to acquire raw speech data;

[0161] The first speech data generation module 20 is used to generate first speech data corresponding to the original speech data according to a preset target accent.

[0162] The slicing module 30 is used to slice the original speech data to obtain various original speech slices, and at the same time slice the first speech data to obtain various first speech slices.

[0163] The spliced ​​speech slice determination module 40 is used to determine each spliced ​​speech slice from each of the original speech slices and each of the first speech slices;

[0164] The accent labeling module 50 is used to input each of the spliced ​​speech slices into a pre-trained accent recognition model to obtain the accent labels of each of the spliced ​​speech slices; the accent recognition model is trained using a speech sample set as training samples and the accent labels of each spliced ​​sample slice in the speech sample set as sample labels.

[0165] The target accent data generation module 60 is used to generate target accent data based on each of the original speech slices and the accent tags of each of the spliced ​​speech slices and the target accent.

[0166] As can be seen from the above technical solution, this application first generates first speech data corresponding to the original speech data according to the preset target accent. This can be regarded as standard speech data under the target accent, which is directly generated speech data. Then, the original speech data and the first speech data are sliced ​​to divide each of them into multiple small parts. This is to reduce the size of the original speech data. Generating accent data from small parts will be more detailed and accurate. Among them, the slices are spliced ​​to form various spliced ​​speech slices. Furthermore, this application pre-trains an accent recognition model to slice the spliced ​​speech. When a speech slice is input into this accent recognition model, it can process the spliced ​​speech slice, predict the accent label of the spliced ​​speech slice, and thus determine the accent information corresponding to the spliced ​​speech slice. This allows the determination of the accent of each original speech slice. Then, based on each original speech slice, accent label, and target accent, the target accent data can be generated. Compared with the existing technology, the generation method of this application is more refined. In the process of directly generating accent data according to the original speech data, it may be affected by the complex and variable global features in the original speech data, resulting in poor generation effect. The slice format of this application can effectively avoid this problem.

[0167] Furthermore, embodiments of this application provide an apparatus for generating accent data. Optionally, Figure 4 The hardware structure block diagram of the accent data generation device is shown, with reference to... Figure 4 The hardware structure of the accent data generation device may include: at least one processor 01, at least one communication interface 02, at least one memory 03, and at least one communication bus 04.

[0168] In this embodiment, the number of processor 01, communication interface 02, memory 03 and communication bus 04 is at least one, and processor 01, communication interface 02 and memory 03 communicate with each other through communication bus 04.

[0169] Processor 01 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0170] Memory 03 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device.

[0171] The memory stores a program that the processor can call. The program is used to execute the following method for generating accent data, including:

[0172] Acquire raw speech data;

[0173] Generate first voice data corresponding to the original voice data according to the preset target accent;

[0174] The original speech data is sliced ​​to obtain various original speech slices, and the first speech data is sliced ​​to obtain various first speech slices.

[0175] Each spliced ​​speech slice is determined by each of the original speech slices and each of the first speech slices;

[0176] Each of the spliced ​​speech segments is input into a pre-trained accent recognition model to obtain accent labels for each of the spliced ​​speech segments; the accent recognition model is trained using a speech sample set as training samples and the accent labels of each spliced ​​speech segment in the speech sample set as sample labels.

[0177] Based on each of the original speech segments, and based on the accent tags of each of the spliced ​​speech segments and the target accent, target accent data is generated.

[0178] Optionally, the refined and extended functions of the program can be found in the description of the method for generating accent data in the method embodiments.

[0179] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor. When the program runs, it controls the device where the storage medium is located to execute the following method for generating accent data, including:

[0180] Acquire raw speech data;

[0181] Generate first voice data corresponding to the original voice data according to the preset target accent;

[0182] The original speech data is sliced ​​to obtain various original speech slices, and the first speech data is sliced ​​to obtain various first speech slices.

[0183] Each spliced ​​speech slice is determined by each of the original speech slices and each of the first speech slices;

[0184] Each of the spliced ​​speech segments is input into a pre-trained accent recognition model to obtain accent labels for each of the spliced ​​speech segments; the accent recognition model is trained using a speech sample set as training samples and the accent labels of each spliced ​​speech segment in the speech sample set as sample labels.

[0185] Based on each of the original speech segments, and based on the accent tags of each of the spliced ​​speech segments and the target accent, target accent data is generated.

[0186] Specifically, the storage medium can be a computer-readable storage medium, which can be an electronic storage device such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM.

[0187] Optionally, the refined and extended functions of the program can be found in the description of the method for generating accent data in the method embodiments.

[0188] Furthermore, the functional modules in the various embodiments of this disclosure can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part. If the function is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a live streaming device, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this disclosure.

[0189] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0190] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0191] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating accent data, characterized in that, include: Acquire raw voice data; Generate first voice data corresponding to the original voice data according to the preset target accent; The original speech data is sliced ​​to obtain various original speech slices, and the first speech data is sliced ​​to obtain various first speech slices. Determining each concatenated speech segment from each of the original speech segments and each of the first speech segments includes: assembling each of the original speech segments into an original speech sequence and determining the sequence position of each of the original speech segments as the original sequence position; assembling each of the first speech segments into a first speech sequence and determining the sequence position of each of the first speech segments as the first sequence position; based on the original sequence position and the first sequence position, corresponding each of the original speech segments and each of the first speech segments; and aligning and concatenating the corresponding original speech segments with the first speech segments to obtain each concatenated speech segment. Each of the spliced ​​speech segments is input into a pre-trained accent recognition model to obtain accent labels for each of the spliced ​​speech segments; the accent recognition model is trained using a speech sample set as training samples and the accent labels of each spliced ​​speech segment in the speech sample set as sample labels. Based on each of the original speech segments, and based on the accent tags of each of the spliced ​​speech segments and the target accent, target accent data is generated.

2. The method according to claim 1, characterized in that, The step of mapping each original speech slice and each first speech slice to another based on the original sequence position and the first sequence position includes: For each of the original speech segments, a first speech segment whose first sequence position is the same as the original sequence position of the original speech segment is determined; Obtain the text data of the first speech slice, and at the same time obtain the text data of the original speech slice; Determine whether the text data of the first speech slice is completely consistent with the text data of the original speech slice; If so, then the first speech slice is taken as the first speech slice corresponding to the original speech slice.

3. The method according to claim 1, characterized in that, The process of establishing the speech sample set includes: Multiple voice sample data are extracted from a pre-defined target accent database and used as the first sample data; Multiple voice sample data are extracted from a pre-determined database of other accents and used as second sample data. Each of the first sample data is sliced ​​to obtain a first sample slice, and each of the second sample data is sliced ​​to obtain a second sample slice. For each of the first sample slices, a first target accent sample corresponding to the first sample slice is generated according to the target accent. Each of the first sample slices is aligned and spliced ​​with its corresponding first target accent sample to obtain each first spliced ​​sample. For each second sample slice, generate a second target accent sample corresponding to the second sample slice according to the target accent; Each second sample slice is aligned and spliced ​​with its corresponding second target accent sample to obtain each spliced ​​second sample. Each of the first spliced ​​samples and each of the second spliced ​​samples are used as spliced ​​sample slices and summarized to obtain a speech sample set.

4. The method according to claim 1, characterized in that, The step of generating target accent data based on each of the original speech segments and the accent tags of each of the concatenated speech segments and the target accent includes: Each concatenated speech slice with the accent label set to the preset first label is used as a first accent concatenated slice; Determine the original speech slice in each of the first accent splicing slices; Generate a target accent slice corresponding to the original speech slice in each of the first accent splicing slices according to the target accent; Each concatenated speech slice with the accent label set to the preset second label is used as a second accent concatenation slice; Identify the original speech slice in each of the second accent splicing slices; The target accent slice corresponding to the original speech slice in each of the first accent splicing slices is combined with the original speech slice in each of the second accent splicing slices to generate target accent data.

5. The method according to any one of claims 1 to 4, characterized in that, The step of slicing the original speech data to obtain individual original speech slices includes: Identify the pause locations in the raw speech data; The original speech data is divided into first original slices according to each of the pause positions; Determine the text data corresponding to the original speech data; Identify each punctuation mark in the text data; Each punctuation mark is matched with the original speech data to determine the position of each punctuation mark in the original speech data, which is then used as the position of each punctuation mark. Determine whether the positions of each punctuation mark in the original speech data coincide with the pause positions; Each non-overlapping punctuation mark is used as a segmentation point. Each of the positions to be segmented in each of the first original slices is segmented to obtain each original speech slice.

6. The method according to any one of claims 1 to 4, characterized in that, The step of generating first speech data corresponding to the original speech data according to a preset target accent includes: Extract the timbre, prosodic, and emotional features from the original speech data; Determine the text data corresponding to the original speech data; Based on the text data, and using the target accent, timbre features, prosodic features, and emotional features as standards, first speech data corresponding to the original speech data is generated.

7. An apparatus for generating accent data, characterized in that, include: The raw speech data acquisition module is used to acquire raw speech data; The first speech data generation module is used to generate first speech data corresponding to the original speech data according to a preset target accent. The slicing module is used to slice the original speech data to obtain various original speech slices, and at the same time slice the first speech data to obtain various first speech slices. A speech segment concatenation determination module is used to determine concatenated speech segments from each of the original speech segments and each of the first speech segments; including: assembling the original speech segments into an original speech sequence and determining the sequence position of each original speech segment as the original sequence position; assembling the first speech segments into a first speech sequence and determining the sequence position of each first speech segment as the first sequence position; based on the original sequence position and the first sequence position, corresponding each of the original speech segments and each of the first speech segments; and aligning and concatenating the corresponding original speech segments with the first speech segments to obtain concatenated speech segments. The accent labeling module is used to input each of the spliced ​​speech segments into a pre-trained accent recognition model to obtain accent labels for each of the spliced ​​speech segments; the accent recognition model is trained using a speech sample set as training samples and the accent labels of each spliced ​​sample segment in the speech sample set as sample labels. The target accent data generation module is used to generate target accent data based on each of the original speech slices and the accent tags of each of the spliced ​​speech slices and the target accent.

8. A device for generating accent data, characterized in that, Including memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement the steps of the method for generating accent data as described in any one of claims 1-6.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for generating accent data as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Speech clone model training method and device, speech synthesis method and device and related equipment

    CN116403558A

  • Voice conversion method and device, electronic equipment and storage medium

    CN117935772A