Voice data screening method and device, electronic equipment and readable storage medium

By quantitatively evaluating and screening dialect speech data and standard speech data, the problem of poor dialect speech conversion effect was solved, achieving efficient speech conversion data acquisition and reducing manual screening costs.

CN114758664BActive Publication Date: 2025-12-12VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210365542.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-06
Publication Date
2025-12-12
Estimated Expiration
2042-04-06

AI Technical Summary

Technical Problem

Existing speech conversion technology has poor conversion results when converting dialects, due to a lack of sufficient dialect speech data.

Method used

Speech conversion is performed based on T dialect speech data and standard speech data corresponding to the target speaker. Objective quantitative evaluation is carried out using speech recognition comparison results and audio information comparison results. Target dialect speech data with better conversion effect is selected, and a large amount of conversion data is obtained through style transfer.

Benefits of technology

It improved the efficiency of dialect speech conversion, reduced the workload of manual screening, saved labor costs, and obtained high-quality conversion data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114758664B_ABST
    Figure CN114758664B_ABST
Patent Text Reader

Abstract

The application discloses a voice data screening method and device, electronic equipment and a readable storage medium, wherein the method comprises: obtaining first conversion data based on T dialect voice data and selected standard voice data corresponding to a target speaker; determining target dialect voice data based on at least one of a first information comparison result of the T dialect voice data and the first conversion data and a first determination result of the target speaker corresponding to the first conversion data; obtaining second conversion data based on P dialect voice data corresponding to the target dialect voice data and K standard voice data corresponding to the target speaker; and screening third conversion data from the second conversion data based on at least one of a second information comparison result of the P dialect voice data and the second conversion data and a second determination result of the target speaker corresponding to the second conversion data. The first information comparison result comprises at least one of a voice recognition comparison result and an audio information comparison result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of speech processing, and particularly relates to a speech data screening method and device, an electronic device and a readable storage medium. BACKGROUND

[0002] Speech conversion technology can retain text-related information of a source speaker, and replace the timbre of the source speaker's speech data with the timbre of another target speaker, which makes the speech conversion technology widely applied in the fields of speech broadcasting, intelligent translation and the like. With the development of speech technology, more and more users hope to provide a dialect version of speech conversion service, so a large amount of dialect speech data is needed. However, the dialect speech data is difficult to collect at present, and therefore the quantity of dialect speech data is usually small, which makes the current speech conversion technology have poor conversion effect when converting speech in a dialect. SUMMARY

[0003] The embodiments of the present application aim to provide a speech data screening method and device, an electronic device and a readable storage medium, which can solve the problem of poor conversion effect when converting speech in a dialect in the related art.

[0004] In a first aspect, the embodiments of the present application provide a speech data screening method, which comprises: obtaining first conversion data based on T pieces of dialect speech data and target speaker corresponding selected standard speech data, T being an integer greater than zero; processing the first conversion data based on at least one of a first information comparison result of the T pieces of dialect speech data and the first conversion data, and a first determination result of the first conversion data corresponding to the target speaker, and determining target dialect speech data from the T pieces of dialect speech data based on a first processing result; obtaining second conversion data based on P pieces of dialect speech data corresponding to the target dialect speech data and K pieces of standard speech data corresponding to the target speaker, P being greater than T, and K being an integer greater than zero; processing the second conversion data based on at least one of a second information comparison result of the P pieces of dialect speech data and the second conversion data, and a second determination result of the second conversion data corresponding to the target speaker, and screening third conversion data from the second conversion data based on a second processing result; wherein the first information comparison result comprises at least one of a speech recognition comparison result and an audio information comparison result.

[0005] In a second aspect, an embodiment of the present application provides a voice data screening device, the device comprising: a first conversion processing module configured to obtain first conversion data based on T pieces of dialect voice data and target speaker corresponding selected standard voice data, T being an integer greater than zero; a first screening processing module configured to process the first conversion data based on at least one of a first information comparison result of the T pieces of dialect voice data and the first conversion data, and a first determination result of the target speaker corresponding to the first conversion data, and determine target dialect voice data from the T pieces of dialect voice data based on a first processing result; a second conversion processing module configured to obtain second conversion data based on P pieces of dialect voice data corresponding to the target dialect voice data and K pieces of standard voice data corresponding to the target speaker, P being greater than T, and K being an integer greater than zero; and a second screening processing module configured to process the second conversion data based on at least one of a second information comparison result of the P pieces of dialect voice data and the second conversion data, and a second determination result of the target speaker corresponding to the second conversion data, and screen third conversion data from the second conversion data based on a second processing result; wherein the first information comparison result comprises at least one of a voice recognition comparison result and an audio information comparison result.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising a processor and a memory, the memory storing programs or instructions executable on the processor, and the programs or instructions being executed by the processor to implement steps of the method according to the first aspect.

[0007] In a fourth aspect, an embodiment of the present application provides a readable storage medium, the readable storage medium storing programs or instructions, and the programs or instructions being executed by a processor to implement steps of the method according to the first aspect.

[0008] In a fifth aspect, an embodiment of the present application provides a chip, the chip comprising a processor and a communication interface, the communication interface being coupled to the processor, and the processor being configured to run programs or instructions to implement the method according to the first aspect.

[0009] In a sixth aspect, an embodiment of the present application provides a computer program product, the program product being stored in a storage medium, and the program product being executed by at least one processor to implement the method according to the first aspect.

[0010] In the embodiment of the present application, according to the T dialect speech data and the selected standard speech data corresponding to the target speaker, first conversion data is obtained, wherein T is an integer greater than zero, the first conversion data retains the text information of the T dialect speech data, and the timbre of the T dialect speech data is changed to the timbre of the target speaker. After obtaining the first conversion data, at least one of the first information comparison result of the T dialect speech data and the first conversion data and the first determination result of the target speaker corresponding to the first conversion data is used to process the first conversion data, and the target dialect speech data is determined from the T dialect speech data according to the first processing result, wherein the first information comparison result includes at least one of the speech recognition comparison result and the audio information comparison result, and the first information comparison result and the first determination result are objective quantitative evaluation indexes, which can accurately evaluate the first conversion data to accurately select the target dialect speech data with better conversion effect from the T dialect speech data. Then, according to the target dialect speech data, P dialect speech data is determined, and the P dialect speech data and K standard speech data of the target speaker are subjected to speech conversion to obtain second conversion data. Further, at least one of the second information comparison result of the P dialect speech data and the second conversion data and the second determination result of the target speaker corresponding to the second conversion data is used to process the second conversion data, and the third conversion data is selected from the second conversion data according to the second processing result, and the second information comparison result and the second determination result are objective quantitative indexes, which can accurately evaluate the second conversion data to select the third conversion data with better conversion effect from the second conversion data. Thus, in the embodiment, by performing style transfer on a small amount of dialect speech data and a large amount of standard speech data, a large amount of conversion data using dialects can be obtained, and the third conversion data with better conversion effect can be automatically selected from the conversion data by using objective quantitative evaluation indexes, the data quality of the third conversion data is high, and the workload of subsequent manual selection of conversion data can be greatly reduced by automatic selection, saving labor costs. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 is a flowchart of the speech data selection method of the embodiment of the present application;

[0012] Figure 2 is a structural schematic diagram of the speech recognition model of the embodiment of the present application;

[0013] Figure 3 is a structural schematic diagram of the speaker recognition model of the embodiment of the present application;

[0014] Figure 4 is a structural schematic diagram of the speech conversion model of the embodiment of the present application;

[0015] Figure 5is a block diagram of a voice data screening device according to an embodiment of the present application;

[0016] Figure 6 is a hardware structure schematic diagram of an electronic device according to an embodiment of the present application Figure 1 ;

[0017] Figure 7 is a hardware structure schematic diagram of an electronic device according to an embodiment of the present application Figure 2 . DETAILED DESCRIPTION

[0018] The technical solutions of the embodiments of the present application will be described clearly below in conjunction with the drawings of the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art belong to the scope of protection of the present application.

[0019] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be exchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a category and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.

[0020] The voice data screening method provided by the embodiments of the present application will be described in detail below in conjunction with the drawings, specific embodiments and application scenarios.

[0021] Figure 1 A flowchart of a voice data screening method according to an embodiment of the present application is shown, the method is applied to an electronic device, comprising:

[0022] Step 101: Based on T dialect voice data and target speaker corresponding selected standard voice data, obtain first conversion data, T is an integer greater than zero.

[0023] In this step, the dialect speech data refers to speech data using the target dialect; the target speaker refers to a speaker object corresponding to the target conversion voice; the standard speech data refers to the Mandarin audio of the target speaker; and the selected standard speech data refers to the Mandarin audio of a specific target speaker. For example, the standard speech data of the target speaker is randomly selected to determine the selected standard speech data, wherein the target dialect and the target speaker can be pre-specified data for the user. Speech conversion is performed on the T pieces of dialect speech data and the selected standard speech data to obtain first conversion data, which retains the text information of the T pieces of dialect speech data and changes the voice to the voice of the target speaker.

[0024] Specifically, the speech conversion model is pre-trained, and the T pieces of dialect speech data and the selected standard speech data corresponding to the target speaker are input into the speech conversion model, and the speech conversion model outputs the first conversion data. The training data of the speech conversion model can be a small amount of dialect speech data and a large amount of standard speech data.

[0025] In an optional implementation, before obtaining the first conversion data based on the T pieces of dialect speech data and the selected standard speech data corresponding to the target speaker, the method further includes: determining N source dialect speakers based on the target dialect; and obtaining the T pieces of dialect speech data based on the N source dialect speakers. The source dialect speaker refers to a speaker object using the target dialect, and the T pieces of dialect speech data are a part of the total dialect speech data of the source dialect speaker. For example, the total dialect speech data corresponding to the N source dialect speakers is randomly extracted to obtain the T pieces of dialect speech data, wherein the T pieces of dialect speech data include T / N pieces of dialect speech data corresponding to each of the N source dialect speakers.

[0026] For example, it is determined that the target dialect is Cantonese, the target speaker is speaker A, N Cantonese speakers in the voice library are used as source dialect speakers, T / N pieces of dialect speech data of each Cantonese speaker are randomly selected, the selected standard speech data is randomly selected in the Mandarin audio of speaker A, the T pieces of dialect speech data corresponding to the N source dialect speakers and the selected standard speech data corresponding to the target speaker are input into the speech conversion model, and the first conversion data is obtained.

[0027] In step 102, the first conversion data is processed based on at least one of a first information comparison result of the T pieces of dialect speech data and the first conversion data and a first determination result that the first conversion data corresponds to the target speaker, and target dialect speech data is determined from the T pieces of dialect speech data based on a first processing result. The first information comparison result includes at least one of a speech recognition comparison result and an audio information comparison result.

[0028] In this step, the first information comparison result refers to a result obtained by comparing the T pieces of dialect speech data and the first converted data, wherein the first information comparison result includes at least one of a speech recognition comparison result and an audio information comparison result, the speech recognition comparison result is obtained by comparing speech recognition results of the T pieces of dialect speech data and the first converted data, and the audio information comparison result is obtained by comparing audio information of the T pieces of dialect speech data and the first converted data, wherein the audio information includes but is not limited to fundamental frequency information and first formant information. Because the timbre is changed in the speech conversion process, the T pieces of dialect speech data and the first converted data are comparable, so that the first information comparison result can be an index for objectively quantitatively evaluating the conversion effect of the first converted data.

[0029] The purpose of performing speech conversion on the T pieces of dialect speech data is to obtain dialect speech data of a target speaker's timbre, that is, without recording a large amount of dialect speech data, different timbre dialect speech data can be obtained through speech conversion, but the conversion effect of different dialect speech data is different, and in this embodiment, whether the first converted data belongs to the target speaker is judged to obtain a first determination result, and the first determination result can be an index for objectively quantitatively evaluating the conversion effect of the first converted data.

[0030] The first information comparison result and the first determination result can both objectively quantitatively evaluate the conversion effect of the first converted data, so that at least one of the first information comparison result and the first determination result can realize the evaluation of the first converted data, that is, when the first converted data is processed, the following optional implementation modes exist:

[0031] Implementation mode one, processing the first converted data based on the first information comparison result of the T pieces of dialect speech data and the first converted data.

[0032] Implementation mode two, processing the first converted data based on the first determination result that the first converted data corresponds to the target speaker.

[0033] Implementation mode three, processing the first converted data based on the first information comparison result of the T pieces of dialect speech data and the first converted data and the first determination result that the first converted data corresponds to the target speaker.

[0034] After processing the first conversion data, a first processing result is obtained, the first processing result is a result of objectively quantitatively evaluating the first conversion data, and therefore the first processing result can display conversion effects of different dialect speech data, and then the target dialect speech data can be determined from the T pieces of dialect speech data according to the first processing result, the target dialect speech data being dialect speech data with better conversion effect in speech conversion. Specifically, evaluation values of the T pieces of dialect speech data in the first processing result are determined, the evaluation values are sorted, and the dialect speech data with the highest evaluation value is determined as the target dialect speech data.

[0035] In step 103, second conversion data is obtained based on the P pieces of dialect speech data corresponding to the target dialect speech data and the K pieces of standard speech data corresponding to the target speaker, P is greater than T, and K is an integer greater than zero.

[0036] In this step, because the conversion effect of the target dialect speech data is better, more dialect speech data, i.e., the P pieces of dialect speech data, is determined according to the target dialect speech data, where P is greater than T. That is, a small amount of dialect speech data, i.e., the T pieces of dialect speech data, is used for speech conversion to preliminarily screen out the target dialect speech data, and then the P pieces of dialect speech data corresponding to the target dialect speech data are focused on, which effectively reduces the amount of dialect speech data for speech conversion and avoids generating a large amount of second conversion data, thereby effectively improving the conversion efficiency.

[0037] In a specific implementation, a target speaker corresponding to the target dialect speech data is determined, and P pieces of dialect speech data corresponding to the target speaker are determined, where the P pieces of dialect speech data can be all data of the target speaker using the target dialect. Different source speakers correspond to different conversion effects, and therefore by determining the target speaker with better conversion effect from the source dialect speakers, subsequent attention is no longer paid to other dialect speakers except the target speaker in the source dialect speakers, which can effectively improve the conversion efficiency and ensure the conversion effect.

[0038] The P pieces of dialect speech data and K pieces of standard speech data corresponding to the target speaker are converted to obtain second conversion data, wherein the second conversion data retains the text information of the P pieces of dialect speech data, and the timbre is changed to the timbre of the target speaker. Specifically, a speech conversion model is pre-trained, the P pieces of dialect speech data and the K pieces of standard speech data corresponding to the target speaker are input into the speech conversion model, and second conversion data output by the speech conversion model is obtained. Wherein K is an integer greater than zero, preferably, the data quantity of the K pieces of standard speech data is greater than the data quantity of the selected standard speech data, and the standard speech data of the target speaker is of the target timbre. However, the speaking style of the target speaker will be different when the standard speech data corresponds to different content, so by selecting more standard speech data, more speaking styles of the target timbre are obtained to obtain second conversion data with rich speaking styles.

[0039] In step 104, at least one of the second information comparison result of the P pieces of dialect speech data and the second conversion data and the second determination result of the second conversion data corresponding to the target speaker is used to process the second conversion data, and third conversion data is selected from the second conversion data based on the second processing result.

[0040] In this step, the second information comparison result refers to the result obtained by comparing the P pieces of dialect speech data and the second conversion data, wherein the second information comparison result can also include at least one of the speech recognition comparison result and the audio information comparison result. The speech recognition comparison result is obtained by comparing the speech recognition result of the P pieces of dialect speech data and the speech recognition result of the second conversion data, and the audio information comparison result is obtained by comparing the audio information of the P pieces of dialect speech data and the audio information of the second conversion data, wherein the audio information includes but is not limited to fundamental frequency information and first formant information. The P pieces of dialect speech data and the second conversion data have comparability, so that the second information comparison result can be an index for objectively quantifying and evaluating the conversion effect of the second conversion data.

[0041] The second conversion data is judged whether it belongs to the target speaker, and the second determination result is obtained. The second determination result can be an index for objectively quantifying and evaluating the conversion effect of the second conversion data.

[0042] The second information comparison result and the second determination result can objectively quantify and evaluate the conversion effect of the second conversion data, so that at least one of the second information comparison result and the second determination result can realize the evaluation of the second conversion data, that is, when the second conversion data is processed, the following optional implementation modes exist:

[0043] In the first implementation, the second conversion data is processed based on the second information comparison result of the P pieces of dialect speech data and the second conversion data.

[0044] In the second implementation, the second conversion data is processed based on the second determination result of the second conversion data corresponding to the target speaker.

[0045] In the third implementation, the second conversion data is processed based on the second information comparison result of the P pieces of dialect speech data and the second conversion data, and the second determination result of the second conversion data corresponding to the target speaker.

[0046] After the second conversion data is processed, a second processing result is obtained, the second processing result being a result of objectively quantitatively evaluating the second conversion data, so that the second processing result can display conversion effects of different dialect speech data, and third conversion data can be selected from the second conversion data according to the second processing result, the third conversion data having a better conversion effect.

[0047] Specifically, a conversion threshold is preset, an evaluation value of the P pieces of dialect speech data in the second processing result is determined, and the second conversion data with an evaluation value greater than the conversion threshold is selected as the third conversion data.

[0048] In the embodiments of the present application, the first conversion data is obtained by performing speech conversion on the T pieces of dialect speech data and the selected standard speech data corresponding to the target speaker. The first conversion data is objectively quantitatively evaluated by using at least one of the first information comparison result and the first determination result, so as to select the target dialect speech data. Then, the P pieces of dialect speech data corresponding to the target dialect speech data and the K pieces of standard speech data of the target speaker are subjected to speech conversion, so as to obtain the second conversion data. The second conversion data is objectively quantitatively evaluated by using at least one of the second information comparison result and the second determination result, so as to select the third conversion data from the second conversion data. In the embodiments, a small amount of dialect speech data is subjected to style transfer on a large amount of standard speech data, so as to obtain a large amount of conversion data using dialects. The third conversion data with a better conversion effect is automatically selected by using the objective quantitative evaluation index, the data quality of the third conversion data is high, and the workload of subsequent manual selection of the conversion data is greatly reduced, and the labor cost is saved.

[0049] In an embodiment of the present application, when the first information comparison result includes a speech recognition comparison result, before the step 102, the method further includes:

[0050] In the step 105, the first speech recognition result of the T pieces of dialect speech data is determined.

[0051] Step 106, determining a second speech recognition result of the first converted data.

[0052] Step 107, determining a speech recognition comparison result of the T pieces of dialect speech data and the first converted data based on the first speech recognition result and the second speech recognition result.

[0053] Wherein, the first speech recognition result is determined by performing speech recognition on the T pieces of dialect speech data, and the second speech recognition result is determined by performing speech recognition on the first conversion. The speech recognition comparison result of the T pieces of dialect speech data and the first converted data is determined by comparing the first speech recognition result and the second speech recognition result. The speech recognition comparison result can be an index for objectively quantifying and evaluating the conversion effect of the first converted data.

[0054] In a specific embodiment, a speech recognition model is pre-trained, and the speech recognition model is used to recognize speech data, that is, the T pieces of dialect speech data are input into the speech recognition model to obtain an output item of the speech recognition model, which is the first speech recognition result. The first speech recognition result is the text information obtained by recognizing the T pieces of dialect speech data. The first converted data is input into the speech recognition model to obtain an output item of the speech recognition model, which is the second speech recognition result. The second speech recognition result is the text information obtained by recognizing the first converted data. Since the voice conversion process is a change in timbre and does not change the speech content, the first speech recognition result and the second speech recognition result are comparable. By comparing the first speech recognition result and the second speech recognition result, the speech recognition comparison result is obtained.

[0055] In a possible implementation, the speech recognition model is obtained by training non-dialect speech training data. A large amount of non-dialect speech training data can be used to train a more accurate speech recognition model. Considering that voice conversion is a change in timbre, even if the speech recognition model is trained using non-dialect speech training data, it can still be used for speech recognition of dialect speech data and first converted data to obtain accurate first speech recognition results and second speech recognition results. The speech recognition model trained based on non-dialect speech training data is ingeniously used in dialect speech data, which provides the possibility of obtaining converted data with good conversion effect using a small amount of dialect speech data.

[0056] Specifically, a speech recognition model of a Transformer structure of CTC-attention (wherein, CTC is Connectionist temporal classification, and attention is an attention model) is built. For example, the model structure of the speech recognition model is as followsFigure 2 As shown, the speech recognition model is composed of an encoding network and a decoding network. The input features of the encoding network are input into a self-attention structure, a feature fusion structure (Concate & LayerNorm), a one-dimensional convolution structure (Conv1D), a feature fusion structure (Concate & LayerNorm), and the encoded features are output by a Softmax. The encoded features are input into the decoding network, and are input into a masked self-attention structure (Masked Self-Attention), a feature fusion structure (Concate & LayerNorm), a self-attention structure (Self-Attention), a feature fusion structure (Concate & LayerNorm), a one-dimensional convolution structure (Conv1D), a feature fusion structure (Concate & LayerNorm), and the recognition result is output by a Softmax. The speech recognition model can include 12 layers of the encoding network and 6 layers of the decoding network. The number of hidden layer neurons of the encoding network can be 2048, and the number of hidden neurons of the decoding part can be 6.

[0057] Further, the training data of the speech recognition model is non-dialectal speech training data. Before the speech training data is input into the speech recognition model, the speech training data is subjected to audio data processing to obtain audio features. For example, the speech training data is subjected to fbank (Filter Bank, a processing algorithm) feature extraction, i.e., pre-emphasis, framing, windowing, short-time Fourier transform, and Mel filtering to obtain fbank features. The dimension of the fbank features can be selected as 80, the frame length window length can be selected as 2048 sampling points, and the frame shift can be selected as 300 sampling points. The speech recognition model is trained with the speech training data until a preset training end condition is met. The preset training end condition includes that the number of training reaches a set value, such as 20w steps, or the loss function value of the validation set decreases to a stable value, or the character error rate of the recognition result is less than a set value, such as 9%.

[0058] Further, the step 107 determines the speech recognition comparison result of the T pieces of dialectal speech data and the first conversion data based on the first speech recognition result and the second speech recognition result, including:

[0059] Step 1071, determining the character error rate between the first speech recognition result and the second speech recognition result.

[0060] Step 1072, in the case that the character error rate is within a preset numerical range, determining the speech recognition comparison result of the T pieces of dialectal speech data and the first conversion data according to the character error rate.

[0061] Step 1073, in the case that the character error rate is not in the preset numerical range, deleting the dialect speech data corresponding to the character error rate not in the preset numerical range in the T pieces of dialect speech data.

[0062] Wherein, the first speech recognition result and the second speech recognition result are both text information, therefore, comparing the characters included in the text information of the first speech recognition result and the second speech recognition result, the character error rate between the first speech recognition result and the second speech recognition result can be determined. A preset numerical range is set in advance, within the preset numerical range, it indicates that the conversion effect is better, therefore, the speech recognition comparison result of the T pieces of dialect speech data and the first conversion data can be further determined according to the character error rate. In the case that the character error rate is not in the preset data range, it indicates that the conversion effect is too poor, therefore, the dialect speech data corresponding to the character error rate not in the preset numerical range is deleted in the T pieces of dialect speech data, effectively reducing the number of dialect speech data.

[0063] In a possible implementation, in the case that the character error rate is in the preset numerical range, the character error rate in the preset numerical range can be directly determined as the speech recognition comparison result of the T pieces of dialect speech data and the first conversion data. Of course, a calculation formula can also be set in advance, and the character error rate is further calculated according to the calculation formula to determine the speech recognition comparison result of the T pieces of dialect speech data and the first conversion data.

[0064] For example, the calculation formula corresponding to the speech recognition comparison result of the T pieces of dialect speech data and the first conversion data is as follows:

[0065]

[0066] Wherein, Score ASR characterizes the speech recognition comparison result of the T pieces of dialect speech data and the first conversion data; X j characterizes the jth piece of dialect speech data of the source dialect speaker X; T is the data amount of the dialect speech data corresponding to the source dialect speaker X; Y ref characterizes the random specific audio corresponding to the target speaker Y, that is, the selected standard speech data; characterizes the first conversion data obtained after the jth piece of dialect speech data of the source dialect speaker X is converted; CER characterizes the average character error rate; M ASR (X j ) characterizes the recognition result of the jth piece of dialect speech data of the source dialect speaker X by the speech recognition model, corresponding to the first speech recognition result; characterizes the recognition result of the first conversion data corresponding to the jth piece of dialect speech data by the speech recognition model, corresponding to the second speech recognition result.

[0067] By Score ASR The language information of the original dialect voice data after voice conversion can be reflected, and the unclear pronunciation of the voice conversion model can be filtered out. In this embodiment, the preset numerical range is set to be less than 1, and the CER is greater than or equal to 1. At this time, it is indicated that the conversion effect is too poor, and the dialect voice data of the source dialect speaker with CER greater than 1 is excluded.

[0068] In this embodiment, the voice recognition comparison result is accurately determined through the first voice recognition result and the second voice recognition result. When the conversion effect is good, the character error rate in the voice recognition comparison result is low, and when the conversion effect is poor, the character error rate in the voice recognition comparison result is high. Therefore, the voice recognition comparison result can be an index for objectively quantitatively evaluating the conversion effect of the first conversion data.

[0069] In an embodiment of the present application, when the first information comparison result includes the audio information comparison result, before step 102, the method further includes:

[0070] Step 108, determining a fundamental frequency comparison result based on the first fundamental frequency segment length of the T pieces of dialect voice data and the second fundamental frequency segment length of the first conversion data.

[0071] Step 109, determining a formant comparison result based on the first formant information of the T pieces of dialect voice data and the first formant information of the first conversion data.

[0072] Step 110, determining an audio information comparison result of the T pieces of dialect voice data and the first conversion data based on the fundamental frequency comparison result and the formant comparison result.

[0073] The first fundamental frequency segment length refers to the length result obtained by segmenting the fundamental frequency in the audio information of the T pieces of dialect voice data, and the second fundamental frequency segment length refers to the length result obtained by segmenting the fundamental frequency in the audio information of the first conversion data. The first fundamental frequency segment length and the second fundamental frequency segment length are compared to determine the fundamental frequency comparison result. The audio information not only includes the fundamental frequency information, but also includes the first formant information. The first formant information of the audio information of the dialect voice data and the first formant information of the audio information of the first conversion data are compared to determine the formant comparison result. The first conversion data is evaluated according to the fundamental frequency comparison result and the formant comparison result to obtain the audio information comparison result. The audio information comparison result can be an index for objectively quantitatively evaluating the conversion effect of the first conversion data.

[0074] For example, the calculation formula of the T pieces of dialect voice data and the first conversion data audio information comparison result is as follows:

[0075]

[0076]

[0077] wherein Score f characterizes the audio information contrast result; f0 characterizes the fundamental frequency; f1 characterizes the first formant; characterizes the length of the kth segment of the L fundamental frequency segments obtained by the dio algorithm, corresponding to the second fundamental frequency segment length; characterizes X i the length of the kth segment of the L fundamental frequency segments obtained by the dio algorithm, corresponding to the first fundamental frequency segment length; in the left formula and the relationship between the two corresponds to the fundamental frequency contrast result, by comparing the fundamental frequency segment lengths corresponding to the dialect speech data and the first converted data, the first converted data is evaluated. In the above formula ensure that x is within the range of 0 to 1, the closer x is to 0, the score used for evaluation presents a trend of nonlinear decline, the closer x is to 1, the closer the second fundamental frequency segment length of the first converted audio and the first fundamental frequency segment length of the dialect speech data, the audio rhythm and tone of the first converted audio are closer to the dialect speech data, and the conversion effect is better. characterizes the left derivative of the mth frequency of the first formant after framing, which has adjacent points; characterizes the right derivative of the mth frequency of the first formant after framing, which has adjacent points; characterizes X j the left derivative of the mth frequency of the first formant after framing, which has adjacent points; characterizes X j the right derivative of the mth frequency of the first formant after framing, which has adjacent points; in the right formula 1-x 2 form to ensure that x is within the range of 0 to 1, the closer to 1, the score used for evaluation presents a trend of nonlinear decline. By calculating the ratio of the difference between the left and right derivatives of the dialect speech data and the first converted data, the formant contrast result is determined, which represents the first resonance peak jitter of the first converted data compared to the dialect speech data at this point. Select M points with derivative difference of the dialect speech data less than the derivative difference of the first converted data for calculation, the closer the relationship value between the obtained derivatives is to 1, the closer the resonance peak waveform is to the dialect speech data, which can be regarded as no resonance peak jitter phenomenon, and the conversion effect is better.

[0078] In an embodiment of the present application, before step 102, the method further comprises:

[0079] Step 111, identifying the spectrum of the first converted data to determine the first prediction result of the first converted data corresponding to the target speaker.

[0080] Step 112, identifying the speaker of the first converted data to determine the second prediction result of the first converted data corresponding to the target speaker.

[0081] Step 113, determining the first determination result of the first converted data corresponding to the target speaker based on the first prediction result, the second prediction result and the true result corresponding to the target speaker.

[0082] In the embodiment, the first converted data is audio data, thus has spectrum information, different speakers correspond to different spectrum information, by identifying the spectrum information of the first converted data, the first prediction result of the target speaker corresponding to the first converted data is determined. Specifically, the similarity between the spectrum information of the first converted data and the spectrum information of the standard voice data of the target speaker can be determined, and the value corresponding to the similarity can be directly determined as the first prediction result. In the embodiment, not only the conversion effect of the first converted data is determined by the first prediction result, but also the second converted data is further identified to determine the second prediction result of the target speaker corresponding to the first converted data, so as to determine whether the first converted data corresponds to the target speaker by using the double verification mode of the first prediction result and the second prediction result, and ensure the accuracy of the determined first determination result.

[0083] In a specific embodiment, the voice conversion module is pre-trained, and the first prediction result of the first converted data corresponding to the target speaker is obtained based on the classifier of the voice conversion model. By inputting the dialect voice data into the voice conversion model, not only the first converted data can be obtained, but also the classifier of the voice conversion model can output the first prediction result.

[0084] In a specific embodiment, the speaker recognition model is pre-trained, and the second prediction result of the first converted data corresponding to the target speaker is obtained based on the speaker recognition model. By inputting the first converted data into the speaker recognition model, the speaker recognition model outputs the second prediction result.

[0085] Specifically, the speaker recognition model based on the self-attention convolution structure is built, and an exemplary model structure of the speaker recognition model is as shown in Figure 3As shown, the speaker recognition model is composed of a mapping structure (Liner & Relu), a one-dimensional convolution structure (Conv1D Block), a mean pooling structure (Mean pooling), a self-attention structure (Self-Attention), and an output structure (Linear and Softmax). The training data of the speaker recognition model includes dialect training data and standard training data. The dialect training data can be a small amount of dialect audio data recorded by voice actors who master Cantonese, Northeastern Mandarin, Sichuanese, etc. The standard training data is a large amount of Mandarin audio data existing in the database. Before inputting the training data of the speaker recognition model into the speaker recognition model, the audio data processing is performed on the training data to obtain the audio features. The fbank feature extraction is performed on the training data to obtain 80-dimensional fbank features. The 80-dimensional fbank features are input into the mapping structure composed of 2 layers of full connection and elu activation function. The 80-dimensional features are mapped to 128 dimensions, and then pass through the convolution layer of 3 layers of 1-dimensional convolution plus GLU (Gated Linear Units) residual structure, and the mean pooling structure for summarizing the speech information. Finally, the speaker dimension features are obtained through the self-attention structure, and the probability of identifying the corresponding speaker is output through the softmax structure to obtain the second prediction result.

[0086] The stargan v2 voice conversion model based on adversarial learning is built. As an example, the model structure of the voice conversion model is as shown in Figure 4 The voice conversion model is composed of four modules A, B, C, and D. Module A (Style Encoder) is a speaker style generation module, which is specifically a pre-trained speaker recognition model. The output of the speaker recognition model is a speaker dimension feature (Speaker Vector) obtained through the self-attention structure. Figure 4 Module B in the middle is a target speaker spectrum conversion generation module. A seq2seq network structure based on self-attention is built, taking audio features (fbank features) and speaker dimension features as input items, and outputting the converted spectrum of the target speaker's voice. Figure 4 Module C in the middle is a spectrum judgment module, which includes a judge and a classifier. The model structure of the judge and the classifier is a pre-trained speaker recognition model. The spectrum judgment module is used for adversarial training to judge and improve the conversion effect of the target speaker spectrum conversion generation module. Figure 4 Module D in the middle is a vocoder module, which is used to convert the spectrum data to audio data based on the HIFI-GAN structure, i.e., to output the first conversion data, and the classifier outputs the first prediction result. The training data of the voice conversion model is dialect training data and standard training data.

[0087] The voice conversion model and the speaker recognition model can obtain accurate first prediction results and second prediction results.

[0088] Further, the step 112 determines a first determination result of the target speaker corresponding to the first conversion data based on the first prediction result, the second prediction result, and a real result corresponding to the target speaker, and the step 112 includes:

[0089] The step 1121 determines a first cross-entropy between the first prediction result and the real result corresponding to the target speaker.

[0090] The step 1122 determines a second cross-entropy between the second prediction result and the real result corresponding to the target speaker.

[0091] The step 1123 determines the first determination result of the target speaker corresponding to the first conversion data based on the first cross-entropy and the second cross-entropy.

[0092] The first cross-entropy can indicate the difference between the first prediction result and the real result corresponding to the target speaker, so that the first cross-entropy can intuitively reflect the conversion effect of the first conversion data. Meanwhile, the second cross-entropy can indicate the difference between the second prediction result and the real result corresponding to the target speaker, so that the second cross-entropy can also intuitively reflect the conversion effect of the first conversion data. The first cross-entropy and the second cross-entropy are calculated according to a preset calculation mode to determine the first determination result of the target speaker corresponding to the first conversion data.

[0093] For example, the calculation formula of the first determination result is as follows:

[0094]

[0095] Score(Y) = λ * CE(Y) + (1 - λ) * P(Y) speaker Score(Y) represents the first determination result, C represents a classifier, Y represents the target speaker, C represents a speaker recognition model, CE represents a cross-entropy, and λ represents a preset weight coefficient. vc Score(Y) represents the first determination result, C represents a classifier, Y represents the target speaker, C represents a speaker recognition model, CE represents a cross-entropy, and λ represents a preset weight coefficient. speaker Score(Y) represents the first determination result, C represents a classifier, Y represents the target speaker, C represents a speaker recognition model, CE represents a cross-entropy, and λ represents a preset weight coefficient. vc Score(Y) represents the first determination result, C represents a classifier, Y represents the target speaker, C represents a speaker recognition model, CE represents a cross-entropy, and λ represents a preset weight coefficient. Score(Y) represents the first determination result, C represents a classifier, Y represents the target speaker, C represents a speaker recognition model, CE represents a cross-entropy, and λ represents a preset weight coefficient. Score(Y) represents the first determination result, C represents a classifier, Y represents the target speaker, C represents a speaker recognition model, CE represents a cross-entropy, and λ represents a preset weight coefficient. Score(Y) represents the first determination result, C represents a classifier, Y represents the target speaker, C represents a speaker recognition model, CE represents a cross-entropy, and λ represents a preset weight coefficient. Score(Y) represents the first determination result, C represents a classifier, Y represents the target speaker, C represents a speaker recognition model, CE represents a cross-entropy, and λ represents a preset weight coefficient. Score(Y) represents the first determination result, C represents a classifier, Y represents the target speaker, C represents a speaker recognition model, CE represents a cross-entropy, and λ represents a preset weight coefficient. Score(Y) represents the first determination result, C represents a classifier, Y represents the target speaker, C represents a speaker recognition model, CE represents a cross-entropy, and λ represents a preset weight coefficient.

[0096] Through the calculation formula of the first determination result, the first determination result can be accurately determined, which can be an index for objectively quantifying and evaluating the conversion effect of the first conversion data.

[0097] In a specific embodiment, when the first conversion data is processed based on the first information comparison result of the T pieces of dialect speech data and the first conversion data, the first determination result that the first conversion data corresponds to the target speaker, and the target dialect speech data is determined from the T pieces of dialect speech data based on the first processing result, the first processing result can be determined according to the following calculation formula:

[0098] Score pick-source = λ ASR * Score ASR + λ speaker * Score speaker ++ λ f * Score f (4)

[0099] Wherein, Score pick-source represents the first processing result; Score ASR represents the speech recognition comparison result obtained by the speech recognition model; Score speaker represents the first determination result obtained by the speech conversion module and the speaker recognition model; Score f represents the audio information comparison result based on the audio information; λ ASR , λ speaker , λ f represent the preset weight coefficients.

[0100] Specifically:

[0101]

[0102]

[0103]

[0104]

[0105] Through the calculation formula of the first processing result, the first processing result can be accurately determined, and the accuracy of the target dialect speech data determined according to the first processing result can be ensured to be high.

[0106] In an embodiment of the present application, the second information comparison result includes at least one of the speech recognition comparison result and the audio information comparison result of the P pieces of dialect speech data and the second conversion data.

[0107] When the second information comparison result comprises P pieces of dialect speech data and the speech recognition comparison result of the second conversion data, before step 104, the method further comprises:

[0108] determining a third speech recognition result of the P pieces of dialect speech data; determining a fourth speech recognition result of the second conversion data; determining the speech recognition comparison result of the P pieces of dialect speech data and the second conversion data based on the third speech recognition result and the fourth speech recognition result.

[0109] When the text information between the P pieces of dialect speech data and the second conversion data is the same, the third speech recognition result and the fourth speech recognition result are comparable, and the speech recognition comparison result of the P pieces of dialect speech data and the second conversion data is determined by comparing the third speech recognition result and the fourth speech recognition result. The speech recognition comparison result can be an index for objectively quantifying and evaluating the conversion effect of the second conversion data.

[0110] In a specific embodiment, a speech recognition model is pre-trained, and the speech recognition model is used to recognize speech data, that is, the P pieces of dialect speech data are input into the speech recognition model to obtain the output item third speech recognition result of the speech recognition model, and the third speech recognition result is the text information obtained by recognizing the P pieces of dialect speech data. The second conversion data is input into the speech recognition model to obtain the output item fourth speech recognition result of the speech recognition model, and the fourth speech recognition result is the text information obtained by recognizing the second conversion data. The model structure of the speech recognition model can be as described above.

[0111] Specifically, based on the third speech recognition result and the fourth speech recognition result, the speech recognition comparison result of the P pieces of dialect speech data and the second conversion data is determined, which comprises: determining the character error rate between the third speech recognition result and the fourth speech recognition result. When the character error rate is within a preset numerical range, the speech recognition comparison result of the P pieces of dialect speech data and the second conversion data is determined according to the character error rate. When the character error rate is not within the preset numerical range, the dialect speech data corresponding to the character error rate not within the preset numerical range is deleted from the P pieces of dialect speech data.

[0112] By comparing the character information included in the text information of the third speech recognition result and the fourth speech recognition result, the character error rate between the third speech recognition result and the fourth speech recognition result is determined. When the character error rate is within the preset numerical range, it indicates that the conversion effect is better, and therefore the speech recognition comparison result of the P pieces of dialect speech data and the second conversion data can be further determined according to the character error rate. In the case where the character error rate is not within the preset data range, it indicates that the conversion effect is too poor, and therefore the dialect speech data corresponding to the character error rate not within the preset numerical range is deleted from the P pieces of dialect speech data, effectively reducing the number of dialect speech data.

[0113] In a specific embodiment, the calculation formula of the speech recognition comparison result of the P pieces of dialect speech data and the second conversion data is as follows:

[0114]

[0115] wherein Score ASR,ij characterizes the i th standard speech data of the target speaker Y, and the speech recognition comparison result of the j th piece of dialect speech data of the source dialect speaker X after speech conversion, Y i characterizes the i th standard speech data of the target speaker Y. characterizes the second conversion data obtained after the j th piece of dialect speech data of the source dialect speaker X is converted into the i th standard speech data of the target speaker Y. By Score ASR,ij It can be reflected whether the second conversion data after speech conversion retains the original language information of the dialect speech data, and the unclear pronunciation of the speech conversion model is filtered out. In this embodiment, the preset numerical range is set to be less than 1, and CER is greater than or equal to 1. At this time, it indicates that the conversion effect is too poor, and the dialect speech data of the source dialect speaker with CER greater than 1 is excluded.

[0116] In an embodiment of the present application, in the case where the second information comparison result includes the audio information comparison result of the P pieces of dialect speech data and the second conversion data, before step 104, the method further includes: determining the pitch comparison result between the P pieces of dialect speech data and the second conversion data based on the third pitch segment length of the P pieces of dialect speech data and the fourth pitch segment length of the second conversion data. Determine the formant comparison result between the P pieces of dialect speech data and the second conversion data based on the first formant information of the P pieces of dialect speech data and the first formant information of the second conversion data. Determine the audio information comparison result between the P pieces of dialect speech data and the second conversion data based on the pitch comparison result and the formant comparison result between the P pieces of dialect speech data and the second conversion data.

[0117] The third fundamental frequency segment length refers to a length result obtained by segmenting the fundamental frequency in the audio information of the P pieces of dialect voice data, and the fourth fundamental frequency segment length refers to a length result obtained by segmenting the fundamental frequency in the audio information of the second converted data. The third fundamental frequency segment length and the fourth fundamental frequency segment length are compared to determine a fundamental frequency comparison result between the P pieces of dialect voice data and the second converted data. The audio information includes not only the fundamental frequency information but also the first formant information, and therefore the first formant information of the audio information of the P pieces of dialect voice data and the first formant information of the audio information of the second converted data are compared to determine a formant comparison result. The second converted data is evaluated according to the fundamental frequency comparison result and the formant comparison result to obtain an audio information comparison result, which can be an index for objectively quantitatively evaluating the conversion effect of the second converted data.

[0118] For example, the calculation formula of the audio information comparison result of the P pieces of dialect voice data and the second converted data is as follows:

[0119]

[0120] wherein, indicates The length of the kth segment of the L pieces of fundamental frequency segments obtained by the dio algorithm corresponds to the third fundamental frequency segment length. indicates X i The length of the kth segment of the L pieces of fundamental frequency segments obtained by the dio algorithm corresponds to the fourth fundamental frequency segment length. In the left formula, and The relationship between the two corresponds to the audio information comparison result of the P pieces of dialect voice data and the second converted data. indicates The left derivative of the mth frequency of the first formant existing between adjacent points after framing; indicates The right derivative of the mth frequency of the first formant existing between adjacent points after framing.

[0121] The audio information comparison result of the P pieces of dialect voice data and the second converted data can be accurately obtained through the above calculation formula (5), which is beneficial to accurately screening the third converted data from the second converted data.

[0122] In an embodiment of the present application, before the step 104, the method further comprises: identifying the spectrum of the second converted data to determine a third prediction result of the second converted data corresponding to the target speaker; identifying the speaker of the second converted data to determine a fourth prediction result of the second converted data corresponding to the target speaker; and determining a second determination result of the second converted data corresponding to the target speaker based on the third prediction result, the fourth prediction result, and a real result corresponding to the target speaker.

[0123] In the embodiment, the conversion effect of the second converted data is determined not only by the third prediction result, but also by a fourth prediction result of the target speaker corresponding to the second converted data determined by speaker identification, so that the second converted data is determined to correspond to the target speaker by double verification of the third prediction result and the fourth prediction result, thereby ensuring the accuracy of the second determination result.

[0124] In a specific implementation, the speech conversion module is pre-trained, and the third prediction result of the second converted data corresponding to the target speaker is obtained based on a classifier of the speech conversion model. By inputting the P dialect speech data into the speech conversion model, the second converted data can be obtained, and the classifier of the speech conversion model can also output the third prediction result.

[0125] In a specific implementation, the speaker identification model is pre-trained, and the fourth prediction result of the second converted data corresponding to the target speaker is obtained based on the speaker identification model. By inputting the second converted data into the speaker identification model, the fourth prediction result is output by the speaker identification model.

[0126] Further, the second determination result of the second converted data corresponding to the target speaker is determined based on the third prediction result, the fourth prediction result, and the real result corresponding to the target speaker, which comprises: determining a third cross-entropy between the third prediction result and the real result corresponding to the target speaker; determining a fourth cross-entropy between the third prediction result and the real result corresponding to the target speaker; and determining the second determination result of the second converted data corresponding to the target speaker based on the third cross-entropy and the fourth cross-entropy.

[0127] For example, the calculation formula of the second determination result is as follows:

[0128]

[0129] By the calculation formula of the second determination result, the second determination result can be accurately determined, which can be an index for objectively and quantitatively evaluating the conversion effect of the second converted data.

[0130] In a specific implementation, in the case that the second conversion data is processed based on the second information comparison result of the P pieces of dialect speech data and the second conversion data and the second determination result of the second conversion data corresponding to the target speaker, the third conversion data is selected from the second conversion data based on a second processing result, and the second processing result can be determined according to the following calculation formula:

[0131]

[0132] Score represents an evaluation result of speech conversion of the jth piece of dialect speech data of the source dialect speaker X according to the ith standard speech data of the target speaker Y, and corresponds to the second processing result, Score ASR,ij Score represents a speech recognition comparison result of the P pieces of dialect speech data and the second conversion data. speaker,ij Score represents a second determination result of the second conversion data corresponding to the target speaker based on the speech conversion module and the speaker recognition model.

[0133] Specifically,

[0134]

[0135]

[0136] Through the calculation formula of the second processing result, the second processing result can be accurately determined, and the third conversion data with better conversion effect can be selected according to the second processing.

[0137] Further, the calculation formula of the second processing result can be as follows:

[0138]

[0139] f Score represents an audio information comparison result of the P pieces of dialect speech data and the second conversion data.

[0140] Specifically,

[0141]

[0142]

[0143]

[0144] ​​By the calculation formula of the second processing result, the more accurate second processing result can be determined by considering the comparison result of the audio information of the P pieces of dialect speech data and the second conversion data, and the third conversion data with better conversion effect can be screened out according to the second processing.

[0145] The speech data screening method provided in the embodiments of the present application can be executed by a speech data screening device. The speech data screening device provided in the embodiments of the present application is described by taking the speech data screening device as an example.

[0146] Figure 5 A block diagram of the speech data screening device of another embodiment of the present application is shown, which includes:

[0147] The first conversion processing module 51 is configured to obtain first conversion data based on T pieces of dialect speech data and selected standard speech data corresponding to the target speaker, where T is an integer greater than zero.

[0148] The first screening processing module 52 is configured to process the first conversion data based on at least one of a first information comparison result of the T pieces of dialect speech data and the first conversion data and a first determination result of the target speaker corresponding to the first conversion data, and determine target dialect speech data from the T pieces of dialect speech data based on a first processing result.

[0149] The second conversion processing module 53 is configured to obtain second conversion data based on P pieces of dialect speech data corresponding to the target dialect speech data and K pieces of standard speech data corresponding to the target speaker, where P is greater than T, and K is an integer greater than zero.

[0150] The second screening processing module 54 is configured to process the second conversion data based on at least one of a second information comparison result of the P pieces of dialect speech data and the second conversion data and a second determination result of the target speaker corresponding to the second conversion data, and screen third conversion data from the second conversion data based on a second processing result.

[0151] The first information comparison result includes at least one of a speech recognition comparison result and an audio information comparison result.

[0152] Optionally, the device further includes a speech result determination module.

[0153] The speech result determination module includes:

[0154] The first recognition processing unit is configured to determine a first speech recognition result of the T pieces of dialect speech data.

[0155] The second recognition processing unit is configured to determine a second speech recognition result of the first converted data.

[0156] The speech result determination unit is configured to determine a speech recognition comparison result of the T pieces of dialect speech data and the first converted data based on the first speech recognition result and the second speech recognition result.

[0157] Optionally, the speech result determination unit comprises:

[0158] The first determination sub-unit is configured to determine a character error rate between the first speech recognition result and the second speech recognition result.

[0159] The second determination sub-unit is configured to determine a speech recognition comparison result of the T pieces of dialect speech data and the first converted data according to the character error rate when the character error rate is within a preset numerical range.

[0160] The third determination sub-unit is configured to delete, from the T pieces of dialect speech data, dialect speech data corresponding to a character error rate that is not within the preset numerical range when the character error rate is not within the preset numerical range.

[0161] Optionally, the apparatus further comprises an audio result determination module.

[0162] The audio result determination module comprises:

[0163] The first comparison processing unit is configured to determine a fundamental frequency comparison result based on a first fundamental frequency segment length of the T pieces of dialect speech data and a second fundamental frequency segment length of the first converted data.

[0164] The second comparison processing unit is configured to determine a formant comparison result based on first formant information of the T pieces of dialect speech data and first formant information of the first converted data.

[0165] The audio result determination unit is configured to determine an audio information comparison result of the T pieces of dialect speech data and the first converted data based on the fundamental frequency comparison result and the formant comparison result.

[0166] Optionally, the apparatus further comprises a determination result determination module.

[0167] The determination result determination module comprises:

[0168] The first prediction processing unit is configured to identify a spectrum of the first converted data to determine a first prediction result of a target speaker corresponding to the first converted data.

[0169] A second prediction processing unit is configured to identify the speaker of the first converted data and determine a second prediction result of the first converted data corresponding to the target speaker.

[0170] A determination result determining unit is configured to determine a first determination result of the first converted data corresponding to the target speaker based on the first prediction result, the second prediction result and a real result corresponding to the target speaker.

[0171] Optionally, the determination result determining unit comprises:

[0172] A fourth determining sub-unit is configured to determine a first cross-entropy between the first prediction result and the real result corresponding to the target speaker.

[0173] A fifth determining sub-unit is configured to determine a second cross-entropy between the second prediction result and the real result corresponding to the target speaker.

[0174] A sixth determining sub-unit is configured to determine the first determination result of the first converted data corresponding to the target speaker based on the first cross-entropy and the second cross-entropy.

[0175] In the embodiments of the present application, a small amount of dialectal speech data is used for style transfer with a large amount of standard speech data, so as to obtain a large amount of converted data using dialects. The converted data is automatically selected by using objective quantitative evaluation indexes, and the third converted data with better conversion effect is screened out. The data quality of the third converted data is high, and the workload of subsequent manual screening of the converted data is greatly reduced, and the labor cost is saved.

[0176] The voice data screening device in the embodiments of the present application can be an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), and can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, and the like, and the embodiments of the present application are not limited in this regard.

[0177] The voice data screening device in the embodiments of the present application can be a device with a motion system. The motion system can be an Android motion system, an ios motion system, or other possible motion systems, and the embodiments of the present application are not limited in this regard.

[0178] The voice data screening device provided in the embodiments of the present application can implement each process implemented by the method embodiments, and thus repeated descriptions are not given herein.

[0179] Optionally, as shown in Figure 6 The embodiments of the present application also provide an electronic device 60, which includes a processor 61, a memory 62, and a program or instruction stored in the memory 62 and executable on the processor 61. When the program or instruction is executed by the processor 61, each step of any voice data screening method embodiment described above is implemented, and the same technical effects are achieved. Thus, repeated descriptions are not given herein.

[0180] It should be noted that the electronic device in the embodiments of the present application includes the mobile electronic device and the non-mobile electronic device described above.

[0181] Figure 7 To implement the hardware structure of an electronic device in the embodiments of the present application.

[0182] The electronic device 700 includes, but is not limited to, a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, and a processor 710, etc.

[0183] Those skilled in the art can understand that the electronic device 700 can also include a power supply (such as a battery) for powering various components, which can be logically connected to the processor 710 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 7 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than shown, or combine certain components, or different component arrangements, which are not described here.

[0184] The processor 710 is configured to: obtain first conversion data based on T pieces of dialect speech data and standard speech data corresponding to a target speaker, T being an integer greater than zero; process the first conversion data based on at least one of a first information comparison result of the T pieces of dialect speech data and the first conversion data, and a first determination result of the first conversion data corresponding to the target speaker, and determine target dialect speech data from the T pieces of dialect speech data based on a first processing result; obtain second conversion data based on P pieces of dialect speech data corresponding to the target dialect speech data and K pieces of standard speech data corresponding to the target speaker, P being greater than T, and K being an integer greater than zero; process the second conversion data based on at least one of a second information comparison result of the P pieces of dialect speech data and the second conversion data, and a second determination result of the second conversion data corresponding to the target speaker, and screen third conversion data from the second conversion data based on a second processing result; and the first information comparison result includes at least one of a speech recognition comparison result and an audio information comparison result.

[0185] In the embodiments of the present application, by performing style transfer on a small amount of dialect speech data and a large amount of standard speech data, a large amount of conversion data using dialects can be obtained, and the conversion data can be automatically selected by objective quantitative evaluation indexes, and third conversion data with better conversion effect can be screened out, the data quality of the third conversion data is higher, and at the same time, the workload of subsequent manual screening of conversion data can be greatly reduced by automatic screening, and the labor cost is saved.

[0186] Optionally, the processor 710 is further configured to determine a first speech recognition result of the T pieces of dialect speech data; determine a second speech recognition result of the first converted data; and determine a speech recognition comparison result of the T pieces of dialect speech data and the first converted data based on the first speech recognition result and the second speech recognition result.

[0187] Optionally, the processor 710 is further configured to determine a character error rate between the first speech recognition result and the second speech recognition result; in a case where the character error rate is in a preset numerical range, determine the speech recognition comparison result of the T pieces of dialect speech data and the first converted data according to the character error rate; and in a case where the character error rate is not in the preset numerical range, delete dialect speech data corresponding to the character error rate that is not in the preset numerical range from the T pieces of dialect speech data.

[0188] Optionally, the processor 710 is further configured to determine a fundamental frequency comparison result based on a first fundamental frequency segment length of the T pieces of dialect speech data and a second fundamental frequency segment length of the first converted data; determine a formant comparison result based on first formant information of the T pieces of dialect speech data and first formant information of the first converted data; and determine an audio information comparison result of the T pieces of dialect speech data and the first converted data based on the fundamental frequency comparison result and the formant comparison result.

[0189] Optionally, the processor 710 is further configured to identify a spectrum of the first converted data to determine a first prediction result of a target speaker corresponding to the first converted data; identify a speaker of the first converted data to determine a second prediction result of the target speaker corresponding to the first converted data; and determine a first determination result of the target speaker corresponding to the first converted data based on the first prediction result, the second prediction result, and a true result corresponding to the target speaker.

[0190] Optionally, the processor 710 is further configured to determine a first cross-entropy between the first prediction result and the true result corresponding to the target speaker; determine a second cross-entropy between the second prediction result and the true result corresponding to the target speaker; and determine the first determination result of the target speaker corresponding to the first converted data based on the first cross-entropy and the second cross-entropy.

[0191] It should be understood that in the embodiments of the present application, the input unit 704 can include a graphics processor (GPU) 7041 and a microphone 7042, and the graphics processor 7041 processes image data of a still picture or a video image obtained by an image capturing device (such as a camera) in a video image capturing mode or an image capturing mode. The display unit 706 can include a display panel 7061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 707 includes at least one of a touch panel 7071 and other input devices 7072. The touch panel 7071 is also called a touch screen. The touch panel 7071 can include two parts of a touch detection device and a touch controller. The other input devices 7072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), a trackball, a mouse, a motion stick, and the like, and will not be described here. The memory 709 can be used to store software programs and various data, including but not limited to application programs and action systems. The processor 710 can integrate an application processor and a modem processor, wherein the application processor mainly processes action systems, user pages and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 710.

[0192] The memory 709 can be used to store software programs and various data. The memory 709 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 709 can include a volatile memory or a non-volatile memory, or the memory 709 can include both a volatile memory and a non-volatile memory. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 709 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0193] The processor 710 can include one or more processing units; optionally, the processor 710 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 710.

[0194] The embodiments of the present application also provide a readable storage medium, the readable storage medium stores programs or instructions, the programs or instructions are executed by a processor to realize each process of the above-mentioned voice data screening method embodiments, and the same technical effects can be achieved. To avoid repetition, details are not described here.

[0195] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0196] The embodiment of the present application further provides a chip, which comprises a processor and a communication interface, the communication interface is coupled with the processor, the processor is used for running programs or instructions to realize the processes of the voice data screening method and achieve the same technical effects. To avoid repetition, details are not described herein.

[0197] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system on chip (SoC), a system chip, a chip system or a system on chip (SoC), etc.

[0198] The embodiment of the present application provides a computer program product, which is stored in a storage medium, and is executed by at least one processor to realize the processes of the voice data screening method and achieve the same technical effects. To avoid repetition, details are not described herein.

[0199] It should be noted that in this document, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "comprises a" does not exclude the presence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiment of the present application is not limited to the order of performing the functions as shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in the opposite order, for example, the described method can be performed in an order different from the described order, and various steps can also be added, omitted or combined. In addition, the features described with reference to some examples can be combined in other examples.

[0200] Through the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned example methods can be realized by means of software and a necessary general hardware platform, and of course, can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product in essence or in the form of a part that contributes to the prior art, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) and includes a plurality of instructions for causing a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0201] The embodiments of the present application are described above in combination with the drawings, but the present application is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are only illustrative and not restrictive. Those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.

Claims

1. A voice data screening method characterized by comprising: The method comprises: Based on T dialect speech data and the selected standard speech data corresponding to the target speaker, obtain first conversion data, T is an integer greater than zero; Based on at least one of the first information comparison result of the T dialect speech data and the first conversion data and the first determination result of the first conversion data corresponding to the target speaker, process the first conversion data, and determine the target dialect speech data from the T dialect speech data based on the first processing result; Based on P dialect speech data corresponding to the target dialect speech data and K standard speech data corresponding to the target speaker, obtain second conversion data, P is greater than T, and K is an integer greater than zero; Based on at least one of the second information comparison result of the P dialect speech data and the second conversion data and the second determination result of the second conversion data corresponding to the target speaker, process the second conversion data, and screen the third conversion data from the second conversion data based on the second processing result; The first information comparison result comprises at least one of a speech recognition comparison result and an audio information comparison result; The first information comparison result comprises at least one of a speech recognition comparison result and an audio information comparison result; The first information comparison result comprises at least one of a speech recognition comparison result and an audio information comparison result; 2. The method of claim 1, wherein, In the case where the first information comparison result comprises a speech recognition comparison result, before processing the first conversion data based on at least one of the first information comparison result of the T dialect speech data and the first conversion data and the determination result of the first conversion data corresponding to the target speaker, the method further comprises: Determine the first speech recognition result of the T dialect speech data; Determine the second speech recognition result of the first conversion data; Based on the first speech recognition result and the second speech recognition result, determine the speech recognition comparison result of the T dialect speech data and the first conversion data.

3. The method of claim 2, wherein, The first information comparison result comprises at least one of a speech recognition comparison result and an audio information comparison result; Determine the character error rate between the first speech recognition result and the second speech recognition result; In the case where the character error rate is within a preset numerical range, determine the speech recognition comparison result of the T dialect speech data and the first conversion data according to the character error rate; In the case where the character error rate is not within the preset numerical range, delete the dialect speech data corresponding to the character error rate not within the preset numerical range in the T dialect speech data.

4. The method of claim 1, wherein, In the case that the first information comparison result includes an audio information comparison result, before processing the first conversion data based on at least one of the information comparison result of the T pieces of dialect speech data and the first conversion data and the first determination result that the first conversion data corresponds to the target speaker, the method further includes: determining a fundamental frequency comparison result based on a first fundamental frequency segment length of the T pieces of dialect speech data and a second fundamental frequency segment length of the first conversion data; determining a formant comparison result based on first formant information of the T pieces of dialect speech data and first formant information of the first conversion data; determining an audio information comparison result of the T pieces of dialect speech data and the first conversion data based on the fundamental frequency comparison result and the formant comparison result.

5. The method of claim 1, wherein, Before processing the first conversion data based on at least one of the first information comparison result of the T pieces of dialect speech data and the first conversion data and the first determination result that the first conversion data corresponds to the target speaker, the method further includes: identifying a spectrum of the first conversion data to determine a first prediction result that the first conversion data corresponds to the target speaker; identifying a speaker of the first conversion data to determine a second prediction result that the first conversion data corresponds to the target speaker; determining the first determination result that the first conversion data corresponds to the target speaker based on the first prediction result, the second prediction result and a real result corresponding to the target speaker.

6. The method of claim 5, wherein, The determination of the first determination result that the first conversion data corresponds to the target speaker based on the first prediction result, the second prediction result and the real result corresponding to the target speaker includes: determining a first cross-entropy between the first prediction result and the real result corresponding to the target speaker; determining a second cross-entropy between the second prediction result and the real result corresponding to the target speaker; determining the first determination result that the first conversion data corresponds to the target speaker based on the first cross-entropy and the second cross-entropy.

7. A voice data screening apparatus characterized by comprising: The apparatus includes: a first conversion processing module configured to obtain first conversion data based on T pieces of dialect speech data and selected standard speech data corresponding to a target speaker, T being an integer greater than zero; a first screening processing module configured to process the first conversion data based on at least one of a first information comparison result of the T pieces of dialect speech data and the first conversion data and a first determination result that the first conversion data corresponds to the target speaker, and determine target dialect speech data from the T pieces of dialect speech data based on a first processing result; the determination of the target dialect speech data from the T pieces of dialect speech data based on the first processing result includes: determining evaluation values of the T pieces of dialect speech data in the first processing result, sorting the evaluation values, and determining a dialect speech data with a highest evaluation value as the target dialect speech data, the dialect speech data with the highest evaluation value being the target dialect speech data. The second conversion processing module is configured to obtain second conversion data based on the P pieces of dialect speech data corresponding to the target dialect speech data and K pieces of standard speech data corresponding to the target speaker, where P is greater than T, and K is an integer greater than zero; The second screening processing module is configured to process the second conversion data based on at least one of a second information comparison result of the P pieces of dialect speech data and the second conversion data and a second determination result of the second conversion data corresponding to the target speaker, and screen third conversion data from the second conversion data based on a second processing result. The first information comparison result includes at least one of a speech recognition comparison result and an audio information comparison result.

8. The apparatus of claim 7, wherein, The device further includes a speech result determination module. The speech result determination module includes: The first recognition processing unit is configured to determine a first speech recognition result of the T pieces of dialect speech data. The second recognition processing unit is configured to determine a second speech recognition result of the first conversion data. The speech result determination unit is configured to determine a speech recognition comparison result of the T pieces of dialect speech data and the first conversion data based on the first speech recognition result and the second speech recognition result.

9. The apparatus of claim 8, wherein, The speech result determination unit includes: The first determination subunit is configured to determine a character error rate between the first speech recognition result and the second speech recognition result. The second determination subunit is configured to, in a case where the character error rate is within a preset numerical range, determine a speech recognition comparison result of the T pieces of dialect speech data and the first conversion data according to the character error rate. The third determination subunit is configured to, in a case where the character error rate is not within the preset numerical range, delete, from the T pieces of dialect speech data, dialect speech data corresponding to a character error rate that is not within the preset numerical range.

10. The apparatus of claim 7, wherein, The device further includes an audio result determination module. The audio result determination module includes: The first comparison processing unit is configured to determine a fundamental frequency comparison result based on a first fundamental frequency segment length of the T pieces of dialect speech data and a second fundamental frequency segment length of the first conversion data. The second comparison processing unit is configured to determine a formant comparison result based on first formant information of the T pieces of dialect speech data and first formant information of the first conversion data. The audio result determination unit is configured to determine an audio information comparison result of the T pieces of dialect speech data and the first conversion data based on the fundamental frequency comparison result and the formant comparison result.

11. The apparatus of claim 7, wherein, The device further includes a determination result determination module. The determination result determination module includes: The first prediction processing unit is configured to identify a spectrum of the first conversion data to determine a first prediction result of the target speaker corresponding to the first conversion data. The second prediction processing unit is configured to identify a speaker of the first conversion data to determine a second prediction result of the target speaker corresponding to the first conversion data. The determination result determination unit is configured to determine a first determination result of the target speaker corresponding to the first conversion data based on the first prediction result, the second prediction result and a real result corresponding to the target speaker.

12. The apparatus of claim 11, wherein, The determination result determination unit comprises: A fourth determination sub-unit configured to determine a first cross-entropy between the first prediction result and the real result corresponding to the target speaker; A fifth determination sub-unit configured to determine a second cross-entropy between the second prediction result and the real result corresponding to the target speaker; A sixth determination sub-unit configured to determine the first determination result of the target speaker corresponding to the first conversion data based on the first cross-entropy and the second cross-entropy.

13. An electronic device, comprising: A processor and a memory are included, the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the voice data screening method according to any one of claims 1-6.

14. A readable storage medium, characterized by, The programs or instructions are stored on the readable storage medium, and the programs or instructions are executed by the processor to implement the steps of the voice data screening method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method and Apparatus for Automatically Converting Voice

    US20090037179A1

  • Modeling method for speech recognition, apparatus and device

    US20200327883A1