Voice Similarity Determination Method, Device, and Program Product

The method extracts evaluation pronunciation features from user audio using standard pronunciation features to analyze voice similarity across multiple languages, addressing the increased module volume and hardware demands in existing language learning software.

JP7717822B2Active Publication Date: 2025-08-04LEMON CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
JP2023547643
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-07
Filing Date
2022-01-31
Publication Date
2025-08-04
Estimated Expiration
2042-01-31

AI Technical Summary

Technical Problem

Existing language learning software can only analyze voice similarity for one language, leading to increased module volume and hardware demands when adding similarity analysis functions for multiple languages.

Method used

A method and device for determining voice similarity by extracting evaluation pronunciation features from user audio based on standard pronunciation features, allowing analysis across multiple languages with reduced computational requirements.

Benefits of technology

Enables voice similarity analysis for multiple languages with a smaller module volume, reducing hardware demands and maintaining efficient performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717822000001
    Figure 0007717822000001
  • Figure 0007717822000002
    Figure 0007717822000002
  • Figure 0007717822000003
    Figure 0007717822000003
Patent Text Reader

Abstract

The voice similarity determination method, device, and program product provided by this embodiment relate to voice technology, and the method includes the steps of: playing a demonstration audio to obtain a user's evaluation audio, where the demonstration audio is an audio of a specified content read aloud in a specified language; obtaining standard pronunciation features corresponding to the demonstration audio, and extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, where the standard pronunciation features are used to reflect a specific pronunciation of the specified content in a specified language; and determining a feature difference between the standard pronunciation features and the evaluation pronunciation features, and determining a similarity between the evaluation audio and the demonstration audio according to the feature difference. In the solution of the present application, by extracting evaluation pronunciation features corresponding to the standard pronunciation features corresponding to the demonstration audio from the evaluation audio, the volume of the module for realizing the similarity analysis function of listen-and-repeat can be made relatively small.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to voice technology, and in particular, to a method and apparatus for determining voice similarity, and a program product.

Background Art

[0002] Many users choose to learn languages online. For example, they use language learning software to learn languages.

[0003] In many language learning software in the prior art, an analysis module is arranged to realize the similarity analysis function of listen-and-repeat. The user can read out the specified content, and the software can analyze the audio generated when the user is reading the specified content, and determine the similarity between the audio and the standard audio corresponding to the specified content. Thus, the user can know the effect of listen-and-repeat.

[0004] However, the analysis module provided by the prior art can generally analyze only one language. When the similarity analysis function of listen-and-repeat for other types of languages is added, the volume of the analysis module becomes large, and the hardware device for executing the analysis module is highly demanded.

Summary of the Invention

Problems to be Solved by the Invention

[0005] Embodiments of the present disclosure provide a method and device for determining voice similarity, and a program product, in order to overcome the problem that the volume of the module for realizing the similarity analysis function of listen-and-repeat is large in the prior art.

Means for Solving the Problems

[0006] In a first aspect, embodiments of the present disclosure provide a method for determining voice similarity based on voice interaction, the method comprising: playing a demonstration audio to obtain a user's evaluation audio, wherein the demonstration audio is an audio that reads specified content in a specified language; obtaining standard pronunciation features corresponding to the demonstration audio and extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, wherein the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content; determining a feature difference between the standard pronunciation features and the evaluation pronunciation features, and determining a similarity between the evaluation audio and the demonstration audio according to the feature difference.

[0007] In a second aspect, embodiments of the present disclosure provide a method for processing a data request instruction applicable to a server, the method comprising: receiving a data request instruction; transmitting, according to the data request instruction, an encoder based on a voice recognition model, a demonstration audio, and standard pronunciation features corresponding to the demonstration audio to a user terminal, wherein the demonstration audio is an audio that reads specified content in a specified language, the encoder is used to extract evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, and the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content.

[0008] In a third aspect, embodiments of the present disclosure provide a voice similarity determination device, the device comprising: An acquisition unit for playing demonstration audio and acquiring user evaluation audio, wherein the demonstration audio is audio for reading specified content in a specified language, and the acquisition unit; A feature extraction unit for acquiring standard pronunciation features corresponding to the demonstration audio and extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, wherein the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content, and the feature extraction unit; An analysis unit for determining a feature difference between the standard pronunciation features and the evaluation pronunciation features and determining a similarity between the evaluation audio and the demonstration audio according to the feature difference, and includes.

[0009] In a fourth aspect, an embodiment of the present disclosure provides a processing device for a data request instruction disposed in a server, and the device includes: A receiving unit for receiving a data request instruction; A transmitting unit for transmitting an encoder based on an audio recognition model, demonstration audio, and standard pronunciation features corresponding to the demonstration audio to a user terminal according to the data request instruction, and includes. The demonstration audio is audio for reading specified content in a specified language, the encoder is used to extract evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, and the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content.

[0010] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, and the electronic device includes: A memory; A processor; A computer program, and includes. The computer program is stored in the memory and configured to be executed by the processor to implement the method for determining voice similarity based on voice interaction described in the first aspect or the method for processing a data request instruction described in the second aspect.

[0011] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium storing a computer program, where the computer program is executed by a processor to implement the method for determining voice similarity based on voice interaction described in the first aspect or the method for processing a data request instruction described in the second aspect.

[0012] In a seventh aspect, an embodiment of the present disclosure provides a computer program product including a computer program, where when the computer program is executed by a processor, the method for determining voice similarity based on voice interaction described in the first aspect or the method for processing a data request instruction described in the second aspect is implemented.

[0013] In an eighth aspect, an embodiment of the present disclosure provides a computer program, where when the computer program is executed by a processor, the method for determining voice similarity based on voice interaction described in the first aspect or the method for processing a data request instruction described in the second aspect is implemented.

Advantages of the Invention

[0014] The method, device, and program product for determining voice similarity provided by this embodiment include the steps of playing a demonstration audio to obtain a user's evaluation audio, where the demonstration audio is an audio that reads specified content in a specified language; obtaining standard pronunciation features corresponding to the demonstration audio and extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, where the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content; determining the feature difference between the standard pronunciation features and the evaluation pronunciation features, and determining the similarity between the evaluation audio and the demonstration audio according to the feature difference. In the solution according to the present application, by extracting evaluation pronunciation features corresponding to the standard pronunciation features corresponding to the demonstration audio from the evaluation audio, the module for realizing the similarity analysis function of listen-and-repeat can be made to have a relatively small volume. Also, since the standard pronunciation features of the demonstration audio can reflect specific pronunciations in the specified language of the specified content, this solution can provide the function of similarity analysis of listen-and-repeat for multiple language categories while having a small calculation volume.

Brief Description of the Drawings

[0015] Hereinafter, in order to more clearly explain the embodiments of the present disclosure and the solutions in the prior art, the drawings necessary for use in the description of the embodiments or the prior art will be briefly described. Of course, the drawings described below are some embodiments of the present disclosure, and those skilled in the art can conceive of other drawings based on these drawings without creative effort.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0016] Hereinafter, in order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, with reference to the drawings related to the embodiments of the present disclosure, the technical solutions will be described clearly and completely. Naturally, the described embodiments are only a part of the embodiments of the present disclosure, not all of them. Those skilled in the art can obtain all other embodiments without creative labor based on the embodiments in the present disclosure, and all of them belong to the protection scope of the present disclosure.

[0017] FIG. 1 is a diagram of an application scenario shown in one exemplary embodiment.

[0018] As shown in FIG. 1, the user terminal can play a demonstration audio (the content of the demonstration audio is indicated by "XXX" in the figure), and the user can listen to and repeat the demonstration audio.

[0019] The user can click button 11 to control the user terminal to record audio 12 for listen-and-repeat. The user terminal can analyze the recorded audio 12 to determine the similarity between the audio 12 and the demonstration audio, so that the user can know the listen-and-repeat effect.

[0020] However, in the solutions provided by the prior art for analyzing the audio recorded during listen-and-repeat to determine the similarity, all of them can only analyze the audio of one language. For example, similarity analysis can be performed only on the audio generated when the user listens and repeats in the common language, or only on the audio generated when the user listens and repeats in English.

[0021] Based on the solutions provided by the prior art, if the similarity analysis function for listen-and-repeat of other categories of languages is directly added, the analysis module for realizing the entire function will be large in volume, and the hardware device for executing the analysis module is highly demanded.

[0022] For example, when it is necessary to analyze the audio recorded by listening and repeating in different dialects and determine the similarity between the audio and the demonstration audio, the volume of the analysis module will become large.

[0023] To solve the above technical problem, the solution provided by the present application is that when analyzing the recorded evaluation audio, only the evaluation pronunciation features corresponding to the standard pronunciation features corresponding to the demonstration audio are extracted from the evaluation audio, thereby reducing the volume of the module for realizing the similarity analysis function of listen-and-repeat. Further, since the demonstration audio is an audio that reads out the specified content in the specified language, and the standard pronunciation features are used to reflect the specific pronunciation in the specified language of the specified content, the solution according to the present application can determine the similarity between the evaluation audio and the demonstration audio based on the standard pronunciation features and the extracted evaluation pronunciation features, and such an embodiment can be applied to demonstration audios with different specified contents and different specified languages, thereby providing the function of similarity analysis of listen-and-repeat for multiple language categories.

[0024] FIG. 2 is a flowchart of a method for determining audio similarity based on voice interaction shown in one exemplary embodiment of the present application.

[0025] As shown in FIG. 2, the method for determining audio similarity based on voice interaction provided by the present application includes steps 201 to 203.

[0026] In step 201, the demonstration audio is played to obtain the user's evaluation audio, and the demonstration audio is an audio that reads out the specified content in the specified language.

[0027] The method provided by the present application can be executed by an electronic device with computing capabilities, and the electronic device can be, for example, a user terminal, and the user terminal can be provided with a microphone. The user terminal can be a device such as a mobile phone or a tablet computer.

[0028] Specifically, the user terminal can play demonstration audio, and the demonstration audio is audio that reads out specified content in a specified language. For example, character content can be preset, and the character content can be set as needed, such as "Happy New Year". Audio that reads out the content in the specified language can be pre-recorded. For example, audio that reads out the content in Cantonese can be pre-recorded. The specific language to be used can also be set as needed.

[0029] Furthermore, since the demonstration audio is for providing references to the user, multiple reference audios that read out the specified content in the specified language can be recorded, and the required demonstration audio can be selected from them as needed. For example, reference audios that read out "Happy New Year" in Cantonese can be recorded using different devices in different environments.

[0030] When actually applied, after the playback of the demonstration audio is completed, the user terminal can turn on the microphone to obtain the user's evaluation audio.

[0031] In one embodiment, a button for triggering the acquisition of evaluation audio can be arranged on the screen of the user terminal. By clicking the button, the user can trigger the user terminal to turn on the microphone to obtain the evaluation audio.

[0032] In other embodiments, after the playback of the demonstration audio is completed, the user terminal can turn on the microphone to obtain the evaluation audio.

[0033] After listening to the demonstration audio, the user can listen and repeat. Specifically, by reading the specified content in the specified language, the user terminal can obtain the evaluation audio generated when the user listens and repeats the demonstration audio.

[0034] In one alternative embodiment, the user may further operate the user terminal to send a command to complete listening and repeating to the user terminal. For example, a button for instructing the completion of listening and repeating can be displayed on the screen of the user terminal, and the user can send a command to complete listening and repeating by clicking the button. In other embodiments, when listening and repeating, the user can long-press a preset button and release the listening and repeating button after completing listening and repeating, thereby sending a command to complete listening and repeating to the user terminal.

[0035] Optionally, when obtaining the evaluation audio, the user terminal can also determine whether the user has completed listening and repeating by detecting the evaluation audio. For example, according to the energy value of the audio, it can be determined whether the user is still continuously listening and repeating.

[0036] In step 202, the standard pronunciation features corresponding to the demonstration audio are obtained, and the evaluation pronunciation features corresponding to the standard pronunciation features are extracted from the evaluation audio. The standard pronunciation features are used to reflect the specific pronunciation of the specified content in the specified language.

[0037] Specifically, the user terminal can obtain the standard pronunciation features corresponding to the demonstration audio.

[0038] Furthermore, the standard pronunciation features can be transmitted by the server to the user terminal, and the user terminal can store the received standard pronunciation features and obtain the standard pronunciation features when analyzing the evaluation audio. For example, when the user operates the user terminal to start an application set according to the method provided in the present application, the user terminal can interact with the server and request the server to send the demonstration audio and its corresponding standard pronunciation features to the user terminal.

[0039] In actual application, since the standard pronunciation features corresponding to different demonstration audios are also different, when the user terminal obtains the standard pronunciation features, it can obtain the corresponding standard pronunciation features based on the played demonstration audio.

[0040] The standard pronunciation features can be preset according to the specified content and the specified language, and can reflect the specific pronunciation in the specified language of the specified content.

[0041] Specifically, a plurality of reference audios for reading the specified content in the specified language can be pre-recorded, and any one of the reference audios can be used as the demonstration audio. The reference pronunciation features of each reference audio that can represent the pronunciation features when reading the specified content in the specified language are extracted, and these reference pronunciation features are fused to obtain the standard pronunciation features corresponding to the demonstration audio. Since the standard pronunciation features are obtained by fusing a plurality of reference pronunciation features, the standard pronunciation features can represent the specific pronunciation in the specified language of the specified content.

[0042] Furthermore, the user terminal can further extract evaluation pronunciation features corresponding to standard pronunciation features from the evaluation audio. In such an embodiment, the evaluation pronunciation features corresponding to the standard pronunciation features can be extracted by narrowing down the target within the evaluation audio, and it is not necessary to extract all the features of the evaluation audio. Therefore, the amount of data to be processed can be reduced, and the hardware requirements necessary for analyzing the evaluation audio can be reduced.

[0043] In one embodiment, when extracting the reference pronunciation features of each reference audio, by collecting the features of preset sampling points in each reference audio, the standard pronunciation features can include the features of these preset sampling points. The preset sampling points in the reference audio can be determined according to the pronunciation positions with the feature points of the language when reading the specified content in the specified language.

[0044] In such an embodiment, according to the positions of the preset sampling points corresponding to the demonstration audio, the evaluation pronunciation features corresponding to the standard pronunciation features can be extracted from the evaluation audio.

[0045] In other embodiments, when extracting the reference pronunciation features of each reference audio, by collecting the features of the preset categories of each reference audio, the standard pronunciation features can include the features of these preset categories. The features of the preset categories of the reference audio can be determined according to the features with the feature points of the language when reading the specified content in the specified language. For example, it may be a feature for representing pitch changes, or a feature for representing the reading methods of all or some characters.

[0046] In such an embodiment, according to the features of the preset categories corresponding to the demonstration audio, the evaluation pronunciation features corresponding to the standard pronunciation features can be extracted from the evaluation audio.

[0047] In step 203, the characteristic difference between the standard pronunciation characteristic and the evaluation pronunciation characteristic is determined, and the similarity between the evaluation audio and the demonstration audio is determined according to the characteristic difference.

[0048] Here, the user terminal can further determine the characteristic difference between the standard pronunciation characteristic and the evaluation pronunciation characteristic by comparing the standard pronunciation characteristic and the evaluation pronunciation characteristic. For example, by comparing each characteristic in the standard pronunciation characteristic with each characteristic in the evaluation pronunciation characteristic, the characteristic difference between the standard pronunciation characteristic and the evaluation pronunciation characteristic can be obtained.

[0049] For example, alignment processing can be performed on the standard pronunciation characteristic and the evaluation pronunciation characteristic, and the characteristic difference between the standard pronunciation characteristic and the evaluation pronunciation characteristic can be obtained by comparing the first characteristic in the standard pronunciation characteristic included in each alignment point with the second characteristic in the evaluation pronunciation characteristic.

[0050] Specifically, the similarity between the evaluation audio and the demonstration audio can also be determined according to the characteristic difference. For example, the characteristic distance between the two can be determined according to the characteristic difference between the standard pronunciation characteristic and the evaluation pronunciation characteristic, and the characteristic distance can be used as the similarity between the evaluation audio and the demonstration audio.

[0051] In one alternative embodiment, the user terminal can further map the determined similarity as a score or evaluation content and display the score, so as to enable the user to recognize the listen-and-repeat effect. For example, by presetting the mapping relationship between the similarity and the score or evaluation content, the corresponding score or evaluation content can be determined according to the determined similarity.

[0052] The method for determining voice similarity based on voice interaction provided by the present application includes the steps of playing a demonstration audio and obtaining a user's evaluation audio, where the demonstration audio is an audio that reads out specified content in a specified language, obtaining standard pronunciation features corresponding to the demonstration audio, and extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, where the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content, determining the feature difference between the standard pronunciation features and the evaluation pronunciation features, and determining the similarity between the evaluation audio and the demonstration audio according to the feature difference. In the solution according to the present application, by extracting evaluation pronunciation features corresponding to the standard pronunciation features corresponding to the demonstration audio from the evaluation audio, the module for realizing the similarity analysis function of listen-and-repeat can be made to have a relatively small volume. Also, since the standard pronunciation features of the demonstration audio can reflect specific pronunciations in the specified language of the specified content, this solution can provide the function of similarity analysis of listen-and-repeat for multiple language categories while having a small calculation volume.

[0053] Figure 3 is a flowchart of the method for determining voice similarity based on voice interaction shown in another exemplary embodiment of the present application.

[0054] As shown in Figure 3, the method for determining voice similarity based on voice interaction provided by the present application includes steps 301 to 309.

[0055] In step 301, in response to a start command, a data request command is sent to the server.

[0056] Here, the method provided by the present application can be executed by an electronic device with computing capabilities. The electronic device can be, for example, a user terminal, and the user terminal can be equipped with a microphone. The user terminal can be a device such as a mobile phone or a tablet computer.

[0057] Specifically, the user can operate the user terminal to send a start command to the user terminal to start the similarity analysis function of listen-and-repeat. For example, the similarity analysis function of listen-and-repeat can be set in the application as one of the items in the application, and the application can be installed on the user terminal. In this case, the user can operate the user terminal to start the application and send a start command to the user terminal by selecting an item with the similarity analysis function of listen-and-repeat in the application.

[0058] Furthermore, the user terminal can send a data request command to the server in response to the start command. The data request command is used to request data for realizing the similarity analysis function of listen-and-repeat.

[0059] In step 302, an encoder, a demonstration audio, and standard pronunciation features corresponding to the demonstration audio are received.

[0060] In actual application, after receiving the data request command sent from the user terminal, the server can send an encoder, a demonstration audio, and standard pronunciation features corresponding to the demonstration audio to the user terminal.

[0061] Here, an encoder, a demonstration audio, and standard pronunciation features corresponding to the demonstration audio are preset in the server.

[0062] Specifically, the encoder sent by the server to the user terminal may be one obtained by pre-training.

[0063] Furthermore, an initial model can be trained using the speech recognition data to obtain a speech recognition model. Then, the encoder in the speech recognition model is trained using the audio of multiple language categories to obtain an encoder for extracting pronunciation features.

[0064] When actually applied, the speech recognition data can be audio data with text labels. Using the speech recognition model obtained by training the speech recognition data, some of the audio data can be processed to obtain the text content corresponding to the audio data.

[0065] The speech recognition model is equipped with an encoder (Encoder). Since the encoder can effectively extract information related to text and pronunciation, in the method provided by the present application, the encoder in the speech recognition model is trained using the audio data of multiple language categories to obtain an encoder that can extract pronunciation features.

[0066] Here, the audio data of multiple language categories includes the audio of multiple language categories, and each audio is with a language category label. For example, if the language used in some audio is Sichuan dialect, the language category label of the audio is a label representing Sichuan dialect. Also, for example, if the language used in some audio is Cantonese, the language category label of the audio is a label representing Cantonese. By training the encoder using the audio data of multiple language categories, the discrimination degree of the pronunciation features of different language categories by the encoder can be improved.

[0067] In order to further reduce the hardware resources required by the method provided by this application, the encoder can use the configuration of a three-layer long short-term memory network, and each layer of the network can be set with 512 nodes.

[0068] Specifically, several demonstration audios and their corresponding standard feature information are set in the server, and one or more of these demonstration audios can be sent to the user terminal, and the standard feature information corresponding to the demonstration audio can also be sent.

[0069] Furthermore, the standard pronunciation feature corresponding to the demonstration audio is obtained by fusing a plurality of reference pronunciation features. Each reference pronunciation feature is obtained by performing feature extraction on each reference audio using an encoder. Each reference audio is an audio that reads out the specified content in the specified language, and the demonstration audio is any one of the reference audios.

[0070] In such an embodiment, a plurality of reference audios that read out the specified content in the specified language can be pre-recorded. Then, the encoder is used to perform feature extraction on each reference audio to obtain the reference pronunciation feature corresponding to each reference audio. Subsequently, each reference pronunciation feature is fused to obtain a standard pronunciation feature.

[0071] Since the reference audio is for generating a standard pronunciation feature, the reference audio can be recorded by a user who uses the specified language as the daily communication language. Thereby, the standard pronunciation feature can accurately reflect the feature points of the specified language.

[0072] In step 303, the demonstration audio is played to obtain the user's evaluation audio, and the demonstration audio is the audio that reads out the specified content in the specified language.

[0073] Specifically, after receiving the demonstration audio, the user terminal can play any one of the demonstration audios.

[0074] Since the implementation form and principle of step 303 are similar to those of step 201, it will not be repeatedly described.

[0075] Optionally, on the user screen of the user terminal, in order to prompt the user with the content to be listened to and repeated, the specified content corresponding to the demonstration audio may be further displayed.

[0076] After the playback of the demonstration audio is completed, the user terminal can also play voice content for prompting the user to listen and repeat, such as "Please listen and repeat". Optionally, after the playback of the prompt content is completed, the user terminal can turn on the microphone to obtain the user's evaluation audio.

[0077] Optionally, by equipping the user terminal with a camera, a user image can be obtained and displayed on the user terminal. In one alternative embodiment, the user terminal may recognize the user image to determine whether the user has completed listening and repeating.

[0078] In step 304, the standard pronunciation features corresponding to the demonstration audio are obtained, and the standard pronunciation features are used to reflect the specific pronunciation in the specified language of the specified content.

[0079] Since the implementation form and principle of obtaining standard pronunciation features in step 304 are similar to those in step 202, they will not be repeatedly described.

[0080] In step 305, based on the encoder of the speech recognition model, evaluation pronunciation features corresponding to the standard pronunciation features are extracted from the evaluation audio.

[0081] Furthermore, the user terminal can use the encoder of the speech recognition model transmitted from the server to extract evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio.

[0082] In actual application, the user terminal can input the evaluation audio into the encoder to obtain evaluation pronunciation features corresponding to the standard pronunciation features.

[0083] Here, since the encoder can distinguish the pronunciation features of different language categories, the encoder can extract evaluation pronunciation features corresponding to the language category from the evaluation audio. Also, since the standard pronunciation features corresponding to the demonstration audio are also obtained using the encoder, by processing the evaluation audio using the same encoder, evaluation pronunciation features corresponding to the standard pronunciation features can be obtained.

[0084] Specifically, by extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio using the encoder, it is possible to narrow down and extract the evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, and it is not necessary to extract all the features of the evaluation audio. Therefore, the amount of data to be processed can be reduced, and the hardware requirements required for analyzing the evaluation audio can be reduced.

[0085] In step 306, a time stretching function is determined according to the standard pronunciation features and the evaluation pronunciation features.

[0086] Furthermore, since audio is time-ordered data, the standard pronunciation features corresponding to the demonstration audio also have time attributes, and the evaluation pronunciation features extracted from the evaluation data also have time attributes. Therefore, according to the standard pronunciation features and the evaluation pronunciation features, a time stretching function can be determined that represents the temporal correspondence between the standard pronunciation features and the evaluation pronunciation features.

[0087] In one embodiment, one time stretching function can be determined so that the evaluation pronunciation features are aligned with the standard pronunciation features on the time axis, and the time axis of the evaluation pronunciation features can be non-linearly mapped onto the time axis of the standard pronunciation features. The aligned standard pronunciation features have a first feature corresponding to the alignment point, the aligned evaluation pronunciation features have a second feature corresponding to the alignment point, and there is a certain degree of feature difference between the first feature and the second feature corresponding to each alignment point. The time stretching function can satisfy that the sum of the feature differences corresponding to each alignment point is the smallest.

[0088] In actual application, the user terminal can determine a time stretching function that satisfies the above conditions according to the standard pronunciation features and the evaluation pronunciation features.

[0089] In step 307, according to the time stretching function, the standard pronunciation features, and the evaluation pronunciation features, combinations of a plurality of alignment points are determined. Each combination of alignment points includes one standard feature point in the standard pronunciation features and one evaluation feature point in the evaluation pronunciation features.

[0090] After determining the time stretching function, the user terminal can determine a combination of multiple alignment points based on the current time stretching function, standard pronunciation features, and evaluated pronunciation features. Each combination of alignment points includes one standard feature point in the standard pronunciation features and one evaluated feature point in the evaluated pronunciation features, and the standard feature point and the evaluated feature point in the combination of alignment points correspond to the same time point.

[0091] In step 308, according to the standard feature point and the evaluated feature point included in each combination of alignment points, a feature difference corresponding to each combination of alignment points is determined.

[0092] Specifically, for each combination of alignment points, the feature difference between the standard feature point and the evaluated feature point in the combination of alignment points can be determined. For example, the distance between the standard feature point and the evaluated feature point can be calculated as the feature difference of the combination of alignment points.

[0093] In step 309, according to the feature difference of each combination of alignment points, the similarity between the evaluated audio and the demonstration audio is determined.

[0094] Furthermore, the sum of the feature differences of each combination of alignment points can be used as the similarity between the evaluated audio and the demonstration audio. By first aligning the feature points and then comparing the features, the feature difference between the evaluated audio and the demonstration audio can be accurately determined, and the similarity between the two can be accurately determined.

[0095] FIG. 4 is a schematic diagram of a similarity determination process shown in one exemplary embodiment of the present application.

[0096] As shown in FIG. 4, the user terminal can obtain the evaluated audio 41 and can also obtain the standard pronunciation features 42 corresponding to the demonstration audio.

[0097] The user terminal inputs the evaluation audio 41 into the encoder 43, and the encoder 43 can output an evaluation pronunciation feature 44 corresponding to the standard pronunciation feature 42 in the evaluation audio 41. In one embodiment, the evaluation audio 41 can be directly input into the encoder 43. In other embodiments, first, the evaluation audio 41 can be filtered, and then the filtered audio can be input into the encoder 43. For example, the mel filter bank can be used to process the evaluation audio 41.

[0098] The user terminal can further compare the standard pronunciation feature 42 and the evaluation pronunciation feature 44 to obtain the similarity 45 between the evaluation audio 41 and the demonstration audio.

[0099] In step 310, the mapping function and the setting information corresponding to the demonstration audio are obtained, and the setting information is Similarity between the evaluation audio and the demonstration audio used to indicate the mapping relationship between the and the score.

[0100] In step 311, based on the mapping function and the setting information corresponding to the demonstration audio, the similarity between the evaluation audio and the demonstration audio is mapped as a score.

[0101] When actually applied, the user terminal can further obtain the mapping function and the setting information corresponding to the demonstration audio.

[0102] In one alternative embodiment, the mapping function and the setting information corresponding to the demonstration audio can be those sent by the server so that they can be obtained by the user terminal. For example, when the server sends the demonstration audio to the user terminal, in addition to the setting information corresponding to the demonstration audio, the server can simultaneously send the mapping function.

[0103] The user terminal can store the received mapping function and the setting information corresponding to the demonstration audio, and can obtain this information when mapping the similarity as a score.

[0104] When the server sends a plurality of demonstration audios to the user terminal, the server can further send the setting information corresponding to each demonstration audio to the user terminal.

[0105] Specifically, the setting information is Similarity between the evaluation audio and the demonstration audio used to indicate the mapping relationship with the score. For example, the setting information may include, in addition to several scores, the mapping relationship corresponding to each score. The mapping function can map the determined similarity as a score based on these setting information.

[0106] Furthermore, the setting information may further include the maximum score, the similarity corresponding to the maximum score, the minimum score, and the similarity corresponding to the minimum score. For example, the maximum score can be 100 and the minimum score can be 0.

[0107] In actual application, the mapping function can be a linear function, and based on the linear function, the maximum score, the similarity corresponding to the maximum score, the minimum score, and the similarity corresponding to the minimum score, the determined similarity can be mapped as a score. By mapping the similarity as a score based on the linear function, the data processing amount can be further reduced, and the hardware requirements of the user terminal for executing the method provided by the present application can be further reduced.

[0108] The method provided by this application has setting information corresponding to different demonstration audios, and the setting information includes a maximum score and a minimum score. The maximum score in each setting information can be set to the same value, for example, all 100, and the minimum score in each setting information can be set to the same value, for example, all 0. In this way, the similarity can be mapped within the score range of the same scale by using the solution means provided by this application.

[0109] Specifically, the user terminal can also make the user recognize the listen-and-repeat effect by displaying the determined score.

[0110] Furthermore, the similarity corresponding to the maximum score is the average value of a plurality of reference similarities, and each reference similarity is the similarity between each reference pronunciation feature and the standard pronunciation feature.

[0111] When actually applied, for each reference audio, the corresponding reference pronunciation feature can be extracted, and the reference pronunciation feature of the reference audio can be extracted by using an encoder. The reference similarity between each reference pronunciation feature and the standard pronunciation feature can be determined. For example, the reference similarity between each reference pronunciation feature and the standard pronunciation feature can be determined based on the dynamic time warping method. Then, the average value of these reference similarities is determined as the similarity corresponding to the maximum score.

[0112] The similarity corresponding to the minimum score is the average value of a plurality of white noise similarities, and each white noise similarity is the similarity between each white noise feature and the standard pronunciation feature. Each white noise feature is obtained by performing feature extraction on each preset white noise audio by using an encoder.

[0113] Specifically, furthermore, several white noise audios can be prepared in advance, and the similarity corresponding to the minimum score can also be determined based on the several white noise audios. An encoder can be used to extract the white noise features of each white noise audio, and then the white noise similarity between each white noise feature and the standard pronunciation feature can be determined, and the average value of the multiple white noise similarities can be set as the similarity corresponding to the minimum value.

[0114] FIG. 5 is a flowchart of a method for processing a data request instruction shown in one exemplary embodiment of the present application.

[0115] As shown in FIG. 5, the method for processing a data request instruction provided by the present application includes step 501 and step 502. In step 501, a data request instruction is received. In step 502, according to the data request instruction, an encoder based on an audio recognition model, a demonstration audio, and standard pronunciation features corresponding to the demonstration audio are transmitted to the user terminal.

[0116] The demonstration audio is an audio that reads out specified content in a specified language. The encoder is used to extract evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, and the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content.

[0117] The method provided by the present application can be applied to the server side that can provide data to the user terminal.

[0118] Specifically, the user terminal can send a data request command to the server based on a user operation. The server is set with an encoder, a demonstration audio, and standard pronunciation features corresponding to the demonstration audio in the embodiments shown in FIG. 2 or FIG. 3. After receiving the data request command sent from the user terminal, the server feeds back the encoder, the demonstration audio, and the standard pronunciation features corresponding to the demonstration audio to the user terminal.

[0119] FIG. 6 is a structural diagram of a voice similarity determination device based on voice interaction shown in one exemplary embodiment of the present application.

[0120] As shown in FIG. 6, the voice similarity determination device 600 based on voice interaction provided by the present application includes an acquisition unit 610 for playing a demonstration audio and acquiring a user's evaluation audio, where the demonstration audio is an audio for reading specified content in a specified language, and the acquisition unit 610; a feature extraction unit 620 for acquiring standard pronunciation features corresponding to the demonstration audio and extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, where the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content, and the feature extraction unit 620; an analysis unit 630 for determining a feature difference between the standard pronunciation features and the evaluation pronunciation features and determining a similarity between the evaluation audio and the demonstration audio according to the feature difference.

[0121] The voice similarity determination device based on voice interaction provided by the present application is similar to that in the embodiment shown in FIG. 2, and thus will not be described repeatedly.

[0122] FIG. 7 is a structural diagram of a voice similarity determination device based on voice interaction shown in another exemplary embodiment of the present application.

[0123] As shown in FIG. 7, in the voice similarity determination device 700 based on voice interaction provided by the present application, the feature extraction unit 620 is specifically used to extract evaluation pronunciation features corresponding to the standard pronunciation features in the evaluation audio based on the encoder of the speech recognition model.

[0124] Optionally, the standard pronunciation features corresponding to the demonstration audio are obtained by fusing a plurality of reference pronunciation features. Each reference pronunciation feature is obtained by performing feature extraction on each reference audio using the encoder. Each of the reference audios is an audio that reads out the specified content in the specified language, and the demonstration audio is any one of the reference audios.

[0125] Optionally, the analysis unit 630 includes a function determination module 631 for determining a time stretching function according to the standard pronunciation features and the evaluation pronunciation features, and an alignment module 632 for determining a combination of a plurality of alignment points according to the time stretching function, the standard pronunciation features, and the evaluation pronunciation features. Each combination of alignment points includes one standard feature point in the standard pronunciation features and one evaluation feature point in the evaluation pronunciation features, and an alignment module 632, a difference determination module 633 for determining a feature difference corresponding to each combination of alignment points according to the standard feature point and the evaluation feature point included in each combination of alignment points, and a similarity determination module 634 for determining the similarity between the evaluation audio and the demonstration audio according to the feature difference of each combination of alignment points.

[0126] Optionally, the apparatus further includes a mapping unit 640, and the mapping unit 640 obtains a mapping function and setting information corresponding to the demonstration audio, and the setting information is Similarity between the evaluation audio and the demonstration audio used to indicate the mapping relationship with the score, and is used to map the similarity between the evaluation audio and the demonstration audio as a score according to the mapping function and the setting information corresponding to the demonstration audio.

[0127] Optionally, the setting information includes a maximum score, a similarity corresponding to the maximum score, a minimum score, and a similarity corresponding to the minimum score.

[0128] Optionally, the similarity corresponding to the maximum score is an average value of a plurality of reference similarities, and each of the reference similarities is a similarity between each of the reference pronunciation features and the standard pronunciation feature.

[0129] Optionally, the similarity corresponding to the minimum score is an average value of a plurality of white noise similarities, and each of the white noise similarities is a similarity between each white noise feature and the standard pronunciation feature, and each of the white noise features is obtained by performing feature extraction on each preset white noise audio using the encoder.

[0130] Optionally, before the acquisition unit 610 plays the demonstration audio, the apparatus sends a data request command to the server in response to a start command, and further includes a transceiver unit 650 used to receive the encoder, the demonstration audio, and the standard pronunciation feature corresponding to the demonstration audio.

[0131] Optionally, the speech recognition model is obtained by training an initial model using speech recognition data. The encoder for extracting pronunciation features is obtained by training the encoder in the speech recognition model using audio data of multiple language categories.

[0132] Optionally, the encoder is a three-layer long short-term memory network.

[0133] Since the speech similarity determination device based on speech interaction provided by the present application is similar to that according to the embodiment shown in FIG. 3, it will not be described repeatedly.

[0134] FIG. 8 is a structural diagram of a processing device for a data request command shown in one exemplary embodiment of the present application.

[0135] As shown in the figure, the processing device 800 for a data request command provided by the present application is arranged in a server, and the device includes a receiving unit 810 for receiving a data request command, and a transmitting unit 820 for transmitting, according to the data request command, an encoder based on a speech recognition model, a demonstration audio, and standard pronunciation features corresponding to the demonstration audio to a user terminal. The demonstration audio is an audio that reads specified content in a specified language, the encoder is used to extract evaluation pronunciation features corresponding to the standard pronunciation features from evaluation audio, and the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content.

[0136] Since the processing device for a data request command provided by the present application is similar to that according to the embodiment shown in FIG. 5, it will not be described repeatedly.

[0137] The present application further provides a computer program product including a computer program, and when the computer program is executed by a processor, the technical solution according to an embodiment of any of the above methods is realized.

[0138] The present application further provides a computer program, and when the computer program is executed by a processor, the technical solution according to an embodiment of any of the above methods is realized.

[0139] The device provided by this embodiment can be used to execute the technical solution according to the embodiment of the above method. Since its realization principle and technical effect are similar, this embodiment will not be repeatedly described here.

[0140] FIG. 9 shows a schematic structural diagram of an electronic device 900 suitable for realizing an embodiment of the present disclosure. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (abbreviated as PDAs), tablet computers (abbreviated as PADs), portable multimedia players (abbreviated as PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG. 9 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure. Multimedia Player, abbreviated as PMP), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs and desktop computers, but is not limited thereto. The electronic device shown in FIG. 9 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0141] As shown in FIG. 9, the electronic device 900 can include a processing device (such as a central processing unit or a graphics processor) 901, and the processing device can execute various appropriate operations and processes according to a program stored in a read-only memory (abbreviated as ROM) 902 or a program loaded from a storage device 908 into a random access memory (abbreviated as RAM) 903. Various programs and data necessary for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (abbreviated as I / O) interface 905 is also connected to the bus 904.

[0142] Normally, an input device 906 including a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc., an output device 907 including a liquid crystal display (abbreviated as LCD), a speaker, a vibrator, etc., a storage device 908 including a magnetic tape, a hard disk, etc., and a communication device 909 can be connected to the I / O interface 905. The communication device 909 can enable the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. FIG. 9 shows an electronic device 900 equipped with various devices, but it should be understood that not all of the illustrated devices need to be implemented or arranged. Alternatively, more or fewer devices can be implemented or arranged.

[0143] In particular, according to the embodiments of the present disclosure, the above-described process described with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product including a computer program loaded on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, the above functions limited by the method according to the embodiments of the present disclosure are executed.

[0144] It should be noted that the computer-readable medium described in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium include an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory ( ErasableProgrammable Read Only Memory (abbreviated as PROM, EPROM, or flash memory), optical fiber, Portable Compact Disk Read Only Memory (abbreviated as CD-ROM), optical storage device, magnetic memory component, or any suitable combination thereof, but not limited thereto. In the present disclosure, a computer-readable storage medium can be any tangible medium that includes or stores a program used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can further be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can transmit, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code included in the computer-readable medium can be transmitted using any suitable medium, including but not limited to wires, optical fiber cables, Radio Frequency (RF), or any suitable combination thereof.

[0145] The above computer-readable medium may be included in the above electronic device or may exist alone without being assembled into the electronic device.

[0146] One or more programs are carried on the above computer-readable medium, and when the one or more programs are executed by the above electronic device, the electronic device executes the method shown in the above embodiments.

[0147] The computer program code for performing the operations of the present disclosure can be written in one or more programming languages, including object-oriented programming languages such as Java (registered trademark), Smalltalk, C++, and conventional procedural programming languages such as the "C" language or similar programming languages, or combinations thereof. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer via any type of network, such as a Local Area Network (LAN) or a Wide Area Network (WAN), or can be connected to an external computer (e.g., connected via the Internet using an Internet service provider).

[0148] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation that can be implemented by systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions shown in the blocks can be executed in an order different from that shown in the figures. For example, two blocks shown connected and displayed can actually be executed substantially in parallel, or, depending on the related functions, the blocks may be executed in reverse order. Additionally, each block of the block diagram and / or flowchart, and combinations of blocks of the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system for performing the specified function or operation, or can also be implemented using a combination of dedicated hardware and computer instructions.

[0149] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. The name of the unit may not be intended to limit the unit itself in a particular situation. For example, the acquisition unit may be described as "the unit for acquiring the user's evaluation audio".

[0150] The functions described above in this specification can be executed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used include, but are not limited to, Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on Chip (SOC), Complex Programm able Logic Devices (CPLD), etc.

[0151] In the context of the present disclosure, a machine-readable medium can be a tangible medium that includes or stores a program that is used by or in combination with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples of machine-readable storage media can include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic memory components, or any suitable combination of the foregoing.

[0152] In a first aspect, according to one or more embodiments of the present disclosure, a method for determining voice similarity based on voice interaction is provided, and the method includes A step of playing demonstration audio to obtain user evaluation audio, wherein the demonstration audio is audio that reads specified content in a specified language. A step of obtaining standard pronunciation features corresponding to the demonstration audio and extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, wherein the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content. Determining a feature difference between the standard pronunciation features and the evaluation pronunciation features, and determining a similarity between the evaluation audio and the demonstration audio according to the feature difference.

[0153] According to one or more embodiments of the present disclosure, the step of extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio includes a step of extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio based on an encoder of an automatic speech recognition model.

[0154] According to one or more embodiments of the present disclosure, the standard pronunciation features corresponding to the demonstration audio are obtained by fusing a plurality of reference pronunciation features. Each reference pronunciation feature is obtained by performing feature extraction on each reference audio using the encoder. Each reference audio is audio that reads the specified content in the specified language, and the demonstration audio is any one of the reference audios.

[0155] According to one or more embodiments of the present disclosure, the step of determining a feature difference between the standard pronunciation features and the evaluation pronunciation features, and determining a similarity between the evaluation audio and the demonstration audio according to the feature difference includes a step of determining a time stretching function according to the standard pronunciation features and the evaluation pronunciation features. Determining a combination of a plurality of alignment points according to the time stretching function, the standard pronunciation feature, and the evaluation pronunciation feature, wherein each combination of the alignment points includes one standard feature point in the standard pronunciation feature and one evaluation feature point in the evaluation pronunciation feature; Determining a feature difference corresponding to each combination of the alignment points according to the standard feature point and the evaluation feature point included in each combination of the alignment points; Determining a similarity between the evaluation audio and the demonstration audio according to the feature difference of each combination of the alignment points.

[0156] According to one or more embodiments of the present disclosure, the method further includes: Obtaining a mapping function and setting information corresponding to the demonstration audio, wherein the setting information is used to indicate a mapping relationship between scores; Similarity between the evaluation audio and the demonstration audio Mapping the similarity between the evaluation audio and the demonstration audio as a score according to the mapping function and the setting information corresponding to the demonstration audio. According to one or more embodiments of the present disclosure, the setting information includes a maximum score, a similarity corresponding to the maximum score, a minimum score, and a similarity corresponding to the minimum score.

[0157] According to one or more embodiments of the present disclosure, the similarity corresponding to the maximum score is an average value of a plurality of reference similarities, and each reference similarity is a similarity between each reference pronunciation feature and the standard pronunciation feature.

[0158]

[0159] ​According to one or more embodiments of the present disclosure, the similarity corresponding to the minimum score is the average value of a plurality of white noise similarities, and each of the white noise similarities is the similarity between each white noise feature and the standard pronunciation feature, and each of the white noise features is obtained by performing feature extraction on each preset white noise audio by using the encoder.

[0160] According to one or more embodiments of the present disclosure, before the step of playing the demonstration audio, the method further includes: responding to a start command and sending a data request command to a server; receiving the encoder, the demonstration audio, and the standard pronunciation feature corresponding to the demonstration audio.

[0161] According to one or more embodiments of the present disclosure, the speech recognition model is obtained by training an initial model using speech recognition data. The encoder for extracting pronunciation features is obtained by training the encoder in the speech recognition model by using audio data of a plurality of language categories.

[0162] According to one or more embodiments of the present disclosure, the encoder is a three-layer long short-term memory network.

[0163] In a second aspect, according to one or more embodiments of the present disclosure, there is provided a method for processing a data request command applied to a server, the method including: receiving a data request command; sending, according to the data request command, an encoder based on a speech recognition model, a demonstration audio, and a standard pronunciation feature corresponding to the demonstration audio to a user terminal. The demonstration audio is audio that reads out specified content in a specified language, the encoder is used to extract evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, and the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content.

[0164] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an audio similarity determination device, the device comprising: An acquisition unit for playing demonstration audio to acquire a user's evaluation audio, wherein the demonstration audio is audio that reads out specified content in a specified language; A feature extraction unit for acquiring standard pronunciation features corresponding to the demonstration audio and extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, wherein the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content; An analysis unit for determining a feature difference between the standard pronunciation features and the evaluation pronunciation features and determining a similarity between the evaluation audio and the demonstration audio according to the feature difference.

[0165] According to one or more embodiments of the present disclosure, specifically, the feature extraction unit is used to extract evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio based on an encoder of an automatic speech recognition model.

[0166] According to one or more embodiments of the present disclosure, the standard pronunciation features corresponding to the demonstration audio are obtained by fusing a plurality of reference pronunciation features, and each reference pronunciation feature is obtained by performing feature extraction on each reference audio using the encoder. Each of the reference audios is an audio that reads out the specified content in the specified language, and the demonstration audio is any one of the reference audios.

[0167] According to one or more embodiments of the present disclosure, the analysis unit a function determination module for determining a time stretching function according to the standard pronunciation feature and the evaluation pronunciation feature, an alignment module for determining a combination of a plurality of alignment points according to the time stretching function, the standard pronunciation feature, and the evaluation pronunciation feature, wherein each combination of alignment points includes one standard feature point in the standard pronunciation feature and one evaluation feature point in the evaluation pronunciation feature a difference determination module for determining a feature difference corresponding to each combination of alignment points according to the standard feature point and the evaluation feature point included in each combination of alignment points, and a similarity determination module for determining the similarity between the evaluation audio and the demonstration audio according to the feature difference of each combination of alignment points.

[0168] According to one or more embodiments of the present disclosure, the apparatus obtains a mapping function and setting information corresponding to the demonstration audio, and the setting information Similarity between the evaluation audio and the demonstration audio is used to indicate the mapping relationship between the score and the score According to the mapping function and the configuration information corresponding to the demonstration audio, a mapping unit is further included, which is used to map the similarity between the evaluation audio and the demonstration audio as a score.

[0169] According to one or more embodiments of the present disclosure, the configuration information includes a maximum score, a similarity corresponding to the maximum score, a minimum score, and a similarity corresponding to the minimum score.

[0170] According to one or more embodiments of the present disclosure, the similarity corresponding to the maximum score is the average value of a plurality of reference similarities, and each of the reference similarities is the similarity between each reference pronunciation feature and the standard pronunciation feature.

[0171] According to one or more embodiments of the present disclosure, the similarity corresponding to the minimum score is the average value of a plurality of white noise similarities, and each of the white noise similarities is the similarity between each white noise feature and the standard pronunciation feature. Each of the white noise features is obtained by performing feature extraction on each preset white noise audio using the encoder.

[0172] According to one or more embodiments of the present disclosure, before the acquisition unit plays the demonstration audio, in response to the start command, the device further includes a transceiver unit, which is used to send a data request command to the server, and receive the encoder, the demonstration audio, and the standard pronunciation feature corresponding to the demonstration audio.

[0173] According to one or more embodiments of the present disclosure, the speech recognition model is obtained by training an initial model using speech recognition data. The encoder for extracting pronunciation features is obtained by training the encoder in the speech recognition model using audio data of a plurality of language categories.

[0174] According to one or more embodiments of the present disclosure, the encoder is a three-layer long short-term memory network.

[0175] In a fourth aspect, according to one or more embodiments of the present disclosure, there is provided a processing device for a data request instruction, which is arranged in a server, and the device includes: a receiving unit for receiving a data request instruction; According to the data request instruction, a voice recognition model Based on an encoder, a demonstration audio, and a transmission unit for transmitting the encoder, the demonstration audio, and standard pronunciation features corresponding to the demonstration audio to a user terminal; wherein the demonstration audio is an audio for reading out specified content in a specified language, the encoder is used to extract evaluation pronunciation features corresponding to the standard pronunciation features from evaluation audio, and the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content.

[0176] In a fifth aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device including at least one processor and a memory, wherein the memory stores computer-executable instructions, when the at least one processor executes the computer-executable instructions stored in the memory, the at least one processor executes a voice similarity determination method based on voice interaction described in the first aspect and various possible designs of the first aspect or a processing method for a data request instruction described in the second aspect and various possible designs of the second aspect.

[0177] In a sixth aspect, according to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, and when a processor executes the computer-executable instructions, a method for determining voice similarity based on voice interaction described in the first aspect and various possible designs of the first aspect or a method for processing a data request instruction described in the second aspect and various possible designs of the second aspect is realized.

[0178] In a seventh aspect, according to one or more embodiments of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, a method for determining voice similarity based on voice interaction described in the first aspect and various possible designs of the first aspect or a method for processing a data request instruction described in the second aspect and various possible designs of the second aspect is realized.

[0179] In an eighth aspect, according to one or more embodiments of the present disclosure, there is provided a computer program, and when the computer program is executed by a processor, a method for determining voice similarity based on voice interaction described in the first aspect and various possible designs of the first aspect or a method for processing a data request instruction described in the second aspect and various possible designs of the second aspect is realized.

[0180] The above description is only an explanation of some preferred embodiments of the present disclosure and an explanation of the applicable technical principles. Those skilled in the art should understand that the disclosure scope of the present disclosure is not limited to the solution formed by a specific combination of the above technical features, and without departing from the above disclosure concept, other solutions formed by any combination of the above technical features or their equivalent features, for example, solutions formed by replacing the above features with technical features having similar functions (but not limited thereto) disclosed in the present disclosure should also be covered.

[0181] Although the operations have been described in a particular order, it should not be understood that these operations are required to be performed in the particular order or sequence shown. Multitasking and parallel processing may be advantageous in certain environments. Similarly, although the above description includes some specific implementation details, these should not be construed as limiting the scope of the present disclosure. The specific features described in the context of individual embodiments may be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may be implemented alone or in any suitable sub-combination in multiple embodiments.

[0182] Although the subject matter has been described in terms of language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Conversely, the above specific features and acts are merely exemplary forms for implementing the claims.

[0183] This application claims priority to a Chinese patent application filed with the China National Intellectual Property Administration on February 7, 2021, with an application number of 202110179824.X and an application title of "Method and Device for Determining Voice Similarity, and Program Product", and all of its contents are incorporated herein by reference.

Claims

1. A method for determining voice similarity based on voice interaction, which is executed by a user terminal, comprising: playing a demonstration audio to obtain a user's evaluation audio, wherein the demonstration audio is an audio that reads specified content using a specified language; obtaining standard pronunciation features corresponding to the demonstration audio, and extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, wherein the standard pronunciation features are used to reflect specific pronunciations of the specified content in the specified language; determining a feature difference between the standard pronunciation features and the evaluation pronunciation features, and determining a similarity between the evaluation audio and the demonstration audio according to the feature difference; The step of extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio comprises extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio based on an encoder of a voice recognition model; The standard pronunciation features corresponding to the demonstration audio are obtained by fusing a plurality of reference pronunciation features, each reference pronunciation feature is obtained by performing feature extraction on each reference audio using the encoder, each reference audio is an audio that reads the specified content using the specified language, and the demonstration audio is any one of the reference audios. A method for determining voice similarity based on voice interaction, characterized by the above.

2. The step of determining a feature difference between the standard pronunciation features and the evaluation pronunciation features, and determining a similarity between the evaluation audio and the demonstration audio according to the feature difference comprises determining a time stretching function according to the standard pronunciation features and the evaluation pronunciation features; Determining a combination of a plurality of alignment points according to the time stretching function, the standard pronunciation feature, and the evaluation pronunciation feature, wherein each combination of the alignment points includes one standard feature point in the standard pronunciation feature and one evaluation feature point in the evaluation pronunciation feature; Determining a feature difference corresponding to each combination of the alignment points according to the standard feature point and the evaluation feature point included in each combination of the alignment points; Determining the similarity between the evaluation audio and the demonstration audio according to the feature difference of each combination of the alignment points; The method according to claim 1, characterized in that.

3. The method includes: Obtaining a mapping function and setting information corresponding to the demonstration audio, wherein the setting information is used to indicate a mapping relationship between the similarity and scores of the evaluation audio and the demonstration audio; Further including mapping the similarity between the evaluation audio and the demonstration audio as scores according to the mapping function and the setting information corresponding to the demonstration audio; The method according to claim 1, characterized in that.

4. The setting information includes a maximum score, a similarity corresponding to the maximum score, a minimum score, and a similarity corresponding to the minimum score. The method according to claim 3, characterized in that.

5. The similarity corresponding to the maximum score is an average value of a plurality of reference similarities, and each of the reference similarities is a similarity between each of the reference pronunciation features and the standard pronunciation feature. The method according to claim 4, characterized in that.

6. The similarity corresponding to the minimum score is an average value of a plurality of white noise similarities, and each of the white noise similarities is a similarity between each white noise feature and the standard pronunciation feature, and each of the white noise features is obtained by performing feature extraction on each preset white noise audio using the encoder. The method according to claim 4, characterized in that.

7. Before the step of playing the demonstration audio, Responding to a start command and sending a data request command to the server; receiving the encoder, the demonstration audio, and standard pronunciation features corresponding to the demonstration audio; The method according to any one of claims 1, 3 to 6, characterized in that.

8. The speech recognition model is obtained by training an initial model using speech recognition data, The encoder for extracting pronunciation features is obtained by training the encoder in the speech recognition model using audio data of a plurality of language categories. The method according to any one of claims 1, 3 to 7, characterized in that.

9. The encoder is a three-layer long short-term memory network, The method according to any one of claims 1, 3 to 8, characterized in that.

10. A method for processing a data request instruction executed by a server, comprising: receiving a data request instruction; transmitting, according to the data request instruction, an encoder based on a speech recognition model, demonstration audio, and standard pronunciation features corresponding to the demonstration audio to a user terminal; The demonstration audio is audio that reads specified content using a specified language, and the encoder is used to extract evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, and the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content. The standard pronunciation features corresponding to the demonstration audio are obtained by fusing a plurality of reference pronunciation features, and each reference pronunciation feature is obtained by performing feature extraction on each reference audio using the encoder. Each reference audio is audio that reads the specified content using the specified language, and the demonstration audio is any one of the reference audios. A method for processing a data request instruction, characterized in that.

11. An audio similarity determination device, An acquisition unit for playing demonstration audio and acquiring user evaluation audio, wherein the demonstration audio is audio that reads specified content using a specified language, the acquisition unit; A feature extraction unit for acquiring standard pronunciation features corresponding to the demonstration audio and extracting evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, wherein the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content, the feature extraction unit; An analysis unit for determining a feature difference between the standard pronunciation features and the evaluation pronunciation features and determining a similarity between the evaluation audio and the demonstration audio according to the feature difference, including; The feature extraction unit is used to extract evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio based on an encoder of a speech recognition model; The standard pronunciation features corresponding to the demonstration audio are obtained by fusing a plurality of reference pronunciation features, each reference pronunciation feature is obtained by performing feature extraction on each reference audio using the encoder, each of the reference audios is audio that reads the specified content using the specified language, and the demonstration audio is any one of the reference audios; A voice similarity determination device characterized by this.

12. A processing device for data request instructions arranged in a server, including; A receiving unit for receiving data request instructions; According to the data request instructions, a transmission unit for transmitting an encoder based on a speech recognition model, demonstration audio, and standard pronunciation features corresponding to the demonstration audio to a user terminal; The demonstration audio is audio that reads specified content using a specified language, and is used by the encoder to extract evaluation pronunciation features corresponding to the standard pronunciation features from the evaluation audio, and the standard pronunciation features are used to reflect specific pronunciations in the specified language of the specified content. The standard pronunciation features corresponding to the demonstration audio are obtained by fusing a plurality of reference pronunciation features, and each reference pronunciation feature is obtained by performing feature extraction on each reference audio using the encoder. Each of the reference audios is audio that reads the specified content using the specified language, and the demonstration audio is any one of the reference audios. A processing device for a data request instruction, characterized by the above.

13. An electronic device, including a memory, a processor, and a computer program. The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 10. An electronic device characterized by this.

14. A computer-readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement the method according to any one of claims 1 to 10. A computer-readable storage medium characterized by this.

15. A computer program for causing a computer to execute the method according to any one of claims 1 to 10, characterized by this.

Citation Information

Patent Citations

  • On-line spoken language pronunciation quality evaluation method and system

    CN104732977A

  • Method, device and system for displaying singing score

    CN104882147A

  • Method and device for voiceprint similarity comparison and application thereof in digital entertainment on-demand system

    CN105989842A

  • Spoken language comparison method

    CN106782609A

  • Automatic oral English marking method based on feature fusion

    CN106847260A