Pronunciation detection method, device, equipment and storage medium
By acquiring and analyzing the pitch arch characteristics of pronunciators in different languages, the problem of inaccurate pronunciation detection caused by unreliable accent marks in the prior art is solved, and more efficient and accurate pronunciation detection is achieved.
Patent Information
- Application Number
- CN202110992147.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-08-27
AI Technical Summary
In the prior art, the pronunciation detection is not accurate enough because the accent marking data is unreliable.
By obtaining the learning text corresponding to the first language and the pitch arch of the input pronunciation generated by the pronunciation in the second language, the effective characteristic parameters corresponding to the pitch arch are counted, and used to distinguish the pitch rhythms of different languages, thereby determining the degree of matching between the input pronunciation and the target language.
It improves the accuracy of pronunciation detection, avoids the impact of the unreliable accented label data and the multi-sound pronunciation mode on the detection, saves manual resources, and reduces the difficulty of obtaining speech sample data.
Smart Images

Figure CN114283846B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of speech technology, and in particular to a pronunciation detection method, apparatus, device and storage medium. Background Art
[0002] With the development of speech technology, people can detect speech data based on the pitch arch of speech to determine the degree of match between the speech data and the pitch rhythm corresponding to the target language.
[0003] Taking English and Chinese as examples, since there is a high correlation between the pitch curve and stress in English, the relevant technology marks the stress in the English text, and then compares the speech data corresponding to the native Chinese speaker and the native English speaker for the stressed syllables to obtain the pitch detection results.
[0004] However, since the stress is marked by the staff based on their subjective perception, the stress marking is easily affected by the staff, and the marking quality is unreliable, which leads to inaccurate detection of related technologies. Summary of the invention
[0005] The embodiments of the present application provide a pronunciation detection method, device, equipment and storage medium, which can improve the accuracy of pronunciation detection. The technical solution is as follows:
[0006] According to one aspect of an embodiment of the present application, a method for pronunciation detection is provided, the method comprising:
[0007] Acquire a learning text corresponding to the first language and an input speech corresponding to the learning text; wherein the input speech refers to the speech produced by a speaker whose native language is the second language reading the learning text aloud;
[0008] Acquiring a pitch curve of the input speech, wherein the pitch curve is used to indicate a change in intonation of the speech;
[0009] Counting at least one effective feature parameter corresponding to the pitch curve, where the effective feature parameter is used to distinguish the pitch rhythm corresponding to the first language from the pitch rhythm corresponding to the second language;
[0010] Based on the at least one effective feature parameter, a pitch detection result of the input speech is determined; wherein the pitch detection result is used to indicate the degree of matching between the pitch rhythm of the input speech and the learning text in the first language.
[0011] According to one aspect of an embodiment of the present application, a pronunciation detection device is provided, the device comprising:
[0012] An input speech acquisition module is used to acquire a learning text corresponding to the first language and an input speech corresponding to the learning text; wherein the input speech refers to the speech produced by a speaker whose native language is the second language reading the learning text aloud;
[0013] A pitch curve acquisition module, used to acquire the pitch curve of the input speech, wherein the pitch curve is used to indicate the intonation change of the speech;
[0014] An effective parameter statistics module, used for counting at least one effective characteristic parameter corresponding to the pitch curve, wherein the effective characteristic parameter is used to distinguish the pitch rhythm corresponding to the first language from the pitch rhythm corresponding to the second language;
[0015] A detection result acquisition module is used to determine the pitch detection result of the input speech based on the at least one effective feature parameter; wherein the pitch detection result is used to indicate the degree of matching between the pitch rhythm of the input speech and the learning text in the first language.
[0016] According to one aspect of an embodiment of the present application, a computer device is provided, which includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned pronunciation detection method.
[0017] Optionally, the computer device is a terminal or a server.
[0018] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above-mentioned pronunciation detection method.
[0019] According to one aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned pronunciation detection method.
[0020] The technical solution provided in the embodiments of the present application can bring the following beneficial effects:
[0021] By obtaining the pitch detection result of the input speech based on effective feature parameters that can be used to distinguish the pitch rhythm between different languages, the pitch rhythm of the input speech is determined to determine the matching degree between the input speech and the pitch rhythm corresponding to the target language, thereby solving the problem of inaccurate pronunciation detection caused by unreliable stress annotation data in the related art. Since the present application does not need to rely on manually annotated stress annotation data, it avoids the influence of unreliable stress annotation data and polyphonic rhythmic patterns on pronunciation detection, thereby improving the accuracy of pronunciation detection.
[0022] In addition, the present application can detect pronunciation through effective feature parameters without the need to compare stressed syllables one by one, thereby improving the efficiency of pronunciation detection. At the same time, since there is no need to manually mark stress, manual resources are saved and the difficulty of obtaining voice sample data is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 It is a schematic diagram of an implementation environment of a solution provided by an embodiment of the present application;
[0025] Figure 2 is a flow chart of a pronunciation detection method provided by an embodiment of the present application;
[0026] Figure 3 is a schematic diagram of a pitch arch provided by an embodiment of the present application;
[0027] Figure 4 is a schematic diagram of a pitch arch provided by another embodiment of the present application;
[0028] Figure 5 is a schematic diagram of a pitch arch provided by another embodiment of the present application;
[0029] Figure 6 is a table of significance test results provided by an embodiment of the present application;
[0030] Figure 7 and Figure 8 is a schematic diagram of a voice input interface provided by an embodiment of the present application;
[0031] Fig. 9 is a statistical diagram of the number of peaks in the second fundamental frequency curve provided by an embodiment of the present application;
[0032] Fig.10is a block diagram of a pronunciation detection device provided by an embodiment of the present application;
[0033] Fig.11 is a block diagram of a pronunciation detection device provided by another embodiment of the present application;
[0034] Fig.12 It is a structural block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0035] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0036] Please refer to Figure 1 , which shows a schematic diagram of a solution implementation environment provided by an embodiment of the present application. The solution implementation environment can be implemented as a system architecture of a pronunciation detection system. The implementation environment may include: a terminal 10 and a server 20.
[0037] The terminal 10 may be an electronic device such as a mobile phone, a tablet computer, a PC (Personal Computer), a wearable device, a vehicle-mounted device, etc. Optionally, the user may obtain the pitch detection result of the input voice through the client of the target application installed in the terminal 10. The target application may be any application that provides pronunciation detection services, such as a language learning application, a pronunciation detection application, an intelligent reading and writing application, a sentiment analysis application, etc., which is not limited in the embodiments of the present application.
[0038] The server 20 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The server 20 is used to provide background services for the client of the target application in the terminal 10. For example, the server 20 may be the background server of the above-mentioned target application.
[0039] The terminal 10 and the server 20 can communicate with each other via the network 30 .
[0040] Exemplarily, the user inputs input speech into the client of the target application (such as the speech produced by a native Chinese speaker reading English text), the client sends the input speech to the server 20, the server 20 performs pronunciation detection on the input speech, generates a pitch detection result, and sends the pitch detection result to the client, which displays the pitch detection result to the user.
[0041] In some embodiments, after acquiring the input voice, the terminal 10 may directly perform pronunciation detection on the input voice, generate a pitch detection result, and display the pitch detection result to the user.
[0042] Please refer to Figure 2 , which shows a flowchart of a pronunciation detection method provided by an embodiment of the present application. The execution subject of each step of the method can be Figure 1 In the server 20 (or terminal 10) in the implementation environment of the solution shown, the method may include the following steps (201-204):
[0043] Step 201, obtaining a learning text corresponding to a first language and an input speech corresponding to the learning text; wherein the input speech refers to the speech produced by a speaker whose native language is a second language reading aloud the learning text.
[0044] The first language and the second language refer to different languages, that is, the first language may refer to any language except the second language, and the second language may refer to any language except the first language. For example, the languages may include but are not limited to Chinese, English, French, German, etc.
[0045] In the embodiment of the present application, the first language and the second language can be distinguished by their respective corresponding pitch curves. When a speaker of a second language and a native language learns the first language, the pitch curve features corresponding to the second language are often applied to the learning process of the first language (the present application refers to the learner of the pronunciation of the first language who learns the second language as a native language as a second language speaker, and the speaker of the first language as a native language as a native speaker). Therefore, the pitch curve corresponding to the input speech generated by the second language speaker reading the learning text combines the pitch curve features of the first language and the second pitch curve features. Therefore, the pronunciation detection of the second language speaker can be realized based on the pitch curve corresponding to the native speaker and the pitch curve corresponding to the second language speaker.
[0046] Optionally, the pitch curve is a curve used to describe the pitch variation law. The pitch can be represented by the fundamental frequency (generally represented by f0), that is, the pitch curve can be represented by the fundamental frequency curve (i.e., a curve composed of a sequence of fundamental frequency values). Sound can be decomposed into many simple sine waves, among which the sine wave with the lowest frequency is the fundamental tone, and the lowest frequency is the fundamental frequency. The fundamental frequency can carry more energy and is the main component to distinguish the pitch. Therefore, the pitch curve can be represented by the fundamental frequency curve.
[0047] The learning text refers to the text content formed by the corresponding text expression in the first language. For example, if the first language is English, the learning text can be a sentence, paragraph, article, etc. expressed in English.
[0048] For example, in the case where the first language is English and the second language is Chinese, taking the learning text "His technique is ample and his musical ideas are projected beautifully" as an example, reference Figure 3 , Chart 301 shows the pitch curve (i.e. fundamental frequency curve) corresponding to the speech produced by a native English speaker reading a learning text, and Chart 302 shows the pitch curve corresponding to the speech produced by a native Chinese speaker reading a learning text. English belongs to a stress-rhythm language, and is a representative of stress languages, and its corresponding pitch curve is in a "big wave shape". For example, the pitch curve in Chart 301 is in a "big wave shape". Chinese belongs to a syllable-rhythm language, and is a representative of tonal languages. The fundamental frequency curve of Chinese is mainly formed by the interaction between word tones and intonations. Therefore, the corresponding pitch curve of a second language speaker with Chinese as his native language and English as his learning language is in a "big wave superimposed on small wave shape". For example, the pitch curve in Chart 302 is in a "big wave superimposed on small wave shape". Among them, Figure 3 The fundamental frequency value in the corresponding pitch curve is calculated in semitones, and the pitch curve is normalized for easy comparison. It can be seen that there is a significant difference between the pitch curve corresponding to native English speakers and the pitch curve corresponding to second language speakers who use Chinese as their native language and learn English. Therefore, the pronunciation detection of second language speakers who use Chinese as their native language and learn English can be achieved by comparing their corresponding pitch curves.
[0049] Optionally, although there are differences between the pitch curves corresponding to the native speaker and the pitch curves corresponding to the second language speaker, not all feature parameters corresponding to the pitch curves can be used for the pronunciation detection of the second language speaker, so it is necessary to further determine the effective feature parameters among all feature parameters, so as to perform the pronunciation detection of the second language speaker through the effective feature parameters, which will be described in detail below. Among them, the effective feature parameters refer to the feature parameters with significant differences between the pitch curves corresponding to different languages, that is, the effective feature parameters can effectively distinguish the pitch curves corresponding to different languages.
[0050] Step 202: Acquire the pitch curve of the input speech, where the pitch curve is used to indicate the intonation change of the speech.
[0051] Optionally, the process of acquiring the pitch curve is also the process of acquiring the fundamental frequency curve. Among them, each fundamental frequency value in the fundamental frequency curve can be in Hertz or in semitone. The fundamental frequency curve in Hertz can be converted into a fundamental frequency curve in semitone, and the fundamental frequency curve in semitone can be converted into a fundamental frequency curve in Hertz. The present application can improve the pronunciation detection effect by simultaneously acquiring the pitch curves in the above two forms of expression.
[0052] In one example, the process of acquiring the pitch curve can be as follows: eliminate pause segments in the input speech whose pause time is greater than a first threshold to obtain processed input speech; extract the fundamental frequency value from the processed input speech at a first interval time to obtain a first fundamental frequency curve, and the first fundamental frequency curve is a fundamental frequency curve in Hertz; convert the first fundamental frequency curve to obtain a second fundamental frequency curve, and the second fundamental frequency curve is a fundamental frequency curve with syllable markings in semitones; determine the first fundamental frequency curve and the second fundamental frequency curve as the pitch curve of the input speech.
[0053] Among them, the pause segment refers to a speech segment without pronunciation, that is, the speech segment corresponding to the crack of the speech signal. The first threshold time can be, for example, 5 milliseconds, 6 milliseconds, etc., which can be adaptively set and adjusted according to actual usage. Optionally, REAPER (Robust Epoch And Pitch EStimatoR, a language processing system) can be used to extract the fundamental frequency value from the processed input speech at a first interval time to generate a first fundamental frequency curve. In this way, the problem of calculating the fundamental frequency value at the crack can be solved.
[0054] Optionally, after obtaining the first fundamental frequency curve, the influence of differences such as gender can be weakened or eliminated by converting the first fundamental frequency curve in Hertz into a second fundamental frequency curve in semitones. The specific conversion process can be as follows: obtaining the initial pitch range of the speaker; adjusting the initial pitch range to obtain an adjusted pitch range; wherein the adjusted pitch range is larger than the initial pitch range; determining a standard pitch value based on the lower limit value of the adjusted pitch range; and converting the first fundamental frequency curve based on the standard pitch value to obtain a second fundamental frequency curve.
[0055] The initial pitch range refers to the general pitch range set by the developer for the speaker, which can cover the fundamental frequency range of all speakers and can be set according to expert experience. For example, the initial pitch range for men can be set to 75-300 Hz, and the initial pitch range for women can be set to 100-400 Hz.
[0056] In order to ensure the accuracy of the pitch range, the initial pitch range needs to be adaptively amplified and adjusted, and the specific adjustment process can be as follows: based on the initial pitch range, the first quartile and the third quartile of the speaker are calculated, based on the first quartile, the lower limit of the adjusted pitch range is determined, and based on the third quartile, the upper limit of the adjusted pitch range is determined. For example, take the initial pitch range of 100-400 Hz as an example. The quartiles of the initial pitch range are obtained respectively: 100, 200, 300 and 400, based on the first quartile 100, the lower limit of the adjusted pitch range is determined to be 100*0.75=75, based on the third quartile 300, the upper limit of the adjusted pitch range is determined to be 300*1.5=450, then the adjusted pitch range is 75-450 Hz. Optionally, the adjustment range of the upper and lower limits can be set according to the actual use situation, and the embodiment of the present application is not limited here.
[0057] In general, the fundamental frequency value at the lowest point of the pitch range often deviates, and the fundamental frequency value at 5% is more representative. Therefore, the present application uses the value at 5% of the adjusted pitch range as the standard pitch value. For example, if the adjusted pitch range is 75-450 Hz, the standard pitch value can be 75*1.05=78.75 Hz. Optionally, the position of the standard pitch in the adjusted pitch range can be adaptively adjusted according to actual needs, and the present application embodiment is not limited here.
[0058] Optionally, the process of obtaining the second fundamental frequency curve can be expressed by the following formula:
[0059]
[0060] Among them, f0[St] is the second fundamental frequency arch, f0[Hz] is the first fundamental frequency arch, and f 0-base is the standard pitch value.
[0061] Optionally, before converting the first fundamental frequency curve into the second fundamental frequency curve, the first fundamental frequency curve may be interpolated and smoothed to improve the quality of the fundamental frequency curve.
[0062] For example, when the first language is English and the second language is Chinese, the learning text is "Outside only a handful of repoerters remained" as an example. After obtaining the input speech generated by a native Chinese speaker reading the learning text, the above technical solution is used to extract the pitch curve corresponding to the input speech. Figure 4 , Figure 4 The pitch arch is shown, which includes a syllable mark in semitone units (ie Figure 4The second fundamental frequency curve 401 in units of Hz (circles in the figure) and the first fundamental frequency curve 402 in units of Hz. Figure 4 The horizontal axis is time, the unit of the left vertical axis is semitone, and the unit of the right vertical axis is Hertz.
[0063] Step 203: Count at least one effective feature parameter corresponding to the pitch curve, where the effective feature parameter is used to distinguish the pitch rhythm corresponding to the first language from the pitch rhythm corresponding to the second language.
[0064] Optionally, multiple characteristic parameters can be extracted from the pitch arch, including but not limited to: the number of peaks in the pitch arch, the number of valleys in the pitch arch, the average value of the distances between each peak in the pitch arch, the standard deviation of the distances between each peak in the pitch arch, the average value of the protrusions of each peak in the pitch arch, the standard deviation of the protrusions of each peak in the pitch arch, the average value of the fundamental frequency value corresponding to the pitch arch, the pitch range corresponding to the pitch arch, etc. The embodiment of the present application does not limit the characteristic parameters. Among them, the peak distance refers to the time width between adjacent peaks, and the protrusion degree refers to the difference between the peak and the set threshold. The set threshold can be the lowest valley value in the corresponding fundamental frequency arch, or it can be the average value of the fundamental frequency value. The present application does not limit it here.
[0065] The effective characteristic parameter refers to a characteristic parameter among the plurality of characteristic parameters that can be used to distinguish the pitch rhythm corresponding to the first language from the pitch rhythm corresponding to the second language. The pitch rhythm can be characterized by a pitch arch.
[0066] Optionally, for the first language and the second language, the effective feature parameters corresponding to the pitch curve can be determined by a significance test. In one example, the process of determining the effective feature parameters can be as follows:
[0067] 1. Acquire speech sample data, where the speech sample data is speech data obtained based on text content corresponding to the first language.
[0068] The speech sample data may include at least the following three types of speech data: speech data corresponding to a speaker whose native language is the first language, speech data with a high pitch rhythm score corresponding to a speaker whose native language is the second language, and speech data with a low pitch rhythm score corresponding to a speaker whose native language is the second language. Among them, the pitch rhythm score is used to indicate the quality of the pitch rhythm, and the higher the pitch rhythm score, the better the pitch rhythm (that is, the more it conforms to the pitch rhythm corresponding to the first language).
[0069] 2. Divide the speech sample data to obtain a first speech data set, a second speech data set, and a third speech data set.
[0070] Among them, the first speech data set includes speech data corresponding to speakers whose native language is the first language, the second speech data set includes speech data with high pitch rhythm scores corresponding to speakers whose native language is the second language, and the third speech data set includes speech data with low pitch rhythm scores corresponding to speakers whose native language is the second language.
[0071] 3. Obtain the pitch curves corresponding to the first speech data set, the second speech data set, and the third speech data set respectively.
[0072] 4. For the target feature parameters in the pitch arch, the target feature parameters corresponding to the first speech data set and the second speech data set are combined into a first target feature parameter set, the target feature parameters corresponding to the first speech data set and the third speech data set are combined into a second target feature parameter set, and the target feature parameters corresponding to the second speech data set and the third speech data set are combined into a third target feature parameter set.
[0073] Optionally, the target characteristic parameter may refer to any characteristic parameter among the above-mentioned multiple characteristic parameters.
[0074] 5. Perform significance tests on the first target feature parameter set, the second target feature parameter set and the third target feature parameter set respectively.
[0075] For example, taking the significance test of the first target feature parameter set as an example, a hypothesis is established: there is no significance between the target feature parameters corresponding to the first speech data set and the target feature parameters corresponding to the second speech data set. The significance is used to indicate that there is a difference between the target feature parameters of the two data sets.
[0076] A threshold number of sample data are randomly extracted from a first target feature parameter set; a first average value corresponding to the sample data and a second average value corresponding to the first target feature parameter set are respectively calculated, a test statistic is calculated based on the first average value and the second average value, and a corresponding boundary value table is queried based on the test statistic to determine a first probability value, wherein the first probability value is used to indicate the possibility that there is no significance between the target feature parameters corresponding to the first speech data set and the target feature parameters corresponding to the second speech data set.
[0077] If the first probability value is greater than or equal to the first threshold, the hypothesis is established; if the first probability value is less than the first threshold, the hypothesis is not established. For example, assuming that the first threshold is 0.05, if the first probability value is less than 0.05, it indicates that the hypothesis is not established, that is, there is significance between the target feature parameters corresponding to the first speech data set and the target feature parameters corresponding to the second speech data set. If the first probability value is greater than or equal to 0.05, it indicates that the hypothesis is established, that is, there is no significance between the target feature parameters corresponding to the first speech data set and the target feature parameters corresponding to the second speech data set. Optionally, if the first probability value is less than 0.01, it indicates that there is extreme significance between the target feature parameters corresponding to the first speech data set and the target feature parameters corresponding to the second speech data set; if the first probability value is less than 0.001, it indicates that there is extremely significant between the target feature parameters corresponding to the first speech data set and the target feature parameters corresponding to the second speech data set.
[0078] Obtain the significance test results corresponding to multiple feature parameters respectively.
[0079] 6. If the target feature parameter is significant in the first target feature parameter set, the second target feature parameter set and the third target feature parameter set, the target feature parameter is determined as a valid feature parameter; wherein significance is used to indicate that there is a difference between the target feature parameters of the two data sets.
[0080] Step 204: determine a pitch detection result of the input speech based on at least one valid feature parameter; wherein the pitch detection result is used to indicate the degree of matching between the pitch rhythm of the input speech and the learning text in the first language.
[0081] Optionally, the process of obtaining the pitch detection result can be as follows: based on at least one valid feature parameter, determine the pitch rhythm score of the input speech; if the pitch rhythm score is greater than a second threshold, determine that the pitch rhythm of the input speech matches the pitch rhythm of the learning text in the first language; if the pitch rhythm score is less than the second threshold, determine that the pitch rhythm of the input speech does not match the pitch rhythm of the learning text in the first language.
[0082] Optionally, the pitch rhythm score is positively correlated with the matching degree, that is, the higher the pitch rhythm score is, the more the pitch rhythm of the input speech matches the pitch rhythm of the first language.
[0083] The second threshold is a detection standard for the matching degree. Only when the second threshold is met can it be determined that the pitch rhythm of the input speech meets the pitch rhythm of the first language. The second threshold can be set to 75 points, 80 points, etc.
[0084] In one example, a method for obtaining a pitch rhythm score may be as follows: calling a logistic regression model, where the logistic regression model is trained based on a corpus sample with expert scoring annotations; and determining the pitch rhythm score of the input speech based on at least one valid feature through the logistic regression model.
[0085] Optionally, the pitch rhythm score can also be calculated by setting a calculation rule for the pitch rhythm score. For example, a weighted sum is performed on at least one valid feature parameter to obtain the pitch rhythm score. A corresponding relationship table can also be set, and the pitch rhythm score is queried from the relationship table based on at least one feature parameter. The embodiment of the present application does not limit the method for obtaining the pitch rhythm score.
[0086] In summary, the technical solution provided by the embodiment of the present application obtains the pitch detection result of the input speech based on the effective feature parameters that can be used to distinguish the pitch rhythm between different languages to determine the degree of matching between the input speech and the pitch rhythm corresponding to the target language, thereby solving the problem of inaccurate pronunciation detection caused by unreliable stress annotation data in the related art. Since the present application does not need to rely on manually annotated stress annotation data, it avoids the influence of unreliable stress annotation data and polyphonic rhythm patterns on pronunciation detection, thereby improving the accuracy of pronunciation detection.
[0087] In addition, the present application can detect pronunciation through effective feature parameters without the need to compare stressed syllables one by one, thereby improving the efficiency of pronunciation detection. At the same time, since there is no need to manually mark stress, manual resources are saved and the difficulty of obtaining voice sample data is reduced.
[0088] In addition, by extracting the fundamental frequency value in the input speech at appropriate intervals, the problem of calculating the fundamental frequency value at the crack can be effectively solved, thereby improving the accuracy of pronunciation detection. In addition, while obtaining the pitch curve in Hertz, the pitch curve in semitone is also obtained, which weakens or eliminates the influence of differences such as gender, further improving the accuracy of pronunciation detection.
[0089] In an exemplary embodiment, taking English as the first language and Chinese as the second language as an example, the process of determining the corresponding valid feature parameters between English and Chinese may be as follows:
[0090] Acquire speech sample data, which is speech data obtained based on text content corresponding to English. The speech sample data may include at least the following three types of speech data: speech data corresponding to a speaker whose native language is English (hereinafter referred to as the native language group data), speech data with high pitch rhythm scores corresponding to a speaker whose native language is Chinese (hereinafter referred to as the high-scoring second language group data), and speech data with low pitch rhythm scores corresponding to a speaker whose native language is Chinese (hereinafter referred to as the low-scoring second language group data).
[0091] The speech sample data is divided into native language group data sets, high-scoring second language group data sets, and low-scoring second language group data sets.
[0092] By adopting the technical solution provided in the above embodiment, the pitch arch corresponding to the native language group data set, the pitch arch corresponding to the high-scoring second language group data set and the pitch arch corresponding to the low-scoring second language group data set are obtained respectively.
[0093] Get multiple feature parameters corresponding to the native language group data set, the high-scoring second language group data set, and the low-scoring second language group data set. Figure 5 , Figure 5 shows that native English speakers read Figure 4 The pitch curve corresponding to the speech generated by the corresponding learning text. The pitch curve includes syllable markings in semitone units (i.e. Figure 5 The second fundamental frequency arch 501 (circles in the figure) and the first fundamental frequency arch 502 in hertz. Each circle in the second fundamental frequency arch 501 represents a syllable, the height of the center of the circle represents the fundamental frequency value corresponding to the syllable, and the diameter of the circle represents the length of the syllable. The second fundamental frequency arch 501 shows a typical declarative sentence intonation and belongs to a stressed beat language. The length of the stressed syllable is much longer than that of the unstressed syllable. The stressed syllable squeezes the length of the unstressed syllable, and the lengths of the unstressed syllables vary. For the same syllable, the variation range of the fundamental frequency value of the syllable is related to the length of the syllable, and longer syllables usually have larger variations in the fundamental frequency value. Figure 4 Compared with the second fundamental frequency arch 401 in the first fundamental frequency arch, the number of peak and trough changes in the second fundamental frequency arch 501 is less, the peak protrusion is greater, and the peak distance is greater. Therefore, the present application sets the multiple characteristic parameters as: the number of peaks in the first fundamental frequency arch corresponding to the pitch arch, the number of peaks in the second fundamental frequency arch corresponding to the pitch arch, the average value of the distances between each peak in the second fundamental frequency arch, the standard deviation of the distances between each peak in the second fundamental frequency arch, the average value of the protrusion degree of each peak in the second fundamental frequency arch, and the standard deviation of the protrusion degree of each peak in the second fundamental frequency arch.
[0094] For the target feature parameters among the multiple feature parameters, the target feature parameters corresponding to the native language group data set and the high-scoring second language group data set are combined into a first target feature parameter set, the target feature parameters corresponding to the native language group data set and the low-scoring second language group data set are combined into a second target feature parameter set, and the target feature parameters corresponding to the high-scoring second language group data set and the low-scoring second language group data set are combined into a third target feature parameter set.
[0095] The significance test is performed on the first target feature parameter set, the second target feature parameter set and the third target feature parameter set respectively.
[0096] For example, reference Figure 6 , which shows a table of significance test results provided by an embodiment of the present application. The average value of the number of peaks in the second fundamental frequency arch corresponding to the pitch arch and the distances between each peak in the second fundamental frequency arch can effectively distinguish the corresponding pitch rhythms between the native language group data set, the high-scoring second language group data set and the low-scoring second language group data set, then the number of peaks in the second fundamental frequency arch corresponding to the pitch arch and the average value of the distances between each peak in the second fundamental frequency arch can be determined as the corresponding effective feature parameters between English and Chinese, that is, the corresponding effective feature parameters between English and Chinese include at least one of the following: the number of peaks in the second fundamental frequency arch corresponding to the pitch arch and the average value of the distances between each peak in the second fundamental frequency arch.
[0097] Optionally, the number of peaks in the first fundamental frequency arch corresponding to the pitch arch can effectively distinguish the corresponding pitch rhythm between the native language group data set and the high-scoring second language group data set, and the corresponding pitch rhythm between the high-scoring second language group data set and the low-scoring second language group data set. The standard deviation of the protrusion degree of each crest in the second fundamental frequency arch can effectively distinguish the corresponding pitch rhythm between the native language group data set and the high-scoring second language group data set, and the corresponding pitch rhythm between the native language group data set and the low-scoring second language group data set. The standard deviation of each peak distance in the second fundamental frequency arch can effectively distinguish the corresponding pitch rhythm between the native language group data set and the low-scoring second language group data set, and the corresponding pitch rhythm between the high-scoring second language group data set and the low-scoring second language group data set. Then the number of peaks in the first fundamental frequency arch corresponding to the pitch arch, the standard deviation of the protrusion degree of each crest in the second fundamental frequency arch, and the standard deviation of each peak distance in the second fundamental frequency arch can be determined as the corresponding important characteristic parameters between English and Chinese.
[0098] Since the average value of the protrusion degree of each peak in the second fundamental frequency curve cannot distinguish the pitch rhythm between any two sets of data sets, the average value of the protrusion degree of each peak in the second fundamental frequency curve can be determined as an invalid feature parameter corresponding to English and Chinese.
[0099] In summary, the technical solution provided by the embodiment of the present application obtains the pitch detection result of the input speech based on the effective feature parameters that can be used to distinguish the pitch rhythm between different languages to determine the degree of matching between the input speech and the pitch rhythm corresponding to the target language, thereby solving the problem of inaccurate pronunciation detection caused by unreliable stress annotation data in the related art. Since the present application does not need to rely on manually annotated stress annotation data, it avoids the influence of unreliable stress annotation data and polyphonic rhythm patterns on pronunciation detection, thereby improving the accuracy of pronunciation detection.
[0100] In addition, the present application can detect pronunciation through effective feature parameters without the need to compare stressed syllables one by one, thereby improving the efficiency of pronunciation detection. At the same time, since there is no need to manually mark stress, manual resources are saved and the difficulty of obtaining voice sample data is reduced.
[0101] In an exemplary embodiment, taking English as the first language and Chinese as the second language as an example, referring to Figure 1 , the pronunciation detection process can be as follows:
[0102] 1. The target application in the terminal 10 obtains a learning text corresponding to English and an input speech generated by a speaker whose native language is Chinese reading the learning text.
[0103] Optionally, the learning text may be a learning text input by the speaker, or a learning text selected by the speaker from a learning text database provided by the target application. The input speech may be a speech recorded by the speaker in real time, or a previously recorded speech. Figure 7 and Figure 8 , the speaker inputs the learning text "I know the fact, do you know" in the voice input interface 701, and starts recording the input voice by triggering the control 702. After reading the learning text, the speaker ends the recording by triggering 703.
[0104] Optionally, the target application may be any application that provides pronunciation detection services, such as language learning applications, pronunciation detection applications, intelligent reading and writing applications, sentiment analysis applications, and the like.
[0105] 2. The target application in the terminal 10 sends the learning text and input speech to the server 20.
[0106] 3. The server 20 scores the input speech and obtains a pitch and rhythm score.
[0107] Optionally, the scoring process may include the following:
[0108] After acquiring the learning text and the input speech, the server 20 sends the learning text and the input speech to the pitch curve analyzer.
[0109] The pitch curve analysis extracts the pitch curve of the input speech and counts at least one effective characteristic parameter corresponding to the pitch curve. Optionally, the at least one effective characteristic parameter is: the number of peaks in the second fundamental frequency curve corresponding to the pitch curve and the average value of the distances between each peak in the second fundamental frequency curve.
[0110] Since the small fluctuations in the fundamental frequency curve usually reflect the situation of the consonant and vowel segments that are difficult to detect, and we are only interested in the fundamental frequency fluctuations that can be perceived above the syllable level, before obtaining the effective feature parameters, the fundamental frequency curve can also be screened to remove the invalid peaks and invalid troughs in the fundamental frequency curve, that is, to retain the valid peaks and valid troughs in the fundamental frequency curve. Among them, the valid peaks and valid troughs can reflect the fundamental frequency fluctuations above the syllable level.
[0111] Optionally, the method for determining effective peaks and effective troughs is the same, and the following description will take the process of determining effective peaks as an example, and the specific content may be as follows: obtain the first peak, the first trough and the second trough in the pitch arch; wherein the first trough refers to the previous trough corresponding to the first peak, and the second trough refers to the next trough corresponding to the first peak; if the fundamental frequency difference between the first peak and the first trough is greater than a third threshold, and / or the fundamental frequency difference between the first peak and the second trough is greater than the third threshold, the first peak is determined as a valid peak; based on the effective peaks in the pitch arch, obtain at least one valid characteristic parameter.
[0112] Exemplarily, when the fundamental frequency value in the pitch curve is in semitones (i.e., the second fundamental frequency curve), the third threshold value can be set to 0.5, 0.6 semitones, etc., and the third threshold value can be set and adjusted according to actual usage. For example, we set the third threshold value to 0.5 semitones, that is, the target peak and the previous or next trough must differ by at least 0.5 semitones before it can be determined as a valid peak. That is, the target trough and the previous or next peak must differ by at least 0.5 semitones before it can be determined as a valid trough. Optionally, the first fundamental frequency curve can be equivalently screened based on the screened second fundamental frequency curve.
[0113] Alternatively, the topographic prominence analysis technique in MATLAB (a mathematical software) may be used to extract the peaks and troughs in the fundamental frequency curve, as well as parameters related to the peaks and troughs, such as the number of peaks, peak values, and peak distances.
[0114] Optionally, for ease of comparison, the time width between the peaks and troughs can be normalized to between 0 and 1 before obtaining the effective characteristic parameters. Optionally, the process of obtaining the peak distance can be as follows: normalize the time width between adjacent peaks and troughs in the second fundamental frequency arch to obtain a processed second fundamental frequency arch; and obtain each peak distance from the processed second fundamental frequency arch.
[0115] For example, reference Fig. 9 , which shows a statistical diagram of the number of peaks in the second fundamental frequency curve provided by an embodiment of the present application. Curve 901 is a low-scoring bilingual group data set (i.e. Fig. 9 The arch of the number of peaks in the second fundamental frequency arch of each sentence corresponding to EnLo in , and the arch 902 is a high-scoring bilingual group data set (i.e. Fig. 9 The arch of the peak number in the second fundamental frequency arch of each sentence corresponding to EnHi in the native language group data set (i.e. Fig. 9 The number of peaks in the second fundamental frequency curve of each sentence corresponding to EnNa in ( ). Among them, the number of peaks in the second fundamental frequency curve of curve 901 in each sentence is the largest, followed by curve 902, and the number of peaks in the second fundamental frequency curve of curve 903 in each sentence is the smallest. It can be seen that the number of peaks in the second fundamental frequency curve can effectively distinguish the low-scoring bilingual group data set, the high-scoring bilingual group data set and the native language group data set.
[0116] The pitch arch analysis inputs at least one effective characteristic parameter corresponding to the pitch arch into a logistic regression function to obtain a pitch rhythm score. The logistic regression function is obtained by fitting expert scoring training.
[0117] 4. The server 20 sends the pitch rhythm score to the target application.
[0118] 5. The target application displays the pitch rhythm score.
[0119] Optionally, the higher the pitch rhythm score is, the better the pitch rhythm corresponding to the input speech is, that is, the more the pitch rhythm of the input speech matches the corresponding pitch rhythm of English.
[0120] Optionally, in some embodiments, while obtaining the valid characteristic parameters corresponding to the pitch arch, the important characteristic parameters corresponding to the pitch arch can also be obtained, such as the number of peaks in the first fundamental frequency arch corresponding to the pitch arch, the standard deviation of the protrusion degree of each peak in the second fundamental frequency arch, and the standard deviation of the distance between each peak in the second fundamental frequency arch. Then, based on the valid characteristic parameters and the important characteristic parameters, the pitch rhythm of the input speech is scored. In this way, the comprehensiveness of the pitch rhythm scoring can be improved, thereby further improving the accuracy of pronunciation detection. For invalid characteristic parameters (such as the average value of the protrusion degree of each peak in the second fundamental frequency arch), they can be added or deleted according to actual usage needs.
[0121] In a feasible example, the target application can be an application that scores the pronunciation rhythm of the input speech. After obtaining the pitch rhythm score, the server 20 can also obtain scores of other rhythmic parameters through the rhyme analyzer, such as the rhythm score (i.e., the score corresponding to the duration of the syllable), the intensity score (i.e., the score corresponding to the loudness of the syllable), the stress score, etc. Then the pitch rhythm score and the scores of other rhythmic parameters are comprehensively scored to obtain the rhythm score. The server 20 sends the rhythm score to the target application, and the target application displays the rhythm score.
[0122] Optionally, the target application may directly display the prosody score, or display the prosody score in the form of a star rating, or display the prosody score and the prosody star rating at the same time. For example, the prosody score ranges from 0 to 100, and the higher the prosody score, the better the prosody of the input speech. The prosody star rating may range from 0 to 5 stars, and the more stars there are, the better the prosody of the input speech.
[0123] Optionally, the target application can also mark incorrectly stressed syllables and correctly stressed syllables in different colors in the learning text, and mark the correct pronunciation and incorrect pronunciation of incorrectly stressed syllables, so that the speaker can understand the incorrect pronunciation and learn the correct pronunciation. Exemplarily, the correctly stressed syllables can be marked in green, the incorrectly stressed syllables can be marked in red, and the incorrect pronunciation can be marked in red, and the correct pronunciation can be marked in orange.
[0124] In summary, the technical solution provided by the embodiment of the present application obtains the pitch detection result of the input speech based on the effective feature parameters that can be used to distinguish the pitch rhythm between different languages to determine the degree of matching between the input speech and the pitch rhythm corresponding to the target language, thereby solving the problem of inaccurate pronunciation detection caused by unreliable stress annotation data in the related art. Since the present application does not need to rely on manually annotated stress annotation data, it avoids the influence of unreliable stress annotation data and polyphonic rhythm patterns on pronunciation detection, thereby improving the accuracy of pronunciation detection.
[0125] In addition, the present application can detect pronunciation through effective feature parameters without the need to compare stressed syllables one by one, thereby improving the efficiency of pronunciation detection. At the same time, since there is no need to manually mark stress, manual resources are saved and the difficulty of obtaining voice sample data is reduced.
[0126] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.
[0127] Please refer to Fig.10 , which shows a block diagram of a pronunciation detection device provided by an embodiment of the present application. The device has the function of implementing the above method example, and the function can be implemented by hardware, or by hardware executing corresponding software. The device can be a computer device, or it can be set in a computer device. The device 1000 may include: an input speech acquisition module 1001, a pitch curve acquisition module 1002, an effective parameter statistics module 1003 and a detection result acquisition module 1004.
[0128] The input speech acquisition module 1001 is used to acquire a learning text corresponding to a first language and an input speech corresponding to the learning text; wherein the input speech refers to the speech generated by a speaker whose native language is a second language reading the learning text.
[0129] The pitch curve acquisition module 1002 is used to acquire the pitch curve of the input speech, where the pitch curve is used to indicate the intonation change of the speech.
[0130] The effective parameter statistics module 1003 is used to count at least one effective feature parameter corresponding to the pitch curve, and the effective feature parameter is used to distinguish the pitch rhythm corresponding to the first language and the pitch rhythm corresponding to the second language.
[0131] The detection result acquisition module 1004 is used to determine the pitch detection result of the input speech based on the at least one valid feature parameter; wherein the pitch detection result is used to indicate the degree of matching between the pitch rhythm of the input speech and the learning text in the first language.
[0132] In an exemplary embodiment, Fig.11 As shown, the pitch curve acquisition module 1002 includes: a speech processing acquisition submodule 1002a, a first curve acquisition submodule 1002b, a second curve acquisition submodule 1002c and a pitch curve acquisition submodule 1002d.
[0133] The processed speech acquisition submodule 1002a is used to remove pause segments with a pause time greater than a first threshold in the input speech to obtain processed input speech.
[0134] The first curve acquisition submodule 1002b is used to extract the fundamental frequency value from the processed input speech at a first interval time to obtain a first fundamental frequency curve, where the first fundamental frequency curve is a fundamental frequency curve in Hertz.
[0135] The second pitch acquisition submodule 1002c is used to convert the first fundamental frequency pitch to obtain a second fundamental frequency pitch, where the second fundamental frequency pitch is a fundamental frequency pitch with syllable markings in semitone units.
[0136] The pitch curve acquisition submodule 1002d is used to determine the first fundamental frequency curve and the second fundamental frequency curve as the pitch curve of the input speech.
[0137] In an exemplary embodiment, the second arch acquisition submodule 1002c is used to:
[0138] Obtaining an initial pitch range of the speaker;
[0139] Adjusting the initial pitch range to obtain an adjusted pitch range; wherein the adjusted pitch range is larger than the initial pitch range;
[0140] Determining a standard pitch value based on the lower limit value of the adjusted pitch range;
[0141] Based on the standard pitch value, the first fundamental frequency curve is converted to obtain the second fundamental frequency curve.
[0142] In an exemplary embodiment, the detection result acquisition module 1004 is used to:
[0143] Determining a pitch rhythm score of the input speech based on the at least one effective feature parameter;
[0144] If the pitch rhythm score is greater than a second threshold, determining that the pitch rhythm of the input speech matches the pitch rhythm of the learning text in the first language;
[0145] If the pitch rhythm score is less than the second threshold, it is determined that the pitch rhythm of the input speech does not match the pitch rhythm of the learning text in the first language.
[0146] In an exemplary embodiment, the detection result acquisition module 1004 is further used to:
[0147] Calling a logistic regression model, wherein the logistic regression model is trained based on a corpus sample with expert scoring annotations;
[0148] The pitch rhythm score of the input speech is determined based on the at least one effective feature through the logistic regression model.
[0149] In an exemplary embodiment, the effective parameter statistics module 1003 is used to:
[0150] Acquire voice sample data, where the voice sample data is voice data obtained based on text content corresponding to the first language;
[0151] The speech sample data is divided to obtain a first speech data set, a second speech data set and a third speech data set; wherein the first speech data set includes speech data corresponding to speakers whose native language is the first language, the second speech data set includes speech data with high pitch rhythm scores corresponding to speakers whose native language is the second language, and the third speech data set includes speech data with low pitch rhythm scores corresponding to speakers whose native language is the second language;
[0152] Acquire the pitch curves corresponding to the first speech data set, the second speech data set, and the third speech data set respectively;
[0153] For the target feature parameters in the pitch arch, the target feature parameters corresponding to the first speech data set and the second speech data set are combined into a first target feature parameter set, the target feature parameters corresponding to the first speech data set and the third speech data set are combined into a second target feature parameter set, and the target feature parameters corresponding to the second speech data set and the third speech data set are combined into a third target feature parameter set;
[0154] performing significance tests on the first target feature parameter set, the second target feature parameter set, and the third target feature parameter set respectively;
[0155] If the target feature parameter is significant in the first target feature parameter set, the second target feature parameter set and the third target feature parameter set, the target feature parameter is determined as the valid feature parameter; wherein the significance is used to indicate that there is a difference between the target feature parameters of two data sets.
[0156] In an exemplary embodiment, the first language is English and the second language is Chinese;
[0157] The at least one effective characteristic parameter includes at least one of the following: the number of peaks in the second fundamental frequency arch corresponding to the pitch arch, and the average value of the distances between each peak in the second fundamental frequency arch; wherein the peak distance refers to the time width between adjacent peaks.
[0158] In an exemplary embodiment, the effective parameter statistics module 1003 is further used to:
[0159] Obtain a first wave crest, a first wave trough, and a second wave trough in the pitch arch; wherein the first wave trough refers to a previous wave trough corresponding to the first wave crest, and the second wave trough refers to a subsequent wave trough corresponding to the first wave crest;
[0160] If the fundamental frequency difference between the first wave peak and the first wave trough is greater than a third threshold, and / or the fundamental frequency difference between the first wave peak and the second wave trough is greater than the third threshold, determining the first wave peak as a valid wave peak;
[0161] Based on the effective peaks in the pitch arch, the at least one effective characteristic parameter is obtained.
[0162] In an exemplary embodiment, the effective parameter statistics module 1003 is further used to:
[0163] Normalizing the time widths between adjacent peaks and troughs in the second fundamental frequency curve to obtain a processed second fundamental frequency curve;
[0164] The respective peak distances are obtained from the processed second fundamental frequency curve.
[0165] In summary, the technical solution provided by the embodiment of the present application obtains the pitch detection result of the input speech based on the effective feature parameters that can be used to distinguish the pitch rhythm between different languages to determine the degree of matching between the input speech and the pitch rhythm corresponding to the target language, thereby solving the problem of inaccurate pronunciation detection caused by unreliable stress annotation data in the related art. Since the present application does not need to rely on manually annotated stress annotation data, it avoids the influence of unreliable stress annotation data and polyphonic rhythm patterns on pronunciation detection, thereby improving the accuracy of pronunciation detection.
[0166] In addition, the present application can detect pronunciation through effective feature parameters without the need to compare stressed syllables one by one, thereby improving the efficiency of pronunciation detection. At the same time, since there is no need to manually mark stress, manual resources are saved and the difficulty of obtaining voice sample data is reduced.
[0167] It should be noted that the device provided in the above embodiment, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the content structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0168] Please refer to Fig.12 , which shows a block diagram of a computer device provided in one embodiment of the present application. The computer device can be used to implement the pronunciation detection method provided in the above embodiment. Specifically:
[0169] The computer device 1200 includes a central processing unit (such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit) and an FPGA (Field Programmable Gate Array)) 1201, a system memory 1204 including a RAM (Random-Access Memory) 1202 and a ROM (Read-Only Memory) 1203, and a system bus 1205 connecting the system memory 1204 and the central processing unit 1201. The computer device 1200 also includes a basic input / output system (Input Output System, I / O system) 1206 for helping various devices in the server to transmit information, and a large-capacity storage device 1207 for storing an operating system 1213, application programs 1214 and other program modules 1215.
[0170] The basic input / output system 1206 includes a display 1208 for displaying information and an input device 1209 such as a mouse and a keyboard for user inputting information. The display 1208 and the input device 1209 are connected to the central processing unit 1201 through an input / output controller 1210 connected to the system bus 1205. The basic input / output system 1206 may also include an input / output controller 1210 for receiving and processing inputs from a plurality of other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1210 also provides output to a display screen, a printer, or other types of output devices.
[0171] The mass storage device 1207 is connected to the central processing unit 1201 via a mass storage controller (not shown) connected to the system bus 1205. The mass storage device 1207 and its associated computer readable medium provide non-volatile storage for the computer device 1200. That is, the mass storage device 1207 may include a computer readable medium (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0172] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, cassettes, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage medium is not limited to the above. The above-mentioned system memory 1204 and mass storage device 1207 can be collectively referred to as memory.
[0173] According to the embodiment of the present application, the computer device 1200 can also be connected to a remote computer on the network through a network such as the Internet. That is, the computer device 1200 can be connected to the network 1212 through the network interface unit 1211 connected to the system bus 1205, or the network interface unit 1211 can be used to connect to other types of networks or remote computer systems (not shown).
[0174] The memory also includes at least one instruction, at least one program, code set or instruction set, which is stored in the memory and configured to be executed by one or more processors to implement the above-mentioned pronunciation detection method.
[0175] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein the storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set, when executed by a processor, implements the above-mentioned pronunciation detection method.
[0176] Optionally, the computer readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives) or optical disks, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0177] In an exemplary embodiment, a computer program product or a computer program is also provided, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned pronunciation detection method.
[0178] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application are not limited to this.
[0179] The above description is only an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A pronunciation detection method, characterized in that: The method comprises: Acquire a learning text corresponding to the first language and an input speech corresponding to the learning text; wherein the input speech refers to the speech produced by a speaker whose native language is the second language reading the learning text aloud; Acquiring a pitch curve of the input speech, wherein the pitch curve is used to indicate a change in intonation of the speech; Counting at least one effective feature parameter corresponding to the pitch curve, where the effective feature parameter is used to distinguish the pitch rhythm corresponding to the first language from the pitch rhythm corresponding to the second language; Based on the at least one effective feature parameter, determining a pitch detection result of the input speech; wherein the pitch detection result is used to indicate a matching degree between the pitch rhythm of the input speech and the learning text in the first language; The process of determining the effective characteristic parameters includes the following: Acquire voice sample data, where the voice sample data is voice data obtained based on text content corresponding to the first language; The speech sample data is divided to obtain a first speech data set, a second speech data set and a third speech data set; wherein the first speech data set includes speech data corresponding to speakers whose native language is the first language, the second speech data set includes speech data with high pitch rhythm scores corresponding to speakers whose native language is the second language, and the third speech data set includes speech data with low pitch rhythm scores corresponding to speakers whose native language is the second language; Acquire the pitch curves corresponding to the first speech data set, the second speech data set, and the third speech data set respectively; For the target feature parameters in the pitch arch, the target feature parameters corresponding to the first speech data set and the second speech data set are combined into a first target feature parameter set, the target feature parameters corresponding to the first speech data set and the third speech data set are combined into a second target feature parameter set, and the target feature parameters corresponding to the second speech data set and the third speech data set are combined into a third target feature parameter set; performing significance tests on the first target feature parameter set, the second target feature parameter set, and the third target feature parameter set respectively; If the target feature parameter is significant in the first target feature parameter set, the second target feature parameter set and the third target feature parameter set, the target feature parameter is determined as the valid feature parameter; wherein the significance is used to indicate that there is a difference between the target feature parameters of two data sets.
2. The method according to claim 1, characterized in that The step of obtaining the pitch curve of the input speech comprises: Eliminate pause segments in the input speech whose pause time is greater than a first threshold to obtain processed input speech; Extracting a fundamental frequency value from the processed input speech at a first interval time to obtain a first fundamental frequency curve, wherein the first fundamental frequency curve is a fundamental frequency curve in Hertz; Converting the first fundamental frequency curve to obtain a second fundamental frequency curve, wherein the second fundamental frequency curve is a fundamental frequency curve with syllable markings in semitone units; The first fundamental frequency curve and the second fundamental frequency curve are determined as the pitch curve of the input speech.
3. The method according to claim 2, characterized in that The converting the first fundamental frequency curve to obtain a second fundamental frequency curve comprises: Obtaining an initial pitch range of the speaker; Adjusting the initial pitch range to obtain an adjusted pitch range; wherein the adjusted pitch range is larger than the initial pitch range; Determining a standard pitch value based on the lower limit value of the adjusted pitch range; Based on the standard pitch value, the first fundamental frequency curve is converted to obtain the second fundamental frequency curve.
4. The method according to claim 1, characterized in that: The step of determining the pitch detection result of the input speech based on the at least one effective feature parameter comprises: Determining a pitch rhythm score of the input speech based on the at least one effective feature parameter; If the pitch rhythm score is greater than a second threshold, determining that the pitch rhythm of the input speech matches the pitch rhythm of the learning text in the first language; If the pitch rhythm score is less than the second threshold, it is determined that the pitch rhythm of the input speech does not match the pitch rhythm of the learning text in the first language.
5. The method according to claim 4, characterized in that The step of determining the pitch rhythm score of the input speech based on the at least one effective feature parameter comprises: Calling a logistic regression model, wherein the logistic regression model is trained based on a corpus sample with expert scoring annotations; The pitch rhythm score of the input speech is determined based on the at least one effective feature through the logistic regression model.
6. The method according to claim 1, characterized in that The first language is English, and the second language is Chinese; The at least one effective characteristic parameter includes at least one of the following: the number of peaks in the second fundamental frequency arch corresponding to the pitch arch, and the average value of the distances between the peaks in the second fundamental frequency arch; The peak distance refers to the time width between adjacent peaks.
7. The method according to claim 6, characterized in that The method further comprises: Obtain a first wave crest, a first wave trough, and a second wave trough in the pitch arch; wherein the first wave trough refers to a previous wave trough corresponding to the first wave crest, and the second wave trough refers to a subsequent wave trough corresponding to the first wave crest; If the fundamental frequency difference between the first wave peak and the first wave trough is greater than a third threshold, and / or the fundamental frequency difference between the first wave peak and the second wave trough is greater than the third threshold, determining the first wave peak as a valid wave peak; Based on the effective peaks in the pitch arch, the at least one effective characteristic parameter is obtained.
8. The method according to claim 7, characterized in that The method further comprises: Normalizing the time widths between adjacent peaks and troughs in the second fundamental frequency curve to obtain a processed second fundamental frequency curve; The respective peak distances are obtained from the processed second fundamental frequency curve.
9. A pronunciation detection device, characterized in that: The device comprises: An input speech acquisition module is used to acquire a learning text corresponding to the first language and an input speech corresponding to the learning text; wherein the input speech refers to the speech produced by a speaker whose native language is the second language reading the learning text aloud; A pitch curve acquisition module, used to acquire the pitch curve of the input speech, wherein the pitch curve is used to indicate the intonation change of the speech; An effective parameter statistics module, used for counting at least one effective characteristic parameter corresponding to the pitch curve, wherein the effective characteristic parameter is used to distinguish the pitch rhythm corresponding to the first language from the pitch rhythm corresponding to the second language; A detection result acquisition module, used to determine a pitch detection result of the input speech based on the at least one valid feature parameter; wherein the pitch detection result is used to indicate a matching degree between the pitch rhythm of the input speech and the learning text in the first language; The process of determining the effective characteristic parameters includes the following: Acquire voice sample data, where the voice sample data is voice data obtained based on text content corresponding to the first language; The speech sample data is divided to obtain a first speech data set, a second speech data set and a third speech data set; wherein the first speech data set includes speech data corresponding to speakers whose native language is the first language, the second speech data set includes speech data with high pitch rhythm scores corresponding to speakers whose native language is the second language, and the third speech data set includes speech data with low pitch rhythm scores corresponding to speakers whose native language is the second language; Acquire the pitch curves corresponding to the first speech data set, the second speech data set, and the third speech data set respectively; For the target feature parameters in the pitch arch, the target feature parameters corresponding to the first speech data set and the second speech data set are combined into a first target feature parameter set, the target feature parameters corresponding to the first speech data set and the third speech data set are combined into a second target feature parameter set, and the target feature parameters corresponding to the second speech data set and the third speech data set are combined into a third target feature parameter set; performing significance tests on the first target feature parameter set, the second target feature parameter set, and the third target feature parameter set respectively; If the target feature parameter is significant in the first target feature parameter set, the second target feature parameter set and the third target feature parameter set, the target feature parameter is determined as the valid feature parameter; wherein the significance is used to indicate that there is a difference between the target feature parameters of two data sets.
10. The device according to claim 9, characterized in that The pitch arch acquisition module comprises: A speech acquisition processing submodule is used to remove pause segments whose pause time is greater than a first threshold in the input speech to obtain processed input speech; A first frequency curve acquisition submodule is configured to extract a fundamental frequency value from the processed input speech at a first interval time to obtain a first fundamental frequency frequency curve, wherein the first fundamental frequency curve is a fundamental frequency curve in Hertz; A second frequency arch acquisition submodule is used to convert the first fundamental frequency arch to obtain a second fundamental frequency arch, where the second fundamental frequency arch is a fundamental frequency arch with syllable markings in semitone units; The pitch curve acquisition submodule is used to determine the first fundamental frequency curve and the second fundamental frequency curve as the pitch curve of the input speech.
11. The device according to claim 10, characterized in that The second arch acquisition submodule is used for: Obtaining an initial pitch range of the speaker; Adjusting the initial pitch range to obtain an adjusted pitch range; wherein the adjusted pitch range is larger than the initial pitch range; Determining a standard pitch value based on the lower limit value of the adjusted pitch range; Based on the standard pitch value, the first fundamental frequency curve is converted to obtain the second fundamental frequency curve.
12. The device according to claim 9, characterized in that The detection result acquisition module is used to: Determining a pitch rhythm score of the input speech based on the at least one effective feature parameter; If the pitch rhythm score is greater than a second threshold, determining that the pitch rhythm of the input speech matches the pitch rhythm of the learning text in the first language; If the pitch rhythm score is less than the second threshold, it is determined that the pitch rhythm of the input speech does not match the pitch rhythm of the learning text in the first language.
13. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, and the processor loads and executes the at least one instruction to implement the pronunciation detection method according to any one of claims 1 to 8.
14. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the pronunciation detection method according to any one of claims 1 to 8.
15. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. A processor reads and executes the computer instructions from the computer-readable storage medium to implement the pronunciation detection method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Language learning device
JP2007147783A