Pronunciation recognition learning auxiliary method and system based on voice processing
By constructing a pronunciation feature extraction model and spectral diagram technology, the problem of inaccurate judgment of user pronunciation in existing pronunciation recognition technology is solved, and more accurate pronunciation correction and learning assistance are achieved.
Patent Information
- Application Number
- CN202510827741.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing pronunciation recognition technology is not reasonable enough when judging whether the user's pronunciation is standard, and lacks characteristic data, which makes it impossible to effectively help users learn and correct pronunciation, especially for children.
Through speech recognition, the input speech is converted into text, a pronunciation feature extraction model is constructed, the pronunciation features in the spectral map is extracted, and the pronunciation features and text features are compared to the pronunciation characteristics to determine whether the user's pronunciation is standard.
It improves the accuracy and effectiveness of pronunciation standards and can help users more accurately correct pronunciation, especially pronunciation problems of children.
Smart Images

Figure CN120356485A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pronunciation recognition, and specifically to a pronunciation recognition learning assistance method and system based on speech processing. Background Art
[0002] The speech pronunciation recognition technology refers to a technology that analyzes the characteristics of speech signals, combines acoustic models and language models, converts human speech into text or instructions, and evaluates the pronunciation accuracy, so as to assist users in optimizing pronunciation learning.
[0003] The speech pronunciation recognition technology is usually applied to users' pronunciation learning. The existing speech pronunciation recognition technology only considers that the pronunciation is correct when the correct text is recognized from the user's speech. The pronunciation learners are usually students or adults, rather than children. Therefore, more strict judgment criteria are needed to help users correct their pronunciation. At the same time, when the existing speech pronunciation recognition technology processes speech, it does not convert the speech into characteristic data and extract its pronunciation features, resulting in inaccurate judgment of whether the user's pronunciation is standard. For example, in the patent application with the publication number CN109461436A, "A method and system for correcting pronunciation errors in speech recognition" is disclosed. This solution does not give a detailed analysis process on how to judge whether the user's pronunciation is standard. It only judges whether the user's pronunciation is standard by recognizing the text of the speech. However, pronunciation learners are usually not children, and more strict judgment criteria are needed to help users correct their pronunciation. The existing speech pronunciation recognition technology also has problems such as unreasonable judgment methods for whether the user's pronunciation is standard and non-characteristic data in the analysis process, resulting in the inability to effectively help users with pronunciation learning and correction. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems in the existing technology to some extent. By performing speech recognition on the input speech of the user, extracting the input text, preprocessing the speech to be analyzed, obtaining the time-domain diagram of the speech to be analyzed, framing it, dividing the time-domain diagram into different frame data, then performing short-time Fourier transform on the frame data to convert it into the spectrum of the frame data, named frame spectrum, converting the input speech into a spectrogram based on the frame spectrum, and then extracting the pronunciation features of the speech to be analyzed based on the spectrogram, extracting the pronunciation features of the input speech and the input text through a pronunciation feature extraction model, obtaining the speech feature and the text feature respectively, and finally training the user's pronunciation, and at the same time comparing the differences between the speech feature and the text feature to judge whether the user's pronunciation is standard, so as to solve the problems that the existing speech pronunciation recognition technology has unreasonable judgment methods for whether the user's pronunciation is standard and non-characteristic data in the analysis process, resulting in the inability to effectively help users with pronunciation learning and correction.
[0005] To achieve the above object, in a first aspect, the present application provides a pronunciation recognition learning assistance method based on speech processing, including the following steps: Recognize the input speech of the user through speech recognition, and extract the input text; Construct a pronunciation feature extraction model, convert the speech to be analyzed into a spectrogram, and extract the pronunciation features in the spectrogram; Extract the pronunciation features of the input speech and the input text through the pronunciation feature extraction model, and obtain the speech features and the text features respectively; Conduct pronunciation training on the user, and at the same time compare the differences between the speech features and the text features to determine whether the user's pronunciation is standard.
[0006] Further, the above-mentioned speech recognition uses speech-to-text technology to convert the input speech into text, which is named the input text.
[0007] Further, constructing a pronunciation feature extraction model, converting the speech to be analyzed into a spectrogram, and extracting the pronunciation features in the spectrogram includes the following sub-steps: After preprocessing the speech to be analyzed, obtain the time-domain graph of the speech to be analyzed, frame it, and divide the time-domain graph into different frame data; Perform short-time Fourier transform on the frame data to convert it into the spectrum of the frame data, named the frame spectrum, and convert the input speech into a spectrogram based on the frame spectrum; Extract the pronunciation features of the speech to be analyzed based on the spectrogram.
[0008] Further, after preprocessing the speech to be analyzed, obtaining the time-domain graph of the speech to be analyzed, framing it, and dividing the time-domain graph into different frame data includes the following sub-steps: Construct a pronunciation feature extraction model, and name the speech input into the pronunciation feature extraction model as the speech to be analyzed; Filter the speech to be analyzed through a first-order high-pass filter, and then obtain the time-domain graph of the speech to be analyzed; Set the frame length, marked as n, obtain the speech duration of the speech to be analyzed, marked as T, calculate T / n, and mark the calculation result as m. The m is rounded up to an integer; Evenly divide the time-domain graph into m parts, each part is a frame data, and mark the frame data as F(n,m). F(n,m) represents the waveform in the time range of [n×m - n, n×m] in the time-domain graph.
[0009] Further, performing short-time Fourier transform on the frame data to convert it into the spectrum of the frame data, named the frame spectrum, and converting the input speech into a spectrogram based on the frame spectrum includes the following sub-steps: Perform a short-time Fourier transform on F(n,m) to convert it into the frame spectrum of F(n,m), denoted as S(n,m). The frame spectrum is specifically a two-dimensional coordinate system with frequency on the X-axis and amplitude value on the Y-axis; Obtain the minimum and maximum values of the amplitude, denoted as A1 and A2 respectively, to form a range [A1, A2], named the amplitude range; Evenly divide the amplitude range into 256 sub-ranges, named amplitude sub-ranges, and sort and number the amplitude sub-ranges in ascending order, denoted by the symbol P i where i is a positive integer and i is the serial number of P; There are color values from 0 to 255, where 0 represents black and 255 represents white. Set the color value of P i to C i where the C i is 256 - i; Name the waveform in S(n,m) as the wave to be converted, mark the length of the wave to be converted on the X-axis as L, construct a line with length L and width n, named the spectrogram line, input the spectrogram line into S(n,m), and align the spectrogram line with the head and tail of the wave to be converted; For any value of X in S(n,m), name the amplitude value corresponding to X as the target amplitude value, and find the C i corresponding to the P i to which the target amplitude value belongs, denoted as CD i Change the color value of the pixel point at X on the spectrogram line to CD i Analyze each value of X, and finally the obtained spectrogram line is a line segment with color changes; Extract the spectrogram line of S(n,m), rotate it counterclockwise by 90°, and obtain Q(n,m); Establish a two-dimensional coordinate system with m as the horizontal axis and frequency as the vertical axis, named the spectrogram, and input Q(n,m) into the spectrogram.
[0010] Furthermore, the extracting the pronunciation features of the speech to be analyzed based on the spectrogram includes the following sub-steps: Number the input text in order from left to right, denoted by the symbol W j where j is a positive integer and j is the serial number of W; When converting the input speech into input text, mark the pronunciation time of W j and find the frame data within the pronunciation time in the spectrogram, denoted as the word pronunciation data; The word pronunciation data is the pronunciation feature of W j ;
[0011] Furthermore, the extracting the pronunciation features of the input speech and the input text through the pronunciation feature extraction model includes the following sub-steps: Find the standard pronunciation of the input text, named standard speech; Extract the pronunciation features of the input speech through a pronunciation feature extraction model, named speech features; Extract the pronunciation features of the standard speech through a pronunciation feature extraction model, named text features.
[0012] Furthermore, the pronunciation training of the user, and at the same time comparing the differences between the speech features and the text features to determine whether the user's pronunciation is standard includes the following sub-steps: The user inputs or selects training content; Broadcast the standard speech of the training content to the user; The user learns the standard speech and uploads the input speech; Determine whether the input speech conforms to the standard speech.
[0013] Furthermore, the determination of whether the input speech conforms to the standard pronunciation includes the following sub-steps: Group the speech features and text features of the same W in the input speech and the standard speech into a set of character features; j For any set of character features, analyze them, and mark the speech features and text features in the set of character features as the first feature and the second feature respectively; Number the frame data in the first feature in order from left to right, represented by the symbol Z1 h h Number the frame data in the second feature in order from left to right, represented by the symbol Z2 h h where h is a positive integer and h is the serial number of Z1 and Z2; Extract the CD in Z1 h h and Z2 h h i in order from bottom to top, and form a one-dimensional matrix, marked as G1 h h h and G2 h h h h h h ; Calculate G1 h h -G2 h h h h h h h h h h h h h h h h h h h h h h h h h h h h h h h h Calculate , name the calculation result as pronunciation accuracy, where max() is the maximum operator; Find a first quantity of volunteers and require them to conduct a second number of pronunciation trainings and count the pronunciation accuracy of each training, named as regular accuracy. Find the minimum value of the regular accuracy and name it as the accuracy threshold; Calculate the pronunciation accuracy of the user's input speech, named as input accuracy. Compare the input accuracy with the accuracy threshold. If the input accuracy is less than the accuracy threshold, mark that the user's input speech does not conform to the standard speech, otherwise mark that the user's input speech conforms to the standard speech.
[0014] In a second aspect, the present application provides a pronunciation recognition learning assistance system based on speech processing, including a speech input module, a pronunciation feature extraction model construction module, a pronunciation feature extraction module, and a pronunciation standard judgment module; the speech input module, the pronunciation feature extraction model construction module, and the pronunciation feature extraction module are respectively connected to the pronunciation standard judgment module for data connection; The speech input module is used to recognize the user's input speech through speech recognition and extract the input text; The pronunciation feature extraction model construction module is used to construct a pronunciation feature extraction model, convert the speech to be analyzed into a spectrogram and extract the pronunciation features in the spectrogram; The pronunciation feature extraction module is used to extract the pronunciation features of the input speech and the input text through the pronunciation feature extraction model, and obtain the speech feature and the text feature respectively; The pronunciation standard judgment module is used to conduct pronunciation training for the user, and at the same time compare the differences between the speech feature and the text feature to judge whether the user's pronunciation is standard.
[0015] The beneficial effects of the present invention: The present invention recognizes the user's input speech through speech recognition, extracts the input text, preprocesses the speech to be analyzed, obtains the time-domain diagram of the speech to be analyzed, frames it, divides the time-domain diagram into different frame data, and then performs short-time Fourier transform on the frame data to convert it into the spectrum of the frame data, named as frame spectrum. Based on the frame spectrum, the input speech is converted into a spectrogram, and then the pronunciation features of the speech to be analyzed are extracted based on the spectrogram. The advantage is that the speech to be analyzed is converted into a spectrogram, and the spectrogram shows the data of the speech to be analyzed in three dimensions of time, frequency, and amplitude value. Then, effective pronunciation features are extracted from the spectrogram to judge whether the user's pronunciation is standard, improving the accuracy and effectiveness of pronunciation standard judgment; The present invention extracts the pronunciation features of the input speech and the input text through a pronunciation feature extraction model, respectively obtaining the speech features and the text features. Finally, pronunciation training is carried out for the user, and at the same time, the differences between the speech features and the text features are compared to determine whether the user's pronunciation is standard. The advantage is that by extracting the speech features and the text features and analyzing the differences between them, if the deviation is within a reasonable range, it is determined that the user's pronunciation is standard; if the deviation is too large, it means that the user's pronunciation is not standard, improving the accuracy and rationality of pronunciation standard judgment. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic block diagram of the system of the present invention; Figure 2 is a time domain diagram of the present invention; Figure 3 is a schematic diagram of the frame data F(25,1) of the present invention; Figure 4 is a schematic diagram of the frame spectrum of the present invention; Figure 5 is a schematic diagram of putting the spectrogram of the present invention into the frame spectrum; Figure 6 is a schematic diagram of the spectrogram after processing of the present invention; Figure 7 is a schematic diagram of the spectrogram of the present invention; Figure 8 is a step flow chart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0018] Embodiment 1, please refer to Figure 1 As shown, the present application provides a pronunciation recognition learning assistance system based on speech processing, including a speech input module, a pronunciation feature extraction model construction module, a pronunciation feature extraction module, and a pronunciation standard judgment module; the speech input module, the pronunciation feature extraction model construction module, and the pronunciation feature extraction module are respectively connected to the pronunciation standard judgment module for data connection; The speech input module is used to recognize the input speech of the user through speech recognition and extract the input text; through speech recognition, the speech-to-text technology is adopted to convert the input speech into text, named the input text; In practical applications, existing speech recognition technologies are adopted to extract the input text.
[0019] The pronunciation feature extraction model construction module is used to construct a pronunciation feature extraction model, convert the speech to be analyzed into a spectrogram, and extract the pronunciation features in the spectrogram. The pronunciation feature extraction model construction module includes a framing unit, a spectrogram conversion unit, and a feature extraction unit; The framing unit is used to preprocess the speech to be analyzed, obtain the time-domain diagram of the speech to be analyzed, frame it, and divide the time-domain diagram into different frame data; The framing unit is configured with a framing strategy, and the framing strategy includes: Construct a pronunciation feature extraction model, and name the speech input into the pronunciation feature extraction model as the speech to be analyzed; Please refer to Figure 2 As shown, after filtering the speech to be analyzed through a first-order high-pass filter, obtain the time-domain diagram of the speech to be analyzed; Set the frame length, marked as n, obtain the speech duration of the speech to be analyzed, marked as T, calculate T / n, and mark the calculation result as m. m is rounded up to an integer; Evenly divide the time-domain diagram into m parts, each part is a frame data, mark the frame data as F(n,m), and F(n,m) represents the waveform in the time range [n×m - n, n×m] in the time-domain diagram; In practical applications, after filtering the speech to be analyzed through a first-order high-pass filter, the obtained time-domain diagram is as Figure 2 shown, where the X-axis is time, the unit is seconds, the Y-axis is the amplitude value. According to the existing frame length standard, the frame length is usually 10ms to 40ms. In this embodiment, the median value is taken, and the frame length is set to 25ms, that is, n is 25ms, and the speech duration T is 0.535s, that is, 535ms. 535 / 25 = 22 frame data are obtained, that is, F(25,1) to F(25,22), and m is rounded up to an integer; The spectrogram conversion unit is used to perform a short-time Fourier transform on the frame data to convert it into the spectrum of the frame data, named the frame spectrum, and convert the input speech into a spectrogram based on the frame spectrum; The spectrogram conversion unit is configured with a spectrogram conversion strategy, and the spectrogram conversion strategy includes: Please refer to Figures 3 to 4 As shown, perform a short-time Fourier transform on F(n,m) to convert it into the frame spectrum of F(n,m), marked as S(n,m). The frame spectrum is specifically a two-dimensional coordinate system with the X-axis as the frequency and the Y-axis as the amplitude value; Obtain the minimum and maximum values of the amplitude values, marked as A1 and A2 respectively, form the range [A1,A2], and name it the amplitude range; Evenly divide the amplitude range into 256 sub-ranges, name them amplitude sub-ranges, sort and number the amplitude sub-ranges in ascending order, and use the symbol P iIt is represented that, where i is a positive integer and i is the serial number of P; There are color values from 0 to 255, where 0 represents black and 255 represents white. Set the color value of P i to C i , and C i is 256 - i; In practical applications, taking F(25,1) as an example, F(25,1) is as Figure 3 shown and is converted to S(25,1) through short-time Fourier transform as Figure 4 shown. In the spectrogram, the amplitude range is [-1,1]. Divide the amplitude range into 256 amplitude sub-ranges. Calculate [1 - (-1)] / 256 = 0.0078125. Thus, the span of each amplitude sub-range is 0.0078125. Number them to get P i , where 1 ≤ i ≤ 256. For example, when the amplitude value is within [-1, -0.9921875], it belongs to P1, and the color value C1 of P1 is 256 - 1 = 255; Please refer to Figure 5 shown. Name the waveform in S(n,m) as the wave to be converted. Mark the length of the wave to be converted on the X-axis as L. Construct a line with length L and width n, named the spectrogram line, and input the spectrogram line into S(n,m), and the spectrogram line is aligned with the head and tail of the wave to be converted; For any value of X in S(n,m), name the amplitude value corresponding to X as the target amplitude value. Search for the P i corresponding to it i , mark it as CD i , and change the color value of the pixel point at X on the spectrogram line to CD i . Analyze each value of X. Finally, the obtained spectrogram line is a line segment with color changes; Please refer to Figure 6 shown. Extract the spectrogram line of S(n,m) and rotate it counterclockwise by 90° to get Q(n,m); Please refer to Figure 7 shown. Establish a two-dimensional coordinate system with m as the horizontal axis and frequency as the vertical axis, named the spectrogram, and input Q(n,m) into the spectrogram; In practical applications, the length of the wave to be converted on the X-axis is the speech duration T. We get L = T = 0.535. Construct a spectrogram line with length 0.535 and width n = 25ms and input it into S(25,1). The spectrogram line in S(25,1) is as Figure 5 shown. From Figure 5 the enlarged part of some spectrogram lines, it can be seen that the spectrogram line is composed of several pixel grids. According to the X-axis coordinate of each pixel grid, query the amplitude value of the wave to be converted at this X value as the target amplitude value, and then search for the P iThe corresponding C i , and finally obtain CD i and change the color value of the corresponding pixel grid to CD i , and the finally obtained spectrogram is a line segment with color changes. Rotate it counterclockwise by 90° to obtain the processed spectrogram Q(25,1), as Figure 6 shown. Convert each frame data into a spectrogram and combine them in ascending order of m. Finally, obtain the spectrogram as Figure 7 shown; The feature extraction unit is used to extract the pronunciation features of the speech to be analyzed based on the spectrogram; The feature extraction unit is configured with a feature extraction strategy, and the feature extraction strategy includes: Number the input text in order from left to right, represented by the symbol W j , where j is a positive integer and j is the serial number of W; When converting the input speech into input text, mark the pronunciation time of W j , and find the frame data within the pronunciation time in the spectrogram and mark it as the word pronunciation data; The word pronunciation data is the pronunciation feature of W j ; In practical applications, for example, the input text is "afraid of the sun", and W1 and W2 are extracted. The pronunciation time can be directly extracted based on speech recognition technology. Among them, the pronunciation time of W1 is from 0s to 0.25s, and the pronunciation time of W2 is from 0.25s to 0.535s. Mark the frame data within 0s to 0.25s as the word pronunciation data of W1, and at the same time mark the frame data within 0.25s to 0.535s as the word pronunciation data of W2.
[0020] The pronunciation feature extraction module is used to extract the pronunciation features of the input speech and the input text through the pronunciation feature extraction model, and obtain the speech feature and the text feature respectively; The pronunciation feature extraction module is configured with a pronunciation feature extraction strategy, and the pronunciation feature extraction strategy includes: Find the standard pronunciation of the input text and name it the standard speech; Extract the pronunciation feature of the input speech through the pronunciation feature extraction model and name it the speech feature; Extract the pronunciation feature of the standard speech through the pronunciation feature extraction model and name it the text feature; In practical applications, the standard speech is provided by the word library in the Internet, and the speech feature and the text feature are obtained through extraction.
[0021] The pronunciation standard judgment module is used to conduct pronunciation training for users, and at the same time compare the differences between speech features and text features to judge whether the users' pronunciation is standard; the pronunciation standard judgment module includes a pronunciation training unit and a standard judgment unit; The pronunciation training unit is configured with a pronunciation training strategy, and the pronunciation training strategy includes: The user inputs or selects training content; Broadcast the standard speech of the training content for the user; The user learns the standard speech and uploads the input speech; In practical applications, the user can choose to input the training content by themselves or select the built-in training content from the model. After the user makes a selection, the standard speech of the training content is broadcast for the user. The user learns and reads according to the standard speech, and then uploads the input speech during reading; The standard judgment unit is used to judge whether the input speech conforms to the standard speech; The standard judgment unit is configured with a standard judgment strategy, and the standard judgment strategy includes: Group the speech features and text features of the same W j in the input speech and the standard speech into a group of character features; Analyze any group of character features, and mark the speech features and text features in the group of character features as the first feature and the second feature respectively; Number the frame data in the first feature in order from left to right, represented by the symbol Z1 h Number the frame data in the second feature in order from left to right, represented by the symbol Z2 h where h is a positive integer and h is the serial number of Z1 and Z2; Extract the CD h from Z1 h and Z2 i in order from bottom to top, and form a one-dimensional matrix, marked as G1 h and G2 h respectively; Calculate G1 h -G2 h , mark the calculation result as G3 h , and take the absolute value of all the values in G3 h to get G4 h . Number the values in G4 h in order from left to right, represented by the symbol K1(h,t). Number the values in G2 h in order from left to right, represented by the symbol K2(h,t). Where t is a positive integer and (h,t) is the serial number of K1 and K2; In practical applications, W exists in both the input speech and the standard speech j , group the two W1s of the input speech and the standard speech into a group of character features, and group the two W2s into a group of character features. In this embodiment, the group of character features corresponding to W1 is taken as an example, which includes a first feature and a second feature. W1 is within the time period from 0s to 0.25s, which includes 10 frame data, that is, the first feature and the second feature each include 10 frame data, and Z1 is marked h and Z2 h , 1 ≤ h ≤ 10, based on the CD of the spectrogram i , extract in the order from bottom to top and construct a one-dimensional matrix. For example, in Z11, the color values of the spectrogram are 213, 189, 177, 205, 192, and 153 from bottom to top in sequence. Thus, G11 is constructed as [213 189 177 205 192 153]. When calculating G1 h -G2 h , it is necessary to ensure that h is equal. For example, G21 is [209 175 163 213 205 164], and G31 is calculated as [4 14 14 -8 -13 -11]. Take the absolute value to get G41 as [4 14 14 8 13 11]. The numbers K1(1,1) to K1(1,6) are 4, 14, 14, 8, 13, and 11 respectively, and at the same time, the numbers K2(1,1) to K2(1,6) are 209, 175, 163, 213, 205, and 164 respectively; Calculate , name the calculation result as pronunciation accuracy, where max() is the maximum operator; Find a first number of volunteers. The volunteers need to have a Putonghua proficiency certificate. Require the volunteers to conduct a second round of pronunciation training and count the pronunciation accuracy of each training, name it the conventional accuracy, and find the minimum value of the conventional accuracy, name it the accuracy threshold; Calculate the pronunciation accuracy of the user's input speech, name it the input accuracy, compare the input accuracy with the accuracy threshold. If the input accuracy is less than the accuracy threshold, then mark that the user's input speech does not conform to the standard speech, otherwise mark that the user's input speech conforms to the standard speech; In practical applications, each frame of data in the character feature group is analyzed, and finally K1(h,t) and K2(h,t) are obtained, where 1 ≤ h ≤ 10, 1 ≤ t ≤ 6. The value range of t is the result exemplified in this embodiment, rather than actual data. When actually calculating, the value of t is relatively large and it is inconvenient to specifically show in this embodiment. Therefore, only part of the data is used as an example for illustration, and then the pronunciation accuracy is calculated through formulas. The pronunciation accuracy reflects the similarity between the user's pronunciation and the standard pronunciation. However, when actually pronouncing, the user cannot achieve 100% identity with the standard pronunciation and there will always be a certain deviation. It is necessary to determine the maximum allowable deviation as the judgment basis. Therefore, a first number of volunteers are sought, and the volunteers need to have a Putonghua proficiency certificate to analyze the deviation range between the user and the standard pronunciation under normal circumstances. There are no requirements for the setting of the first number and the second number, and the developer can set them by himself. The minimum value of the conventional accuracy is the maximum deviation allowed for the deviation between the user's pronunciation and the standard voice under normal circumstances. Therefore, it is set as the accuracy threshold to judge whether the user's pronunciation is standard, so as to help the user achieve the purpose of learning the standard pronunciation of Putonghua.
[0022] Embodiment 2, please refer to Figure 8 As shown, the present application provides a pronunciation recognition learning assistance method based on speech processing, including the following steps: Step S1, recognize the input speech of the user through speech recognition and extract the input text; through speech recognition, use the speech-to-text technology to convert the input speech into text, named input text; Step S2, construct a pronunciation feature extraction model, convert the speech to be analyzed into a spectrogram and extract the pronunciation features in the spectrogram; Step S2 includes the following sub-steps: Step S201, after preprocessing the speech to be analyzed, obtain the time-domain diagram of the speech to be analyzed, frame it, and divide the time-domain diagram into different frame data; Step S201 includes the following sub-steps: Step S201.1, construct a pronunciation feature extraction model, and name the speech input into the pronunciation feature extraction model as the speech to be analyzed; Step S201.2, filter the speech to be analyzed through a first-order high-pass filter, and obtain the time-domain diagram of the speech to be analyzed; Step S201.3, set the frame length, marked as n, obtain the speech duration of the speech to be analyzed, marked as T, calculate T / n, and mark the calculation result as m. m is rounded up to an integer; Step S201.4, evenly divide the time-domain diagram into m parts, each part is a frame of data, and mark the frame data as F(n,m). F(n,m) represents the waveform in the time range of [n×m - n, n×m] in the time-domain diagram; Step S202: Perform a short-time Fourier transform on the frame data to convert it into the spectrum of the frame data, named the frame spectrum, and convert the input speech into a spectrogram based on the frame spectrum; Step S202 includes the following sub-steps: Step S202.1: Perform a short-time Fourier transform on F(n,m) to convert it into the frame spectrum of F(n,m), marked as S(n,m). The frame spectrum is specifically a two-dimensional coordinate system with frequency on the X-axis and amplitude value on the Y-axis; Step S202.2: Obtain the minimum and maximum values of the amplitude values, marked as A1 and A2 respectively, to form a range [A1, A2], named the amplitude range; Step S202.3: Evenly divide the amplitude range into 256 sub-ranges, named amplitude sub-ranges, sort and number the amplitude sub-ranges in ascending order, and represent them by the symbol P i where i is a positive integer and i is the serial number of P; Step S202.4: There are color values from 0 to 255, where 0 represents black and 255 represents white. Set the color value of P i to C i where C i is 256 - i; Step S202.5: Name the waveform in S(n,m) as the wave to be converted, mark the length of the wave to be converted on the X-axis as L, construct a line with length L and width n, named the spectrogram line, enter the spectrogram line into S(n,m), and align the spectrogram line with the head and tail of the wave to be converted; Step S202.6: For any value of X in S(n,m), name the amplitude value corresponding to X as the target amplitude value, find the C i corresponding to the P i to which the target amplitude value belongs, marked as CD i , and change the color value of the pixel point on the spectrogram line at X to CD i . Analyze each value of X, and the finally obtained spectrogram line is a line segment with color changes; Step S202.7: Extract the spectrogram line of S(n,m), rotate it counterclockwise by 90°, and obtain Q(n,m); Step S202.8: Establish a two-dimensional coordinate system with m as the horizontal axis and frequency as the vertical axis, named the spectrogram, and enter Q(n,m) into the spectrogram; Step S203: Extract the pronunciation features of the speech to be analyzed based on the spectrogram; Step S203 includes the following sub-steps: Step S203.1: Number the input text in order from left to right, represented by the symbol W jIt is represented that, where j is a positive integer and j is the serial number of W; Step S203.2: When converting the input speech into input text, mark the pronunciation time of W, and search for the frame data within the pronunciation time in the spectrogram, and mark it as the word pronunciation data; j Step S203.3: The word pronunciation data is the pronunciation feature of W; j Step S3: Extract the pronunciation features of the input speech and the input text through the pronunciation feature extraction model, and obtain the speech feature and the text feature respectively; Step S3 includes the following sub-steps: Step S301: Search for the standard pronunciation of the input text and name it the standard speech; Step S302: Extract the pronunciation feature of the input speech through the pronunciation feature extraction model and name it the speech feature; Step S303: Extract the pronunciation feature of the standard speech through the pronunciation feature extraction model and name it the text feature; Step S4: Conduct pronunciation training for the user, and at the same time compare the differences between the speech feature and the text feature to judge whether the user's pronunciation is standard; Step S4 includes the following sub-steps: Step S401: The user inputs or selects the training content; Step S402: Broadcast the standard speech of the training content for the user; Step S403: The user learns the standard speech and uploads the input speech; Step S404: Judge whether the input speech conforms to the standard speech; Step S404 includes the following sub-steps: Step S404.1: Group the speech feature and the text feature of the same W in the input speech and the standard speech into a group of word features; j Step S404.2: Analyze any group of word features, and mark the speech feature and the text feature in the group of word features as the first feature and the second feature respectively; Step S404.3: Number the frame data in the first feature in the order from left to right, and represent it by the symbol Z1, number the frame data in the second feature in the order from left to right, and represent it by the symbol Z2, where h is a positive integer and h is the serial number of Z1 and Z2; h h Step S404.4: Extract the CD in Z1 and Z2 in the order from bottom to top, and form a one-dimensional matrix, which are marked as G1 and G2 respectively; h h i h h ; Step S404.5, calculate G1 h - G2 h , mark the calculation result as G3 h , and take the absolute value of all the numerical values in G3 h to obtain G4 h , number the numerical values in G4 h in the order from left to right, which is represented by the symbol K1(h, t), number the numerical values in G2 h in the order from left to right, which is represented by the symbol K2(h, t), where t is a positive integer and (h, t) is the serial number of K1 and K2; Step S404.6, calculate , name the calculation result as pronunciation accuracy, where max() is the maximum operator; Step S404.7, find the first quantity of volunteers. The volunteers need to have a Putonghua proficiency certificate. Require the volunteers to conduct the second round of pronunciation training and count the pronunciation accuracy of each training, name it as regular accuracy, and find the minimum value of the regular accuracy, name it as the accuracy threshold; Step S404.8, calculate the pronunciation accuracy of the user's input speech, name it as input accuracy, compare the input accuracy with the accuracy threshold. If the input accuracy is less than the accuracy threshold, then mark that the user's input speech does not conform to the standard speech, otherwise mark that the user's input speech conforms to the standard speech.
[0023] Embodiment 3, the present application provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete the communication with each other through the communication bus. The memory stores computer-readable instructions, and the processor can call the instructions in the memory. When the computer-readable instructions are executed by the processor, the steps in the pronunciation recognition learning assistance method based on speech processing are run to implement the following functions: recognize the user's input speech through speech recognition, and extract the input text; construct a pronunciation feature extraction model; extract the pronunciation features of the input speech and the input text through the pronunciation feature extraction model to obtain the speech feature and the text feature respectively; conduct pronunciation training on the user, and at the same time compare the differences between the speech feature and the text feature to judge whether the user's pronunciation is standard.
[0024] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0025] Embodiment 4, this application also provides a computer-readable storage medium. This application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, it runs the steps in the pronunciation recognition learning assistance method based on speech processing as described above to achieve the following functions: recognize the input speech of the user through speech recognition and extract the input text; construct a pronunciation feature extraction model; extract the pronunciation features of the input speech and the input text through the pronunciation feature extraction model to obtain the speech feature and the text feature respectively; conduct pronunciation training on the user, and at the same time compare the differences between the speech feature and the text feature to determine whether the user's pronunciation is standard.
[0026] Through the description of the above embodiments, the embodiments of the present invention can be provided as a method, a system, or a computer program product. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical discs, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.
[0027] In the embodiments provided by this application, it should be understood that the disclosed system or method can be implemented in other ways. The embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division, and there can be other division methods in actual implementation. Another example is that multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some communication interfaces. The indirect coupling or communication connection of systems, modules, and units can be electrical, mechanical, or other forms.
[0028] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A pronunciation recognition learning assistance method based on speech processing, characterized in that It includes the following steps: Recognize the input speech of the user through speech recognition and extract the input text; Construct a pronunciation feature extraction model, convert the speech to be analyzed into a spectrogram and extract the pronunciation features in the spectrogram; Extract the pronunciation features of the input speech and the input text through the pronunciation feature extraction model, and obtain the speech features and text features respectively; Conduct pronunciation training for the user, and at the same time compare the differences between the speech features and the text features to determine whether the user's pronunciation is standard.
2. The pronunciation recognition learning assistance method based on speech processing according to claim 1, wherein The above-mentioned speech recognition adopts speech-to-text technology to convert the input speech into text, which is named input text.
3. The pronunciation recognition learning assistance method based on speech processing according to claim 2, wherein Construct a pronunciation feature extraction model, and the steps of converting the speech to be analyzed into a spectrogram and extracting the pronunciation features in the spectrogram include the following sub-steps: After preprocessing the speech to be analyzed, obtain the time-domain diagram of the speech to be analyzed, frame it, and divide the time-domain diagram into different frame data; Perform short-time Fourier transform on the frame data to convert it into the spectrum of the frame data, named frame spectrum, and convert the input speech into a spectrogram based on the frame spectrum; Extract the pronunciation features of the speech to be analyzed based on the spectrogram.
4. The pronunciation recognition learning assistance method based on speech processing according to claim 3, characterized in that The steps of obtaining the time-domain diagram of the speech to be analyzed, framing it, and dividing the time-domain diagram into different frame data after preprocessing the speech to be analyzed include the following sub-steps: Construct a pronunciation feature extraction model, and name the speech input into the pronunciation feature extraction model as the speech to be analyzed; Filter the speech to be analyzed through a first-order high-pass filter to obtain the time-domain diagram of the speech to be analyzed; Set the frame length, marked as n, obtain the speech duration of the speech to be analyzed, marked as T, calculate T / n, and mark the calculation result as m. The m is rounded up to an integer; Evenly divide the time-domain diagram into m parts, each part is a frame data, and mark the frame data as F(n,m). F(n,m) represents the waveform in the time range [n×m - n, n×m] in the time-domain diagram.
5. The pronunciation recognition learning assistance method based on speech processing according to claim 4, characterized in that, The steps of performing short-time Fourier transform on the frame data to convert it into the spectrum of the frame data, named frame spectrum, and converting the input speech into a spectrogram based on the frame spectrum include the following sub-steps: Perform short-time Fourier transform on F(n,m) to convert it into the frame spectrum of F(n,m), marked as S(n,m). The frame spectrum is specifically a two-dimensional coordinate system with the X-axis as the frequency and the Y-axis as the amplitude value; Obtain the minimum value and the maximum value of the amplitude value, marked as A1 and A2 respectively, and form a range [A1, A2], named amplitude range; The amplitude range is evenly divided into 256 sub-ranges, named amplitude sub-ranges, and the amplitude sub-ranges are sorted and numbered in ascending order. Through the symbol P i It is represented that i is a positive integer and i is the serial number of P; There are color values from 0 to 255, where 0 represents black and 255 represents white. Set the color value of P i to C i , and the C i is 256 - i; Name the waveform in S(n,m) as the wave to be converted, mark the length of the wave to be converted on the X-axis as L, construct a line with a length of L and a width of n, named spectrogram line, input the spectrogram line into S(n,m), and align the spectrogram line with the wave to be converted at the head and tail; For any value of X in S(n,m), name the amplitude value corresponding to X as the target amplitude value, and search for the P to which the target amplitude value belongs i The corresponding C i , and mark it as CD i , change the color value of the pixel point at X on the spectrogram line to CD i , analyze each value of X, and finally the obtained spectrogram line is a line segment with color changes; Extract the spectrogram line of S(n,m) and rotate it counterclockwise by 90°, obtaining Q(n,m); Establish a two-dimensional coordinate system with m as the horizontal axis and frequency as the vertical axis, named spectrogram, and input Q(n,m) into the spectrogram.
6. The pronunciation recognition learning assistance method based on speech processing according to claim 5, characterized in that The steps of extracting the pronunciation features of the speech to be analyzed based on the spectrogram include the following sub-steps: Number the input text in order from left to right, indicated by the symbol W j where j is a positive integer and j is the serial number of W; When converting the input speech into input text, mark the pronunciation time of W j In the spectrogram, search for the frame data within the pronunciation time and mark it as the word pronunciation data; The pronunciation data of the character is W j 's pronunciation feature.
7. The pronunciation recognition learning assistance method based on speech processing according to claim 6, characterized in that, The steps of extracting the pronunciation features of the input speech and the input text through the pronunciation feature extraction model include the following sub-steps: Find the standard pronunciation of the input text, named the standard speech; Extract the pronunciation features of the input speech through a pronunciation feature extraction model, named the speech features; Extract the pronunciation features of the standard speech through a pronunciation feature extraction model, named the text features.
8. The pronunciation recognition learning assistance method based on speech processing according to claim 7, characterized in that, The step of training the user's pronunciation and at the same time comparing the differences between the speech features and the text features to determine whether the user's pronunciation is standard includes the following sub-steps: The user inputs or selects the training content; Broadcast the standard speech of the training content to the user; The user learns the standard speech and uploads the input speech; Judge whether the input speech conforms to the standard speech.
9. The pronunciation recognition learning assistance method based on speech processing according to claim 8, wherein The step of judging whether the input speech conforms to the standard pronunciation includes the following sub-steps: Induce the speech features and text features of the same W in the input speech and the standard speech into a group of character feature groups; j Analyze any character feature group, and mark the speech feature and the text feature in the character feature group as the first feature and the second feature respectively; Number the frame data in the first feature in the order from left to right, represented by the symbol Z1 h Number the frame data in the second feature in the order from left to right, represented by the symbol Z2 h where h is a positive integer and h is the serial number of Z1 and Z2; Extract Z1 in the order from bottom to top h and Z2 h from the CD i , and form a one-dimensional matrix, labeled as G1 h and G2 h ; Calculate G1 h -G2 h , mark the calculation result as G3 h , and for all the numerical values in G3 h , take the absolute value to obtain G4 h , number the numerical values in G4 h in the order from left to right, denoted by the symbol K1(h, t), and number the numerical values in G2 h in the order from left to right, denoted by the symbol K2(h, t), where t is a positive integer and (h, t) is the serial number of K1 and K2; Calculation , name the calculation result as pronunciation accuracy, where max() is the maximum operator; Find a first number of volunteers, require the volunteers to conduct a second round of pronunciation training and count the pronunciation accuracy of each training, named the conventional accuracy, and find the minimum value of the conventional accuracy, named the accuracy threshold; Calculate the pronunciation accuracy of the user's input speech, named the input accuracy, compare the input accuracy with the accuracy threshold. If the input accuracy is less than the accuracy threshold, mark that the user's input speech does not conform to the standard speech, otherwise mark that the user's input speech conforms to the standard speech.
10. A pronunciation recognition learning assistance system based on speech processing, which is used to implement the pronunciation recognition learning assistance method based on speech processing described in any one of claims 1-9, and is characterized in that, It includes a speech input module, a pronunciation feature extraction model construction module, a pronunciation feature extraction module, and a pronunciation standard judgment module; the speech input module, the pronunciation feature extraction model construction module, and the pronunciation feature extraction module are respectively connected to the pronunciation standard judgment module for data connection; The speech input module is used to recognize the user's input speech through speech recognition and extract the input text; The pronunciation feature extraction model construction module is used to construct a pronunciation feature extraction model, convert the speech to be analyzed into a spectrogram and extract the pronunciation features in the spectrogram; The pronunciation feature extraction module is used to extract the pronunciation features of the input speech and the input text through the pronunciation feature extraction model, and obtain the speech features and the text features respectively; The pronunciation standard judgment module is used to train the user's pronunciation and at the same time compare the differences between the speech features and the text features to judge whether the user's pronunciation is standard.
Citation Information
Patent Citations
Mispronounce correcting method and system through voice recognition
CN109461436A
Voice evaluation scoring method and device, electronic equipment and storage medium
CN112802456A
Intelligent foreign language spoken language training method
CN119028325A
Voice recognition / synthesis systems based on standardpronunciation analysis methodology and methods therefor
KR1020010106696A
Speaking practice system with reliable pronunciation evaluation
US20240347054A1