A multi-dimensional oral ability identification and quantification method and system
Through a multi-dimensional oral ability identification and quantification method, combined with audio preprocessing, fluency analysis and vocabulary analysis, the problem of incomplete oral ability evaluation in the existing technology is solved, and all-round evaluation and differentiated identification of oral ability are achieved.
Patent Information
- Application Number
- CN202510733487.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Existing oral ability recognition methods lack effective means to judge the fluency of oral expression and the rationality of word structure, resulting in the inability to perform differentiated recognition and the inability to make a comprehensive assessment of the user's oral ability.
Through a multi-dimensional oral ability identification and quantification method, including audio preprocessing, fluency analysis and vocabulary analysis, the fluency and vocabulary structure scores of the spoken audio are obtained, and a comprehensive analysis is performed in combination with the basic clarity score to obtain the oral ability score.
It achieves a comprehensive evaluation of oral ability and can perform differentiated recognition when faced with audio with the same semantics, thereby improving the effectiveness of oral ability evaluation.
Smart Images

Figure CN120431969B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spoken language ability recognition, and in particular to a multi-dimensional spoken language ability recognition quantification method and system. Background Art
[0002] Oral ability refers to a person's ability to communicate through listening and speaking, and is the external manifestation of language ability; it includes two aspects: receiving language information and transmitting language information; receiving language information involves attention, sound recognition, comprehension and judgment, etc.; transmitting language information includes the ability to use voice, language organization and conversion, and language control; oral ability identification refers to measuring an individual's ability to understand and use spoken language for communication, mainly including two aspects: speech recognition rate and oral communication ability.
[0003] Existing methods for identifying oral ability usually identify oral ability based on the recognition of spoken vocabulary and semantics. For example, by establishing a trained language model, the words in the user's spoken language are recognized, and then the recognition accuracy of the vocabulary is improved through the context of each word, thereby achieving accurate recognition of the vocabulary in the spoken language and then identifying the oral ability. Although this improved method can improve the efficiency of spoken recognition, it can only judge the oral ability through vocabulary recognition. In terms of the fluency of oral expression and the rationality of the word structure in spoken ability, there is a lack of effective evaluation methods. This will lead to a relatively one-sided recognition of oral ability, and thus it is impossible to differentiate the oral ability under the same semantics, and thus it is impossible to accurately and comprehensively judge the user's oral ability. For example, in the patent application with publication number CN108829894A, a method and apparatus for spoken word recognition and semantic recognition is disclosed. This solution extracts contextual features of each word in the recognition sentence and performs contextual feature recognition to determine whether each word is a spoken word, thereby improving the efficiency and accuracy of spoken word recognition. Other improvements for oral ability recognition are usually improvements in phonetic characteristics, which still cannot solve the problem of lack of evaluation methods for the fluency of oral expression and the rationality of word structure in oral ability, resulting in a relatively one-sided recognition of oral ability, making it impossible to differentiate the oral ability under the same semantics, and thus unable to accurately and comprehensively evaluate the user's oral ability. In view of this, it is necessary to improve the existing oral ability recognition method. Summary of the Invention
[0004] The present invention aims to solve, at least to a certain extent, one of the technical problems in the prior art by proposing a multi-dimensional oral ability identification and quantification method and system to solve the problem that the existing oral ability identification methods lack an evaluation method for the fluency of oral expression and the rationality of word structure in oral ability, resulting in a relatively one-sided identification of oral ability, making it impossible to differentiate oral abilities under the same semantics, and thus unable to accurately and comprehensively evaluate the user's oral ability.
[0005] To achieve the above objectives, in a first aspect, the present application provides a multi-dimensional oral ability identification and quantification method, comprising the following steps:
[0006] Acquire audio corresponding to spoken language and record it as spoken audio; perform recognition preprocessing on the spoken audio; obtain the language type and basic clarity score of the spoken audio based on the processing results of the recognition preprocessing;
[0007] Analyze the fluency and vocabulary structure of the spoken audio using fluency analysis and vocabulary analysis based on the language type, and obtain a fluency score and vocabulary structure score for the spoken audio based on the analysis results;
[0008] The fluency score and vocabulary score are comprehensively analyzed based on the basic clarity score, and the oral ability score is obtained based on the analysis results.
[0009] Furthermore, recognition preprocessing includes:
[0010] Use AI to recognize spoken audio, obtain the language corresponding to the spoken audio based on the recognition result, and record it as the language type of the spoken audio; convert the spoken audio into text based on the language type of the spoken audio and record it as audio text;
[0011] The audio text is segmented and the words obtained are recorded as audio words YC1 to audio words YC based on the order of recognition in the spoken audio. n .
[0012] Furthermore, the recognition preprocessing also includes:
[0013] For any audio word YC c , the audio words YC in the spoken audio c The audio segments recognized by AI are recorded as audio sub-segments, where c is a positive integer less than or equal to n and greater than or equal to 1;
[0014] Obtain multiple audio recognition software, and respectively identify the audio sub-segments and convert the audio sub-segments into text; obtain the recognition results of all audio recognition software, and when the language recognized in the recognition results of all audio recognition software is the language type of spoken audio, convert the audio word YC cThe language clarity score is recorded as 1; when the language recognized in the recognition result of any audio recognition software is not the language type of spoken audio, the audio word YC c The intelligibility score was recorded as 0.
[0015] Furthermore, the recognition preprocessing also includes:
[0016] When the audio recognition software converts the audio sub-segment into text and the audio word YC c If the audio words YC are the same, c The vocabulary clarity score is recorded as 1; when the result of any audio recognition software converting the audio sub-segment into text is consistent with the audio word YC c If they are different, the audio words YC c The vocabulary clarity score of the audio word YC is recorded as 0, where the initial values of the language clarity score and vocabulary clarity score of each audio word YC are both 0;
[0017] Use the basic clarity algorithm to obtain the basic clarity score of the spoken audio. The basic clarity algorithm is: , where F1 is the basic clarity score, f i is the language clarity score corresponding to the i-th audio word YC in all audio words YC, g j is the vocabulary clarity score corresponding to the j-th audio word YC in all audio words YC.
[0018] Furthermore, the fluency analysis method includes:
[0019] The audio duration corresponding to the spoken audio is recorded as T, and a plane rectangular coordinate system is established, recorded as the fluency analysis coordinate system, where the unit of the X-axis of the fluency analysis coordinate system is time, and the unit of the Y-axis is decibel; based on the decibel of the sound corresponding to the speech in the spoken audio, a corresponding curve is drawn in the interval from X=0 to X=T in the fluency analysis coordinate system, and recorded as the fluency analysis curve;
[0020] Based on the audio sub-segment corresponding to each audio word YC, the curves corresponding to all audio sub-segments are marked in the fluency analysis curve and recorded as the fluency sub-curve of the audio sub-segment;
[0021] For any smooth sub-curve α, the horizontal coordinates of the rightmost point and the leftmost point of the smooth sub-curve α are recorded as the right point duration and the left point duration of the smooth sub-curve α respectively; when there is a smooth sub-curve on the left side of the smooth sub-curve α, the difference between the left point duration of the smooth sub-curve α and the right point duration of the smooth sub-curve closest to the left side of the smooth sub-curve α is recorded as the left interval value of the smooth sub-curve α; when there is no smooth sub-curve on the left side of the smooth sub-curve α, the left interval value of the smooth sub-curve α is recorded as 0.
[0022] Furthermore, the fluency analysis method also includes:
[0023] When there is a smooth sub-curve on the right side of the smooth sub-curve α, the difference between the duration of the right point of the smooth sub-curve α and the duration of the left point of the smooth sub-curve closest to the right side of the smooth sub-curve α is recorded as the right pause value of the smooth sub-curve α; when there is no smooth sub-curve on the right side of the smooth sub-curve α, the right pause value of the smooth sub-curve α is recorded as 0;
[0024] The average of the left and right interval values of the fluency sub-curve α is recorded as the fluency interval value of the audio word YC corresponding to the fluency sub-curve α; the interval fluency values of all audio words YC are obtained, and the fluency score of the spoken audio is obtained based on the interval difference algorithm. The interval difference algorithm is: , where F2 is the fluency score, R u is the intermittent fluency value of the u-th audio word YC in all audio words YC, R sq is the average of all intermittent flow values.
[0025] Furthermore, the lexical analysis method includes:
[0026] Establishing an audio language structure library and an audio associated word library, wherein the audio language structure library and the audio associated word library are used to store the formats corresponding to the languages; based on the language type of the spoken audio, using big data to obtain the language structure and associated words corresponding to the language type of the spoken audio, and storing them in the audio language structure library and the audio associated word library respectively;
[0027] The language structure of the audio text is obtained based on AI; the language structure of the audio text is matched with the language structures in the audio language structure library in turn. When any language structure in the audio language structure library is equal to the language structure of the audio text, the language structure score of the spoken audio is recorded as 1; when all language structures in the audio language structure library are not equal to the language structure of the audio text, the language structure score of the spoken audio is recorded as 0.
[0028] Furthermore, the lexical analysis method also includes:
[0029] When any associated word in the audio associated vocabulary exists in the audio text, the language association score of the spoken audio is recorded as 1; when any associated word in the audio associated vocabulary does not exist in the audio text, the word association score of the spoken audio is recorded as 0;
[0030] The average of the language structure score and the word association score of the spoken audio is recorded as the lexical structure score of the spoken audio.
[0031] Furthermore, based on the basic clarity score, the fluency score and vocabulary score are comprehensively analyzed, and based on the analysis results, the oral ability score is obtained, including:
[0032] Establish a plane rectangular coordinate system and record it as the comprehensive evaluation coordinate system, where the X-axis and Y-axis of the comprehensive evaluation coordinate system are both number axes; record the point (F2, 0) as the smooth reference point, and record the line connecting the coordinate origin and the smooth reference point as the smooth edge;
[0033] Obtain a straight line in the first quadrant of the comprehensive evaluation coordinate system with an angle of (90×F1)° to the smooth edge, and record it as the vocabulary structure candidate line; obtain a point in the first quadrant of the vocabulary structure candidate line with a distance from the coordinate origin equal to the vocabulary structure score, and record it as the vocabulary structure reference point;
[0034] The area of the triangle formed by the coordinate origin, the vocabulary structure reference point, and the fluency reference point is recorded as the speaking ability score of the spoken audio.
[0035] In a second aspect, the present application also provides a multi-dimensional oral ability identification and quantification system, including an audio basic analysis module, an oral ability analysis module, and an oral comprehensive evaluation module;
[0036] The audio basic analysis module is used to obtain the audio corresponding to the spoken language and record it as spoken audio; perform recognition preprocessing on the spoken audio; and obtain the language type and basic clarity score based on the processing results of the recognition preprocessing;
[0037] The speaking ability analysis module is used to analyze the fluency and vocabulary structure of the spoken audio using the fluency analysis method and the vocabulary analysis method based on the language type, and obtain the fluency score and vocabulary structure score of the spoken audio based on the analysis results;
[0038] The oral comprehensive evaluation module is used to conduct a comprehensive analysis of the fluency score and vocabulary score based on the basic clarity score, and obtain the oral ability score based on the analysis results.
[0039] Beneficial effects of the present invention: The present application first obtains audio corresponding to spoken language and records it as spoken audio; performs recognition preprocessing on the spoken audio; and obtains the language type and basic clarity score of the spoken audio based on the processing results of the recognition preprocessing. The advantage of this is that, by performing recognition preprocessing, the type of language in the spoken audio can be effectively identified, thereby providing a clear analysis direction for subsequent language analysis; and by obtaining the basic clarity score, a preliminary score can be made in terms of the clarity of the spoken expression, which helps to make an assessment based on the clarity of the spoken expression in the subsequent comprehensive analysis, thereby improving the effectiveness of the assessment of spoken language ability;
[0040] This application also analyzes the fluency and vocabulary structure of spoken audio by using fluency analysis and vocabulary analysis based on language type, and obtains the fluency score and vocabulary structure score of the spoken audio based on the analysis results; finally, the fluency score and vocabulary score are comprehensively analyzed based on the basic clarity score, and the oral ability score is obtained based on the analysis results. The advantage of this is that by obtaining the fluency score and vocabulary structure score through the fluency analysis and vocabulary analysis methods, the fluency of oral expression and the rationality of the oral word structure in the spoken audio can be judged respectively, and combined with the basic clarity score obtained above, it can ensure that when conducting a comprehensive analysis, the oral ability can be comprehensively judged, and at the same time, it can ensure that when facing multiple audios with the same semantics, the oral ability of each audio can still be differentiated and identified through the basic clarity score, fluency score and vocabulary structure score. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a principle block diagram of the system of the present invention;
[0042] Figure 2 is a flow chart of the steps of the method of the present invention;
[0043] Figure 3 A schematic diagram of the smooth analysis coordinate system of the present invention;
[0044] Figure 4 A schematic diagram of obtaining a speaking ability score according to the present invention;
[0045] Figure 5 Schematic diagram of the structure of the electronic device of the present invention. DETAILED DESCRIPTION
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0047] Example 1, please refer to Figure 1 As shown, the present application provides a multi-dimensional oral ability recognition and quantification system, including an audio basic analysis module, an oral ability analysis module, and an oral comprehensive evaluation module;
[0048] The audio basic analysis module is used to obtain the audio corresponding to the spoken language and record it as spoken audio; perform recognition preprocessing on the spoken audio; and obtain the language type and basic clarity score based on the processing results of the recognition preprocessing;
[0049] The audio basic analysis module includes a recognition preprocessing unit, which is configured with a recognition preprocessing strategy for implementing recognition preprocessing. The recognition preprocessing includes: using AI to recognize spoken audio, obtaining the language corresponding to the spoken audio based on the recognition result, and recording it as the language type of the spoken audio; converting the spoken audio into text based on the language type of the spoken audio, and recording it as audio text;
[0050] In the specific implementation process, for example, during a data processing, after AI recognizes spoken audio, it is determined that the language corresponding to the spoken audio is English. In this case, the language type of the spoken audio can be recorded as English, and in subsequent analysis, English-related grammar, conjunctions, or audio recognition software can be called to analyze the spoken audio;
[0051] The audio text is segmented and the words obtained are recorded as audio words YC1 to audio words YC based on the order of recognition in the spoken audio. n ;
[0052] In a specific implementation process, for example, during a data processing, the audio text obtained is "Oralproficiency test", then through analysis and processing, the audio words YC can be recorded as "Oral", "proficiency" and "test", and the number n is 3;
[0053] For any audio word YC c , the audio words YC in the spoken audio c The audio segments recognized by AI are recorded as audio sub-segments, where c is a positive integer less than or equal to n and greater than or equal to 1;
[0054] Obtain multiple audio recognition software, and respectively recognize the audio sub-segments and convert the audio sub-segments into text; obtain the recognition results of all audio recognition software, and when the language recognized in the recognition results of all audio recognition software is the language type of spoken audio, convert the audio word YC c The language clarity score is recorded as 1; when the language recognized in the recognition result of any audio recognition software is not the language type of spoken audio, the audio word YC c The speech intelligibility score was recorded as 0;
[0055] During the specific implementation process, for example, during a data analysis, the language type of the spoken audio is English, the number of audio recognition software is 5, and the audio sub-segment is the audio corresponding to the audio word YC "proficiency". Through the recognition of multiple audio recognition software, it is found that 4 of the audio recognition software recognize the audio word YC "proficiency" as English, and 1 audio recognition software recognizes the audio word YC "proficiency" as French. This means that there are flaws in the spoken expression of the audio word YC "proficiency", which makes it impossible to be correctly recognized. Therefore, the language clarity score of the audio word YC "proficiency" can be recorded as 0.
[0056] When the audio recognition software converts the audio sub-segment into text and the audio word YC c If the audio words YC are the same, c The vocabulary clarity score is recorded as 1; when the result of any audio recognition software converting the audio sub-segment into text is consistent with the audio word YC c If they are different, the audio words YC c The vocabulary clarity score of the audio word YC is recorded as 0, where the initial values of the language clarity score and vocabulary clarity score of each audio word YC are both 0;
[0057] Use the basic clarity algorithm to obtain the basic clarity score of the spoken audio. The basic clarity algorithm is: , where F1 is the basic clarity score, f i is the language clarity score corresponding to the i-th audio word YC in all audio words YC, g j is the vocabulary clarity score corresponding to the j-th audio word YC in all audio words YC; in the specific implementation process, for example, in a data processing, the value of n is 3, and the language clarity scores corresponding to the three audio words YC are 1, 1 and 1 respectively, and the vocabulary clarity scores corresponding to the three audio words YC are 1, 0 and 0, then it can be obtained by calculation that the basic clarity score is approximately 0.67. By obtaining the basic clarity score, a preliminary score can be made on the clarity of oral expression, which is helpful for making judgments based on the clarity of oral expression in the subsequent comprehensive analysis, so as to improve the effectiveness of oral ability judgment.
[0058] The speaking ability analysis module is used to analyze the fluency and vocabulary structure of the spoken audio using the fluency analysis method and the vocabulary analysis method based on the language type, and obtain the fluency score and vocabulary structure score of the spoken audio based on the analysis results;
[0059] The oral ability analysis module includes a fluency analysis unit and a vocabulary analysis unit. The fluency analysis unit is configured with a fluency analysis strategy. The fluency analysis strategy is used to implement a fluency analysis method. The fluency analysis method includes: recording the audio duration corresponding to the spoken audio as T, establishing a plane rectangular coordinate system, and recording it as a fluency analysis coordinate system, wherein the unit of the X-axis of the fluency analysis coordinate system is time, and the unit of the Y-axis is decibel; based on the sound decibel corresponding to the speech in the spoken audio, drawing a corresponding curve in the interval X=0 to X=T in the fluency analysis coordinate system, and recording it as a fluency analysis curve;
[0060] Based on the audio sub-segment corresponding to each audio word YC, the curves corresponding to all audio sub-segments are marked in the fluency analysis curve and recorded as the fluency sub-curve of the audio sub-segment;
[0061] For any smooth sub-curve α, the horizontal coordinates of the rightmost point and the leftmost point of the smooth sub-curve α are recorded as the right point duration and left point duration of the smooth sub-curve α respectively; when there is a smooth sub-curve to the left of the smooth sub-curve α, the difference between the left point duration of the smooth sub-curve α and the right point duration of the smooth sub-curve closest to the left of the smooth sub-curve α is recorded as the left pause value of the smooth sub-curve α; when there is no smooth sub-curve to the left of the smooth sub-curve α, the left pause value of the smooth sub-curve α is recorded as 0;
[0062] In the specific implementation process, for example, during a data analysis, the smooth analysis coordinate system obtained is as follows Figure 3 As shown, the curve between the dotted lines LL1 and LL2 is the smooth sub-curve α, the curve between the dotted lines JJ1 and JJ2 is the smooth sub-curve closest to the left of the smooth sub-curve α, and the curve between the dotted lines KK1 and KK2 is the smooth sub-curve closest to the right of the smooth sub-curve α. Then, through analysis, it can be obtained that the left interval value of the smooth sub-curve α is the difference between the ordinates of point DD2 and point DD1, and the right interval value of the smooth sub-curve α is the difference between the ordinates of point DD3 and point DD4.
[0063] When there is a smooth sub-curve on the right side of the smooth sub-curve α, the difference between the duration of the right point of the smooth sub-curve α and the duration of the left point of the smooth sub-curve closest to the right side of the smooth sub-curve α is recorded as the right pause value of the smooth sub-curve α; when there is no smooth sub-curve on the right side of the smooth sub-curve α, the right pause value of the smooth sub-curve α is recorded as 0;
[0064] The average of the left and right interval values of the fluency sub-curve α is recorded as the fluency interval value of the audio word YC corresponding to the fluency sub-curve α; the interval fluency values of all audio words YC are obtained, and the fluency score of the spoken audio is obtained based on the interval difference algorithm. The interval difference algorithm is: , where F2 is the fluency score, R uis the intermittent fluency value of the u-th audio word YC in all audio words YC, R sq is the average of all intermittent fluency values;
[0065] In the specific implementation process, for example, during a data processing, the intermittent fluency values of all audio words YC obtained are 10, 20, 14, 16 and 15 respectively. Then, by calculation, it can be obtained that the fluency score is 0.417. By obtaining the fluency score, the fluency of the spoken language in the spoken audio can be analyzed, which is helpful for the subsequent oral ability evaluation to combine the fluency of the spoken language for the final score. At the same time, in this embodiment, the larger the fluency score, the smoother the decibel change corresponding to the spoken language in the spoken audio, that is, the higher the fluency.
[0066] The vocabulary ability analysis unit is configured with a vocabulary ability analysis strategy, which is used to implement a vocabulary analysis method. The vocabulary analysis method includes: establishing an audio language structure library and an audio associated word library, wherein the audio language structure library and the audio associated word library are used to store language corresponding formats; based on the language type of the spoken audio, using big data to obtain the language structure and associated words corresponding to the language type of the spoken audio, and storing them in the audio language structure library and the audio associated word library respectively;
[0067] Obtain the language structure of the audio text based on AI; use the language structures in the audio language structure library to match the language structure of the audio text in sequence. When any language structure in the audio language structure library is equal to the language structure of the audio text, the language structure score of the spoken audio is recorded as 1; when all language structures in the audio language structure library are not equal to the language structure of the audio text, the language structure score of the spoken audio is recorded as 0;
[0068] In a specific implementation process, for example, during a data processing, if an audio structure in the audio language structure library is "As far as...is concerned", and the language structure of the obtained audio text contains "As far as...isconcerned", the language structure score of the spoken audio can be recorded as 1;
[0069] When any associated word in the audio associated vocabulary exists in the audio text, the language association score of the spoken audio is recorded as 1; when any associated word in the audio associated vocabulary does not exist in the audio text, the word association score of the spoken audio is recorded as 0;
[0070] The average of the language structure score and the word association score of the spoken audio is recorded as the lexical structure score of the spoken audio.
[0071] The oral comprehensive evaluation module is used to comprehensively analyze the fluency score and vocabulary score based on the basic clarity score, and obtain the oral ability score based on the analysis results. The oral comprehensive evaluation module includes a scoring integration unit, which is equipped with a scoring integration strategy. The scoring integration strategy includes:
[0072] Establish a plane rectangular coordinate system and record it as the comprehensive evaluation coordinate system, where the X-axis and Y-axis of the comprehensive evaluation coordinate system are both number axes; record the point (F2, 0) as the smooth reference point, and record the line connecting the coordinate origin and the smooth reference point as the smooth edge;
[0073] Obtain a straight line in the first quadrant of the comprehensive evaluation coordinate system with an angle of (90×F1)° to the smooth edge, and record it as the vocabulary structure candidate line; obtain a point in the first quadrant of the vocabulary structure candidate line with a distance from the coordinate origin equal to the vocabulary structure score, and record it as the vocabulary structure reference point;
[0074] In the specific implementation process, for example, in a data processing, the obtained F1 is 0.67 and the vocabulary structure score is 1. Then the straight line in the first quadrant of the comprehensive evaluation coordinate system and with an angle of 60.3° to the smooth edge can be recorded as the vocabulary structure candidate line, and the point in the vocabulary structure candidate line that is in the first quadrant and has a distance of 1 from the coordinate origin can be recorded as the vocabulary structure reference point. Figure 4 As shown, the straight line GG1 is the vocabulary structure candidate line, the point GG0 is the fluency reference point, and the point GG2 is the vocabulary structure reference point. The area of region QY is the oral ability score.
[0075] By using the area of the triangle formed by the coordinate origin, the vocabulary structure reference point and the fluency reference point as the oral ability score, the clarity, fluency and vocabulary structure of the spoken audio can be combined to ensure that when conducting a comprehensive analysis, an all-round evaluation of the oral ability can be made. At the same time, when faced with multiple audios with the same semantics, the oral ability of each audio can still be differentiated through the basic clarity score, fluency score and vocabulary structure score, so as to ensure the purpose of improving the effectiveness of oral ability identification.
[0076] The area of the triangle formed by the coordinate origin, the vocabulary structure reference point, and the fluency reference point is recorded as the speaking ability score of the spoken audio.
[0077] Example 2, please refer to Figure 2 As shown, the present application also provides a multi-dimensional oral ability identification and quantification method, comprising the following steps:
[0078] Step S1, obtaining audio corresponding to spoken language and recording it as spoken audio; performing recognition preprocessing on the spoken audio; obtaining the language type and basic clarity score of the spoken audio based on the processing results of the recognition preprocessing;
[0079] The recognition preprocessing includes: step S101, using AI to recognize the spoken audio, obtaining the language corresponding to the spoken audio based on the recognition result, and recording it as the language type of the spoken audio; converting the spoken audio into text based on the language type of the spoken audio, and recording it as audio text;
[0080] Step S102: segment the audio text and record the obtained words in the order of recognition in the spoken audio as audio words YC1 to audio words YC n ;
[0081] Step S103: for any audio word YC c , the audio words YC in the spoken audio c The audio segments recognized by AI are recorded as audio sub-segments, where c is a positive integer less than or equal to n and greater than or equal to 1;
[0082] Step S104: obtain multiple audio recognition software, and respectively recognize the audio sub-segments and convert the audio sub-segments into text; obtain the recognition results of all audio recognition software, and when the language recognized in the recognition results of all audio recognition software is the language type of spoken audio, convert the audio word YC c The language clarity score is recorded as 1; when the language recognized in the recognition result of any audio recognition software is not the language type of spoken audio, the audio word YC c The speech intelligibility score was recorded as 0;
[0083] Step S105, when the audio recognition software converts the audio sub-segment into text and the audio word YC c If the audio words YC are the same, c The vocabulary clarity score is recorded as 1; when the result of any audio recognition software converting the audio sub-segment into text is consistent with the audio word YC c If they are different, the audio words YC c The vocabulary clarity score of the audio word YC is recorded as 0, where the initial values of the language clarity score and vocabulary clarity score of each audio word YC are both 0;
[0084] Step S106: Obtain a basic clarity score of the spoken audio using a basic clarity algorithm. The basic clarity algorithm is: , where F1 is the basic clarity score, f i is the language clarity score corresponding to the i-th audio word YC in all audio words YC, g j is the vocabulary clarity score corresponding to the j-th audio word YC in all audio words YC.
[0085] Step S2, analyzing the fluency and vocabulary structure of the spoken audio using a fluency analysis method and a vocabulary analysis method based on the language type, and obtaining a fluency score and a vocabulary structure score of the spoken audio based on the analysis results;
[0086] Step S201, the fluency analysis method includes: Step S2011, recording the audio duration corresponding to the spoken audio as T, establishing a plane rectangular coordinate system, recorded as the fluency analysis coordinate system, wherein the unit of the X-axis of the fluency analysis coordinate system is time, and the unit of the Y-axis is decibel; based on the decibel of the sound corresponding to the speech in the spoken audio, drawing a corresponding curve in the interval X=0 to X=T in the fluency analysis coordinate system, and recording it as the fluency analysis curve;
[0087] Step S2012: Based on the audio sub-segment corresponding to each audio word YC, the curves corresponding to all audio sub-segments are marked in the fluency analysis curve and recorded as fluency sub-curves of the audio sub-segment;
[0088] Step S2013: For any smooth sub-curve α, the horizontal coordinates of the rightmost point and the leftmost point of the smooth sub-curve α are recorded as the right point duration and the left point duration of the smooth sub-curve α, respectively. When there is a smooth sub-curve to the left of the smooth sub-curve α, the difference between the left point duration of the smooth sub-curve α and the right point duration of the smooth sub-curve closest to the left of the smooth sub-curve α is recorded as the left pause value of the smooth sub-curve α. When there is no smooth sub-curve to the left of the smooth sub-curve α, the left pause value of the smooth sub-curve α is recorded as 0.
[0089] Step S2014: When there is a smooth sub-curve to the right of the smooth sub-curve α, the difference between the duration of the right point of the smooth sub-curve α and the duration of the left point of the smooth sub-curve closest to the right of the smooth sub-curve α is recorded as the right pause value of the smooth sub-curve α; when there is no smooth sub-curve to the right of the smooth sub-curve α, the right pause value of the smooth sub-curve α is recorded as 0;
[0090] Step S2015: Record the average of the left and right interval values of the fluency sub-curve α as the fluency interval value of the audio word YC corresponding to the fluency sub-curve α; obtain the interval fluency values of all audio words YC, and obtain the fluency score of the spoken audio based on the interval difference algorithm. The interval difference algorithm is: , where F2 is the fluency score, R u is the intermittent fluency value of the u-th audio word YC in all audio words YC, R sq is the average of all intermittent flow values.
[0091] Step S202, the lexical analysis method includes: Step S2021, establishing an audio language structure library and an audio associated word library, wherein the audio language structure library and the audio associated word library are used to store language formats; based on the language type of the spoken audio, using big data to obtain the language structure and associated words corresponding to the language type of the spoken audio, and storing them in the audio language structure library and the audio associated word library respectively;
[0092] Step S2022: Obtain the language structure of the audio text based on AI; sequentially match the language structure of the audio text with the language structures in the audio language structure library; when any language structure in the audio language structure library is equal to the language structure of the audio text, the language structure score of the spoken audio is recorded as 1; when all language structures in the audio language structure library are not equal to the language structure of the audio text, the language structure score of the spoken audio is recorded as 0;
[0093] Step S2023: When any associated word in the audio associated word library exists in the audio text, the language associated score of the spoken audio is recorded as 1; when any associated word in the audio associated word library does not exist in the audio text, the word associated score of the spoken audio is recorded as 0;
[0094] Step S2024: Record the average of the language structure score and the word association score of the spoken audio as the vocabulary structure score of the spoken audio.
[0095] Step S3, comprehensively analyzing the fluency score and the vocabulary score based on the basic clarity score, and obtaining a speaking ability score based on the analysis results;
[0096] Step S301: Establish a plane rectangular coordinate system and record it as a comprehensive evaluation coordinate system, where the X-axis and Y-axis of the comprehensive evaluation coordinate system are both number axes; record the point (F2, 0) as a smooth reference point, and record the line connecting the coordinate origin and the smooth reference point as a smooth edge;
[0097] Step S302: Obtain a straight line in the first quadrant of the comprehensive evaluation coordinate system, which has an angle of (90×F1)° with the smooth edge, and record it as the vocabulary structure selection line; obtain a point in the first quadrant of the vocabulary structure selection line, which has a distance from the coordinate origin equal to the vocabulary structure score, and record it as the vocabulary structure reference point;
[0098] Step S303: The area of the triangle formed by the coordinate origin, the vocabulary structure reference point, and the fluency reference point is recorded as the speaking ability score of the spoken audio.
[0099] Example 3, please refer to Figure 5 As shown, Figure 5A schematic diagram of the structure of an electronic device is provided, which may include: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus. The memory stores computer-readable instructions, and the processor can call the instructions in the memory. When the computer-readable instructions are executed by the processor, the steps of a multi-dimensional oral proficiency recognition and quantification method are executed to implement the following functions: first, audio corresponding to spoken language is obtained and recorded as spoken audio; recognition preprocessing is performed on the spoken audio; based on the processing results of the recognition preprocessing, the language type and basic clarity score of the spoken audio are obtained; then, based on the language type, the fluency and vocabulary structure of the spoken audio are analyzed using a fluency analysis method and a vocabulary analysis method, and based on the analysis results, a fluency score and a vocabulary structure score of the spoken audio are obtained; finally, a comprehensive analysis is performed on the fluency score and vocabulary score based on the basic clarity score, and based on the analysis results, a speaking proficiency score is obtained.
[0100] In addition, the logical instructions in the above-mentioned memory can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0101] Example 4. The present application also provides a computer-readable storage medium. The present application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above multi-dimensional oral ability recognition and quantification method are executed to achieve the following functions: first, the audio corresponding to the spoken language is obtained and recorded as spoken audio; the spoken audio is pre-processed for recognition; the language type and basic clarity score of the spoken audio are obtained based on the processing results of the recognition pre-processing; then, the fluency and vocabulary structure of the spoken audio are analyzed using a fluency analysis method and a vocabulary analysis method based on the language type, and the fluency score and vocabulary structure score of the spoken audio are obtained based on the analysis results; finally, the fluency score and vocabulary score are comprehensively analyzed based on the basic clarity score, and the oral ability score is obtained based on the analysis results.
[0102] Through the description of the above embodiments, the embodiments of the present invention can be provided as methods, systems, or computer program products. Based on this understanding, the essence of the above technical solutions or the portion that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (such as a personal computer, server, or network device) to execute the methods described in various embodiments or certain portions of the embodiments.
[0103] In the embodiments provided in this application, it should be understood that the disclosed system or method can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of systems, modules and units can be electrical, mechanical or other forms.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A multi-dimensional oral ability identification and quantification method, characterized by: The steps include: Acquire audio corresponding to spoken language and record it as spoken audio; perform recognition preprocessing on the spoken audio; obtain the language type and basic clarity score of the spoken audio based on the processing results of the recognition preprocessing; Analyze the fluency and vocabulary structure of the spoken audio using fluency analysis and vocabulary analysis based on the language type, and obtain a fluency score and vocabulary structure score for the spoken audio based on the analysis results; Comprehensively analyze the fluency score and vocabulary score based on the basic clarity score, and obtain the oral ability score based on the analysis results; Recognition preprocessing includes: Use AI to recognize spoken audio, obtain the language corresponding to the spoken audio based on the recognition result, and record it as the language type of the spoken audio; convert the spoken audio into text based on the language type of the spoken audio and record it as audio text; The audio text is segmented and the words obtained are recorded as audio words YC1 to audio words YC based on the order of recognition in the spoken audio. n ; For any audio word YC c , the audio words YC in the spoken audio c The audio segments recognized by AI are recorded as audio sub-segments, where c is a positive integer less than or equal to n and greater than or equal to 1; Obtain multiple audio recognition software, and respectively recognize the audio sub-segments and convert the audio sub-segments into text; obtain the recognition results of all audio recognition software, and when the language recognized in the recognition results of all audio recognition software is the language type of spoken audio, convert the audio word YC c The language clarity score is recorded as 1; when the language recognized in the recognition result of any audio recognition software is not the language type of spoken audio, the audio word YC c The speech intelligibility score was recorded as 0; When the audio recognition software converts the audio sub-segment into text and the audio word YC c If the audio words YC are the same, c The vocabulary clarity score is recorded as 1; when the result of any audio recognition software converting the audio sub-segment into text is consistent with the audio word YC c If they are different, the audio words YC c The vocabulary clarity score of the audio word YC is recorded as 0, where the initial values of the language clarity score and vocabulary clarity score of each audio word YC are both 0; Use the basic clarity algorithm to obtain the basic clarity score of the spoken audio. The basic clarity algorithm is: , where F1 is the basic clarity score, f i is the language clarity score corresponding to the i-th audio word YC in all audio words YC, g j is the vocabulary clarity score corresponding to the j-th audio word YC in all audio words YC.
2. A multi-dimensional oral ability identification and quantification method according to claim 1, characterized in that: Fluency analysis methods include: The audio duration corresponding to the spoken audio is recorded as T, and a plane rectangular coordinate system is established, recorded as the fluency analysis coordinate system, where the unit of the X-axis of the fluency analysis coordinate system is time, and the unit of the Y-axis is decibel; based on the decibel of the sound corresponding to the speech in the spoken audio, a corresponding curve is drawn in the interval from X=0 to X=T in the fluency analysis coordinate system, and recorded as the fluency analysis curve; Based on the audio sub-segment corresponding to each audio word YC, the curves corresponding to all audio sub-segments are marked in the fluency analysis curve and recorded as the fluency sub-curve of the audio sub-segment; For any smooth sub-curve α, the horizontal coordinates of the rightmost point and the leftmost point of the smooth sub-curve α are recorded as the right point duration and the left point duration of the smooth sub-curve α respectively; when there is a smooth sub-curve on the left side of the smooth sub-curve α, the difference between the left point duration of the smooth sub-curve α and the right point duration of the smooth sub-curve closest to the left side of the smooth sub-curve α is recorded as the left interval value of the smooth sub-curve α; when there is no smooth sub-curve on the left side of the smooth sub-curve α, the left interval value of the smooth sub-curve α is recorded as 0.
3. A multi-dimensional oral ability identification and quantification method according to claim 1, characterized in that: Fluency analysis also includes: When there is a smooth sub-curve on the right side of the smooth sub-curve α, the difference between the duration of the right point of the smooth sub-curve α and the duration of the left point of the smooth sub-curve closest to the right side of the smooth sub-curve α is recorded as the right pause value of the smooth sub-curve α; when there is no smooth sub-curve on the right side of the smooth sub-curve α, the right pause value of the smooth sub-curve α is recorded as 0; The average of the left and right interval values of the fluency sub-curve α is recorded as the fluency interval value of the audio word YC corresponding to the fluency sub-curve α; the interval fluency values of all audio words YC are obtained, and the fluency score of the spoken audio is obtained based on the interval difference algorithm. The interval difference algorithm is: , where F2 is the fluency score, R u is the intermittent fluency value of the u-th audio word YC in all audio words YC, R sq is the average of all intermittent flow values.
4. A multi-dimensional oral ability identification and quantification method according to claim 1, characterized in that: Lexical analysis methods include: Establishing an audio language structure library and an audio associated word library, wherein the audio language structure library and the audio associated word library are used to store the formats corresponding to the languages; based on the language type of the spoken audio, using big data to obtain the language structure and associated words corresponding to the language type of the spoken audio, and storing them in the audio language structure library and the audio associated word library respectively; The language structure of the audio text is obtained based on AI; the language structure of the audio text is matched with the language structures in the audio language structure library in turn. When any language structure in the audio language structure library is equal to the language structure of the audio text, the language structure score of the spoken audio is recorded as 1; when all language structures in the audio language structure library are not equal to the language structure of the audio text, the language structure score of the spoken audio is recorded as 0.
5. The multi-dimensional oral ability identification and quantification method according to claim 1, characterized in that: Lexical analysis also includes: When any associated word in the audio associated vocabulary exists in the audio text, the language association score of the spoken audio is recorded as 1; when any associated word in the audio associated vocabulary does not exist in the audio text, the word association score of the spoken audio is recorded as 0; The average of the language structure score and the word association score of the spoken audio is recorded as the lexical structure score of the spoken audio.
6. A multi-dimensional oral ability identification and quantification method according to claim 5, characterized in that: Based on the basic clarity score, the fluency score and vocabulary score are comprehensively analyzed, and the oral ability score is obtained based on the analysis results, including: Establish a plane rectangular coordinate system and record it as the comprehensive evaluation coordinate system, where the X-axis and Y-axis of the comprehensive evaluation coordinate system are both number axes; record the point (F2, 0) as the smooth reference point, and record the line connecting the coordinate origin and the smooth reference point as the smooth edge; Obtain a straight line in the first quadrant of the comprehensive evaluation coordinate system with an angle of (90×F1)° to the smooth edge, and record it as the vocabulary structure candidate line; obtain a point in the first quadrant of the vocabulary structure candidate line with a distance from the coordinate origin equal to the vocabulary structure score, and record it as the vocabulary structure reference point; The area of the triangle formed by the coordinate origin, the vocabulary structure reference point, and the fluency reference point is recorded as the speaking ability score of the spoken audio.
7. A multi-dimensional oral ability recognition and quantification system, used to implement the multi-dimensional oral ability recognition and quantification method according to any one of claims 1 to 6, characterized in that: It includes audio basic analysis module, oral ability analysis module and oral comprehensive evaluation module; The audio basic analysis module is used to obtain the audio corresponding to the spoken language and record it as spoken audio; perform recognition preprocessing on the spoken audio; and obtain the language type and basic clarity score based on the processing results of the recognition preprocessing; The speaking ability analysis module is used to analyze the fluency and vocabulary structure of the spoken audio using the fluency analysis method and the vocabulary analysis method based on the language type, and obtain the fluency score and vocabulary structure score of the spoken audio based on the analysis results; The oral comprehensive evaluation module is used to conduct a comprehensive analysis of the fluency score and vocabulary score based on the basic clarity score, and obtain the oral ability score based on the analysis results.
Citation Information
Patent Citations
Colloquial word identification and semantic identification method and device
CN108829894A
Spoken language pronunciation evaluation method and system for minority language, and storage medium
CN112967711A