Method, system, and storage medium for evaluating oral proficiency

By preprocessing and analyzing the environmental audio signals, an environmental noise impact assessment index is generated. Combined with the speaking speed assessment index, the problems of inconsistent speaking speed and noise impact in speech recognition technology are solved, and a more accurate evaluation of oral ability is achieved.

CN120412650BActive Publication Date: 2025-10-17YUFENG CULTURE TECHNOLOGY (NANTONG) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510500934.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-10-17
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

When evaluating oral ability, existing speech recognition technology is affected by the inconsistent speaking speed of the person being evaluated and environmental noise, resulting in insufficient evaluation accuracy.

Method used

By collecting environmental audio signals, performing preprocessing and short-time Fourier transform, generating discrete time and frequency domain signals, analyzing the environmental noise power, combining the speaking speed and noise impact assessment index, generating a comprehensive oral evaluation index, and conducting a comprehensive evaluation.

Benefits of technology

It improves the accuracy of oral ability evaluation, reduces the impact of environmental noise and inconsistent speaking speed on the evaluation, and provides more objective and accurate evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412650B_ABST
    Figure CN120412650B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system and storage medium for evaluating oral ability, and relates to the technical field of oral evaluation. The present invention collects and preprocesses ambient audio signals from a recording environment to generate an ambient noise power for reflecting the power value of ambient noise, performs correlation analysis on the signals, and generates an ambient noise impact assessment index reflecting the degree of influence of ambient noise on the recording signal. At the same time, a standard oral set is set, and oral parameters of oral ability training subjects are collected. After analysis, a speech speed accuracy assessment index is generated for evaluating the relationship between the speaking speed of a person to be tested and the output accuracy of a device. After comprehensive analysis, a comprehensive oral evaluation index is generated for reflecting a comprehensive evaluation based on the current environment and the speaking speed of the person to be tested. After comparison with an accuracy threshold, the accuracy level of the oral ability evaluation of the person to be tested is output to evaluate the accuracy of the oral ability evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of oral evaluation, in particular to a method and system for evaluating oral ability and a storage medium. BACKGROUND

[0002] Speech recognition technology has made significant progress in recent years and has become an important tool for evaluating oral ability. It converts human speech signals into readable text, thereby analyzing and evaluating an individual's language expression ability. In the field of education, especially in foreign language learning, speech recognition systems can provide real-time feedback on learners' pronunciation, speech rate, and fluency, helping them identify and correct pronunciation errors, and improve the accuracy and confidence of oral expression. At the same time, this technology can automatically score oral tests, providing more objective and fair evaluation results and reducing the workload of teachers. Through in-depth analysis of speech data, speech recognition systems can not only recognize the content of speech, but also assess the fluency and naturalness of oral speech by analyzing elements such as prosody, intonation, and stress. In addition, speech recognition technology is adaptable and can be optimized according to the pronunciation habits and language background of different learners, thereby providing personalized learning experiences. The application of this technology not only improves the efficiency and accuracy of oral ability evaluation, but also provides learners with more abundant learning resources and promotes the motivation for self-directed learning and continuous progress.

[0003] Generally, when evaluating oral ability based on speech recognition, the accuracy of the evaluator whose speech rate is close to the standard oral speech rate is relatively high, while the accuracy of the evaluator whose speech rate deviates from the standard oral speech rate is relatively low. At the same time, environmental noise can affect the collected audio signals and also affect the evaluation accuracy.

[0004] The above information disclosed in the background section is only used to enhance the understanding of the background of the present disclosure, and therefore it can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY

[0005] The purpose of the present application is to provide a method and system for evaluating oral ability and a storage medium to solve the problems raised in the background.

[0006] To achieve the above-mentioned purpose, the present application provides the following technical solutions:

[0007] The method for evaluating oral ability, the specific steps of the present application include:

[0008] S1, collecting the environmental audio signal of the input environment;

[0009] S2, pre-processing the environmental audio signal to generate a discrete time signal, and performing a short-time Fourier transform on the environmental audio signal in combination with the discrete time signal to generate an environmental audio frequency domain signal, the discrete time signal being used to reflect the signal intensity at each discrete time point, and the environmental audio frequency domain signal being used to reflect the amplitude of the environmental audio at the corresponding frequency and time;

[0010] S3, performing a correlation analysis on the environmental audio frequency domain signal to generate an environmental audio power spectrum, and performing a correlation analysis on the environmental audio power spectrum to generate an environmental noise power, the environmental audio power spectrum being used to reflect the power value of the environmental audio at the corresponding frequency and time, and the environmental noise power being used to reflect the power value of the environmental noise;

[0011] S4, performing a correlation analysis on the environmental noise power to generate an environmental noise impact evaluation index HZP, the environmental noise impact evaluation index HZP being used to reflect the influence degree of the environmental noise on the recording signal;

[0012] S5, setting a standard oral set, setting 20 oral ability training objects to read against the standard oral set, collecting oral parameters in the reading process, the oral parameters including the speech speed and the recognition accuracy, performing a correlation analysis on the oral parameters in the reading process to generate an accuracy expression, performing oral test on the to-be-tested person using a speech recognition device, and generating a speech speed and accuracy evaluation index YZP of the to-be-tested person based on the speech speed and the accuracy expression of the to-be-tested person, the speech speed and accuracy evaluation index YZP being the recognition accuracy of the speech recognition device for the speech speed of the to-be-tested person, the to-be-tested person being a person who will use speech recognition to evaluate oral ability;

[0013] S6, performing a correlation analysis on the speech speed and accuracy evaluation index YZP and the environmental noise impact evaluation index HZP to generate an oral evaluation comprehensive index KPZ, the oral evaluation comprehensive index KPZ being used to reflect the comprehensive evaluation combining the current environment and the speech speed of the to-be-tested person;

[0014] S7, comparing the oral evaluation comprehensive index KPZ with an accuracy threshold to output the oral ability evaluation accuracy level of the to-be-tested person.

[0015] Further, in the recording environment, one microphone is arranged on each side of the recording point, the microphones are connected through an audio acquisition device, and the environmental audio signals collected by the two microphones are stored as time sequence signals.

[0016] Further, the environmental audio signal is pre-processed to generate a discrete time signal , the formula being:

[0017] in, , is the sampling frequency, n is a positive integer, used to index discrete sampling time points, discrete time signal Used to reflect the signal strength at each discrete time point.

[0018] Furthermore, the ambient audio signal is subjected to short-time Fourier transform in combination with the discrete time signal to generate the ambient audio frequency domain signal , based on the formula:

[0019] in, represents an imaginary number, is a window function used to limit the time range. , where N is the window size, , ambient audio frequency domain signal Used to reflect the amplitude of ambient audio at frequency f and time t.

[0020] Furthermore, the ambient audio frequency domain signal is subjected to correlation analysis to generate the ambient audio power spectrum. , based on the formula:

[0021] Ambient audio power spectrum Used to reflect the power value of the ambient audio at the corresponding frequency f and time t. It is a modular operation.

[0022] Furthermore, the ambient audio power spectrum is subjected to correlation analysis to generate the ambient noise power , based on the formula:

[0023] Among them, F is the maximum frequency value, T is the time for the microphone to collect the ambient audio signal, in seconds, and the ambient noise power Used to reflect the power value of environmental noise;

[0024] Ambient noise power Correlation analysis is performed to generate the environmental noise impact assessment index HZP, based on the following formula:

[0025]

[0026] in, is the input power of the microphone, The microphone's input sensitivity weighting factor, The value range is ,The environmental noise impact assessment index HZP is used to reflect the impact of environmental noise on the recorded signal.

[0027] Further, the standard spoken language set is a segment of spoken language contrast audio, including a standard speech rate SR, and the speech rate The word quantity of each minute of the repeated reading of the i th spoken language ability training object The accuracy rate of the word recognition output by the speech recognition device for the i th spoken language ability training object is analyzed in relation to the spoken language parameters in the repeated reading process, and an accuracy rate expression AR is generated, so the formula is:

[0028]

[0029] All the spoken language parameters of the collected spoken language ability training objects are substituted into the above formula, the speech rate is substituted into v, and the recognition accuracy rate is substituted into AR, and the values of a and b are fitted and generated by MATLAB software;

[0030] The accuracy rate expression AR is analyzed in relation to the correlation, and a speech rate accuracy rate evaluation index YZP is generated, and the formula is:

[0031]

[0032] V is the speech rate of the to-be-tested person, and the speech rate accuracy rate evaluation index YZP is used to evaluate the relationship between the speech rate of the to-be-tested person and the accuracy rate of the device output;

[0033] The speech rate accuracy rate evaluation index YZP and the environmental noise influence evaluation index HZP are analyzed in relation to the correlation, and a spoken language evaluation comprehensive index KPZ is generated, and the formula is:

[0034]

[0035] is a speech rate accuracy rate weight factor, The value range of The spoken language evaluation comprehensive index KPZ is used to reflect the accuracy degree of the comprehensive evaluation of the spoken language ability test in combination with the current environment and the speech rate of the to-be-tested person.

[0036] Further, the accuracy threshold The value of is 15.7, the spoken language evaluation comprehensive index KPZ is compared with the accuracy threshold When , it is indicated that the accuracy degree of the spoken language ability evaluation of the to-be-tested person in the current environment is level two, and the spoken language ability evaluation is inaccurate; when , it is indicated that the accuracy degree of the spoken language ability evaluation of the to-be-tested person in the current environment is level one, and the spoken language ability evaluation is accurate.

[0037] The application further provides a system for evaluating spoken language ability, which is used for executing the method for evaluating spoken language ability, and comprises: ​

[0038] a signal collection module, configured to collect an environmental audio signal of an input environment;

[0039] an environmental audio signal processing module, configured to pre-process the environmental audio signal, generate a discrete-time signal, and perform a short-time Fourier transform on the environmental audio signal in combination with the discrete-time signal to generate an environmental audio frequency domain signal;

[0040] a frequency domain signal analysis module, configured to perform a correlation analysis on the environmental audio frequency domain signal to generate an environmental audio power spectrum, and perform a correlation analysis on the environmental audio power spectrum to generate an environmental noise power;

[0041] an environmental noise power analysis module, configured to perform a correlation analysis on the environmental noise power to generate an environmental noise influence evaluation index;

[0042] a speech speed analysis module, configured to set a standard spoken language set, set 20 spoken language ability training objects to read aloud against the standard spoken language set, collect spoken language parameters in the reading aloud process, perform a correlation analysis on the spoken language parameters in the reading aloud process to generate an accuracy expression, and perform a correlation analysis on the accuracy expression to generate a speech speed accuracy evaluation index;

[0043] a comprehensive analysis module, configured to perform a correlation analysis on the speech speed accuracy evaluation index YZP and the environmental noise influence evaluation index HZP to generate a spoken language evaluation comprehensive index KPZ, which is used to reflect a comprehensive evaluation in combination with a current environment and a speech speed of a to-be-tested person.

[0044] The application further provides a storage medium having computer program instructions stored thereon, and the computer program instructions are executed by a processor to implement the method.

[0045] Compared with the prior art, the application has the following beneficial effects:

[0046] The application collects an environmental audio signal of an input environment, pre-processes the environmental audio signal, generates an environmental noise power used to reflect a power value of environmental noise, performs a correlation analysis on the environmental noise power, generates an environmental noise influence evaluation index reflecting an influence degree of environmental noise on an input signal, sets a standard spoken language set, collects spoken language parameters of spoken language ability training objects, generates a speech speed accuracy evaluation index used to evaluate a relationship between a speech speed of a to-be-tested person and an output accuracy of a device after analysis, generates a spoken language evaluation comprehensive index used to reflect a comprehensive evaluation in combination with a current environment and a speech speed of a to-be-tested person after comprehensive analysis, and outputs a spoken language ability evaluation accuracy level of the to-be-tested person after comparison with an accuracy threshold, thereby evaluating the accuracy of the spoken language ability evaluation. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 It is a schematic diagram of the overall method of the application.

[0048] Figure 2 The whole system flowchart of the present application. DETAILED DESCRIPTION

[0049] For the purpose, technical solutions and advantages of the present application to be more clearly and intelligibly understood, the present application is further described in detail below with specific embodiments.

[0050] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present application should be understood as the usual meanings understood by those with ordinary skills in the art to which the present application belongs. The terms "first", "second", and similar terms used in the present application do not represent any order, number, or importance, but are only used to distinguish different components. The terms "include" or "contain" and similar terms mean that the elements or objects appearing before the terms cover the elements or objects listed after the terms and their equivalents, without excluding other elements or objects. The terms "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "up", "down", "left", "right", and the like only represent relative positional relationships, which can change accordingly when the absolute position of the described object changes.

[0051] EMBODIMENT

[0052] Referring to Figure 1 and Figure 2 , the present application provides a technical solution:

[0053] Referring to Figure 1 , the method for evaluating oral English ability is used to evaluate the accuracy of oral English evaluation based on speech recognition. In the prior art, due to the limitation of human resources, speech recognition is often used to evaluate the accuracy of oral English evaluation. However, due to the different speaking speeds of the evaluators, the accuracy of the evaluators whose speaking speed is close to the standard oral English speed is relatively high, while the accuracy of the evaluators whose speaking speed deviates from the standard oral English speed is relatively low. At the same time, environmental noise will affect the collected audio signals and also affect the accuracy. In order to evaluate the influence of speaking speed and environment on oral English evaluation accuracy, the specific steps include:

[0054] Step 1, collecting the environmental audio signals of the input environment;

[0055] In order to collect the audio signals of the environment, one microphone is arranged on each side of the input point in the input environment, and the microphones are connected through an audio collection device. The environmental audio signals collected by the two microphones are stored as time series signals, and the audio signals of the environment are collected through the two microphones and are mean-processed, thereby improving the accuracy.

[0056] Step 2, in order to better analyze the environmental audio signal, it is necessary to carry out short-time Fourier transform, so as to more intuitively show the relationship among time, frequency and signal strength. The environmental audio signal is preprocessed to generate a discrete time signal, and the environmental audio signal is combined with the discrete time signal to carry out short-time Fourier transform to generate an environmental audio frequency domain signal. The discrete time signal is used to reflect the signal strength of each discrete time point, and the environmental audio frequency domain signal is used to reflect the amplitude of the environmental audio at the corresponding frequency and time.

[0057] The environmental audio signal is preprocessed to generate a discrete time signal , and the formula is:

[0058] wherein, , is the sampling frequency, n is a positive integer, and is used to index the discrete sampling time point. The interval of the time point of the environmental audio signal sampled by the microphone is , starting from zero, the discrete time signal is used to reflect the signal strength of each discrete time point.

[0059] The environmental audio signal is combined with the discrete time signal to carry out short-time Fourier transform to generate an environmental audio frequency domain signal , and the formula is:

[0060] wherein, represents an imaginary number, is a window function, which is used to limit the time range, , wherein N is the size of the window, , the environmental audio frequency domain signal is used to reflect the amplitude of the environmental audio at the frequency f and the time t.

[0061] Step 3, in order to analyze the power size of the environmental audio, the environmental audio frequency domain signal is subjected to correlation analysis to generate an environmental audio power spectrum, and the environmental audio power spectrum is subjected to correlation analysis to generate an environmental noise power. The environmental audio power spectrum is used to reflect the power value of the environmental audio at the corresponding frequency and time, and the environmental noise power is used to reflect the power value of the environmental noise.

[0062] The environmental audio frequency domain signal is subjected to correlation analysis to generate an environmental audio power spectrum , and the formula is:

[0063] The environmental audio power spectrum is used to reflect the power value of the environmental audio at the corresponding frequency f and the time t. is a modulo operation.

[0064] The environmental audio power spectrum is subjected to correlation analysis to generate an environmental noise power , and the formula is:

[0065] wherein F is a maximum frequency value, T is the time for the microphone to collect the environmental audio signal, in seconds, and the environmental noise power is used to reflect the power value of the environmental noise; the environmental noise power The greater the value of the environmental noise power

[0066] Step 4, the environmental noise power is subjected to correlation analysis to generate an environmental noise influence evaluation index HZP, which is used to reflect the influence degree of the environmental noise on the input signal;

[0067] The environmental noise power is subjected to correlation analysis to generate an environmental noise influence evaluation index HZP, and the formula is:

[0068]

[0069] wherein is the input power of the microphone, is the input sensitivity weight factor of the microphone, The value range of the environmental noise influence evaluation index HZP is The environmental noise influence evaluation index HZP is used to reflect the influence degree of the environmental noise on the input signal, and the greater the value of the environmental noise influence evaluation index HZP, the higher the degree of the environmental noise participating in the input audio, and thus the higher the proportion of the environmental noise, i.e., the lower the accuracy of the speech recognition.

[0070] Since the speech speed of the to-be-evaluated person is not the same, the accuracy of the to-be-evaluated person who is close to the standard spoken language speed will be higher, and the accuracy of the to-be-evaluated person who deviates from the standard spoken language speed will be lower, so the speech speed of the to-be-evaluated person will affect the accuracy of the spoken language evaluation, and the relationship is analyzed through the following steps:

[0071] Step 5, set a standard oral language set, set 20 oral language ability training objects to read against the standard oral language set, collect oral language parameters in the reading process, the oral language parameters include speech speed and recognition accuracy, the recognition accuracy is a ratio, that is, the ratio of the number of words recognized correctly by the speech recognition device to the total number of words in the reading content of the oral language ability training object, perform correlation analysis on the oral language parameters in the reading process, generate an accuracy expression, use a speech recognition device to test the oral language of the to-be-tested person, based on the speech speed and accuracy expression of the to-be-tested person, generate a speech speed and accuracy evaluation index YZP of the to-be-tested person, the speech speed and accuracy evaluation index YZP is the recognition accuracy of the speech recognition device for the speech speed of the to-be-tested person, the to-be-tested person is a person who will use speech recognition to evaluate oral language ability;

[0072] Step 6, perform correlation analysis on the speech speed and accuracy evaluation index YZP and the environmental noise influence evaluation index HZP, and generate an oral language evaluation comprehensive index KPZ, the oral language evaluation comprehensive index KPZ is used to reflect the comprehensive evaluation combining the current environment and the speech speed of the to-be-tested person;

[0073] The standard oral language set is a piece of oral language contrast audio, including a standard speech speed SR, the speech speed is the number of words read per minute by the i th oral language ability training object, the recognition accuracy is the accuracy of word recognition output by the speech recognition device of the i th oral language ability training object, perform correlation analysis on the oral language parameters in the reading process, and generate an accuracy expression AR, so the formula is:

[0074]

[0075] Substitute all the oral language parameters of the oral language ability training objects into the above formula, substitute the speech speed into v, and substitute the recognition accuracy into AR, and generate the values of a and b by fitting through MATLAB software;

[0076] Perform correlation analysis on the accuracy expression AR to generate a speech speed and accuracy evaluation index YZP, and the formula is:

[0077]

[0078] Wherein, V is the speech speed of the to-be-tested person, and the speech speed and accuracy evaluation index YZP is used to evaluate the relationship between the speech speed of the to-be-tested person and the output accuracy of the device;

[0079] Perform correlation analysis on the speech speed and accuracy evaluation index YZP and the environmental noise influence evaluation index HZP to generate an oral language evaluation comprehensive index KPZ, and the formula is:

[0080]

[0081] wherein, is a speech speed accuracy rate weight factor, the value range of is , the speech speed accuracy rate weight factor is used for adjusting the influence degree of the speech speed on the accuracy rate, the greater the value is, the higher the weight is, the speech evaluation comprehensive index KPZ is used for reflecting the accuracy degree of the comprehensive evaluation of the oral ability test combined with the current environment and the speech speed of the person to be tested, the greater the value of the speech evaluation comprehensive index KPZ is, the lower the accuracy degree of the comprehensive evaluation of the oral ability test is.

[0082] Step 7, comparing the speech evaluation comprehensive index KPZ with the accuracy threshold , and outputting the oral ability evaluation accuracy level of the person to be tested.

[0083] The accuracy threshold is 15.7, the speech evaluation comprehensive index KPZ is compared with the accuracy threshold , when , it is indicated that the oral ability evaluation accuracy degree of the person to be tested under the current environment is level two, and the oral ability evaluation is inaccurate; when , it is indicated that the oral ability evaluation accuracy degree of the person to be tested under the current environment is level one, and the oral ability evaluation is accurate, the greater the value of the speech evaluation comprehensive index KPZ is, the lower the accuracy degree of the oral ability test is, when the oral ability evaluation accuracy degree of the person to be tested is level two, it needs to be re-determined.

[0084] With reference to Figure 2 , the application further provides a system for evaluating oral ability, which is used for executing the method for evaluating oral ability, and comprises:

[0085] A signal collection module is used for collecting an environmental audio signal of an input environment;

[0086] An environmental audio signal processing module is used for pre-processing the environmental audio signal, generating a discrete time signal, and performing short-time Fourier transform on the environmental audio signal in combination with the discrete time signal to generate an environmental audio frequency domain signal;

[0087] A frequency domain signal analysis module is used for performing correlation analysis on the environmental audio frequency domain signal to generate an environmental audio power spectrum, and performing correlation analysis on the environmental audio power spectrum to generate an environmental noise power;

[0088] An environmental noise power analysis module is used for performing correlation analysis on the environmental noise power to generate an environmental noise influence evaluation index;

[0089] The speech speed analysis module is configured to set a standard spoken language set, set 20 spoken language ability training objects to read aloud against the standard spoken language set, collect spoken language parameters during the reading aloud process, perform correlation analysis on the spoken language parameters during the reading aloud process, generate an accuracy expression, perform correlation analysis on the accuracy expression, and generate a speech speed accuracy evaluation index;

[0090] The comprehensive analysis module is configured to perform correlation analysis on the speech speed accuracy evaluation index YZP and the environmental noise influence evaluation index HZP, and generate a spoken language evaluation comprehensive index KPZ.

[0091] The present application also provides a storage medium having computer program instructions stored thereon, and the computer program instructions are executed by a processor to implement the method.

[0092] The above formulas are all dimensionless numerical calculations, and the formulas are obtained by software simulation of a large amount of data to obtain a formula of the most recent real situation, and the preset parameters in the formula are set by a person skilled in the art according to the actual situation.

[0093] The above embodiments can be realized wholly or partially by software, hardware, firmware or any other combination. When realized by software, the above embodiments can be realized in the form of a computer program product. Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized by hardware or software methods depends on the specific application and design constraints of the technical solutions.

[0094] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, and can be located in one place or distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0095] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application.

Claims

1. A method for evaluating oral proficiency, characterized in that: The specific steps include: S1, collecting the ambient audio signal of the recording environment; S2. Preprocess the ambient audio signal to generate a discrete time signal, and perform a short-time Fourier transform on the ambient audio signal in combination with the discrete time signal to generate an ambient audio frequency domain signal. The discrete time signal is used to reflect the signal strength at each discrete time point, and the ambient audio frequency domain signal is used to reflect the amplitude of the ambient audio at the corresponding frequency and time. S3. Performing a correlation analysis on the ambient audio frequency domain signal to generate an ambient audio power spectrum, and performing a correlation analysis on the ambient audio power spectrum to generate an ambient noise power. The ambient audio power spectrum is used to reflect the power value of the ambient audio at the corresponding frequency and time, and the ambient noise power is used to reflect the power value of the ambient noise. S4. Performing a correlation analysis on the ambient noise power to generate an ambient noise impact assessment index HZP, where the ambient noise impact assessment index HZP is used to reflect the degree of impact of the ambient noise on the recorded signal; S5. Set a standard spoken language set, and have 20 oral proficiency training subjects repeat the language according to the standard spoken language set. Collect spoken language parameters during the repetition process, including speaking rate and recognition accuracy. Perform correlation analysis on the spoken language parameters during the repetition process to generate an accuracy expression. Conduct a spoken language test on the subject using a speech recognition device. Based on the speaking rate and accuracy expression of the subject, generate a speaking rate accuracy evaluation index YZP of the subject. The speaking rate accuracy evaluation index YZP is the recognition accuracy rate of the speech recognition device for the speaking rate of the subject. The subject is the subject for whom speech recognition will be used to evaluate oral proficiency. S6. Perform a correlation analysis on the speech speed accuracy evaluation index YZP and the environmental noise impact evaluation index HZP to generate a comprehensive oral evaluation index KPZ. The comprehensive oral evaluation index KPZ is used to reflect a comprehensive evaluation based on the current environment and the speaking speed of the person being tested. S7. Combine the oral evaluation comprehensive index KPZ with the accuracy threshold A comparison is performed and the accuracy level of the oral ability evaluation of the person to be tested is output. When the accuracy level of the oral ability evaluation of the person to be tested is level one, there is no need to re-evaluate the oral ability through voice recognition. When the accuracy level of the oral ability evaluation of the person to be tested is level two, it is necessary to re-evaluate the oral ability through voice recognition.

2. The method for evaluating oral proficiency according to claim 1, wherein: In the recording environment, a microphone is set on both sides of the recording point, and the microphones are connected through the audio collection device. The ambient audio signals collected by the two microphones Stored as a time series signal.

3. The method for evaluating oral proficiency according to claim 2, wherein: Ambient audio signals Perform preprocessing to generate discrete-time signals , based on the formula: in, , is the sampling frequency, n is a positive integer, used to index discrete sampling time points, discrete time signal Used to reflect the signal strength at each discrete time point.

4. The method for evaluating oral proficiency according to claim 3, wherein: Combined with the discrete time signal, the ambient audio signal is subjected to short-time Fourier transform to generate the ambient audio frequency domain signal. , based on the formula: in, represents an imaginary number, is a window function used to limit the time range. , where N is the window size, , ambient audio frequency domain signal Used to reflect the amplitude of ambient audio at frequency f and time t.

5. The method for evaluating oral proficiency according to claim 4, wherein: Perform correlation analysis on the ambient audio frequency domain signal to generate the ambient audio power spectrum , based on the formula: Ambient audio power spectrum Used to reflect the power value of the ambient audio at the corresponding frequency f and time t. It is a modular operation.

6. The method for evaluating oral proficiency according to claim 5, wherein: Perform correlation analysis on the ambient audio power spectrum to generate ambient noise power , based on the formula: Among them, F is the maximum frequency value, T is the time for the microphone to collect the ambient audio signal, in seconds, and the ambient noise power Used to reflect the power value of environmental noise; Ambient noise power Correlation analysis is performed to generate the environmental noise impact assessment index HZP, based on the following formula: in, is the input power of the microphone, The microphone's input sensitivity weighting factor, The value range is ,The environmental noise impact assessment index HZP is used to reflect the impact of environmental noise on the recorded signal.

7. The method for evaluating oral proficiency according to claim 1, wherein: The standard spoken language set is a spoken language comparison audio, including a standard speaking speed SR. is the number of words repeated per minute by the i-th oral proficiency training subject, and the recognition accuracy The accuracy rate of word recognition output by the speech recognition device for the i-th oral ability training subject is used. The correlation analysis of the oral parameters in the repetition process is performed to generate the accuracy rate expression AR. The formula is: Substitute the spoken language parameters of all the oral proficiency training subjects collected into the above formula, substitute the speaking speed into v, and the recognition accuracy into AR, and generate the values ​​of a and b through MATLAB software fitting; Correlation analysis is performed on the accuracy expression AR to generate the speech speed accuracy evaluation index YZP based on the following formula: Where V is the speaking speed of the person under test, and the speaking speed accuracy evaluation index YZP is used to evaluate the relationship between the speaking speed of the person under test and the accuracy of the device output; A correlation analysis was conducted on the speech speed accuracy evaluation index YZP and the environmental noise impact evaluation index HZP to generate the comprehensive oral evaluation index KPZ. The formula is as follows: in, is the speech speed accuracy weight factor, The value range is The comprehensive oral evaluation index KPZ is used to reflect the accuracy of the oral ability test by comprehensively evaluating the current environment and the speaking speed of the test taker.

8. The method for evaluating oral proficiency according to claim 7, wherein: Accurate threshold The value of is 15.7, and the comprehensive index of oral evaluation KPZ is compared with the accuracy threshold To compare, when When , it means that the accuracy of the oral ability evaluation of the person under test in the current environment is level 2, and the oral ability evaluation is inaccurate; when , it indicates that the accuracy of the oral ability evaluation of the person under test in the current environment is level one, and the oral ability evaluation is accurate.

9. A system for evaluating oral proficiency, configured to execute the method for evaluating oral proficiency according to claim 1, characterized in that: include: A signal acquisition module is used to collect ambient audio signals from the recording environment; An ambient audio signal processing module is used to pre-process the ambient audio signal to generate a discrete time signal, and perform short-time Fourier transform on the ambient audio signal in combination with the discrete time signal to generate an ambient audio frequency domain signal; The frequency domain signal analysis module is used to perform correlation analysis on the ambient audio frequency domain signal to generate the ambient audio power spectrum, and to perform correlation analysis on the ambient audio power spectrum to generate the ambient noise power; Environmental noise power analysis module, used to perform correlation analysis on environmental noise power and generate an environmental noise impact assessment index; The speech rate analysis module is used to set a standard spoken language set, set 20 oral proficiency training subjects to repeat according to the standard spoken language set, collect spoken language parameters during the repetition process, perform correlation analysis on the spoken language parameters during the repetition process, generate an accuracy expression, perform correlation analysis on the accuracy expression, and generate a speech rate accuracy evaluation index; The comprehensive analysis module is used to perform correlation analysis on the speech speed accuracy evaluation index YZP and the environmental noise impact evaluation index HZP to generate the oral evaluation comprehensive index KPZ. The oral evaluation comprehensive index KPZ is used to reflect the comprehensive evaluation of the current environment and the speaking speed of the person being tested.

10. A storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the method described in claim 1 when executed by a processor.

Citation Information

Patent Citations

  • Method and system for evaluating spoken language skill

    CN101782941A

  • Oral English practicing method and system

    CN109300339A