Calibration data set acquisition method, spoken language test scoring method and related device
By acquiring abnormal voice data and building a target calibration data set with the initial calibration data set, the problem that the calibration data set in the prior art fails to cover abnormal situations is solved, and a more accurate oral examination score is achieved.
Patent Information
- Application Number
- CN202510017099.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-13
AI Technical Summary
The existing calibration data set fails to effectively cover the abnormal situation, resulting in the oral examination scoring system being unable to accurately score when answering abnormalities.
By obtaining exception voice data and combining the initial calibration data set, the target calibration data set is constructed to ensure that the calibration data set covers abnormal situations.
The coverage of the calibration data set is enhanced, so that the training-derived oral examination scoring system can handle the answer scores of abnormal situations more accurately.
Smart Images

Figure CN119993200A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a calibration data set acquisition method, an oral test scoring method and related devices. Background Art
[0002] Since there are differences in the scoring scales of teachers or experts in oral exams in different regions, experts or teachers are needed to score a part of the answer voice data to obtain a calibration data set, and then use the calibration data set to train the oral exam scoring system so that the scores given by the oral exam scoring system for the answer voice data obtained by voice answers to the oral exam are as close as possible to the scores given by experts or teachers.
[0003] However, if the calibration data set does not take abnormal situations into account, that is, when a part of the answer voice data scored by experts or teachers does not contain abnormal situations, the oral test scoring system trained based on the calibration data set will not be able to accurately score when encountering answer voice data with abnormal situations in the future, such as marking high scores as low and low scores as high. Summary of the invention
[0004] The main technical problem solved by the present application is to provide a calibration data set acquisition method, an oral test scoring method and related devices, which can make the calibration data set cover a wider range.
[0005] In order to solve the above technical problems, a technical solution adopted in the present application is: to provide a calibration data set acquisition method, the method comprising: acquiring at least one abnormal voice data; wherein the abnormal voice data is the first answer voice data with an abnormal sound event, and the first answer voice data is obtained by the respondent making a voice answer for an oral test; using at least one abnormal voice data and an initial calibration data set to obtain a target calibration data set; wherein the target calibration data set is used to train the oral test scoring system.
[0006] The abnormal voice data is first answering voice data in which an abnormal sound event occurs and the answer is abnormal.
[0007] Among them, before obtaining at least one abnormal voice data, the calibration data set acquisition method also includes: obtaining a number of first answer voice data; recognizing the number of first answer voice data to obtain a voice recognition result of each first answer voice data; wherein the voice recognition result includes abnormal sound events existing in the first answer voice data; and treating the first answer voice data that has abnormal sound events and is an abnormal answer as abnormal voice data.
[0008] Among them, the speech recognition result also includes speech recognition text, and the language corresponding to the oral test is the target language; the method for determining abnormal answers includes: for the speech recognition text of each first answer speech data, in response to the speech recognition text satisfying the abnormal condition, determining that the first answer speech data belongs to an abnormal answer; wherein the abnormal condition includes at least one of the following: there is no text content corresponding to the target language in the speech recognition text, and the accuracy of the speech recognition text is lower than a preset threshold.
[0009] Among them, abnormal conditions include that the accuracy of the speech recognition text is lower than a preset threshold, and the characterization value of the speech recognition text accuracy includes the hit rate of the speech recognition text. The hit rate of the speech recognition text is the ratio between the number of first answer words and the number of second answer words. The number of first answer words is the number of words that exist in both the speech recognition text and the target answer text, and the number of second answer words is the number of all words in the target answer text. The target answer text is the answer text of the oral test corresponding to the first answer voice data.
[0010] Among them, the abnormal sound event includes answering by others; before the speech recognition text of each first answer voice data, in response to the speech recognition text satisfying the abnormal condition, determining that the first answer voice data belongs to an abnormal answer, the calibration data set acquisition method also includes: for each first answer voice data, in response to the abnormal sound event existing in the first answer voice data, answering for others; from the speech recognition text of the first answer voice data, extracting the recognition text belonging to the corresponding answerer, and using the recognition text as the speech recognition text corresponding to the first answer voice data.
[0011] Among them, before recognizing a number of first answer voice data and obtaining the voice recognition results of each first answer voice data, the calibration data set acquisition method also includes: for each first answer voice data, performing feature extraction on the first answer voice data to obtain initial acoustic features; performing pre-emphasis processing on the initial acoustic features to obtain target acoustic features; recognizing a number of first answer voice data to obtain the voice recognition results of each first answer voice data, including: recognizing the target acoustic features of each first answer voice data to obtain the voice recognition results of each first answer voice data.
[0012] Wherein, obtaining at least one abnormal voice data includes: obtaining at least one first answer voice data corresponding to each type of abnormal sound event as at least one abnormal voice data.
[0013] Among them, before obtaining at least one first answer voice data corresponding to each type of abnormal sound event as at least one abnormal voice data, the calibration data set acquisition method also includes: obtaining the first answer voice data for at least two different types of abnormal sound events as the second answer voice data; using the second answer voice data as the first answer voice data corresponding to the target type abnormal sound event, wherein the target type abnormal sound event is the abnormal sound event category that appears in a preset order in the second answer voice data.
[0014] The method of obtaining a target calibration data set by using at least one abnormal speech data and an initial calibration data set includes: obtaining the score of the target object on the abnormal speech data to obtain abnormal calibration data; and combining the abnormal calibration data and the initial calibration data set to obtain the target calibration data set.
[0015] To solve the above technical problems, another technical solution adopted in the present application is: to provide an oral test scoring method, the method comprising: obtaining voice data to be scored; wherein the voice data to be scored is obtained by the target respondent giving a voice answer to the target oral test; using an oral test scoring system to score the voice data to be scored to obtain a scoring result; wherein the oral test scoring system is trained using a calibration data set, and the calibration data set is obtained using the above calibration data set acquisition method.
[0016] In order to solve the above technical problems, another technical solution adopted by the present application is: to provide a calibration data set acquisition device, which includes an acquisition module and a calibration module; the acquisition module is used to acquire at least one abnormal voice data; wherein the abnormal voice data is the first answer voice data with an abnormal sound event, and the first answer voice data is obtained by the respondent's voice answer for the oral test; the calibration module is used to use at least one abnormal voice data and an initial calibration data set to obtain a target calibration data set; wherein the target calibration data set is used to train the oral test scoring system.
[0017] In order to solve the above technical problems, another technical solution adopted by the present application is: to provide an oral test scoring device, which includes an acquisition module and a scoring module; the acquisition module is used to obtain voice data to be scored; wherein the voice data to be scored is obtained by the target respondent giving a voice answer to the target oral test; the scoring module is used to score the voice data to be scored using an oral test scoring system to obtain a scoring result; wherein the oral test scoring system is trained using a calibration data set, and the calibration data set is obtained using the above-mentioned calibration data set acquisition method.
[0018] In order to solve the above technical problems, another technical solution adopted in the present application is: to provide an electronic device, the electronic device includes a processor and a memory, the memory stores program instructions, and the processor is used to execute the program instructions to implement the above method.
[0019] In order to solve the above technical problem, another technical solution adopted by the present application is: providing a computer-readable storage medium, which is used to store program instructions, and the program instructions can be executed to implement the above method.
[0020] The above technical solution uses at least one abnormal speech data and an initial calibration data set to obtain a target calibration data set. Therefore, the initial calibration data set is supplemented with at least one abnormal speech data, so that the target calibration data set obtained after the supplementation covers the calibration data corresponding to the abnormal speech data, that is, the target calibration data set obtained after the supplementation has a more complete coverage.
[0021] Furthermore, the target calibration data set is subsequently used to train the oral test scoring system, so that the trained oral test scoring system can perform oral test scoring more accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flow chart of an embodiment of a calibration data set acquisition method provided by the present application;
[0023] Figure 2 is a flow chart of another embodiment of the calibration data set acquisition method provided by the present application;
[0024] Figure 3 It is a flowchart of an embodiment of a method for scoring an oral test provided by the present application;
[0025] Figure 4 It is a structural schematic diagram of an embodiment of a calibration data set acquisition device provided by the present application;
[0026] Figure 5 It is a structural schematic diagram of an embodiment of an oral test scoring device provided by the present application;
[0027] Figure 6 It is a structural schematic diagram of an embodiment of an electronic device provided by the present application;
[0028] Figure 7 It is a structural schematic diagram of an embodiment of a computer-readable storage medium provided by the present application. DETAILED DESCRIPTION
[0029] The scheme of the embodiment of the present application is described in detail below in conjunction with the drawings of the specification.
[0030] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.
[0031] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the objects associated before and after are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of, for example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C.
[0032] See also Figure 1 , Figure 1 is a flow chart of an embodiment of a calibration data set acquisition method provided in this application. It should be noted that if there are substantially the same results, this embodiment does not Figure 1 The process sequence shown is limited. Figure 1 As shown, this embodiment includes:
[0033] Step S11: Acquire at least one abnormal voice data.
[0034] Since there are differences in the scoring scales of teachers or experts in oral exams in different regions, experts or teachers are needed to score a part of the answer voice data to obtain a calibration data set, and then use the calibration data set to train the oral exam scoring system so that the scores given by the oral exam scoring system for the answer voice data obtained by voice answers to the oral exam are as close as possible to the scores given by experts or teachers.
[0035] However, if the calibration data set does not take into account abnormal situations, that is, when a part of the answer voice data scored by the expert or teacher does not have abnormal situations, the oral test scoring system trained based on the calibration data set will be unable to accurately score when encountering answer voice data with abnormal situations, such as scoring a high score as a low score and a low score as a high score. It should be noted that the existence of abnormal situations in the answer voice data can be abnormal sound events (such as coughing, blowing, etc.) in the answer voice data.
[0036] Therefore, in this embodiment, at least one abnormal voice data is obtained; wherein the abnormal voice data is the first answer voice data with abnormal sound events, and the first answer voice data is obtained by the respondent answering the oral test by voice. By obtaining at least one abnormal voice data, the initial calibration data set is supplemented, so that the target calibration data set obtained after the supplementation is covered with the calibration data corresponding to the abnormal voice data, that is, the target calibration data set obtained after the supplementation is more fully covered. Furthermore, the target calibration data set is subsequently used to train the oral test scoring system, so that the trained oral test scoring system can perform oral test scoring more accurately.
[0037] There is no limitation on the language involved in the oral test, and it can be set according to actual use needs. For example, if the oral test is an English oral test, then the language involved in the oral test is English. For another example, if the oral test is a Chinese oral test, then the language involved in the oral test is Chinese. In addition, there is no limitation on the respondent, and it can be set according to actual use needs. For example, the respondent is a teacher or a student. Furthermore, a recording receiving device, such as a head-mounted microphone, can be used to receive the respondent's voice answer to the oral test.
[0038] It should be noted that at least one abnormal voice data obtained is the first answer voice data with an abnormal sound event. The oral tests corresponding to different first answer voice data can be the same oral test or different oral tests, which is not limited here.
[0039] In one embodiment, the abnormal sound event may be blowing, coughing, knocking on the table, touching the microphone, incomprehensible pronunciation, poor pronunciation appearance, answering by others, or poor sound quality. It should be noted that, in the case where the abnormal sound event is blowing, there is a blowing sound in the first answer voice data. In the case where the abnormal sound event is coughing, there is a coughing sound in the first answer voice data. In the case where the abnormal sound event is knocking on the table, there is a knocking sound in the first answer voice data. In the case where the abnormal sound event is touching the microphone, there is a touching microphone sound in the first answer voice data. In the case where the abnormal sound event is incomprehensible pronunciation, there is completely incomprehensible voice data in the first answer voice data. In the case where the abnormal sound event is poor pronunciation appearance, there is voice data in the first answer voice data whose pronunciation is not particularly accurate. In the case where the abnormal sound event is poor sound quality, there is voice data in the first answer voice data with large background noise, large background voices, etc. In the case where the abnormal sound event is answered by others, there is voice data of other answerers in the first answer voice data.
[0040] In one embodiment, the abnormal voice data is the first answer voice data with abnormal sound events and abnormal answers. In other words, the first answer voice data with abnormal sound events and abnormal answers is taken as the abnormal voice data.
[0041] In one embodiment, obtaining at least one abnormal voice data is specifically: obtaining at least one first answer voice data corresponding to each type of abnormal sound event as at least one abnormal voice data. In other words, at least one first answer voice data corresponding to each type of abnormal sound event is selected as the acquired abnormal voice data, so that the target calibration data set constructed subsequently covers a wider and more complete range of abnormal voice data types. Furthermore, the target calibration data set is subsequently used to train the oral test scoring system, so that the trained oral test scoring system can perform oral test scoring more accurately.
[0042] There is no limit on the number of first answer voice data corresponding to each type of abnormal sound event obtained, and it can be specifically set according to actual use needs.
[0043] For example, taking the four categories of abnormal sound events including blowing, coughing, other people's answers, and touching the microphone as examples: obtain at least one first answer voice data corresponding to the abnormal sound event-blowing, at least one first answer voice data corresponding to the abnormal sound event-coughing, at least one first answer voice data corresponding to the abnormal sound event-other people's answers, and at least one first answer voice data corresponding to the abnormal sound event-touching the microphone as the acquired abnormal voice data.
[0044] Of course, in other implementations, at least one first answer voice data corresponding to some types of abnormal sound events may also be obtained, or only at least one first answer voice data corresponding to one type of abnormal sound event may be obtained as the acquired abnormal voice data.
[0045] In a specific embodiment, before obtaining at least one first answer voice data corresponding to each type of abnormal sound event as at least one abnormal voice data, first answer voice data for at least two different types of abnormal sound events are also obtained as second answer voice data; the second answer voice data is used as the first answer voice data corresponding to the target type abnormal sound event, wherein the target type abnormal sound event is the abnormal sound event category that appears in a preset order in the second answer voice data.
[0046] That is to say, when there are two types of abnormal sound events in a first answering voice data, the first answering voice data is classified as answering voice data corresponding to the types of abnormal sound events that appear in a preset order.
[0047] The preset order may be the first, second, last, etc., and is not limited here.
[0048] For example, taking the preset order as the first time: the first answering voice data A contains abnormal sound events - coughing, abnormal sound event - blowing and abnormal sound event - touching the microphone. Since the first abnormal sound event that occurs is blowing, the first answering voice data A is classified as the answering voice data corresponding to the abnormal sound event - blowing.
[0049] Step S12: using at least one abnormal speech data and an initial calibration data set to obtain a target calibration data set.
[0050] In this embodiment, at least one abnormal speech data and an initial calibration data set are used to obtain a target calibration data set. That is, at least one abnormal speech data is used to supplement the initial calibration data set, so that the target calibration data set obtained after the supplementation covers the calibration data corresponding to the abnormal speech data, that is, the target calibration data set obtained after the supplementation covers more completely. Furthermore, the target calibration data set is subsequently used to train the oral test scoring system, so that the trained oral test scoring system can perform oral test scoring more accurately.
[0051] It should be noted that the initial calibration data set is obtained by using a conventional calibration data set selection method. For example, a general oral test scoring system is first used to perform machine score prediction on a number of answer voice data; then, according to the machine score distribution, a certain amount of data is randomly selected in each score segment to form an initial calibration data set. The initial calibration data set may include calibration data corresponding to abnormal voice data, or it may not include calibration data corresponding to abnormal voice data. The target calibration data set obtained by using at least one abnormal voice data and the initial calibration data set must include calibration data corresponding to abnormal voice data; therefore, the target calibration data set is obtained by using at least one abnormal voice data and the initial calibration data set, and the initial calibration data set can be supplemented, so that the target calibration data set obtained after the supplementation is covered with the calibration data corresponding to the abnormal voice data, that is, the target calibration data set obtained after the supplementation is more fully covered. Furthermore, the target calibration data set is subsequently used to train the oral test scoring system, so that the trained oral test scoring system can perform oral test scoring more accurately.
[0052] In one embodiment, a target calibration data set is obtained by using at least one abnormal speech data and an initial calibration data set, specifically: obtaining the score of the target object on the abnormal speech data to obtain abnormal calibration data; combining the abnormal calibration data with the initial calibration data set to obtain the target calibration data set. The calibration data is the answering speech data with existing scores, so in order to obtain the target calibration data set by combining with the initial calibration data set, it is necessary to first obtain the score of the target object on the abnormal speech data, that is, it is necessary to first obtain the abnormal calibration data corresponding to the abnormal speech data.
[0053] See also Figure 2 , Figure 2 FIG. 1 is a flow chart of another embodiment of the calibration data set acquisition method provided in the present application. It should be noted that if there are substantially the same results, this embodiment does not necessarily refer to the calibration data set acquisition method. Figure 2 The process sequence shown is limited. Figure 2 As shown, the abnormal voice data is the first answer voice data with abnormal sound events and abnormal answers. Before obtaining at least one abnormal voice data, the following sub-steps are also included:
[0054] Step S21: Acquire some first answering voice data.
[0055] In this implementation, a number of first answering voice data are obtained.
[0056] In one implementation, the first answering voice data may be obtained from local storage or cloud storage. Of course, in other implementations, the first answering voice data may also be obtained in real time, which is not limited here.
[0057] Step S22: Recognize a plurality of first answering voice data to obtain a voice recognition result of each first answering voice data.
[0058] In this embodiment, a number of first answer voice data are recognized to obtain voice recognition results of each first answer voice data; wherein the voice recognition results include abnormal sound events present in the first answer voice data. In other words, by recognizing a number of first answer voice data, it is possible to determine abnormal sound events present in each first answer voice data. It should be noted that not all of the first answer voice data have abnormal sound events, so the recognition results corresponding to some of the first answer voice data may also be that there are no abnormal voice events.
[0059] In one embodiment, the VADFree model can be used to recognize a number of first answer voice data to obtain a voice recognition result of each first answer voice data. The VADFree model is different from the VAD model. The conventional VAD model aims to distinguish the voiced segments, which will cause some abnormal sound events involved in this application (such as blowing, coughing, etc.) to be recognized as silent segments, and will cause other abnormal sound events involved in this application (such as others answering) to be recognized as voiced segments.
[0060] In one embodiment, before recognizing a plurality of first answer voice data and obtaining a voice recognition result of each first answer voice data, for each first answer voice data, feature extraction is performed on the first answer voice data to obtain an initial acoustic feature; the initial acoustic feature is pre-emphasized to obtain a target acoustic feature. At this time, recognizing a plurality of first answer voice data and obtaining a voice recognition result of each first answer voice data is specifically: recognizing the target acoustic feature of each first answer voice data to obtain a voice recognition result of each first answer voice data.
[0061] That is to say, the speech recognition result of the first answering speech data is obtained by performing recognition based on the target acoustic features of the first answering speech data. The target acoustic features are obtained by performing pre-emphasis processing based on the initial acoustic features corresponding to the first answering speech data. The pre-emphasis processing can enhance the high frequency part of the initial acoustic features and reduce spectrum distortion.
[0062] Of course, in other implementations, the initial acoustic features of each first answer voice data may be directly recognized to obtain the voice recognition results of each first answer voice data, which is not limited here.
[0063] In a specific implementation, the acoustic features of the first answering speech data are spectrum features, generally using FilterBank features.
[0064] In one embodiment, the speech recognition result also includes the speech recognition text corresponding to the first answering speech data and the timestamp (eg, start time, end time) corresponding to the abnormal sound event.
[0065] Step S23: The first answering voice data having an abnormal sound event and being an abnormal answer is regarded as abnormal voice data.
[0066] In this embodiment, the first answer voice data that contains an abnormal sound event and is an abnormal answer is regarded as the abnormal voice data.
[0067] In one embodiment, the speech recognition result also includes a speech recognition text, the language corresponding to the oral test is the target language, and the abnormal answer determination step specifically includes: for the speech recognition text of each first answer speech data, in response to the speech recognition text satisfying the abnormal condition, determining that the first answer speech data belongs to an abnormal answer; wherein the abnormal condition includes at least one of the following: there is no text content corresponding to the target language in the speech recognition text, and the accuracy of the speech recognition text is lower than a preset threshold. wherein, the size of the preset threshold is not limited, and can be specifically set according to actual use.
[0068] The abnormal condition is that the speech recognition text does not contain text content in the target language. In an oral test in the target language, the speech recognition text corresponding to the first answer speech data must include text content in the target language; and if the speech recognition text corresponding to the first answer speech data does not include text content in the target language, then the first answer speech data must be the answer speech data under the abnormal answer of the respondent.
[0069] The abnormal condition is that the accuracy of the speech recognition text is lower than the preset threshold. If the accuracy of the speech recognition text corresponding to the first answer voice data is lower than the preset threshold, it means that the respondent has not fully mastered the oral test knowledge, and the first answer voice data must be the answer voice data under the respondent's abnormal answer.
[0070] In a specific implementation, the abnormal condition includes that the accuracy of the speech recognition text is lower than a preset threshold, and the representation value of the accuracy of the speech recognition text includes the hit rate of the speech recognition text, and the hit rate of the speech recognition text is the ratio between the number of first answer words and the number of second answer words, the number of first answer words is the number of words existing in both the speech recognition text and the target answer text, the number of second answer words is the number of all words in the target answer text, and the target answer text is the answer text of the oral test corresponding to the answer voice data. In other words, the answer hit rate of the speech recognition text corresponding to the first answer voice data is used as the accuracy of the speech recognition text corresponding to the first answer voice data.
[0071] In a specific embodiment, the abnormal sound event includes an answer from another person; before determining that the first answer voice data is an abnormal answer in response to the voice recognition text satisfying the abnormal condition for the voice recognition text of each first answer voice data, for each first answer voice data, in response to the abnormal sound event present in the first answer voice data, an answer is made for another person; from the voice recognition text of the first answer voice data, a recognition text belonging to the corresponding answerer is extracted, and the recognition text is used as the voice recognition text corresponding to the first answer voice data.
[0072] That is to say, in the case where an abnormal sound event exists in a first answer voice data, it is not determined whether the first answer voice data is an abnormal answer based on the overall voice recognition text corresponding to the first answer voice data, but the voice recognition text belonging to the answerer in the overall voice recognition text corresponding to the first answer voice data is extracted, and based on the voice recognition text belonging to the answerer, it is determined whether the first answer voice data is an abnormal answer.
[0073] See also Figure 3 , Figure 3 is a flow chart of an embodiment of the oral test scoring method provided by the present application. It should be noted that if there are substantially the same results, this embodiment does not Figure 3 The process sequence shown is limited. Figure 3 As shown, this embodiment includes:
[0074] Step S31: Acquire speech data to be rated.
[0075] Acquire the speech data to be scored; wherein the speech data to be scored is obtained by the target respondent performing speech answers for the target oral test.
[0076] The target respondent may be a teacher or a student, etc., and there is no limitation here.
[0077] Step S32: using the oral test scoring system to score the speech data to be scored, and obtaining a scoring result.
[0078] In this implementation, the oral test scoring system is used to score the speech data to be scored to obtain a scoring result; wherein the oral test scoring system is trained using a calibration data set, and the calibration data set is obtained using the calibration data set acquisition method described above.
[0079] The calibration data set used for training the oral test scoring system is obtained by using the calibration data set acquisition method described above; therefore, the calibration data set used for training the oral test scoring system is calibration data corresponding to the abnormal speech data, that is, a calibration data set with more complete data coverage. The oral test scoring system is trained using the calibration data set, so the oral test scoring system can score the scoring speech data more accurately.
[0080] See also Figure 4 , Figure 4It is a structural schematic diagram of an embodiment of a calibration data set acquisition device provided by the present application. The calibration data set acquisition device 40 includes an acquisition module 41 and a calibration module 42. The acquisition module 41 is used to acquire at least one abnormal voice data; wherein the abnormal voice data is the first answer voice data with an abnormal sound event, and the first answer voice data is obtained by the respondent's voice answer for the oral test; the calibration module 42 is used to use at least one abnormal voice data and an initial calibration data set to obtain a target calibration data set; wherein the target calibration data set is used to train the oral test scoring system.
[0081] The abnormal voice data is the first answer voice data in which the abnormal sound event occurs and the abnormal answer occurs.
[0082] Among them, the acquisition module is also used to, before acquiring at least one abnormal voice data, include: acquiring several first answer voice data; identifying several first answer voice data to obtain voice recognition results of each first answer voice data; wherein the voice recognition results include abnormal sound events existing in the first answer voice data; and treating the first answer voice data with abnormal sound events and being an abnormal answer as abnormal voice data.
[0083] Among them, the above-mentioned speech recognition results also include speech recognition text, and the language corresponding to the oral test is the target language; the method for determining abnormal answers includes: for the speech recognition text of each first answer speech data, in response to the speech recognition text satisfying the abnormal condition, determining that the first answer speech data belongs to an abnormal answer; wherein the abnormal condition includes at least one of the following: there is no text content corresponding to the target language in the speech recognition text, and the accuracy of the speech recognition text is lower than a preset threshold.
[0084] Among them, the above-mentioned abnormal conditions include that the accuracy of the speech recognition text is lower than a preset threshold, and the characterization value of the speech recognition text accuracy includes the hit rate of the speech recognition text. The hit rate of the speech recognition text is the ratio between the number of first answer words and the number of second answer words. The number of first answer words is the number of words that exist in both the speech recognition text and the target answer text, and the number of second answer words is the number of all words in the target answer text. The target answer text is the answer text of the oral test corresponding to the first answer voice data.
[0085] Among them, the abnormal sound event includes answers from others; the acquisition module 41 is also used for, for the voice recognition text of each first answer voice data, in response to the voice recognition text satisfying the abnormal condition, before determining that the first answer voice data belongs to an abnormal answer, including: for each first answer voice data, in response to the abnormal sound event existing in the first answer voice data, answering for others; from the voice recognition text of the first answer voice data, extracting the recognition text belonging to the corresponding answerer, and using the recognition text as the voice recognition text corresponding to the first answer voice data.
[0086] Among them, the acquisition module 41 is also used to recognize a number of first answer voice data and obtain the voice recognition results of each first answer voice data, including: for each first answer voice data, performing feature extraction on the first answer voice data to obtain initial acoustic features; performing pre-emphasis processing on the initial acoustic features to obtain target acoustic features; the acquisition module 41 is used to recognize a number of first answer voice data and obtain the voice recognition results of each first answer voice data, including: recognizing the target acoustic features of each first answer voice data to obtain the voice recognition results of each first answer voice data.
[0087] The acquisition module 41 is used to acquire at least one abnormal voice data, including: acquiring at least one first answer voice data corresponding to each type of abnormal sound event as at least one abnormal voice data.
[0088] Among them, the acquisition module 41 is also used to obtain at least one first answer voice data corresponding to each type of abnormal sound event as at least one abnormal voice data, including: obtaining the first answer voice data of at least two different types of abnormal sound events as the second answer voice data; using the second answer voice data as the first answer voice data corresponding to the target type abnormal sound event, wherein the target type abnormal sound event is the abnormal sound event category that appears in a preset order in the second answer voice data.
[0089] Among them, the calibration module 42 is used to use at least one abnormal speech data and an initial calibration data set to obtain a target calibration data set, including: obtaining the target object's score on the abnormal speech data to obtain abnormal calibration data; combining the abnormal calibration data and the initial calibration data set to obtain the target calibration data set.
[0090] See also Figure 5 , Figure 51 is a schematic diagram of the structure of an embodiment of an oral test scoring device provided by the present application. The oral test scoring device 50 includes an acquisition module 51 and a scoring module 52; the acquisition module 51 is used to acquire speech data to be scored; wherein the speech data to be scored is obtained by the target respondent performing a speech answer for the target oral test; the scoring module 52 is used to score the speech data to be scored using the oral test scoring system to obtain a scoring result; wherein the oral test scoring system is obtained by training with a calibration data set, and the calibration data set is obtained using the calibration data set acquisition method as described above.
[0091] See also Figure 6 , Figure 6 6 is a schematic diagram of the structure of an embodiment of an electronic device provided by the present application. The electronic device 60 includes a memory 61 and a processor 62 coupled to each other, and the processor 62 is used to execute program instructions stored in the memory 61 to implement the steps of any of the above-mentioned calibration data set acquisition methods and / or oral test scoring method embodiments. In a specific implementation scenario, the electronic device 60 may include but is not limited to: a microcomputer, a server, and in addition, the electronic device 60 may also include a mobile device such as a laptop computer and a tablet computer, which is not limited here.
[0092] Specifically, the processor 62 is used to control itself and the memory 61 to implement the steps of any of the above-mentioned calibration data set acquisition methods and / or oral test scoring method embodiments. The processor 62 can also be called a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 62 can be implemented by an integrated circuit chip.
[0093] See also Figure 7 , Figure 7It is a structural diagram of an embodiment of a computer-readable storage medium provided by the present application. The computer-readable storage medium 70 of the embodiment of the present application stores a program instruction 71, and when the program instruction 71 is executed, the method provided by any embodiment of the calibration data set acquisition method and / or the oral test scoring method of the present application and any non-conflicting combination is implemented. Among them, the program instruction 71 can form a program file and be stored in the above-mentioned computer-readable storage medium 70 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) executes all or part of the steps of the various implementation methods of the present application. The aforementioned computer-readable storage medium 70 includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk, or a terminal device such as a computer, a server, a mobile phone, and a tablet.
[0094] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0095] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A calibration data set acquisition method, characterized in that: The method comprises: Acquire at least one abnormal voice data; wherein the abnormal voice data is first answer voice data with abnormal sound events, and the first answer voice data is obtained by the respondent giving a voice answer to an oral test; The at least one abnormal speech data and the initial calibration data set are used to obtain a target calibration data set; wherein the target calibration data set is used to train an oral test scoring system.
2. The method according to claim 1, characterized in that The abnormal voice data is first answering voice data in which the abnormal sound event occurs and the answer is abnormal.
3. The method according to claim 2, characterized in that Before acquiring at least one abnormal voice data, the method further includes: Acquire some first answer voice data; Recognize the plurality of first answering voice data to obtain voice recognition results of each of the first answering voice data; wherein the voice recognition results include abnormal sound events present in the first answering voice data; The first answering voice data which contains the abnormal sound event and is an abnormal answer is used as the abnormal voice data.
4. The method according to claim 3, characterized in that The speech recognition result also includes speech recognition text, and the language corresponding to the oral test is the target language; the method for determining the abnormal answer includes: For the speech recognition text of each of the first answer voice data, in response to the speech recognition text satisfying an abnormal condition, it is determined that the first answer voice data belongs to an abnormal answer; wherein the abnormal condition includes at least one of the following: there is no text content corresponding to the target language in the speech recognition text, and the accuracy of the speech recognition text is lower than a preset threshold.
5. The method according to claim 4, characterized in that The abnormal condition includes that the accuracy of the speech recognition text is lower than a preset threshold value, and the characterization value of the speech recognition text accuracy includes the hit rate of the speech recognition text, and the hit rate of the speech recognition text is the ratio between the number of first answer words and the number of second answer words, the first number of answer words is the number of words existing in both the speech recognition text and the target answer text, the second number of answer words is the number of all words in the target answer text, and the target answer text is the answer text of the oral test corresponding to the first answer voice data.
6. The method according to claim 4, characterized in that The abnormal sound event includes another person answering; before the speech recognition text of each of the first answer voice data is generated, in response to the speech recognition text satisfying the abnormal condition, determining that the first answer voice data belongs to an abnormal answer, the method further includes: For each of the first answering voice data, answering for the other person in response to an abnormal sound event present in the first answering voice data; From the speech recognition text of the first answering speech data, the recognition text corresponding to the answerer is extracted, and the recognition text is used as the speech recognition text corresponding to the first answering speech data.
7. The method according to claim 3, characterized in that Before recognizing the plurality of first answering voice data to obtain voice recognition results of each of the first answering voice data, the method further includes: For each of the first answering voice data, extracting features of the first answering voice data to obtain initial acoustic features; Performing pre-emphasis processing on the initial acoustic feature to obtain a target acoustic feature; The recognizing the plurality of first answering voice data to obtain a voice recognition result of each of the first answering voice data includes: The target acoustic features of each of the first answering voice data are recognized to obtain a voice recognition result of each of the first answering voice data.
8. The method according to claim 1, characterized in that The obtaining of at least one abnormal voice data comprises: At least one first answer voice data corresponding to each type of abnormal sound event is obtained as the at least one abnormal voice data.
9. The method according to claim 8, characterized in that Before obtaining at least one first answer voice data corresponding to each type of abnormal sound event as the at least one abnormal voice data, the method further includes: Acquire first answering voice data in which at least two different types of abnormal sound events exist as second answering voice data; The second answering voice data is used as the first answering voice data corresponding to the target class abnormal sound event, wherein the target class abnormal sound event is the abnormal sound event category that appears in a preset order in the second answering voice data.
10. The method according to claim 1, characterized in that The step of obtaining a target calibration data set by using the at least one abnormal speech data and the initial calibration data set includes: Obtaining the score of the target object on the abnormal speech data to obtain abnormal calibration data; The abnormal calibration data and the initial calibration data set are combined to obtain the target calibration data set.
11. A method for scoring an oral test, characterized in that: The method comprises: Acquire speech data to be scored; wherein the speech data to be scored is obtained by the target respondent performing speech answers for the target oral test; The speech data to be scored is scored using an oral test scoring system to obtain a scoring result; wherein the oral test scoring system is trained using a calibration data set, and the calibration data set is obtained using the calibration data set acquisition method according to any one of claims 1 to 10.
12. A calibration data set acquisition device, characterized in that: The device comprises: An acquisition module, configured to acquire at least one abnormal voice data; wherein the abnormal voice data is first answer voice data with abnormal sound events, and the first answer voice data is obtained by the respondent giving a voice answer to an oral test; The calibration module is used to obtain a target calibration data set using the at least one abnormal speech data and the initial calibration data set; wherein the target calibration data set is used to train the oral test scoring system.
13. An oral test scoring device, characterized in that: The device comprises: An acquisition module is used to acquire speech data to be scored; wherein the speech data to be scored is obtained by the target respondent performing speech answers for the target oral test; A scoring module is used to score the speech data to be scored using an oral test scoring system to obtain a scoring result; wherein the oral test scoring system is trained using a calibration data set, and the calibration data set is obtained using the calibration data set acquisition method according to any one of claims 1 to 10.
14. An electronic device, characterized in that: The electronic device comprises a processor and a memory, wherein the memory stores program instructions, and the processor is configured to execute the program instructions to implement the method according to any one of claims 1 to 11.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program instructions, and the program instructions can be executed to implement the method according to any one of claims 1-11.
Citation Information
Patent Citations
Pronunciation quality assessment and error detection method based on fusion of multiple characteristics and multiple systems
CN101727903A
Method and system capable of detecting oral test cheating
CN103065642A
Scoring method and training method for spoken language questions and answers, computer equipment and storage medium
CN114360537A
Voice evaluation method and device, electronic equipment and storage medium
CN115295020A
Chinese speech recognition text error correction method and device based on semi-supervised mode
CN117809656A
Cited By
Oral answer detection method, device and equipment and storage medium
CN121054030A