A Spoken Language Assessment Method and System Based on Text and Speech Recognition

Through multi-dimensional scoring and speech text recognition technology, combined with reading and free statement scoring methods, the problem that the existing English speaking evaluation system cannot comprehensively evaluate users' oral ability, and improves the accuracy of the evaluation results.

CN114842875BActive Publication Date: 2025-07-08INTERNATIONAL INNOVATION CENTER OF TSINGHUA UNIVERSITY SHANGHAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210402853.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-18
Publication Date
2025-07-08
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

The existing English oral evaluation system usually only uses a single-dimensional scoring, which causes students to only practice fixed texts, lose their practical significance, and the recording cost of each oral material is high, so they cannot fully evaluate the user's oral ability.

Method used

By obtaining the user's reading voice, free statement voice and corresponding text, a multi-dimensional scoring method is used, combining the ratings of the reading and free statement parts, a comprehensive evaluation is used using voice and text recognition technology, and a secondary scoring is performed through the evaluation system of different users, adjusting the scoring weight to improve accuracy.

Benefits of technology

A comprehensive assessment of the user's oral ability is achieved, the accuracy of the evaluation results is improved, and the problem that the existing system cannot fully evaluate the content of free statements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842875B_ABST
    Figure CN114842875B_ABST
Patent Text Reader

Abstract

The present application provides a spoken language evaluation method and system based on text and speech recognition, including: respectively recognizing the text and speech of the user to be evaluated to obtain a first evaluation score and a second evaluation score; determining a first evaluation score difference and a second evaluation score difference under each evaluation dimension based on the first evaluation score and the second evaluation score; when the first evaluation score difference is greater than a first threshold score or any second evaluation score difference is greater than a second threshold score, determining a target user from the spoken language evaluation database; re-evaluating the user to be evaluated through the spoken language evaluation system used by the target user to obtain a third evaluation score and a fourth evaluation score; determining the final spoken language evaluation score of the user to be evaluated based on the first evaluation score, the second evaluation score, the third evaluation score, and the fourth evaluation score. In this way, by combining the speech and documents of the user to be evaluated for spoken language evaluation, the spoken language ability of the user to be evaluated can be evaluated more comprehensively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of speech evaluation, and in particular, to an oral English evaluation method and system based on text and speech recognition. Background Art

[0002] As a standard for oral English evaluation in language learning, with the popularization of oral English examinations, more and more schools need to use oral English training systems in daily teaching to evaluate and score the oral English pronunciation of students, so as to help students improve their oral English level. In the selection test, an oral English examination system is used, and the students' oral English examination scores are used as part of the English subject scores in the selection test.

[0003] Currently, the mainstream oral English training systems on the market usually only evaluate the oral English level of subjects using a single dimension of recognizability. Each oral evaluation material is recorded by corresponding oral personnel through standard reading, and the speech similarity between the subject and the standard recording is calculated for scoring. However, it is too costly for each oral personnel to perform a standard reading for each oral material, and the obtained oral training materials are too single, so that more and more students can only practice fixed texts, and what they learn is only "dumb English", losing practical significance. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide an oral English evaluation method and system based on text and speech recognition. By evaluating the reading speech, free statement speech, reading text, and free statement text, the oral English ability of the user to be evaluated can be more comprehensively evaluated, and through the evaluation system used by different users, the user to be evaluated can be re-evaluated, which can improve the accuracy of the evaluation results.

[0005] The embodiment of this application provides an oral English evaluation method based on text and speech recognition. The oral English evaluation method includes:

[0006] Obtain the reading speech, free statement speech, the reading text corresponding to the reading speech, and the free statement text corresponding to the free statement speech of the user to be evaluated; the reading speech is the speech of the user to be evaluated reading the standard evaluation text, and the free statement speech is the speech of the user to be evaluated making a free statement for the evaluation topic;

[0007] Based on the reading speech and the reading text, determine the first evaluation score for the reading part, and based on the free statement speech and the free statement text, determine the second evaluation score for the free statement part; both the first evaluation score and the second evaluation score are composed of evaluation sub-scores under multiple scoring dimensions; different scoring dimensions are used to represent different aspects of the oral English ability of the user to be evaluated;

[0008] Based on the first evaluation score and the second evaluation score, determine the first evaluation score difference and the second evaluation score difference under each evaluation dimension; the first evaluation score difference is the score difference between the first evaluation score and the second evaluation score;

[0009] When the first evaluation score difference is greater than the first threshold score or any second evaluation score difference is greater than the second threshold score, query the oral evaluation scores that meet the preset requirements from the oral evaluation database according to the first evaluation score and the second evaluation score, and determine the target users corresponding to the oral evaluation scores;

[0010] Use the oral evaluation system used by the target users to re-evaluate the reading voice and the free statement voice of the user to be evaluated respectively, and obtain the third evaluation score and the fourth evaluation score of the user to be evaluated; the reading voice is the voice corresponding to the reading text before the voice-text conversion, and the free statement voice is the voice corresponding to the free statement text before the voice-text conversion;

[0011] Based on the first evaluation score, the second evaluation score, the third evaluation score and the fourth evaluation score of the user to be evaluated, determine the final oral evaluation score of the user to be evaluated.

[0012] Optionally, determine the second evaluation score of the user to be evaluated and the evaluation sub-scores under each evaluation dimension included in the second evaluation score through the following steps:

[0013] Conduct a preliminary evaluation on the free statement text to determine the reference scores of each scoring paragraph included in the free statement text;

[0014] Extract evaluation features from each scoring paragraph respectively to determine the evaluation parameters of various evaluation features included in each scoring paragraph;

[0015] For each scoring paragraph, based on the evaluation parameters of various evaluation features included in the scoring paragraph, the corresponding part of the free statement voice, the initial scoring weight under each evaluation dimension, and the reference score of the scoring paragraph, determine the initial paragraph evaluation score of the scoring paragraph under each evaluation dimension;

[0016] For each scoring paragraph, adjust the initial scoring weight under the corresponding evaluation dimension respectively based on the evaluation parameters of each evaluation feature included in the scoring paragraph to determine the target scoring weight of each evaluation dimension;

[0017] For each scoring paragraph, based on the initial paragraph evaluation score of the scoring paragraph under each evaluation dimension, the initial scoring weight under each evaluation dimension, and the target scoring weight, determine the target paragraph evaluation score of the scoring paragraph under each evaluation dimension;

[0018] Based on the target paragraph evaluation scores of each scoring segment under each scoring dimension, determine the second evaluation score of the user to be evaluated and the evaluation sub-scores under each scoring dimension included in the second evaluation score.

[0019] Optionally, the scoring dimension includes at least one of the following: recognizability, tone, fluency, and intonation accuracy.

[0020] Optionally, the evaluation features include at least one of the following: the number of text events, the relevance of the answer content to the topic of the question, the number of word vectors, and the number of stressed syllables in the vocabulary.

[0021] Optionally, for each scoring segment, based on the evaluation parameters of each evaluation feature included in the scoring segment, adjust the initial scoring weight under the corresponding scoring dimension to determine the target scoring weight of each scoring dimension, including:

[0022] Based on the number of text events included in the scoring segment and the mapping relationship between the number of text events and the weight, adjust the initial scoring weight of recognizability to determine the target scoring weight of recognizability;

[0023] Based on the relevance of the answer content of the scoring segment to the topic of the question and the mapping relationship between the relevance and the weight, adjust the initial scoring weight of tone to determine the target scoring weight of tone;

[0024] Based on the number of word vectors included in the scoring segment and the mapping relationship between the number of word vectors and the weight, adjust the initial scoring weight of fluency to determine the target scoring weight of fluency;

[0025] Based on the number of stressed syllables in the vocabulary included in the scoring segment and the mapping relationship between the number of stressed syllables in the vocabulary and the weight, adjust the initial scoring weight of intonation accuracy to determine the target scoring weight of intonation accuracy.

[0026] Optionally, determine the relevance of the answer content to the topic of the question through the following steps:

[0027] Obtain the question word vector corresponding to the question text and the paragraph word vector corresponding to the scoring segment; the question text is the text obtained according to the evaluation question;

[0028] Perform clustering processing on the question word vector and the paragraph word vector respectively to obtain at least one first feature cluster corresponding to the question word vector and at least one second feature cluster corresponding to the paragraph word vector;

[0029] Extract the central vector of each first feature cluster as the first topic vector, and extract the central vector of each second feature cluster as the second topic vector;

[0030] Perform a weighted sum of all the first topic vectors to obtain the topic vector of the question, and perform a weighted sum of all the second topic vectors to obtain the topic vector of the paragraph;

[0031] Based on the topic vector of the question and the topic vector of the paragraph, determine the relevance of the answer content to the topic of the question.

[0032] Optionally, the method for querying the oral evaluation scores that meet the preset requirements from the oral evaluation database according to the first evaluation score and the second evaluation score, and determining the target users corresponding to the oral evaluation scores is as follows:

[0033] Query from the oral score database the reading oral scores whose difference between the reading part and the first evaluation score is less than the third threshold score and the difference between the test sub-scores under the same scoring dimension is less than the fourth threshold score;

[0034] Query from the oral score database the free statement oral scores whose difference between the free statement part and the second evaluation score is less than the third threshold score and the difference between the test sub-scores under the same scoring dimension is less than the fourth threshold score;

[0035] Determine the users corresponding to the retrieved reading oral scores and free statement oral scores as the target users.

[0036] Optionally, the initial evaluation of the free statement text to determine the reference scores of each scoring paragraph included in the free statement text includes:

[0037] According to the first scoring rule, based on the number of words or characters included in the free statement text, determine the initial evaluation score of the user to be evaluated;

[0038] For each scoring paragraph, according to the second scoring rule, based on the initial evaluation score and the text content of the scoring paragraph, determine the reference score of the scoring paragraph; the sum of the reference scores of all scoring paragraphs is equal to the initial evaluation score.

[0039] The embodiment of the present application also provides an oral evaluation system based on text and speech recognition. The oral evaluation system includes:

[0040] An acquisition module, configured to acquire the reading voice, free statement voice, the reading text corresponding to the reading voice, and the free statement text corresponding to the free statement voice of the user to be evaluated; the reading voice is the voice of the user to be evaluated reading the standard evaluation text, and the free statement voice is the voice of the user to be evaluated making a free statement for the evaluation question;

[0041] An identification module, configured to determine a first evaluation score for the reading part based on the reading speech and the reading text, and determine a second evaluation score for the free statement part based on the free statement speech and the free statement text; both the first evaluation score and the second evaluation score are composed of evaluation sub-scores under multiple scoring dimensions; different scoring dimensions are used to represent different aspects of the oral language ability of the user to be evaluated;

[0042] A first determination module, configured to determine a first evaluation score difference and a second evaluation score difference under each scoring dimension based on the first evaluation score and the second evaluation score; the first evaluation score difference is the score difference between the first evaluation score and the second evaluation score;

[0043] A query module, configured to, when the first evaluation score difference is greater than a first threshold score or any second evaluation score difference is greater than a second threshold score, query, according to the first evaluation score and the second evaluation score, an oral evaluation score that meets the preset requirements from an oral evaluation database, and determine the target user corresponding to the oral evaluation score;

[0044] An evaluation module, configured to re-evaluate the reading speech and the free statement speech of the user to be evaluated respectively through the oral evaluation system used by the target user, to obtain a third evaluation score and a fourth evaluation score of the user to be evaluated; the reading speech is the speech corresponding to the reading text before speech-text conversion, and the free statement speech is the speech corresponding to the free statement text before speech-text conversion;

[0045] A second determination module, configured to determine the final oral evaluation score of the user to be evaluated based on the first evaluation score, the second evaluation score, the third evaluation score, and the fourth evaluation score of the user to be evaluated.

[0046] Optionally, when the identification module is used to determine the second evaluation score of the user to be evaluated and the evaluation sub-scores under each scoring dimension included in the second evaluation score through the following steps, the identification module is used to:

[0047] Conduct a preliminary evaluation on the free statement text to determine the reference scores of each scoring paragraph included in the free statement text;

[0048] Extract evaluation features from each scoring paragraph respectively to determine the evaluation parameters of multiple evaluation features included in each scoring paragraph;

[0049] For each scoring paragraph, based on the evaluation parameters of multiple evaluation features included in the scoring paragraph, the partial free statement speech corresponding to the scoring paragraph, the initial scoring weight under each scoring dimension, and the reference score of the scoring paragraph, determine the initial paragraph evaluation score of the scoring paragraph under each scoring dimension;

[0050] For each scoring paragraph, based on the evaluation parameters of each evaluation feature included in the scoring paragraph, adjust the initial scoring weight under the corresponding scoring dimension to determine the target scoring weight for each scoring dimension;

[0051] For each scoring paragraph, based on the initial paragraph evaluation score of the scoring paragraph under each scoring dimension, the initial scoring weight under each scoring dimension, and the target scoring weight, determine the target paragraph evaluation score of the scoring paragraph under each scoring dimension;

[0052] Based on the target paragraph evaluation scores of each scoring paragraph under each scoring dimension, determine the second evaluation score of the user to be evaluated and the evaluation sub-scores under each scoring dimension included in the second evaluation score.

[0053] Optionally, the scoring dimension includes at least one of the following: recognizability, tone, fluency, and intonation accuracy.

[0054] Optionally, the evaluation feature includes at least one of the following: the number of text events, the relevance of the answer content to the topic of the question, the number of word vectors, and the number of lexical stressed syllables.

[0055] Optionally, when the recognition module is used to, for each scoring paragraph, adjust the initial scoring weight under the corresponding scoring dimension based on the evaluation parameters of each evaluation feature included in the scoring paragraph to determine the target scoring weight for each scoring dimension, the recognition module is used to:

[0056] Based on the number of text events included in the scoring paragraph and the mapping relationship between the number of text events and the weight, adjust the initial scoring weight of recognizability to determine the target scoring weight of recognizability;

[0057] Based on the relevance of the answer content of the scoring paragraph to the topic of the question and the mapping relationship between the relevance and the weight, adjust the initial scoring weight of tone to determine the target scoring weight of tone;

[0058] Based on the number of word vectors included in the scoring paragraph and the mapping relationship between the number of word vectors and the weight, adjust the initial scoring weight of fluency to determine the target scoring weight of fluency;

[0059] Based on the number of lexical stressed syllables included in the scoring paragraph and the mapping relationship between the number of lexical stressed syllables and the weight, adjust the initial scoring weight of intonation accuracy to determine the target scoring weight of intonation accuracy.

[0060] Optionally, when the recognition module is used to determine the relevance of the answer content to the topic of the question through the following steps, the recognition module is used to:

[0061] Obtain the question word vectors corresponding to the question text and the passage word vectors corresponding to the scoring passages; the question text is the text obtained according to the assessment questions;

[0062] Perform clustering processing on the question word vectors and passage word vectors respectively to obtain at least one first feature cluster corresponding to the question word vectors and at least one second feature cluster corresponding to the passage word vectors;

[0063] Extract the central vector of each first feature cluster as the first topic vector, and extract the central vector of each second feature cluster as the second topic vector;

[0064] Perform weighted summation on all the first topic vectors to obtain the question topic vector, and perform weighted summation on all the second topic vectors to obtain the passage topic vector;

[0065] Based on the question topic vector and the passage topic vector, determine the relevance between the answer content and the question topic.

[0066] Optionally, the oral assessment system further includes a third determination module, and the third determination module is used for:

[0067] Query from the oral score database the reading oral scores where the difference between the reading part and the first assessment score is less than the third threshold score and the difference between the test sub-scores in the same scoring dimension is less than the fourth threshold score;

[0068] Query from the oral score database the free statement oral scores where the difference between the free statement part and the second assessment score is less than the third threshold score and the difference between the test sub-scores in the same scoring dimension is less than the fourth threshold score;

[0069] Determine the users corresponding to the found reading oral scores and free statement oral scores as the target users.

[0070] Optionally, when the recognition module is used to conduct a preliminary assessment on the free statement text and determine the reference score of each scoring passage included in the free statement text, the recognition module is used for:

[0071] According to the first scoring rule, based on the number of words or characters included in the free statement text, determine the initial assessment score of the user to be evaluated;

[0072] For each scoring passage, according to the second scoring rule, based on the initial assessment score and the text content of this scoring passage, determine the reference score of this scoring passage; the sum of the reference scores of all scoring passages is equal to the initial assessment score.

[0073] An embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the above-mentioned oral evaluation method are executed.

[0074] An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the above-mentioned oral evaluation method are executed.

[0075] An oral evaluation method and system based on text and speech recognition provided by an embodiment of the present application, the oral evaluation method includes:

[0076] Obtain the reading voice, free statement voice, reading text corresponding to the reading voice, and free statement text corresponding to the free statement voice of the user to be evaluated; the reading voice is the voice of the user to be evaluated reading a standard evaluation text, and the free statement voice is the voice of the user to be evaluated making a free statement for an evaluation question; based on the reading voice and reading text, determine the first evaluation score for the reading part, and based on the free statement voice and free statement text, determine the second evaluation score for the free statement part; both the first evaluation score and the second evaluation score are composed of evaluation sub-scores under multiple scoring dimensions; different scoring dimensions are used to represent different aspects of the oral ability of the user to be evaluated; based on the first evaluation score and the second evaluation score, determine the first evaluation score difference and the second evaluation score difference under each scoring dimension; the first evaluation score difference is the score difference between the first evaluation score and the second evaluation score; when the first evaluation score difference is greater than the first threshold score or any second evaluation score difference is greater than the second threshold score, query the oral evaluation scores that meet the preset requirements from the oral evaluation database according to the first evaluation score and the second evaluation score, and determine the target user corresponding to the oral evaluation score; use the oral evaluation system used by the target user to re-evaluate the reading voice and free statement voice of the user to be evaluated respectively, and obtain the third evaluation score and the fourth evaluation score of the user to be evaluated; the reading voice is the voice corresponding to the reading text before voice-text conversion, and the free statement voice is the voice corresponding to the free statement text before voice-text conversion; based on the first evaluation score, second evaluation score, third evaluation score, and fourth evaluation score of the user to be evaluated, determine the final oral evaluation score of the user to be evaluated.

[0077] In this way, by evaluating the reading speech, free statement speech, reading text, and free statement text, the present application can more comprehensively evaluate the oral English ability of the user to be evaluated, and through the evaluation system used by different users to conduct a secondary evaluation of the user to be evaluated, the accuracy of the evaluation results can be improved. In addition, the present application also discloses a technical solution for multi-dimensional evaluation of free oral statements based on text recognition, which solves the problem that the existing oral evaluation system cannot score the free statement content of users.

[0078] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0080] Figure 1 It is a flowchart of an oral evaluation method based on text and speech recognition provided by an embodiment of the present application;

[0081] Figure 2 It is one of the structural diagrams of an oral evaluation system based on text and speech recognition provided by an embodiment of the present application;

[0082] Figure 3 It is another structural diagram of an oral evaluation system based on text and speech recognition provided by an embodiment of the present application;

[0083] Figure 4 It is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0084] To make the objects, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all of them. Usually, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by those of ordinary skill in the art without creative efforts belongs to the scope of protection of the present application.

[0085] As the voice evaluation scoring serves as the oral evaluation standard for language learning, with the popularization of oral English tests, more and more schools need to use an oral English training system in their daily teaching to evaluate and score the oral English pronunciation of students, so as to help students improve their oral English level. In the selection test, when using the oral English test system, the oral English test scores of students are used as part of the English subject scores in the selection test.

[0086] Currently, the mainstream oral English training systems on the market usually only evaluate the oral English level of subjects using a single dimension of recognizability. Moreover, each oral evaluation material is recorded by corresponding oral personnel through standard reading, and the speech similarity between the subject and the standard recording is calculated for scoring. However, it is too costly to require oral personnel to perform a standard reading for each oral material, and the obtained oral training materials are too single, which makes it so that more and more students can only practice fixed texts, and what they learn is only "dumb English", losing its practical significance.

[0087] Based on this, an oral evaluation method and system based on text and speech recognition can more comprehensively evaluate the oral English ability of the user to be evaluated by evaluating the reading speech, free statement speech, reading text, and free statement text, and can perform a secondary evaluation of the user to be evaluated through the evaluation systems used by different users, which can improve the accuracy of the evaluation results.

[0088] Please refer to Figure 1 , Figure 1 which is a flowchart of an oral evaluation method based on text and speech recognition provided by an embodiment of the present application. As shown Figure 1 in it, the oral evaluation method provided by the embodiment of the present application includes:

[0089] S101. Obtain the reading speech of the user to be evaluated, the free statement speech, the reading text corresponding to the reading speech, and the free statement text corresponding to the free statement speech; the reading speech is the speech of the user to be evaluated reading the standard evaluation text, and the free statement speech is the speech of the user to be evaluated making a free statement regarding the evaluation question.

[0090] When conducting an oral evaluation, it generally includes two parts. One part is the standard evaluation text provided during the reading test (i.e., the reading part), and the other part is the free statement regarding relevant content for the evaluation question (i.e., the free statement part).

[0091] For the reading part, obtain the reading speech and the reading text corresponding to the reading speech; for the free statement part, obtain the free statement speech and the free statement text corresponding to the free statement speech.

[0092] Here, when performing speech-text conversion, it can be converted through a pre-trained speech-text conversion model, or an existing speech-text conversion tool can be used for conversion, which is not limited here.

[0093] In this way, by obtaining two different texts and speeches, and performing subsequent text and speech recognition to obtain the oral score of the user to be evaluated, the oral expression ability of the user to be evaluated can be evaluated more comprehensively.

[0094] S102. Based on the reading speech and the reading text, determine the first evaluation score for the reading part, and based on the free statement speech and the free statement text, determine the second evaluation score for the free statement part; both the first evaluation score and the second evaluation score are composed of evaluation sub-scores under multiple scoring dimensions; different scoring dimensions are used to represent different aspects of the oral ability of the user to be evaluated.

[0095] Here, the first evaluation score is the score obtained by pre-scoring the reading speech and then adjusting the score through recognizing the reading text. The second evaluation score is the score obtained by pre-scoring the free statement speech and then adjusting the score through recognizing the free statement text. The number and type of the scoring dimensions corresponding to the first evaluation score and the second evaluation score are the same, but the evaluation sub-scores under each scoring dimension can be the same or different.

[0096] The first evaluation score represents the score of the reading part of the user to be evaluated, and the second evaluation score represents the score of the free statement part of the user to be evaluated.

[0097] In an implementation manner of the present application, the second evaluation score of the user to be evaluated and the evaluation sub-scores under each evaluation dimension included in the second evaluation score are determined through the following steps: conduct a preliminary evaluation on the free statement text to determine the reference scores of each scoring paragraph included in the free statement text; extract evaluation features for each scoring paragraph respectively to determine the evaluation parameters of various evaluation features included in each scoring paragraph; for each scoring paragraph, based on the evaluation parameters of various evaluation features included in this scoring paragraph, the corresponding part of the free statement speech, the initial scoring weight under each evaluation dimension, and the reference score of this scoring paragraph, determine the initial paragraph evaluation score of this scoring paragraph under each evaluation dimension; for each scoring paragraph, respectively adjust the initial scoring weight under the corresponding evaluation dimension based on the evaluation parameters of each evaluation feature included in this scoring paragraph to determine the target scoring weight of each evaluation dimension; for each scoring paragraph, based on the initial paragraph evaluation score of this scoring paragraph under each evaluation dimension, the initial scoring weight under each evaluation dimension, and the target scoring weight, determine the target paragraph evaluation score of this scoring paragraph under each evaluation dimension; based on the target paragraph evaluation scores of each scoring paragraph under each evaluation dimension, determine the second evaluation score of the user to be evaluated and the evaluation sub-scores under each evaluation dimension included in the second evaluation score.

[0098] In another implementation manner provided by the present application, the conduct of a preliminary evaluation on the free statement text to determine the respective reference scores of each scoring paragraph included in the free statement text includes: according to the first scoring rule, based on the number of words or characters included in the free statement text, determine the initial evaluation score of the user to be evaluated; for each scoring paragraph, according to the second scoring rule, based on the initial evaluation score and the text content of this scoring paragraph, determine the reference score of this scoring paragraph; the sum of the reference scores of all scoring paragraphs is equal to the initial evaluation score.

[0099] Here, by conducting a preliminary evaluation on the free statement text, the number of words or characters included in the free statement text can be determined; different numbers of words or different numbers of characters correspond to different initial evaluation scores in the first scoring rule; the free statement text includes at least one scoring paragraph, and among them, the specific scoring paragraphs included in the free statement text can be determined by identifying specific characters (such as full stops) included in the free statement text.

[0100] For example, assume that the number of words included in the recognized free statement text is 260 words. For a text of 100 - 200 words as stipulated in the first scoring rule, an initial evaluation score of 60 points (full score) is assigned; for a text of 201 - 300 words, an initial evaluation score of 80 points (full score) is assigned; for a text of more than 300 words, an initial evaluation score of 100 points (full score) is assigned. Therefore, the initial evaluation score of the user to be evaluated determined according to the first scoring rule is 80 points.

[0101] In the second scoring rule, it is stipulated that based on the initial evaluation score determined from the free statement text according to the respective space proportion of each scoring paragraph or the relevance of the paragraph to the test questions, the reference score of each scoring paragraph is determined.

[0102] To better understand the second scoring rule, it is illustrated by the following example. When it is stipulated in the second scoring rule that the reference score of the scoring paragraph is determined according to the space proportion, assume that the initial evaluation score of the free statement text is 80 points. The free statement text includes three paragraphs: scoring paragraph 1, scoring paragraph 2, and scoring paragraph 3. Among them, the number of words included in scoring paragraph 1 accounts for 20% of the total number of words in the free statement text, then the reference score of scoring paragraph 1 is determined to be 16 points (full score); the number of words included in scoring paragraph 2 accounts for 30% of the total number of words in the free statement text, then the reference score of scoring paragraph 2 is determined to be 24 points (full score); the number of words included in scoring paragraph 3 accounts for 50% of the total number of words in the free statement text, then the reference score of scoring paragraph 3 is determined to be 40 points (full score).

[0103] After determining the reference score of each scoring paragraph, in order to determine the actual score of each scoring paragraph (i.e., the initial paragraph evaluation score), for each scoring paragraph, it is necessary to extract evaluation features of the paragraph, determine various evaluation features included in the scoring paragraph and the evaluation parameters of each evaluation feature, and at the same time, determine the corresponding part of the free statement voice of the scoring paragraph. The evaluation features may include at least one of the following: the number of text events, the relevance of the answer content to the topic of the question, the number of word vectors, and the number of lexical stress syllables.

[0104] Among them, the number of text events refers to the number of text events included in the scoring paragraph. The text events included in each scoring paragraph can be determined through a text event extraction model, so as to determine the number of text events.

[0105] In another implementation provided by the present application, the relevance between the answer content and the topic of the question is determined through the following steps: obtaining the question word vector corresponding to the question text and the paragraph word vector corresponding to the scoring paragraph; the question text is the text obtained according to the evaluation question; clustering the question word vector and the paragraph word vector respectively to obtain at least one first feature cluster corresponding to the question word vector and at least one second feature cluster corresponding to the paragraph word vector; extracting the central vector of each first feature cluster as the first topic vector, and extracting the central vector of each second feature cluster as the second topic vector; performing weighted summation on all the first topic vectors to obtain the question topic vector, and performing weighted summation on all the second topic vectors to obtain the paragraph topic vector; determining the relevance between the answer content and the question topic based on the question topic vector and the paragraph topic vector.

[0106] Here, the question word vector corresponding to the question text and the paragraph word vector corresponding to the scoring paragraph can be obtained through a word vector extraction model; wherein, before obtaining the paragraph word vector, the answer word vector of the free statement text can be obtained first, so as to determine the paragraph word vector corresponding to each scoring paragraph based on the answer word vector; for the word vector extraction model, the word2vec model can be preferably used.

[0107] For clustering the question word vector and the paragraph word vector respectively, the KMeans method can be used to cluster the word vectors to determine at least one first feature cluster corresponding to the question word vector and the central vector of each first feature cluster; determining at least one second feature cluster corresponding to the paragraph word vector and the central vector of each second feature cluster. Thus, after performing weighted summation processing, the question topic vector and the paragraph topic vector can be obtained, and finally, based on the calculation of vector similarity, the relevance between the answer content in the scoring paragraph and the question topic can be determined.

[0108] Among them, after obtaining the paragraph word vector of the scoring paragraph, the number of word vectors and the number of lexical stress syllables included in the scoring paragraph can also be determined.

[0109] When determining the initial paragraph evaluation score of the scoring paragraph under each scoring dimension, specifically, it can be: first, based on the initial scoring weight under each scoring dimension corresponding to the scoring paragraph and the reference score of the scoring paragraph, determine the reference score of the scoring paragraph under each scoring dimension, and then respectively based on the evaluation parameters of the evaluation features included in the paragraph, the part of the free statement speech corresponding to the scoring paragraph, and the preset paragraph scoring rule, determine the initial paragraph evaluation score of the scoring paragraph under each scoring dimension.

[0110] Here, the scoring dimensions include at least one of the following: recognizability, tone, fluency, and pitch accuracy. It should be noted that there is a one-to-one correspondence between the scoring dimensions and the evaluation features. The number of text events corresponds to recognizability, the relevance of the answer content to the topic of the question corresponds to tone, the number of word vectors corresponds to fluency, and the number of lexical stressed syllables corresponds to pitch accuracy. In this way, through the specific evaluation parameters of the number of text events, the reference score and the initial paragraph evaluation score of this assigned paragraph under recognizability can be determined. The evaluation scores corresponding to other scoring dimensions are similar and will not be elaborated here.

[0111] To better understand how to determine the initial paragraph evaluation score of a certain assigned paragraph under each scoring dimension, the following example is used for illustration. Assume that the reference score of this paragraph is 24 points, and the initial scoring weights corresponding to each scoring dimension of recognizability, tone, fluency, and pitch accuracy are all 0.25. The number of text events included in this assigned paragraph is 5, the relevance of the answer content to the topic of the question is 80%, the number of word vectors is 78, and the number of lexical stressed syllables is 10. Based on the initial scoring weights of each scoring dimension corresponding to this assigned paragraph and the reference score of this assigned paragraph, it is determined that the reference score of this assigned paragraph under each scoring dimension is 6 points. Then, according to the paragraph scoring rule, based on the number of text events included in this paragraph being 5, the initial paragraph evaluation score under recognizability is determined to be 4 points. Based on the relevance of the answer content included in this paragraph to the topic of the question being 80%, the initial paragraph evaluation score under tone is determined to be 5 points. Based on the number of word vectors included in this paragraph being 78, the initial paragraph evaluation score under fluency is determined to be 6 points. Based on the number of lexical stressed syllables included in this paragraph being 10, the initial paragraph evaluation score under pitch accuracy is determined to be 4 points. In this way, the initial paragraph evaluation score of this assigned paragraph under each scoring dimension is also determined. Among them, the paragraph scoring rule stipulates the corresponding relationship between the evaluation parameters of each evaluation feature and the ratio of the initial paragraph evaluation score and the reference score under the corresponding scoring dimension.

[0112] In addition, the initial paragraph evaluation score can also be determined by using a pre-trained scoring model. Input the partial free-form statement speech corresponding to this assigned paragraph into the scoring model, and the initial paragraph evaluation score of this assigned paragraph is output by the scoring model. Then multiply it by the initial scoring weights under each scoring dimension, that is, the initial paragraph evaluation score of this assigned paragraph under each scoring dimension is determined.

[0113] After determining the initial paragraph evaluation scores of the scored paragraphs under each scoring dimension, it is also necessary to adjust the initial scoring weights under each scoring dimension. In another implementation manner provided in the present application, for each scored paragraph, based on the evaluation parameters of each evaluation feature included in the scored paragraph, the initial scoring weight under the corresponding scoring dimension is adjusted to determine the target scoring weight of each scoring dimension, including: based on the number of text events included in the scored paragraph and the mapping relationship between the number of text events and the weight, adjusting the initial scoring weight of recognizability to determine the target scoring weight of recognizability; based on the relevance between the answer content included in the scored paragraph and the topic of the question and the mapping relationship between the relevance and the weight, adjusting the initial scoring weight of tone to determine the target scoring weight of tone; based on the number of word vectors included in the scored paragraph and the mapping relationship between the number of word vectors and the weight, adjusting the initial scoring weight of fluency to determine the target scoring weight of fluency; based on the number of lexical stressed syllables included in the scored paragraph and the mapping relationship between the number of lexical stressed syllables and the weight, adjusting the initial scoring weight of intonation accuracy to determine the target scoring weight of intonation accuracy.

[0114] Here, the mapping relationship between the number of text events and the weight defines the relationship between the specific evaluation parameters of the number of text events and the corresponding target weight or weight adjustment parameter. The other three mapping relationships are similar to the mapping relationship between the number of text events and the weight, and will not be elaborated here.

[0115] To better understand the method of adjusting the initial scoring weight, continue with the above example to illustrate the process of adjusting the initial scoring weight. Here, taking the adjustment of the initial scoring weight of recognizability as an example, the initial scoring weight corresponding to recognizability is 0.25. The mapping relationship between the number of text events and the weight stipulates that when the number of text events is 4 - 6, the weight is reduced by 0.05. The determined number of text events is 5, so the initial scoring weight corresponding to recognizability is adjusted from 0.25 to the target scoring weight of 0.2. The adjustment processes of the initial scoring weights of other scoring dimensions are similar to that of recognizability and will not be elaborated here.

[0116] After determining the initial paragraph evaluation scores of the scored paragraphs under each scoring dimension, the initial scoring weights under each scoring dimension, and the target scoring weights, for the target paragraph evaluation scores of the scored paragraphs under each scoring dimension, the specific determination method can be: for the initial paragraph evaluation scores of each scored paragraph under each scoring dimension, divide the initial paragraph evaluation score by the initial scoring weight under the corresponding dimension, and then multiply by the target scoring weight under the corresponding dimension. The score obtained at this time is the target paragraph evaluation score of the scored paragraph under this scoring dimension.

[0117] In this way, after determining the target paragraph evaluation scores of each assigned paragraph under each scoring dimension, the score obtained by adding up the target paragraph evaluation scores of each assigned paragraph under the same dimension is the sub-evaluation score under this scoring dimension included in the second evaluation score. Adding up the target paragraph evaluation scores of all assigned paragraphs under each scoring dimension can obtain the second evaluation score.

[0118] In addition, it should be noted that the determination method of the first evaluation score corresponding to the reading part is similar to that of the second evaluation score of the free statement part. It also performs operations such as scoring from four scoring dimensions and modifying weights. Therefore, it will not be elaborated here.

[0119] S103. Based on the first evaluation score and the second evaluation score, determine the first evaluation score difference and the second evaluation score difference under each scoring dimension; the first evaluation score difference is the score difference between the first evaluation score and the second evaluation score.

[0120] Here, the first evaluation score difference can be determined by subtracting the second evaluation score from the first evaluation score. The second evaluation score difference under each scoring dimension can be determined by the following method: subtracting the sub-evaluation score included in the second evaluation score from the sub-evaluation score included in the first evaluation score under the same scoring dimension. Among them, if the obtained score difference is negative, an absolute value change can be performed, and the score after taking the absolute value is determined as the obtained first evaluation score difference or second evaluation score difference.

[0121] S104. When the first evaluation score difference is greater than the first threshold score or any second evaluation score difference is greater than the second threshold score, query the oral evaluation scores that meet the preset requirements from the oral evaluation database according to the first evaluation score and the second evaluation score, and determine the target user corresponding to the oral evaluation score.

[0122] Here, the specific values of the first threshold score and the second threshold score can be selected according to applicability.

[0123] Among them, when the first evaluation score difference is not greater than the first threshold score and any second evaluation score difference is not greater than the second threshold score, the total score of the first evaluation score and the second evaluation score of the user to be evaluated can be determined as the final oral evaluation score of the user to be evaluated.

[0124] Exemplarily, the preset requirement is that the proportion deviation of each sub-evaluation score of the first evaluation score is less than or equal to a preset similarity threshold (for example, it can be set to 10%. If not retrieved, it can be gradually set to 15%, 20%); or the preset requirement is that the proportion deviation of each sub-evaluation score of the second evaluation score is less than or equal to a preset similarity threshold (for example, it can be set to 10%. If not retrieved, it can be gradually set to 15%, 20%);

[0125] In an implementation provided by this application, querying the spoken language assessment scores that meet the preset requirements from the spoken language assessment database according to the first assessment score and the second assessment score, and determining the target users corresponding to the spoken language assessment scores includes: querying from the spoken language score database the reading spoken language scores whose difference from the first assessment score in the reading part is less than the third threshold score and the difference in the test sub-scores under the same scoring dimension is less than the fourth threshold score; querying from the spoken language score database the free statement spoken language scores whose difference from the second assessment score in the free statement part is less than the third threshold score and the difference in the test sub-scores under the same scoring dimension is less than the fourth threshold score; and determining the users corresponding to the found reading spoken language scores and free statement spoken language scores respectively as the target users.

[0126] Here, the spoken language score database stores the spoken language assessment scores of multiple users in different regions. Each user's spoken language assessment score includes the assessment score of the reading part, the assessment sub-scores of each scoring dimension in the reading part, the assessment score of the free statement part, and the assessment sub-scores of each scoring dimension in the free statement part.

[0127] Among them, when determining the target users, the target users corresponding to the reading part and the target users corresponding to the free statement part are determined respectively, and the number of target users determined for each part is at least one.

[0128] Exemplarily, when determining the target users corresponding to the reading part, specifically: traverse the assessment scores (reading spoken language scores) corresponding to the reading part in the spoken language score database. When there are reading spoken language scores whose difference from the first assessment score is less than the third threshold score, the difference in the test sub-scores under recognizability is less than the fourth threshold score, the difference in the test sub-scores under fluency is less than the fourth threshold score, the difference in the test sub-scores under tone is less than the fourth threshold score, and the difference in the test sub-scores under pronunciation accuracy is less than the fourth threshold score, determine the user corresponding to the reading spoken language score as the target user. The method of determining the target users for the free statement part is similar to that of the reading part and will not be elaborated here.

[0129] S105. Re-score the reading voice and the free statement voice of the user to be assessed respectively through the spoken language assessment system used by the target user, and obtain the third assessment score and the fourth assessment score of the user to be assessed; the reading voice is the voice corresponding to the reading text before the voice-text conversion, and the free statement voice is the voice corresponding to the free statement text before the voice-text conversion.

[0130] It should be noted that the reason for re-evaluating the reading speech and free statement speech of the user to be evaluated through the oral evaluation system used by the target user is that there may be slight differences in setting the initial scoring weights corresponding to different scoring dimensions in the oral evaluation systems used in different regions. Depending on the equipment in each place, perhaps in one place the microphone is good and the acquisition environment is good, so the examiner will set higher requirements for voice-text conversion, or the acquisition environment is noisy, the equipment is old, then the accuracy of voice-text conversion will be set lower, such as increasing text error correction, or setting certain voices to be automatically recognized as correct, etc. There are many ways of such calibration, so there will also be some differences in the oral evaluation systems used by different users to be evaluated. In addition, in different places, the same piece of speech may be converted into different texts, resulting in different final oral evaluation scores.

[0131] Here, re-score the reading speech of the user to be evaluated through the oral evaluation system used by the target user corresponding to the reading oral score to determine the third evaluation score of the user to be evaluated; re-score the free statement speech of the user to be evaluated through the oral evaluation system used by the target user corresponding to the free statement oral score to determine the fourth evaluation score of the user to be evaluated.

[0132] It should also be noted that by transmitting the voice information with abnormal scores to the evaluation systems in different regions for re-evaluation and mutual verification respectively, the present application can improve the accuracy and objectivity of the voice evaluation system.

[0133] S106. Determine the final oral evaluation score of the user to be evaluated based on the first evaluation score, the second evaluation score, the third evaluation score and the fourth evaluation score of the user to be evaluated.

[0134] Here, determining the oral evaluation score of the user to be evaluated based on the first evaluation score, the second evaluation score, the third evaluation score and the fourth evaluation score of the user to be evaluated can be: taking the average of the two evaluation scores as the final oral evaluation score of the user to be evaluated. That is, perform mean processing on the first evaluation score and the third evaluation score to determine the fifth evaluation score, perform mean processing on the second evaluation score and the fourth evaluation score to determine the sixth evaluation score, and determine the score after adding the fifth evaluation score and the sixth evaluation score as the final oral evaluation score of the user to be evaluated.

[0135] In addition, a difference threshold for the difference between the first evaluation result and the second evaluation result can also be set. When the difference between the scores of the two evaluation results is less than the preset difference threshold, select the average of the first evaluation result and the second evaluation result as the final evaluation result. When the difference between the scores of the two evaluation results is greater than or equal to the preset difference threshold, select the highest score as the final evaluation result.

[0136] In this way, by evaluating the reading speech, free statement speech, reading text, and free statement text, the present application can more comprehensively evaluate the oral ability of the user to be evaluated. And by using different evaluation systems used by different users to conduct a secondary evaluation of the user to be evaluated, the accuracy of the evaluation results can be improved. In addition, the present application also discloses a technical solution for multi-dimensional evaluation of free oral statements based on text recognition, which solves the problem that the existing oral evaluation system cannot score the free statement content of users.

[0137] Please refer to Figure 2 、 Figure 3 , Figure 2 which is one of the structural schematic diagrams of an oral evaluation system based on text and speech recognition provided by an embodiment of the present application. Figure 3 which is the second structural schematic diagram of an oral evaluation system based on text and speech recognition provided by an embodiment of the present application. As Figure 2 shown in

[0138] The oral evaluation system 200 includes:

[0139] An acquisition module 210, configured to acquire the reading speech, free statement speech, reading text corresponding to the reading speech, and free statement text corresponding to the free statement speech of the user to be evaluated; the reading speech is the speech of the user to be evaluated reading a standard evaluation text, and the free statement speech is the speech of the user to be evaluated making a free statement for an evaluation question;

[0140] An identification module 220, configured to determine a first evaluation score for the reading part based on the reading speech and reading text, and determine a second evaluation score for the free statement part based on the free statement speech and free statement text; both the first evaluation score and the second evaluation score are composed of evaluation sub-scores under multiple scoring dimensions; different scoring dimensions are used to represent different aspects of the oral ability of the user to be evaluated;

[0141] A first determination module 230, configured to determine a first evaluation score difference and a second evaluation score difference under each scoring dimension based on the first evaluation score and the second evaluation score; the first evaluation score difference is the score difference between the first evaluation score and the second evaluation score;

[0142] The evaluation module 250 is used to re-evaluate the reading speech and the free statement speech of the user to be evaluated respectively through the oral evaluation system used by the target user, and obtain the third evaluation score and the fourth evaluation score of the user to be evaluated; the reading speech is the speech corresponding to the reading text before the speech-text conversion, and the free statement speech is the speech corresponding to the free statement text before the speech-text conversion;

[0143] The second determination module 260 is used to determine the final oral evaluation score of the user to be evaluated based on the first evaluation score, the second evaluation score, the third evaluation score, and the fourth evaluation score of the user to be evaluated.

[0144] Optionally, when the recognition module 220 is used to determine the second evaluation score of the user to be evaluated and the evaluation sub-scores under each evaluation dimension included in the second evaluation score through the following steps, the recognition module 220 is used to:

[0145] Extract evaluation features from each scoring paragraph respectively, and determine the evaluation parameters of various evaluation features included in each scoring paragraph;

[0146] For each scoring paragraph, based on the evaluation parameters of various evaluation features included in the scoring paragraph, the corresponding part of the free statement speech of the scoring paragraph, the initial scoring weight under each evaluation dimension, and the reference score of the scoring paragraph, determine the initial paragraph evaluation score of the scoring paragraph under each evaluation dimension;

[0147] For each scoring paragraph, based on the evaluation parameters of each evaluation feature included in the scoring paragraph, adjust the initial scoring weight under the corresponding evaluation dimension respectively, and determine the target scoring weight of each evaluation dimension;

[0148] For each scoring paragraph, based on the initial paragraph evaluation score of the scoring paragraph under each evaluation dimension, the initial scoring weight under each evaluation dimension, and the target scoring weight, determine the target paragraph evaluation score of the scoring paragraph under each evaluation dimension;

[0149] Based on the target paragraph evaluation scores of each scoring paragraph under each evaluation dimension, determine the second evaluation score of the user to be evaluated and the evaluation sub-scores under each evaluation dimension included in the second evaluation score.

[0150] Optionally, the evaluation dimension includes at least one of the following: recognizability, tone, fluency, and pitch accuracy.

[0151] Optionally, the evaluation feature includes at least one of the following: the number of text events, the relevance of the answer content to the topic of the question, the number of word vectors, and the number of lexical stress syllables.

[0152] Optionally, when the recognition module 220 is used to adjust the initial scoring weights in each corresponding scoring dimension respectively based on the evaluation parameters of each evaluation feature included in the scoring paragraph, and determine the target scoring weights for each scoring dimension, the recognition module 220 is used to:

[0153] Based on the number of text events included in the scoring paragraph and the mapping relationship between the number of text events and weights, adjust the initial scoring weight of recognizability, and determine the target scoring weight of recognizability;

[0154] Based on the relevance between the answer content of the scoring paragraph and the topic of the question and the mapping relationship between the relevance and weights, adjust the initial scoring weight of tone, and determine the target scoring weight of tone;

[0155] Based on the number of word vectors included in the scoring paragraph and the mapping relationship between the number of word vectors and weights, adjust the initial scoring weight of fluency, and determine the target scoring weight of fluency;

[0156] Based on the number of lexical stressed syllables included in the scoring paragraph and the mapping relationship between the number of lexical stressed syllables and weights, adjust the initial scoring weight of intonation accuracy, and determine the target scoring weight of intonation accuracy.

[0157] Optionally, when the recognition module 220 is used to determine the relevance between the answer content and the topic of the question through the following steps, the recognition module 220 is used to:

[0158] Obtain the question word vectors corresponding to the question text and the paragraph word vectors corresponding to the scoring paragraph; the question text is the text obtained according to the evaluation question;

[0159] Perform clustering processing on the question word vectors and the paragraph word vectors respectively to obtain at least one first feature cluster corresponding to the question word vectors and at least one second feature cluster corresponding to the paragraph word vectors;

[0160] Extract the central vector of each first feature cluster as the first topic vector, and extract the central vector of each second feature cluster as the second topic vector;

[0161] Perform weighted summation on all the first topic vectors to obtain the question topic vector, and perform weighted summation on all the second topic vectors to obtain the paragraph topic vector;

[0162] Based on the question topic vector and the paragraph topic vector, determine the relevance between the answer content and the topic of the question.

[0163] Optionally, as Figure 3 shown, the oral evaluation system 200 further includes a third determination module 270, and the third determination module 270 is used to:

[0164] Query the reading and speaking scores in the speaking test score database where the difference between the reading part and the first evaluation score is less than the third threshold score, and the difference between the test sub-scores in the same scoring dimension is less than the fourth threshold score;

[0165] Query the free statement speaking scores in the speaking test score database where the difference between the free statement part and the second evaluation score is less than the third threshold score, and the difference between the test sub-scores in the same scoring dimension is less than the fourth threshold score;

[0166] Determine the users corresponding to the found reading and speaking scores and free statement speaking scores as the target users.

[0167] Optionally, when the recognition module 220 is used to conduct a preliminary evaluation of the free statement text and determine the reference score for each scoring paragraph included in the free statement text, the recognition module 220 is used to:

[0168] According to the first scoring rule, based on the number of words or characters included in the free statement text, determine the initial evaluation score of the user to be evaluated;

[0169] For each scoring paragraph, according to the second scoring rule, based on the initial evaluation score and the text content of this scoring paragraph, determine the reference score of this scoring paragraph; the sum of the reference scores of all scoring paragraphs is equal to the initial evaluation score.

[0170] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 4 shown in, the electronic device 400 includes a processor 410, a memory 420, and a bus 430.

[0171] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 runs, the processor 410 communicates with the memory 420 through the bus 430. When the machine-readable instructions are executed by the processor 410, the steps of the speaking test method in the above Figure 1 as shown method embodiments can be executed. The specific implementation manner can refer to the method embodiments and will not be elaborated here.

[0172] An embodiment of the present application also provides a computer-readable storage medium. A computer program is stored on this computer-readable storage medium. When the computer program is run by a processor, the steps of the speaking test method in the above Figure 1 as shown method embodiments can be executed. The specific implementation manner can refer to the method embodiments and will not be elaborated here.

[0173] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0174] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0175] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0176] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0177] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0178] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the technical field can still modify the technical solutions recorded in the foregoing embodiments or easily conceive of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A spoken language assessment method based on text and speech recognition, characterized in that, The oral evaluation method includes: Obtaining the reading voice of the user to be evaluated, the free statement voice, the reading text corresponding to the reading voice, and the free statement text corresponding to the free statement voice; the reading voice is the voice of the user to be evaluated reading the standard evaluation text, and the free statement voice is the voice of the user to be evaluated making a free statement for the evaluation questions; Based on the reading voice and the reading text, determining the first evaluation score for the reading part, and based on the free statement voice and the free statement text, determining the second evaluation score for the free statement part; both the first evaluation score and the second evaluation score are composed of evaluation sub-scores under multiple scoring dimensions; different scoring dimensions are used to represent different aspects of the oral ability of the user to be evaluated; Based on the first evaluation score and the second evaluation score, determining the first evaluation score difference and the second evaluation score difference under each scoring dimension; the first evaluation score difference is the score difference between the first evaluation score and the second evaluation score; When the first evaluation score difference is greater than the first threshold score or any second evaluation score difference is greater than the second threshold score, querying the oral evaluation scores that meet the preset requirements from the oral evaluation database according to the first evaluation score and the second evaluation score, and determining the target user corresponding to the oral evaluation score; Re-scoring the reading voice and the free statement voice of the user to be evaluated respectively through the oral evaluation system used by the target user, and obtaining the third evaluation score and the fourth evaluation score of the user to be evaluated; the reading voice is the voice corresponding to the reading text before voice-text conversion, and the free statement voice is the voice corresponding to the free statement text before voice-text conversion; Based on the first evaluation score, the second evaluation score, the third evaluation score, and the fourth evaluation score of the user to be evaluated, determining the final oral evaluation score of the user to be evaluated.

2. The oral evaluation method according to claim 1, wherein Determining the second evaluation score of the user to be evaluated and the evaluation sub-scores under each scoring dimension included in the second evaluation score through the following steps: Conducting a preliminary evaluation on the free statement text to determine the reference scores of each scoring paragraph included in the free statement text; Extracting evaluation features for each scoring paragraph respectively to determine the evaluation parameters of various evaluation features included in each scoring paragraph; For each scoring paragraph, based on the evaluation parameters of various evaluation features included in this scoring paragraph, the partial free statement voice corresponding to this scoring paragraph, the initial scoring weight under each scoring dimension, and the reference score of this scoring paragraph, determining the initial paragraph evaluation score of this scoring paragraph under each scoring dimension; For each scoring paragraph, respectively adjusting the initial scoring weight under the corresponding scoring dimension based on the evaluation parameters of each evaluation feature included in this scoring paragraph to determine the target scoring weight under each scoring dimension; For each scoring paragraph, based on the initial paragraph evaluation score of this scoring paragraph under each scoring dimension, the initial scoring weight under each scoring dimension, and the target scoring weight, determining the target paragraph evaluation score of this scoring paragraph under each scoring dimension; Based on the target paragraph evaluation scores of each assigned paragraph under each scoring dimension, determine the second evaluation score of the user to be evaluated and the evaluation sub-scores under each scoring dimension included in the second evaluation score.

3. The oral evaluation method according to claim 2, characterized in that, The scoring dimension includes at least one of the following: recognizability, tone, fluency, and intonation accuracy.

4. The oral evaluation method according to claim 3, wherein The evaluation features include at least one of the following: the number of text events, the relevance of the answer content to the topic of the question, the number of word vectors, and the number of lexical stressed syllables.

5. The oral evaluation method according to claim 4, characterized in that For each assigned paragraph, based on the evaluation parameters of each evaluation feature included in the assigned paragraph, adjust the initial scoring weights under the corresponding scoring dimension to determine the target scoring weights for each scoring dimension, including: Based on the number of text events included in the assigned paragraph and the mapping relationship between the number of text events and the weights, adjust the initial scoring weight of recognizability to determine the target scoring weight of recognizability; Based on the relevance of the answer content of the assigned paragraph to the topic of the question and the mapping relationship between the relevance and the weights, adjust the initial scoring weight of tone to determine the target scoring weight of tone; Based on the number of word vectors included in the assigned paragraph and the mapping relationship between the number of word vectors and the weights, adjust the initial scoring weight of fluency to determine the target scoring weight of fluency; Based on the number of lexical stressed syllables included in the assigned paragraph and the mapping relationship between the number of lexical stressed syllables and the weights, adjust the initial scoring weight of intonation accuracy to determine the target scoring weight of intonation accuracy.

6. The oral evaluation method according to claim 5, characterized in that, Determine the relevance of the answer content to the topic of the question through the following steps: Obtain the question word vectors corresponding to the question text and the paragraph word vectors corresponding to the assigned paragraph; the question text is the text obtained according to the evaluation question; Perform clustering processing on the question word vectors and the paragraph word vectors respectively to obtain at least one first feature cluster corresponding to the question word vectors and at least one second feature cluster corresponding to the paragraph word vectors; Extract the central vector of each first feature cluster as the first topic vector, and extract the central vector of each second feature cluster as the second topic vector; Perform weighted summation on all the first topic vectors to obtain the question topic vector, and perform weighted summation on all the second topic vectors to obtain the paragraph topic vector; Based on the question topic vector and the paragraph topic vector, determine the relevance of the answer content to the topic of the question.

7. The oral evaluation method according to claim 1, wherein The method of querying the oral evaluation scores that meet the preset requirements from the oral evaluation database according to the first evaluation score and the second evaluation score and determining the target users corresponding to the oral evaluation scores includes: Query from the oral score database the reading oral scores whose difference from the first evaluation score in the reading part is less than the third threshold score and the difference in the test sub-scores under the same scoring dimension is less than the fourth threshold score; Query from the oral score database the free statement oral scores whose difference from the second evaluation score in the free statement part is less than the third threshold score and the difference in the test sub-scores under the same scoring dimension is less than the fourth threshold score; Determine the users corresponding to the found reading oral scores and free statement oral scores as the target users.

8. The oral evaluation method according to claim 2, characterized in that The initial evaluation of the free statement text is performed to determine the reference score of each scoring paragraph included in the free statement text, including: According to the first scoring rule, based on the number of words or characters included in the free statement text, the initial evaluation score of the user to be evaluated is determined; For each scoring paragraph, according to the second scoring rule, based on the initial evaluation score and the text content of the scoring paragraph, the reference score of the scoring paragraph is determined; the sum of the reference scores of all scoring paragraphs is equal to the initial evaluation score.

9. A spoken language assessment system based on text and speech recognition, characterized in that, The oral evaluation system includes: An acquisition module for acquiring the reading voice, free statement voice, the reading text corresponding to the reading voice, and the free statement text corresponding to the free statement voice of the user to be evaluated; the reading voice is the voice of the user to be evaluated reading the standard evaluation text, and the free statement voice is the voice of the user to be evaluated making a free statement for the evaluation question; An identification module for determining the first evaluation score of the reading part based on the reading voice and the reading text, and determining the second evaluation score of the free statement part based on the free statement voice and the free statement text; both the first evaluation score and the second evaluation score are composed of evaluation sub-scores under multiple scoring dimensions; different scoring dimensions are used to represent different aspects of the oral language ability of the user to be evaluated; A first determination module for determining the first evaluation score difference and the second evaluation score difference under each scoring dimension based on the first evaluation score and the second evaluation score; the first evaluation score difference is the score difference between the first evaluation score and the second evaluation score; A query module for, when the first evaluation score difference is greater than the first threshold score or any second evaluation score difference is greater than the second threshold score, querying the oral evaluation score that meets the preset requirements from the oral evaluation database according to the first evaluation score and the second evaluation score, and determining the target user corresponding to the oral evaluation score; An evaluation module for re-scoring the reading voice and the free statement voice of the user to be evaluated respectively through the oral evaluation system used by the target user to obtain the third evaluation score and the fourth evaluation score of the user to be evaluated; the reading voice is the voice corresponding to the reading text before the voice-text conversion, and the free statement voice is the voice corresponding to the free statement text before the voice-text conversion; A second determination module for determining the final oral evaluation score of the user to be evaluated based on the first evaluation score, the second evaluation score, the third evaluation score, and the fourth evaluation score of the user to be evaluated.

10. An electronic device, characterized in that, Including: A processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are run by the processor, the steps of the oral evaluation method according to any one of claims 1 to 8 are executed.

11. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by the processor, the steps of the oral evaluation method according to any one of claims 1 to 8 are executed.

Citation Information

Patent Citations

  • Oral retelling marking method and system

    CN108428382A

  • Methods for measuring speech intelligibility, and related systems and apparatus

    US20210225389A1