Data processing method, system and apparatus
By employing a multi-dimensional scoring model based on the CAF framework in intelligent language assessment, the initial speech data is preprocessed and analyzed using data processing modules across multiple dimensions. This solves the problem of weak interpretability of assessment results in traditional solutions and enables a comprehensive and accurate assessment of users' language abilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YUANLI WEILAI SCI & TECH CO LTD
- Filing Date
- 2026-03-31
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies are insufficient to comprehensively and accurately reflect a user's overall language ability in intelligent language assessment. Traditional solutions often focus on single-dimensional indicators and lack multi-dimensional collaborative analysis mechanisms, resulting in weak interpretability and insufficient granularity of assessment results.
A multi-dimensional scoring model based on the CAF framework is adopted. The initial speech data is preprocessed through speech recognition and text analysis technology. Data processing modules with multiple dimensions such as complexity, accuracy and fluency are used to process the target speech data and target spoken text to generate comprehensive spoken test data.
It enables multi-dimensional assessment of users' language abilities, improves the interpretability and accuracy of assessment results, and can truly reflect users' comprehensive language abilities.
Smart Images

Figure CN122290589A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification relate to the technical field of data processing, in particular to a data processing method, system and device. BACKGROUND
[0002] In the field of intelligent language assessment, oral ability assessment usually relies on speech recognition and text analysis technology. However, traditional solutions focus on single-dimensional indicators (such as pronunciation accuracy or vocabulary richness), which are difficult to comprehensively and truly reflect the comprehensive language ability of users. Although some solutions combine speech and text information, the processing flow is fragmented and lacks a multi-dimensional collaborative analysis mechanism, resulting in weak interpretability and insufficient granularity of the evaluation results. Therefore, there is an urgent need for an effective data processing method to solve the above problems. SUMMARY In view of this, the embodiments of the present specification provide a data processing method. One or more embodiments of the present specification also relate to a data processing system, a data processing device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects in the prior art.
[0003] According to a first aspect of the embodiments of the present specification, a data processing method is provided, comprising: determining initial speech data of a user for an oral test task, and preprocessing the initial speech data to obtain target speech data and target oral text; processing the target speech data and the target oral text using at least two-dimensional data processing modules in a data processing model, respectively, to obtain oral analysis data corresponding to the at least two-dimensional data processing modules, respectively; generating oral test data corresponding to the oral test task based on at least two groups of oral analysis data.
[0004] Optionally, the preprocessing of the initial speech data to obtain target speech data and target oral text comprises: preprocessing the initial speech data to obtain the target speech data, and performing speech recognition on the target speech data to obtain the target oral text.
[0005] Optionally, the processing of the target speech data and the target oral text using at least two-dimensional data processing modules in a data processing model, respectively, to obtain oral analysis data corresponding to the at least two-dimensional data processing modules, respectively, comprises: determining a complexity data processing module, an accuracy data processing module, and a fluency data processing module in the data processing model; The target spoken text is input into the complexity data processing module to obtain complexity spoken language analysis data; the target speech data and the target spoken text are input into the accuracy data processing module to obtain accuracy spoken language analysis data; and the target speech data and the target spoken text are input into the fluency data processing module to obtain fluency spoken language analysis data. The complexity spoken language analysis data, the accuracy spoken language analysis data, and the fluency spoken language analysis data are used as the spoken language analysis data.
[0006] Optionally, inputting the target spoken text into the complexity data processing module to obtain complex spoken language analysis data includes: The target spoken text is input into the complexity data processing module to obtain lexical complexity data and grammatical complexity data. The lexical complexity data includes lexical richness information and lexical level information. The lexical complexity data and the grammatical complexity data are used as the complexity spoken language analysis data.
[0007] Optionally, the step of inputting the target speech data and the target spoken text into the accuracy data processing module to obtain accurate spoken analysis data includes: The syntactic accuracy unit and pronunciation accuracy unit included in the accuracy data processing module are determined. The target spoken text is input into the syntactic accuracy unit to obtain syntactic analysis data, and the target speech data is input into the pronunciation accuracy unit to obtain pronunciation analysis data; The syntactic analysis data and the pronunciation analysis data are used as the accuracy spoken language analysis data.
[0008] Optionally, the step of inputting the target speech data and the target spoken text into the fluency data processing module to obtain fluency spoken analysis data includes: The fluency data processing module is defined as including the expression fluency unit and the sentence coherence unit. The target speech data and the target spoken text are input into the expression fluency unit to obtain expression fluency data, and the target spoken text is input into the sentence coherence unit to obtain sentence coherence data. The fluency data and the coherence data are used as the fluency spoken language analysis data.
[0009] Optionally, the step of inputting the target speech data and the target spoken text into the expression fluency unit to obtain expression fluency data includes: The target speech data and the target spoken text are input into the expression fluency unit to obtain target speech rate information and expression fluency information; The target speech rate information and the expression fluency information are weighted and calculated to obtain the expression fluency data.
[0010] Optionally, inputting the target spoken text into the sentence coherence unit to obtain sentence coherence data includes: The target spoken text is input into the sentence coherence unit to obtain conjunction information, sentence length information and syntactic information; The sentence coherence data is obtained by weighting the conjunction information, sentence length information, and syntactic information.
[0011] According to a second aspect of the embodiments of this specification, a data processing system is provided, comprising: a client and a server, including: The client is used to submit initial voice data related to the oral test task to the server. The server is configured to preprocess the initial speech data to obtain target speech data and target spoken text; process the target speech data and target spoken text respectively using data processing modules of at least two dimensions in the data processing model to obtain spoken language analysis data corresponding to the at least two dimensions of the data processing module; generate spoken language test data corresponding to the spoken language test task based on at least two sets of spoken language analysis data, and feed the spoken language test data back to the client.
[0012] According to a third aspect of the embodiments of this specification, a data processing apparatus is provided, comprising: The determination module is configured to determine the user's initial speech data for the oral test task, and preprocess the initial speech data to obtain target speech data and target spoken text; The processing module is configured to process the target speech data and the target spoken text respectively using at least two dimensions of the data processing model to obtain spoken analysis data corresponding to the at least two dimensions of the data processing module. The generation module is configured to generate oral test data corresponding to the oral test task based on at least two sets of oral analysis data.
[0013] According to a fourth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above-described data processing method.
[0014] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of the data processing method described above.
[0015] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the data processing method described above.
[0016] This specification provides a data processing method in one embodiment that determines the user's initial speech data for a spoken language test task, preprocesses the initial speech data to obtain target speech data and target spoken text. The target speech data and target spoken text are processed using at least two data processing modules in a data processing model to obtain spoken language analysis data corresponding to each of the at least two data processing modules. Spoken language test data corresponding to the spoken language test task is generated based on at least two sets of spoken language analysis data. By analyzing the target speech data and target spoken text in multiple dimensions, accurate spoken language test analysis results are obtained. The spoken language test data can realistically reflect the user's language ability. Analyzing the target speech data and target spoken text in at least two dimensions improves the interpretability of the spoken language test data. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a data processing method provided in one embodiment of this specification; Figure 2 This is a flowchart illustrating the processing procedure of a data processing method provided in one embodiment of this specification. Figure 3 This is a data processing flowchart of a data processing method provided in one embodiment of this specification; Figure 4 This is a schematic diagram of the structure of a data processing system provided in one embodiment of this specification; Figure 5 This is a schematic diagram of the structure of a data processing device provided in one embodiment of this specification; Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0018] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0019] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0020] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0021] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0022] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0023] The CAF framework refers to the theoretical framework of Complexity, Accuracy, and Fluency. This is a core theory in the field of second language acquisition. This invention constructs a multi-dimensional scoring model based on this framework.
[0024] The Common European Framework of Reference for Languages (CEFR) is an internationally recognized standard for describing language proficiency, dividing language levels into six levels from A1 (beginner) to C2 (proficient). In this invention, this framework serves as a benchmark for quantifying difficulty, used to determine the difficulty coefficients of vocabulary and grammar (e.g., defining A2 level vocabulary as low-weight and C2 level vocabulary as high-weight), thereby achieving accurate calculation of language complexity.
[0025] Automatic Speech Recognition (ASR): A technology that converts human speech signals into text sequences. In this invention, the ASR module acts as a data input, responsible for receiving students' spoken recordings and transcribing them into a text format suitable for analysis, serving as the basic data source for CAF calculations.
[0026] Natural Language Processing (NLP) is a technical field involving the understanding and generation of human natural language by computers. In this invention, NLP technology is mainly used to perform word segmentation, part-of-speech tagging, dependency parsing, and grammatical error correction on the text transcribed from ASR, in order to support the calculation of complexity and accuracy.
[0027] S-shaped mapping (S-curve mapping): Based on the theory of non-linear growth in language acquisition, learners progress slowly in the initial stage, accelerate in the middle stage, and stabilize in the later stage. This invention utilizes an S-shaped function to fit this pattern to some indicators.
[0028] The trade-off effect refers to the fact that in second language production (such as speaking or writing), learners' cognitive resources (attention / working memory) are limited, and they cannot simultaneously and perfectly focus on all dimensions of language. Therefore, in pursuing high-quality performance in one language dimension, learners often have to sacrifice performance in other dimensions, thus creating a competitive relationship of "one gaining at the expense of the other."
[0029] On-the-spot performance is the concrete externalization of a learner's latent language knowledge in a real-world environment. It is not merely what a student "knows," but rather what a student "does" at a specific moment. It is instantaneous, observable behavior (such as a spoken response or an essay). It is highly dependent on the external environment at the time (such as exam pressure, noise, and familiarity with the topic) and internal state (such as fatigue, anxiety, motivation, and attention allocation).
[0030] This specification provides a data processing method, and also relates to a data processing system, a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0031] See Figure 1 , Figure 1 A flowchart of a data processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0032] Step 102: Determine the user's initial speech data for the oral test task, and preprocess the initial speech data to obtain target speech data and target spoken text.
[0033] Specifically, users can be language proficiency test takers or learners (students). The oral test task is the language proficiency test task that users need to perform. The method of testing users' language proficiency is to collect the speech data output by users during the oral test task; that is, users provide speech content, i.e., the initial speech data. Oral test tasks can also be applied to oral examination scenarios and daily oral task scenarios. The purpose of preprocessing the initial speech data is to obtain data that can be used to analyze users' language proficiency. The data format includes, but is not limited to, audio data and text data. The target speech data can be the audio obtained after preprocessing the initial speech data, which can be directly used for user language proficiency analysis. The target spoken text can be the text content obtained by transcribing the target speech data.
[0034] Therefore, when conducting language proficiency tests on users, users can provide initial speech data by performing oral test tasks. After determining the initial speech data submitted by users for the oral test task, considering that the initial speech data may contain noisy data and that the audio format of the initial speech data may not be suitable for subsequent language proficiency assessment, it is necessary to preprocess the initial speech data to obtain target speech data and target spoken text corresponding to the initial speech data.
[0035] Furthermore, considering that the initial speech data may contain noisy data, and that the audio format of the initial speech data may not be suitable for subsequent language ability assessment, it is necessary to preprocess the initial speech data and convert the target speech data into text form of target spoken language through speech recognition, so as to facilitate subsequent analysis and assessment of the user's language ability. The specific implementation is as follows: The initial speech data is preprocessed to obtain the target speech data, and the target speech data is then subjected to speech recognition to obtain the target spoken text.
[0036] Specifically, the preprocessing of the initial speech data includes, but is not limited to, adjusting the audio format of the initial speech data, denoising the initial speech data, and processing the volume of the initial speech data. The purpose of speech recognition on the target audio data is to convert the target speech data into text, obtain the target spoken text, and retain the silence segment information in the target spoken text, that is, the timestamp information of the silence segment, which serves as the pause time of the user's speech.
[0037] Based on this, the initial speech data undergoes preprocessing such as format conversion, noise reduction and enhancement, and volume normalization to obtain the target speech data. Automatic speech recognition technology is then used to perform speech recognition on the target speech data, identifying the text within it and obtaining the target spoken text containing silence segment information. The silence segment information can then be used to determine the user's pause times.
[0038] For example, in a scenario where language proficiency is tested orally, the user, i.e., the test taker, obtains initial audio data for the oral test task by recording audio in spoken form. Preprocessing the initial audio data, such as noise reduction, enhancement, and audio format adjustment, yields target audio data. Automatic speech recognition technology is then used to perform text recognition on the target audio data, converting it into target spoken text.
[0039] In summary, by preprocessing the initial speech data, noise reduction and formatting are achieved. Automatic speech recognition technology is then used to recognize the target speech data, converting it into text, thus enabling rapid audio transcription.
[0040] Step 104: Use at least two data processing modules in the data processing model to process the target speech data and the target spoken text respectively, to obtain spoken analysis data corresponding to the at least two data processing modules respectively.
[0041] Specifically, after determining the initial speech data for the user's oral test task and preprocessing the initial speech data to obtain the target speech data and target spoken text, at least two data processing modules in the data processing model can be used to process the target speech data and target spoken text respectively, obtaining oral analysis data corresponding to at least two data processing modules. The data processing model is a multi-dimensional scoring model built on the CAF framework, including but not limited to assessment dimensions such as complexity, accuracy, and fluency. Each assessment dimension corresponds to one data processing module, and each data processing module corresponds to one language ability assessment dimension. The oral analysis data corresponding to each data processing module is the assessment result obtained by evaluating the target speech data and target spoken text in the assessment dimension. The assessment result corresponds to high or low language ability. The oral analysis data can be represented in the form of a score or a level; the higher the score, the stronger the language ability; the higher the level, the stronger the language ability.
[0042] Based on this, after determining the initial speech data of the user for the oral test task and preprocessing the initial speech data to obtain the target speech data and target spoken text, the target speech data and target spoken text are processed by the data processing modules of at least two dimensions in the data processing model. This enables the analysis of the target speech data and target spoken text in at least two language ability assessment dimensions, obtaining the spoken analysis data corresponding to the data processing modules of at least two dimensions. The spoken analysis data corresponding to each data processing module represents the assessment results of different language ability assessment dimensions, thus realizing multi-dimensional language ability assessment.
[0043] Furthermore, considering that user language proficiency assessments are based on target speech data and target spoken text, it is necessary to ensure the accuracy and comprehensiveness of the assessment. This can be achieved by analyzing user language proficiency across multiple dimensions, including complexity, accuracy, and fluency. The specific implementation is as follows: The data processing model is defined by three modules: a complexity data processing module, an accuracy data processing module, and a fluency data processing module. The target spoken text is input into the complexity data processing module to obtain complexity spoken language analysis data. The target speech data and the target spoken text are input into the accuracy data processing module to obtain accuracy spoken language analysis data. The target speech data and the target spoken text are input into the fluency data processing module to obtain fluency spoken language analysis data. The complexity spoken language analysis data, the accuracy spoken language analysis data, and the fluency spoken language analysis data are used as the spoken language analysis data.
[0044] Specifically, the complexity data processing module is used to evaluate the user's language ability in the dimension of language complexity based on target speech data and / or target spoken text, and the evaluation results are represented as complex spoken language analysis data; the accuracy data processing module is used to evaluate the user's language ability in the dimension of language accuracy based on target speech data and / or target spoken text, and the evaluation results are represented as accurate spoken language analysis data; the fluency data processing module is used to evaluate the user's language ability in the dimension of language fluency based on target speech data and / or target spoken text, and the evaluation results are represented as fluency spoken language analysis data.
[0045] Based on this, the data processing model identifies a complexity data processing module, an accuracy data processing module, and a fluency data processing module. The target spoken text is input into the complexity data processing module, which evaluates the user's language ability in terms of language complexity based on the target speech data and / or the target spoken text, obtaining complexity spoken language analysis data. The target speech data and target spoken text are input into the accuracy data processing module, which evaluates the user's language ability in terms of language accuracy based on the target speech data and / or the target spoken text, obtaining accuracy spoken language analysis data. Finally, the target speech data and target spoken text are input into the fluency data processing module, which evaluates the user's language ability in terms of language fluency based on the target speech data and / or the target spoken text, obtaining fluency spoken language analysis data. By using the complexity spoken language analysis data, accuracy spoken language analysis data, and fluency spoken language analysis data as spoken language analysis data, the user's language ability can be quantified across multiple dimensions, improving the interpretability of the spoken language analysis data.
[0046] In summary, using complexity-based spoken language analysis data, accuracy-based spoken language analysis data, and fluency-based spoken language analysis data as spoken language analysis data ensures the multi-dimensional interpretability of spoken language analysis data and accurately assesses users' acquired language abilities.
[0047] Furthermore, when evaluating a user's language ability in terms of language complexity based on target speech data and / or target spoken text, the data input type for the complexity data processing module can be determined. At least one type of data can be selected from the target speech data and the target spoken text to be input into the complexity data processing module, as specifically implemented below: The target spoken text is input into the complexity data processing module to obtain lexical complexity data and grammatical complexity data. The lexical complexity data includes lexical richness information and lexical level information. The lexical complexity data and the grammatical complexity data are used as the complexity spoken language analysis data.
[0048] Specifically, lexical complexity data represents the lexical complexity of the target spoken text provided by the user, while grammatical complexity data represents the grammatical complexity of the target spoken text. The lexical richness information included in the lexical complexity data represents the lexical richness of the target spoken text. The lexical level information included in the lexical complexity data represents the lexical level of the target spoken text. This lexical level information can be obtained by substituting the lexical level score determined based on the vocabulary in the target spoken text into a sigmoid function with a preset inflection point parameter, mapping it to the 0-1 range, and using this information to represent the user's ability to use higher-level vocabulary. The sigmoid function is used to positively incentivize vocabulary beyond the user's basic level.
[0049] Based on this, the target spoken text is input into the complexity data processing module to obtain lexical complexity data and grammatical complexity data. The lexical complexity data includes lexical richness information and lexical level information. This lexical complexity data can be obtained by weighting the lexical richness information and lexical level information, and is used to represent language ability in the dimension of lexical complexity. Using lexical complexity data and grammatical complexity data as complexity spoken language analysis data can be achieved by generating complexity spoken language analysis data based on these two data points. That is, weighting the lexical complexity data and grammatical complexity data yields complexity spoken language analysis data representing the user's language ability in the dimension of complexity.
[0050] Continuing with the previous example, the target spoken text is input into the complexity data processing module. The lexical complexity unit processes the target spoken text to obtain lexical complexity data, and the grammatical complexity unit processes the target spoken text to obtain grammatical complexity data. Finally, processing the target spoken text using the lexical complexity unit yields lexical complexity data containing both lexical richness and lexical level information.
[0051] The weight of lexical richness information can be 40%. Lexical richness is used to measure the diversity of user vocabulary usage, excluding the influence of simple repetition. It is determined by removing high-frequency meaningless function words (such as articles and pronouns) from the target spoken text based on a pre-defined stop word list. The ratio of the cleaned independent words to the total number of effective words is calculated to obtain the original lexical richness value. To avoid excessive penalty for low-scoring segments and maintain sensitivity to key segment distinctions, the original value is non-linearly mapped to the 0-1 standard interval using a sigmoid function. , where c is a preset inflection point constant.
[0052] The vocabulary level information corresponds to a weight of 60%. The vocabulary level is used to assess a user's ability to use advanced vocabulary, providing positive incentives for vocabulary beyond the basic level. The determination method is as follows: all valid words used by the user in the current ability assessment (excluding high-frequency meaningless functional words and deduplication) are compared with the CEFR vocabulary list to identify the level of each word. After removing basic level (A1) and unleveled words, other levels (i.e., A2, B1, B2, C1, C2 levels) are assigned exponentially increasing weight scores, and the sum of the weight scores of all words is calculated. The average vocabulary level weight score is then calculated. The calculated average score is substituted into an S-shaped function containing a preset inflection point parameter and mapped onto the 0-1 interval to obtain vocabulary level information. By weighting the vocabulary richness information and vocabulary level information, vocabulary complexity data can be obtained.
[0053] Metric: Syntactic complexity data refers to grammatical complexity, used to evaluate the richness of sentence structure and the sophistication of grammar provided by the user. It is determined by using NLP technology to match valid sentences output by the user with a CEFR grammar level database, identifying correctly used grammar points. Based on the level of the identified grammar points, exponentially increasing weights are assigned, and the sum of the weighted scores for all grammar points is calculated. The average grammar level score is then calculated. The calculated average score is substituted into an S-shaped function with a preset inflection point parameter and mapped onto the 0-1 interval to obtain grammatical complexity data. The weight of lexical complexity can be 60%, and the weight of grammatical complexity can be 40%. The lexical complexity data and grammatical complexity data are weighted and calculated to obtain complex spoken language analysis data representing the user's language ability in the complexity dimension.
[0054] In conclusion, using lexical complexity data and grammatical complexity data as spoken language complexity analysis data improves the interpretability of users' language proficiency in the lexical complexity dimension of language ability assessment.
[0055] Furthermore, when evaluating a user's language ability in terms of language accuracy based on target speech data and / or target spoken text, the data input type of the accuracy data processing module can be determined. Based on the data input type, the input data corresponding to the syntactic accuracy unit and pronunciation accuracy unit included in the data processing module can be determined, as specifically implemented below: The accuracy data processing module includes a syntactic accuracy unit and a pronunciation accuracy unit; the target spoken text is input into the syntactic accuracy unit to obtain syntactic analysis data, and the target speech data is input into the pronunciation accuracy unit to obtain pronunciation analysis data; the syntactic analysis data and the pronunciation analysis data are used as the accuracy spoken language analysis data.
[0056] Specifically, a syntactic accuracy unit is used to determine the user's language ability in the dimension of syntactic accuracy by analyzing the target oral text in terms of syntactic accuracy, and the syntactic analysis data is used to represent the user's language ability in this ability dimension of syntactic accuracy; the pronunciation accuracy unit is used to determine the user's language ability in pronunciation accuracy by analyzing the target speech data in the dimension of pronunciation accuracy, and the pronunciation analysis data is used to represent the user's language ability in this ability dimension of pronunciation accuracy.
[0057] Based on this, the accuracy data processing module includes a syntactic accuracy unit and a pronunciation accuracy unit. The target oral text is input into the syntactic accuracy unit, and the syntactic analysis data obtained by analyzing the target oral text in the dimension of syntactic accuracy represents the syntactic accuracy degree of the user's language. The target speech data is input into the pronunciation accuracy unit, and the pronunciation analysis data obtained by analyzing the target oral text in the dimension of pronunciation accuracy represents the pronunciation accuracy degree of the user's language. The syntactic analysis data and the pronunciation analysis data are used as the accuracy oral analysis data, which can be to perform a weighted calculation on the syntactic analysis data and the pronunciation analysis data to obtain the accuracy oral analysis data representing the user's language accuracy ability.
[0058] Continuing with the above example, the accuracy data processing module includes a syntactic accuracy unit and a pronunciation accuracy unit. The target oral text is input into the syntactic accuracy unit, and the obtained syntactic analysis data identifies the grammar correctness rate in the user's language ability, which is used to evaluate the proportion of grammatically correct sentences and reflects the normativity of the language. The determination method of the syntactic analysis data can be: using an AI grammar model to analyze the target oral text and counting the number of grammatically correct sentences. A non-linear exponential decay model is used to calculate the score, , where n is a preset exponential parameter (0 < n < 1), so that the score shows a gentle decay trend as the correctness rate decreases, constructing a tolerance interval for errors, and the grammar correctness rate is the syntactic analysis data.
[0059] The target speech data and target spoken text are input into the pronunciation accuracy unit. The resulting pronunciation analysis data can be a speech evaluation score, representing pronunciation accuracy, used to evaluate the clarity, intonation, and naturalness of word and sentence pronunciation. The pronunciation analysis data can be determined by inputting the target speech data into an AI speech evaluation model to obtain pronunciation evaluation scores (0-100 range) for selected words and sentences. When processing word pronunciation, the highest score for each word across multiple attempts or multi-dimensional model outputs is taken, and the average highest score for all words is calculated. When processing sentence pronunciation, the highest score for each sentence in the evaluation is taken, and the average highest score for all sentences is calculated. The word pronunciation scores and sentence pronunciation scores are then weighted and summed, with the word pronunciation score accounting for 40% and the sentence pronunciation score accounting for 60%. The calculated scores are then substituted into an S-shaped function containing a preset inflection point parameter, mapped to the 0-1 range, to obtain the pronunciation accuracy. The weight of syntactic accuracy can be 60%, and the weight of pronunciation accuracy can be 40%. By weighting the syntactic analysis data and the pronunciation analysis data, accurate spoken language analysis data representing the user's language accuracy ability can be obtained.
[0060] In summary, using syntactic analysis data and pronunciation analysis data as accurate spoken language analysis data allows for weighted calculation of these two data to obtain accurate spoken language analysis data that represents the user's language accuracy ability, thus evaluating the user's language ability in the dimension of language accuracy.
[0061] Furthermore, when evaluating a user's language proficiency in the fluency dimension based on target speech data and / or target spoken text, the fluency data processing module can be defined as including the expression fluency unit and the sentence coherence unit. At least one type of data can be selected from the target speech data and the target spoken text and input into the fluency data processing module, as specifically implemented below: The fluency data processing module includes an expression fluency unit and a sentence coherence unit; the target speech data and the target spoken text are input into the expression fluency unit to obtain expression fluency data, and the target spoken text is input into the sentence coherence unit to obtain sentence coherence data; the expression fluency data and the sentence coherence data are used as the fluency spoken language analysis data.
[0062] Specifically, the fluency unit analyzes the target spoken text in terms of fluency to determine the user's language proficiency in this dimension, and uses fluency data to represent the user's language proficiency in this dimension. The coherence unit analyzes the target spoken text in terms of coherence to determine the user's language proficiency in this dimension, and uses coherence data to represent the user's language proficiency in this dimension.
[0063] Based on this, the fluency data processing module is defined to include an expressive fluency unit and a sentence coherence unit. Target speech data and target spoken text are input into the expressive fluency unit, and the expressive fluency data represents the user's spoken language fluency. Target spoken text is input into the sentence coherence unit, and the target spoken text and target speech data are analyzed in terms of sentence coherence, resulting in sentence coherence data representing the user's sentence coherence. Using the expressive fluency data and sentence coherence data as fluency spoken language analysis data can be achieved by weighting the expressive fluency data and sentence coherence data. The weight of expressive fluency data can be 40%, and the weight of sentence coherence data can be 60%.
[0064] In summary, by weighting the fluency data and sentence coherence data, we obtain fluency spoken language analysis data, which is used to represent the smoothness, naturalness, and coherence of users' language output, and accurately assess the language abilities that users have acquired.
[0065] Furthermore, the target speech data and target spoken text are input into the fluency unit to analyze the fluency of the user's language use. The specific implementation is as follows: The target speech data and the target spoken text are input into the expression fluency unit to obtain target speech rate information and expression fluency information; the target speech rate information and the expression fluency information are weighted and calculated to obtain the expression fluency data.
[0066] Specifically, the target speech rate information represents the optimal speech rate during the user's language proficiency test, used to measure the user's speech rate in continuous speech segments. The fluency information represents the fluency of the user's speech during the language proficiency test, used to measure unnatural pauses and corrections during the user's speech.
[0067] Based on this, target speech data and target spoken text are input into the fluency unit to obtain target speech rate information for measuring the user's speech rate in continuous speech segments, and fluency information for measuring unnatural pauses and corrections during the user's speech expression. The target speech rate information and fluency information are weighted and calculated to obtain fluency data; this can be done by weighting the target speech rate information and fluency information according to preset weights.
[0068] Continuing with the previous example, the target speech data and target spoken text are input into the fluency unit to obtain target speech rate information used to measure the user's speech rate in continuous speech segments. Target speech rate information can be represented as the optimal speech rate, with a weight of 70%. The target speech rate information is determined by calculating the speech rate of the user's longest continuous speech segment. A preset optimal speech rate range (including a lower and upper threshold) is set. A sigmoid function is used to map the speech rate: a lower score is achieved when the speech rate is below the lower limit or above the upper limit, and a score approaches full marks when the speech rate is within the optimal range.
[0069] The target speech data and target spoken text are input into the fluency unit to obtain fluency information, which measures unnatural pauses and corrections during the user's speech expression. This fluency information can be represented as fluency score. The frequency of filler words (e.g., uh, um) and the frequency of repeated corrections in the target speech data and target spoken text are detected. The filler rate and self-correction rate are calculated separately. , Fluency score is calculated by weighted subtraction to obtain expression fluency information (expression fluency). The weight of expression fluency can be 30%. By weighting the optimal speaking speed and expression fluency, the expression fluency data can be obtained.
[0070] In summary, by weighting the target speech rate information and the fluency information, we can obtain fluency data, achieve fine-grained segmentation of user fluency, and improve the interpretability of fluency data.
[0071] Furthermore, the target spoken text is input into the sentence coherence unit to analyze the sentence coherence of the user's language ability and obtain sentence coherence data. The specific implementation is as follows: The target spoken text is input into the sentence coherence unit to obtain conjunction information, sentence length information, and syntactic information; the conjunction information, sentence length information, and syntactic information are weighted and calculated to obtain the sentence coherence data.
[0072] Specifically, conjunction information represents the richness of conjunctions in the user's language, used to assess the diversity of logical conjunction usage. Sentence length information represents the average sentence length of the user's language, which can indirectly reflect the complexity and development level of the user's language. Syntactic information represents the peak syntactic performance of the user's language, used to assess the peak performance of the user's language, reflecting the highest level of language organization ability the user has acquired.
[0073] Based on this, the target spoken text is input into the sentence coherence unit to obtain conjunction information, sentence length information, and syntactic information. Conjunction information represents the conjunction richness of the user's language, sentence length information represents the average sentence length of the user's language, and syntactic information represents the peak syntactic performance of the user's language. The conjunction information, sentence length information, and syntactic information are weighted and calculated to obtain sentence coherence data.
[0074] Continuing with the previous example, conjunction information corresponds to conjunction richness, with a weight of 40%. Conjunction information is used to evaluate the diversity of a user's use of logical conjunctions. The calculation of conjunction information can be done by counting the number of non-repeating logical conjunctions (such as and, but, however, etc.) in the target spoken text. Substituting the number of non-repeating conjunctions into a sigmoid function with a preset inflection point parameter, it is mapped to the 0-1 interval. Sentence length information corresponds to average sentence length, with a weight of 30%. Sentence length information indirectly reflects the user's language complexity and development level. The calculation of sentence length information can be done by calculating the average number of words in the user's output sentences (excluding simple sentences with fewer than 3 words). Substituting the original average sentence length value into a sigmoid function with a preset inflection point parameter, it is mapped to the 0-1 interval. Syntax information corresponds to peak syntactic performance, with a weight of 30%. Syntax information is used to evaluate the user's peak performance, reflecting the highest level of language organization ability the user has acquired. The calculation of syntactic information can be done by extracting the maximum number of words the user speaks in a single instance and mapping it using a sigmoid function. Extract the word count of the longest sentence and use an S-shaped function for mapping and scoring. Then, sum the two peak metrics mentioned above with a weighted average of 50%.
[0075] In summary, weighted calculations are performed on conjunction information, sentence length information, and syntactic information to obtain sentence coherence data. This enables fine-grained segmentation of user sentence coherence and improves the interpretability of sentence coherence data.
[0076] Step 106: Generate oral test data corresponding to the oral test task based on at least two sets of oral analysis data.
[0077] Specifically, after processing the target speech data and target spoken text using at least two dimensions of the data processing model to obtain spoken language analysis data corresponding to each of the at least two data processing modules, spoken language test data corresponding to the spoken language test task can be generated based on at least two sets of spoken language analysis data. Each set of spoken language analysis data corresponds to different dimensions of language ability assessment information. Generating spoken language test data based on at least two sets of spoken language analysis data can be achieved by dynamically weighting the at least two sets of spoken language analysis data. The weights corresponding to each set of spoken language analysis data can be dynamically adjusted according to the user's language ability characteristics. That is, if the spoken language analysis data determines that the user exhibits high-order or high-level language features, the weight of that assessment dimension is adjusted. Specifically, if the user exhibits high-order language features in the complexity dimension, the weight of the complexity dimension is adjusted.
[0078] Based on this, after processing the target speech data and target spoken text using at least two dimensions of the data processing model to obtain spoken analysis data corresponding to at least two data processing modules, spoken test data corresponding to the spoken test task is generated based on at least two sets of spoken analysis data. The spoken test data corresponding to the user's current spoken test task is obtained by dynamically weighting the at least two sets of spoken analysis data.
[0079] In practical applications, at least two sets of spoken language analysis data can be used, or even three sets. Each set of spoken language analysis data corresponds to the evaluation information (dynamic weights) for the evaluation dimensions of complexity, accuracy, and fluency. If, based on the spoken language analysis data corresponding to complexity, it is determined that a user demonstrates strong language ability in this dimension, the weight of this dimension can be increased. A weight threshold can be set for each dimension to represent the upper limit of weight adjustment. The weights corresponding to complexity, accuracy, and fluency can be set to 40%, 25%, and 30%, respectively. The weight threshold for complexity can be set to 60%. It should be noted that the weight thresholds corresponding to complexity, accuracy, and fluency can be set according to actual needs; this embodiment does not impose any limitations on this.
[0080] The data processing method provided in one embodiment of this specification achieves high-precision, structured, multi-dimensional ability diagnosis, improving the interpretability of the assessment. By constructing a structured assessment system based on CAF theory, the general oral assessment is refined into more than ten quantitative indicators. It not only outputs a comprehensive index but also accurately identifies learners' strengths and weaknesses in specific dimensions, providing clearly targeted data support for teaching interventions and solving the problem of traditional assessments lacking fine-grained diagnostic value.
[0081] Users can obtain continuous oral proficiency assessments by taking ongoing oral tests. This process-based evaluation quantifies even small improvements between two levels, effectively reducing frustration in the initial stages. Addressing the issue of traditional linear scoring potentially undermining beginners' confidence, an exponential decay model and S-shaped mapping are introduced. This significantly reduces the penal sensitivity of lower score brackets, not only accommodating learners' trial-and-error errors and protecting their learning motivation, but also keenly capturing and amplifying even small signs of progress during the initial "0 to 1" stage, thus solving the problem of insufficient motivation in lower score brackets caused by linear scoring.
[0082] In the embodiments described in this specification, by dynamically adjusting the weights corresponding to complexity, accuracy, and fluency, the assessment bias caused by the trade-off effect can be reduced, and the acquired language ability can be accurately assessed. By introducing peak performance indicators and a dynamic weight adjustment mechanism, the algorithm compensates for fluency fluctuations that occur when learners attempt higher-level language expressions. This mechanism allows the assessment results to no longer be limited by conservative on-the-spot performance, but to penetrate the surface fluency interference, accurately identify and quantify the learner's acquired language ability (Competence), and objectively affirm the student's initiative in challenging high-difficulty expressions. A unified scale is constructed to achieve longitudinal tracking throughout the entire learning cycle. A unified calculation system is established that runs through different learning stages. The underlying data of all dimensions are directly aligned with the CEFR international standard and adopt consistent CAF calculation logic. This ensures that the assessment data of different learning stages are mathematically rigorously comparable, thereby generating a continuous longitudinal growth trajectory, supporting long-term effect tracking and ability prediction for learners, and solving the problem that traditional assessments cannot make long-term comparisons due to fragmented standards.
[0083] This specification provides a data processing method in one embodiment that determines the user's initial speech data for a spoken language test task, preprocesses the initial speech data to obtain target speech data and target spoken text. The target speech data and target spoken text are processed using at least two data processing modules in a data processing model to obtain spoken language analysis data corresponding to each of the at least two data processing modules. Spoken language test data corresponding to the spoken language test task is generated based on at least two sets of spoken language analysis data. By analyzing the target speech data and target spoken text in multiple dimensions, accurate spoken language test analysis results are obtained. The spoken language test data can realistically reflect the user's language ability. Analyzing the target speech data and target spoken text in at least two dimensions improves the interpretability of the spoken language test data.
[0084] The following is in conjunction with the appendix Figure 2 Taking the application of the data processing method provided in this specification in oral proficiency assessment as an example, the data processing method will be further explained. Figure 2A flowchart illustrating the processing procedure of a data processing method according to an embodiment of this specification is shown, specifically including the following steps.
[0085] Step 202: Determine the user's initial speech data for the oral test task.
[0086] Step 204: Preprocess the initial speech data to obtain the target speech data, and perform speech recognition on the target speech data to obtain the target spoken text.
[0087] Step 206: Determine the complexity data processing module, accuracy data processing module, and fluency data processing module in the data processing model.
[0088] Step 208: Input the target spoken text into the complexity data processing module to obtain complexity spoken analysis data; input the target speech data and target spoken text into the accuracy data processing module to obtain accuracy spoken analysis data; and input the target speech data and target spoken text into the fluency data processing module to obtain fluency spoken analysis data.
[0089] Step 210: Perform dynamic weighted calculation on the complexity spoken language analysis data, accuracy spoken language analysis data, and fluency spoken language analysis data to obtain the spoken language test data corresponding to the spoken language test task.
[0090] In practical applications, the data processing flow is as follows: Figure 3 As shown, the input data is the user's initial speech data. After preprocessing the initial speech data, target speech data and target spoken text are obtained. The CAF three-dimensional parallel computing module is used to process the target speech data and target spoken text, specifically, the complexity processing module, the accuracy processing module, and the fluency processing module are used to process the target speech data and target spoken text to obtain spoken language analysis data for each dimension. The spoken language analysis data (scores) for each dimension are dynamically weighted and calculated to obtain spoken language test data. The spoken language test data is the output result, representing the user's spoken language ability assessment result.
[0091] The data processing method provided in one embodiment of this specification fully considers the impact of trade-offs, introducing a dynamic weight adjustment mechanism and peak feature extraction technology to accurately assess the user's (student's) actual language proficiency. A high-precision oral assessment system is built based on CAF, deconstructing oral proficiency into three dimensions: complexity, accuracy, and fluency, and further refining them into more than ten independent quantitative indicators. This system breaks the "black box" limitation of end-to-end models, achieving fine-grained, structured diagnosis of oral proficiency. Specific nonlinear functions (such as sigmoid and exponential functions) are applied to adapt to the nonlinear curve of language learning.
[0092] Figure 4 This specification shows a schematic diagram of the structure of a data processing system according to one embodiment. Figure 4 As shown, the data processing system 400 includes a client 410 and a server 420. The client 410 is used to submit initial speech data related to a spoken language test task to the server 420. The server 420 is used to preprocess the initial speech data to obtain target speech data and target spoken language text. The target speech data and target spoken language text are processed by data processing modules of at least two dimensions in the data processing model to obtain spoken language analysis data corresponding to the data processing modules of at least two dimensions respectively. Spoken language test data corresponding to the spoken language test task is generated based on at least two sets of spoken language analysis data, and the spoken language test data is fed back to the client 410.
[0093] In practical applications, the user's client can record audio for a spoken language test task, obtaining initial voice data. The client sends this initial voice data to the server, which analyzes it to assess the user's language ability. The server preprocesses the initial voice data to obtain target voice data and target spoken text. At least two dimensions of the data processing model are used to process the target voice data and target spoken text separately, obtaining spoken language analysis data corresponding to each of the at least two data processing modules. Based on at least two sets of spoken language analysis data, spoken language test data corresponding to the spoken language test task is generated and fed back to the client, informing the user of the test results for their language ability. By analyzing the target voice data and target spoken text across multiple dimensions, accurate spoken language test analysis results are obtained. The spoken language test data can accurately reflect the user's language ability, and analyzing the target voice data and target spoken text across at least two dimensions improves the interpretability of the spoken language test data.
[0094] Corresponding to the above method embodiments, this specification also provides data processing apparatus embodiments. Figure 5 A schematic diagram of the structure of a data processing apparatus according to one embodiment of this specification is shown. Figure 5 As shown, the device includes: The determining module 502 is configured to determine the user's initial speech data for the oral test task, and preprocess the initial speech data to obtain target speech data and target spoken text. The processing module 504 is configured to process the target speech data and the target spoken text respectively using at least two dimensions of the data processing model to obtain spoken analysis data corresponding to the at least two dimensions of the data processing module. The generation module 506 is configured to generate oral test data corresponding to the oral test task based on at least two sets of oral analysis data.
[0095] In an optional embodiment, the preprocessing of the initial speech data to obtain target speech data and target spoken text includes: The initial speech data is preprocessed to obtain the target speech data, and the target speech data is then subjected to speech recognition to obtain the target spoken text.
[0096] In an optional embodiment, the step of processing the target speech data and the target spoken text using at least two dimensions of the data processing model to obtain spoken analysis data corresponding to each of the at least two dimensions of the data processing module includes: The complexity data processing module, accuracy data processing module, and fluency data processing module in the data processing model are determined. The target spoken text is input into the complexity data processing module to obtain complexity spoken language analysis data; the target speech data and the target spoken text are input into the accuracy data processing module to obtain accuracy spoken language analysis data; and the target speech data and the target spoken text are input into the fluency data processing module to obtain fluency spoken language analysis data. The complexity spoken language analysis data, the accuracy spoken language analysis data, and the fluency spoken language analysis data are used as the spoken language analysis data.
[0097] In an optional embodiment, inputting the target spoken text into the complexity data processing module to obtain complex spoken language analysis data includes: The target spoken text is input into the complexity data processing module to obtain lexical complexity data and grammatical complexity data. The lexical complexity data includes lexical richness information and lexical level information. The lexical complexity data and the grammatical complexity data are used as the complexity spoken language analysis data.
[0098] In an optional embodiment, inputting the target speech data and the target spoken text into the accuracy data processing module to obtain accurate spoken analysis data includes: The syntactic accuracy unit and pronunciation accuracy unit included in the accuracy data processing module are determined. The target spoken text is input into the syntactic accuracy unit to obtain syntactic analysis data, and the target speech data is input into the pronunciation accuracy unit to obtain pronunciation analysis data; The syntactic analysis data and the pronunciation analysis data are used as the accuracy spoken language analysis data.
[0099] In an optional embodiment, the step of inputting the target speech data and the target spoken text into the fluency data processing module to obtain fluency spoken analysis data includes: The fluency data processing module is defined as including the expression fluency unit and the sentence coherence unit. The target speech data and the target spoken text are input into the expression fluency unit to obtain expression fluency data, and the target spoken text is input into the sentence coherence unit to obtain sentence coherence data. The fluency data and the coherence data are used as the fluency spoken language analysis data.
[0100] In an optional embodiment, the step of inputting the target speech data and the target spoken text into the fluency unit to obtain fluency data includes: The target speech data and the target spoken text are input into the expression fluency unit to obtain target speech rate information and expression fluency information; The target speech rate information and the expression fluency information are weighted and calculated to obtain the expression fluency data.
[0101] In an optional embodiment, inputting the target spoken text into the sentence coherence unit to obtain sentence coherence data includes: The target spoken text is input into the sentence coherence unit to obtain conjunction information, sentence length information and syntactic information; The sentence coherence data is obtained by weighting the conjunction information, sentence length information, and syntactic information.
[0102] This specification provides a data processing apparatus in one embodiment that determines initial speech data of a user for a spoken language test task and preprocesses the initial speech data to obtain target speech data and target spoken text. The target speech data and target spoken text are processed separately using data processing modules in at least two dimensions of the data processing model to obtain spoken language analysis data corresponding to each of the at least two data processing modules. Spoken language test data corresponding to the spoken language test task is generated based on at least two sets of spoken language analysis data. By analyzing the target speech data and target spoken text in multiple dimensions, accurate spoken language test analysis results are obtained. The spoken language test data can realistically reflect the user's language ability. Analyzing the target speech data and target spoken text in at least two dimensions improves the interpretability of the spoken language test data.
[0103] The above is an illustrative scheme of a data processing apparatus according to this embodiment. It should be noted that the technical solution of this data processing apparatus and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the data processing apparatus, please refer to the description of the technical solution of the data processing method described above.
[0104] Figure 6 A structural block diagram of a computing device 600 according to one embodiment of this specification is shown. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.
[0105] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0106] In one embodiment of this specification, the above-described components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0107] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.
[0108] The processor 620 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-described data processing method.
[0109] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the data processing method described above.
[0110] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described data processing method.
[0111] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the data processing method described above.
[0112] An embodiment of this specification also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described data processing method.
[0113] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the data processing method described above.
[0114] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0115] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0116] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0117] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0118] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A data processing method, characterized in that, include: Determine the user's initial speech data for the oral test task, and preprocess the initial speech data to obtain target speech data and target spoken text; The target speech data and the target spoken text are processed by data processing modules of at least two dimensions in the data processing model to obtain spoken analysis data corresponding to the data processing modules of at least two dimensions respectively; The oral test data corresponding to the oral test task is generated based on at least two sets of oral analysis data.
2. The data processing method according to claim 1, characterized in that, The preprocessing of the initial speech data to obtain target speech data and target spoken text includes: The initial speech data is preprocessed to obtain the target speech data, and the target speech data is then subjected to speech recognition to obtain the target spoken text.
3. The data processing method according to claim 1, characterized in that, The process of using at least two data processing modules in the data processing model to process the target speech data and the target spoken text respectively, to obtain spoken analysis data corresponding to each of the at least two data processing modules, includes: The complexity data processing module, accuracy data processing module, and fluency data processing module in the data processing model are determined. The target spoken text is input into the complexity data processing module to obtain complexity spoken language analysis data; the target speech data and the target spoken text are input into the accuracy data processing module to obtain accuracy spoken language analysis data; and the target speech data and the target spoken text are input into the fluency data processing module to obtain fluency spoken language analysis data. The complexity spoken language analysis data, the accuracy spoken language analysis data, and the fluency spoken language analysis data are used as the spoken language analysis data.
4. The data processing method according to claim 3, characterized in that, The step of inputting the target spoken text into the complexity data processing module to obtain complex spoken analysis data includes: The target spoken text is input into the complexity data processing module to obtain lexical complexity data and grammatical complexity data. The lexical complexity data includes lexical richness information and lexical level information. The lexical complexity data and the grammatical complexity data are used as the complexity spoken language analysis data.
5. The data processing method according to claim 3, characterized in that, The step of inputting the target speech data and the target spoken text into the accuracy data processing module to obtain accurate spoken analysis data includes: The syntactic accuracy unit and pronunciation accuracy unit included in the accuracy data processing module are determined. The target spoken text is input into the syntactic accuracy unit to obtain syntactic analysis data, and the target speech data is input into the pronunciation accuracy unit to obtain pronunciation analysis data; The syntactic analysis data and the pronunciation analysis data are used as the accuracy spoken language analysis data.
6. The data processing method according to claim 3, characterized in that, The step of inputting the target speech data and the target spoken text into the fluency data processing module to obtain fluency spoken analysis data includes: The fluency data processing module is defined as including the expression fluency unit and the sentence coherence unit. The target speech data and the target spoken text are input into the expression fluency unit to obtain expression fluency data, and the target spoken text is input into the sentence coherence unit to obtain sentence coherence data. The fluency data and the coherence data are used as the fluency spoken language analysis data.
7. The data processing method according to claim 6, characterized in that, The step of inputting the target speech data and the target spoken text into the expression fluency unit to obtain expression fluency data includes: The target speech data and the target spoken text are input into the expression fluency unit to obtain target speech rate information and expression fluency information; The target speech rate information and the expression fluency information are weighted and calculated to obtain the expression fluency data.
8. The data processing method according to claim 6, characterized in that, The step of inputting the target spoken text into the sentence coherence unit to obtain sentence coherence data includes: The target spoken text is input into the sentence coherence unit to obtain conjunction information, sentence length information and syntactic information; The sentence coherence data is obtained by weighting the conjunction information, sentence length information, and syntactic information.
9. A data processing system, characterized in that, Including both client and server sides, including: The client is used to submit initial voice data related to the oral test task to the server. The server is configured to preprocess the initial speech data to obtain target speech data and target spoken text; process the target speech data and target spoken text respectively using data processing modules of at least two dimensions in the data processing model to obtain spoken language analysis data corresponding to the at least two dimensions of the data processing module; generate spoken language test data corresponding to the spoken language test task based on at least two sets of spoken language analysis data, and feed the spoken language test data back to the client.
10. A data processing apparatus, characterized in that, include: The determination module is configured to determine the user's initial speech data for the oral test task, and preprocess the initial speech data to obtain target speech data and target spoken text; The processing module is configured to process the target speech data and the target spoken text respectively using at least two dimensions of the data processing model to obtain spoken analysis data corresponding to the at least two dimensions of the data processing module. The generation module is configured to generate oral test data corresponding to the oral test task based on at least two sets of oral analysis data.
11. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the data processing method according to any one of claims 1 to 8.
12. A computer-readable storage medium, characterized in that, It stores computer-executable instructions that, when executed by a processor, implement the steps of the data processing method according to any one of claims 1 to 8.
13. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processor, implement the steps of the data processing method according to any one of claims 1 to 8.