Vocabulary prediction method, device, storage medium and computer equipment

By combining the hierarchical evaluation model and the vocabulary prediction model, the text rating and vocabulary of the target text are evaluated, and the problem of inaccurate vocabulary prediction in the existing technology is solved and more accurate vocabulary prediction is achieved.

CN116070625BActive Publication Date: 2025-08-26GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111280487.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-29
Publication Date
2025-08-26
Estimated Expiration
2041-10-29

AI Technical Summary

Technical Problem

Existing vocabulary prediction methods cannot accurately reflect the user's real vocabulary mastery level, resulting in a low prediction accuracy.

Method used

The hierarchical evaluation model is used to evaluate the text level of the target text, and combined with the vocabulary prediction model, the predicted vocabulary of the target user is determined through the vocabulary rank and order relationship in the target text.

Benefits of technology

It improves the accuracy of vocabulary prediction, reflects the user's understanding of vocabulary, and ensures the rationality of the prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116070625B_ABST
    Figure CN116070625B_ABST
Patent Text Reader

Abstract

This application discloses a vocabulary prediction method, apparatus, storage medium, and computer device, wherein the method comprises: obtaining a target text for a target user, obtaining each first target vocabulary in the target text and the order relationship between each first target vocabulary; using a grade evaluation model to obtain a target text grade of the target text based on each first target vocabulary and the order relationship between each first target vocabulary; using a vocabulary prediction model to obtain a text vocabulary corresponding to the target text based on each first target vocabulary and the vocabulary grade corresponding to each first target vocabulary; and determining a predicted vocabulary of the target user based on the target grade vocabulary grade and the text vocabulary. Using this application, the accuracy of vocabulary prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to a vocabulary prediction method, apparatus, storage medium, and computer equipment. Background Art

[0002] With the rise of online education, more and more users are using learning software to learn languages. To test their learning results, users use the vocabulary prediction function of learning software to predict their vocabulary level.

[0003] Existing vocabulary prediction methods randomly select hierarchical words from a vocabulary library, then generate test questions for the user based on the selected hierarchical words. Vocabulary prediction is then performed based on the test results. The resulting prediction results only reflect the user's vocabulary memorization level. According to the characteristics of language, true mastery of a word is achieved only when it is used correctly (e.g., connecting words into sentences). Simply memorizing or writing a word fails to reflect the user's vocabulary mastery level, and further, it is impossible to determine whether the words memorized or written by the user are truly mastered. Therefore, existing vocabulary prediction methods are unable to determine the user's true vocabulary level, and the accuracy of vocabulary prediction is low. Summary of the Invention

[0004] The present application provides a vocabulary prediction method, apparatus, storage medium, and computer device, which can solve the technical problem of how to improve the accuracy of vocabulary prediction.

[0005] In a first aspect, an embodiment of the present application provides a vocabulary prediction method, the method comprising:

[0006] Acquire a target text for a target user, acquire first target words in the target text and a sequence relationship between the first target words;

[0007] Using a grade evaluation model, based on the first target words and the order relationship between the first target words, obtain a target text grade of the target text;

[0008] Using a vocabulary prediction model, based on each of the first target words and the vocabulary level corresponding to each of the first target words, obtain a text vocabulary corresponding to the target text;

[0009] The predicted vocabulary size of the target user is determined based on the target level vocabulary size corresponding to the target text level and the text vocabulary size.

[0010] In a second aspect, an embodiment of the present application provides a vocabulary prediction device, comprising:

[0011] An acquisition module, configured to acquire a target text for a target user, and acquire first target words in the target text and a sequence relationship between the first target words;

[0012] a grade acquisition module, configured to acquire a target text grade of the target text based on the first target words and the order relationship between the first target words by using a grade evaluation model;

[0013] A vocabulary acquisition module, configured to acquire a text vocabulary corresponding to the target text based on the first target words and the vocabulary levels corresponding to the first target words using a vocabulary prediction model;

[0014] A determination module is configured to determine a predicted vocabulary size of the target user based on a target level vocabulary size corresponding to the target text level and the text vocabulary size.

[0015] In a third aspect, an embodiment of the present application provides a storage medium, wherein the storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the above method.

[0016] In a fourth aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned method when executing the program.

[0017] In an embodiment of the present application, the target text of the target user is evaluated for its text level by adopting a level evaluation model, and the text vocabulary corresponding to the target text is determined in combination with a vocabulary prediction model. Finally, the predicted vocabulary of the target user is determined based on the level vocabulary corresponding to the text level of the target text and the text vocabulary corresponding to the target text. By adopting the user's text, starting from the text perspective that reflects the user's understanding of vocabulary, and combining the level evaluation model and the vocabulary prediction model to predict the user's vocabulary level, the rationality of the vocabulary prediction is guaranteed and the accuracy of the vocabulary prediction is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 A flowchart of a vocabulary prediction method provided in an embodiment of the present application;

[0020] Figure 2 A flowchart of a vocabulary prediction method provided in an embodiment of the present application;

[0021] Figure 3 A flowchart of a vocabulary prediction method provided in an embodiment of the present application;

[0022] Figure 4 A flowchart of a vocabulary prediction method provided in an embodiment of the present application;

[0023] Figure 5 A flowchart of a vocabulary prediction method provided in an embodiment of the present application;

[0024] Figure 6 A flowchart of a vocabulary prediction method provided in an embodiment of the present application;

[0025] Figure 7 A schematic diagram of the structure of a vocabulary prediction device provided in an embodiment of the present application;

[0026] Figure 8 A schematic diagram of the structure of a vocabulary prediction device provided in an embodiment of the present application;

[0027] Figure 9 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] To make the features and advantages of this application more obvious and easy to understand, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.

[0029] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. The flowcharts shown in the accompanying drawings are only exemplary illustrations and do not have to be executed according to the steps shown. For example, some steps are parallel and do not have a strict logical order, so the actual execution order is variable. In addition, the terms "first", "second", "third", "fourth", "fifth", "sixth", "seventh", and "eighth" are only for the purpose of distinction and should not be used as limitations on the content of this disclosure.

[0030] The following will be combined Figure 1-Figure 7 , a detailed introduction to the vocabulary prediction method provided in the embodiments of the present application is given.

[0031] See Figure 1 , which is a flowchart of a vocabulary prediction method according to an embodiment of the present application. Figure 1 As shown, the method may include the following steps S101 to S104.

[0032] S101: Acquire a target text for a target user, and acquire first target words in the target text and a sequence relationship between the first target words.

[0033] Specifically, the target text is text material used to predict the target user's vocabulary. The text material may be an article written by the target user. Optionally, the article includes a title and body content. Since text is a form of written language, specifically a sentence or a combination of sentences, and sentences are composed of words and the sequential relationships between words, it can be understood that the target text includes at least one first target word and the sequential relationships between the first target words.

[0034] The vocabulary prediction device first obtains a target text for a target user, and then obtains first target words in the target text and a sequence relationship between the first target words based on the target text.

[0035] S102 : Using a grade evaluation model, based on the first target words and the order relationship between the first target words, obtain the target text grade of the target text.

[0036] Specifically, the grade assessment model is an assessment model for assessing the text grade of a text, wherein the text grade is used to indicate the text level of the text; illustratively, this embodiment divides the text grades into primary school level, junior high school level, senior high school level, level 4 and level 6, etc., which are not limited here. It should be noted that the text grade of a text can reflect the vocabulary range of the words mastered by the user who wrote the text; illustratively, the vocabulary range corresponding to each text grade can be 0-1000 for primary school level, 1000-3000 for junior high school level, 3000-4000 for senior high school level, 4000-5500 for level 4 and 5500-7000 for level 6, which are not limited here.

[0037] The vocabulary prediction device inputs the target text into the level evaluation model. Exemplarily, the first target words acquired by the vocabulary prediction device are sequentially input into the level evaluation model according to the order relationship between the first target words.

[0038] The grade evaluation model identifies text features such as word dependencies and sentence structures within sentences of the target text based on the received first target words and the sequential relationship between the first target words, and then determines the text grade corresponding to the currently input text based on the first target words and the text features, and finally outputs the obtained text grade.

[0039] The vocabulary prediction device uses the text level output by the level evaluation model as the target text level of the target text.

[0040] Optionally, the vocabulary prediction device can directly input a text file corresponding to the target text into the grade assessment model. After receiving the text file, the grade assessment model directly assesses the target text grade based on the text content stored in the text file (i.e., the first target words in the target text and the order relationship between the first target words), without limitation herein.

[0041] S103 , using a vocabulary prediction model, based on the first target words and the vocabulary levels corresponding to the first target words, to obtain the text vocabulary corresponding to the target text.

[0042] Specifically, the vocabulary prediction model is a prediction model used to predict the vocabulary of a target user based on a target text corresponding to the target user, and the vocabulary predicted by the vocabulary prediction model is a specific value.

[0043] The vocabulary prediction device inputs the target text into the vocabulary prediction model. Exemplarily, the vocabulary prediction device may obtain each first target word input level evaluation model.

[0044] The vocabulary prediction model obtains the vocabulary level corresponding to each first target vocabulary based on the received first target vocabulary, and then performs vocabulary prediction based on each first target vocabulary and the vocabulary level corresponding to each first target vocabulary to obtain the text vocabulary, and finally outputs the obtained text vocabulary.

[0045] The vocabulary prediction device uses the text vocabulary output by the vocabulary prediction model as the text vocabulary corresponding to the target text.

[0046] Optionally, the vocabulary prediction device can directly input a text file corresponding to the target text into the vocabulary prediction model. After receiving the text file, the vocabulary prediction model directly predicts the text vocabulary corresponding to the target text based on the text content stored in the text file (i.e., each first target word in the target text), which is not limited here.

[0047] S104: Determine the predicted vocabulary size of the target user based on the target level vocabulary size corresponding to the target text level and the text vocabulary size.

[0048] Specifically, the level vocabulary corresponding to the text level can be the upper limit of the vocabulary range corresponding to the text level. For example, the level vocabulary corresponding to each text level can be 1000 for primary school level, 3000 for junior high school level, 4000 for high school level, 5500 for CET-4 level, and 7000 for CET-6 level. It can also be calculated based on the vocabulary range corresponding to the text level, which is not limited here.

[0049] The vocabulary prediction device first obtains target-level vocabulary corresponding to the target text level, and then determines the predicted vocabulary of the target user based on the target-level vocabulary and the text vocabulary.

[0050] Optionally, the target level vocabulary can be compared with the text vocabulary to determine whether the text vocabulary obtained by the vocabulary prediction model is too high, and then the final predicted vocabulary can be obtained based on the judgment result. Specifically, since the text level of the target text written by the target user directly reflects the vocabulary level of the target user, the text vocabulary predicted based on the target text must fall within the vocabulary range corresponding to the text level. Therefore, if the predicted text vocabulary falls within the vocabulary range corresponding to the text level, the predicted text vocabulary is determined to be the final predicted vocabulary; if the predicted text vocabulary exceeds the vocabulary range corresponding to the text level, the final predicted vocabulary is determined based on the vocabulary range corresponding to the text level.

[0051] In an embodiment of the present application, the target text of the target user is evaluated for its text level by adopting a level evaluation model, and the text vocabulary corresponding to the target text is determined in combination with a vocabulary prediction model. Finally, the predicted vocabulary of the target user is determined based on the level vocabulary corresponding to the text level of the target text and the text vocabulary corresponding to the target text. By adopting the user's text, starting from the text perspective that reflects the user's understanding of vocabulary, and combining the level evaluation model and the vocabulary prediction model to predict the user's vocabulary level, the rationality of the vocabulary prediction is guaranteed and the accuracy of the vocabulary prediction is improved.

[0052] In the embodiments of the present application, before obtaining the target text for the target user, it is necessary to first train the grade assessment model and the vocabulary prediction model to implement the process of using the grade assessment model and the vocabulary prediction model to perform grade assessment and vocabulary prediction on the target text. The following embodiments focus on how to train the grade assessment model and the vocabulary prediction model.

[0053] See Figure 2 , which is a flowchart of a vocabulary prediction method according to an embodiment of the present application. Figure 2 As shown, the method may include the following steps S201 to S209.

[0054] S201: Receive a first number of first training texts and a first text level corresponding to the first training texts, and obtain first training words in the first training texts and a sequence relationship between the first training words.

[0055] Specifically, the first training text includes at least one first training vocabulary and the order relationship between each first training vocabulary. It can be understood that each first training vocabulary is divided into at least one first training sentence according to the punctuation marks in the text. In the first training text, the word dependency relationship within each first training sentence is legal and meets the first text level corresponding to the first training text; similarly, the sentence structure within each first training sentence is legal and meets the first text level corresponding to the first training text. Optionally, in order to ensure the training results of the grade evaluation model, a large number of first pre-training texts can be obtained for training. Exemplarily, the first number of first training texts can be 100,000 non-repetitive excellent texts, which is not limited here. Exemplarily, the first text levels corresponding to the first training text include elementary school level, junior high school level, high school level, level 4 level and level 6 level, etc.

[0056] The vocabulary prediction device first obtains a first number of first training texts and a first text level corresponding to the first training texts, and then obtains first training words in the first training texts and a sequence relationship between the first training words.

[0057] S202 : Based on the first training words in the first training text and the order relationship between the first training words, and the first text level corresponding to the first training text, a level evaluation model is trained to obtain the trained level evaluation model.

[0058] Specifically, the vocabulary prediction device takes the first training text and the first text level corresponding to each first training text as a group of first training data, wherein the first training text in the first training data can be each first training vocabulary and the sequential relationship between each first training vocabulary, or it can be a text file of the first training text. It should be noted that if the input first training text is input into the level evaluation model in the format of a text file, the level evaluation model has the ability to recognize the text content in the text file.

[0059] The vocabulary prediction device sequentially inputs each first training data into the pre-trained grade evaluation model to train the grade evaluation model using each first training data to obtain a trained grade evaluation model.

[0060] Exemplarily, the pre-training grade evaluation model can be a pre-trained language model ELECTRAL, and correspondingly, the grade evaluation model can be a five-layer model, namely the Muti-Head Attention layer, the first ADD&Norm layer, the FeedForward layer, the second ADD&Norm layer and the softmax layer; then the Muti-Head Attention layer is used to learn the word dependencies and sentence structure in the first training text, and then the first ADD&Norm layer is used to perform residual connections to prevent the model from solving the problem of gradient disappearance and the problem of weight matrix degradation, and the obtained data is normalized; then the FeedForward layer is used to train the weight parameters in the activation function; finally, the second ADD&Norm layer is used to perform the process as performed by the first ADD&Norm layer to obtain the vector representation of the first training text, which represents the semantic information of the first training text, and then the softmax layer is used to classify the vector representation of the first training text based on the text level of the first training text. It should be noted that the above description of the grade evaluation model is only an example and does not impose any limitation on the grade evaluation model.

[0061] Thus, by obtaining a large amount of first training data to train the initial grade evaluation model (i.e., the grade evaluation model before training), a trained grade evaluation model is obtained to improve the accuracy of the grade evaluation results output by the grade evaluation model based on the target text, thereby improving the prediction accuracy of the vocabulary size.

[0062] S203: Receive a second number of second training texts and second text levels corresponding to the second training texts, and obtain each second training vocabulary in the second training texts and a vocabulary level corresponding to each second training vocabulary.

[0063] Specifically, the second training text includes at least one second training vocabulary. The vocabulary level refers to the level information of the corresponding vocabulary in the vocabulary grading table. Optionally, the vocabulary grading table can be a frequency level table of words, that is, the frequency of use of words in daily life is counted, and then the words are graded according to the frequency of use; illustratively, the vocabulary grading table can have 5 levels: level 0, level 1, level 2, level 3 and level 4, and the number of words corresponding to each vocabulary level is: level 0 is 1000, level 1 is 2000, level 2 is 4000, level 3 is 8000, and level 4 is 38000.

[0064] The vocabulary prediction device obtains a second number of second training texts and a text grade corresponding to each second training text, then obtains second training words in each second training text, and determines a vocabulary grade corresponding to each second training word based on a vocabulary grade table.

[0065] Optionally, to ensure the training results of the vocabulary prediction model, a large number of second pre-training texts can be obtained for training. Exemplarily, the second number of second training texts can be 100,000 non-repetitive texts, which is not limited here. Exemplarily, the second training texts correspond to second text levels including elementary school level, junior high school level, high school level, CET-4 level, and CET-6 level.

[0066] S204: Acquire a vocabulary of a training level corresponding to the second training text based on a second text level corresponding to the second training text.

[0067] Specifically, the text levels can be divided into primary school level, junior high school level, senior high school level, CET-4 level and CET-6 level, etc. The vocabulary range corresponding to each text level can be 0-1000 for primary school level, 1000-3000 for junior high school level, 3000-4000 for senior high school level, 4000-5500 for CET-4 level, and 5500-7000 for CET-6 level, which are not limited here. For example, the training level vocabulary corresponding to each text level can be the upper limit of the range of the vocabulary range corresponding to each text level, that is, 1000 for primary school level, 3000 for junior high school level, 4000 for senior high school level, 5500 for CET-4 level, and 7000 for CET-6 level, which are not limited here.

[0068] When obtaining the second text level corresponding to the second training text, the vocabulary prediction device obtains the training level vocabulary corresponding to the second text level.

[0069] S205 , based on each second training word in the second training text, the vocabulary level corresponding to each second training word, and the vocabulary of the training level corresponding to the second training text, perform model training on the vocabulary prediction model to obtain the trained vocabulary prediction model.

[0070] Specifically, the vocabulary prediction device takes the vocabulary level corresponding to each second training vocabulary in the second training text and the training level vocabulary corresponding to the second training text as a set of second training data, and then inputs each second training data into the pre-trained vocabulary prediction model in turn, so as to train the vocabulary prediction model with the second training data and obtain a trained vocabulary prediction model.

[0071] Exemplarily, parameter training is performed on prediction parameters in the vocabulary prediction model using each second training data, and the prediction parameters are shown as W in the following formula.

[0072]

[0073] Among them, V(arg max(C(i))-1) refers to the number of words corresponding to the previous vocabulary level of the vocabulary level with the largest number of words in the second training text. For example, V(-1)=0, V(0)=1000, V(2)=2000, V(3)=8000; C(i) is the number of words of each vocabulary level in the second training text, i is the vocabulary level, for example, C(0) refers to the number of words of level 0 in the second training text; argmax() is a function used to determine the vocabulary level with the largest number of words; N is the total number of words of each second training vocabulary in the second training text; W is the prediction parameter of the vocabulary prediction model.

[0074] Thus, by obtaining a large amount of second training data to train the initial vocabulary prediction model (i.e., the vocabulary prediction model before training), a trained vocabulary prediction model is obtained to improve the accuracy of the vocabulary prediction results output by the vocabulary model based on the target text, thereby improving the accuracy of vocabulary prediction.

[0075] S206: Acquire a target text for a target user, and acquire first target words in the target text and a sequence relationship between the first target words.

[0076] Please refer to step S101 for details, which will not be described again here.

[0077] S207 , using a grade evaluation model to obtain a target text grade of the target text based on the first target words and the order relationship between the first target words.

[0078] Please refer to step S102 for details, which will not be described again here.

[0079] S208 : Using a vocabulary prediction model, based on the first target words and the vocabulary levels corresponding to the first target words, obtain the text vocabulary corresponding to the target text.

[0080] Please refer to step S103 for details, which will not be described again here.

[0081] S209 : Determine the predicted vocabulary size of the target user based on the target level vocabulary size corresponding to the target text level and the text vocabulary size.

[0082] Please refer to step S104 for details, which will not be described again here.

[0083] In an embodiment of the present application, a target text of a target user is evaluated for its text level by adopting a level evaluation model, and a text vocabulary corresponding to the target text is determined in combination with a vocabulary prediction model. Finally, a predicted vocabulary of the target user is determined based on the level vocabulary corresponding to the text level of the target text and the text vocabulary corresponding to the target text. By adopting the user's text, starting from a text perspective that reflects the user's understanding of vocabulary, and combining the level evaluation model and the vocabulary prediction model to predict the user's vocabulary level, the rationality of the vocabulary prediction is ensured and the accuracy of the vocabulary prediction is improved. A large amount of first training data is obtained to train the initial level evaluation model (i.e., the level evaluation model before training) to obtain a trained level evaluation model, so as to improve the accuracy of the level evaluation results output by the level evaluation model based on the target text, thereby improving the prediction accuracy of the vocabulary. A large amount of second training data is obtained to train the initial vocabulary prediction model (i.e., the vocabulary prediction model before training) to obtain a trained vocabulary prediction model, so as to improve the accuracy of the vocabulary prediction results output by the vocabulary model based on the target text, thereby improving the prediction accuracy of the vocabulary.

[0084] In an embodiment of the present application, when obtaining a second training text, the second training text can be first subjected to lexical processing to improve the training effect of the vocabulary prediction model. Furthermore, when training the vocabulary prediction model, data statistics of each second training word can be first collected to reduce the data processing process of the vocabulary prediction model, thereby improving the training efficiency of the vocabulary prediction model. The following embodiment focuses on how to perform lexical processing on the second training text and how to collect data statistics on the second training words.

[0085] See Figure 3 , provides a flow chart of obtaining a vocabulary prediction model for the embodiment of the present application. Figure 3 As shown, the method may include the following steps S301 to S306.

[0086] S301: Receive a second number of second training texts and a second text level corresponding to the second training texts.

[0087] Specifically, the vocabulary prediction device obtains a second number of second training texts and text levels corresponding to the second training texts, and then obtains initial training vocabulary in each second training text.

[0088] Optionally, to ensure the training results of the vocabulary prediction model, a large number of second pre-training texts can be obtained for training. Exemplarily, the second number of second training texts can be 100,000 non-repetitive texts, which is not limited here. Exemplarily, the second training texts correspond to second text levels including elementary school level, junior high school level, high school level, CET-4 level, and CET-6 level.

[0089] S302: Performing vocabulary processing on each initial training vocabulary in the second training text to obtain each second training vocabulary corresponding to the second training text, wherein the vocabulary processing includes at least one of vocabulary normalization processing, keyword removal processing, and vocabulary deduplication processing.

[0090] Specifically, the second training text includes at least one initial training vocabulary. Vocabulary normalization refers to the process of modifying different vocabulary with the same root, such as different tenses, different parts of speech, singular and plural forms of nouns, into the stems corresponding to the vocabulary. For example, "needs" is stemmed / normalized to "need", "got" is stemmed / normalized to "get", "interesting" and "interestingly" are stemmed / normalized to "interest", etc. Keyword removal refers to the removal of stop words and low-level common words, etc. For example, a, an, the, etc. are all low-level common words. Vocabulary deduplication refers to removing the same vocabulary and retaining one of the vocabulary.

[0091] Optionally, different vocabulary processing methods have different priority orders, that is, keyword removal processing has the highest priority, vocabulary normalization processing has a medium priority, and vocabulary deduplication processing has the lowest priority. For example, keyword removal processing can be used to first reduce the number of initial training vocabulary in the second training text, and then vocabulary normalization processing can be used to normalize the initial training vocabulary after keyword removal processing to vocabulary with the same stem, so as to more accurately remove repeated vocabulary in the initial training vocabulary after vocabulary normalization processing in the vocabulary deduplication processing.

[0092] The vocabulary prediction device selects a vocabulary processing method for each initial training vocabulary in the second training text, then sequentially obtains the second training text, and performs vocabulary processing on each initial training vocabulary in the second training text to obtain second training vocabulary.

[0093] S303: Obtain the vocabulary level corresponding to each second training vocabulary.

[0094] Specifically, vocabulary level refers to the level information of the corresponding vocabulary in the vocabulary level table. Optionally, the vocabulary level table can be a word usage frequency level table, which counts the frequency of word usage in daily life and then classifies the words according to their frequency of use. For example, the vocabulary level table can have five levels: level 0, level 1, level 2, level 3, and level 4, and the number of words corresponding to each vocabulary level is: level 0 is 1000, level 1 is 2000, level 2 is 4000, level 3 is 8000, and level 4 is 38,000.

[0095] The vocabulary prediction device sequentially obtains second training words in the second training text, and determines the vocabulary level corresponding to each second training word based on the vocabulary level table.

[0096] Therefore, by performing vocabulary preprocessing on the second training text to remove the initial training vocabulary in the second training text that will affect the vocabulary prediction results, the impact on the training process of the vocabulary prediction model is reduced, so as to improve the accuracy of the vocabulary prediction results output by the vocabulary model based on the target text, thereby improving the vocabulary prediction accuracy.

[0097] S304: Acquire a vocabulary of a training level corresponding to the second training text based on a second text level corresponding to the second training text.

[0098] Specifically, the text levels can be divided into primary school level, junior high school level, senior high school level, CET-4 level and CET-6 level, etc. The vocabulary range corresponding to each text level can be 0-1000 for primary school level, 1000-3000 for junior high school level, 3000-4000 for senior high school level, 4000-5500 for CET-4 level, and 5500-7000 for CET-6 level, which are not limited here. For example, the training level vocabulary corresponding to each text level can be the upper limit of the range of the vocabulary range corresponding to each text level, that is, 1000 for primary school level, 3000 for junior high school level, 4000 for senior high school level, 5500 for CET-4 level, and 7000 for CET-6 level, which are not limited here.

[0099] When obtaining the second text level corresponding to the second training text, the vocabulary prediction device obtains the training level vocabulary corresponding to the second text level.

[0100] S305 : Based on each second training vocabulary in the second training text and the vocabulary level corresponding to each second training vocabulary, obtain the total number of training vocabulary in the second training text and the number of training vocabulary at each vocabulary level.

[0101] Specifically, the vocabulary prediction device sequentially counts the total number of training words of the second training words in the second training text, and then counts the number of training words of each vocabulary level in the second training text based on the vocabulary level corresponding to each second training word.

[0102] S306: Based on the total number of training words in the second training text, the number of training words at each vocabulary level, and the vocabulary of the training level corresponding to the second training text, a vocabulary prediction model is trained to obtain the trained vocabulary prediction model.

[0103] Specifically, the vocabulary prediction device takes the total number of training words of the second training vocabulary in the second training text, the number of training words of each vocabulary level in the second training text, and the training level vocabulary corresponding to the second training text as a set of second training data, and then inputs each second training data into the pre-trained vocabulary prediction model in turn, so as to train the vocabulary prediction model with the second training data and obtain a trained vocabulary prediction model.

[0104] Exemplarily, parameter training is performed on prediction parameters in the vocabulary prediction model using each second training data, and the prediction parameters are shown as W in the following formula.

[0105]

[0106] Among them, V(arg max(C(i))-1) refers to the number of words corresponding to the previous vocabulary level of the vocabulary level with the largest number of words in the second training text. For example, V(-1)=0, V(0)=1000, V(2)=2000, V(3)=8000; C(i) is the number of words of each vocabulary level in the second training text, i is the vocabulary level, for example, C(0) refers to the number of words of level 0 in the second training text; argmax() is a function used to determine the vocabulary level with the largest number of words; N is the total number of words of each second training vocabulary in the second training text; W is the prediction parameter of the vocabulary prediction model.

[0107] Therefore, the initial vocabulary prediction model (i.e., the vocabulary prediction model before training) is trained through the second training data after vocabulary preprocessing to obtain a trained vocabulary prediction model, so as to improve the accuracy of the vocabulary prediction results output by the vocabulary model based on the target text, thereby improving the vocabulary prediction accuracy.

[0108] In an embodiment of the present application, a target text of a target user is evaluated for its text grade by using a grade evaluation model, and a text vocabulary corresponding to the target text is determined in combination with a vocabulary prediction model. Finally, a predicted vocabulary of the target user is determined based on the grade vocabulary corresponding to the text grade of the target text and the text vocabulary corresponding to the target text. By using the user's text, starting from a text perspective that reflects the user's degree of understanding of vocabulary, and combining the grade evaluation model and the vocabulary prediction model to predict the user's vocabulary level, the rationality of the vocabulary prediction is ensured and the accuracy of the vocabulary prediction is improved. By performing vocabulary preprocessing on the second training text to remove the initial training vocabulary in the second training text that will affect the vocabulary prediction result, the impact on the training process of the vocabulary prediction model is reduced, so as to improve the accuracy of the vocabulary prediction result output by the vocabulary prediction model based on the target text, thereby improving the vocabulary prediction accuracy. The initial vocabulary prediction model (i.e., the vocabulary prediction model before training) is trained using the second training data after vocabulary preprocessing to obtain a trained vocabulary prediction model, so as to improve the accuracy of the vocabulary prediction result output by the vocabulary prediction model based on the target text, thereby improving the vocabulary prediction accuracy.

[0109] In the embodiments of the present application, when predicting the text vocabulary corresponding to the target text, the target text can be first lexically processed to improve the prediction accuracy of the vocabulary prediction model. When obtaining the predicted vocabulary, the final predicted vocabulary can be determined by comparing the target level vocabulary with the text vocabulary. The following embodiments focus on how to perform lexical processing on the target text and how to obtain the predicted vocabulary.

[0110] See Figure 4 , which is a flowchart of a vocabulary prediction method according to an embodiment of the present application. Figure 4 As shown, the method may include the following steps S401 to S407.

[0111] S401: Acquire a target text for a target user, and acquire first target words in the target text and a sequence relationship between the first target words.

[0112] Please refer to step S101 for details, which will not be repeated here.

[0113] S402 : Using a grade evaluation model, based on the first target words and the order relationship between the first target words, obtain the target text grade of the target text.

[0114] Please refer to step S102 for details, which will not be repeated here.

[0115] S403 , performing lexical processing on each first target word in the target text to obtain a second target word corresponding to the target text, wherein the lexical processing includes at least one of lexical normalization processing, keyword removal processing, and lexical deduplication processing.

[0116] Specifically, the target training text includes at least one first target vocabulary. Vocabulary normalization refers to the process of modifying different vocabulary with different tenses, different parts of speech, singular and plural forms, etc., which belong to the same root, into the stems corresponding to the vocabulary. For example, "needs" is stemmed / normalized to "need", "got" is stemmed / normalized to "get", "interesting" and "interestingly" are stemmed / normalized to "interest", etc. Keyword removal refers to the removal of stop words and low-level common words, etc. For example, a, an, the, etc. are all low-level common words. Vocabulary deduplication refers to removing the same vocabulary and retaining one vocabulary.

[0117] Optionally, different vocabulary processing methods have different priority orders, that is, keyword removal processing has the highest priority, vocabulary normalization processing has a medium priority, and vocabulary deduplication processing has the lowest priority. For example, keyword removal processing can be used to first reduce the number of first target words in the second training text, and then vocabulary normalization processing can be used to normalize the first target words after keyword removal processing to words with the same stem, so as to more accurately remove repeated words in the first target words after vocabulary normalization processing in the vocabulary deduplication processing.

[0118] The vocabulary prediction device selects a vocabulary processing method required for each first target vocabulary in the target text, then sequentially obtains the target text, and performs vocabulary processing on each first target vocabulary in the target text to obtain a second target vocabulary.

[0119] S404: Obtain the vocabulary level corresponding to each second target vocabulary word.

[0120] Specifically, vocabulary level refers to the level information of the corresponding vocabulary in the vocabulary level table. Optionally, the vocabulary level table can be a word usage frequency level table, which counts the frequency of word usage in daily life and then classifies the words according to their frequency of use. For example, the vocabulary level table can have five levels: level 0, level 1, level 2, level 3, and level 4, and the number of words corresponding to each vocabulary level is: level 0 is 1000, level 1 is 2000, level 2 is 4000, level 3 is 8000, and level 4 is 38,000.

[0121] The vocabulary prediction device sequentially obtains second target words in the target text and determines the vocabulary level corresponding to each second target word based on the vocabulary level table.

[0122] Therefore, by performing vocabulary preprocessing on the target text to remove the first target vocabulary in the target text that will affect the vocabulary prediction results, the impact on the training process of the vocabulary prediction model is reduced, so as to improve the accuracy of the vocabulary prediction results output by the vocabulary model based on the target text, thereby improving the vocabulary prediction accuracy.

[0123] S405 , using the vocabulary prediction model, based on the second target words and the vocabulary levels corresponding to the second target words, obtain the text vocabulary corresponding to the target text.

[0124] Specifically, the vocabulary prediction device counts the total number of target words of the second target vocabulary in the target text, and then counts the number of target words of each vocabulary level in the target text based on the vocabulary level corresponding to the second target vocabulary, and then calculates the number of target words of each vocabulary level in the target text, and then calculates the number of target words of the second target vocabulary in the target text and the vocabulary prediction model.

[0125] The vocabulary prediction device sequentially inputs the total number of words corresponding to each second target word in the target text and the number of target words corresponding to each vocabulary level into the vocabulary prediction model to predict the text vocabulary corresponding to the target text through the vocabulary prediction model.

[0126] Exemplarily, the vocabulary prediction model predicts the vocabulary of the target text based on the following formula.

[0127]

[0128] Wherein, V(arg max(C(i))-1) refers to the number of words corresponding to the previous vocabulary level of the vocabulary level with the largest number of words in the target text. For example, V(-1)=0, V(0)=1000, V(2)=2000, V(3)=8000; C(i) is the number of target words of each vocabulary level in the target text, i is the vocabulary level, for example, C(0) refers to the number of level 0 words in the target text; argmax() is a function used to determine the vocabulary level with the largest number of words; N is the total number of words of each second target vocabulary in the target text; W is the prediction parameter of the vocabulary prediction model.

[0129] Therefore, by performing vocabulary preprocessing on the target text, the first target vocabulary in the target text that will affect the vocabulary prediction results is removed, thereby reducing the impact on the vocabulary prediction results of the vocabulary prediction model, so as to improve the accuracy of the vocabulary prediction results output by the vocabulary model based on the target text, thereby improving the vocabulary prediction accuracy.

[0130] S406: If the target level vocabulary is greater than or equal to the text vocabulary, determine the text vocabulary as the predicted vocabulary.

[0131] Specifically, the target level vocabulary is the level vocabulary corresponding to the target text level, which can be the upper limit of the vocabulary range corresponding to the target text level. For example, the level vocabulary corresponding to each text level can be 1000 for primary school level, 3000 for junior high school level, 4000 for high school level, 5500 for level 4, and 7000 for level 6. It can also be calculated based on the vocabulary range corresponding to the text level, which is not limited here.

[0132] The vocabulary prediction device first obtains the target-level vocabulary corresponding to the target text level, then compares the target text level with the text vocabulary to determine whether the text vocabulary size predicted by the vocabulary prediction model is too high. If the target-level vocabulary size is greater than or equal to the text vocabulary size, the text vocabulary size predicted by the vocabulary prediction model is determined to be a reasonable value, and the text vocabulary size is output as the final predicted vocabulary size.

[0133] Specifically, because the text level of the target text written by the target user directly reflects the target user's vocabulary level, the text vocabulary predicted based on the target text must fall within the vocabulary range corresponding to the text level. If the text vocabulary predicted based on the target text falls within the vocabulary range obtained based on the target text's text level, it means that the target text does not contain, or only contains a small number of advanced vocabulary that would increase the predicted vocabulary level but is beyond the target user's grasp.

[0134] S407: If the target level vocabulary is smaller than the text vocabulary, determine the target level vocabulary as the predicted vocabulary.

[0135] The vocabulary prediction device first obtains the target-level vocabulary corresponding to the target text level, then compares the target text level with the text vocabulary to determine whether the text vocabulary size predicted by the vocabulary prediction model is too high. If the target-level vocabulary size is less than the text vocabulary size, the text vocabulary size predicted by the vocabulary prediction model is determined to be unreasonable, and the target-level vocabulary size is output as the final predicted vocabulary size.

[0136] Specifically, because the text level of the target text written by the target user directly reflects the target user's vocabulary level, the text vocabulary predicted based on the target text must fall within the vocabulary range corresponding to the text level. If the text vocabulary predicted based on the target text exceeds the vocabulary range obtained based on the text level of the target text, it means that the target text contains a large number of advanced words that could increase the vocabulary prediction result but are beyond the target user's grasp.

[0137] Therefore, by comparing the target level vocabulary and the text vocabulary, it is determined whether the predicted text vocabulary exceeds the vocabulary range corresponding to the text level, so as to identify whether there are a large number of pieced-together high-level words in the target text, and obtain the predicted vocabulary based on the recognition results, thereby avoiding the influence of pieced-together high-level words on the vocabulary prediction results, and then obtaining a vocabulary prediction result that can better reflect the actual vocabulary level of the target user, thereby improving the accuracy of vocabulary prediction.

[0138] In an embodiment of the present application, a target text of a target user is evaluated for its text grade by adopting a grade evaluation model, and a text vocabulary corresponding to the target text is determined in combination with a vocabulary prediction model. Finally, a predicted vocabulary of the target user is determined based on the grade vocabulary corresponding to the text grade of the target text and the text vocabulary corresponding to the target text. By adopting the user's text, starting from a text perspective that reflects the user's degree of understanding of vocabulary, and combining the grade evaluation model and the vocabulary prediction model to predict the user's vocabulary level, the rationality of the vocabulary prediction is ensured and the accuracy of the vocabulary prediction is improved. By performing vocabulary preprocessing on the target text to remove the first target vocabulary in the target text that will affect the vocabulary prediction result, the influence on the vocabulary prediction result of the vocabulary prediction model is reduced, so as to improve the vocabulary output of the vocabulary model based on the target text. The accuracy of vocabulary prediction results is improved, thereby improving the accuracy of vocabulary prediction; by performing vocabulary preprocessing on the target text to remove the first target vocabulary in the target text that will affect the vocabulary prediction results, thereby reducing the impact on the vocabulary prediction results of the vocabulary prediction model, so as to improve the accuracy of vocabulary prediction results output by the vocabulary model based on the target text, thereby improving the accuracy of vocabulary prediction; by comparing the target level vocabulary and the text vocabulary, it is judged whether the predicted text vocabulary exceeds the vocabulary range corresponding to the text level, so as to identify whether there are a large number of spliced ​​high-level words in the target text, and obtain the predicted vocabulary based on the recognition result, thereby avoiding the impact of the spliced ​​high-level words on the vocabulary prediction results, and then obtaining a vocabulary prediction result that can better reflect the real vocabulary level of the target user, thereby improving the accuracy of vocabulary prediction.

[0139] In the embodiments of the present application, before performing vocabulary estimation based on the target text, it is necessary to first obtain the text perplexity of the target text, thereby determining whether to perform vocabulary estimation based on the target text based on the text perplexity. The following embodiments focus on how to train a perplexity estimation model, how to obtain a perplexity threshold based on the perplexity estimation model, how to obtain the text perplexity of the target text, and how to determine whether to perform vocabulary estimation based on the target text based on the perplexity threshold and text perplexity.

[0140] See Figure 5 , which is a flowchart of a vocabulary prediction method according to an embodiment of the present application. Figure 5 As shown, the method may include the following steps S501-S510.

[0141] S501 : Obtain a third number of third training texts, obtain third training words in the third training texts and the order relationship between the third training words.

[0142] Specifically, the third training text includes at least one third training vocabulary and a sequential relationship between each third training vocabulary. It is understood that each third training vocabulary is segmented into at least one third training sentence based on punctuation marks in the text. In the third training text, the word dependency relationship within each third training sentence is valid.

[0143] Optionally, in order to ensure the training results of the perplexity assessment model, a large amount of third pre-training texts can be obtained for training. Exemplarily, the third amount of third training texts can be 100,000 non-repetitive excellent texts, which is not limited here.

[0144] The vocabulary prediction device first obtains a third number of third training texts, and then obtains each third training word in the third training texts and the order relationship between each third training word.

[0145] S502 : Based on the third training words in the third training text and the order relationship between the third training words, a perplexity assessment model is trained to obtain the trained perplexity assessment model.

[0146] Specifically, the vocabulary prediction device uses each third training word in the third training text and the order relationship between the third training words as a set of third training data. The third training data are then sequentially input into a pre-trained perplexity assessment model. The perplexity assessment model is then trained based on the third training data to obtain the probability of a subsequent word that may appear after each third training word. It should be noted that the probabilities of the subsequent words are used to calculate the text perplexity of the input text.

[0147] Among them, the text perplexity is used to indicate the fluency of the text content of the input text. The higher the text perplexity, the less fluent the text content is. It may be a text that is randomly pieced together. It is not worthwhile to predict the vocabulary based on the text. In other words, it is impossible to predict the accurate vocabulary based on the text. On the contrary, the lower the text perplexity, the more fluent the text content is. It may be a text that is carefully written by the user. It is worthwhile to predict the vocabulary based on the text. In other words, it is possible to predict the accurate vocabulary based on the text.

[0148] Exemplarily, the perplexity assessment model before training can be a statistical language model or a GPT-2 model, which is not limited here. It should be noted that in the training process of the perplexity assessment model, the probability of the subsequent words that may appear after each third training word is calculated based on a large number of third training words and the sequential relationship between each third training word. For example, based on a large number of third training texts, the subsequent words that may appear after "last" and their probabilities are calculated as follows: "day" 10%, "week" 8%, "month" 7%, "year" 5%, etc., which will not be repeated here. It should be emphasized that the above examples are only for illustrative purposes and do not impose any limitations on this solution.

[0149] Therefore, a large amount of third training data is obtained to train the perplexity assessment model, and a trained perplexity assessment model is obtained to improve the accuracy of the perplexity assessment results output by the perplexity assessment model based on the target text. The target text used to evaluate the vocabulary is screened through the text perplexity output by the perplexity assessment model, thereby avoiding the situation where vocabulary prediction is performed on randomly pieced together texts, thereby improving the accuracy of vocabulary prediction.

[0150] S503: Obtain a fourth number of disordered texts, and obtain disordered words in the disordered texts and the order relationship between the disordered words.

[0151] Specifically, the disordered text includes at least one disordered word and the order relationship between the disordered words. It should be noted that, in the disordered text, the word dependency relationship within each disordered sentence is illegal. Exemplarily, a fourth number of texts to be disordered can be randomly selected from the third training text, and then the third training words in the texts to be disordered can be randomly shuffled to obtain the disordered text. For example, assuming that the text content of the text to be disordered is "Last year, I visited the GreatWall with my friends on May Day holiday.", then after randomly shuffling the text to be disordered, the disordered text is obtained, and the text content of the disordered text is "friends Wall holiday May year Great the Last, I withon my Day visited". It should be noted that the above description is only for illustrative purposes and is not limited here.

[0152] S504: Using a perplexity evaluation model, based on the perplexity of the perplexity words in the perplexity text and the order relationship between the perplexity words, obtain the perplexity of the perplexity text.

[0153] Specifically, the vocabulary prediction device sequentially inputs each disordered text into the perplexity evaluation model to obtain the disordered text perplexity of each disordered text through the perplexity evaluation model.

[0154] It should be noted that the confusion evaluation model obtains the probability value of the subsequent word of each disordered word based on the order relationship between each disordered word in the disordered text.

[0155] For example, if the current scrambled word is "last," and the subsequent word in the scrambled text is "year," the probability of "year" is 5%. If the subsequent word in the scrambled text is "I," the probability of "I" is 0%. For example, if the scrambled text is "friends Wall holiday May year Great the Last, I with on my Day visited," the perplexity assessment model sequentially calculates the probability of "Wall" following "friends," then the probability of "holiday" following "Wall," then the probability of "May" following "holiday," and so on.

[0156] The perplexity assessment model then calculates the perplexity of the scrambled text based on the probability values ​​of each successive word. For example, the perplexity assessment model obtains the logarithmic probability value L1(i) of the probability value L0(i) of each successive word, then obtains the absolute value L2(i) of the logarithmic probability value L1(i), adds each L2(i) together to obtain a total probability value L3, then takes the average of the total probability value L3 and the number of successive words to obtain a probability mean L4, and finally exponentiates the total probability value L3 with e as the base to obtain the perplexity of the scrambled text PPL.

[0157] S505: Obtain a perplexity threshold based on the scrambled text perplexity of the scrambled text.

[0158] Specifically, after the vocabulary prediction model obtains the random text perplexity PPL of each random text, it adds up the perplexity PPL of each random text to obtain the total random text perplexity, and then calculates the average based on the number of texts in the random text and the total random text perplexity to obtain the perplexity threshold.

[0159] It should be noted that the perplexity threshold is the basis for determining whether the input text is a randomly assembled text. Specifically, when the text perplexity of the input text is greater than the perplexity threshold, the input text is determined to be a randomly assembled text; when the text perplexity of the input text is less than or equal to the perplexity threshold, the input text is determined to be not a randomly assembled text / fluent text.

[0160] Therefore, by obtaining a large amount of disordered text to calculate the perplexity threshold, that is, to determine the perplexity range of randomly pieced-together text, the vocabulary prediction device can judge whether the input text is a randomly pieced-together text through the text perplexity of the input text and the perplexity threshold, thereby screening the target text used to evaluate the vocabulary, avoiding the situation of vocabulary prediction for randomly pieced-together text, and thus improving the accuracy of vocabulary prediction.

[0161] S506: Acquire a target text for a target user, and acquire first target words in the target text and a sequence relationship between the first target words.

[0162] Please refer to step S101 for details, which will not be repeated here.

[0163] S507 , using a perplexity evaluation model to obtain a target text perplexity of the target text based on the first target words and the order relationship between the first target words.

[0164] Specifically, the vocabulary prediction device inputs the target text into the perplexity evaluation model to obtain the text perplexity of the target text through the perplexity evaluation model.

[0165] It should be noted that the confusion assessment model obtains the probability value of the subsequent words of each target word based on the sequential relationship between each target word in the target text.

[0166] For example, if the current target word is "last," and the following word in the target text is "year," the probability of "year" is 5%. If the following word in the target text is "I," the probability of "I" is 0%. For example, if the target text is "Last year, I visited the Great Wall with my friends on May Day holiday," the perplexity assessment model sequentially calculates the probability of "year" following "Last," then the probability of "visited" following "I," then the probability of "the" following "visited," and so on.

[0167] The perplexity assessment model then calculates the target text perplexity of the target text based on the probability values ​​of each subsequent word. For example, the perplexity assessment model obtains the logarithmic probability value L1(i) of the probability value L0(i) of each target word's subsequent word, then obtains the absolute value L2(i) of the logarithmic probability value L1(i), adds each L2(i) together to obtain a total probability value L3, then takes the average of the total probability value L3 and the number of subsequent words to obtain a probability mean L4. Finally, the perplexity is exponentially calculated with base e to obtain the target text perplexity.

[0168] S508: If the target text perplexity is less than or equal to the perplexity threshold, a grade evaluation model is used to obtain a target text grade of the target text based on the first target words and the order relationship between the first target words.

[0169] S509: If the perplexity of the target text is less than or equal to the perplexity threshold, a vocabulary prediction model is used to obtain the text vocabulary corresponding to the target text based on the first target words and the vocabulary levels corresponding to the first target words.

[0170] Specifically, the target text perplexity is compared with the perplexity threshold. If the target text perplexity is less than or equal to the perplexity threshold, the target text is determined to be fluent, and a highly accurate vocabulary can be predicted based on the target text. A rank assessment model is then used to determine the target text rank based on the order of each first target vocabulary word and the order of each first target vocabulary word. A vocabulary prediction model is also used to determine the text vocabulary of the target text based on each first target vocabulary word and the vocabulary rank corresponding to each first target vocabulary word.

[0171] Furthermore, the vocabulary prediction device compares the perplexity of the target text with the perplexity threshold. If the perplexity of the target text is greater than the perplexity threshold, the target text is determined to be a randomly pieced-together text, and only a vocabulary prediction result with a low accuracy can be obtained based on the target text. Then, an error prompt is directly output, or 0 is output as the vocabulary prediction result.

[0172] S510: Determine a predicted vocabulary size of the target user based on the target level vocabulary size corresponding to the target text level and the text vocabulary size.

[0173] Please refer to step S104 for details, which will not be described again here.

[0174] Therefore, by comparing the target text perplexity and the perplexity threshold of the target text, it is determined whether the target text is a randomly assembled text, and the target text used to evaluate the vocabulary is screened, avoiding the situation of vocabulary prediction for randomly assembled texts, thereby improving the accuracy of vocabulary prediction.

[0175] In an embodiment of the present application, a target text of a target user is evaluated for its text grade by adopting a grade evaluation model, and a text vocabulary corresponding to the target text is determined in combination with a vocabulary prediction model. Finally, a predicted vocabulary of the target user is determined based on the grade vocabulary corresponding to the text grade of the target text and the text vocabulary corresponding to the target text. By adopting the user's text, starting from a text perspective that reflects the user's degree of understanding of vocabulary, and combining the grade evaluation model and the vocabulary prediction model to predict the user's vocabulary level, the rationality of the vocabulary prediction is ensured and the accuracy of the vocabulary prediction is improved. A large amount of third training data is obtained to train the perplexity evaluation model, and a trained perplexity evaluation model is obtained to improve the accuracy of the perplexity evaluation result output by the perplexity evaluation model based on the target text, thereby improving the accuracy of the text perplexity evaluation result output by the perplexity evaluation model. The perplexity of the target text used to evaluate the vocabulary is screened to avoid the situation where vocabulary prediction is performed on randomly assembled texts, thereby improving the prediction accuracy of the vocabulary; the perplexity threshold is calculated by obtaining a large amount of disordered text, that is, the perplexity range of the randomly assembled text is determined, so that the vocabulary prediction device can judge whether the input text is a randomly assembled text by the text perplexity of the input text and the perplexity threshold, thereby screening the target text used to evaluate the vocabulary, avoiding the situation where vocabulary prediction is performed on randomly assembled texts, thereby improving the prediction accuracy of the vocabulary; by comparing the target text perplexity of the target text and the perplexity threshold, it is judged whether the target text is a randomly assembled text, thereby screening the target text used to evaluate the vocabulary, avoiding the situation where vocabulary prediction is performed on randomly assembled texts, thereby improving the prediction accuracy of the vocabulary.

[0176] In the embodiments of this application, it is necessary to first train the perplexity assessment model, the level assessment model, and the vocabulary prediction model, and then perform corresponding evaluation and prediction based on the trained models to determine the final predicted vocabulary size based on the output results of each model. The following embodiments focus on how to perform model training and how to obtain the predicted vocabulary size.

[0177] See Figure 6 , which is a flowchart of a vocabulary prediction method according to an embodiment of the present application. Figure 6 As shown, the method may include the following steps S1 to S12.

[0178] S1, obtain model training data and threshold calculation data.

[0179] A certain amount of first model training data, second model training data, third model training data and threshold calculation data are obtained, wherein the first model training data, the second model training data and the third model training data can be the same training text, and the threshold calculation data is a disordered text.

[0180] S2, train the perplexity evaluation model.

[0181] The initial perplexity assessment model is trained using the first training data to obtain a trained perplexity assessment model.

[0182] S3, calculate the perplexity threshold.

[0183] The disordered texts in the threshold calculation data are input into the perplexity evaluation model in sequence to obtain the disordered text perplexity of each disordered text, and then the mean of the perplexity of each disordered text is calculated, and the obtained mean is used as the perplexity threshold.

[0184] S4, training grade evaluation model.

[0185] The initial grade evaluation model is trained using the second training data to obtain a trained grade evaluation model.

[0186] S5, train vocabulary prediction model.

[0187] The initial vocabulary prediction model is trained using the third training data to obtain a trained vocabulary prediction model.

[0188] S6, obtain the target text.

[0189] Get the target text written by the target user.

[0190] S7, obtain the target text perplexity.

[0191] The target text is input into the perplexity evaluation model to obtain the target text perplexity of the target text through the perplexity evaluation model.

[0192] S8, judging whether the target text is a qualified text.

[0193] If the text perplexity of the target text is greater than the perplexity threshold, the target text is judged to be an unqualified text, that is, the target text is randomly pieced together and has no value for vocabulary prediction for the target text. Therefore, the vocabulary prediction process ends; if the text perplexity of the target text is less than or equal to the perplexity threshold, the target text is judged to be a qualified text, and then S9 and S11 are executed.

[0194] S9, obtaining the target text level.

[0195] The target text is input into the grade evaluation model to obtain a target text grade of the target text through the grade evaluation model.

[0196] S10, obtaining the target level vocabulary.

[0197] Based on the vocabulary range corresponding to the text level, the upper limit of the vocabulary range corresponding to the target text level is obtained, and the upper limit of the vocabulary range corresponding to the target text level is used as the target level vocabulary.

[0198] S11, obtain text vocabulary.

[0199] The target text is input into the vocabulary prediction model to obtain the text vocabulary of the target text through the vocabulary prediction model. It should be noted that before obtaining the text vocabulary, the target text can be subjected to vocabulary processing.

[0200] S12, determining whether the target level vocabulary is smaller than the text vocabulary.

[0201] If the target level vocabulary size is smaller than the text vocabulary size, then S13 is executed. If the target level vocabulary size is greater than or equal to the text vocabulary size, then S14 is executed.

[0202] S13, outputting the text vocabulary as the predicted vocabulary.

[0203] S14, outputting the target level vocabulary as the predicted vocabulary.

[0204] In an embodiment of the present application, the target text of the target user is evaluated for its text level by adopting a level evaluation model, and the text vocabulary corresponding to the target text is determined in combination with a vocabulary prediction model. Finally, the predicted vocabulary of the target user is determined based on the level vocabulary corresponding to the text level of the target text and the text vocabulary corresponding to the target text. By adopting the user's text, starting from the text perspective that reflects the user's understanding of vocabulary, and combining the level evaluation model and the vocabulary prediction model to predict the user's vocabulary level, the rationality of the vocabulary prediction is guaranteed and the accuracy of the vocabulary prediction is improved.

[0205] The following will be combined with the Figure 7 -Attached Figure 8 The vocabulary prediction device provided in the embodiment of the present application is introduced in detail. Figure 7 -Attached Figure 8 Vocabulary prediction device, used to execute this application Figures 1-6 For the convenience of explanation, only the part related to the embodiment of the present application is shown. For the specific technical details not disclosed, please refer to the present application. Figures 1-6 The embodiment shown.

[0206] See Figure 7 , is a schematic diagram of the structure of a vocabulary prediction device provided in an embodiment of the present application. Figure 7 As shown, the vocabulary prediction device 1 of the embodiment of the present application may include: a text acquisition module 11 , a level acquisition module 12 , a vocabulary acquisition module 13 , and a determination module 14 .

[0207] A text acquisition module 11 is configured to acquire a target text for a target user, and acquire first target words in the target text and a sequence relationship between the first target words;

[0208] A grade acquisition module 12 is configured to acquire a target text grade of the target text based on the first target words and the order relationship between the first target words using a grade evaluation model;

[0209] A vocabulary acquisition module 13 is configured to acquire a text vocabulary corresponding to the target text based on the first target words and the vocabulary levels corresponding to the first target words using a vocabulary prediction model;

[0210] The determination module 14 is configured to determine the predicted vocabulary size of the target user based on the target level vocabulary size corresponding to the target text level and the text vocabulary size.

[0211] In an embodiment of the present application, a text level of a target text is evaluated by using a level evaluation model to obtain the vocabulary range of the target user, and a text vocabulary corresponding to the target text is predicted by using a vocabulary model. Finally, a predicted vocabulary of the target user is obtained based on the vocabulary range of the target user and the text vocabulary corresponding to the target text, thereby avoiding the impact of piecing together high-level vocabulary in the target text and memorizing but not being able to master it on the vocabulary prediction results, and thus obtaining a vocabulary prediction result that can better reflect the actual vocabulary level of the target user, thereby improving the accuracy of vocabulary prediction.

[0212] Optional, please refer to Figure 8 , is a schematic diagram of the structure of a vocabulary prediction device provided in an embodiment of the present application. Figure 8 As shown, the vocabulary prediction device 1 further includes: a first training module 15 .

[0213] The first training module 15 is specifically used to:

[0214] receiving a first number of first training texts and first text levels corresponding to the first training texts, and obtaining first training words in the first training texts and a sequence relationship between the first training words;

[0215] Based on the first training words in the first training text and the order relationship between the first training words, and the first text level corresponding to the first training text, the level evaluation model is trained to obtain the trained level evaluation model.

[0216] In an embodiment of the present application, a large amount of first training data is obtained to perform model training on the initial grade evaluation model (i.e., the grade evaluation model before training) to obtain a trained grade evaluation model, so as to improve the accuracy of the grade evaluation results output by the grade evaluation model based on the target text, thereby improving the prediction accuracy of the vocabulary.

[0217] Optional, please refer to Figure 8 , is a schematic diagram of the structure of a vocabulary prediction device provided in an embodiment of the present application. Figure 8 As shown, the vocabulary prediction device 1 further includes: a second training module 16.

[0218] The second training module 16 is specifically used to:

[0219] receiving a second number of second training texts and second text levels corresponding to the second training texts, and obtaining each second training vocabulary in the second training texts and a vocabulary level corresponding to each second training vocabulary;

[0220] Acquire a training level vocabulary corresponding to the second training text based on a second text level corresponding to the second training text;

[0221] Based on each second training word in the second training text, the vocabulary level corresponding to each second training word, and the training level vocabulary corresponding to the second training text, the vocabulary prediction model is trained to obtain the trained vocabulary prediction model.

[0222] In an embodiment of the present application, a large amount of second training data is obtained to perform model training on the initial vocabulary prediction model (i.e., the vocabulary prediction model before training) to obtain a trained vocabulary prediction model, so as to improve the accuracy of the vocabulary prediction results output by the vocabulary model based on the target text, thereby improving the vocabulary prediction accuracy.

[0223] Optionally, the second training module 16 is specifically configured to:

[0224] Based on each second training vocabulary in the second training text and the vocabulary level corresponding to each second training vocabulary, obtaining the total number of training vocabulary in the second training text and the number of training vocabulary at each vocabulary level;

[0225] Based on the total number of training words in the second training text, the number of training words at each vocabulary level, and the vocabulary of the training level corresponding to the second training text, the vocabulary prediction model is trained to obtain the trained vocabulary prediction model.

[0226] In an embodiment of the present application, the initial vocabulary prediction model (i.e., the vocabulary prediction model before training) is trained using the second training data after vocabulary preprocessing to obtain a trained vocabulary prediction model, so as to improve the accuracy of the vocabulary prediction results output by the vocabulary model based on the target text, thereby improving the accuracy of vocabulary prediction.

[0227] Optionally, the second training module 16 is specifically configured to:

[0228] Performing vocabulary processing on each initial training vocabulary in the second training text to obtain each second training vocabulary corresponding to the second training text, wherein the vocabulary processing includes at least one of vocabulary normalization processing, key vocabulary removal processing, and vocabulary duplication removal processing;

[0229] The vocabulary level corresponding to each second training vocabulary is obtained.

[0230] In an embodiment of the present application, vocabulary preprocessing is performed on the second training text to remove initial training vocabulary in the second training text that may affect the vocabulary prediction results, thereby reducing the impact on the training process of the vocabulary prediction model, thereby improving the accuracy of the vocabulary prediction results output by the vocabulary model based on the target text, and further improving the vocabulary prediction accuracy.

[0231] Optionally, the vocabulary acquisition module 13 is specifically configured to:

[0232] Performing lexical processing on each first target word in the target text to obtain a second target word corresponding to the target text, wherein the lexical processing includes at least one of lexical normalization processing, key word removal processing, and lexical duplicate removal processing;

[0233] Obtaining vocabulary levels corresponding to the second target vocabulary words;

[0234] The vocabulary prediction model is used to obtain the text vocabulary corresponding to the target text based on the second target words and the vocabulary levels corresponding to the second target words.

[0235] In an embodiment of the present application, vocabulary preprocessing is performed on the target text to remove the first target vocabulary in the target text that will affect the vocabulary prediction results, thereby reducing the impact on the vocabulary prediction results of the vocabulary prediction model, so as to improve the accuracy of the vocabulary prediction results output by the vocabulary model based on the target text, thereby improving the vocabulary prediction accuracy.

[0236] Optionally, the determination module 14 is specifically configured to:

[0237] If the target level vocabulary is greater than or equal to the text vocabulary, determining the text vocabulary as the predicted vocabulary;

[0238] If the target-level vocabulary size is smaller than the text vocabulary size, the target-level vocabulary size is determined to be the predicted vocabulary size.

[0239] In an embodiment of the present application, by comparing the target level vocabulary and the text vocabulary, it is determined whether the predicted text vocabulary exceeds the vocabulary range corresponding to the text level, so as to identify whether there are a large number of pieced-together high-level words in the target text, and obtain a predicted vocabulary based on the recognition result, thereby avoiding the influence of the pieced-together high-level words on the vocabulary prediction results, and then obtaining a vocabulary prediction result that can better reflect the actual vocabulary level of the target user, thereby improving the vocabulary prediction accuracy.

[0240] Optional, please refer to Figure 8 , is a schematic diagram of the structure of a vocabulary prediction device provided in an embodiment of the present application. Figure 8 As shown, the vocabulary prediction device 1 further includes: a perplexity acquisition module 17.

[0241] The perplexity acquisition module 17 is specifically used to:

[0242] Using a perplexity assessment model, based on the first target words and the order relationship between the first target words, obtain a target text perplexity of the target text;

[0243] If the target text perplexity is less than or equal to the perplexity threshold, the grade evaluation model is executed to obtain the target text grade of the target text based on the first target words and the order relationship between the first target words.

[0244] In an embodiment of the present application, by comparing the target text perplexity and the perplexity threshold of the target text to determine whether the target text is a randomly assembled text, the target text used to evaluate the vocabulary is screened, avoiding the situation of predicting the vocabulary of randomly assembled texts, thereby improving the accuracy of vocabulary prediction.

[0245] Optional, please refer to Figure 8 , is a schematic diagram of the structure of a vocabulary prediction device provided in an embodiment of the present application. Figure 8 As shown, the vocabulary prediction device 1 further includes: a third training module 18.

[0246] The third training module 18 is specifically used to:

[0247] Obtaining a third number of third training texts, obtaining third training words in the third training texts and a sequence relationship between the third training words;

[0248] Based on the third training words in the third training text and the sequential relationship between the third training words, the perplexity assessment model is trained to obtain the trained perplexity assessment model.

[0249] In an embodiment of the present application, a large amount of third training data is obtained to train the perplexity assessment model to obtain a trained perplexity assessment model, so as to improve the accuracy of the perplexity assessment results output by the perplexity assessment model based on the target text. The target text used to evaluate the vocabulary is screened by the text perplexity output by the perplexity assessment model, thereby avoiding the situation where vocabulary prediction is performed on randomly pieced together texts, thereby improving the accuracy of vocabulary prediction.

[0250] Optional, please refer to Figure 8 , is a schematic diagram of the structure of a vocabulary prediction device provided in an embodiment of the present application. Figure 8 As shown, the vocabulary prediction device 1 further includes a threshold acquisition module 19 .

[0251] The threshold acquisition module 19 is specifically used to:

[0252] Obtaining a fourth number of random texts, obtaining random words in the random texts and a sequence relationship between the random words;

[0253] Using a perplexity evaluation model, based on the perplexity of the scrambled text and the order relationship between the scrambled words in the scrambled text, obtaining the perplexity of the scrambled text;

[0254] A perplexity threshold is obtained based on the scrambled text perplexity of the scrambled text.

[0255] In an embodiment of the present application, a large amount of disordered text is obtained to calculate the perplexity threshold, that is, the perplexity range of randomly pieced-together text is determined, so that the vocabulary prediction device can judge whether the input text is a randomly pieced-together text through the text perplexity of the input text and the perplexity threshold, thereby screening the target text used to evaluate the vocabulary, avoiding the situation of vocabulary prediction for randomly pieced-together text, and thus improving the accuracy of vocabulary prediction.

[0256] The embodiment of the present application also provides a storage medium, which can store multiple program instructions, which are suitable for being loaded and executed by a processor as described above. Figures 1-6 The method steps of the embodiment shown, the specific execution process can be found in Figures 1-6 The detailed description of the illustrated embodiment will not be repeated here.

[0257] See Figure 9 , provides a schematic diagram of the structure of a computer device according to an embodiment of the present application. Figure 9 As shown, the computer device 1000 may include: at least one processor 1001, at least one memory 1002, at least one network interface 1003, at least one input and output interface 1004, at least one communication bus 1005 and at least one display unit 1006. Among them, the processor 1001 may include one or more processing cores. The processor 1001 uses various interfaces and lines to connect the various parts of the entire computer device 1000, and executes various functions and processes data of the terminal 1000 by running or executing instructions, programs, code sets or instruction sets stored in the memory 1002, and calling data stored in the memory 1002. The memory 1002 can be a high-speed RAM memory or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 1002 can optionally be at least one storage device located away from the aforementioned processor 1001. Among them, the network interface 1003 can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface). The communication bus 1005 is used to realize the connection and communication between these components. As Figure 9 As shown, the memory 1002 as a storage medium of a terminal device may include an operating system, a network communication module, an input and output interface module, and a vocabulary prediction program.

[0258] exist Figure 9 In the computer device 1000 shown, the input and output interface 1004 is mainly used to provide an input interface for users and access devices, and to obtain data input by users and access devices.

[0259] In one embodiment.

[0260] The processor 1001 may be configured to call the vocabulary prediction program stored in the memory 1002 and specifically perform the following operations:

[0261] Acquire a target text for a target user, acquire first target words in the target text and a sequence relationship between the first target words;

[0262] Using a grade evaluation model, based on the first target words and the order relationship between the first target words, obtain a target text grade of the target text;

[0263] Using a vocabulary prediction model, based on each of the first target words and the vocabulary level corresponding to each of the first target words, obtain a text vocabulary corresponding to the target text;

[0264] The predicted vocabulary size of the target user is determined based on the target level vocabulary size corresponding to the target text level and the text vocabulary size.

[0265] Optionally, before executing the step of obtaining the target text for the target user, the processor 1001 further performs the following operations:

[0266] receiving a first number of first training texts and first text levels corresponding to the first training texts, and obtaining first training words in the first training texts and a sequence relationship between the first training words;

[0267] Based on the first training words in the first training text and the order relationship between the first training words, and the first text level corresponding to the first training text, the level evaluation model is trained to obtain the trained level evaluation model.

[0268] Optionally, before executing the step of obtaining the target text for the target user, the processor 1001 further performs the following operations:

[0269] receiving a second number of second training texts and second text levels corresponding to the second training texts, and obtaining each second training vocabulary in the second training texts and a vocabulary level corresponding to each second training vocabulary;

[0270] Acquire a training level vocabulary corresponding to the second training text based on a second text level corresponding to the second training text;

[0271] Based on each second training word in the second training text, the vocabulary level corresponding to each second training word, and the training level vocabulary corresponding to the second training text, the vocabulary prediction model is trained to obtain the trained vocabulary prediction model.

[0272] Optionally, when the processor 1001 performs the model training on the vocabulary prediction model based on each second training vocabulary in the second training text and the vocabulary level corresponding to each second training vocabulary, and the training level vocabulary corresponding to the second training text, to obtain the trained vocabulary prediction model, the processor 1001 specifically performs the following operations:

[0273] Based on each second training vocabulary in the second training text and the vocabulary level corresponding to each second training vocabulary, obtaining the total number of training vocabulary in the second training text and the number of training vocabulary at each vocabulary level;

[0274] Based on the total number of training words in the second training text, the number of training words at each vocabulary level, and the vocabulary of the training level corresponding to the second training text, the vocabulary prediction model is trained to obtain the trained vocabulary prediction model.

[0275] Optionally, when executing the step of obtaining each second training vocabulary word and the vocabulary level corresponding to each second training vocabulary word in the second training text, the processor 1001 specifically performs the following operations:

[0276] Performing vocabulary processing on each initial training vocabulary in the second training text to obtain each second training vocabulary corresponding to the second training text, wherein the vocabulary processing includes at least one of vocabulary normalization processing, key vocabulary removal processing, and vocabulary duplication removal processing;

[0277] The vocabulary level corresponding to each second training vocabulary is obtained.

[0278] Optionally, when executing the vocabulary prediction model to obtain the text vocabulary corresponding to the target text based on the first target words and the vocabulary level corresponding to the first target words, the processor 1001 specifically performs the following operations:

[0279] Performing lexical processing on each first target word in the target text to obtain a second target word corresponding to the target text, wherein the lexical processing includes at least one of lexical normalization processing, key word removal processing, and lexical duplicate removal processing;

[0280] Obtaining vocabulary levels corresponding to the second target vocabulary words;

[0281] The vocabulary prediction model is used to obtain the text vocabulary corresponding to the target text based on the second target words and the vocabulary levels corresponding to the second target words.

[0282] Optionally, when the processor 1001 determines the predicted vocabulary of the target user based on the target level vocabulary corresponding to the target text level and the text vocabulary, it specifically performs the following operations:

[0283] If the target level vocabulary is greater than or equal to the text vocabulary, determining the text vocabulary as the predicted vocabulary;

[0284] If the target-level vocabulary size is smaller than the text vocabulary size, the target-level vocabulary size is determined to be the predicted vocabulary size.

[0285] Optionally, after executing the steps of obtaining the target text for the target user and obtaining the first target words in the target text and the order relationship between the first target words, the processor 1001 further performs the following operations:

[0286] Using a perplexity assessment model, based on the first target words and the order relationship between the first target words, obtain a target text perplexity of the target text;

[0287] If the target text perplexity is less than or equal to the perplexity threshold, the grade evaluation model is executed to obtain the target text grade of the target text based on the first target words and the order relationship between the first target words.

[0288] Optionally, before executing the step of obtaining the target text for the target user, the processor 1001 further performs the following operations:

[0289] Obtaining a third number of third training texts, obtaining third training words in the third training texts and a sequence relationship between the third training words;

[0290] Based on the third training words in the third training text and the sequential relationship between the third training words, the perplexity assessment model is trained to obtain the trained perplexity assessment model.

[0291] Optionally, after performing the model training on the perplexity assessment model based on the third training words in the third training text and the sequential relationship between the third training words to obtain the trained perplexity assessment model, the processor 1001 further performs the following operations:

[0292] Obtaining a fourth number of random texts, obtaining random words in the random texts and a sequence relationship between the random words;

[0293] Using a perplexity evaluation model, based on the perplexity of the scrambled text and the order relationship between the scrambled words in the scrambled text, obtaining the perplexity of the scrambled text;

[0294] A perplexity threshold is obtained based on the scrambled text perplexity of the scrambled text.

[0295] In an embodiment of the present application, the target text of the target user is evaluated for its text level by adopting a level evaluation model, and the text vocabulary corresponding to the target text is determined in combination with a vocabulary prediction model. Finally, the predicted vocabulary of the target user is determined based on the level vocabulary corresponding to the text level of the target text and the text vocabulary corresponding to the target text. By adopting the user's text, starting from the text perspective that reflects the user's understanding of vocabulary, and combining the level evaluation model and the vocabulary prediction model to predict the user's vocabulary level, the rationality of the vocabulary prediction is guaranteed and the accuracy of the vocabulary prediction is improved.

[0296] It should be noted that for the aforementioned method embodiments, for ease of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0297] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0298] The above is a description of a vocabulary prediction method, vocabulary prediction device, storage medium and equipment provided by this application. For those skilled in the art, based on the ideas of the embodiments of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A vocabulary prediction method, characterized in that: The method comprises: Acquire a target text for a target user, acquire first target words in the target text and a sequence relationship between the first target words; Using a grade evaluation model, based on the first target words and the order relationship between the first target words, obtain a target text grade of the target text; Using a vocabulary prediction model, based on each of the first target words and the vocabulary level corresponding to each of the first target words, obtain a text vocabulary corresponding to the target text; If the target level vocabulary corresponding to the target text level is greater than or equal to the text vocabulary, determining the text vocabulary as the predicted vocabulary of the target user; If the target-level vocabulary size is smaller than the text vocabulary size, the target-level vocabulary size is determined to be the predicted vocabulary size.

2. The method according to claim 1, characterized in that Before obtaining the target text for the target user, the method further includes: receiving a first number of first training texts and first text levels corresponding to the first training texts, and obtaining first training words in the first training texts and a sequence relationship between the first training words; Based on the first training words in the first training text and the order relationship between the first training words, and the first text level corresponding to the first training text, the level evaluation model is trained to obtain the trained level evaluation model.

3. The method according to claim 1, characterized in that Before obtaining the target text for the target user, the method further includes: receiving a second number of second training texts and second text levels corresponding to the second training texts, and obtaining each second training vocabulary in the second training texts and a vocabulary level corresponding to each second training vocabulary; Acquire a training level vocabulary corresponding to the second training text based on a second text level corresponding to the second training text; Based on each second training word in the second training text, the vocabulary level corresponding to each second training word, and the training level vocabulary corresponding to the second training text, the vocabulary prediction model is trained to obtain the trained vocabulary prediction model.

4. The method according to claim 3, characterized in that The method of performing model training on a vocabulary prediction model based on each second training vocabulary in the second training text, the vocabulary level corresponding to each second training vocabulary, and the training level vocabulary corresponding to the second training text to obtain the trained vocabulary prediction model includes: Based on each second training vocabulary in the second training text and the vocabulary level corresponding to each second training vocabulary, obtaining the total number of training vocabulary in the second training text and the number of training vocabulary at each vocabulary level; Based on the total number of training words in the second training text, the number of training words at each vocabulary level, and the vocabulary of the training level corresponding to the second training text, the vocabulary prediction model is trained to obtain the trained vocabulary prediction model.

5. The method according to claim 3, characterized in that The obtaining of each second training word in the second training text and a vocabulary level corresponding to each second training word includes: Performing vocabulary processing on each initial training vocabulary in the second training text to obtain each second training vocabulary corresponding to the second training text, wherein the vocabulary processing includes at least one of vocabulary normalization processing, key vocabulary removal processing, and vocabulary duplication removal processing; The vocabulary level corresponding to each second training vocabulary is obtained.

6. The method according to claim 1, characterized in that The adopting of the vocabulary prediction model to obtain the text vocabulary corresponding to the target text based on the first target words and the vocabulary level corresponding to the first target words includes: Performing lexical processing on each first target word in the target text to obtain a second target word corresponding to the target text, wherein the lexical processing includes at least one of lexical normalization processing, key word removal processing, and lexical duplicate removal processing; Obtaining the vocabulary level corresponding to each second target vocabulary; The vocabulary prediction model is adopted to obtain the text vocabulary corresponding to the target text based on the second target words and the vocabulary levels corresponding to the second target words.

7. The method according to claim 1, characterized in that After obtaining the target text for the target user and obtaining the first target words in the target text and the order relationship between the first target words, the method further includes: Using a perplexity assessment model, based on the first target words and the order relationship between the first target words, obtain a target text perplexity of the target text; If the target text perplexity is less than or equal to the perplexity threshold, the grade evaluation model is executed to obtain the target text grade of the target text based on the first target words and the order relationship between the first target words.

8. The method according to claim 7, characterized in that Before obtaining the target text for the target user, the method further includes: Obtaining a third number of third training texts, obtaining third training words in the third training texts and a sequence relationship between the third training words; Based on the third training words in the third training text and the sequential relationship between the third training words, the perplexity assessment model is trained to obtain the trained perplexity assessment model.

9. The method according to claim 8, characterized in that After performing model training on the perplexity assessment model based on the third training words in the third training text and the order relationship between the third training words to obtain the trained perplexity assessment model, the method further includes: Obtaining a fourth number of random texts, obtaining random words in the random texts and a sequence relationship between the random words; Using a perplexity evaluation model, based on the perplexity of the scrambled text and the order relationship between the scrambled words in the scrambled text, obtaining the perplexity of the scrambled text; A perplexity threshold is obtained based on the scrambled text perplexity of the scrambled text.

10. A vocabulary prediction device, characterized in that: include: A text acquisition module, configured to acquire a target text for a target user, and acquire first target words in the target text and a sequence relationship between the first target words; a grade acquisition module, configured to acquire a target text grade of the target text based on the first target words and the order relationship between the first target words by using a grade evaluation model; A vocabulary acquisition module, configured to acquire a text vocabulary corresponding to the target text based on the first target words and the vocabulary levels corresponding to the first target words using a vocabulary prediction model; a determination module, configured to determine the text vocabulary as the predicted vocabulary of the target user if the target level vocabulary corresponding to the target text level is greater than or equal to the text vocabulary; If the target-level vocabulary size is smaller than the text vocabulary size, the target-level vocabulary size is determined to be the predicted vocabulary size.

11. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the vocabulary prediction method according to any one of claims 1 to 9 is implemented.

12. A computer device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the vocabulary prediction method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Content recommendation method and device and computing equipment

    CN111241397A

  • Vocabulary level test processing method and system based on reading understanding practice

    CN113065334A