Text difficulty grading assessment method, device, equipment and storage medium

By obtaining the character and word characteristics of the text to be graded and calculating the weighted scores of Chinese characters and word segmentation, the accuracy and efficiency of text difficulty grading are solved, and more accurate text matching and reading ability evaluation are achieved.

CN114416977BActive Publication Date: 2025-08-29SHANGHAI LIULISHUO INFORMATION TECH CO LTD
View PDF -1 Cites 0 Cited by

Patent Information

Application Number
CN202111645082.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-08-29
Estimated Expiration
2041-12-29

Smart Images

  • Figure CN114416977B_ABST
    Figure CN114416977B_ABST
Patent Text Reader

Abstract

A text difficulty grading and assessment method, apparatus, device, and storage medium, the method comprising: obtaining a text to be graded; preprocessing the text to be graded to obtain a feature set, including multiple features related to text granularity, wherein the feature types in the feature set include at least characters and words; obtaining difficulty assessment values ​​corresponding to each feature, including character difficulty values ​​and word difficulty values; obtaining the sum of the character difficulty values ​​of all characters in the feature set of the text to be graded as a Chinese character weighted score, and obtaining the sum of the word difficulty values ​​of all words in the feature set of the text to be graded as a word segmentation weighted score; using the Chinese character weighted score and the word segmentation weighted score to obtain a predicted difficulty level of the text to be graded; judging the predicted difficulty level and the difficulty assessment value, and adjusting the predicted difficulty level when the predicted difficulty level difficulty assessment value meets a threshold condition. The present invention is advantageous in improving the accuracy of text difficulty grading while improving assessment efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of natural language processing technology, and in particular to a text difficulty grading assessment method and apparatus, device, and storage medium. Background Art

[0002] Graded reading is based on the laws of children's physical and mental development. It selects reading materials of corresponding difficulty levels for users of different age groups or reading levels, thereby gradually improving users' reading ability.

[0003] In the modern information society, the quantity and types of reading materials (such as textbooks, magazines, newspapers, etc.) are growing exponentially. How to select texts that match the age group or reading level from such a large amount of reading materials has become a difficult problem.

[0004] Therefore, in order to select texts that match the age group or reading level from a large amount of reading materials, it is necessary to grade the reading difficulty of the texts so that users can obtain texts that match their reading ability. Summary of the Invention

[0005] The problem solved by the embodiments of the present invention is to provide a text difficulty grading evaluation method and apparatus, device and storage medium, which improve the accuracy of text difficulty grading while improving the evaluation efficiency.

[0006] To solve the above problems, an embodiment of the present invention provides a method for assessing text difficulty grading, comprising: obtaining a text to be graded; preprocessing the text to be graded to obtain a feature set of the text to be graded, the feature set including multiple features related to text granularity, and the feature types in the feature set including at least characters and words; obtaining a difficulty assessment value corresponding to each feature in the feature set of the text to be graded, the difficulty assessment value including a character difficulty value corresponding to each character and a word difficulty value corresponding to each word, the character difficulty value being negatively correlated with character frequency, and the word difficulty value being negatively correlated with word frequency; obtaining the sum of the character difficulty values ​​of all characters in the feature set of the text to be graded as a Chinese character weighted score, and obtaining the sum of the word difficulty values ​​of all words in the feature set of the text to be graded as a word segmentation weighted score; obtaining a predicted difficulty level of the text to be graded using the Chinese character weighted score and the word segmentation weighted score; judging the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded, and adjusting the predicted difficulty level when the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet a threshold condition.

[0007] Accordingly, an embodiment of the present invention further provides a text difficulty grading and evaluation device, comprising: a text acquisition module for acquiring a text to be graded; a preprocessing module for preprocessing the text to be graded and acquiring a feature set of the text to be graded, wherein the feature set includes a plurality of features related to text granularity, and the feature types in the feature set include at least characters and words; a difficulty evaluation value acquisition module for acquiring a difficulty evaluation value corresponding to each feature in the feature set of the text to be graded, wherein the difficulty evaluation value includes a character difficulty value corresponding to each character and a word difficulty value corresponding to each word, wherein the character difficulty value is negatively correlated with the character frequency, and the word difficulty value is negatively correlated with the word frequency; and a ranking score acquisition module. , used to obtain the sum of the character difficulty values ​​of all characters in the feature set of the text to be graded as the weighted score of Chinese characters, and obtain the sum of the word difficulty values ​​of all words in the feature set of the text to be graded as the weighted score of word segmentation; a difficulty level prediction module, used to use the weighted score of Chinese characters and the weighted score of word segmentation to obtain the predicted difficulty level of the text to be graded; a difficulty level adjustment module, used to judge the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded, and adjust the predicted difficulty level when the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded meet the threshold condition.

[0008] Accordingly, an embodiment of the present invention further provides a device comprising at least one memory and at least one processor, wherein the memory stores one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the text difficulty grading assessment method described in an embodiment of the present invention.

[0009] Correspondingly, an embodiment of the present invention further provides a storage medium, wherein the storage medium stores one or more computer instructions, and the one or more computer instructions are used to implement the text difficulty grading assessment method described in the embodiment of the present invention.

[0010] Compared with the prior art, the technical solution of the embodiment of the present invention has the following advantages:

[0011] In the text difficulty grading evaluation method provided by the embodiment of the present invention, after obtaining the character difficulty value negatively correlated with the character frequency and the word difficulty value negatively correlated with the word frequency, the sum of the character difficulty values ​​of all characters in the feature set of the text to be graded is obtained as the Chinese character weighted score, the sum of the word difficulty values ​​of all words in the feature set of the text to be graded is obtained as the word segmentation weighted score, and then the Chinese character weighted score and the word segmentation weighted score are used to obtain the predicted difficulty level of the text to be graded, and when the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded meet the threshold condition, the predicted difficulty level is adjusted; wherein, the character difficulty The value is negatively correlated with the word frequency. The lower the word frequency, the greater the difficulty of the word, and the higher the word difficulty value. The word difficulty value will affect the difficulty of the text to be graded. The Chinese character weighted score is obtained by summing the word difficulty values ​​of all the words. The sum of the word difficulty values ​​can characterize the overall difficulty of the words in the text to be tested. Therefore, compared with the difficulty evaluation values ​​corresponding to other types of features, the Chinese character weighted score can better characterize the difficulty of the words in the entire text, thereby characterizing the difficulty of the text to be graded. Similarly, the word difficulty value is negatively correlated with the word frequency. The lower the word frequency, the greater the difficulty of the word, and the higher the word difficulty value. The word difficulty value will also affect the difficulty of the text to be graded. The difficulty of the text, the word segmentation weighted score is obtained by summing the word difficulty values ​​of all words, and the sum of the word difficulty values ​​can represent the overall difficulty of the words in the text to be tested. Compared with the difficulty assessment values ​​corresponding to other types of features, the word segmentation weighted score can better represent the difficulty of the words in the entire text, thereby representing the difficulty of the text to be graded. Accordingly, the predicted difficulty level of the text to be graded is obtained by using the Chinese character weighted score and the word segmentation weighted score, which can obtain the overall difficulty of characters and words. Characters and words are the two types of features with the smallest text granularity, so it is beneficial to improve the accuracy of text difficulty grading and facilitate users to select suitable reading texts. ; Moreover, the embodiment of the present invention only needs to obtain the predicted difficulty level through the weighted score of Chinese characters and the weighted score of word segmentation, which is beneficial to reducing the amount of data calculation in the process of predicting the difficulty level, thereby improving the evaluation efficiency of the text difficulty grading evaluation method; in addition, after obtaining the predicted difficulty level, the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded are also judged. When the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded meet the threshold condition, the predicted difficulty level is adjusted, which can further improve the accuracy of the difficulty level finally obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a flow chart of an embodiment of a method for assessing text difficulty grading according to the present invention;

[0013] Figure 2 This is a functional block diagram of an embodiment of a device for evaluating text difficulty levels according to the present invention;

[0014] Figure 3 This is a hardware structure diagram of a device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0015] As can be seen from the background technology, accurately assessing the difficulty of text is one of the difficult issues in current graded reading research.

[0016] In order to solve the above technical problems, an embodiment of the present invention provides a method for evaluating text difficulty levels. Figure 1 , shows a flowchart of an embodiment of a text difficulty grading evaluation method of the present invention.

[0017] In an embodiment of the present invention, the text difficulty grading evaluation method includes the following basic steps:

[0018] Step S1: Obtain the text to be graded;

[0019] Step S2: pre-processing the text to be classified to obtain a feature set of the text to be classified, wherein the feature set includes a plurality of features related to text granularity, and the feature types in the feature set include at least characters and words;

[0020] Step S3: Obtaining a difficulty evaluation value corresponding to each feature in the feature set of the text to be graded, wherein the difficulty evaluation value includes a character difficulty value corresponding to each character and a word difficulty value corresponding to each word, wherein the character difficulty value is negatively correlated with the character frequency, and the word difficulty value is negatively correlated with the word frequency;

[0021] Step S4: obtaining the sum of the character difficulty values ​​of all characters in the feature set of the text to be graded as a weighted score of the Chinese characters, and obtaining the sum of the word difficulty values ​​of all words in the feature set of the text to be graded as a weighted score of the word segmentation;

[0022] Step S5: using the weighted scores of Chinese characters and the weighted scores of word segments to obtain the predicted difficulty level of the text to be graded;

[0023] Step S6: judging the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded, and adjusting the predicted difficulty level when the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded meet a threshold condition.

[0024] The embodiment of the present invention selects difficulty assessment values ​​corresponding to features that can better characterize the overall difficulty of the text to be tested (i.e., the weighted scores of Chinese characters and the weighted scores of word segments) to obtain the predicted difficulty level of the text to be graded, which is beneficial to improving the accuracy of text difficulty grading. Moreover, the predicted difficulty level only needs to be obtained through the weighted scores of Chinese characters and the weighted scores of word segments, which is beneficial to reducing the amount of data calculation in the process of predicting the difficulty level, thereby improving the evaluation efficiency of the text difficulty grading evaluation method; in addition, after obtaining the predicted difficulty level, adjusting the predicted difficulty level can further improve the accuracy of the difficulty level finally obtained.

[0025] In order to make the above-mentioned objects, features and advantages of the embodiments of the present invention more obvious and easy to understand, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0026] refer to Figure 1 , execute step S1 to obtain the text to be graded.

[0027] The text to be graded is used as an object for difficulty grading, and the difficulty level of the text to be graded is subsequently determined by grading the text to be graded so that the user can obtain a text that matches their reading ability. Specifically, the text to be graded can be obtained by manual input or automatically obtained in response to a received processing instruction.

[0028] In this embodiment, the text to be classified is a Chinese text. Specifically, the language of the text to be classified is Simplified Chinese. It should be noted that the text to be classified can be reading material from any source. Specifically, it can be any text from the Internet (for example, it can be text obtained from various document libraries or book websites), or it can be any book or publication. It is understandable that in the process of obtaining the text to be classified, some content that is irrelevant to the text content (for example, watermarks, signatures, etc.) can be directly removed.

[0029] Continue to refer Figure 1 , execute step S2, pre-process the text to be classified, and obtain a feature set of the text to be classified, wherein the feature set includes multiple features related to text granularity, and the feature types in the feature set include at least characters and words.

[0030] By performing preprocessing, features related to text granularity are obtained, and the features are correspondingly related to the difficulty level of the text, so that the difficulty evaluation value corresponding to each feature can be used as an evaluation factor for calculating the predicted difficulty level.

[0031] In this embodiment, a feature set for measuring the reading difficulty level is constructed from the language level of the text to be tested. Therefore, the feature set includes multiple features related to text granularity, so that the reading difficulty of the text to be tested can be characterized by using features of different dimensions.

[0032] In this embodiment, the text to be graded is Chinese text, so the feature types in the feature set include at least characters and words. Character-level features and word-level features are the two types of features with the smallest text granularity, and therefore help more accurately characterize the reading difficulty of the text to be graded, thereby improving the accuracy of subsequent text difficulty grading.

[0033] In this embodiment, preprocessing the text to be classified includes: splitting the text to be classified according to different text granularities, and obtaining a feature set corresponding to the text to be classified.

[0034] Specifically, segmenting the text to be classified according to different text granularities includes performing word segmentation on the text to be classified to obtain word-level features of the text to be classified. For example, if one of the sentences in the text to be classified is {I am Chinese}, it can be segmented to obtain {I am Chinese}. Similarly, the feature set corresponding to the text to be classified can be obtained using the above-mentioned word segmentation method.

[0035] As an example, a word segmenter can be used to perform word segmentation on the text to be tested. For example, the word segmenter includes jieba. It should be noted that in other embodiments, other word segmenters with better word segmentation effects can also be used for word segmentation.

[0036] It can be understood that the word-level features can be directly obtained from the text to be graded.

[0037] In this embodiment, in the step of preprocessing the text to be graded, the feature type in the feature set also includes sentences. Subsequently, by obtaining the difficulty assessment value corresponding to the sentence-level features, the predicted difficulty level is adjusted according to the difficulty assessment value corresponding to the sentence-level features.

[0038] Correspondingly, preprocessing the text to be classified further includes: performing sentence segmentation processing on the text to be classified.

[0039] As an example, segmenting the text to be classified includes: traversing the text to be classified, determining the positions of segmentation symbols in the text to be classified, and breaking the text to be classified into a plurality of separate sentences according to the positions of the segmentation symbols. For example, the segmentation symbols may be punctuation marks.

[0040] In this embodiment, pre-processing the text to be classified further includes: before splitting the text to be classified according to different text granularities, performing a first filtering process on the text to be classified to filter out content that does not meet text requirements.

[0041] The text to be graded can be a text from any source, and the type of the text to be tested is prone to diversity. Therefore, by performing a first filtering process, a text for removing the text whose difficulty level cannot be predicted by the features of the word level and the word level is removed. For example, when the type of the text to be tested is an ancient Chinese text (for example, ancient poetry or classical Chinese, etc.), there are large differences between ancient Chinese and simplified Chinese in syntax and word meanings, and non-Chinese texts do not have the features of the word level. Therefore, ancient Chinese texts and non-Chinese texts should be eliminated. As an example, since ancient poetry is usually regular, such as the number of words, the number of words can be used as a screening condition to eliminate ancient poetry texts. As another example, non-Chinese texts can be filtered out by character encoding.

[0042] It is understandable that since punctuation marks (e.g., period, semicolon) and special characters (e.g., #, &) and other features do not affect the reading difficulty of the text, punctuation marks, special characters and other features are not taken into account when subsequently obtaining the difficulty assessment value corresponding to each feature.

[0043] Execute step S3 to obtain a difficulty evaluation value corresponding to each feature in the feature set of the text to be graded, wherein the difficulty evaluation value includes a character difficulty value corresponding to each character and a word difficulty value corresponding to each word. The character difficulty value is negatively correlated with the character frequency, and the word difficulty value is negatively correlated with the word frequency.

[0044] Subsequently, the character difficulty value is used to obtain a weighted score for the Chinese characters, and the word difficulty value is used to obtain a weighted score for the word segmentation. The weighted scores for the Chinese characters and word segmentation are used as evaluation factors for assessing the difficulty level of the text to be tested. The character difficulty value is negatively correlated with the character frequency. A lower character frequency indicates a greater difficulty for the character, and a correspondingly higher character difficulty value. Similarly, the word difficulty value is negatively correlated with the word frequency. A lower word frequency indicates a greater difficulty for the word, and a correspondingly higher word difficulty value. Therefore, by obtaining the character difficulty value and the word difficulty value, it is easier to more accurately characterize the difficulty level of the entire text.

[0045] In this embodiment, before obtaining the difficulty evaluation value corresponding to each feature, the method further includes: obtaining a dictionary library, including a dictionary library and a lexicon library, wherein the dictionary library is marked with the difficulty value of each character, and the lexicon library is marked with the difficulty value of each word.

[0046] Subsequently, the word difficulty value of each character and the word difficulty value of each word in the text to be tested are obtained through the dictionary. On the one hand, by obtaining the dictionary, the corresponding word difficulty value and word difficulty value are directly obtained from the dictionary, which reduces the complexity of obtaining the word difficulty value and word difficulty value. On the other hand, compared with large-scale corpus, the number of characters and words in the text to be graded is relatively small. If the difficulty evaluation value corresponding to each feature is set directly according to the frequency of occurrence of each feature in the feature type of the text to be graded, it is easy to cause the difficulty value of some features to be much greater than or much less than its true difficulty value due to insufficient sample quantity, thereby reducing the accuracy of text difficulty grading. Therefore, obtaining the difficulty value of each character as the word difficulty value through the dictionary library and obtaining the difficulty value of each word as the word difficulty value through the lexicon library is conducive to improving the accuracy of the word difficulty value and word difficulty value in the text to be graded.

[0047] Therefore, in a specific embodiment, a reference text set consisting of a large number of reference texts can be first obtained, and the features related to the difficulty value of each text in the reference text set can be extracted to establish a dictionary. Based on the dictionary, the difficulty assessment value corresponding to each feature in the feature set of the text to be graded can be determined.

[0048] Specifically, obtaining a dictionary library includes: obtaining a reference text set and extracting features of each text in the reference text set; counting the frequency of occurrence of each feature according to the feature type; sorting according to the frequency of occurrence of each feature to obtain the dictionary library, wherein the dictionary library includes a dictionary library and a lexicon library; determining the difficulty value corresponding to each character in the dictionary library and the difficulty value corresponding to each word in the lexicon library based on the sorting results of the feature types corresponding to the dictionary library and the lexicon library, respectively, and normalizing the difficulty value corresponding to each character and the difficulty value corresponding to each word to obtain the character difficulty value of each character and the word difficulty value of each word.

[0049] In a specific embodiment, the reference text set can be derived from any unannotated text on the internet. For example, a large number of texts can be obtained from various document libraries and book websites to serve as the reference text set. Furthermore, the reference text set can be continuously updated with new texts or outdated texts removed over time to ensure the timeliness of each text in the reference text set.

[0050] It is understandable that, in the process of obtaining the reference text set, some contents irrelevant to the text content, such as watermarks, signatures, etc., can be directly removed.

[0051] As an example, the splitting process described in the foregoing embodiments can be used to split each text in the reference text set, obtain multiple features of the reference text set, construct a language dictionary based on the obtained multiple features, calculate the difficulty values corresponding to each feature in the language dictionary, and build a solid corpus resource foundation for obtaining the difficulty evaluation values corresponding to each feature in the feature set of the text to be graded later.

[0052] It should be noted that the difficulty values corresponding to each feature in the language dictionary can be obtained based on multiple factors. In this embodiment, the frequency of occurrence of a feature can be used as a factor to measure the difficulty of the feature. In specific applications, the difficulty values of each feature can be determined according to the frequency of occurrence of the feature, and the greater the frequency of occurrence of the feature, the smaller the corresponding difficulty value. Therefore, in this embodiment, during the process of obtaining the language dictionary, the frequency of occurrence of each feature is counted according to the feature type, and the language dictionary is obtained by sorting according to the frequency of occurrence of each feature, where the language dictionary includes a character dictionary and a word dictionary.

[0053] In this embodiment, the difficulty values corresponding to each character in the character dictionary and the difficulty values corresponding to each word in the word dictionary are determined respectively according to the sorting results of the corresponding feature types of the character dictionary and the word dictionary, and the difficulty values corresponding to each character and the difficulty values corresponding to each word are respectively normalized.

[0054] For the convenience of understanding, the following takes the determination of the difficulty value of a character as an example for illustration. For example, if there are 10,000 characters in the character dictionary, among which, the character "了" has the highest frequency of 500, the character "牖" has the lowest frequency of 1, and the frequencies of the remaining 9,998 characters are between 2 and 499. The greater the frequency of occurrence of a character, the lower the corresponding difficulty value. According to the frequency of occurrence of the characters, the difficulty values of the 10,000 characters are set from the difficulty value range [0, 1].

[0055] For example, when performing the normalization process, the difficulty value of the character ranked first in the frequency sorting (i.e., "了") can be set to 0, the difficulty value of the character ranked 10,000th in the frequency sorting (i.e., "牖") can be set to 1, and the difficulty values of the remaining characters are set in sequence according to their frequencies of occurrence.

[0056] For other features in the language dictionary, such as the difficulty values of words in the word dictionary, the determination process of the difficulty values of characters can be referred to, which will not be elaborated here.

[0057] It should be noted that for the difficulty value of a sentence, it can be set according to the length and number of sentences. Among them, the longer the sentence and the more the number, correspondingly, it can be considered that the difficulty value of this sentence is greater.

[0058] By establishing a dictionary library and a lexicon library, and setting difficulty values ​​for characters and words according to their frequencies, difficulty values ​​corresponding to different feature types under multiple dimensions of each text in the reference text set can be obtained. Moreover, by calculating the difficulty values ​​corresponding to different feature types, the obtained character difficulty values ​​and word difficulty values ​​are made more accurate, thereby improving the accuracy of the character difficulty values ​​and word difficulty values ​​of the text to be graded.

[0059] It should be noted that before normalizing the difficulty values ​​corresponding to each character and each word, the acquisition of the dictionary library also includes: performing a second filtering process to filter out the characters with the highest frequency before the first preset value, and removing the remaining characters; performing a third filtering process to filter out the words with the highest frequency before the second preset value, and removing the remaining words.

[0060] When the frequency of appearance of a certain character is too low, the character is usually an uncommon character, and uncommon characters usually cannot be used to represent the actual difficulty level of the text to be tested. If the uncommon character is retained, it is easy to cause errors in the character difficulty values ​​of each character after normalization. The character difficulty value is used to obtain the weighted score of the Chinese character, and the weighted score of the Chinese character is used as an evaluation factor for predicting the difficulty level, which leads to prediction errors in the process of predicting the difficulty level. Therefore, by performing a second filtering process to remove characters with a lower frequency of appearance, the noise that may cause prediction errors can be removed, thereby improving the accuracy of predicting the difficulty level.

[0061] The first preset value should not be too small or too large. If the first preset value is too large, it is likely to retain too many uncommon characters, resulting in prediction errors during the difficulty level prediction process. If the first preset value is too small, it is likely to cause misoperation, eliminating characters that can be used to assess the actual difficulty of the test text, which is also likely to result in prediction errors during the difficulty level prediction process. To this end, in this embodiment, the first preset value is between 3500 and 4500. For example, the first preset value is 4000.

[0062] Similarly, by performing the third filtering process, uncommon words can be removed, and the noise that may cause prediction errors can be removed accordingly, thereby improving the accuracy of the prediction difficulty level.

[0063] The second preset value should not be too small or too large. If the second preset value is too large, it is likely to retain too many uncommon words, resulting in prediction errors during the difficulty level prediction process. If the second preset value is too small, it is likely to cause erroneous operation, eliminating words that can be used to assess the actual difficulty of the test text, which is also likely to result in prediction errors during the difficulty level prediction process. Therefore, in this embodiment, the second preset value is between 45,000 and 55,000.

[0064] It should be noted that in the text to be measured, the same words are grouped into the same category, and the same characters are grouped into the same category. Usually, the number of word types is greater than the number of character types. Therefore, the first preset value is less than the second preset value.

[0065] Correspondingly, in this embodiment, obtaining the difficulty evaluation values corresponding to each feature in the feature set of the text to be classified includes: obtaining the difficulty values of each character in the feature set of the text to be classified as character difficulty values according to the dictionary library; obtaining the difficulty values of each word in the feature set of the text to be classified as word difficulty values according to the dictionary.

[0066] Specifically, obtaining the difficulty values of each character in the feature set of the text to be classified as character difficulty values according to the dictionary library includes: retrieving the characters in the feature set of the text to be classified in the dictionary library; when the same character is retrieved, taking the difficulty value of the corresponding character in the dictionary library as the difficulty value of the corresponding character in the feature set of the text to be classified, and when the same character is not retrieved, setting the difficulty value of the character not retrieved in the feature set of the text to be classified as the maximum value of the difficulty values in the dictionary library.

[0067] As an example, the same character '了' appears in the dictionary library and the feature set of the text to be classified. As described above, the character difficulty value of the character '了' in the text to be classified can be set to 0.

[0068] It should be noted that for some characters in the text to be classified, the dictionary library may not include these characters, which means that the number of occurrences of these characters is small and their difficulty values are large. The difficulty values of these characters can be set to the maximum value of the difficulty values. Specifically, in this embodiment, the maximum value of the difficulty values in the dictionary library is 1.

[0069] Similarly, obtaining the difficulty values of each word in the feature set of the text to be classified as word difficulty values according to the dictionary includes: retrieving the words in the feature set of the text to be classified in the dictionary; when the same word is retrieved, taking the difficulty value of the corresponding word in the dictionary as the difficulty value of the corresponding word in the feature set of the text to be classified, and when the same word is not retrieved, setting the difficulty value of the word not retrieved in the feature set of the text to be classified as the maximum value of the difficulty values in the dictionary.

[0070] It should be noted that for some words in the text to be classified, the dictionary may not include these words, which means that the number of occurrences of these words is small and their difficulty values are large. The difficulty values of these words can be set to the maximum value of the difficulty values. Specifically, in this embodiment, the maximum value of the difficulty values in the dictionary is 1.

[0071] In this embodiment, obtaining the difficulty evaluation value corresponding to each feature further includes: obtaining the number of sentences and the average length of sentences in the text to be graded.

[0072] The number of sentences and the average length of sentences can also be used to characterize the difficulty level of the text to be tested. The longer the sentence, the greater the difficulty value of the sentence, and the greater the difficulty of the text to be tested. The more sentences there are, the greater the difficulty of the text to be tested. By obtaining the number of sentences and the average length of sentences of the text to be graded, the predicted difficulty level can be adjusted based on the difficulty evaluation value of the sentence-level features, thereby further improving the accuracy of difficulty grading of the text to be graded.

[0073] Execute step S4 to obtain the sum of the character difficulty values ​​of all characters in the feature set of the text to be graded as the Chinese character weighted score, and obtain the sum of the word difficulty values ​​of all words in the feature set of the text to be graded as the word segmentation weighted score.

[0074] The character difficulty value is negatively correlated with the character frequency. The lower the character frequency, the greater the difficulty of the character, and the higher the character difficulty value. The weighted score of Chinese characters is obtained by summing the character difficulty values ​​of all characters. The sum of the character difficulty values ​​can represent the overall difficulty of the characters in the text to be tested. Therefore, compared with the difficulty assessment values ​​corresponding to other types of features, the weighted score of Chinese characters can better represent the difficulty of the characters in the entire text, that is, the weighted score of Chinese characters can be used to represent the overall difficulty value of the characters. In other words, the weighted score of Chinese characters is more strongly correlated with the actual difficulty of the text to be tested, thereby being able to represent the difficulty of the text to be graded.

[0075] Similarly, the word difficulty value is negatively correlated with the word frequency. The lower the word frequency, the greater the difficulty of the word, and the higher the word difficulty value. The weighted score of the word segmentation is obtained by summing the word difficulty values ​​of all words. The sum of the word difficulty values ​​can represent the overall difficulty of the words in the text to be tested. Therefore, compared with the difficulty assessment values ​​corresponding to other types of features, the weighted score of the word segmentation can better represent the difficulty of the words in the entire text, that is, the weighted score of the word segmentation can be used to represent the overall difficulty value of the word. In other words, the weighted score of the word segmentation is more strongly correlated with the actual difficulty of the text to be tested, and thus can also represent the difficulty of the text to be graded.

[0076] Correspondingly, by using the Chinese character weighted score and word segmentation weighted score to obtain the predicted difficulty level of the text to be graded, the overall difficulty of characters and words can be obtained, and characters and words are the two types of features with the smallest text granularity, so it is beneficial to improve the accuracy of text difficulty grading and facilitate users to select suitable reading texts.

[0077] Moreover, this embodiment only needs to obtain the predicted difficulty level through the weighted scores of Chinese characters and word segmentation, which is beneficial to reducing the amount of data calculation in the process of predicting the difficulty level, thereby improving the evaluation efficiency of the text difficulty grading evaluation method.

[0078] At the same time, the predicted difficulty level is obtained by the weighted score of Chinese characters and the weighted score of word segmentation, which means that the weighted score of Chinese characters and the weighted score of word segmentation are used to construct a model for predicting the difficulty level, which correspondingly simplifies the complexity of the model for predicting the difficulty level, and is also conducive to improving the evaluation efficiency of the text difficulty grading evaluation method.

[0079] Execute step S5 to obtain the predicted difficulty level of the text to be graded using the weighted scores of Chinese characters and the weighted scores of word segments.

[0080] It can be seen from the above records that compared with the difficulty assessment values ​​corresponding to other types of features, the Chinese character weighted score is more capable of representing the difficulty of the characters in the entire text, and the word segmentation weighted score is more capable of representing the difficulty of the words in the entire text. Therefore, by using the Chinese character weighted score and the word segmentation weighted score, the predicted difficulty level of the text to be graded is obtained, thereby improving the accuracy of text difficulty grading.

[0081] Specifically, using the Chinese character weighted score and the word segmentation weighted score to obtain the predicted difficulty level of the text to be graded includes: normalizing the Chinese character weighted score and the word segmentation weighted score according to the preset difficulty level, obtaining a first initial predicted difficulty level corresponding to the Chinese character weighted score, and a second initial predicted difficulty level corresponding to the word segmentation weighted score; weighting the first initial predicted difficulty level and the second initial predicted difficulty level, and rounding the weighted result to obtain the predicted difficulty level of the text to be graded.

[0082] By weighting the first initial predicted difficulty level and the second initial predicted difficulty level, the influence of the difficulty values ​​of characters and words on the difficulty level can be comprehensively considered.

[0083] Among them, the Chinese character weighted scores and word segmentation weighted scores are first normalized according to the preset difficulty levels, so that the mapping relationship between the Chinese character weighted scores and the predicted difficulty levels, as well as the mapping relationship between the word segmentation weighted scores and the predicted difficulty levels can be obtained more intuitively, that is, the correlation between the Chinese character weighted scores and the accuracy of the predicted difficulty levels, as well as the correlation between the word segmentation weighted scores and the accuracy of the predicted difficulty levels can be obtained more directly, so that it is easy to determine the distribution of weights between the two when weighting the Chinese character weighted scores and the word segmentation weighted scores.

[0084] Moreover, after the normalization process, the first initial predicted difficulty level and the second initial predicted difficulty level are further weighted, and the weighted results are rounded to obtain the predicted difficulty level of the text to be graded, thereby utilizing the influence of the difficulty evaluation value at the character level and the difficulty evaluation value at the word level on the predicted difficulty level to obtain a predicted difficulty level with higher accuracy.

[0085] In the process of normalizing the weighted scores of Chinese characters and word segmentation according to the preset difficulty levels, the intervals of the preset difficulty levels can be set according to actual needs.

[0086] In this embodiment, the grade and semester are mapped to different difficulty levels to obtain a preset difficulty level. As an example, the interval of the preset difficulty level can be set according to the grade information (grade 1 to grade 12 of high school) and the semester information (first semester and second semester). For example, the difficulty level of the first semester of the first grade can be set to 1, the difficulty level of the second semester of the first grade can be set to 2, ..., and so on. The difficulty level of the second semester of the third year of high school is set to 24. Accordingly, the interval of the preset difficulty level is 1 to 24. Among them, the greater the difficulty level, the higher the reading difficulty of the text to be tested.

[0087] Specifically, the preset difficulty level can be obtained by formula (1):

[0088] S=(C-1)*2+T (1)

[0089] Among them, S is used to indicate the preset difficulty level, C is used to indicate grade information, and T is used to indicate semester information.

[0090] In this embodiment, the normalization process of the weighted scores of Chinese characters includes: performing a first initial normalization process using a first normalization model, and rounding the result of the first initial normalization process. The first normalization model is

[0091]

[0092] Among them, L char is used to represent the result of the first initial normalization, x is used to represent the weighted score of the Chinese character, A char Used to indicate the minimum value of the reference value of the weighted score of Chinese characters, B char Used to indicate the maximum value of the reference value of the weighted score of Chinese characters, N min It is used to indicate the lowest level of the preset difficulty level, N max Used to indicate the highest level of the preset difficulty levels.

[0093] After obtaining the weighted score of the Chinese character, the weighted score of the Chinese character is compared with the minimum value and the maximum value of the weighted score reference value of the Chinese character. When the weighted score of the Chinese character is less than or equal to the minimum value of the weighted score reference value of the Chinese character, the result of the first initial normalization is set to the minimum value of the weighted score reference value of the Chinese character. Accordingly, the result after rounding the result of the first initial normalization is still the minimum value of the weighted score reference value of the Chinese character; when the weighted score of the Chinese character is greater than or equal to the maximum value of the weighted score reference value of the Chinese character, the result of the first initial normalization is set to the maximum value of the weighted score reference value of the Chinese character. Accordingly, the result after rounding the result of the first initial normalization is still the maximum value of the weighted score reference value of the Chinese character.

[0094] When the Chinese character weighted score is between the minimum value and the maximum value of the Chinese character weighted score reference value, by dividing the difference between the Chinese character weighted score and the minimum value of the Chinese character weighted score reference value by the difference between the maximum value and the minimum value of the Chinese character weighted score reference value, the position of the Chinese character weighted score in the entire Chinese character weighted score reference value interval can be obtained, and then mapped to the position in the entire preset difficulty level interval, thereby obtaining the first initial normalization result.

[0095] Similarly, the normalization process for the weighted word segmentation score includes: performing a second initial normalization using a second normalization model and rounding the result of the second initial normalization, wherein the second normalization model is

[0096]

[0097] Among them, L wor d is used to represent the result of the second initial normalization, y is used to represent the weighted score of the word segmentation, A word Used to represent the minimum value of the reference value of the word segmentation weighted score, B word Used to indicate the maximum value of the reference value of the word segmentation weighted score, N min It is used to indicate the lowest level of the preset difficulty level, N max Used to indicate the highest level of the preset difficulty levels.

[0098] For a detailed description of the second normalized model, please refer to the relevant description of the second normalized model, which will not be repeated here.

[0099] It should be noted that, in this embodiment, grades and semesters are mapped to different difficulty levels to obtain preset difficulty levels. Therefore, the lowest level of the preset difficulty level is 1, and the highest level of the preset difficulty level is 24.

[0100] Specifically, the weighted processing process can be expressed by formula (2):

[0101] Pred=w1×Pred char +w2×Pred word (2)

[0102] Among them, Pred is used to represent the result after weighted processing. char It is used to represent the result after rounding the result of the first initial normalization. word It is used to represent the result after rounding the result of the second initial normalization, and w1 and w2 represent weights respectively.

[0103] In this embodiment, when the result after weighting processing (ie, the value of Pred) is a decimal, it is rounded to obtain a prediction difficulty level with an integer value. For example, the result after weighting processing can be rounded to an integer.

[0104] It should be noted that, in one embodiment, before conducting the actual evaluation, the minimum value of the Chinese character weighted score reference value, the maximum value of the Chinese character weighted score reference value, the minimum value of the word segmentation weighted score reference value, the maximum value of the word segmentation weighted score reference value, and the weights corresponding to the first initial predicted difficulty level and the second initial predicted difficulty level can be adjusted and determined by using relevant data of the training text marked with the actual difficulty level, thereby improving the accuracy of the predicted difficulty level.

[0105] In other embodiments, the minimum and maximum values ​​of the Chinese character weighted score reference values, the minimum and maximum values ​​of the word segmentation weighted score reference values, and the weights corresponding to the first and second initial predicted difficulty levels may be adjusted over time based on actual circumstances. For example, the difficulty classification criteria for characters and words may change over time.

[0106] Execute step S6 to judge the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded. When the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded meet a threshold condition, adjust the predicted difficulty level.

[0107] After obtaining the predicted difficulty level, the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded are also judged. When the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the threshold conditions, the predicted difficulty level is adjusted, which can further improve the accuracy of the difficulty level finally obtained.

[0108] In this embodiment, when the predicted difficulty level and the difficulty assessment values ​​corresponding to the features in the feature set of the text to be graded meet threshold conditions, the predicted difficulty level is adjusted based on rules. By adjusting the predicted difficulty level based on rules, the amount of data computation required is effectively reduced, thereby ensuring the high efficiency of the text difficulty grading and assessment method. Furthermore, even if the adjustment amount for the predicted difficulty level and the difficulty assessment values ​​corresponding to the features in the feature set of the text to be graded do not satisfy a linear relationship, the predicted difficulty level can still be adjusted based on the rules.

[0109] Specifically, adjusting the predicted difficulty level using the predicted difficulty level and the difficulty assessment values ​​corresponding to each feature in the feature set of the text to be graded includes: calculating an average of the character difficulty values ​​corresponding to a first number of characters with the largest difficulty values ​​in the feature set of the text to be graded as the character difficulty average, and calculating an average of the word difficulty values ​​corresponding to a second number of characters with the largest difficulty values ​​in the feature set of the text to be graded as the word difficulty average; accordingly, determining when the predicted difficulty level and the difficulty assessment values ​​corresponding to each feature in the feature set of the text to be graded meet a threshold condition, the threshold condition including a first threshold condition, the first threshold condition including: the predicted difficulty level is less than a first preset difficulty prediction range value, the character difficulty average is greater than or equal to a first preset threshold value, and the word difficulty average is greater than or equal to a second preset threshold value; when the predicted difficulty level and the difficulty assessment values ​​corresponding to each feature in the feature set of the text to be graded meet the first threshold condition, adding the first preset difficulty level value to the predicted difficulty level; otherwise, maintaining the predicted difficulty level.

[0110] In this embodiment, the first number and the second number are both integers greater than 1. As an example, since the total number of characters in the test text is generally greater than the total number of words, the first number is greater than the second number. It is understood that the relationship between the first number and the second number is not limited to the above case.

[0111] In the process of determining the minimum value of the Chinese character weighted score reference value, the maximum value of the Chinese character weighted score reference value, the minimum value of the word segmentation weighted score reference value, the maximum value of the word segmentation weighted score reference value, and the weights corresponding to the first initial predicted difficulty level and the second initial predicted difficulty level, due to the influence of the text actually used, there may be a situation where the difficulty grading accuracy of the test text with a higher actual difficulty level is not high, that is, after the test text is graded for difficulty, it is easy to have a predicted difficulty level lower than its actual difficulty level. Therefore, when the difficulty assessment value and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the first threshold condition, the predicted difficulty level is added to the first preset difficulty level value.

[0112] In the process of determining the minimum value of the Chinese character weighted score reference value, the maximum value of the Chinese character weighted score reference value, the minimum value of the word segmentation weighted score reference value, the maximum value of the word segmentation weighted score reference value, and the weights corresponding to the first initial predicted difficulty level and the second initial predicted difficulty level, if the actual difficulty level of the text used is too low, it will easily lead to inaccurate difficulty grading of the test text with a high actual difficulty level. Therefore, it is necessary to add the predicted difficulty level to the first preset difficulty level value.

[0113] Correspondingly, in another embodiment, it may also be that: the threshold condition includes a second threshold condition, the second threshold condition includes: the predicted difficulty level is greater than the second preset difficulty prediction range value, and the average difficulty value of the character is less than or equal to the third preset threshold value, and the average difficulty value of the word is less than or equal to the fourth preset threshold value. When the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the second threshold condition, the predicted difficulty level is subtracted from the second preset difficulty level value, otherwise the predicted difficulty level is maintained.

[0114] Based on similar reasons as mentioned above, in the process of determining the minimum value of the Chinese character weighted score reference value, the maximum value of the Chinese character weighted score reference value, the minimum value of the word segmentation weighted score reference value, the maximum value of the word segmentation weighted score reference value, and the weights corresponding to the first initial predicted difficulty level and the second initial predicted difficulty level, if the actual difficulty level of the text used is too high, it will easily lead to inaccurate difficulty grading of the test text with a low actual difficulty level. Therefore, it is necessary to subtract the second preset difficulty level value from the predicted difficulty level.

[0115] In addition, since the weighted score of Chinese characters is obtained by summing the character difficulty values ​​of all characters, and the weighted score of word segmentation is obtained by summing the word difficulty values ​​of all words, even if the predicted difficulty level deviates due to the character difficulty value of individual characters, or the predicted difficulty level deviates due to the word difficulty value of individual words, the accuracy of the final difficulty level can be improved by adjusting the predicted difficulty level.

[0116] It should be noted that the first preset difficulty prediction range value, the first preset threshold value, the second preset threshold value, the second preset difficulty prediction range value, the third preset threshold value and the fourth preset threshold value can be set according to actual conditions.

[0117] It should also be noted that the average difficulty value of the first number of characters with the highest difficulty values ​​in the feature set of the text to be graded is calculated as the average difficulty value of the characters. If the value of the first number is too small, it will easily lead to excessive noise, which is not conducive to ensuring the effectiveness of adjusting the predicted difficulty level. If the value of the first number is too large, it will easily lose the meaning of calculating the average difficulty value of the first number of characters with the highest difficulty values. Therefore, in this embodiment, the first number is any integer between 3 and 15.

[0118] Based on similar reasons as above, in this embodiment, the second number is any integer from 2 to 10.

[0119] It should be noted that there may be some texts to be tested whose overall difficulty values ​​are not high, but some of their words or characters appear less frequently, and the corresponding difficulty values ​​are high. Therefore, when calculating the average difficulty values ​​of the first number of characters with the highest difficulty values ​​in the text to be graded, and calculating the average difficulty values ​​of the first number of characters with the highest difficulty values ​​in the text to be graded, the average difficulty values ​​of the first number of characters with the highest difficulty values ​​and the average difficulty values ​​of the first number of characters with the highest difficulty values ​​are both greater than their corresponding preset thresholds, thereby causing the adjusted predicted difficulty level to be greater than its actual difficulty level, and correspondingly obtaining an incorrect difficulty level. For this reason, in addition to using characters and words to measure the difficulty level of the text to be tested, this embodiment can also use sentences to measure the difficulty level of the text to be tested, thereby further improving the accuracy of the difficulty level finally obtained.

[0120] In this embodiment, the threshold condition includes a third threshold condition, which further includes: the number of sentences is greater than or equal to a fifth preset threshold, the average sentence length is greater than or equal to a sixth preset threshold, and the predicted difficulty level is less than a third preset difficulty level value. When the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the third threshold condition, the predicted difficulty level is set to the third preset difficulty level value.

[0121] Accordingly, in another embodiment, the threshold condition includes a fourth threshold condition, which further includes: the number of sentences is less than or equal to a seventh preset threshold, the average sentence length is less than or equal to an eighth preset threshold, and the predicted difficulty level is greater than a fourth preset difficulty level value. When the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the fourth threshold condition, the predicted difficulty level is set to the fourth preset difficulty level value.

[0122] As an example, when the number of sentences in the text to be tested is greater than or equal to 70, the average length of the sentences is greater than or equal to 30, and the predicted difficulty level is less than 18, the predicted difficulty level calculated by the weighted score of Chinese characters and the weighted score of word segmentation can be set to 18.

[0123] In other embodiments, when the number of sentences in the text to be tested is less than or equal to 30, the average length of the sentences is less than or equal to 15, and the predicted difficulty level is greater than 10, the predicted difficulty level calculated by the weighted score of Chinese characters and the weighted score of word segmentation can be set to 10.

[0124] It should be noted that, in a specific implementation, the fifth preset threshold, the sixth preset threshold, the seventh preset threshold, and the eighth preset threshold can be flexibly set according to actual conditions.

[0125] Correspondingly, an embodiment of the present invention also provides a device for evaluating text difficulty levels. Figure 3 This is a functional block diagram of an embodiment of a text difficulty grading evaluation device of the present invention.

[0126] The text difficulty grading and evaluation device comprises: a text acquisition module 10 for acquiring a text to be graded; a preprocessing module 20 for preprocessing the text to be graded and acquiring a feature set of the text to be graded, wherein the feature set comprises a plurality of features related to text granularity, and the feature types in the feature set comprise at least characters and words; a difficulty evaluation value acquisition module 30 for acquiring a difficulty evaluation value corresponding to each feature in the feature set of the text to be graded, wherein the difficulty evaluation value comprises a character difficulty value corresponding to each character and a word difficulty value corresponding to each word, wherein the character difficulty value is negatively correlated with the character frequency, and the word difficulty value is negatively correlated with the word frequency; a ranking score acquisition module 40 for acquiring The sum of the character difficulty values ​​of all characters in the feature set of the text to be graded is used as a weighted score for the Chinese characters, and the sum of the word difficulty values ​​of all words in the feature set of the text to be graded is obtained as a weighted score for the word segmentation; a difficulty level prediction module 50 is used to obtain a predicted difficulty level of the text to be graded using the weighted scores for the Chinese characters and the weighted scores for the word segmentation; a difficulty level adjustment module 60 is used to judge the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded, and adjust the predicted difficulty level when the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded meet a threshold condition.

[0127] The text to be graded serves as the subject for difficulty grading. The text difficulty grading device is used to grade the text to be graded and determine its difficulty level so that the user can obtain a text that matches their reading ability. In this embodiment, the text to be graded is a Chinese text, and the language of the text to be graded is Simplified Chinese. For a detailed description of the text to be graded, please refer to the corresponding description in the previous embodiment and will not be repeated here.

[0128] Through the preprocessing module 20, features related to the text granularity are obtained. The features are related to the difficulty level of the text and can represent the reading difficulty of the text to be tested from features of different dimensions, so that the difficulty evaluation value corresponding to each feature can be used as an evaluation factor for calculating the predicted difficulty level.

[0129] In this embodiment, the feature types in the feature set include at least characters and words. Character-level features and word-level features are the two types of features with the smallest text granularity, which are conducive to more accurately characterizing the reading difficulty of the test text, thereby improving the accuracy of subsequent text difficulty grading.

[0130] In this embodiment, the preprocessing module 20 includes a text segmentation unit for segmenting the text to be classified according to different text granularities and obtaining a feature set corresponding to the text to be classified. Specifically, the text segmentation unit includes a word segmentation subunit for segmenting the text to be classified and obtaining word-level features of the text to be classified. As an example, the word segmentation subunit can be a word segmenter, such as Jieba.

[0131] In this embodiment, the feature type in the feature set also includes sentences. Accordingly, the text splitting unit also includes a sentence segmentation unit for performing sentence segmentation processing on the text to be classified. The sentence segmentation unit is used to traverse the text to be classified, determine the position of the sentence segmentation symbols in the text to be classified, and decompose the text to be classified into multiple separate sentences based on the position of the sentence segmentation symbols. For example, the sentence segmentation symbols can be punctuation marks.

[0132] In this embodiment, the preprocessing module 20 further includes a first filtering unit configured to perform a first filtering process on the text to be graded before the splitting process to filter out content that does not meet the text requirements. This first filtering process is used to remove text for which a difficulty level cannot be predicted using character-level and word-level features. For example, ancient Chinese text and non-Chinese text should be removed.

[0133] The difficulty assessment value acquisition module 30 is used to obtain the difficulty assessment value corresponding to each feature. The difficulty assessment value includes a character difficulty value corresponding to each character and a word difficulty value corresponding to each word. Subsequently, a Chinese character weighted score is obtained based on the character difficulty value, and a word segmentation weighted score is obtained based on the word difficulty value. The Chinese character weighted score and word segmentation weighted score are used as evaluation factors for evaluating the difficulty level of the test text. The character difficulty value is negatively correlated with the character frequency. The lower the character frequency, the greater the difficulty of the character, and the higher the character difficulty value. Similarly, the word difficulty value is negatively correlated with the word frequency. The lower the word frequency, the greater the difficulty of the word, and the higher the word difficulty value. Therefore, by obtaining the character difficulty value and word difficulty value, it is easier to more accurately characterize the difficulty level of the entire text.

[0134] In this embodiment, the difficulty assessment value acquisition module 30 includes a dictionary acquisition unit configured to acquire a dictionary. The dictionary includes a dictionary and a lexicon. The dictionary contains difficulty values ​​for each character, while the lexicon contains difficulty values ​​for each word. Accordingly, the dictionary is used to obtain the character difficulty value for each character and the word difficulty value for each word in the test text.

[0135] On the one hand, directly obtaining the corresponding character and word difficulty values ​​from the dictionary reduces the complexity of obtaining character and word difficulty values. On the other hand, compared with large-scale corpora, the number of characters and words in the text to be graded is relatively small. If the difficulty assessment value corresponding to each feature is directly set according to the frequency of occurrence of each feature in the feature type of the text to be graded, it is easy for the difficulty value of some features to be much greater or much less than their true difficulty value due to insufficient sample size, thus reducing the accuracy of the text difficulty classification. Therefore, obtaining the difficulty value of each character from the dictionary as the character difficulty value, and obtaining the difficulty value of each word from the lexicon as the word difficulty value, is conducive to improving the accuracy of the character and word difficulty values ​​in the text to be tested.

[0136] Specifically, the dictionary acquisition unit includes: a reference text set acquisition subunit, which is used to acquire a large number of reference texts to form a reference text set; a feature extraction subunit, which is used to extract features related to the difficulty value of each text in the reference text set; a statistical subunit, which is used to count the frequency of occurrence of each feature according to the feature type; a sorting subunit, which is used to sort according to the frequency of occurrence of each feature to obtain the dictionary, and the dictionary includes a dictionary library and a lexicon library; a processing subunit, which is used to determine the difficulty value corresponding to each character in the dictionary library and the difficulty value corresponding to each word in the lexicon library according to the sorting results of the corresponding feature types of the dictionary library and the lexicon library, and normalize the difficulty value corresponding to each character and the difficulty value corresponding to each word.

[0137] In a specific embodiment, the reference text set can be sourced from any unannotated text on the Internet. For example, a large amount of text can be obtained from various libraries and book websites and jointly used as the reference text set. At the same time, the reference text set can be continuously updated with new text or obsolete text removed over time to ensure the timeliness of each text in the reference text set. It can be understood that during the process of obtaining the reference text set, some content irrelevant to the text content, such as watermarks and signatures, can be directly removed.

[0138] For ease of understanding, the following takes the determination of the difficulty value of a character as an example for illustration. For example, if there are 10,000 characters in the dictionary library, among which the frequency of "了" is the highest, at 500, and the frequency of "牖" is the lowest, at 1, and the frequencies of the remaining 9,998 characters are between 2 and 499. Among them, the greater the frequency of a character appears, the lower the corresponding difficulty value. According to the frequency of the character appearance, the difficulty values of the 10,000 characters are set from the difficulty value range [0, 1].

[0139] For example, when performing normalization processing, the difficulty value of the character ranked first in the frequency order (i.e., "了") can be set to 0, and the difficulty value of the character ranked 10,000th in the frequency order (i.e., "牖") can be set to 1. The difficulty values of the remaining characters are set in sequence according to their appearance frequencies.

[0140] For other features in the language dictionary library, such as the difficulty values of words in the dictionary library, the determination process for the difficulty values of characters can be referred to, and details will not be elaborated here.

[0141] It should be noted that the language dictionary library acquisition unit further includes: a second filtering processing sub-unit, used to screen out the characters with the highest frequencies among the first preset value before normalizing the difficulty values corresponding to each character, and remove the remaining characters; a third filtering processing sub-unit, used to screen out the words with the highest frequencies among the second preset value before normalizing the difficulty values corresponding to each word, and remove the remaining words.

[0142] By the second filtering processing sub-unit, removing the characters with lower appearance frequencies can remove the noise that may cause prediction errors, thereby improving the accuracy of predicting the difficulty level. Similarly, by the third filtering processing sub-unit, it can play a role in removing rare words, and correspondingly remove the noise that may cause prediction errors, and further improve the accuracy of predicting the difficulty level. In this embodiment, the first preset value is from 3,500 to 4,500, and the second preset value is from 45,000 to 55,000.

[0143] Accordingly, in this embodiment, the difficulty assessment value acquisition module 30 further includes: a first retrieval unit, configured to obtain, based on a dictionary library, the difficulty value of each character in the feature set of the text to be graded as a character difficulty value; and a second retrieval unit, configured to obtain, based on the dictionary library, the difficulty value of each word in the feature set of the text to be graded as a word difficulty value.

[0144] Specifically, the first retrieval unit is used to search for words in the feature set of the text to be graded in the dictionary library. When the same word is retrieved, the difficulty value of the corresponding word in the dictionary library is used as the difficulty value of the corresponding word in the feature set of the text to be graded. When the same word is not retrieved, the difficulty value of the unretrieved word in the feature set of the text to be graded is set to the maximum difficulty value in the dictionary library.

[0145] Similarly, the second retrieval unit is used to search the dictionary library for words in the feature set of the text to be graded. When the same word is retrieved, the difficulty value of the corresponding word in the dictionary library is used as the difficulty value of the corresponding word in the feature set of the text to be graded. When the same word is not retrieved, the difficulty value of the unretrieved word in the feature set of the text to be graded is set to the maximum difficulty value in the dictionary library.

[0146] In this embodiment, the difficulty evaluation value acquisition module 30 further includes: a sentence number extraction unit for acquiring the number of sentences in the text to be graded; and a sentence length acquisition unit for acquiring the average length of sentences in the text to be graded.

[0147] The number of sentences and average sentence length can also be used to characterize the difficulty level of the test text. Longer sentences are associated with a higher difficulty level, and the text itself is more difficult to read. A greater number of sentences indicates a higher difficulty level.

[0148] By obtaining the number of sentences and the average length of sentences in the text to be graded, the predicted difficulty level can be adjusted based on the difficulty evaluation value of sentence-level features, thereby further improving the accuracy of difficulty grading of the text to be graded.

[0149] The ranking score acquisition module 40 is used to obtain a Chinese character weighted score based on the sum of the character difficulty values ​​of all characters in the feature set of the text to be graded, and to obtain a word segmentation weighted score based on the sum of the word difficulty values ​​of all words in the feature set of the text to be graded.

[0150] The character difficulty value is negatively correlated with the character frequency. The lower the character frequency, the greater the difficulty of the character, and the higher the character difficulty value. The weighted score of Chinese characters is obtained by summing the character difficulty values ​​of all characters. The sum of the character difficulty values ​​can represent the overall difficulty of the characters in the text to be tested. Therefore, compared with the difficulty assessment values ​​corresponding to other types of features, the weighted score of Chinese characters can better represent the difficulty of the characters in the entire text. In other words, the weighted score of Chinese characters is more strongly correlated with the actual difficulty of the text to be tested, and thus can represent the difficulty of the text to be graded.

[0151] Similarly, the word difficulty value is negatively correlated with the word frequency. The lower the word frequency, the greater the difficulty of the word, and the higher the word difficulty value. The weighted score of word segmentation is obtained by summing the word difficulty values ​​of all words. The average value of the word difficulty value can represent the overall difficulty of the words in the text to be tested. Therefore, the weighted score of word segmentation can better represent the difficulty of the words in the entire text. In other words, the weighted score of word segmentation is more strongly correlated with the actual difficulty of the text to be tested, and thus can also represent the difficulty of the text to be graded.

[0152] Correspondingly, using the weighted scores of Chinese characters and word segmentation to obtain the predicted difficulty level of the text to be graded can obtain the overall difficulty of characters and words. Characters and words are the two types of features with the smallest text granularity, so they are conducive to improving the accuracy of text difficulty grading and making it easier for users to select suitable reading texts.

[0153] Furthermore, this embodiment only requires the use of weighted Chinese character scores and weighted word segmentation scores to obtain the predicted difficulty level, which helps reduce the amount of data calculations during the difficulty level prediction process, thereby improving the evaluation efficiency of the text difficulty grading assessment method. At the same time, using weighted Chinese character scores and weighted word segmentation scores to obtain the predicted difficulty level also means that the weighted Chinese character scores and weighted word segmentation scores are used to construct a model for predicting the difficulty level, which correspondingly simplifies the complexity of the model for predicting the difficulty level and also helps improve the evaluation efficiency of the text difficulty grading assessment method.

[0154] In this embodiment, the difficulty level prediction module 50 includes: a preset difficulty level configuration unit for determining a preset difficulty level; a normalization processing unit for normalizing the weighted scores of Chinese characters and the weighted scores of word segments according to the preset difficulty levels, respectively, to obtain a first initial predicted difficulty level corresponding to the weighted scores of Chinese characters, and a second initial predicted difficulty level corresponding to the weighted scores of word segments; a weighted processing unit for performing weighted processing on the first initial predicted difficulty level and the second initial predicted difficulty level; and a second data processing unit for rounding the result after weighted processing to obtain the predicted difficulty level of the text to be graded.

[0155] By weighting the first and second initial predicted difficulty levels, the impact of the overall difficulty values ​​of characters and words on the difficulty level can be comprehensively considered. Specifically, by first normalizing the character weighted scores and word segmentation weighted scores according to the preset difficulty level, the mapping relationship between the character weighted scores and the predicted difficulty level, as well as the mapping relationship between the word segmentation weighted scores and the predicted difficulty level, can be more intuitively obtained. This means that the correlation between the character weighted scores and the accuracy of the predicted difficulty level, as well as the correlation between the word segmentation weighted scores and the accuracy of the predicted difficulty level, can be more directly obtained. This makes it easier to determine the weight distribution between the character weighted scores and the word segmentation weighted scores when weighting them. Furthermore, after normalization, the first and second initial predicted difficulty levels are further weighted, and the weighted results are rounded to obtain the predicted difficulty level of the text to be graded. This allows the influence of the character-level difficulty assessment value and the word-level difficulty assessment value on the predicted difficulty level to obtain a more accurate predicted difficulty level.

[0156] It should be noted that the interval of the preset difficulty level can be set according to actual needs. In this embodiment, the preset difficulty level has a mapping relationship with the grade and the semester. As an example, the preset difficulty level configuration unit sets the interval of the preset difficulty level according to the grade information (grade 1 of primary school to grade 12 of high school) and the semester information (first semester and second semester). For example, the difficulty level of the first semester of the first grade can be set to 1, the difficulty level of the second semester of the first grade can be set to 2, ..., and so on. The difficulty level of the second semester of the third year of high school is set to 24. Correspondingly, the interval of the preset difficulty level is 1 to 24. Among them, the greater the difficulty level, the higher the reading difficulty of the text to be tested.

[0157] Specifically, the preset difficulty level configuration unit obtains the preset difficulty level through formula (1):

[0158] S=(C-1)*2+T (1)

[0159] Among them, S is used to indicate the preset difficulty level, C is used to indicate grade information, and T is used to indicate semester information.

[0160] In this embodiment, the normalization processing unit includes a first normalization processing unit sub-unit, which is used to perform a first initial normalization using a first normalization model and round the result of the first initial normalization. The first normalization model is

[0161]

[0162] Among them, L char is used to represent the result of the first initial normalization, x is used to represent the weighted score of the Chinese character, Achar Used to indicate the minimum value of the reference value of the weighted score of Chinese characters, B char Used to indicate the maximum value of the reference value of the weighted score of Chinese characters, N min It is used to indicate the lowest level of the preset difficulty level, N max Used to indicate the highest level of the preset difficulty levels.

[0163] After obtaining the Chinese character weighted score, the Chinese character weighted score is compared with the minimum and maximum values ​​of the Chinese character weighted score reference value. When the Chinese character weighted score is less than or equal to the minimum value of the Chinese character weighted score reference value, the result of the first initial normalization is set to the minimum value of the Chinese character weighted score reference value. Accordingly, the result of rounding the result of the first initial normalization is still the minimum value of the Chinese character weighted score reference value. When the Chinese character weighted score is greater than or equal to the maximum value of the Chinese character weighted score reference value, the result of the first initial normalization is set to the maximum value of the Chinese character weighted score reference value. Accordingly, the result of rounding the result of the first initial normalization is still the maximum value of the Chinese character weighted score reference value. When the Chinese character weighted score is between the minimum and maximum values ​​of the Chinese character weighted score reference value, by dividing the difference between the Chinese character weighted score and the minimum value of the Chinese character weighted score reference value by the difference between the maximum and minimum values ​​of the Chinese character weighted score reference value, the position of the Chinese character weighted score in the entire Chinese character weighted score reference value interval can be obtained, and then mapped to the position in the entire preset difficulty level interval, thereby obtaining the first initial normalization result.

[0164] Similarly, the normalization processing unit further includes a second normalization processing unit subunit, which is used to perform a second initial normalization using a second normalization model and round the result of the second initial normalization. The second normalization model is

[0165]

[0166] Among them, L word is used to represent the result of the second initial normalization, y is used to represent the weighted score of the word segmentation, and A word Used to represent the minimum value of the reference value of the word segmentation weighted score, B word Used to indicate the maximum value of the reference value of the word segmentation weighted score, N min It is used to indicate the lowest level of the preset difficulty level, N max Used to indicate the highest level of the preset difficulty levels.

[0167] For a detailed description of the second normalized model, please refer to the relevant description of the second normalized model, which will not be repeated here.

[0168] It should be noted that, in this embodiment, grades and semesters are mapped to different difficulty levels to obtain preset difficulty levels. Therefore, the lowest level of the preset difficulty level is 1, and the highest level of the preset difficulty level is 24.

[0169] Specifically, the weighted processing unit performs the weighted processing according to formula (2):

[0170] Pred=w1×Pred char +w2×Pred word (2)

[0171] Among them, Pred is used to represent the result after weighted processing. char It is used to represent the result after rounding the result of the first initial normalization. word It is used to represent the result after rounding the result of the second initial normalization, and w1 and w2 represent weights respectively.

[0172] In this embodiment, the second data processing unit performs a rounding operation on the weighted result (ie, the value of Pred) to obtain a prediction difficulty level with an integer value. For example, the weighted result can be rounded off.

[0173] It should be noted that, in one embodiment, before conducting the actual evaluation, the minimum value of the Chinese character weighted score reference value, the maximum value of the Chinese character weighted score reference value, the minimum value of the word segmentation weighted score reference value, the maximum value of the word segmentation weighted score reference value, and the weights corresponding to the first initial predicted difficulty level and the second initial predicted difficulty level can be adjusted and determined by using relevant data of the training text marked with the actual difficulty level, thereby improving the accuracy of the predicted difficulty level. In other embodiments, the minimum value of the Chinese character weighted score reference value, the maximum value of the Chinese character weighted score reference value, the minimum value of the word segmentation weighted score reference value, the maximum value of the word segmentation weighted score reference value, and the weights corresponding to the first initial predicted difficulty level and the second initial predicted difficulty level can also be continuously adjusted according to actual conditions over time. For example, the difficulty classification standards for characters and words have changed over time.

[0174] After obtaining the predicted difficulty level, the difficulty level adjustment module 60 is used to judge the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded. When the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the threshold conditions, the predicted difficulty level is adjusted, which can further improve the accuracy of the difficulty level finally obtained.

[0175] In this embodiment, the difficulty level adjustment module 60 is used to adjust the predicted difficulty level based on rules. Specifically, the difficulty level adjustment module 60 includes: a third data processing unit for calculating the average of the character difficulty values ​​corresponding to the first number of characters with the highest difficulty values ​​in the feature set of the text to be graded, as the average character difficulty value; and calculating the average of the word difficulty values ​​corresponding to the first number of characters with the highest difficulty values ​​in the feature set of the text to be graded, as the average word difficulty value; and a judgment unit for judging whether the predicted difficulty level and the difficulty assessment values ​​corresponding to each feature in the feature set of the text to be graded meet a threshold condition.

[0176] In one embodiment, the threshold condition includes a first threshold condition, which includes: the predicted difficulty level is less than a first preset difficulty prediction range value, and the average difficulty value of the character is greater than or equal to the first preset threshold value, and the average difficulty value of the word is greater than or equal to a second preset threshold value; the adjustment unit is used to add the first preset difficulty level value to the predicted difficulty level when the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the first threshold condition, otherwise maintain the predicted difficulty level.

[0177] In another embodiment, the threshold condition may include a second threshold condition, the second threshold condition including: the predicted difficulty level is greater than a second preset difficulty prediction range value, and the average difficulty value of the character is less than or equal to a third preset threshold value, and the average difficulty value of the word is less than or equal to a fourth preset threshold value, and the adjustment unit is configured to, when the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the second threshold condition, subtract the second preset difficulty level value from the predicted difficulty level, otherwise maintain the predicted difficulty level.

[0178] It should also be noted that there may be some texts to be tested whose overall difficulty values ​​are not high, but some of their words or characters appear less frequently, and the corresponding difficulty values ​​are high. Therefore, when calculating the average difficulty values ​​of the first number of characters with the highest difficulty values ​​in the text to be graded, and calculating the average difficulty values ​​of the first number of characters with the highest difficulty values ​​in the text to be graded, the average difficulty values ​​of the first number of characters with the highest difficulty values ​​and the average difficulty values ​​of the first number of characters with the highest difficulty values ​​are both greater than their corresponding preset thresholds, thereby causing the adjusted predicted difficulty level to be greater than its actual difficulty level, and correspondingly obtaining an erroneous difficulty level value. For this reason, in addition to using characters and words to measure the difficulty level of the text to be graded, this embodiment can also use sentences to measure the difficulty level of the text to be graded, thereby further improving the accuracy of the difficulty level finally obtained.

[0179] In this embodiment, the threshold condition includes a third threshold condition, which further includes: the number of sentences is greater than or equal to a fifth preset threshold, the average sentence length is greater than or equal to a sixth preset threshold, and the predicted difficulty level is less than a third preset difficulty level value. When the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the third threshold condition, the adjustment unit sets the predicted difficulty level to the third preset difficulty level value.

[0180] Accordingly, in another embodiment, the threshold condition includes a fourth threshold condition, which further includes: the number of sentences is less than or equal to a seventh preset threshold, the average sentence length is less than or equal to an eighth preset threshold, and the predicted difficulty level is greater than a fourth preset difficulty level value. When the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the fourth threshold condition, the adjustment unit sets the predicted difficulty level to the fourth preset difficulty level value.

[0181] An embodiment of the present invention further provides a device, which can implement the text difficulty grading assessment method provided by the embodiment of the present invention by loading the above-mentioned text difficulty grading assessment method in the form of a program.

[0182] refer to Figure 3 , which shows a hardware structure diagram of a device provided by an embodiment of the present invention. The device of this embodiment includes: at least one processor 01, at least one communication interface 02, at least one memory 03 and at least one communication bus 04.

[0183] In this embodiment, the number of the processor 01 , the communication interface 02 , the memory 03 and the communication bus 04 is at least one, and the processor 01 , the communication interface 02 and the memory 03 communicate with each other through the communication bus 04 .

[0184] The communication interface 02 may be an interface of a communication module for network communication, such as an interface of a GSM module.

[0185] The processor 01 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the text difficulty grading and evaluation method described in this embodiment.

[0186] The memory 03 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0187] The memory 03 stores one or more computer instructions, and the one or more computer instructions are executed by the processor 01 to implement the text difficulty grading assessment method provided in the above embodiment.

[0188] It should be noted that the above-mentioned terminal device may also include other devices (not shown) that may not be necessary for understanding the contents disclosed in the embodiments of the present invention; given that these other devices may not be necessary for understanding the contents disclosed in the embodiments of the present invention, the embodiments of the present invention will not introduce them one by one.

[0189] An embodiment of the present invention further provides a storage medium storing one or more computer instructions, wherein the one or more computer instructions are used to implement the text difficulty grading assessment method provided in the aforementioned embodiment.

[0190] In the text difficulty grading and evaluation method provided by this embodiment, the sum of the character difficulty values ​​of all characters is obtained as the Chinese character weighted score, and the sum of the word difficulty values ​​of all words is obtained as the word segmentation weighted score. The Chinese character weighted score and the word segmentation weighted score are then used to obtain the predicted difficulty level of the text to be graded, and the predicted difficulty level is adjusted. Among them, the character difficulty value is negatively correlated with the character frequency. The lower the character frequency, the greater the difficulty of the character, and the higher the character difficulty value. The character difficulty value will affect the difficulty of the text to be graded. The Chinese character weighted score is obtained by summing the character difficulty values ​​of all characters. The sum of the character difficulty values ​​obtained can represent the overall difficulty of the characters in the text to be tested. Therefore, compared with the difficulty evaluation values ​​corresponding to other types of features, the weighted score of Chinese characters can better represent the difficulty of the characters in the entire text, thereby representing the difficulty of the text to be graded. Similarly, the word difficulty value is negatively correlated with the word frequency. The lower the word frequency, the greater the difficulty of the word, and the higher the word difficulty value. The word difficulty value will also affect the difficulty of the text to be graded. The weighted score of word segmentation is obtained by summing the word difficulty values ​​of all words. The sum of the word difficulty values ​​can represent the word difficulty of the text to be graded. Overall difficulty: Compared with the difficulty evaluation values ​​corresponding to other types of features, the weighted score of the word segmentation is more capable of characterizing the difficulty of the words in the entire text, thereby characterizing the difficulty of the text to be graded. Accordingly, using the weighted score of Chinese characters and the weighted score of the word segmentation to obtain the predicted difficulty level of the text to be graded can obtain the overall difficulty of characters and words, and characters and words are the two types of features with the smallest text granularity, so it is beneficial to improve the accuracy of text difficulty grading and facilitate users to select suitable reading texts; moreover, the embodiment of the present invention only needs to obtain the predicted difficulty level through the weighted score of Chinese characters and the weighted score of the word segmentation, which is beneficial to reduce the amount of data calculation in the process of predicting the difficulty level, thereby improving the evaluation efficiency of the text difficulty grading evaluation method; in addition, after obtaining the predicted difficulty level, the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded is also judged. When the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded meet the threshold condition, the predicted difficulty level is adjusted, which can further improve the accuracy of the final difficulty level.

[0191] The embodiments of the present invention described above are combinations of elements and features of the present invention. Unless otherwise mentioned, the elements or features may be considered as optional. Each element or feature may be put into practice without being combined with other elements or features. In addition, the embodiments of the present invention may be constructed by combining some elements and / or features. The order of operations described in the embodiments of the present invention may be rearranged. Some configurations of any one embodiment may be included in another embodiment and may be replaced by the corresponding configuration of another embodiment. It is obvious to those skilled in the art that claims that do not have a clear reference relationship to each other in the appended claims may be combined into embodiments of the present invention, or may be included as new claims in amendments after submitting this application.

[0192] The embodiments of the present invention may be implemented by various means such as hardware, firmware, software, or a combination thereof. In a hardware configuration, the method according to the exemplary embodiment of the present invention may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, etc.

[0193] In a firmware or software configuration, the embodiments of the present invention may be implemented in the form of modules, procedures, functions, and the like. Software codes may be stored in a memory unit and executed by a processor. The memory unit may be located inside or outside the processor and may send and receive data to and from the processor via various known means.

[0194] The above description of the disclosed embodiments will enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

[0195] Although the embodiments of the present invention are disclosed above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be based on the scope defined by the claims.

Claims

1. A text difficulty grading assessment method, characterized in that: include: Get the text to be graded; Preprocessing the text to be classified to obtain a feature set of the text to be classified, wherein the feature set includes multiple features related to text granularity, and the feature types in the feature set include at least characters and words; Obtaining a difficulty assessment value corresponding to each feature in the feature set of the text to be graded, wherein the difficulty assessment value includes a character difficulty value corresponding to each character and a word difficulty value corresponding to each word, wherein the character difficulty value is negatively correlated with the character frequency, and the word difficulty value is negatively correlated with the word frequency; Obtaining the sum of the character difficulty values ​​of all characters in the feature set of the text to be classified as a weighted score of the Chinese characters, and obtaining the sum of the word difficulty values ​​of all words in the feature set of the text to be classified as a weighted score of the word segmentation; Obtaining a predicted difficulty level of the text to be graded using the weighted scores of the Chinese characters and the weighted scores of the word segments; The predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded are judged, and when the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded meet a threshold condition, the predicted difficulty level is adjusted.

2. The text difficulty grading evaluation method according to claim 1, characterized in that: Preprocessing the text to be classified includes: splitting the text to be classified according to different text granularities, and obtaining a feature set corresponding to the text to be classified.

3. The text difficulty grading evaluation method according to claim 2, characterized in that: The pre-processing of the text to be classified further includes: before the text to be classified is split according to different text granularities, a first filtering process is performed on the text to be classified to filter out content that does not meet text requirements.

4. The text difficulty grading evaluation method according to claim 1, wherein: Before obtaining the difficulty evaluation value corresponding to each feature, the method further includes: obtaining a dictionary library, including a dictionary library and a lexicon library, wherein the dictionary library is marked with the difficulty value of each character, and the lexicon library is marked with the difficulty value of each word; Obtaining the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded includes: obtaining the difficulty value of each character in the feature set of the text to be graded as a character difficulty value according to the dictionary library; obtaining the difficulty value of each word in the feature set of the text to be graded as a word difficulty value according to the dictionary library.

5. The text difficulty grading evaluation method according to claim 4, characterized in that: Obtaining, based on the dictionary library, a difficulty value of each character in the feature set of the text to be classified as a character difficulty value comprises: searching the dictionary library for the characters in the feature set of the text to be classified; when identical characters are retrieved, using the difficulty value of the corresponding character in the dictionary library as the difficulty value of the corresponding character in the feature set of the text to be classified; and when identical characters are not retrieved, setting the difficulty value of the unretrieved character in the feature set of the text to be classified as the maximum difficulty value in the dictionary library; Obtaining, based on the dictionary library, the difficulty value of each word in the feature set of the text to be graded as the word difficulty value includes: searching the word in the feature set of the text to be graded in the dictionary library; when the same word is retrieved, using the difficulty value of the corresponding word in the dictionary library as the difficulty value of the corresponding word in the feature set of the text to be graded; when the same word is not retrieved, setting the difficulty value of the unretrieved word in the feature set of the text to be graded to the maximum value of the difficulty values ​​in the dictionary library.

6. The text difficulty grading evaluation method according to claim 4, characterized in that: The acquiring of the dictionary includes: acquiring a reference text set and extracting features of each text in the reference text set; Count the frequency of each feature according to the feature type; Sorting the features according to their frequencies of occurrence to obtain the dictionary library, which includes a dictionary library and a lexicon library; The difficulty value corresponding to each character in the dictionary library and the difficulty value corresponding to each word in the dictionary library are determined according to the sorting results of the dictionary library and the feature types corresponding to the dictionary library, and the difficulty value corresponding to each character and the difficulty value corresponding to each word are normalized respectively.

7. The text difficulty grading evaluation method according to claim 6, characterized in that: Before normalizing the difficulty values ​​corresponding to each character and each word, respectively, obtaining the dictionary further includes: Performing a second filtering process to select the words with the highest frequency before the first preset value and removing the remaining words; Perform a third filtering process to filter out the words with the highest frequencies before the second preset value, and remove the remaining words.

8. The text difficulty grading evaluation method according to claim 7, characterized in that: The first preset value is 3500 to 4500, and the second preset value is 45000 to 55000.

9. The text difficulty grading evaluation method according to claim 1, wherein: Obtaining the predicted difficulty level of the to-be-graded text using the Chinese character weighted scores and the word segmentation weighted scores includes: normalizing the Chinese character weighted scores and the word segmentation weighted scores according to preset difficulty levels to obtain a first initial predicted difficulty level corresponding to the Chinese character weighted scores and a second initial predicted difficulty level corresponding to the word segmentation weighted scores; The first initial predicted difficulty level and the second initial predicted difficulty level are weighted, and the weighted result is rounded to obtain the predicted difficulty level of the text to be graded.

10. The text difficulty grading evaluation method according to claim 9, characterized in that: The normalization process of the weighted score of the Chinese character includes: performing a first initial normalization using a first normalization model, wherein the first normalization model is Among them, L char is used to represent the result of the first initial normalization, x is used to represent the weighted score of the Chinese character, A char Used to indicate the minimum value of the reference value of the weighted score of Chinese characters, B char Used to indicate the maximum value of the reference value of the weighted score of Chinese characters, N min It is used to indicate the lowest level of the preset difficulty level, N max Used to indicate the highest level of the preset difficulty levels.

11. The text difficulty grading evaluation method according to claim 9, wherein: The normalization process of the word segmentation weighted score includes: performing a second initial normalization using a second normalization model, wherein the second normalization model is Among them, L word is used to represent the result of the second initial normalization, y is used to represent the weighted score of the word segmentation, and A word Used to represent the minimum value of the reference value of the word segmentation weighted score, B word Used to indicate the maximum value of the reference value of the word segmentation weighted score, N min It is used to indicate the lowest level of the preset difficulty level, N max Used to indicate the highest level of the preset difficulty levels.

12. The text difficulty grading evaluation method according to any one of claims 1 to 11, wherein: Adjusting the prediction difficulty level includes: Calculating an average of the character difficulty values ​​corresponding to the first number of characters with the largest character difficulty values ​​in the feature set of the text to be classified as the average character difficulty value, and calculating an average of the word difficulty values ​​corresponding to the first second number of words with the largest word difficulty values ​​in the feature set of the text to be classified as the average word difficulty value, wherein the first number and the second number are both integers greater than 1; The threshold condition includes a first threshold condition, wherein the first threshold condition includes: the predicted difficulty level is less than a first preset difficulty prediction range value, the average difficulty value of the characters is greater than or equal to the first preset threshold value, and the average difficulty value of the words is greater than or equal to a second preset threshold value; when the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the first threshold condition, the predicted difficulty level is increased by the first preset difficulty level value, otherwise the predicted difficulty level is maintained; Alternatively, the threshold condition includes a second threshold condition, which includes: the predicted difficulty level is greater than a second preset difficulty prediction range value, and the average difficulty value of the characters is less than or equal to a third preset threshold value, and the average difficulty value of the words is less than or equal to a fourth preset threshold value; when the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the second threshold condition, the predicted difficulty level is subtracted from the second preset difficulty level value, otherwise the predicted difficulty level is maintained.

13. The text difficulty grading evaluation method according to claim 12, wherein: The first number is any integer from 3 to 15, and the second number is any integer from 2 to 10.

14. The text difficulty grading evaluation method according to claim 12, wherein: In the step of preprocessing the text to be classified, the feature types in the feature set further include sentences; Obtaining the difficulty evaluation value corresponding to each feature further includes: obtaining the number of sentences and the average length of sentences in the text to be graded; The threshold condition includes a third threshold condition, wherein the third threshold condition includes: the number of sentences is greater than or equal to a fifth preset threshold, the average length of the sentences is greater than or equal to a sixth preset threshold, and the predicted difficulty level is less than a third preset difficulty level value; when the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the third threshold condition, the predicted difficulty level is set to the third preset difficulty level value; Alternatively, the threshold condition includes a fourth threshold condition, and the fourth threshold condition also includes: the number of sentences is less than or equal to a seventh preset threshold, and the average length of the sentences is less than or equal to an eighth preset threshold, and the predicted difficulty level is greater than a fourth preset difficulty level value; when the predicted difficulty level and the difficulty assessment value corresponding to each feature in the feature set of the text to be graded meet the fourth threshold condition, the predicted difficulty level is set to the fourth preset difficulty level value.

15. A text difficulty grading assessment device, characterized in that: include: A text acquisition module is used to obtain the text to be graded; A preprocessing module, configured to preprocess the text to be classified to obtain a feature set of the text to be classified, wherein the feature set includes a plurality of features related to text granularity, and the feature types in the feature set include at least characters and words; A difficulty evaluation value acquisition module is used to obtain a difficulty evaluation value corresponding to each feature in the feature set of the text to be graded, wherein the difficulty evaluation value includes a character difficulty value corresponding to each character and a word difficulty value corresponding to each word, wherein the character difficulty value is negatively correlated with the character frequency, and the word difficulty value is negatively correlated with the word frequency; A ranking score acquisition module is used to obtain the sum of the character difficulty values ​​of all characters in the feature set of the text to be graded as a weighted Chinese character score, and obtain the sum of the word difficulty values ​​of all words in the feature set of the text to be graded as a weighted word segmentation score; A difficulty level prediction module, configured to obtain a predicted difficulty level of the text to be graded using the weighted scores of the Chinese characters and the weighted scores of the word segments; The difficulty level adjustment module is used to judge the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded, and adjust the predicted difficulty level when the predicted difficulty level and the difficulty evaluation value corresponding to each feature in the feature set of the text to be graded meet a threshold condition.

16. A device, characterized in that The method comprises at least one memory and at least one processor, wherein the memory stores one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the text difficulty grading assessment method according to any one of claims 1 to 14.

17. A storage medium, characterized in that: The storage medium stores one or more computer instructions, and the one or more computer instructions are used to implement the text difficulty grading assessment method according to any one of claims 1 to 14.