A Chinese Text Readability Evaluation Method Based on Information Gain and Hierarchical Classification
By using information gain and hierarchical classification methods in text readability evaluation, a language feature representation module, a depth representation module guided by information gain and hierarchical classification module was established, which solved the problem that the existing technology failed to fully explore the depth characteristics of the article and ignored the hierarchical sequence relationship, and achieved a more accurate and fine-grained text readability evaluation.
Patent Information
- Application Number
- CN202411787057.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-06
AI Technical Summary
The existing technology has failed to fully explore the deep characteristics of the article and ignores the hierarchical sequence relationships of the difficulty level of text readability, resulting in the difficulty difference relationship between different categories not being fully reflected.
A readability evaluation model is established based on information gain and hierarchical classification methods, including language feature representation module, information gain-guided depth representation module and hierarchical classification module. The information gain principle combines the representation of multiple encoding layers, captures hierarchical structure information, and learns the relationship between different readability difficulty categories.
Fully explore the depth representation of the text, accurately capture the relationship between different readability difficulty categories, and improve the accuracy and fine-grainedness of text readability evaluation.
Smart Images

Figure CN119293238B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a Chinese text readability evaluation method based on information gain and hierarchical classification. Background Art
[0002] Text readability refers to the ease with which the content of a text can be understood and absorbed by readers. Automatic Readability Assessment (ARA) is to quantitatively and automatically evaluate and analyze the difficulty of a text, which plays a crucial role in many fields such as education, publishing, information retrieval, and natural language processing.
[0003] In early readability research, a direct connection between readability and language features was established by using readability formulas. Subsequently, researchers used machine learning methods to automatically learn the relationship between language features and readability. However, the evaluation method based on language features can only simply utilize the shallow features of the text and cannot capture more valuable semantic information, resulting in low evaluation accuracy. In contrast, pre-trained language models represented by BERT, with their powerful semantic modeling capabilities, have far exceeded previous model methods in many natural language processing tasks and also shown great advantages in the readability evaluation task. However, deep models also have some problems, such as poor interpretability and often ignoring the effective utilization of text language features.
[0004] In recent years, automatic readability evaluation research has gradually shifted to a method that combines manually extracted language features and deep features extracted by deep neural networks, namely the Hybrid Automatic Readability Assessment (HARA). However, most current models still have the following problems: First, when current pre-trained language models extract text features, they focus on semantic-level features and ignore other-level features such as vocabulary and syntax, and fail to fully explore the deep features of the article. Second, most current models regard the text readability evaluation task as a classification task with independent categories and ignore the sequential relationship of difficulty levels, resulting in the inability to fully reflect the difficulty difference relationship between different categories. Summary of the Invention
[0005] The purpose of the present invention is to provide a Chinese text readability evaluation method based on information gain and hierarchical classification, which is used to solve the technical problems in the prior art that the deep features of the article cannot be fully explored and the difficulty difference relationship between different categories cannot be fully reflected due to ignoring the hierarchical sequence relationship of readability difficulty levels.
[0006] The described Chinese text readability evaluation method based on information gain and hierarchical classification includes: This method establishes a readability evaluation model based on information gain and hierarchical classification, including a language feature representation module, an information gain-guided depth representation module, and a hierarchical classification module; among them, the language feature representation module is used to extract language features and map these features to the deep semantic space in the form of natural language descriptions to obtain corresponding language feature representations; the information gain-guided depth representation module uses the encoded output layer containing semantic information as the anchor layer, and based on the information gain principle, according to the difference in information contained in different encoding layers, fuses the representations of multiple encoding layers to fully mine the depth representation of the text; the hierarchical classification module is modeled as a hierarchical tree structure, and by capturing the hierarchical structure information, the mutual relationship between different readability difficulty categories is learned.
[0007] Preferably, in the information gain-guided depth representation module, first divide the text into several paragraph blocks, and then input them into the pre-trained model together to generate their respective encoded representations. Use the inter-block residual Transformer module to fuse the encoded representations of these paragraph blocks to obtain the feature representations of all paragraph blocks in the encoding layer; the depth representation module selects to perform a separate classification evaluation task on the last encoding layer and calculates the information entropy of this layer alone; successively splice the depth representations of the previous encoding layers with the depth representation of the last encoding layer, perform separate classification evaluation tasks respectively, and calculate the conditional entropy obtained by fusing each encoding layer with the last encoding layer. Use the information entropy and conditional entropy as the loss function of the feature extraction loss; this method makes all gain coefficients form a gain coefficient vector, and the adaptive weights of each layer form a weight vector; calculate the KL divergence between the gain coefficient vector and the weight vector as the guiding loss, and the information gain-guided depth representation module is jointly trained by the feature extraction loss and the guiding loss.
[0008] Preferably, i the feature representations of all paragraph blocks in the X i layer encoding layer are expressed as follows: X i = IsR-Transformer ( s i1 , s i2 ,⋯, s in ) where, for n paragraph blocks s 1, s 2,⋯, s n , each encoding layer of the pre-trained model can generate its own encoded representations for these n paragraph blocks s i1 , si2 , ⋯, s in , IsR-Transformer represents that the inter-block residual Transformer module fuses the encoded representations of these passage blocks; the information entropy is expressed as follows:
[0009] ,
[0010] wherein, P i represents the probability of being classified as the i th class, C represents the number of classes, y i represents the true label, H(X L ) represents the information entropy of the L th encoding layer.
[0011] The conditional entropy is expressed as follows:
[0012] ,
[0013] wherein, P i represents the probability of being classified as the L th class after classification evaluation after concatenating with the i th encoding layer, C represents the number of classes, y i represents the true label, H ( X L │ X i ) represents the conditional entropy obtained by fusing the i th encoding layer and the L th encoding layer.
[0014] The loss function of the depth representation module is expressed as follows:
[0015] ,
[0016] wherein, represents the loss of feature extraction for the depth representation.
[0017] Preferably, calculate the information gain generated by each layer: IG ( X L , X i ) = H ( X L ) - H (X L │ X i ),in, IG ( X L , X i ) indicates the L The information gain generated by the fusion of the feature information of the i-th coding layer by the i-th coding layer; the calculated gain coefficient is used to represent the weight of guiding the fusion of the i-th coding layer of the pre-trained language model, which is expressed as:
[0018] ,
[0019] in, IG ( X L , X j ) indicates the L Layer coding layer fusion j The information gain generated after the feature information of the layer encoding layer is fused j Layer encoding layer from the previous L -1 encoding layer to choose from, τ represents the temperature coefficient, represents the first i The weights when layer encoding layers are fused.
[0020] forward L When the feature representation of the -1 layer is fused, weighted fusion is used to obtain the fused representation, and then this fused representation is combined with the L The depth representation of the layer is concatenated and the depth representation is calculated by linear transformation. The specific representation is as follows:
[0021] ,
[0022] ,
[0023] in, X i For the i The feature representation of the layer encoding layer, β i For the i The weights of the feature representation of the layer encoding layer are learnable adaptive weights; X fused For the front L The fused representation obtained by weighted fusion of the depth representation of -1 layer, [ X L , X fused ] is the fusion representation with the L The depth of the layer stitching,W and b are the weight matrix and offset of the linear transformation respectively, f deep is the finally obtained depth representation.
[0024] Preferably, let all gain coefficients form a gain coefficient vector α , and the adaptive weights of each layer β i form a weight vector β ; calculate the divergence between the gain coefficient vector α and the weight vector β as the guiding loss, and the calculation process is as follows: KL ,
[0025] ,
[0026] wherein, represents the guiding loss, represents the calculation of the divergence between the gain coefficient vector α and the weight vector β . KL The total loss function of the depth representation module guided by information gain can be expressed as follows:
[0027] ,
[0028] ,
[0029] wherein, is the guiding loss, is the feature extraction loss, is the total loss of the depth representation module guided by information gain.
[0030] Preferably, project the obtained language feature representation into the depth feature space by means of an orthogonal projection vector, and the new language feature representation obtained is as follows:
[0031] ,
[0032] wherein, f' Ling is the new language feature representation, f Ling is the language feature representation output by the language feature representation module; finally, the new language feature representation f' Ling and the depth feature representation f deep are concatenated to obtain the final text representation f all .
[0033] Preferably, in the hierarchical classification module, the sequence relationship based on the readability difficulty category is modeled as a hierarchical tree, and each level is regarded as a difficulty perception block. In the difficulty perception block, this method realizes the perception of difficulty through classification operations. Through the operations of multiple difficulty perception blocks, the difficulty perception granularity can be refined from the high-level nodes to the low-level nodes in sequence; when a classification error occurs, different penalties are given according to the distance between the error result and the target difficulty level, so as to reflect the difficulty relationship between different categories. A hierarchical constraint penalty function is designed as the loss function to measure the impact on its subsequent nodes if the node is classified incorrectly. If a prediction error occurs in the coarse-grained classification at the high level, a greater penalty will be imposed.
[0034] Preferably, the difficulty perception block includes a feature transformation block, a classifier, and a residual fusion gate. This method first inputs the text representation f all into the feature transformation block for feature transformation operations, which is expressed as follows:
[0035] ,
[0036] where, σ 1 is the Sigmoid activation function, σ 2 is the GELU activation function, W 1 and b 1 represent the weight matrix and offset of the first linear transformation in the feature transformation block, W 2 and b 2 represent the weight matrix and offset of the second linear transformation in the feature transformation block, and is the text representation after transformation.
[0037] In order to enable the model to obtain a preliminary difficulty perception, the classifier performs the first coarse-grained classification operation, which is expressed as follows: P i = classifier i ( f all ), P i is the probability distribution of the i th readability difficulty level under the coarse-grained level, classifier i represents the classification of the text features into the i th readability difficulty level; in order to alleviate the vanishing gradient, the hierarchical classification module uses a residual fusion gate to enhance the feature representation, which is specifically expressed as follows:
[0038] ,
[0039] ,
[0040] where,f' all , f all represents the result of concatenating the transformed text representation and the previous text representation, W and b are the weight matrix and offset of the corresponding linear transformation, σ 1 is the Sigmoid activation function, G represents the new features obtained after linear transformation and activation of the result of concatenating text representations; ⊙ represents the dot product operation, f'' all is the new feature and the transformed text representation f' all and the previous text representation f all after being fused, that is, the enhanced features output by the residual fusion gate.
[0041] Preferably, the hierarchical constraint penalty function is expressed as follows:
[0042] ,
[0043] wherein, n is the height of the hierarchical tree, , represents the influence factor, where N i represents the number of all descendant nodes of the i th node in the current layer, n is the total number of corresponding nodes, p i represents the predicted probability value, y i represents the true label.
[0044] Preferably, the final loss function of this method is expressed as:
[0045] ,
[0046] wherein, Loss is the total loss of the model in this method, λ is a hyperparameter that balances the losses of the depth representation module guided by information gain and the hierarchical classification module.
[0047] The present invention has the following advantages: By using the depth representation module guided by information gain, this method takes the encoding output layer containing semantic information as the anchor layer, and utilizes the principle of information gain to fuse the representations of multiple encoding layers according to the difference in information contained in different encoding layers to fully explore the depth representation of the text. The hierarchical classification module is modeled as a hierarchical tree structure, and by capturing the hierarchical structure information, the mutual relationship between different readability difficulty categories can be learned.
[0048] This method uses an inter-block residual Transformer module to fuse these paragraph block features, enrich the encoding layer of semantic representation, eliminate the uncertainty of information, and thereby represent the information gain degree of this information. In order to make the adaptive weights gradually converge to the actual importance of the features of each encoding layer, this method uses the KL divergence (information divergence) between the gain coefficient vector and the weight vector as the guiding loss. On the other hand, in order to enable these encoding layers to extract depth information more fully, the depth representation module uses the information entropy and conditional entropy calculated above as the loss function to train the further feature extraction ability of each encoding layer.
[0049] In the difficulty perception block, this method realizes the perception of difficulty through classification operations. Since when modeling as a hierarchical tree, the difficulty differences in high-level nodes are the largest, and the model can more easily distinguish the difficulty relationships between them. Through the operations of multiple difficulty perception blocks, the difficulty perception granularity can be refined from high-level nodes to low-level nodes in sequence. Therefore, when performing the final evaluation, this model can obtain the difficulty level of the text more accurately.
[0050] The language feature representation module adopted by this method extracts language features and maps these features to the deep semantic space in the form of natural language descriptions to solve the possible semantic space misalignment problem between traditional numerical language features and deep features. Brief Description of the Drawings
[0051] Figure 1 It is a flowchart of the Chinese text readability evaluation method based on information gain and hierarchical classification of the present invention.
[0052] Figure 2 It is a schematic diagram of the hierarchical tree formed by the hierarchical classification module modeling in the present invention. Detailed Embodiments
[0053] The following is a more detailed description of the specific embodiments of the present invention with reference to the accompanying drawings through the description of the embodiments, so as to help those skilled in the art have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.
[0054] As Figure 1 - Figure 2As shown in the figure, the present invention provides a Chinese text readability evaluation method based on information gain and hierarchical classification, including: This method establishes a new readability evaluation model, which is a readability evaluation model based on information gain and hierarchical classification (IGHC-ARA, where IG represents information gain, that is, information gain, and HC represents hierarchical classification, that is, hierarchical classification). The new readability evaluation model (IGHC-ARA) includes a language feature representation module, an information gain-guided depth representation module, and a hierarchical classification module. Among them, the language feature representation module is used to extract language features and map these features to the deep semantic space in the form of natural language descriptions to solve the possible semantic space misalignment problem between traditional numerical language features and deep features. The information gain-guided depth representation module uses the encoded output layer containing semantic information as the anchor layer, and utilizes the information gain principle to fuse the representations of multiple encoding layers according to the difference in information contained in different encoding layers to fully mine the deep representation of the text. The hierarchical classification module is modeled as a hierarchical tree structure, and by capturing the hierarchical structure information, it learns the mutual relationship between different readability difficulty categories.
[0055] The language feature representation module includes: a language feature extractor and a language feature interpreter. The language feature extractor is used to extract corresponding language features from the text, and the language feature interpreter maps the extracted language features to the deep semantic space in the form of natural language descriptions to solve the possible semantic space misalignment problem between traditional numerical language features and deep features.
[0056] For the language feature representation module, this method first selects a series of representative language features as inputs, and these language features include complex characters and vocabulary. The specific language features are shown in Table 1.
[0057] Table 1: Language features for the language feature interpreter.
[0058]
[0059] This method sets that the language features in the text include language features l 1, l 2, ⋯, l m , where m represents the total number of language features. This method obtains the corresponding language features based on statistical methods, and uses "the number of [feature name] in the text is [value]" as the interpretation template for language features, and wraps all language features with the interpretation template. The obtained result is denoted as T 1, T 2, ⋯, T mThe obtained result is input into the language feature interpreter to obtain the corresponding language feature representation. The calculation formula is as follows: f Ling = Interpretor ( T 1 ,T 2 ,⋯,T m ), where Interpretor represents processing the language features within the parentheses by the language feature interpreter, f Ling represents the language feature representation output by this module.
[0060] Information gain-guided depth representation module: According to the differences in the encoded information of different encoding layers of the pre-trained language model, capture depth features at different levels, and finally fuse the depth features at different levels using the information gain principle.
[0061] For the information gain-guided depth representation module, this method first divides the text into n paragraph blocks s 1, s 2, ⋯, s n , and then inputs them into the pre-trained model together. Each encoding layer of the pre-trained model can generate its own encoded representation for these n paragraph blocks s i1 , s i2 , ⋯, s in , where i = 1, ⋯, L , L is the number of encoding layers. The depth representations obtained by different encoding layers contain feature representations with different meanings. To better perform information interaction between the paragraph blocks, this method uses an Inter-section R-Transformer (IsR-Transformer) module to fuse the encoded representations of these paragraph blocks to obtain the feature representations of all paragraph blocks in the i th encoding layer X i , which is expressed as follows:
[0062] X i = IsR-Transformer ( s i1 , s i2 , ⋯, s in ),
[0063] To enrich the encoding layer of semantic representation, the deep representation module calculates the information gain degree of the information by fusing the results of the previous L -layer encoding layer in sequence based on the principle of information gain to eliminate the uncertainty of the information. L -1 encoding layers.
[0064] Since the last encoding layer of the pre-trained model contains the richest semantic information, the deep representation module selects to perform a separate classification evaluation task on the feature representation of the L -layer encoding layer and calculates the information entropy of this layer alone. The information entropy is expressed as follows:
[0065] ,
[0066] where P i represents the probability of being classified as the i -th class, C represents the number of classes, y i represents the true label, H(X L ) represents the L -layer encoding layer.
[0067] After that, to analyze how the information in the previous L -1 encoding layers can most effectively complement the L -layer encoding layer, the deep representation module concatenates the deep representations of the previous L -1 encoding layers with the deep representation of the L -layer encoding layer in sequence, performs a classification evaluation task for each, and calculates the conditional entropy obtained by fusing each encoding layer with the L -layer encoding layer. The conditional entropy is expressed as follows:
[0068] ,
[0069] where P i represents the probability of being classified as the L -th class after concatenating with the i -layer encoding layer and performing a classification evaluation, C represents the number of classes, y i represents the true label, H ( X L │ X i ) represents the conditional entropy obtained by fusing the i -layer encoding layer with the L -layer encoding layer.
[0070] To enable these encoding layers to extract depth information more fully, the depth representation module uses the information entropy and conditional entropy calculated above as the loss function to train the further feature extraction ability of each encoding layer. This loss function is expressed as follows:
[0071] ,
[0072] where represents the loss of feature extraction for depth representation.
[0073] Meanwhile, this method uses the information entropy generated by separately classifying and evaluating the L -th layer encoding layer, subtracts the conditional entropy obtained after fusing the feature information of the i -th layer encoding layer, and obtains the information gain generated by each layer:
[0074] IG ( X L , X i ) = H ( X L ) - H ( X L │ X i ),
[0075] where IG ( X L , X i ) represents the information gain generated by the L -th layer encoding layer after fusing the feature information of the i -th layer encoding layer.
[0076] To highlight important encoding layer features for comparison, this method calculates a gain coefficient, which is used to represent the weight when fusing the i -th layer encoding layer that guides the pre-trained language model, and is expressed as:
[0077] ,
[0078] where IG ( X L , X j ) represents the information gain generated by the L -th layer encoding layer after fusing the feature information of the j -th layer encoding layer. The j -th layer encoding layer to be fused is selected from the previous L - 1 encoding layers.τ Represents the temperature coefficient, Represents the weight when the i -th layer encoding layer of the pre-trained language model is fused.
[0079] When actually fusing the feature representations of the first L -1 layers, this method first uses a common weighted fusion to obtain a fused representation, and then concatenates this fused representation with the depth representation of the L -th layer, and calculates the depth representation through a linear transformation, which is specifically expressed as follows:
[0080] ,
[0081] ,
[0082] where, X i is the feature representation of the i -th layer encoding layer, β i is the weight of the feature representation of the i -th layer encoding layer, and is a learnable adaptive weight; X fused is the fused representation obtained by weighted fusion of the depth representations of the first L -1 layers, X L , X fused is the concatenation of the fused representation and the depth of the L -th layer, W and b are respectively the weight matrix and offset of the said linear transformation, f deep is the finally obtained depth representation.
[0083] Finally, in order to make the adaptive weights gradually converge to the actual importance of the feature of each encoding layer, this method makes all gain coefficients form a gain coefficient vector α , and the adaptive weights β i of each layer form a weight vector β ; calculates the α divergence (information divergence) between the gain coefficient vector β and the weight vector KL as the guiding loss, and the calculation process is expressed as follows:
[0084] ,
[0085] where, represents the guiding loss, represents calculating the gain coefficient vectorα with the weight vector β between KL divergence
[0086] This module is jointly trained by the feature extraction loss and the guiding loss. Therefore, the total loss function of the information gain-guided deep representation module can be expressed as follows:
[0087] ,
[0088] where is the guiding loss, is the feature extraction loss, is the total loss of the information gain-guided deep representation module
[0089] Fusion of language feature representation and deep representation: Next, the language feature representation and the deep representation are fused to obtain the final text representation. The specific method of this step is to project the obtained language feature representation into the deep feature space by using an orthogonal projection vector to remove redundant information. The new language feature representation obtained is as follows:
[0090] ,
[0091] where f' Ling is the new language feature representation. Finally, the new language feature representation f' Ling and the deep feature representation f deep are concatenated to obtain the final text representation f all .
[0092] Hierarchical classification module: It is used to implement the hierarchical text classification task, so as to capture the hierarchical structure information and enable the model to better learn the mutual relationship between categories. The hierarchical classification module models the sequence relationship of readability difficulty categories as a hierarchical tree to better capture the difficulty differences between different categories, as Figure 2 shown
[0093] For the described hierarchical tree, this method takes each level as a Difficulty Aware Block (DAB for short). The specific structure includes a Feature Transform Block (FTB for short), a classifier, and a residual fusion gate. Through the difficulty aware block, the perception of different granularity difficulties can be gradually realized
[0094] In the difficulty aware block, in order to obtain a better feature representation, this method first represents the text f allIt is input into the Feature Transformation Block (FTB) for feature transformation operations, which is expressed as follows:
[0095] ,
[0096] where, σ 1 is the Sigmoid activation function, σ 2 is the GELU activation function, W 1 and b 1 represent the weight matrix and offset of the first linear transformation in the feature transformation block, W 2 and b 2 represent the weight matrix and offset of the second linear transformation in the feature transformation block, which is the text representation after transformation. In order to enable the model to obtain a preliminary difficulty perception, the classifier performs the first coarse-grained classification operation, which is expressed as follows: P i = classifier i ( f all ), P i is the probability distribution of the i th readability difficulty level at the coarse-grained level. classifier i represents the classification of the text features into the i th readability difficulty level. In order to alleviate the vanishing gradient, the hierarchical classification module uses a residual fusion gate to enhance the feature representation, which is specifically expressed as follows:
[0097] ,
[0098] ,
[0099] where, f' all , f all represents the result of concatenating the text representation after transformation and the previous text representation, W and b are the weight matrix and offset of the corresponding linear transformation, σ 1 is the Sigmoid activation function, G represents the new feature obtained after the result of concatenating the text representations undergoes linear transformation and activation; ⊙ represents the dot product operation, f'' all is the result of fusing the new feature with the text representation after transformation f' all and the previous text representation f all , that is, the enhanced feature output by the residual fusion gate.
[0100] In the difficulty perception block, the method realizes the perception of difficulty through classification operations. Since the difficulty differences in high-level nodes are the largest when modeled as a hierarchical tree, the model can more easily distinguish the difficulty relationships between them. Through the operations of multiple difficulty perception blocks, the difficulty perception granularity can be refined from high-level nodes to low-level nodes in sequence. Therefore, when conducting the final evaluation, this model can obtain the difficulty level of the text more accurately.
[0101] To distinguish the "closeness and distance" in prediction errors, that is, when a classification error occurs, different penalties are given according to the distance between the error result and the target difficulty level, so as to reflect the difficulty relationships between different categories. A hierarchical constraint penalty function is also designed as the loss function in the hierarchical classification module, which is expressed as follows:
[0102] ,
[0103] Among them, n is the height of the hierarchical tree, , represents the influence factor, where N i represents the i th node in the current layer, and n is the total number of corresponding nodes, p i represents the predicted probability value, y i represents the true label. The hierarchical classification module uses the hierarchical constraint penalty function to measure the impact on its subsequent nodes if the classification of this node is incorrect. If the model makes a prediction error in the high-level coarse-grained classification, it will be subject to a greater penalty.
[0104] When training the model, this method combines the loss of the information gain-guided depth representation module and the loss of the hierarchical classification module, and uses it as the total loss function to train the entire model. The final loss function is expressed as:
[0105] ,
[0106] Among them, Loss is the total loss of the model in this method, λ is a hyperparameter that balances the losses of the information gain-guided depth representation module and the hierarchical classification module.
[0107] The above has described the present invention by way of example in conjunction with the accompanying drawings. Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various non-substantive improvements are made by adopting the inventive concept and technical solutions of the present invention, or the inventive concept and technical solutions of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. A Chinese text readability evaluation method based on information gain and hierarchical classification, characterized by: include: This method establishes a readability evaluation model based on information gain and hierarchical classification, including a language feature representation module, a deep representation module guided by information gain, and a hierarchical classification module; wherein the language feature representation module is used to extract language features and map these features to a deep semantic space in a natural language description manner to obtain corresponding language feature representations; the deep representation module guided by information gain uses the encoding output layer containing semantic information as the anchor layer, and uses the information gain principle to fuse the representations of multiple encoding layers according to the differences in information contained in different encoding layers to fully mine the deep representation of the text; the hierarchical classification module is modeled as a hierarchical tree structure, and by capturing the hierarchical structure information, the relationship between different readability difficulty categories can be learned; In the deep representation module guided by information gain, the text is first divided into several paragraph blocks, which are then input into the pre-trained model together to generate their own encoding representations. The encoding representations of these paragraph blocks are fused using the inter-block residual Transformer module to obtain the feature representations of all paragraph blocks in the encoding layer. The deep representation module selects the last encoding layer for fusion. Perform a separate classification evaluation task and calculate the separate information entropy of the layer; splice the depth representations of the previous coding layers with the depth representation of the last coding layer in turn, perform classification evaluation tasks separately, calculate the conditional entropy obtained by fusion of each coding layer with the last coding layer, and use the information entropy and conditional entropy as the loss function of feature extraction loss; this method makes all gain coefficients form a gain coefficient vector, and the adaptive weights of each layer form a weight vector; Calculate the relationship between the gain coefficient vector and the weight vector KL Divergence is used as the guided loss, and the information gain guided deep representation module is trained jointly by the feature extraction loss and the guided loss.
2. The Chinese text readability evaluation method based on information gain and hierarchical classification according to claim 1 is characterized by: X i For the i The feature representation of the layer encoding layer is expressed as follows: X i = IsR-Transformer ([ s i1 , s i2 ,⋯, s in ]), where for n paragraph blocks s 1, s 2,⋯, s n , each encoding layer of the pre-trained model can generate its own encoding representation for these n paragraph blocks s i1 , s i2 ,⋯, s in , IsR-Transformer The inter-block residual Transformer module fuses the encoded representations of these paragraph blocks; the information entropy is expressed as follows: , in, P i The features of the Lth coding layer are classified as i The probability of the class, C Indicates the number of categories, y i represents the true label, H(X L ) Indicates L Information entropy of the layer coding layer; The conditional entropy is expressed as follows: , in, q i Indicates i The coding layer and L The features fused by the layer coding layer are classified into i The probability of the class, C Indicates the number of categories, y i represents the true label, H ( X L │ X i ) indicates the i The coding layer and L The conditional entropy obtained by fusion of layer coding layers; The loss function of the deep representation module is expressed as follows: , in, represents the loss for feature extraction on deep representation.
3. The Chinese text readability evaluation method based on information gain and hierarchical classification according to claim 2 is characterized by: Calculate the information gain generated by each layer: IG ( X L , X i )= H ( X L )- H ( X L │ X i ),in, IG ( X L , X i ) indicates the L The information gain generated by the fusion of the feature information of the i-th coding layer by the i-th coding layer; the calculated gain coefficient is used to represent the weight of guiding the fusion of the i-th coding layer of the pre-trained language model, which is expressed as: , in, IG ( X L , X j ) indicates the L Layer coding layer fusion j The information gain generated after the feature information of the layer encoding layer is fused j Layer encoding layer from the previous L -1 encoding layer to choose from, τ represents the temperature coefficient, represents the first i The weights when layer coding layers are fused; forward L When the feature representation of the -1 layer is fused, weighted fusion is used to obtain the fused representation, and then this fused representation is combined with the L The depth representation of the layer is concatenated and the depth representation is calculated by linear transformation. The specific representation is as follows: , , in, X i For the i The feature representation of the layer encoding layer, β i For the i The weights of the feature representation of the layer encoding layer are learnable adaptive weights; X fused For the front L The fused representation obtained by weighted fusion of the depth representation of -1 layer, [ X L , X fused ] is the fusion representation with the L The depth of the layer stitching, W and b are the weight matrix and offset of the linear transformation respectively, f deep is the final depth representation.
4. The Chinese text readability evaluation method based on information gain and hierarchical classification according to claim 3 is characterized by: Let all gain coefficients Composition gain coefficient vector α , and the adaptive weights of each layer β i Composition weight vector β ; Calculate the gain coefficient vector α With the weight vector β Between KL Divergence is used as the guided loss, and the calculation process is expressed as follows: , in, represents the guidance loss, Represents the calculation gain coefficient vector α With the weight vector β Between KL Divergence; The total loss function of the information gain-guided deep representation module can be expressed as follows: , in, To guide the loss, is the feature extraction loss, The depth guided by information gain represents the total loss of the module.
5. The Chinese text readability evaluation method based on information gain and hierarchical classification according to claim 4 is characterized by: The obtained language feature representation is projected into the deep feature space by using an orthogonal projection vector, and the new language feature representation obtained is as follows: , in, f' Ling For new language features, f Ling is the language feature representation output by the language feature representation module; finally, the new language feature representation f' Ling and deep feature representation f deep Splice them together to get the final text representation f all .
6. The Chinese text readability evaluation method based on information gain and hierarchical classification according to claim 5 is characterized by: In the hierarchical classification module, the sequence relationship based on readability difficulty categories is modeled as a hierarchical tree, and each level is regarded as a difficulty perception block. In the difficulty perception block, this method realizes the perception of difficulty through classification operations. Through the operation of multiple difficulty perception blocks, the difficulty perception granularity can be refined from high-level nodes to low-level nodes in sequence; when a classification error occurs, different penalties are given according to the distance between the error result and the target difficulty level, thereby reflecting the difficulty relationship between different categories. A hierarchical constraint penalty function is designed as a loss function to measure the impact of a node misclassified on its subsequent nodes. If the prediction is wrong in the case of high-level coarse-grained classification, it will be punished more severely.
7. The Chinese text readability evaluation method based on information gain and hierarchical classification according to claim 6 is characterized by: The difficulty perception block includes a feature transformation block, a classifier, and a residual fusion gate. This method first represents the text f all Input to the feature transformation block for feature transformation operation, which is expressed as follows: , in, σ 1 is the Sigmoid activation function, σ 2 is the GELU activation function, W 1 and b 1 represents the weight matrix and offset of the first linear transformation in the feature transformation block, W 2 and b 2 represents the weight matrix and offset of the second linear transformation in the feature transformation block, which is the text representation after the transformation; In order to give the model a preliminary sense of difficulty, the classifier performs the first coarse-grained classification operation, which is expressed as follows: Q j = classifier j ( f all ), Q j For coarse-grained j The probability distribution of readability difficulty levels, classifier j Indicates the first j readability difficulty level classification; in order to alleviate the gradient disappearance, the hierarchical classification module uses a residual fusion gate to enhance the feature representation, which is specifically expressed as follows: , , in,[ f' all , f all ] represents the result of concatenating the transformed text representation with the previous text representation. W and b is the weight matrix and offset of the corresponding linear transformation, σ 1 is the Sigmoid activation function, G represents the new feature obtained after the concatenation of text representations is linearly transformed and activated; ⊙ represents the dot product operation, f'' all Represents the new features and the transformed text f' all And the previous text representation f all The fused result is the enhanced feature output by the residual fusion gate.
8. The Chinese text readability evaluation method based on information gain and hierarchical classification according to claim 7 is characterized by: The hierarchical constraint penalty function is expressed as follows: , in, m is the tree height of the hierarchical tree, , represents the impact factor, where N i Indicates the current layer i The number of all descendant nodes of a node, t is the total number of nodes in the current layer; p i represents the predicted probability value, y i represents the true label.
9. The Chinese text readability evaluation method based on information gain and hierarchical classification according to claim 8 is characterized by: The final loss function of this method is expressed as: , in, Loss is the total loss of the model in this method, λ is a hyperparameter that balances the losses of both the information gain-guided deep representation module and the hierarchical classification module.
Citation Information
Patent Citations
Hybrid readability evaluation method and system based on public and private feature decomposition
CN117313704A