Text readability classification method and system based on large model and comparative learning

By adopting a method based on big model and contrast learning in medical text readability classification, and using multilingual features and contrast text vectors, the problems of insufficient feature capture and real-time requirements in the prior art are solved, and higher classification accuracy and real-time performance are achieved.

CN120144762AInactive Publication Date: 2025-06-13ANHUI MEDICAL UNIV

Patent Information

Application Number
CN202510592946.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture medical-related features such as term density and diagnostic logic complexity of medical texts, resulting in insufficient accuracy in medical text readability classification, and existing large models are difficult to meet real-time requirements in medical text scenarios.

Method used

The text readability classification method based on large model and contrast learning is adopted, and multilingual features are manually extracted, and the original text vector and contrast text vector are generated using two pre-trained models, the similarity score is calculated, the classification matrix is ​​generated, and the model is trained through the loss function to achieve readability level prediction.

Benefits of technology

It significantly enhances the model's discrimination and distinction ability of medical text features, avoids feature confusion, improves the accuracy of medical text readability classification, and meets the real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144762A_ABST
    Figure CN120144762A_ABST
Patent Text Reader

Abstract

The invention discloses a text readability classification method and system based on a large model and comparative learning, and relates to the field of natural language processing. The method comprises the following specific steps: inputting an original text into a BERT encoder to generate a text vector; meanwhile, a plurality of independent BERT encoders are adopted to directly generate comparison text vectors of different grades; original text vectors and comparison text vectors are calculated through cosine similarity, a classification matrix is generated, and model parameters are optimized based on comparison learning. Compared with a traditional method, the method not only can utilize the strong feature extraction capability of a deep learning model, but also can effectively distinguish medical texts with different difficulty levels through a contrast learning strategy, can effectively capture the complex features of the medical texts, remarkably improves the accuracy and interpretability of readability classification in the field, and improves the recognition efficiency. And powerful support is provided for personalized medical health services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and specifically to a text readability classification method and system based on large models and contrastive learning. Background Art

[0002] Text readability classification is to classify texts according to predetermined readability levels, and is widely applied in scenarios such as education, content recommendation, and automatic evaluation. Current readability classification research mostly focuses on general-domain texts and often relies on surface language features or traditional readability formulas. However, medical text readability classification faces the following challenges due to dense professional terms (such as "coronary atherosclerosis"), complex logical structures (such as evidence-based medical reasoning chains), and significant differences in audience cognition (different needs of patients, medical students, and doctors): Traditional readability methods rely on traditional language features such as sentence length and word frequency, and cannot capture medical-related features such as the term density and diagnostic logic complexity of medical texts.

[0003] Most traditional medical readability measurement methods are formula-based or linear regression-based. From the text features extracted by them, the features they encompass are limited measurable surface language features. On the one hand, the traditional language features they rely on cannot capture medical-related features such as the term density and diagnostic logic complexity of medical texts. On the other hand, these features are limited in number and cannot represent the deep language features of texts. From the method perspective, the relationship expressions between the feature variables in the formula method and the linear regression model are relatively single. Therefore, if more comprehensive and accurate text features are to be encompassed, it is not enough to use traditional readability formulas alone.

[0004] The reading difficulty differences in medical texts are often reflected in the implicit knowledge density (such as the evidence level of clinical guidelines) rather than the explicit language structure, resulting in insufficient classification accuracy of traditional models for adjacent difficulty levels.

[0005] In recent years, the Transformer architecture (such as GPT, BERT, etc.) has made breakthrough progress in the field of natural language processing, and researchers have gradually applied it to text readability classification tasks. However, the existing methods face the following challenges in the medical text scenario: Although GPT series models achieve context awareness through the autoregressive generation mechanism (such as GPT-3 contains 175 billion parameters), their huge computational cost results in a long inference time for a single text, making it difficult to meet the real-time requirements of the healthcare scenario (such as online consultation recommendations require a <100ms response).

[0006] Due to the high density of implicit knowledge in medical texts and the more complex cross-level differences, texts at different readability levels show differences in language expression, use of professional vocabulary, etc. If all data is mixed into the same model for training, "confusion" of features may occur. Therefore, although a single BERT model has high inference efficiency, its homogenized feature extraction mechanism makes it difficult to fully distinguish adjacent-level texts (such as "high school level" and "college level"), thus affecting the accuracy of the contrastive learning model. Summary of the Invention

[0007] The object of the present invention is to solve the above existing problems.

[0008] To this end, the technical solution adopted by the present invention is as follows: A text readability classification method based on a large model and contrastive learning, Extract multi-lingual features from the original text manually; Use a first pre-trained model to encode the original text according to N readability levels, extract the last layer [CLS] token vector corresponding to the readability level as the semantic representation, and splice it with the multi-lingual features to generate N original text vectors, Use a second pre-trained model to generate N contrast text vectors for the original text according to N readability levels respectively; Calculate the similarity scores between the original text vectors and the contrast text vectors to generate an N×N dimensional classification matrix; Train the classification matrix based on a loss function, optimize the parameters of the first pre-trained model and the second pre-trained model through gradient descent, and predict the readability level of the text to be predicted based on the trained first pre-trained model and second pre-trained model.

[0009] Further, the prediction process is as follows: Input the text to be predicted into the trained first pre-trained model to generate 1 original text vector, Input the text to be predicted into the trained second pre-trained model to generate N contrast text vectors, Calculate the similarity scores between the original text vector and the N contrast text vectors in sequence, The readability level corresponding to the maximum similarity score is the readability level of the text to be predicted.

[0010] Preferably, the generation method of the contrast text vector is one of the following two methods: 1) Use a second pre-trained model to encode the original text according to N readability levels, extract the last layer [CLS] token vector corresponding to the readability level as the semantic representation, and splice it with the multi-lingual features to generate N contrast text vectors, 2) Use N of the second pre-trained models to encode the original text according to N readability levels respectively, extract the last-layer [CLS] token vectors corresponding to the readability levels as semantic representations, and concatenate them with the multi-lingual features to generate N comparison text vectors; where the parameters of each second pre-trained model are independently initialized and share the same vocabulary.

[0011] Furthermore, the gradient descent uses the Adam gradient optimization algorithm, the first pre-trained model is a BERT encoder, and the second pre-trained model is a Qwen2.5-14B large language model or a BERT encoder.

[0012] Furthermore, the multi-lingual features include numerical language features and proportional language features; Before the concatenation, perform standardized preprocessing on the multi-lingual features; Perform Z-score standardization on the numerical language features and Min-Max standardization on the proportional language features.

[0013] Furthermore, the similarity score is calculated based on cosine similarity. For the original text vector at the i-th readability level and the comparison text vector at the j-th readability level the similarity score is calculated by the formula: , , represents the set of integers; The similarity score is the element in the i-th row and j-th column of the N×N-dimensional classification matrix.

[0014] Furthermore, the diagonal elements of the classification matrix represent the similarity of the original text at the same readability level, and the off-diagonal elements represent the similarity across readability levels. Design the loss function to maximize the diagonal elements of the classification matrix and minimize the off-diagonal elements; The formula of the loss function is as follows: where the parameter τ is set to a value in the interval [0.05, 0.12] to adjust the sharpness of the similarity distribution; Furthermore, the type of the original text includes at least medical texts; The multi-lingual features include at least character complexity features, lexical complexity features, syntactic complexity features, text structure features, and medical features.

[0015] A text readability classification system based on large models and contrastive learning, characterized by including the following modules: Manual extraction module: Extract multilingual features from the original text manually; Original feature generation module: Use the first pre-trained model to encode the original text according to N readability levels, and splice it with the multilingual features to generate N original text vectors. Comparison feature generation module: Use the second pre-trained model to encode the original text according to N readability levels, and splice it with the multilingual features to generate N comparison text vectors, calculate the similarity scores between the text vectors, and generate an N×N-dimensional classification matrix; Level prediction module: Train the classification matrix based on the loss function, optimize the parameters of the N pre-trained language models through gradient descent, and realize the readability level prediction of the original text.

[0016] Compared with the prior art, the advantages of the present invention are as follows: (1) By configuring a pre-trained model for each readability level respectively in the present invention to form parallel independent encoders, each model focuses on learning the features of medical texts at the specified readability level, avoiding the feature confusion problem of traditional single pre-trained models; (2) The present invention constructs a control of "texts at the same level are positive examples, and texts at different levels are negative examples". By using the loss function, the model is forced to learn the feature space where texts at the same level are clustered and texts at different levels are far away, significantly enhancing the discriminability of the model for features and the discrimination ability in adjacent readability levels.

[0017] (3) Traditional text readability methods often rely on traditional explicit language features such as sentence length and word frequency, and cannot capture the medical-related features and implicit semantic features unique to medical texts. This results in insufficient accuracy of traditional text readability methods in the task of automatic grading of medical texts. The present invention constructs an automatic readability grading model suitable for medical texts by extracting and splicing explicit traditional text features, medical text features, and implicit semantic features, etc., and can effectively solve the problem of automatic readability grading of texts in the medical field. Brief description of the drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart of Embodiment 1 of the present invention; Figure 2Confusion matrix for predicting 5 readability levels in Embodiment 1 of the present invention; Figure 3 Flow diagram of Embodiment 2 of the present invention; Figure 4 Confusion matrix for predicting 5 readability levels in Embodiment 2 of the present invention; Figure 5 Confusion matrix for predicting 3 readability levels in Embodiment 3 of the present invention. Detailed implementation manners

[0020] To achieve the above object, the present invention is implemented through the following technical solutions.

[0021] Embodiment 1, the overall process is shown in Figure 1 : S101. Extract multi - language features from the original text manually; The input text in this embodiment is medical text; In this embodiment, the dimension of the multi - language features is 22 - dimensional, including character complexity features, vocabulary complexity features, syntactic complexity features, text structure features, and medical features. As shown in Table 1: Table 1 Manually extracted multi - language feature table In this embodiment, the multi - language features are divided into numerical - type language features and ratio - type language features; Before splicing, pre - process the multi - language features by standardization; For numerical - type language features, perform Z - score standardization, which is expressed as follows: Where, represents the mean value, represents the variance.

[0022] For ratio - type language features, perform Min - Max standardization, which is expressed as follows: Where, represents the minimum value, represents the maximum value.

[0023] S102. Use the first pre - trained model to encode the original text according to 5 readability levels, extract the last - layer [CLS] token vector corresponding to the readability level as the semantic representation, and splice it with the multi - language features to generate 5 original text vectors, The readability level numbers in this embodiment are 1 - 5 respectively.

[0024] The first pre-trained model in this embodiment is a BERT encoder, generating a 768-dimensional semantic representation , concatenate the semantic representation with the multilingual features in Table 1 to generate a 768 + 22-dimensional original text vector, which is expressed as follows: The i-th original text vector in this embodiment is expressed as ; S103. Use the second pre-trained model to encode the original text according to 5 readability levels, extract the last layer [CLS] token vector corresponding to the readability level as the semantic representation, and concatenate it with the multilingual features to generate 5 comparison text vectors, The second pre-trained model in this embodiment is the Qwen2.5 - 14B large language model, generating a 768-dimensional semantic representation, concatenating the semantic representation with the multilingual features to generate a 768 + 22-dimensional comparison text vector, referring to the concatenation formula in step S102 of this embodiment.

[0025] The j-th comparison text vector in this embodiment is expressed as .

[0026] S104. Calculate the similarity score between the original text vector and the comparison text vector to generate an N×N-dimensional classification matrix; In this embodiment, the similarity score is calculated based on cosine similarity. For the original text vector of the i-th readability level and the comparison text vector of the j-th readability level The similarity score The calculation formula is: , , represents the set of integers; The similarity score is the element in the i-th row and j-th column of the 5×5-dimensional classification matrix.

[0027] S105. Train the classification matrix based on the loss function, and optimize the parameters of the first pre-trained model and the second pre-trained model through gradient descent to realize the readability level prediction of the original text.

[0028] The loss function in this embodiment is InfoNCE, and the formula is as follows: The temperature parameter τ is set to a value in the range of 0.05 - 0.12 to adjust the sharpness of the similarity distribution. Under the condition of fixing other model parameters, experiments show that when τ = 0.08, the classification accuracy reaches a peak of 89.2%, and the accuracy fluctuation does not exceed ±1.5% within the range of τ ∈ [0.05, 0.12]. The experimental results are as Figure 2 shown.

[0029] Example 2, the overall process is shown in Figure 3 : S201. Same as S101; S202. Same as S102; S203. Five second pre-trained models are used to encode the original text according to five readability levels respectively, and the last layer [CLS] token vector corresponding to the readability level is extracted as the semantic representation, and it is concatenated with the multi-language features to generate five comparison text vectors; Among them, the parameters of each second pre-trained model are independently initialized and share the same vocabulary.

[0030] The second pre-trained model in this embodiment is a BERT encoder, and an independently initialized BERT encoder (BERT 1 ~BERT 5 ) is assigned to each of the five readability levels. The independent BERT encoders respectively correspond to the feature space mappings of different readability levels. A linear projection layer is connected after the output layer of each encoder to map the features of each level to a unified comparison space. The formula is: where is a learnable parameter matrix, is a bias term.

[0031] In the multi-encoder architecture of this embodiment, each BERT encoder uses independently initialized parameters to ensure that the model can learn the specific representations of different readability levels; at the same time, all encoders are forced to share the same tokenization vocabulary to ensure the tokenization consistency of the input text in different encoders and avoid introducing noise due to tokenization differences. This design, through the cooperation of parameter independence and vocabulary consistency, improves the contrast learning effect while maintaining the semantic space alignment across encoders.

[0032] The j-th comparison text vector in this embodiment is represented as .

[0033] S204. Same as S104; S205. Train the classification matrix based on the loss function, and optimize the parameters of all the BERT encoders through gradient descent to achieve the readability level prediction of the original text. The loss function is the same as that in Embodiment 1. The experimental results are as Figure 4 shown.

[0034] The model hyperparameters of this embodiment are set as follows: (1) Learning rate adaptation strategy: During the training process, if the loss reduction amplitude of 100 consecutive batches is less than , then the learning rate decays by 0.8 times the current value, and the minimum learning rate threshold is . The optimizer uses Adam, and the update formula is: (2) BERT model configuration: The embedding layer size is 768, the hidden layer size is 768, the number of Transformer layers is 12, and the activation function is GELU.

[0035] Embodiment 3: Redivide the 5 readability levels in Embodiment 2 into 3 readability levels. Both the first pre-training model and the second pre-training model are BERT encoders, the number of the second pre-training models is 3, and the numbers of the original text vectors and the comparison text vectors are both 3. The experimental results are as Figure 5 shown.

[0036] Embodiment 4 (Medical text readability level prediction scenario): To accurately predict medical text classification, in this embodiment, a medical popular science text dataset is built by itself. This database includes 8,552 medical popular science related texts, and the data comes from popular science books, medical health related web pages, official accounts and other channels. Perform readability annotation on the texts in this dataset. Referring to the grade division method of the "Chinese Curriculum Standards" and the actual difficulty of the texts, each text is labeled with a difficulty level from 1 to 7: Level 1 corresponds to the difficulty of the first stage of primary school Chinese textbooks (grades 1-2), Level 2 corresponds to the difficulty of the second stage of primary school Chinese textbooks (grades 3-4), Level 3 corresponds to the difficulty of the first stage of primary school Chinese textbooks (grades 5-6), Level 4 corresponds to the difficulty of junior high school Chinese textbooks, Level 5 corresponds to the difficulty of high school Chinese textbooks, Level 6 corresponds to the difficulty of university Chinese textbooks, and Level 7 corresponds to the difficulty of professional papers. Construct a 7-level readability level annotation system according to this rule, and each text is assigned a unique readability level label according to the above rules.

[0037] The 8,552 pieces of readability-graded dataset selected are divided into a training set of 3,080 pieces, a validation set of 385 pieces, and a test set of 385 pieces according to a ratio of 8:1:1. The validation set is used to drive the update and optimization of model parameters. The validation set is used to dynamically adjust hyperparameters. The test set is used to independently evaluate the generalization performance of the model, and its test results are the core empirical basis for the effectiveness of the technical solution.

[0038] In this embodiment, the process of predicting the readability level of medical texts is as follows: Input a medical text to be predicted, which may come from a patient's online question, a medical popular science article, or a professional paper. The text usually contains a certain number of medical terms, proper nouns, and long sentence segments, such as: "The mechanism of myopia involves the synergistic effect of a multi-gene regulatory network and epigenetic modification. Genome-wide association studies (GWAS) have confirmed that the polymorphism at the rs634990 locus of the GJD2 gene on chromosome 15q14 can increase the risk of myopia by 1.89 times (95% CI: 1.32 - 2.71). In terms of environmental factors, continuous blue light exposure (wavelength 450 - 480 nm) activates the melanopsin protein (OPN4) in retinal ganglion cells, inhibits the activity of dopaminergic interneurons, and leads to hypoxic remodeling of the sclera (Ophthalmology, 2023)..." Extract multi-lingual features from the original text manually, as shown in Table 2; (1) Word segmentation and term recognition: First, segment the text according to the word segmentation tool, and call the medical term dictionary and other rule bases to determine whether the text contains specific medical terms (such as disease names, drug names, anatomical terms, etc.), and at the same time preprocess special characters.

[0039] (2) Statistical linguistic features: Extract character features (character type diversity, number of rare characters, proportion of common characters), lexical features (total number of words, TTR, MTLD, proportion of words of different levels, etc.), syntactic features (average sentence length, coefficient of variation of sentence length, adjacent sentence repetition rate, etc.), and overall text structure features (text length, total number of sentences) one by one.

[0040] (3) Statistical medical features: Include the number of medical terms, term density (ratio of the number of terms to the total number of words), term reuse density (proportion of repeated terms), term length index (ratio of the total number of characters of all terms to the number of terms), and positional dispersion (discrete measure of the distribution of terms in the full text).

[0041] Table 2 Manually extracted features and their calculation rules (example text) Unify the multi - language features and medical features obtained from the above statistics for normalization or standardization processing.

[0042] In this embodiment, both the first pre - trained model and the second pre - trained model are BERT encoders. The number of the second pre - trained models is 7, the number of original text vectors is 1, and the number of comparison text vectors is 7.

[0043] It should be noted that the number of original text vectors is different from the training scenario, and it is 7 in the training scenario.

[0044] Load the trained first pre - trained model (BERT encoder), and obtain the last - layer [CLS] token vector of the input text as the semantic representation, for example, [0.18, 0.02,..., 0.41]; Concatenate the processed artificially extracted multi - language features (17 - dimensional) and medical features (5 - dimensional) with the semantic representation extracted by BERT to obtain the original text vector.

[0045] Call the 7 second pre - trained models (BERT encoders) that have been fine - tuned in the training phase. Each model has "learned" to capture features of its respective readability level during the training phase. During inference, these encoders are all in the inference mode.

[0046] Input the text to be predicted into the second pre - trained models (BERT encoders) corresponding to each level respectively. After combining the multi - language features of the text, obtain the corresponding comparison text vectors. In this way, "comparison text vectors" corresponding to levels 1 - 7 one by one are formed.

[0047] Calculate the cosine similarity between the original text vector and the 7 comparison text vectors.

[0048] Calculate the cosine similarity between the original text vector and the comparison text vectors output by BERT for each level one by one in the vector space, forming a similarity list with a length of 7, such as: [0.0611, 0.2150, 0.2130, 0.2029, 0.3109, 0.3509, 0.6592]; Decision rule: Among the above similarity scores, select the level corresponding to the maximum value as the prediction result of the readability level of the text. Taking 7 levels as an example, if the similarity of level 7 (professional level) in the similarity list is the highest, it means that the text has a higher matching degree in the sub - model space of the corresponding professional level, so it is determined that its readability is the most difficult (i.e., more professional); if the best match is level 1 (primary school level), it means that the text is the easiest to read.

[0049] Grading Results: Taking the example of grading into 7 levels, the system will return the readability level (such as "professional level") and the corresponding eigenvalue to the user or downstream module for recommendation, screening, or other scenarios.

[0050] Experimental Comparison: To verify the effectiveness of the present invention, CNN, LSTM, RNN, TextRCNN, MLF-BERT, and LLAMA-BERT are compared with Example 2 and Example 3.

[0051] The evaluation metrics are accuracy, precision, recall, and f1-score. The comparison results of the performance of the 5 readability level classification models are shown in Table 3. Combining Table 2 and Figure 4 、 Figure 5 It can be seen that the method proposed by the present invention is superior to the existing methods in all evaluation metrics, and the cross-level contrast learning is particularly effective in distinguishing adjacent difficulty levels.

[0052] Table 3 Comparison Experimental Results of Each Model By configuring a pre-trained language model for each readability level respectively, the present invention forms parallel independent encoders, enabling each model to focus on learning the features of medical texts at the specified readability level, and avoiding the feature confusion problem of the traditional single BERT encoder on the premise of utilizing the efficiency advantage of BERT; The present invention constructs the first Chinese medical text corpus containing readability level labels, which includes a total of 8,552 medical texts. Compared with the general readability corpus (such as Chinese textbooks), the corpus sources included in this corpus are rich, covering the long-tail data in the medical field, and can effectively make up for the limitations of traditional models in content modeling in the professional field. At the same time, the medical-related features, language features extracted manually are concatenated with the bert semantic features, making this readability classification method more suitable for medical scenarios.

[0053] The present invention constructs a control of "texts at the same level as positive examples, and texts at different levels as negative examples" within the same batch or the same dataset range. By using InfoNCE or similar contrast losses, the model is forced to learn the feature space where texts at the same level (such as "junior and intermediate" texts) are clustered and texts at different levels (such as "junior and intermediate" texts vs. "senior and advanced" texts) are separated, significantly enhancing the discriminability of the model for features and the distinguishing ability in adjacent readability levels.

[0054] As described above, it is only the specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.

Claims

1. A text readability classification method based on large model and contrastive learning, characterized in that: Extracting multi-linguistic features from the original text manually; Using the first pre-trained model, encoding the original text according to N readability levels, extracting the last layer [CLS] tag vector corresponding to the readability level as the semantic representation, and concatenating it with the multi-linguistic features to generate N original text vectors, Using the second pre-trained model, generate N comparison text vectors for the original text according to N readability levels; Calculating the similarity score between the original text vector and the comparison text vector to generate an N×N dimensional classification matrix; The classification matrix is ​​trained based on the loss function, the parameters of the first pre-trained model and the second pre-trained model are optimized by gradient descent, and the readability level prediction of the text to be predicted is realized based on the trained first pre-trained model and the second pre-trained model.

2. The method according to claim 1, characterized in that The prediction process is as follows: Input the text to be predicted into the first pre-trained model to generate an original text vector. Input the text to be predicted into the trained second pre-trained model to generate N comparison text vectors. Calculate the similarity scores between the original text vector and the N comparison text vectors in turn. The readability level corresponding to the maximum similarity score is the readability level of the text to be predicted.

3. The method according to claim 1, characterized in that The comparison text vector is generated in one of the following two ways: 1) Using the second pre-trained model, encode the original text according to N readability levels, extract the last layer [CLS] tag vector corresponding to the readability level as the semantic representation, and concatenate it with the multi-linguistic features to generate N comparison text vectors, 2) Using N of the second pre-trained models, respectively encode the original text according to N readability levels, extract the last layer [CLS] tag vector corresponding to the readability level as the semantic representation, and concatenate it with the multi-linguistic features to generate N comparative text vectors; wherein each second pre-trained model parameter is initialized independently and shares the same vocabulary.

4. The method according to claim 1, characterized in that The gradient descent adopts the Adam gradient optimization algorithm, the first pre-trained model is a BERT encoder, and the second pre-trained model is a Qwen2.5-14B large language model or a BERT encoder.

5. The method according to claim 1, characterized in that The multi-linguistic features include numerical language features and proportional language features; Before the splicing, the multi-lingual features are subjected to standardized preprocessing; Z-score standardization is used for numerical language features, and Min-Max standardization is used for proportional language features.

6. The method according to claim 1, characterized in that The similarity score is calculated based on the cosine similarity, and for the original text vector of the i-th readability level and the comparison text vector of the jth readability level Similarity score The calculation formula is: , , represents a set of integers; The similarity score is the i-th row and j-th column element of the N×N dimensional classification matrix.

7. The method according to claim 1, characterized in that The original text type at least includes medical text; The multi-linguistic features include at least character complexity features, vocabulary complexity features, syntactic complexity features, text structure features and medical features.

8. The method according to claim 6, characterized in that The diagonal elements of the classification matrix represent the similarity of the original texts at the same readability level, and the off-diagonal elements represent the similarity across readability levels. The loss function is designed to maximize the diagonal elements of the classification matrix and minimize the off-diagonal elements; The formula of the loss function is as follows: The parameter τ is set to a value in the interval [0.05, 0.12] to adjust the sharpness of the similarity distribution.

9. A text readability classification system based on large models and contrastive learning, characterized in that: Includes the following modules: Manual extraction module: extracts multi-language features from the original text manually; The original feature generation module uses the first pre-trained model to encode the original text according to N readability levels, and concatenates it with the multi-linguistic features to generate N original text vectors. A contrast feature generation module: using the second pre-trained model, encoding the original text according to N readability levels, and concatenating it with the multi-linguistic features to generate N contrast text vectors, calculating the similarity scores between the text vectors, and generating an N×N dimensional classification matrix; Level prediction module: The classification matrix is ​​trained based on the loss function, and the parameters of the N pre-trained language models are optimized by gradient descent to achieve readability level prediction of the original text.

Citation Information

Patent Citations

  • Chinese text readability evaluation method and system fusing text distribution law characteristics

    CN113934850A

  • Text readability evaluation method and system based on international Chinese education Chinese level grade standard

    CN115859962A

  • Medical term standardization method based on fusion multi-strategy comparative learning

    CN117633148A

  • Dangerous behavior identification and early warning method based on multi-modal analysis

    CN119360278A

  • Text readability classification method, system and equipment based on comparative learning and medium

    CN119474971A

Cited By

  • Self-adaptive medical text classification method based on loss threshold and dynamic weight

    CN120561308A

  • An adaptive medical text classification method based on loss threshold and dynamic weight

    CN120561308B

  • AI generation text detection method based on sentence length distribution and text predictability characteristics

    CN122433705A

  • Ai-generated text detection method based on sentence length distribution and text predictability features

    CN122433705B