A text relevance scoring and medical entity standardization method and system

By constructing an adaptive boundary-focused loss function and a dynamic weight matrix, the training process of the relevance test model is optimized, the class imbalance problem is solved, and more accurate relevance scoring and medical entity standardization are achieved.

CN120561273BActive Publication Date: 2025-10-03SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511037546.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-03
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

Existing correlation test models cannot distinguish differences in correlation levels when the categories are unbalanced, resulting in inaccurate correlation scores between fuzzy medical entities and entities in the knowledge base, and reducing the accuracy of medical entity standardization.

Method used

An adaptive boundary focusing loss function is constructed to dynamically adjust the decision boundary of the correlation scoring category through boundary sparsity terms and accidental penalty terms. The training process of the correlation test model is optimized by combining the dynamic weight matrix and data balance sensitivity coefficient.

Benefits of technology

The accuracy of correlation scores and the standardization precision of medical entities have been improved, and it can accurately distinguish differences in correlation levels when facing data imbalance, thereby improving the adaptability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561273B_ABST
    Figure CN120561273B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and in particular to a text relevance scoring and medical entity standardization method and system. The present invention constructs an adaptive boundary focusing loss function to train a relevance test model, constructs a boundary sparsity term based on the absolute value of the learnable boundary offset between each adjacent relevance score category; combines the dynamic weight matrix and the boundary-aware probability distribution of the predicted relevance scores of all training set samples belonging to each relevance score category under the actual decision boundary, and constructs a contingency penalty term; the trained model can obtain an effective relevance score for two text entities. When standardizing medical entities, the model is used to calculate the relevance score of the entity to be standardized and each related entity, screen the target related entity, and use a large-scale pre-trained language model to determine the standardization result from the screening result. The present invention effectively improves the accuracy of medical entity standardization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a text relevance scoring and medical entity standardization method and system. Background Art

[0002] As a core technology for measuring the degree of association between texts, text relevance scoring is widely used in scenarios such as intelligent question-answering, information recommendation, and entity standardization. For example, in medical entity standardization, due to differences in physician experience and writing habits, medical entities in clinical texts contain various non-standard expressions, such as synonyms, abbreviations, spelling errors, and text omissions, which may affect the performance of medical text mining tasks. Therefore, by calculating the relevance scores of fuzzy medical entities in clinical texts and various entities in the knowledge base, the degree of association between fuzzy medical entities and entities in the knowledge base is measured, and the fuzzy medical entities in the original medical text are then mapped to standard concepts in the knowledge base. This can effectively unify the expression of medical terminology, eliminate semantic ambiguity, and improve the level of structured medical data.

[0003] Calculating text relevance scores is typically based on embedding techniques. This involves mapping text into dense vectors in a low-dimensional vector space using a pre-trained language model. Metrics such as cosine similarity and Euclidean distance are then used to calculate the similarity between these vectors, thereby quantifying text relevance. However, vector similarity is merely a quantitative representation based on semantics and cannot fully and accurately understand the precise meaning of text in a specific context. For example, in clinical cases, "hyperglycemia" and "abnormally elevated blood sugar levels" refer to the same medical concept but differ significantly in their expressions. Vector similarity makes it difficult to accurately identify the underlying semantic equivalence between the two, hindering effective standardization to the correct medical terminology, severely impacting medical data analysis. To address the limitations of embedding techniques in understanding context, relevance verification models have been introduced. These models deeply analyze the context and semantics of the original text, more accurately determining the true associations between fuzzy medical entities, effectively filtering out irrelevant entities, and improving the accuracy and efficiency of entity standardization. This ensures that each fuzzy medical entity is accurately mapped to the most appropriate standard concept, providing a reliable data foundation for subsequent tasks such as clinical decision support and disease prediction analysis.

[0004] Correlation testing models often use a cross-entropy loss function or a mean squared error loss function. However, when the categories are imbalanced, the cross-entropy loss function assigns equal weight to training set samples of different categories. This leads to an overemphasis on the more numerous categories and an inability to accurately identify the characteristic differences in different correlation scores. This can cause systematic deviations in the correlation scores of fuzzy medical entities and entities in the knowledge base, making it impossible to accurately map fuzzy medical entities with fewer category training set samples to the most appropriate standard concepts. The main drawback of the mean squared error loss function is its "symmetrical and uniform penalty" for prediction errors. The mean squared error loss function assumes that the differences in scores at different levels are linear and equidistant, which is contrary to actual clinical needs. Taking tumor marker testing as an example, misclassifying a strong correlation (4 points) between "elevated alpha-fetoprotein" and "liver cancer" as a moderate correlation (3 points) has a more serious impact on clinical diagnostic decisions than misclassifying a very strong correlation (5 points) as a strong correlation (4 points). However, the mean squared error loss function cannot distinguish this difference, causing the correlation test model to only pursue numerical proximity during training, ignoring the actual semantic association of the correlation level. This makes it impossible to accurately reflect the correlation scores between fuzzy medical entities and entities in the knowledge base, reducing the standardization accuracy of medical entities. Summary of the Invention

[0005] To this end, the technical problem to be solved by the present invention is to overcome the defect that the loss function of the existing correlation test model treats the weights of training set samples equally when the categories are unbalanced, and cannot distinguish the differences in correlation levels. The obtained correlation score is difficult to accurately reflect the correlation between fuzzy medical entities and entities in the knowledge base, which reduces the standardization accuracy of medical entities.

[0006] To solve the above technical problems, the present invention provides a text relevance scoring method, comprising:

[0007] Construct boundary sparsity terms based on the absolute values ​​of learnable boundary offsets between all adjacent relevance score categories;

[0008] Calculate the actual decision boundary between each adjacent relevance score category based on the learnable boundary offset between each adjacent relevance score category, the preset initial boundary between each adjacent relevance score category, and the scaling factor;

[0009] Based on the actual decision boundary between each adjacent relevance score category, the boundary-aware probability distribution of each training set sample belonging to each relevance score category under the actual decision boundary is calculated; each training set sample includes: a pair of text entities and their true relevance score labels;

[0010] Based on the data balance sensitivity coefficient and the number of training set samples corresponding to each correlation score, a dynamic weight matrix is ​​constructed;

[0011] Based on the dynamic weight matrix and the boundary-aware probability distribution of all training set samples belonging to each relevance score category under the actual decision boundary, a contingency penalty term is constructed;

[0012] Based on the boundary sparsity term and the accidental penalty term, an adaptive boundary focus loss function is constructed to train the correlation test model;

[0013] Input the two texts to be scored into the trained relevance test model and output the relevance score between the two texts to be scored.

[0014] Preferably, the actual decision boundary between each adjacent correlation score category is calculated based on the learnable boundary offset between each adjacent correlation score category, the preset initial boundary between each adjacent correlation score category, and the scaling factor, and the formula is:

[0015] ,

[0016] in, For the The actual decision boundary between adjacent relevance score categories, For the The preset initial boundaries between adjacent correlation score categories, is the scaling factor, For the The learnable boundary offset between adjacent relevance score categories, Scoring category interval index for adjacent correlations.

[0017] Preferably, the boundary-aware probability distribution of each training set sample belonging to each correlation score category under the actual decision boundary is calculated based on the actual decision boundary between each adjacent correlation score category, and the formula is:

[0018] ,

[0019] in, For the training set samples Belongs to the relevance score category under the actual decision boundary The boundary perception probability of For the training set samples, For the training set samples Boundary-aware probability distribution of belonging to arbitrary relevance score categories under the actual decision boundary, is an exponential function with a natural constant as its base, For the training set samples The probability distribution of belonging to each relevance score category, The number of categories to score relevance for, The index of the interval between adjacent correlation scoring categories, For the The actual decision boundary between adjacent relevance score categories, is the first relevance score category index, Score category index for the second relevance.

[0020] Preferably, the dynamic weight matrix Rank The elements of the column are:

[0021] ,

[0022] in, is the first in the dynamic weight matrix Rank Elements of the column, Score the true relevance label as the relevance score category The number of training set samples, Score the true relevance label as the relevance score category The number of training set samples, is the data balance sensitivity coefficient, is the first relevance score category index, Score category index for the second relevance, is an absolute value.

[0023] Preferably, the adaptive boundary focusing loss function is constructed based on the boundary sparsity term and the accidental penalty term, and the formula is:

[0024] ,

[0025] in, is the adaptive boundary focusing loss function, is the boundary sparse item weight, is the boundary sparse term, is the weight of the accidental penalty term, is the accidental penalty term, is the first in the dynamic weight matrix Rank Elements of the column, The number of categories to score relevance for, The index of the interval between adjacent correlation scoring categories, is the learnable boundary offset between the nth adjacent correlation score categories, for The absolute value of For the training set samples Belongs to the relevance score category under the actual decision boundary The boundary perception probability of For the training set samples Belongs to the relevance score category under the actual decision boundary The boundary perception probability of For the training set samples, For the training set samples Boundary-aware probability distribution of belonging to arbitrary relevance score categories under the actual decision boundary, is the first relevance score category index, Score category index for the second relevance.

[0026] Preferably, the relevance verification model is any one of a DeBERTa-BiLSTM model, a DeBERTa-Transformer model, a BERT-GRU model, and an entity interpretation enhanced interaction network.

[0027] Preferably, the two texts to be scored are subjected to an entity interpretation enhanced interaction network to obtain a relevance score of the two texts to be scored, including:

[0028] Obtain the context of each text to be scored, combine each text to be scored and its context, and pass them through the embedding layer to obtain the embedded information features of each text to be scored;

[0029] The embedded information features of the two texts to be rated are subjected to cross attention to extract the fusion attention features;

[0030] The fused attention features are passed through a fully connected neural network layer to obtain the relevance score of the two texts to be rated.

[0031] The present invention also provides a medical entity standardization method, comprising:

[0032] Matching corresponding standardized entities for the text to be standardized in the medical language database;

[0033] Using one of the above-mentioned text relevance scoring methods, respectively calculate the relevance score between the text to be standardized and each standardized entity;

[0034] Based on the relevance score between the text to be standardized and each standardized entity, the standardized entities of the text to be standardized are screened;

[0035] Through large-scale pre-trained language models, the standardized results of the text to be standardized are obtained from the screening results.

[0036] Preferably, matching the to-be-standardized text with the corresponding standardized entity in the medical language database includes:

[0037] The similarity between the embedding vector of the text to be standardized and the embedding vector of each entity in the medical language database is calculated respectively. According to the similarity, the entities in the medical language database are sorted in descending order, and the first preset number of entities in the sorting result are used as the corresponding standardized entities to match the text to be standardized.

[0038] The present invention also provides a medical entity standardization system, comprising:

[0039] A matching module, configured to match corresponding standardized entities for the text to be standardized in the medical language database;

[0040] A relevance scoring module, configured to calculate the relevance scores of the text to be standardized and each standardized entity respectively using one of the above-mentioned text relevance scoring methods;

[0041] A screening module, configured to screen the standardized entities of the text to be standardized based on a relevance score between the text to be standardized and each standardized entity;

[0042] The result acquisition module is used to obtain the standardized results of the text to be standardized from the screening results through a large-scale pre-trained language model.

[0043] The above technical solution of the present invention has the following beneficial effects compared with the prior art:

[0044] The present invention describes a method and system for text relevance scoring and medical entity standardization that introduces learnable boundary offsets between adjacent relevance scoring categories, allowing the relevance verification model to flexibly adjust the decision boundaries of adjacent relevance scoring categories based on data characteristics, thus breaking through the limitations of linear equidistance. The absolute values ​​of the learnable boundary offsets between all adjacent relevance scoring categories are summed to construct a boundary sparse term; this enables the relevance verification model to learn non-equidistant, more discriminative boundaries during training, allowing the relevance verification model to focus on subtle grade differences when judging the relevance of fuzzy medical entities and entities in medical language databases, thereby improving the accuracy of medical entity standardization.

[0045] The accidental penalty term is constructed based on the dynamic weight matrix and the boundary-aware probability distribution of all training set samples belonging to each correlation score category under the actual decision boundary; the dynamic weight matrix comprehensively considers the score difference distance and sample quantity difference between categories, and uses the data balance sensitivity coefficient to adjust the data balance sensitivity, which solves the problem of unreasonable sample weights when the categories are unbalanced. It also significantly enhances the model's ability to recognize differences in correlation score levels by combining the probability distribution recalculated under the actual decision boundary. Compared with the traditional loss function that treats sample weights equally, the accidental penalty term is based on the boundary-aware probability distribution under the actual decision boundary, and can adaptively adjust the penalty intensity for misclassification of samples of different categories, so that when the model handles the data-imbalanced medical entity correlation scoring task, it neither overly favors the majority class nor can it accurately distinguish differences in correlation levels, thereby improving the accuracy of correlation scoring and the accuracy of medical entity standardization. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:

[0047] Figure 1 It is a flow chart of a text relevance scoring method of the present invention.

[0048] Figure 2 It is a structural diagram of the entity interpretation enhanced interaction network.

[0049] Figure 3 It is a structural diagram of a medical entity standardization method of the present invention. DETAILED DESCRIPTION

[0050] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.

[0051] Reference Figure 1 As shown, this embodiment provides a text relevance scoring method, including the following steps:

[0052] Step S11: constructing a boundary sparsity term based on the absolute value of the learnable boundary offset between all adjacent relevance score categories;

[0053] In this embodiment, preferably, the boundary sparse term is constructed based on the absolute value of the learnable boundary offset between all adjacent correlation score categories, and the boundary sparse term is: ,in, The number of categories to score relevance for, The index of the interval between adjacent correlation scoring categories, For the The learnable boundary offset between adjacent relevance score categories, for The absolute value of .

[0054] In this embodiment, the correlation scoring categories include: no correlation (0 points), very weak correlation (1 point), weak correlation (2 points), moderate correlation (3 points), strong correlation (4 points), and very strong correlation (5 points). There are five adjacent category intervals. For example, when hour, Indicates the learnable boundary offset between no correlation (0 point) and very weak correlation (1 point). hour, Indicates the learnable boundary offset between very weak correlation (score of 1) and weak correlation (score of 2).

[0055] The boundary sparsity term accumulates the absolute values ​​of the learnable boundary offsets between adjacent relevance score categories. This allows the model to proactively widen the decision boundary between adjacent relevance score categories during learning, avoiding misjudgments due to similar features between categories. This mechanism not only effectively reduces overlap and confusion between categories, but also enhances the model's robustness to long-tail and marginal data, ensuring more accurate and reliable relevance scoring results when processing the complex semantic relationships in medical texts.

[0056] Step S12: Calculating the actual decision boundary between each adjacent relevance score category based on the learnable boundary offset between each adjacent relevance score category, the preset initial boundary between each adjacent relevance score category, and the scaling factor;

[0057] In this embodiment, preferably, the actual decision boundary between each adjacent correlation score category is calculated based on the learnable boundary offset between each adjacent correlation score category, the preset initial boundary between each adjacent correlation score category, and the scaling factor, and the formula is:

[0058] ,

[0059] in, For the The actual decision boundary between adjacent relevance score categories, For the The preset initial boundaries between adjacent correlation score categories, is the scaling factor, For the The learnable boundary offset between adjacent relevance score categories, Scoring category interval index for adjacent correlations.

[0060] By calculating the actual decision boundary between each adjacent relevance score category, the present invention enables the model to dynamically adjust the dividing threshold between different relevance levels based on the semantic distribution characteristics of entity pairs in the training data, thereby achieving refined control of semantic boundaries and significantly improving the model's ability to understand the semantic relationships of medical texts. This dynamic adjustment mechanism is particularly critical in medical text processing: on the one hand, it can accurately capture subtle semantic differences between entities such as clinical terms and disease names, for example, distinguishing the relevance boundary between the highly similar concepts of "myocardial infarction" and "angina pectoris"; on the other hand, the boundary adjustment amplitude is flexibly controlled through the scaling coefficient, ensuring the model's adaptability to specific semantic patterns in the medical field while maintaining the stability of the scoring criteria, effectively avoiding category confusion caused by data bias.

[0061] Step S13: Based on the actual decision boundary between each adjacent relevance score category, calculate the boundary-aware probability distribution of each training set sample belonging to each relevance score category under the actual decision boundary; wherein each training set sample includes: a pair of text entities and their true relevance score labels;

[0062] In this embodiment, preferably, the boundary-aware probability distribution of each training set sample belonging to each correlation score category under the actual decision boundary is calculated based on the actual decision boundary between each adjacent correlation score category, and the formula is:

[0063] ,

[0064] in, For the training set samples Belongs to the relevance score category under the actual decision boundary The boundary perception probability of For the training set samples, For the training set samples Boundary-aware probability distribution of belonging to arbitrary relevance score categories under the actual decision boundary, is an exponential function with a natural constant as its base, For the training set samples The probability distribution of belonging to each relevance score category, The number of categories to score relevance for, The index of the interval between adjacent correlation scoring categories, For the The actual decision boundary between adjacent relevance score categories, is the first relevance score category index, Score category index for the second relevance.

[0065] In classification tasks, the probability distribution of the current sample belonging to each correlation score category usually corresponds to the logits output by the neural network, and its numerical size reflects the model's "confidence" that the sample belongs to a certain category. However, the correlation test model assumes that the boundaries between categories are fixed and equidistant, and does not consider the subtle differences in the true distribution between categories. It ignores the complexity of the data distribution, resulting in inaccurate category distinction. This application combines a learnable boundary offset with a scaling factor to adjust the initial boundary so that the correlation category interval can adapt to the data distribution and obtain a boundary-aware probability distribution that reflects the true correlation category interval.

[0066] For example, when the relevance score categories include: no relevance (0 points), very weak relevance (1 point), weak relevance (2 points), moderate relevance (3 points), strong relevance (4 points), and very strong relevance (5 points), there are five adjacent relevance score category intervals. The initial boundaries between adjacent relevance score categories are set to 1.0, and the boundary offset between adjacent relevance score categories can be learned. 、 、 、 、 , scaling factor ;

[0067] The actual decision boundary between each adjacent relevance score category is calculated as: 、 、 、 、 ;

[0068] If the probability distribution of the current sample output by the fully connected neural network layer of the entity interpretation enhanced interaction network belongs to each relevance score category is , then according to the actual decision boundary between each adjacent correlation score category, the initial boundary perception probability distribution of the current sample belonging to each correlation score category under the actual decision boundary can be recalculated;

[0069] The current sample belongs to the first Initial boundary-aware probability of relevance score categories The current sample belongs to The probability of the relevance score category is the same as the The difference between the cumulative boundaries of the relevance score categories, where Cumulative boundaries of relevance score categories The calculation formula is: ;

[0070] For the irrelevant (0 points) category, its cumulative boundary is 0, and the initial boundary perception probability that the current sample belongs to the irrelevant (0 points) category under the actual decision boundary is 0.5 ;

[0071] For the very weak correlation (1 point) category, the cumulative boundary is , the initial boundary perception probability that the current sample belongs to the very weakly correlated (1 point) category under the actual decision boundary is ;

[0072] For the weakly correlated (2-point) category, the cumulative boundary is , the initial boundary perception probability of the current sample belonging to the weakly correlated (2 points) category under the actual decision boundary is ;

[0073] For the moderately correlated (3 points) category, the cumulative boundary is , the initial boundary perception probability that the current sample belongs to the moderately relevant (3 points) category under the actual decision boundary is ;

[0074] For the strong correlation (4 points) category, the cumulative boundary is , the initial boundary perception probability of the current sample belonging to the strong correlation (4 points) category under the actual decision boundary is ;

[0075] For the very strong correlation (5 points) category, the cumulative boundary is , the initial boundary perception probability that the current sample belongs to the very strong correlation (5 points) category under the actual decision boundary is ;

[0076] Softmax is used to calculate the boundary-aware probability distribution of the current sample belonging to each relevance score category under the actual decision boundary:

[0077] ,

[0078] in, For the current sample Belongs to the relevance score category under the actual decision boundary The boundary-aware probability distribution of For the current sample The initial boundary-aware probability of belonging to the relevance score category under the actual decision boundary, For the current sample The boundary-aware probability distribution of each relevance score category under the actual decision boundary, For the current sample, is the number of relevance scoring categories. In this embodiment, .

[0079] Calculate the boundary perception probability distribution of the current sample belonging to each relevance score category under the actual decision boundary .

[0080] The actual decision boundary is generated by dynamically adjusting the learnable boundary offset, the preset initial boundary, and the scaling factor. The probability distribution calculated based on this can reflect the subtle differences between data distribution characteristics and categories in real time. When the model deviates during training, such as inaccurately distinguishing between the "weakly correlated (2 points)" and "moderately correlated (3 points)" categories, the recalculated probability distribution will generate corresponding feedback in the loss function, prompting the correlation test model to adjust the parameters to optimize the boundary and avoid training falling into a local optimal solution. This dynamic feedback mechanism allows the model to continuously correct its judgment of category boundaries during training, gradually optimizing in a direction that is more consistent with the actual distribution of the data, significantly improving the effectiveness and pertinence of training.

[0081] During training, by adjusting the probability distribution based on the actual decision boundary and constructing a contingency penalty term, the model gains a deeper understanding of the boundary characteristics of different relevance score levels. This effectively avoids overfitting, especially when dealing with fuzzy intervals and data imbalance. For example, when training relevance scores for medical knowledge related to rare diseases, the model recalculates the probability distribution based on the actual decision boundary, fully exploiting the characteristics of rare, highly relevant samples. Even when faced with unseen similar data, the model can accurately determine the relevance score category based on the patterns learned during training, thereby improving the model's generalization ability.

[0082] Step S14: constructing a dynamic weight matrix based on the data balance sensitivity coefficient and the number of training set samples corresponding to each correlation score;

[0083] In this embodiment, preferably, the first Rank The elements of the column are:

[0084] ,

[0085] in, is the first in the dynamic weight matrix Rank Elements of the column, Score the true relevance label as the relevance score category The number of training set samples, Score the true relevance label as the relevance score category The number of training set samples, is the data balance sensitivity coefficient, is the first relevance score category index, Score category index for the second relevance, is an absolute value.

[0086] The present invention effectively solves the data imbalance and category misjudgment problems in medical text relevance scoring by constructing an accidental penalty term based on a dynamic weight matrix. The dynamic weight matrix is ​​based on the number of samples in the category training set of the real relevance score label. 、 Based on the set data balance sensitivity coefficient Dynamically adjust weights to accurately measure the misclassification costs between different relevance score categories, and adaptively give higher weights to categories with fewer training set samples. For example, in the relevance scoring of medical terms for rare diseases, misclassification due to data scarcity can be avoided by using the square of the inter-class distance. and training set sample size The dual constraints of differences strengthen the model's ability to distinguish category boundaries.

[0087] Step S15: Based on the dynamic weight matrix, the predicted relevance scores of all training set samples in the training set belong to the boundary perception probability of each relevance score category, and the accidental penalty term is constructed. ;

[0088] The accidental penalty term combines the prediction probability and dynamic weight to perform targeted penalties on accidental misclassifications in model predictions, effectively suppressing the model's tendency to overfit to most categories and effectively improving the accuracy and reliability of text relevance scores.

[0089] Step S16: Based on the boundary sparsity term and the accidental penalty term, an adaptive boundary focus loss function is constructed to train a correlation test model;

[0090] In this embodiment, preferably, an adaptive boundary focusing loss function is constructed based on the boundary sparsity term and the accidental penalty term, and the formula is:

[0091] ,

[0092] in, is the adaptive boundary focusing loss function, is the boundary sparse item weight, is the boundary sparse term, is the weight of the accidental penalty term, is the accidental penalty term, is the first in the dynamic weight matrix Rank Elements of the column, The number of categories to score relevance for, The index of the interval between adjacent correlation scoring categories, is the learnable boundary offset between the nth adjacent correlation score categories, for The absolute value of For the training set samples Belongs to the relevance score category under the actual decision boundary The boundary perception probability of For the training set samples Belongs to the relevance score category under the actual decision boundary The boundary perception probability of For the training set samples, For the training set samples Boundary-aware probability distribution of belonging to arbitrary relevance score categories under the actual decision boundary, is the first relevance score category index, Score category index for the second relevance.

[0093] The present invention innovatively introduces adaptive boundary focusing loss By integrating boundary sparsity terms and accidental penalty terms, the classification boundaries and penalty mechanisms can be dynamically adjusted according to data distribution, improving the model's adaptability and ability to distinguish entity descriptions of different complexities, and enhancing the accuracy and robustness of medical text relevance scoring. The boundary sparsity term strengthens boundary discrimination by constraining the learnable boundary offset between adjacent relevance scoring categories, effectively avoiding misjudgments in semantically ambiguous areas; the accidental penalty term is based on a dynamic weight matrix and adaptively penalizes the training set sample distribution for misclassifications between different categories, solving the common class imbalance problem in medical data. This multi-dimensional constraint mechanism can accurately capture subtle semantic differences between medical entities, while dynamically optimizing the decision boundary through learnable boundary offsets, boundary sparsity term weights, and accidental penalty term weights, enabling the relevance test model to achieve the optimal balance between generalization and specificity, effectively improving the prediction accuracy of the relevance test model.

[0094] Step S17: Input the two texts to be scored into the trained correlation test model, and output the correlation score between the two texts to be scored.

[0095] In this embodiment, specifically, the process of obtaining the relevance score of two texts to be scored is as follows:

[0096] The score corresponding to the highest probability correlation score category in the probability distribution of the two texts to be rated belonging to each correlation score category is used as the correlation score of the two texts to be rated.

[0097] In this embodiment, specifically, the relevance verification model is any one of a DeBERTa-BiLSTM model, a DeBERTa-Transformer model, a BERT-GRU model, and an entity interpretation enhanced interactive network.

[0098] Current medical entity standardization frameworks generally lack a sufficient understanding of context. Most methods only use the text entity itself as input, ignoring contextual information such as the patient's basic information, examination indicators, and disease course description. This contextual information is often crucial for correct standardization to precise terminology. The lack of context modeling can lead to mismatches in the standardization of ambiguous entities, reducing the reliability and clinical application value of the overall system. To address these technical issues, the present invention proposes an entity interpretation enhanced interactive network.

[0099] like Figure 2 As shown, Figure 2 Structural diagram of interaction networks enhanced for entity interpretation.

[0100] In this embodiment, preferably, the two texts to be scored are subjected to an entity interpretation enhanced interaction network to obtain a relevance score of the two texts to be scored, including:

[0101] Obtain the context of each text to be scored, combine each text to be scored and its context, and pass them through the embedding layer to obtain the embedded information features of each text to be scored;

[0102] In this embodiment, specifically, the context information of the text to be scored is obtained in the following manner:

[0103] For the text to be rated that originates from clinical text, the clinical text in which the text to be rated is directly used as its context to preserve the original medical context information;

[0104] If the text to be scored is not associated with clinical text, Baidu Encyclopedia will be searched and the Baidu definition corresponding to the text to be scored will be used as its context supplement to ensure that each text to be scored can obtain background information related to its semantics, thereby providing complete semantic support for subsequent text relevance analysis and medical entity standardization processing.

[0105] In this embodiment, specifically, the formula for combining the text to be scored and its context is:

[0106] ,

[0107] in, is the combined text to be scored and its context, Indicates the start of a sequence. Indicates the separation of sequences, is the text to be rated. The context of the text to be scored.

[0108] The embedded information features of the two texts to be rated are subjected to cross attention to extract the fusion attention features;

[0109] The fused attention features are passed through a fully connected neural network layer to obtain the relevance score of the two texts to be rated.

[0110] The entity interpretation enhanced interactive network provided by the present invention adopts a dual-pathway structure, which can obtain the context of the text to be scored in combination with clinical text or Baidu interpretation, ensure the provision of semantically relevant background information for the text to be scored, and provide a semantic basis for subsequent text relevance analysis. After splicing the text to be scored and its context, the deep representation of the entity name information and the contextual semantic information is extracted through the embedding layer, wherein the embedding layer can be composed of a pre-trained language model or a customized encoder. Taking into account the limited amount of information in the entity name itself, the present invention combines the contextual information of the text to be scored to form an information source, and realizes the semantic alignment and fusion between the two channels through the cross-attention mechanism. The fused semantic representation is further input into the fully connected neural network layer for feature integration and regression calculation, and finally outputs the relevance score of the candidate entity for sorting and filtering, which comprehensively considers the semantics of the entity itself, the contextual context and the interaction relationship between entities, and effectively solves the problem that traditional methods are difficult to handle complex semantic dependencies.

[0111] Most existing medical entity standardization methods use a single-round retrieval mechanism, which only searches for the most relevant candidates in the standard entity library based on the initial query. They lack a multi-round, fine-grained information deepening mechanism. When there are subtle but critical semantic differences between the candidate entities retrieved initially and the text to be standardized, the system lacks the ability to further trace and verify, which makes it easy to have entity matching errors or standardization failures. Especially in fine-grained medical classification systems such as diseases, drugs, and surgical procedures, the roughness of this single-layer retrieval is particularly prominent. Therefore, this embodiment 2 provides a medical entity standardization method.

[0112] like Figure 3 As shown, Figure 3 This is a structural diagram of a medical entity standardization method of the present invention.

[0113] Step S21: matching the corresponding standardized entity for the text to be standardized in the medical language database;

[0114] In this embodiment, specifically, the medical language database is any one of the Unified Medical Language System (UMLS) and the International Classification of Diseases (ICD) databases.

[0115] In this embodiment, specifically, matching the to-be-standardized text with the corresponding standardized entity in the medical language database includes:

[0116] The similarity between the embedding vector of the text to be standardized and the embedding vector of each entity in the medical language database is calculated respectively. According to the similarity, the entities in the medical language database are sorted in descending order, and the first preset number of entities in the sorting result are used as the corresponding standardized entities to match the text to be standardized.

[0117] In this embodiment, optionally, a medical entity embedding model obtained through contrastive learning training is used to perform vector encoding on each entity in the standardized text and the medical language database.

[0118] In this embodiment, preferably, the text to be standardized is passed through a large-scale pre-trained language model to generate multiple related texts of the text to be standardized, and the similarity between each related text and the embedding vector of each entity in the medical language database is calculated to match the corresponding standardized entity for the text to be standardized.

[0119] Large-scale pre-trained language models (LLMs) are used to analyze the text to be standardized, assist in understanding its implicit knowledge needs, and generate several auxiliary queries, thereby improving recall coverage.

[0120] In the task of medical entity standardization, the standardized entities retrieved in the first stage through surface embedding vector similarity retrieval often contain a large number of semantically irrelevant items, which affects the subsequent matching accuracy. Therefore, the present invention further screens the standardized entities through a relevance test model.

[0121] Step S22: using one of the above-mentioned text relevance scoring methods, respectively calculating the relevance score between the text to be standardized and each standardized entity;

[0122] Step S23: screening the standardized entities of the text to be standardized based on the relevance score between the text to be standardized and each standardized entity;

[0123] In this embodiment, specifically, the screening of the standardized entities of the text to be standardized includes: according to the correlation score between the text to be standardized and each standardized entity, eliminating the standardized entities with a correlation score less than or equal to 2, and retaining the remaining entities for subsequent screening.

[0124] Step S24: Obtain the standardized results of the text to be standardized from the screening results through a large-scale pre-trained language model.

[0125] For example, when the text to be standardized is “rheumatic heart disease”, the corresponding standardized entities are matched for the text to be standardized in the medical language database. The standardized entities include: “rheumatic heart disease”, “rheumatic heart valve disease”, “rheumatic pericarditis”, and “rheumatic myocarditis”. Finally, through a large-scale pre-trained language model, the standardized result of “chronic obstructive pulmonary disease” is “rheumatic heart disease”.

[0126] In this embodiment, specifically, obtaining the normalized result of the text to be normalized from the screening results by using a large-scale pre-trained language model includes:

[0127] Construct prompt words (Prompt), input the entity to be standardized and the filtered target-related entities into the large-scale pre-trained language model, and based on the powerful semantic reasoning and comprehensive capabilities of the large model, select the most semantically consistent entity from multiple target-related entities as the standardization result of the entity to be standardized.

[0128] By constructing prompt words and calling a large-scale language model to perform the final standardized entity selection, the present invention can combine the large model's capabilities in complex semantic reasoning and polysemy disambiguation, effectively make up for the shortcomings of traditional retrieval methods in deep semantic understanding, and further improve the consistency and accuracy of standardized results.

[0129] In this embodiment, specifically, the large-scale pre-trained language model is any one of Deepseek, GPT-4, and Qwen models.

[0130] The present invention effectively alleviates the matching error problems of traditional methods in the case of ambiguity, heterogeneous expressions, and context dependence through a multi-level screening mechanism of vector retrieval, correlation testing, and large-scale model reasoning, significantly improves the standardization accuracy, and can effectively deal with practical problems such as complex medical entity representation, strong ambiguity, and fine-grained classification, achieving medical entity standardization with low time consumption, high accuracy, and strong robustness.

[0131] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0132] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0133] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0134] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0135] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A text relevance scoring method, characterized in that: include: Construct boundary sparsity terms based on the absolute values ​​of learnable boundary offsets between all adjacent relevance score categories; Calculate the actual decision boundary between each adjacent relevance score category based on the learnable boundary offset between each adjacent relevance score category, the preset initial boundary between each adjacent relevance score category, and the scaling factor; Based on the actual decision boundary between each adjacent relevance score category, the boundary-aware probability distribution of each training set sample belonging to each relevance score category under the actual decision boundary is calculated; each training set sample includes: a pair of text entities and their true relevance score labels; Based on the data balance sensitivity coefficient and the number of training set samples corresponding to each correlation score, a dynamic weight matrix is ​​constructed; Rank The elements of the column are: , in, is the first in the dynamic weight matrix Rank Elements of the column, Score the true relevance label as the relevance score category The number of training set samples, Score the true relevance label as the relevance score category The number of training set samples, is the data balance sensitivity coefficient, is the first relevance score category index, Score category index for the second relevance, is the absolute value; Based on the dynamic weight matrix and the boundary-aware probability distribution of all training set samples belonging to each relevance score category under the actual decision boundary, a contingency penalty term is constructed; Based on the boundary sparsity term and the accidental penalty term, an adaptive boundary focus loss function is constructed to train the correlation test model; the formula of the adaptive boundary focus loss function is: , in, is the adaptive boundary focusing loss function, is the boundary sparse item weight, is the boundary sparse term, is the weight of the accidental penalty term, is the accidental penalty term, The number of categories to score relevance for, The index of the interval between adjacent correlation scoring categories, is the learnable boundary offset between the nth adjacent correlation score categories, for The absolute value of For the training set samples Belongs to the relevance score category under the actual decision boundary The boundary perception probability of For the training set samples Belongs to the relevance score category under the actual decision boundary The boundary perception probability of For the training set samples, For the training set samples Boundary-aware probability distribution of belonging to arbitrary relevance score categories under the actual decision boundary; Input the two texts to be scored into the trained relevance test model and output the relevance score between the two texts to be scored.

2. A text relevance scoring method according to claim 1, characterized in that: The actual decision boundary between each adjacent correlation score category is calculated based on the learnable boundary offset between each adjacent correlation score category, the preset initial boundary between each adjacent correlation score category, and the scaling factor. The formula is: , in, For the The actual decision boundary between adjacent relevance score categories, For the The preset initial boundaries between adjacent correlation score categories, is the scaling factor, For the The learnable boundary offset between adjacent relevance score categories, Scoring category interval index for adjacent correlations.

3. A text relevance scoring method according to claim 1, characterized in that: According to the actual decision boundary between each adjacent relevance score category, the boundary-aware probability distribution of each training set sample belonging to each relevance score category under the actual decision boundary is calculated, and the formula is: , in, For the training set samples Belongs to the relevance score category under the actual decision boundary The boundary perception probability of For the training set samples, For the training set samples Boundary-aware probability distribution of belonging to arbitrary relevance score categories under the actual decision boundary, is an exponential function with a natural constant as its base, For the training set samples The probability distribution of belonging to each relevance score category, The number of categories to score relevance for, The index of the interval between adjacent correlation scoring categories, For the The actual decision boundary between adjacent relevance score categories, is the first relevance score category index, Score category index for the second relevance.

4. A text relevance scoring method according to claim 1, characterized in that: The relevance verification model is any one of the DeBERTa-BiLSTM model, the DeBERTa-Transformer model, the BERT-GRU model, and the entity interpretation enhanced interaction network.

5. A text relevance scoring method according to claim 4, characterized in that: The two texts to be rated are enhanced through the entity interpretation interaction network to obtain the relevance scores of the two texts to be rated, including: Obtain the context of each text to be scored, combine each text to be scored and its context, and pass them through the embedding layer to obtain the embedded information features of each text to be scored; The embedded information features of the two texts to be rated are subjected to cross attention to extract the fusion attention features; The fused attention features are passed through a fully connected neural network layer to obtain the relevance score of the two texts to be rated.

6. A medical entity standardization method, characterized in that: include: Matching corresponding standardized entities for the text to be standardized in the medical language database; Utilizing the text relevance scoring method according to any one of claims 1 to 5, respectively calculating the relevance score between the text to be standardized and each standardized entity; Based on the relevance score between the text to be standardized and each standardized entity, the standardized entities of the text to be standardized are screened; Through large-scale pre-trained language models, the standardized results of the text to be standardized are obtained from the screening results.

7. A medical entity standardization method according to claim 6, characterized in that: The matching of the to-be-standardized text to the corresponding standardized entity in the medical language database includes: The similarity between the embedding vector of the text to be standardized and the embedding vector of each entity in the medical language database is calculated respectively. According to the similarity, the entities in the medical language database are sorted in descending order, and the first preset number of entities in the sorting result are used as the corresponding standardized entities to match the text to be standardized.

8. A medical entity standardization system, characterized in that: include: A matching module, configured to match corresponding standardized entities for the text to be standardized in the medical language database; A relevance scoring module, configured to calculate the relevance score between the text to be standardized and each standardized entity using the text relevance scoring method described in any one of claims 1 to 5; A screening module, configured to screen the standardized entities of the text to be standardized based on a relevance score between the text to be standardized and each standardized entity; The result acquisition module is used to obtain the standardized results of the text to be standardized from the screening results through a large-scale pre-trained language model.

Citation Information

Patent Citations

  • Medical term automatic standardization system and method integrating self-supervision and active learning

    CN113436698A

  • Reading type examination question generation system and method based on commonsense reasoning

    WO2023225858A1