Intelligent labeling and quality control method and system for clinical data
By building a medical knowledge graph and a multi-head attention mechanism, combined with reinforcement learning methods, clinical data labeling rules are automatically generated, which solves the time-consuming and labor-intensive problems of traditional labeling and achieves intelligent labeling with high accuracy and consistency.
Patent Information
- Application Number
- CN202510957863.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional clinical data annotation relies on manual work, which is time-consuming and labor-intensive, and the annotation quality is difficult to guarantee, and semantic understanding is insufficient.
We construct a medical knowledge graph, extract clinical text features through a multi-head attention mechanism, combine reinforcement learning methods for knowledge reasoning, generate a hierarchical semantic linking network, and automatically discover implicit association paths and transform them into annotation rules.
It improves the accuracy and completeness of annotation, realizes the intelligence and automation of the annotation process, reduces labor costs, and ensures the semantic consistency and adaptability of annotation.
Smart Images

Figure CN120853971A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to intelligent annotation and quality control methods and systems for clinical data. Background Technology
[0002] In the medical field, the annotation and quality control of clinical data are fundamental to medical research, clinical decision support, and the construction of intelligent medical systems. With the widespread adoption of electronic medical record systems, a vast amount of clinical text data is generated and stored, containing rich medical information such as disease diagnoses, symptom descriptions, and treatment plans. To fully utilize this data, unstructured clinical text needs to be transformed into structured medical knowledge, which requires high-quality data annotation. Traditional clinical data annotation mainly relies on manual work by medical experts. This method is not only time-consuming and labor-intensive, but also suffers from difficulty in guaranteeing annotation quality due to the complexity and specialization of medical knowledge. Summary of the Invention
[0003] The embodiments of the present invention provide a method and system for intelligent annotation and quality control of clinical data, which can solve the problems in the prior art.
[0004] A first aspect of the present invention provides a method for intelligent annotation and quality control of clinical data, comprising:
[0005] The system receives clinical text data and constructs a medical knowledge graph, which includes disease nodes, symptom nodes, and semantic relationships between nodes. The nodes in the medical knowledge graph are then vectorized to generate node feature vectors.
[0006] Semantic recognition is performed on the clinical text data. Contextual features in the clinical text data are extracted through a multi-head attention mechanism to identify medical entities and clinical event information in the clinical text data, obtain temporal dependencies between medical entities, and generate structured semantic representations.
[0007] Calculate the similarity between the structured semantic representation and the node feature vector, establish semantic links between semantic representations with similarity exceeding a preset similarity threshold and corresponding nodes, and form semantic association paths in the medical knowledge graph; classify the associations based on the strength of the semantic association paths to generate a hierarchical semantic link network.
[0008] Knowledge reasoning is performed in the hierarchical semantic linking network. Based on the reinforcement learning method, multi-step reasoning is carried out according to the semantic link strength to obtain the implicit association path between nodes; the implicit association path is transformed into clinical data annotation rules.
[0009] Validation samples are selected from the clinical text data. The clinical data annotation rules are validated and optimized using the validation samples to generate different combinations of annotation rules. The newly added clinical text data is automatically annotated using the combinations of annotation rules to generate annotation results.
[0010] Semantic recognition is performed on the clinical text data. A multi-head attention mechanism is used to extract contextual features from the clinical text data, identify medical entities and clinical event information within the clinical text data, obtain temporal dependencies between medical entities, and generate structured semantic representations, including:
[0011] Multi-head attention is calculated on the contextual features of the clinical text data, and a query matrix, key matrix, and value matrix are generated through linear transformation. Attention scores are calculated based on the query matrix, key matrix, and value matrix. The outputs of multiple attention heads are concatenated and subjected to linear transformation to obtain fused attention features.
[0012] A bidirectional propagation conditional random field is constructed based on the fused attention features. The conditional random field uses a learnable label transition probability matrix to label the boundary positions of medical entities. The labeled boundary positions of medical entities are combined with the fused attention features to obtain the contextual semantic representation of the medical entities. Based on the contextual semantic representation, the types of medical entities are classified to generate medical entity recognition results.
[0013] Temporal information is extracted from the medical entity identification results, and a temporal dependency graph is constructed by combining the occurrence order of the medical entities; the temporal dependency graph is subjected to temporal consistency verification to eliminate temporal conflicts and generate a temporal relationship network.
[0014] The medical entity identification results and the temporal relationship network are integrated to generate a structured semantic representation.
[0015] Semantic links are established between semantic representations with similarity exceeding a preset similarity threshold and their corresponding nodes, forming semantic association paths in the medical knowledge graph. Based on the strength of these semantic association paths, the associations are graded to generate a hierarchical semantic link network, including:
[0016] Semantic representations with similarity exceeding a preset similarity threshold are linked to nodes in the corresponding medical knowledge graph, and the similarity is recorded as the initial strength of the semantic link.
[0017] Based on the semantic link relationship, a breadth-first search is performed to obtain a set of adjacent nodes; the Euclidean distance between the nodes in the set of adjacent nodes is calculated sequentially; a semantic association path is constructed based on the Euclidean distance between the nodes, and the semantic association path contains a sequence of nodes from the source node to the target node;
[0018] The path score of the semantically associated path is calculated by multiplying the Euclidean distances between adjacent nodes on the path and introducing a path length penalty factor.
[0019] The semantic association path is divided into multi-level association paths based on the numerical range of the initial intensity, and the multi-level association paths are hierarchically integrated to construct a hierarchical semantic link network.
[0020] Knowledge reasoning is performed in the hierarchical semantic linking network. Based on reinforcement learning, multi-step reasoning is conducted using semantic link strength to obtain the implicit association paths between nodes, including:
[0021] A knowledge reasoning environment is constructed by combining the current node information, historical reasoning path information, and target node information to form a state vector, and an action space is constructed based on the set of adjacent nodes of the current node; a combined reward function is constructed based on the semantic link strength and the similarity of the target node.
[0022] The state vector and action space are subjected to depth Q-value calculation to obtain the state-action value; the strategy is iterated based on the state-action value and the combined reward function to obtain the optimized reasoning strategy; the optimized reasoning strategy is used as the decision basis for knowledge reasoning.
[0023] Multi-step reasoning is performed in the medical knowledge graph, and the state-action values obtained from each step of reasoning are multiplied together to generate candidate association paths; the semantic consistency of the candidate association paths is verified, and implicit association paths that meet the verification conditions are selected.
[0024] Transforming the implicit association paths into clinical data annotation rules includes:
[0025] The implicit association path is mapped to a rule template, which includes a method for converting path node features to predicates and a mapping method for annotation operations;
[0026] Based on the rule template, path features are extracted to generate the preconditions and execution actions for annotation rules; the reliability score of the implicit associated path is used as the confidence level of the annotation rule.
[0027] The annotation rules are simplified and similar rules are merged to generate an optimized set of annotation rules, which serves as the clinical data annotation rules that meet the evaluation criteria.
[0028] The clinical data annotation rules are validated and optimized using the validation samples to generate different combinations of annotation rules. The automatic annotation of newly added clinical text data using these combinations of annotation rules includes:
[0029] The system receives a validation sample dataset and clinical data annotation rules to be validated; extracts text features and standard annotation information from the validation sample dataset; annotates the validation sample dataset according to the clinical data annotation rules to be validated, and obtains the rule annotation results; compares the rule annotation results with the standard annotation information, and calculates the annotation precision and recall.
[0030] Based on the annotation accuracy and recall, a rule evaluation index is constructed to quantitatively assess the annotation capability of the clinical data annotation rules; an evaluation index threshold is set, and the optimal rule combination that meets the evaluation index threshold is selected.
[0031] New clinical text data is input into the optimal rule combination, and the new clinical text data is matched based on the rule annotation conditions and execution actions; the annotation execution order is determined according to the confidence and priority of the rules, and automatic annotation results are generated.
[0032] A second aspect of the present invention provides an intelligent annotation and quality control system for clinical data, comprising:
[0033] The first unit is used to receive clinical text data, construct a medical knowledge graph, which includes disease nodes, symptom nodes, and semantic relationships between nodes; and to vectorize the nodes in the medical knowledge graph to generate node feature vectors.
[0034] The second unit is used to perform semantic recognition on the clinical text data, extract contextual features from the clinical text data through a multi-head attention mechanism, identify medical entities and clinical event information in the clinical text data, obtain temporal dependencies between medical entities, and generate structured semantic representations.
[0035] The third unit is used to calculate the similarity between the structured semantic representation and the node feature vector, establish semantic links between semantic representations with similarity exceeding a preset similarity threshold and corresponding nodes, and form semantic association paths in the medical knowledge graph; classify the associations based on the strength of the semantic association paths to generate a hierarchical semantic link network.
[0036] The fourth unit is used to perform knowledge reasoning in the hierarchical semantic linking network. Based on the reinforcement learning method, it performs multi-step reasoning based on the semantic linking strength to obtain the implicit association paths between nodes; and transforms the implicit association paths into clinical data annotation rules.
[0037] The fifth unit is used to select validation samples from the clinical text data, use the validation samples to validate and optimize the clinical data annotation rules, generate different combinations of annotation rules, and use the combinations of annotation rules to automatically annotate the newly added clinical text data to generate annotation results.
[0038] A third aspect of the embodiments of the present invention,
[0039] An electronic device is provided, comprising:
[0040] processor;
[0041] Memory used to store processor-executable instructions;
[0042] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0043] Fourth aspect of the present invention,
[0044] A computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0045] The beneficial effects of this application are as follows:
[0046] This invention constructs a medical knowledge graph and combines it with a multi-head attention mechanism for semantic recognition, which can accurately capture medical entities and their temporal dependencies in clinical texts, significantly improving the accuracy and completeness of clinical data annotation and solving the technical problem of insufficient semantic understanding in traditional annotation methods.
[0047] This invention employs a knowledge reasoning method based on reinforcement learning to automatically discover implicit connection paths between nodes and transform them into annotation rules, thereby realizing the intelligent and automated annotation process, significantly reducing the cost of manual annotation, improving annotation efficiency, and ensuring semantic consistency of annotation through a hierarchical semantic linking network.
[0048] This invention validates and optimizes annotation rules through validation samples, generating annotation rule combinations adapted to different clinical scenarios. This gives the annotation system good scalability and adaptability, enabling it to cope with complex and ever-changing clinical data types and improving the quality control level and clinical application value of annotation results. Attached Figure Description
[0049] Figure 1 This is a flowchart illustrating the intelligent annotation and quality control method for clinical data according to an embodiment of the present invention.
[0050] Figure 2 A flowchart illustrating the process of generating a hierarchical semantic linking network. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0052] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0053] Figure 1 This is a flowchart illustrating the intelligent annotation and quality control method for clinical data according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0054] The system receives clinical text data and constructs a medical knowledge graph, which includes disease nodes, symptom nodes, and semantic relationships between nodes. The nodes in the medical knowledge graph are then vectorized to generate node feature vectors.
[0055] Semantic recognition is performed on the clinical text data. Contextual features in the clinical text data are extracted through a multi-head attention mechanism to identify medical entities and clinical event information in the clinical text data, obtain temporal dependencies between medical entities, and generate structured semantic representations.
[0056] Calculate the similarity between the structured semantic representation and the node feature vector, establish semantic links between semantic representations with similarity exceeding a preset similarity threshold and corresponding nodes, and form semantic association paths in the medical knowledge graph; classify the associations based on the strength of the semantic association paths to generate a hierarchical semantic link network.
[0057] Knowledge reasoning is performed in the hierarchical semantic linking network. Based on the reinforcement learning method, multi-step reasoning is carried out according to the semantic link strength to obtain the implicit association path between nodes; the implicit association path is transformed into clinical data annotation rules.
[0058] Validation samples are selected from the clinical text data. The clinical data annotation rules are validated and optimized using the validation samples to generate different combinations of annotation rules. The newly added clinical text data is automatically annotated using the combinations of annotation rules to generate annotation results.
[0059] In one optional implementation, semantic recognition is performed on the clinical text data. A multi-head attention mechanism is used to extract contextual features from the clinical text data, identify medical entities and clinical event information within the clinical text data, obtain temporal dependencies between medical entities, and generate a structured semantic representation, including:
[0060] Multi-head attention is calculated on the contextual features of the clinical text data, and a query matrix, key matrix, and value matrix are generated through linear transformation. Attention scores are calculated based on the query matrix, key matrix, and value matrix. The outputs of multiple attention heads are concatenated and subjected to linear transformation to obtain fused attention features.
[0061] A bidirectional propagation conditional random field is constructed based on the fused attention features. The conditional random field uses a learnable label transition probability matrix to label the boundary positions of medical entities. The labeled boundary positions of medical entities are combined with the fused attention features to obtain the contextual semantic representation of the medical entities. Based on the contextual semantic representation, the types of medical entities are classified to generate medical entity recognition results.
[0062] Temporal information is extracted from the medical entity identification results, and a temporal dependency graph is constructed by combining the occurrence order of the medical entities; the temporal dependency graph is subjected to temporal consistency verification to eliminate temporal conflicts and generate a temporal relationship network.
[0063] The medical entity identification results and the temporal relationship network are integrated to generate a structured semantic representation.
[0064] This invention provides a semantic recognition method for clinical text data. It extracts contextual features from clinical text data through a multi-head attention mechanism, identifies medical entities and clinical event information, obtains temporal dependencies between medical entities, and generates structured semantic representations.
[0065] When performing semantic recognition on clinical text data, the clinical text data is first received as input. For example, consider the sentence: "The patient developed fever symptoms three days ago, with a temperature of 38.5℃. Ibuprofen was administered to reduce the fever, and the temperature has now returned to normal." This text data is then segmented and tokenized, converted into a sequence of words: ["patient", "at", "three days ago", "developed", "fever", "symptoms", "", "temperature", "38.5℃", "", "administered", "ibuprofen", "reduced fever", "treatment", "now", "temperature", "already", "recovered", "normal"].
[0066] Multi-head attention computation is performed on the contextual features of clinical text data, generating query, key, and value matrices through linear transformations. Assuming a lexical embedding dimension of 512, eight attention heads are set, each with a dimension of 64. Three different linear transformation layers are applied to the embedding representation of each lexical, generating query, key, and value vectors respectively. For example, for the lexical "fever," the query transformation matrix generates the query vector Q_fever, the key transformation matrix generates the key vector K_fever, and the value transformation matrix generates the value vector V_fever.
[0067] Attention scores are calculated based on the query matrix, key matrix, and value matrix. For each attention head, the dot product of the query vector and all key vectors is calculated to obtain the raw attention score. Taking the first attention head as an example, the dot product of the query vector for "hot" and the key vectors of all terms is calculated, resulting in a series of scores [0.2, 0.1, 0.3, 0.8, 1.0, 0.9, 0.1, 0.7, 0.5, 0.1, 0.2, 0.4, 0.8, 0.3, 0.1, 0.2, 0.6, 0.1, 0.2, 0.4]. These scores are divided by 8 (i.e., the square root of each head dimension) and then normalized using the softmax function to obtain the attention weights. The value vectors are then weighted and summed using these weights to obtain the output of the attention head.
[0068] The outputs of the eight attention heads are concatenated to form a vector of dimension 512, which is then passed through a linear transformation layer to obtain the fused attention feature. For the word "fever", the resulting fused attention feature vector contains contextual information related to the word, with particular emphasis on its association with words such as "symptoms", "body temperature", and "38.5℃".
[0069] A bidirectional propagation conditional random field (CRF) is constructed based on fused attention features. This CRF uses a learnable label transition probability matrix to label the boundary positions of medical entities. The label set is defined as {B-symptom, I-symptom, B-drug, I-drug, B-treatment, I-treatment, B-sign, I-sign, O}, where B represents the entity's starting position, I represents a position within the entity, and O represents a non-entity position. The CRF model calculates the label probability distribution for each position based on the fused attention features of the current word and the label information of neighboring words. For example, for "fever" in the sequence, the model predicts a probability of 0.92 for the label B-symptom, with lower probabilities for other labels; for "symptom," the model predicts a probability of 0.89 for the label I-symptom. The optimal label path was found throughout the sequence using the Viterbi algorithm. "Fever symptoms" was labeled as a symptom entity, "ibuprofen" as a drug entity, "antipyretic treatment" as a treatment entity, and "body temperature 38.5℃" and "body temperature has returned to normal" as vital signs entities.
[0070] The boundary positions of the labeled medical entities are combined with the fusion attention features to obtain the contextual semantic representation of the medical entities. For each identified medical entity, the fusion attention features of all words within its boundary are extracted, and the vector representation of the entity is obtained through average pooling. For example, for the entity "fever symptoms", the fusion attention features of "fever" and "symptoms" are averaged to obtain the semantic representation vector of the entity [0.42, 0.35, ..., 0.51].
[0071] The system classifies medical entities based on contextual semantic representations, generating medical entity identification results. A fully connected layer maps the semantic representation vectors of entities to the entity type space, and a softmax function is used to obtain the type probability distribution. Ultimately, "fever symptoms" is identified as a symptom type, "ibuprofen" as a drug type, "antipyretic treatment" as a treatment type, and "body temperature 38.5℃" and "body temperature has returned to normal" as signs types.
[0072] Temporal information is extracted from the medical entity recognition results, and a temporal dependency graph is constructed by combining the occurrence order of medical entities. For the example text, the time words "three days ago" and "now" are identified, corresponding to time points t1 and t2 respectively. By analyzing the syntactic structure and contextual relationships, it is determined that "fever symptoms" and "body temperature 38.5℃" occur at t1, "ibuprofen antipyretic treatment" occurs after t1 and before t2, and "body temperature has returned to normal" occurs at t2. A temporal dependency graph is constructed, where nodes are medical entities and edges represent temporal relationships. For example, the edge from "fever symptoms" to "ibuprofen antipyretic treatment" is labeled "before," indicating that the former occurs before the latter.
[0073] Perform temporal consistency verification on the temporal dependency graph, eliminate temporal conflicts, and generate a temporal relationship network. Check the graph for loops or contradictory temporal relationships. If conflicts exist, adjust the conflicting relationships based on contextual information and medical knowledge. In the example, there are no temporal conflicts, so the original dependency graph structure is maintained as the final temporal relationship network.
[0074] The results of medical entity recognition and the temporal relationship network are integrated to generate a structured semantic representation. Each medical entity is represented as a structured object, containing entity text, entity type, entity location, and time information. Temporal relationships are represented as associations between entities. The final structured semantic representation is generated in JSON format, including the symptom "fever" (time: t1), the signs "body temperature 38.5℃" (time: t1) and "body temperature has returned to normal" (time: t2), the drug "ibuprofen" (time: between t1 and t2), the treatment "antipyretic treatment" (time: between t1 and t2), and the temporal relationships between them, such as "fever" preceding "ibuprofen antipyretic treatment," and "ibuprofen antipyretic treatment" preceding "body temperature has returned to normal."
[0075] In one optional implementation, semantic links are established between semantic representations with similarity exceeding a preset similarity threshold and their corresponding nodes, forming semantic association paths in the medical knowledge graph; the associations are then graded based on the strength of these semantic association paths to generate a tiered semantic link network, including:
[0076] Semantic representations with similarity exceeding a preset similarity threshold are linked to nodes in the corresponding medical knowledge graph, and the similarity is recorded as the initial strength of the semantic link.
[0077] Based on the semantic link relationship, a breadth-first search is performed to obtain a set of adjacent nodes; the Euclidean distance between the nodes in the set of adjacent nodes is calculated sequentially; a semantic association path is constructed based on the Euclidean distance between the nodes, and the semantic association path contains a sequence of nodes from the source node to the target node;
[0078] The path score of the semantically associated path is calculated by multiplying the Euclidean distances between adjacent nodes on the path and introducing a path length penalty factor.
[0079] The semantic association path is divided into multi-level association paths based on the numerical range of the initial intensity, and the multi-level association paths are hierarchically integrated to construct a hierarchical semantic link network.
[0080] Figure 2 This embodiment provides a method for constructing a hierarchical semantic link network in a medical knowledge graph, illustrating the process of generating a hierarchical semantic link network. The method establishes links between semantic representations and knowledge graph nodes, and classifies these links based on the strength of the semantic association paths, ultimately generating a hierarchical semantic link network.
[0081] Specifically, the semantic representation of each node in the medical knowledge graph is first obtained. For each node, a pre-trained language model is used to extract its semantic features, forming a high-dimensional vector representation. For example, for the node "hypertension", it can be converted into a 768-dimensional semantic vector [0.21, -0.15, 0.32, ..., 0.08] using a language model.
[0082] For input medical queries or symptom descriptions, the same language model is used to extract their semantic representations. For example, for the input "headache accompanied by dizziness and nausea", the semantic vector [0.18, -0.22, 0.30, ..., 0.12] is extracted through the language model.
[0083] The cosine similarity between the semantic representation of the input query and the semantic representations of each node in the medical knowledge graph is calculated. In this embodiment, the preset similarity threshold is set to 0.75. When the calculated similarity is greater than 0.75, a semantic link is established between the input query and the corresponding node, and the similarity value is recorded as the initial strength of the semantic link.
[0084] For example, the input query "headache accompanied by dizziness and nausea" has a similarity of 0.82 with the node "migraine", 0.78 with the node "meningitis", and 0.76 with the node "vestibular neuronitis". Since these similarities all exceed the preset threshold of 0.75, semantic links are established between the input query and these nodes, and the initial strengths are recorded as 0.82, 0.78, and 0.76, respectively.
[0085] Based on the established semantic links, a breadth-first search algorithm is executed to obtain the set of neighboring nodes. Starting from the nodes that have established semantic links with the input query, these nodes are regarded as source nodes, and all nodes directly connected to them in the medical knowledge graph are searched to form the set of neighboring nodes.
[0086] In this embodiment, the neighboring nodes of the node "migraine" include "vasodilation", "neuritis", "photosensitivity", etc.; the neighboring nodes of the node "meningitis" include "neck stiffness", "fever", "abnormal cerebrospinal fluid", etc.; and the neighboring nodes of the node "vestibular neuronitis" include "vertigo", "balance disorder", "deafness", etc.
[0087] Calculate the Euclidean distance between nodes in the set of adjacent nodes. The semantic representation of a node is a high-dimensional vector, and the Euclidean distance between two nodes is obtained by calculating the Euclidean distance between two high-dimensional vectors. For example, the Euclidean distance between the node "migraine" and its neighbor "vasodilation" is 0.45, the Euclidean distance between it and "neuritis" is 0.38, and the Euclidean distance between it and "photosensitivity" is 0.52.
[0088] Semantic association paths are constructed based on the Euclidean distance between nodes. A semantic association path contains a sequence of nodes from the source node to the target node. In this embodiment, starting from the source node "migraine," multiple semantic association paths can be constructed, such as "migraine-neuritis-pain receptor activation-persistent headache" and "migraine-vasodilation-blood pressure fluctuation-hypertension-cerebrovascular accident," etc.
[0089] Calculate the path score for semantically related paths. The path score is obtained by multiplying the Euclidean distances between adjacent nodes on the path and introducing a path length penalty factor. The path length penalty factor is set to 0.9 raised to the power of the path length to reduce the score of excessively long paths. For example, for the path "migraine-neuritis-pain receptor activation-persistent headache", the Euclidean distances between adjacent nodes are 0.38, 0.42, and 0.35, respectively, and the path length is 4, then the path score is:
[0090] 0.38×0.42×0.35×(0.9^4)=0.38×0.42×0.35×0.6561=0.0362.
[0091] The semantic association paths are divided into multi-level association paths based on the numerical range of the initial strength. In this embodiment, the initial strength is divided into three levels: high association (0.85-1.0), medium association (0.75-0.85), and low association (below 0.75 but an association derived through other paths). For example, the initial strength of the input query with "migraine" is 0.82, which is a medium association; the initial strength with "meningitis" is 0.78, also a medium association; and the association strength derived through the path with "persistent headache" is 0.0362, which is a low association.
[0092] A hierarchical semantic linking network is constructed by hierarchically integrating multi-level associated paths. First, all direct links with high and medium relevance are retained. Then, some low-relevance paths are selected and retained based on path scores. In this embodiment, a path score threshold of 0.03 is set for low-relevance paths; low-relevance paths with scores higher than this threshold will be retained in the hierarchical semantic linking network.
[0093] The resulting hierarchical semantic linking network contains semantic association paths of varying strengths, visualized as a radial structure with the input query at the center and associated nodes of different levels surrounding it. Nodes are connected by directed edges, the thickness of which indicates the strength of the association. This hierarchical semantic linking network can be used in clinical decision support scenarios such as medical diagnostic assistance, etiological analysis, and treatment plan recommendation.
[0094] In one optional implementation, knowledge reasoning is performed in the hierarchical semantic linking network. Based on a reinforcement learning method, multi-step reasoning is conducted using semantic link strength to obtain the implicit association paths between nodes, including:
[0095] A knowledge reasoning environment is constructed by combining the current node information, historical reasoning path information, and target node information to form a state vector, and an action space is constructed based on the set of adjacent nodes of the current node; a combined reward function is constructed based on the semantic link strength and the similarity of the target node.
[0096] The state vector and action space are subjected to depth Q-value calculation to obtain the state-action value; the strategy is iterated based on the state-action value and the combined reward function to obtain the optimized reasoning strategy; the optimized reasoning strategy is used as the decision basis for knowledge reasoning.
[0097] Multi-step reasoning is performed in the medical knowledge graph, and the state-action values obtained from each step of reasoning are multiplied together to generate candidate association paths; the semantic consistency of the candidate association paths is verified, and implicit association paths that meet the verification conditions are selected.
[0098] When performing knowledge reasoning in a hierarchical semantic linking network, a knowledge reasoning environment is first constructed. This environment includes information about the current node, historical reasoning paths, and the target node, which are combined to form a state vector. For example, for the starting node "headache" and the target node "cerebral hemorrhage" in a medical knowledge graph, the system combines the semantic feature vector of "headache" (such as a 768-dimensional vector obtained through a pre-trained language model), the empty historical path vector, and the semantic feature vector of "cerebral hemorrhage" into a state vector. Simultaneously, the system constructs an action space based on the neighboring nodes of the current node "headache," such as "migraine," "temporal lobe epilepsy," and "hypertension," where actions represent possible next directions of reasoning.
[0099] The system constructs a combined reward function based on semantic link strength and target node similarity. Semantic link strength is a pre-calculated semantic association between two nodes, ranging from 0 to 1; for example, the semantic link strength between "headache" and "hypertension" is 0.75. Target node similarity is obtained by calculating the semantic similarity between the current node and the target node; for example, using cosine similarity, the similarity between "hypertension" and "cerebral hemorrhage" is 0.62. The combined reward function is: R = 0.7 × semantic link strength + 0.3 × target node similarity. This reward function design considers both the reliability of the current reasoning step and the relevance of the reasoning direction to the final goal.
[0100] The system uses a deep neural network architecture to calculate the deep Q-values of the state vector and action space. The input layer receives the state vector and extracts features through two hidden layers (each containing 256 neurons using the ReLU activation function). The output layer corresponds to the Q-value of each possible action. For example, when the state is "headache", the system calculates a Q-value of 0.82 for the action "choose hypertension" and a Q-value of 0.65 for the action "choose migraine".
[0101] The system iterates its policy based on the value of state-action and the combined reward function, employing an experience replay mechanism to store state transition samples during the inference process. Each sample includes the current state, the selected action, the obtained reward, and the next state. The system randomly selects batches of samples (e.g., batches of size 32) from the experience pool and updates the Q-network parameters according to the principle of temporal difference learning. The system uses an ε-greedy policy to balance exploration and exploitation, with an initial ε value set at 0.9, gradually decreasing to 0.1 during training. After 10,000 rounds of iterative training, the system obtains an optimized inference policy, which serves as the decision-making basis for knowledge inference.
[0102] When performing multi-step reasoning in a medical knowledge graph, the system progressively selects the next node based on an optimized strategy. For example, starting with "headache," the system selects "hypertension" as the first step (Q value 0.82), then "vascular wall injury" (Q value 0.78), and finally reaches "cerebral hemorrhage" (Q value 0.91). The system multiplies the state-action values obtained from each step of reasoning to generate the overall confidence of the candidate association path. For the path "headache → hypertension → vascular wall injury → cerebral hemorrhage," its confidence is 0.82 × 0.78 × 0.91 = 0.58.
[0103] The system performs semantic consistency verification on candidate association paths to ensure that the paths are logically sound in medical terms. The verification process includes two parts: first, a path integrity check to ensure direct semantic links exist between nodes in the path; and second, a semantic fluency check to ensure that the path representation conforms to medical knowledge. The system evaluates the semantic fluency score of the paths using a pre-trained medical knowledge model, setting a threshold of 0.65, and filtering out paths with scores higher than the threshold. For example, the semantic fluency score of the path mentioned above is 0.78, which is greater than the threshold, and is therefore confirmed as a valid implicit association path.
[0104] In practical applications, the system successfully identified multiple implicit correlation pathways in its analysis of the association between a patient's symptoms of "persistent headache accompanied by hypertension" and the disease "cerebral hemorrhage," including "persistent headache → hypertension → vascular wall damage → cerebral hemorrhage" (confidence level 0.58) and "persistent headache → blood pressure fluctuations → increased vascular pressure → cerebral hemorrhage" (confidence level 0.49). Doctors can use this information to assess the patient's risk of cerebral hemorrhage, promptly arrange further examinations, and improve diagnostic accuracy.
[0105] Through the above implementation methods, the present invention realizes a knowledge reasoning method based on reinforcement learning in hierarchical semantic linking networks, effectively mining implicit association paths between nodes, and providing decision support for fields such as medical diagnosis.
[0106] In one optional implementation, converting the implicit association path into clinical data annotation rules includes:
[0107] The implicit association path is mapped to a rule template, which includes a method for converting path node features to predicates and a mapping method for annotation operations;
[0108] Based on the rule template, path features are extracted to generate the preconditions and execution actions for annotation rules; the reliability score of the implicit associated path is used as the confidence level of the annotation rule.
[0109] The annotation rules are simplified and similar rules are merged to generate an optimized set of annotation rules, which serves as the clinical data annotation rules that meet the evaluation criteria.
[0110] The implicit association paths obtained based on reinforcement learning consist of three parts: path sequence, node features, and path score. In the association path example, the path sequence is "symptom description - disease manifestation - diagnosis result," the node features include "persistent headache - migraine symptoms - tension headache," and the path score is 0.85. The system constructs a rule template structure, which contains four fields: rule identifier, conditional predicate, execution action, and rule confidence. The path information is mapped to the template structure through a mapping function. The specific mapping process includes: extracting the semantic features of the path nodes and converting them into rule conditional predicates; extracting the path directionality and determining the annotation operation type; extracting the path score and setting the rule confidence.
[0111] For the example above, the system extracts the "persistent headache" node feature and converts it into the conditional predicate "symptom_duration>24h AND pain_location=head"; extracts the "migraine symptom" feature and converts it into the conditional predicate "pain_type=migraine"; the path points to "tension headache" and is converted into the annotation action "annotation_type=tension_headache"; the path score of 0.85 is used as the rule confidence. A complete rule example is: rule identifier "R001", conditional predicate "symptom_duration>24h AND pain_location=head AND pain_type=migraine", execution action "set_label(tension_headache)", confidence 0.85.
[0112] After rule generation, optimization is performed. During the structural simplification process, similar conditional predicates are merged, such as "symptom_duration>24h" and "symptom_lasting>1d" being combined into a unified expression "duration>24h". During rule merging, rules with inclusion relationships are combined. For example, rules R001 and R002 contain the conditions "head_pain" and "chronic_head_pain" respectively. Since the latter is a special case of the former, the system merges the two rules, retaining the more specific conditional description. Example of the optimized rule set: Rule R001', conditional predicate "duration>24h AND location=head AND type=migraine", action "set_label(tension_headache)", confidence level 0.85.
[0113] Rule evaluation uses a test dataset to validate rule effectiveness. The test data contains 100 standardized clinical records, each with a symptom description and a standard diagnostic label. The system applies the generated rules to annotate the test data and calculates the annotation accuracy. In the example, rule R001' achieves 85% precision and 75% recall on the test data. Based on the evaluation results, rules are filtered, retaining those with precision exceeding 80% and recall exceeding 70% as the final output.
[0114] The final set of rules is applied to the automatic annotation of newly added clinical texts. Example input text: "The patient reported persistent headache for more than 48 hours, located on one side, accompanied by nausea." The system sequentially matches the rule conditions, finds that it meets the requirements of rule R001', executes the corresponding annotation action, generates the annotation result "tension headache," and records the rule confidence score of 0.85 as the annotation reliability index.
[0115] Through the above processing flow, the system transforms implicit correlation paths into executable annotation rules, and ensures the usability of the annotation rules through rule optimization and evaluation. In practical applications, the system can process complex clinical texts containing multiple symptom features and diagnostic types, generating corresponding annotation rule sets, and providing reliable rule support for the automated annotation of clinical data.
[0116] In one optional implementation, the clinical data annotation rules are validated and optimized using the validation samples to generate different combinations of annotation rules. The automatic annotation of newly added clinical text data using these combinations of annotation rules includes:
[0117] The system receives a validation sample dataset and clinical data annotation rules to be validated; extracts text features and standard annotation information from the validation sample dataset; annotates the validation sample dataset according to the clinical data annotation rules to be validated, and obtains the rule annotation results; compares the rule annotation results with the standard annotation information, and calculates the annotation precision and recall.
[0118] Based on the annotation accuracy and recall, a rule evaluation index is constructed to quantitatively assess the annotation capability of the clinical data annotation rules; an evaluation index threshold is set, and the optimal rule combination that meets the evaluation index threshold is selected.
[0119] New clinical text data is input into the optimal rule combination, and the new clinical text data is matched based on the rule annotation conditions and execution actions; the annotation execution order is determined according to the confidence and priority of the rules, and automatic annotation results are generated.
[0120] In one embodiment of an automated clinical data annotation system, the system first receives a validation sample dataset and clinical data annotation rules to be validated. The validation sample dataset contains multiple clinical record texts and their standard annotation information, which is manually annotated by professional physicians. For example, the validation sample dataset may contain a clinical record: "The patient's blood pressure measurement was 160 / 95 mmHg, with a history of hypertension," where "hypertension" is marked as a disease entity.
[0121] The system extracts text features from the validation sample dataset, including lexical features, grammatical features, and contextual features. Lexical features are achieved by segmenting clinical text into words or phrases using word segmentation techniques. For example, "The patient's blood pressure measurement was 160 / 95 mmHg, and he has a history of hypertension" is segmented into words such as "patient," "blood pressure," "measurement value," "160 / 95 mmHg," "hypertension," and "medical history." Grammatical features identify the grammatical role of each word through part-of-speech tagging, such as tagging "hypertension" as a noun. Contextual features consider the surrounding word environment of the target word; for example, the words "has" before "hypertension" and "medical history" after it suggest that it may be a disease.
[0122] The system annotates the validation sample dataset according to the clinical data annotation rules to be validated. Suppose there is a rule: "If the word 'hypertension' appears in the text and is immediately followed by 'medical history,' then 'hypertension' will be labeled as a disease entity." After applying this rule, the system will label "hypertension" in the example text as a disease entity, thus obtaining the rule-based annotation result.
[0123] The system compares the rule annotation results with the standard annotation information to calculate the annotation precision and recall. Annotation precision is the ratio of the number of entities correctly annotated by the rule to the total number of entities annotated by the rule; recall is the ratio of the number of entities correctly annotated by the rule to the total number of entities in the standard annotation. In the example, if the rule correctly annotates "hypertension" as a disease entity, and this is the only disease entity in the standard annotation, then both precision and recall are 100%.
[0124] Based on annotation accuracy and recall, the system constructs rule evaluation metrics to quantitatively assess the annotation capability of clinical data annotation rules. A commonly used evaluation metric is the F1 score, which is the harmonic mean of accuracy and recall. The system can set specific thresholds for evaluation metrics, such as an F1 score greater than 0.8, to filter out rules that meet the criteria.
[0125] The system combines multiple rules to form rule combinations and evaluates the performance of each combination. For example, combining the aforementioned rule with another rule: "If the word 'hypertension' appears in the text and is preceded by 'having', then label 'hypertension' as a disease entity," may yield a higher F1 score. By trying different rule combinations, the system ultimately selects the combination with the highest evaluation index as the optimal rule combination.
[0126] To optimize rule combinations, the system can adjust rule parameters or modify rule logic based on verification results. For example, if a rule is found to be incorrectly labeling "low blood pressure" as hypertension, the rule can be modified to add a negative condition: "does not contain the word 'low'", thereby improving accuracy.
[0127] When new clinical text data needs to be labeled, the system inputs it into the optimal rule combination. For each new text, the system matches the labeling conditions and actions of the rules. For example, in the new text "The patient was diagnosed with essential hypertension and antihypertensive treatment is recommended," the system detects that "hypertension" meets the conditions in the optimal rule combination and labels it as a disease entity.
[0128] The system determines the annotation execution order based on the rule's confidence level and priority. Confidence level reflects the rule's reliability, typically based on the accuracy during the verification phase; priority determines which rule's annotation result is adopted first when multiple rules match successfully. For example, if rule A has an accuracy rate of 95%, rule B has an accuracy rate of 85%, and both match a certain text segment, the system will prioritize the annotation result of rule A.
[0129] The automatically generated annotation results include entity type, location information, and confidence score. For example, for the text "The patient was diagnosed with essential hypertension and antihypertensive treatment is recommended", the automatic annotation result might be: {"entity":"hypertension","type":"disease","position":[7,10],"confidence":0.95}, indicating that "hypertension" is annotated as a disease entity, located at the 7th to 10th character position of the text, with an annotation confidence score of 0.95.
[0130] The system can also provide a visual representation of the annotation results, such as using different colors to mark different types of entities in the original text, helping medical professionals quickly review the annotation results. Annotations that the system cannot determine can be marked as pending confirmation for manual review.
[0131] Through the above implementation methods, the system can effectively and automatically annotate clinical text data, reduce the manual annotation burden on medical professionals, improve the efficiency of clinical data processing, and provide support for subsequent clinical research and medical decision-making.
[0132] The present invention provides an intelligent annotation and quality control system for clinical data, comprising:
[0133] The first unit is used to receive clinical text data, construct a medical knowledge graph, which includes disease nodes, symptom nodes, and semantic relationships between nodes; and to vectorize the nodes in the medical knowledge graph to generate node feature vectors.
[0134] The second unit is used to perform semantic recognition on the clinical text data, extract contextual features from the clinical text data through a multi-head attention mechanism, identify medical entities and clinical event information in the clinical text data, obtain temporal dependencies between medical entities, and generate structured semantic representations.
[0135] The third unit is used to calculate the similarity between the structured semantic representation and the node feature vector, establish semantic links between semantic representations with similarity exceeding a preset similarity threshold and corresponding nodes, and form semantic association paths in the medical knowledge graph; classify the associations based on the strength of the semantic association paths to generate a hierarchical semantic link network.
[0136] The fourth unit is used to perform knowledge reasoning in the hierarchical semantic linking network. Based on the reinforcement learning method, it performs multi-step reasoning based on the semantic linking strength to obtain the implicit association paths between nodes; and transforms the implicit association paths into clinical data annotation rules.
[0137] The fifth unit is used to select validation samples from the clinical text data, use the validation samples to validate and optimize the clinical data annotation rules, generate different combinations of annotation rules, and use the combinations of annotation rules to automatically annotate the newly added clinical text data to generate annotation results.
[0138] A third aspect of the embodiments of the present invention,
[0139] An electronic device is provided, comprising:
[0140] processor;
[0141] Memory used to store processor-executable instructions;
[0142] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0143] Fourth aspect of the present invention,
[0144] A computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0145] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligent annotation and quality control of clinical data, characterized in that, include: The system receives clinical text data and constructs a medical knowledge graph, which includes disease nodes, symptom nodes, and semantic relationships between nodes. The nodes in the medical knowledge graph are then vectorized to generate node feature vectors. Semantic recognition is performed on the clinical text data. Contextual features in the clinical text data are extracted through a multi-head attention mechanism to identify medical entities and clinical event information in the clinical text data, obtain temporal dependencies between medical entities, and generate structured semantic representations. Calculate the similarity between the structured semantic representation and the node feature vector, establish semantic links between semantic representations with similarity exceeding a preset similarity threshold and corresponding nodes, and form semantic association paths in the medical knowledge graph; classify the associations based on the strength of the semantic association paths to generate a hierarchical semantic link network. Knowledge reasoning is performed in the hierarchical semantic linking network. Based on the reinforcement learning method, multi-step reasoning is carried out according to the semantic link strength to obtain the implicit association path between nodes; the implicit association path is transformed into clinical data annotation rules. Validation samples are selected from the clinical text data. The clinical data annotation rules are validated and optimized using the validation samples to generate different combinations of annotation rules. The newly added clinical text data is automatically annotated using the combinations of annotation rules to generate annotation results.
2. The method according to claim 1, characterized in that, Semantic recognition is performed on the clinical text data. A multi-head attention mechanism is used to extract contextual features from the clinical text data, identify medical entities and clinical event information within the clinical text data, obtain temporal dependencies between medical entities, and generate structured semantic representations, including: Multi-head attention is calculated on the contextual features of the clinical text data, and a query matrix, key matrix, and value matrix are generated through linear transformation. Attention scores are calculated based on the query matrix, key matrix, and value matrix. The outputs of multiple attention heads are concatenated and subjected to linear transformation to obtain fused attention features. A bidirectional propagation conditional random field is constructed based on the fused attention features. The conditional random field uses a learnable label transition probability matrix to label the boundary positions of medical entities. The labeled boundary positions of medical entities are combined with the fused attention features to obtain the contextual semantic representation of the medical entities. Based on the contextual semantic representation, the types of medical entities are classified to generate medical entity recognition results. Temporal information is extracted from the medical entity identification results, and a temporal dependency graph is constructed by combining the occurrence order of the medical entities; the temporal dependency graph is subjected to temporal consistency verification to eliminate temporal conflicts and generate a temporal relationship network. The medical entity identification results and the temporal relationship network are integrated to generate a structured semantic representation.
3. The method according to claim 1, characterized in that, Semantic links are established between semantic representations whose similarity exceeds a preset similarity threshold and their corresponding nodes, forming semantic association paths in the medical knowledge graph. Based on the strength of the semantic association paths, the association relationships are classified into hierarchical levels to generate a hierarchical semantic link network, including: Semantic representations with similarity exceeding a preset similarity threshold are linked to nodes in the corresponding medical knowledge graph, and the similarity is recorded as the initial strength of the semantic link. Based on the semantic link relationship, a breadth-first search is performed to obtain a set of adjacent nodes; the Euclidean distance between the nodes in the set of adjacent nodes is calculated sequentially; a semantic association path is constructed based on the Euclidean distance between the nodes, and the semantic association path contains a sequence of nodes from the source node to the target node; The path score of the semantically associated path is calculated by multiplying the Euclidean distances between adjacent nodes on the path and introducing a path length penalty factor. The semantic association path is divided into multi-level association paths based on the numerical range of the initial intensity, and the multi-level association paths are hierarchically integrated to construct a hierarchical semantic link network.
4. The method according to claim 1, characterized in that, Knowledge reasoning is performed in the hierarchical semantic linking network. Based on reinforcement learning, multi-step reasoning is conducted using semantic link strength to obtain the implicit association paths between nodes, including: A knowledge reasoning environment is constructed by combining the current node information, historical reasoning path information, and target node information to form a state vector, and an action space is constructed based on the set of adjacent nodes of the current node; a combined reward function is constructed based on the semantic link strength and the similarity of the target node. The state vector and action space are subjected to depth Q-value calculation to obtain the state-action value; the strategy is iterated based on the state-action value and the combined reward function to obtain the optimized reasoning strategy; the optimized reasoning strategy is used as the decision basis for knowledge reasoning. Multi-step reasoning is performed in the medical knowledge graph, and the state-action values obtained from each step of reasoning are multiplied together to generate candidate association paths; the semantic consistency of the candidate association paths is verified, and implicit association paths that meet the verification conditions are selected.
5. The method according to claim 4, characterized in that, Transforming the implicit association paths into clinical data annotation rules includes: The implicit association path is mapped to a rule template, which includes a method for converting path node features to predicates and a mapping method for annotation operations; Based on the rule template, path features are extracted to generate the preconditions and execution actions for annotation rules; the reliability score of the implicit associated path is used as the confidence level of the annotation rule. The annotation rules are simplified and similar rules are merged to generate an optimized set of annotation rules, which serves as the clinical data annotation rules that meet the evaluation criteria.
6. The method according to claim 1, characterized in that, The clinical data annotation rules are validated and optimized using the validation samples to generate different combinations of annotation rules. The automatic annotation of newly added clinical text data using these combinations of annotation rules includes: The system receives a validation sample dataset and clinical data annotation rules to be validated; extracts text features and standard annotation information from the validation sample dataset; annotates the validation sample dataset according to the clinical data annotation rules to be validated, and obtains the rule annotation results; compares the rule annotation results with the standard annotation information, and calculates the annotation precision and recall. Based on the annotation accuracy and recall, a rule evaluation index is constructed to quantitatively assess the annotation capability of the clinical data annotation rules; an evaluation index threshold is set, and the optimal rule combination that meets the evaluation index threshold is selected. New clinical text data is input into the optimal rule combination, and the new clinical text data is matched based on the rule annotation conditions and execution actions; the annotation execution order is determined according to the confidence and priority of the rules, and automatic annotation results are generated.
7. A clinical data intelligent annotation and quality control system, used to implement the method as described in any one of claims 1-6, characterized in that, include: The first unit is used to receive clinical text data, construct a medical knowledge graph, which includes disease nodes, symptom nodes, and semantic relationships between nodes; and to vectorize the nodes in the medical knowledge graph to generate node feature vectors. The second unit is used to perform semantic recognition on the clinical text data, extract contextual features from the clinical text data through a multi-head attention mechanism, identify medical entities and clinical event information in the clinical text data, obtain temporal dependencies between medical entities, and generate structured semantic representations. The third unit is used to calculate the similarity between the structured semantic representation and the node feature vector, establish semantic links between semantic representations with similarity exceeding a preset similarity threshold and corresponding nodes, and form semantic association paths in the medical knowledge graph; classify the associations based on the strength of the semantic association paths to generate a hierarchical semantic link network. The fourth unit is used to perform knowledge reasoning in the hierarchical semantic linking network. Based on the reinforcement learning method, it performs multi-step reasoning based on the semantic linking strength to obtain the implicit association paths between nodes; and transforms the implicit association paths into clinical data annotation rules. The fifth unit is used to select validation samples from the clinical text data, use the validation samples to validate and optimize the clinical data annotation rules, generate different combinations of annotation rules, and use the combinations of annotation rules to automatically annotate the newly added clinical text data to generate annotation results.
8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.
Citation Information
Cited By
Corpus construction method and system
CN121122547A
Patient screening method and system based on artificial intelligence in clinical research
CN121938657A
Artificial intelligence-based patient screening method and system in clinical research
CN121938657B