Clinical research cohort screening and construction method and system based on intelligent labeling

CN121790022BActive Publication Date: 2026-09-11NEWLINK TECH INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610271299.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-06
Publication Date
2026-09-11
Estimated Expiration
2046-03-06

AI Technical Summary

Benefits of technology

[0046] By constructing a mapping network based on a multi-source knowledge base to standardize terminology, the problem of inconsistent terminology in multi-source heterogeneous medical data is solved, improving the accuracy and efficiency of data integration. Terminology ambiguity constraints and temporal feature verification are used to perform temporal alignment and feature fusion of medical data, effectively solving common problems of temporal disorder and feature conflict in clinical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121790022B_ABST
    Figure CN121790022B_ABST
Patent Text Reader

Abstract

The application provides a clinical research cohort screening and construction method and system based on intelligent labeling, relates to the technical field of medical data analysis and clinical research, and comprises the following steps: performing term standardization, time sequence alignment and feature fusion on multi-source heterogeneous medical data, constructing a knowledge graph, integrating multi-scale constraints, realizing sample typing by using graph reasoning and spectral clustering, identifying and labeling abnormalities by means of medical ontology and statistical dual-space projection, tracing conflict points, generating correction suggestions to optimize the cohort structure, and realizing the intelligentization, standardization and high accuracy of clinical cohort construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data analysis and clinical research technology, and in particular to a method and system for screening and constructing clinical research cohorts based on intelligent annotation. Background Technology

[0002] The selection and construction of clinical research cohorts is a core component of medical research, playing a crucial role in disease mechanism research, treatment protocol evaluation, and the implementation of precision medicine. Traditional clinical cohort construction primarily relies on manual screening and annotation, where medical staff review patients' electronic medical records, laboratory reports, and medical imaging data to determine research subjects based on inclusion and exclusion criteria. With the rapid increase in the volume and diversification of medical data, large-scale clinical studies have placed higher demands on the efficiency and accuracy of cohort construction. In recent years, with the development of artificial intelligence technology, intelligent screening methods based on natural language processing and machine learning have been increasingly applied to clinical research cohort construction. These methods automate the analysis of medical data, assisting researchers in completing sample screening more efficiently. Summary of the Invention

[0003] This invention provides a method and system for screening and constructing clinical research cohorts based on intelligent annotation, which can solve the problems in the prior art.

[0004] A first aspect of this invention provides a method for screening and constructing clinical research cohorts based on intelligent annotation, comprising:

[0005] Acquire multi-source heterogeneous medical data, including structured medical records and text records;

[0006] A mapping network is constructed based on a multi-source knowledge base to achieve terminology standardization. Through terminology ambiguity constraints and temporal feature verification, the medical data is temporally aligned and feature fused to obtain a candidate sample set.

[0007] The clinical events in the candidate sample set are transformed into entity nodes, and directed relation edges are established. Multi-scale temporal constraints and spatial topological constraints are embedded to obtain a knowledge graph. Based on the knowledge graph, implicit evolutionary patterns and abnormal evolutionary paths are identified through graph reasoning algorithms, and the candidate sample set is subjected to pattern classification and spectral clustering to obtain multiple sample clusters.

[0008] Samples are allocated and intelligently labeled according to sample clusters and clinical research stratification requirements to obtain the labeling results of each sample and form a preliminary cohort structure.

[0009] The annotation results of each sample are simultaneously projected onto the clinical semantic space constructed based on medical ontology relationships and the statistical space constructed based on sample feature distribution, and the spatial deviation of each sample is calculated and the annotated abnormal samples are identified.

[0010] By tracing the evolution path of the labeled abnormal samples in the knowledge graph, identifying the feature conflict points that cause projection deviation, generating correction suggestions based on the iteration direction of the conflict points in the dual space, and adjusting the sample labels in the initial queue structure according to the correction suggestions, a clinical research queue is obtained.

[0011] Terminology standardization is achieved by constructing a mapping network based on a multi-source knowledge base. Through terminology ambiguity constraints and temporal feature verification, the medical data is temporally aligned and feature fused to obtain a candidate sample set, including:

[0012] Standardization rules for terms are extracted from the medical ontology database, clinical guideline database, and drug knowledge base, respectively. Three independent mapping sub-networks are constructed. The text records in the medical data are transformed in the three sub-networks. When the mapping results are consistent, they are directly adopted. When conflicts occur, the semantic matching confidence of each mapping sub-network is calculated. The result with the highest semantic matching confidence is selected as the standardized term, and the conflict mapping is stored in the term ambiguity database.

[0013] The standardized terms are input into the semantic understanding model, and a penalty term based on the term ambiguity library is introduced into the model loss function. When the extracted feature vector corresponds to a combination of terms and there is a conflict record in the ambiguity library, the loss weight is corrected.

[0014] The structured medical records in the medical data are sorted by timestamp, the time transition probability between events is calculated, the time expression of unstructured text is extracted as time anchors, semantic feature association mapping is constructed, consistency verification of the two time relationships is performed, when contradictions are found, the original records are traced back for manual review and marking, after contradictory samples are removed, the remaining features are aligned and fused by time, and the fused features are standardized and mapped to obtain a candidate sample set with unified semantic representation.

[0015] The clinical events in the candidate sample set are transformed into entity nodes, and directed relation edges are established. Multi-scale temporal constraints and spatial topological constraints are embedded to obtain a knowledge graph, including:

[0016] The clinical events of each candidate sample in the candidate sample set are divided into multiple levels. Within each level, the clinical events are transformed into entity nodes containing attribute features. Cross-level causal directed edges and intra-level temporal directed edges are established to form a two-dimensional graph topology.

[0017] Extract the time interval between adjacent entity nodes, divide the time interval into multiple time scales, establish a mapping relationship between time scales and time constraint functions, limit the reasonable range of time interval between the nodes at both ends of the relationship edge, and embed the time constraint function into the edge attributes to form multi-scale time constraints.

[0018] Calculate the in-degree ratio of each entity node, identify target entity nodes that exceed the centrality threshold, calculate the shortest path length between other entity nodes as the topological distance, establish an inverse mapping from topological distance to constraint strength, embed constraint strength into edge attributes to form spatial topological constraints, combine the relationship edges with embedded double constraints with entity nodes to obtain the knowledge graph.

[0019] Based on the knowledge graph, implicit evolutionary patterns and abnormal evolutionary paths are identified through graph reasoning algorithms, and the candidate sample set is subjected to pattern classification and spectral clustering to obtain multiple sample clusters, including:

[0020] The subgraph structure of each candidate sample is extracted from the knowledge graph and multiple rounds of random walks are performed. Starting from the node with zero in-degree, probability transitions are performed according to the edge weights of the directed relation edges. The node sequence and edge sequence of the walk path are recorded as the sampling path. High-frequency combination patterns are extracted by sequence alignment. Patterns that appear more frequently than the pattern recognition threshold are identified as hidden evolutionary patterns.

[0021] Examine whether there is a dangling starting point with an in-degree greater than zero but no predecessor node, or a dangling ending point with an out-degree greater than zero but no successor node in the evolution path. Search the knowledge graph globally for edges connected to the dangling nodes and determine whether the other end node of the edge is within the sample time span but not included. If it is satisfied, mark the path as an abnormal evolution path with missing nodes.

[0022] Target evolution nodes are selected based on the topological centrality sorting of nodes in the implicit evolution law, the inclusion of non-abnormal evolution paths with them is statistically analyzed, and the evolution structure features are generated by combining the average time interval of the edges in the path. Spectral clustering is performed on the evolution structure features of all candidate samples to obtain multiple sample clusters.

[0023] The annotation results of each sample are simultaneously projected onto a clinical semantic space constructed based on medical ontology relationships and a statistical space constructed based on sample feature distribution. The spatial deviation of each sample is calculated and abnormally annotated samples are identified, including:

[0024] Based on the annotation results of each sample, the corresponding disease category, hierarchical identifier and clinical endpoint event are identified. Based on the disease-symptom-drug ternary relationship in the medical ontology, a semantic relationship graph is constructed. The annotation results of each sample are mapped to coordinate vectors in the semantic relationship graph. The semantic similarity between the annotation results and the standard disease phenotype is calculated, and the projection coordinates of the sample in the clinical semantic space are determined.

[0025] Principal component analysis is performed on the feature vectors of all samples in the preliminary queue structure to extract the basis vectors of the statistical space composed of multiple principal components. The feature vectors corresponding to the labeling results of each sample are linearly projected onto the basis vectors of the statistical space to obtain the projected coordinates of each sample in the statistical space.

[0026] The Euclidean distance between the projected coordinates of the same sample in the clinical semantic space and the statistical space is calculated as the spatial deviation. The spatial deviation distribution is constructed and the mean and standard deviation of the deviation are calculated. The deviation threshold is obtained by weighted combination. Samples whose spatial deviation exceeds the deviation threshold are identified as labeled abnormal samples.

[0027] Tracing the evolution path of the labeled anomalous samples in the knowledge graph, identifying feature conflict points that lead to projection deviation, generating correction suggestions based on the iteration direction of the conflict points in the dual space, and adjusting the sample labeling in the initial cohort structure according to the correction suggestions to obtain the clinical research cohort, including:

[0028] The evolution path of labeled abnormal samples is extracted from the knowledge graph. The attribute features of each entity node are projected onto the clinical semantic space and the statistical space respectively. Nodes whose projection coordinate differences exceed the difference threshold are marked as candidate conflict nodes. A local subgraph centered on the candidate conflict nodes is constructed and the structural entropy difference is calculated. The candidate conflict node with the largest structural entropy difference is selected as the feature conflict point.

[0029] Starting from the projection coordinates of the feature conflict point in the clinical semantic space, and ending with the projection coordinates of the abnormal sample in the statistical space, the connection vector is calculated and iterated in different directions in the two spaces. When the projection distance of the iteration trajectory converges, the convergence point is mapped to the candidate correction category and the hierarchical label adjustment target.

[0030] The standard evolution path corresponding to the candidate correction category is retrieved in the knowledge graph. The similarity with the abnormal sample path is calculated. The category with the highest similarity is selected as the final correction category and the hierarchical label is updated. After the annotation adjustment is performed on all abnormal samples, the distribution balance and category consistency of the cohort are re-evaluated to obtain a clinical research cohort that meets the requirements.

[0031] A second aspect of this invention provides a system for screening and constructing clinical research cohorts based on intelligent annotation, comprising:

[0032] The medical data acquisition unit is used to acquire multi-source heterogeneous medical data, including structured medical records and text records;

[0033] The data standardization and fusion unit is used to construct a mapping network based on a multi-source knowledge base to achieve terminology standardization. Through terminology ambiguity constraints and temporal feature verification, the medical data is time-series aligned and feature-fused to obtain a candidate sample set.

[0034] The knowledge graph construction and clustering unit is used to transform the clinical events of the candidate sample set into entity nodes, establish directed relation edges, embed multi-scale temporal constraints and spatial topological constraints to obtain a knowledge graph; based on the knowledge graph, the implicit evolutionary rules and abnormal evolutionary paths are identified through the graph reasoning algorithm, and the candidate sample set is subjected to pattern classification and spectral clustering to obtain multiple sample clusters;

[0035] The queue construction and intelligent annotation unit is used to allocate samples and perform intelligent annotation according to sample clusters and clinical research stratification requirements, so as to obtain the annotation results of each sample and form a preliminary queue structure.

[0036] The dual spatial projection and anomaly detection unit is used to simultaneously project the annotation results of each sample onto a clinical semantic space constructed based on medical ontology relationships and a statistical space constructed based on sample feature distribution, calculate the spatial deviation of each sample and identify annotated abnormal samples.

[0037] The evolution path tracing and queue optimization unit is used to trace the evolution path of the labeled abnormal samples in the knowledge graph, identify the feature conflict points that cause projection deviation, generate correction suggestions based on the iteration direction of the conflict points in the dual space, and adjust the sample labeling in the initial queue structure according to the correction suggestions to obtain the clinical research queue.

[0038] A third aspect of the embodiments of the present invention,

[0039] An electronic device is provided, comprising:

[0040] processor;

[0041] Memory used to store processor-executable instructions;

[0042] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0043] Fourth aspect of the embodiments of the present invention,

[0044] A computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0045] The beneficial effects of this application are as follows:

[0046] By constructing a mapping network based on a multi-source knowledge base to standardize terminology, the problem of inconsistent terminology in multi-source heterogeneous medical data is solved, improving the accuracy and efficiency of data integration. Terminology ambiguity constraints and temporal feature verification are used to perform temporal alignment and feature fusion of medical data, effectively solving common problems of temporal disorder and feature conflict in clinical data.

[0047] Transforming clinical events into entity nodes of a knowledge graph and embedding multi-scale temporal and spatial topological constraints enables a more comprehensive capture of the complex relationships between clinical events. By identifying implicit evolutionary patterns and abnormal evolutionary paths through graph reasoning algorithms, refined pattern classification and spectral clustering of clinical samples are achieved, improving the scientific rigor of sample grouping.

[0048] By tracing the evolution path of labeled abnormal samples in the knowledge graph, identifying feature conflict points and generating correction suggestions, self-correction and optimization in the cohort construction process are achieved, improving the quality and reliability of the final clinical research cohort, significantly reducing the workload of manual screening and labeling, and improving the efficiency and accuracy of clinical research cohort construction. Attached Figure Description

[0049] Figure 1 This is a flowchart illustrating the method for screening and constructing a clinical research cohort based on intelligent annotation, as described in an embodiment of the present invention.

[0050] Figure 2 This is a flowchart illustrating the method for generating a candidate sample set with unified semantic representation according to an embodiment of the present invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0053] Figure 1 This is a flowchart illustrating the method for screening and constructing a clinical research cohort based on intelligent annotation, as described in an embodiment of the present invention. Figure 1 As shown, the method includes:

[0054] Acquire multi-source heterogeneous medical data, including structured medical records and text records;

[0055] A mapping network is constructed based on a multi-source knowledge base to achieve terminology standardization. Through terminology ambiguity constraints and temporal feature verification, the medical data is temporally aligned and feature fused to obtain a candidate sample set.

[0056] The clinical events in the candidate sample set are transformed into entity nodes, and directed relation edges are established. Multi-scale temporal constraints and spatial topological constraints are embedded to obtain a knowledge graph. Based on the knowledge graph, implicit evolutionary patterns and abnormal evolutionary paths are identified through graph reasoning algorithms, and the candidate sample set is subjected to pattern classification and spectral clustering to obtain multiple sample clusters.

[0057] Samples are allocated and intelligently labeled according to sample clusters and clinical research stratification requirements to obtain the labeling results of each sample and form a preliminary cohort structure.

[0058] The annotation results of each sample are simultaneously projected onto the clinical semantic space constructed based on medical ontology relationships and the statistical space constructed based on sample feature distribution, and the spatial deviation of each sample is calculated and the annotated abnormal samples are identified.

[0059] By tracing the evolution path of the labeled abnormal samples in the knowledge graph, identifying the feature conflict points that cause projection deviation, generating correction suggestions based on the iteration direction of the conflict points in the dual space, and adjusting the sample labels in the initial queue structure according to the correction suggestions, a clinical research queue is obtained.

[0060] Figure 2 This is a flowchart illustrating the method for generating a candidate sample set with unified semantic representation according to an embodiment of the present invention. In one optional implementation, a mapping network is constructed based on a multi-source knowledge base to achieve terminology standardization. Through terminology ambiguity constraints and temporal feature verification, the medical data is temporally aligned and feature fused to obtain a candidate sample set, including:

[0061] Standardization rules for terms are extracted from the medical ontology database, clinical guideline database, and drug knowledge base, respectively. Three independent mapping sub-networks are constructed. The text records in the medical data are transformed in the three sub-networks. When the mapping results are consistent, they are directly adopted. When conflicts occur, the semantic matching confidence of each mapping sub-network is calculated. The result with the highest semantic matching confidence is selected as the standardized term, and the conflict mapping is stored in the term ambiguity database.

[0062] The standardized terms are input into the semantic understanding model, and a penalty term based on the term ambiguity library is introduced into the model loss function. When the extracted feature vector corresponds to a combination of terms and there is a conflict record in the ambiguity library, the loss weight is corrected.

[0063] The structured medical records in the medical data are sorted by timestamp, the time transition probability between events is calculated, the time expression of unstructured text is extracted as time anchors, semantic feature association mapping is constructed, consistency verification of the two time relationships is performed, when contradictions are found, the original records are traced back for manual review and marking, after contradictory samples are removed, the remaining features are aligned and fused by time, and the fused features are standardized and mapped to obtain a candidate sample set with unified semantic representation.

[0064] In this specific embodiment, terminology standardization rules are extracted from a medical ontology database, a clinical guideline database, and a drug knowledge base to construct three independent mapping sub-networks. The medical ontology database mainly includes standardized medical terminology systems such as SNOMED CT and ICD-10, from which synonym mappings, hypernym relationships, and terminology normalization rules are extracted. The clinical guideline database covers standard terminology definitions and their clinical expression variations in various specialty treatment guidelines. The drug knowledge base contains the correspondence between generic drug names, brand names, and common abbreviations. For the medical ontology database, a graph-based mapping model is used to transform the semantic relationships between terminology nodes into distance constraints in a vector space. For the clinical guideline database, a rule-based pattern matching method is used to extract standard terminology expressions. For the drug knowledge base, a combination of character edit distance and transliteration similarity is used to establish mapping relationships.

[0065] For the text records in the input medical data, word segmentation and entity recognition are performed. Mapping transformation is then performed in three pre-constructed sub-networks. For example, for the word "heart attack", the medical ontology database will map it as "myocardial infarction", the clinical guideline database will map it as "acute myocardial infarction", and the drug knowledge base will not generate a mapping result. When the mapping results of the three sub-networks are consistent or partially consistent and there is no conflict, the consistent result is directly adopted as the standardized term. When the mapping results conflict, the semantic matching confidence of each mapping result is calculated.

[0066] The semantic matching confidence score is calculated based on the following formula: for each mapping result M i Calculate the cosine similarity S between the original term and the original term in the word vector space. i Simultaneously consider the weights W of the mapping subnetwork i (Based on domain applicability presets), and the frequency F of this mapping in historical mapping records. i Taking into account these factors, the final confidence level is obtained, and the result with the highest confidence level is selected as the standardized term. The conflict mapping relationship and its contextual information are stored in the term ambiguity library, including the original term, each mapping result, context and the final selection result.

[0067] Standardized terms are input into a pre-trained semantic understanding model (such as the BERT model adapted for the medical field) to extract semantic representation vectors of the text. During model training, a penalty term based on a term ambiguity corpus is introduced to modify the loss function. When the term combination corresponding to the feature vector extracted by the model has conflicting records in the ambiguity corpus, the corresponding loss weight is increased according to the severity of the conflict. Specifically, for a term combination represented by feature vector V, if there is a set C of conflicting records in the ambiguity corpus, the modified loss function is supplemented with a penalty term λ·∑(similarity(V, C)). i)), where λ is the balance factor and similarity is the vector similarity metric function. This penalty mechanism makes the model more cautious when dealing with ambiguous terms, thereby improving the accuracy of semantic understanding.

[0068] Structured medical records in the medical data are sorted by timestamps to construct a time-series event chain. For each type of medical event (such as examination, medication, surgery, etc.), the time-series transition probability between events is calculated to form a Markov transition matrix. For example, for a "diabetic" patient, the event sequence is "blood glucose test → oral hypoglycemic drug → insulin injection". The corresponding time-series transition probability can be used for subsequent time-series consistency verification.

[0069] Meanwhile, time expressions in unstructured text are extracted and transformed into standardized time anchors. The time expression extraction adopts a hybrid method combining rules and deep learning to identify absolute time expressions (such as "January 5, 2023") and relative time expressions (such as "three days ago" and "last month"), and map them to a unified time coordinate system.

[0070] Based on the extracted time anchors and the timestamps of the structured records, a semantic feature association mapping is constructed to establish an association between events mentioned in the unstructured text and corresponding events in the structured records. The temporal relationship is verified by calculating the consistency index of the temporal relationship between the two sources. When a temporal contradiction is found, such as the unstructured text mentioning "the patient started taking drug A last year" while the structured record shows that the first use of drug A was three months ago, the system will mark the contradiction and trace back to the original medical records for manual review.

[0071] After manual review, samples with temporal inconsistencies are removed. The remaining features are aligned and fused along the timeline. For features from different sources at the same time point, a weighted fusion strategy is adopted, with the weights dynamically adjusted based on data reliability. Finally, the fused features are standardized and mapped to a unified semantic representation, forming a candidate sample set that provides high-quality basic data for subsequent medical data analysis and applications.

[0072] In one optional implementation, the clinical events of the candidate sample set are transformed into entity nodes, and directed relational edges are established, embedding multi-scale temporal constraints and spatial topological constraints to obtain a knowledge graph, including:

[0073] The clinical events of each candidate sample in the candidate sample set are divided into multiple levels. Within each level, the clinical events are transformed into entity nodes containing attribute features. Cross-level causal directed edges and intra-level temporal directed edges are established to form a two-dimensional graph topology.

[0074] Extract the time interval between adjacent entity nodes, divide the time interval into multiple time scales, establish a mapping relationship between time scales and time constraint functions, limit the reasonable range of time interval between the nodes at both ends of the relationship edge, and embed the time constraint function into the edge attributes to form multi-scale time constraints.

[0075] Calculate the in-degree ratio of each entity node, identify target entity nodes that exceed the centrality threshold, calculate the shortest path length between other entity nodes as the topological distance, establish an inverse mapping from topological distance to constraint strength, embed constraint strength into edge attributes to form spatial topological constraints, combine the relationship edges with embedded double constraints with entity nodes to obtain the knowledge graph.

[0076] In this specific embodiment, the clinical events in the candidate sample set are divided into multiple levels, such as a patient basic information layer, a disease diagnosis layer, a treatment intervention layer, and an examination result layer. In the patient basic information layer, information such as age, gender, and medical history are converted into attribute features. In the disease diagnosis layer, information such as disease name, diagnosis time, and severity are converted into attribute features. In the treatment intervention layer, information such as drug name, dosage, and treatment plan are converted into attribute features. In the examination result layer, information such as examination items, examination results, and examination time are converted into attribute features. In this way, each clinical event is transformed into an entity node containing the corresponding attribute features.

[0077] Next, we establish causal directed edges across levels, such as pointing from the "Type 2 Diabetes" entity node in the disease diagnosis layer to the "Insulin Injection" entity node in the treatment intervention layer, indicating that the disease leads to the corresponding treatment measures. At the same time, we establish temporal directed edges within the level, such as pointing from the "Initial Insulin Injection" entity node to the "Insulin Dosage Adjustment" entity node in the treatment intervention layer, indicating the chronological order of treatment. In this way, we form a two-dimensional graph topology structure that includes causal and temporal relationships.

[0078] To embed multi-scale time constraints, the time intervals between adjacent entity nodes are extracted. For example, the time interval from "hypertension diagnosis" to "use of antihypertensive medication" is three days, and the time interval from "use of antihypertensive medication" to "blood pressure follow-up" is two weeks. The time intervals are divided into multiple time scales, such as short-term (less than one week), medium-term (one week to one month), and long-term (more than one month). For different time scales, corresponding time constraint functions are established. For the short-term time scale, the time constraint function can be set as a strong constraint, requiring the time interval to be within a specific range; for the medium-term time scale, a medium constraint can be set, allowing a certain range of time deviation; for the long-term time scale, a weak constraint can be set, allowing a larger range of time deviation. These time constraint functions are embedded into the attributes of directed relation edges to form multi-scale time constraints.

[0079] To implement spatial topological constraints, the in-degree ratio of each entity node in the graph is calculated. The in-degree ratio equals the ratio of a node's out-degree to its in-degree, reflecting the information flow characteristics of the node in the graph. A centrality threshold is set; nodes with an in-degree ratio greater than 1.5 or less than 0.5 are considered target entity nodes. For example, the node "myocardial infarction" has an out-degree of 8 and an in-degree of 2, resulting in an in-degree ratio of 4, exceeding the centrality threshold and thus being identified as a target entity node. The shortest path length between these target entity nodes and other entity nodes in the graph is calculated and used as the topological distance. An inverse mapping relationship is established between the topological distance and the constraint strength; that is, the smaller the topological distance, the stronger the constraint. For example, when the topological distance is 1, the constraint strength is 1; when the topological distance is 2, the constraint strength is 0.5; and when the topological distance is 3, the constraint strength is 0.33. These constraint strength values ​​are embedded into the attributes of the corresponding relation edges to form spatial topological constraints. By combining directed relation edges embedded with multi-scale temporal constraints and spatial topological constraints with entity nodes, a complete knowledge graph is obtained. This knowledge graph not only contains entity information and relational structures of clinical events, but also incorporates constraint information in the temporal and spatial dimensions, which can more accurately express the complex relationships between clinical events.

[0080] In practical applications, such as analyzing the diagnosis and treatment process of diabetic patients, this knowledge graph can clearly show the time pattern from the initial diagnosis to the occurrence of various complications, as well as the impact path of different treatment plans on disease progression. This information is of great reference value for doctors to formulate personalized treatment plans and predict disease development trends.

[0081] The knowledge graph constructed using the above methods can effectively capture the temporal sequence relationships and topological features among clinical events, providing a solid foundation for subsequent medical decision support and clinical prediction.

[0082] In one optional implementation, based on the knowledge graph, implicit evolutionary patterns and abnormal evolutionary paths are identified through a graph reasoning algorithm, and the candidate sample set is subjected to pattern classification and spectral clustering to obtain multiple sample clusters, including:

[0083] The subgraph structure of each candidate sample is extracted from the knowledge graph and multiple rounds of random walks are performed. Starting from the node with zero in-degree, probability transitions are performed according to the edge weights of the directed relation edges. The node sequence and edge sequence of the walk path are recorded as the sampling path. High-frequency combination patterns are extracted by sequence alignment. Patterns that appear more frequently than the pattern recognition threshold are identified as hidden evolutionary patterns.

[0084] Examine whether there is a dangling starting point with an in-degree greater than zero but no predecessor node, or a dangling ending point with an out-degree greater than zero but no successor node in the evolution path. Search the knowledge graph globally for edges connected to the dangling nodes and determine whether the other end node of the edge is within the sample time span but not included. If it is satisfied, mark the path as an abnormal evolution path with missing nodes.

[0085] Target evolution nodes are selected based on the topological centrality sorting of nodes in the implicit evolution law, the inclusion of non-abnormal evolution paths with them is statistically analyzed, and the evolution structure features are generated by combining the average time interval of the edges in the path. Spectral clustering is performed on the evolution structure features of all candidate samples to obtain multiple sample clusters.

[0086] In this specific embodiment, it is necessary to extract the subgraph structure of each candidate sample in the knowledge graph and perform multiple rounds of random walks to discover the hidden evolutionary patterns. Specifically, starting from the node with an in-degree of zero, these nodes usually represent the initial point of an event or state. During the random walk, the probability transition is carried out according to the weight of the directed relation edges. That is, the higher the edge weight, the greater the transition probability. For example, if starting from node A, there are edges pointing to nodes B and C with weights of 0.7 and 0.3 respectively, then the probability of walking to B in the next step is 70%, and the probability of walking to C is 30%.

[0087] Each walk records the node sequence and edge sequence as a sampling path. For example, a sampling path might be: Node A - (Relation r1) -> Node B - (Relation r2) -> Node C. After a preset number of walks (e.g., 1000), sequence alignment analysis is performed on all sampling paths to extract high-frequency combination patterns. The sequence alignment uses a local sequence alignment algorithm to identify common subsequences in the paths. If a combination pattern (e.g., node sequence "ABC" or relation sequence "r1-r2") appears more frequently than a preset pattern recognition threshold (e.g., 30%) in all sampling paths, then that pattern is identified as a hidden evolutionary pattern.

[0088] Each evolution path is examined for anomalies, which are mainly of two types: first, there is a dangling starting point with an in-degree greater than zero but no predecessor node; second, there is a dangling ending point with an out-degree greater than zero but no successor node. For each extracted evolution path, the predecessor and successor nodes of each node are checked to see if they are complete. If a dangling node is found, the edges connected to the dangling node are searched in the global scope of the knowledge graph, and it is determined whether the other end node of the edge is within the sample time span but is not included in the current evolution path.

[0089] For example, if node P in the evolution path has an edge pointing outward, but the node Q connected by that edge does not appear in the path, and the time attribute of node Q is within the time span of the sample (i.e. should be included), then the path is marked as an abnormal evolution path with "missing anomaly". This anomaly usually reflects incomplete data collection or that there are branches in the evolution path that have not been captured.

[0090] Based on the identified implicit evolutionary patterns, target evolutionary nodes are selected for subsequent analysis. The selection criteria are based on the centrality of nodes in the topological structure, including degree centrality, betweenness centrality, and eigenvector centrality. All nodes are sorted from high to low according to their centrality index, and the top-ranked nodes (e.g., the top 20%) are selected as target evolutionary nodes.

[0091] For non-abnormal evolution paths, the inclusion of target evolution nodes is statistically analyzed. Specifically, the proportion of target nodes included in each path and the average time interval between adjacent nodes in the path are calculated. These statistical features are combined into an evolution structure feature vector, with each vector containing two parts: the target node inclusion rate and the time interval.

[0092] Spectral clustering is performed on the evolutionary structural features of all candidate samples to construct a similarity matrix between samples. A Gaussian kernel function is used to calculate the similarity between eigenvectors. The similarity matrix is ​​then normalized, and its Laplacian matrix is ​​calculated. Eigenvectors from the Laplacian matrix are extracted, and the k smallest eigenvectors (where k is the expected number of clusters) are selected to form a new feature matrix. The K-means algorithm is applied to the row vectors of the feature matrix to divide the samples into k clusters.

[0093] In practical applications, spectral clustering results can be used to identify sample groups with similar evolutionary patterns. For example, in disease evolution analysis, it may be found that a certain group of patients has similar disease development paths; in enterprise development research, it may be possible to identify groups of enterprises with similar growth trajectories. These classification results help to develop differentiated strategies for different types of samples, improving the accuracy of decision-making.

[0094] The above methods enable the mining of sample evolution patterns, identification of abnormal paths, and sample clustering based on knowledge graphs, providing a foundation for subsequent predictive analysis and decision support.

[0095] In one optional implementation, the annotation results of each sample are simultaneously projected onto a clinical semantic space constructed based on medical ontology relationships and a statistical space constructed based on sample feature distributions. The spatial deviation of each sample is calculated, and abnormally annotated samples are identified, including:

[0096] Based on the annotation results of each sample, the corresponding disease category, hierarchical identifier and clinical endpoint event are identified. Based on the disease-symptom-drug ternary relationship in the medical ontology, a semantic relationship graph is constructed. The annotation results of each sample are mapped to coordinate vectors in the semantic relationship graph. The semantic similarity between the annotation results and the standard disease phenotype is calculated, and the projection coordinates of the sample in the clinical semantic space are determined.

[0097] Principal component analysis is performed on the feature vectors of all samples in the preliminary queue structure to extract the basis vectors of the statistical space composed of multiple principal components. The feature vectors corresponding to the labeling results of each sample are linearly projected onto the basis vectors of the statistical space to obtain the projected coordinates of each sample in the statistical space.

[0098] The Euclidean distance between the projected coordinates of the same sample in the clinical semantic space and the statistical space is calculated as the spatial deviation. The spatial deviation distribution is constructed and the mean and standard deviation of the deviation are calculated. The deviation threshold is obtained by weighted combination. Samples whose spatial deviation exceeds the deviation threshold are identified as labeled abnormal samples.

[0099] In this specific embodiment, for each sample cluster, samples are allocated according to the stratification requirements of the clinical study. These stratification requirements typically include age group, gender ratio, disease severity distribution, and comorbidities. For each sample cluster, the distribution of various stratification indicators within its samples is calculated, and a match assessment is performed against the stratification requirements of the study design. The match assessment uses a weighted Euclidean distance, with the weights of each stratification indicator set according to the clinical study objectives.

[0100] Based on the matching degree assessment results, the samples are intelligently labeled. In the process of intelligent labeling, it is necessary to determine the dominant disease category of each cluster. By analyzing the clinical manifestations, biochemical indicators and imaging characteristics of samples within the cluster, and combining them with medical knowledge graphs, disease inference is made. For each sample, based on the dominant disease category of its cluster and combined with individual characteristic differences, personalized disease category labeling is performed. At the same time, based on the correlation between sample characteristics and clinical stratification indicators, stratification labels are assigned to each sample, such as disease severity grade, treatment response type, etc.

[0101] In the intelligent labeling process, a combination of rule-based and machine learning methods is used for disease category labeling. A rule base is constructed based on medical diagnostic standards, which includes diagnostic conditions and exclusion criteria for various diseases. A support vector machine classifier is trained, with sample features as input and disease category probability distribution as output. The result of rule inference and classifier prediction is combined and fused through a Bayesian network to obtain the final disease category label.

[0102] For the labeling of hierarchical labels, a multi-label classification method is adopted, considering the dependencies between labels, constructing a conditional random field model to capture the association constraints between different hierarchical labels, such as the association pattern between certain severity levels and specific comorbidities. Through the inference of the conditional random field, a set of mutually coordinated hierarchical labels is generated for each sample.

[0103] For the prediction and labeling of clinical endpoint events, a survival analysis model is constructed. Based on the historical data of the samples, a Cox proportional hazards model is trained to predict the risk probability and expected time of a specific clinical endpoint event (such as disease progression, relapse, death, etc.) in the samples. At the same time, through a competitive risk model, the competitive relationship between multiple possible endpoint events is considered, and the most likely endpoint event type and its occurrence risk are labeled for each sample.

[0104] After intelligent annotation is completed, all annotation information is integrated to form a complete annotation profile for each sample, constructing a preliminary cohort structure. This cohort structure includes basic information for each sample, disease category annotation, stratification identifier, and predicted clinical endpoint events. To validate the rationality of the cohort structure, a cohort balance assessment is performed. The differences in feature distributions between different stratified groups are calculated to ensure balanced distribution across groups on key confounding factors.

[0105] Based on the annotation results of each sample in the preliminary queue structure, a clinical semantic space is constructed. Specifically, disease category information, such as "type 2 diabetes" and "hypertension", is extracted from the sample annotations; disease stratification markers, such as "early", "late", and "moderate", are identified; clinical endpoint events, such as "diabetic nephropathy" and "cardiovascular events", are determined; and a pre-established medical ontology database is accessed, which stores a ternary relationship network of disease-symptom-drug. For example, "type 2 diabetes" is associated with symptoms such as "thirst" and "polyuria", and with drugs such as "insulin" and "metformin". Based on these relationships, a semantic relationship graph G=(V, E) is constructed, where node V represents disease, symptom, or drug entities, and edge E represents the strength of the association between them.

[0106] For each sample's annotation result, it is mapped to a coordinate vector v in the semantic relation graph. i The specific method involves activating corresponding nodes in the semantic graph based on information such as diseases, symptoms, and medications mentioned in the sample annotations. This activation pattern is then transformed into a low-dimensional vector representation using a graph embedding algorithm (such as DeepWalk or Node2Vec). Furthermore, the semantic similarity s between the sample annotation results and the standard disease phenotype is calculated. i Similarity can be measured using metrics such as Jaccard similarity coefficient or cosine similarity, and the projected coordinates L of the sample in the clinical semantic space can be used. i From vector v i and similarity s i To be determined jointly.

[0107] A statistical space based on sample feature distribution is constructed. Feature vectors of all samples are extracted from the initial cohort. These features may include demographic characteristics (such as age and gender), laboratory test results (such as blood glucose and blood pressure), and imaging features. Principal component analysis is performed on these feature vectors to extract the principal components that can explain the data variation as the basis vectors of the statistical space {PC1, PC2, ..., PC...}. k}, where k is generally chosen as the smallest principal component number that can explain at least 85% of the variation in the data, and for each sample's feature vector x i The coordinates S in the statistical space are obtained by calculating its projection onto each principal component. i = (x i PC1, x i PC2, ..., x i PC k ).

[0108] Calculate the spatial deviation of the projected coordinates of each sample in the two spaces. For sample i, calculate its coordinates L in the clinical semantic space. i With coordinates S in statistical space i The Euclidean distance D between them i = ||L i - S i To ensure comparability of coordinates between two spaces, it is necessary to normalize the coordinates beforehand and collect the spatial deviations of all samples {D1, D2, ..., D...}. n Construct the deviation distribution and calculate the mean deviation μ and standard deviation σ.

[0109] Determine the deviation threshold T = μ + w·σ, where w is a weighting parameter that can be adjusted according to the specific application scenario, generally ranging from 2 to 3. For spatial deviation D... i Samples exceeding the threshold T are identified as anomaly samples. For example, in a tumor staging study, a sample may be labeled as "early-stage tumor" but its clinical features show multiple metastatic lesions. In this case, the projection of the sample in the two spaces will be significantly different, and it will be identified as anomaly.

[0110] For the identified anomaly samples, the reasons for the anomalies can be further analyzed. They may be labeling errors, such as mislabeling a stage III tumor as a stage I tumor; or they may be real medical anomalies, such as diseases with atypical manifestations. In addition, the deviation contribution of each sample in each dimension can be calculated to identify the key factors that cause the anomalies and provide guidance for subsequent labeling corrections.

[0111] This dual-spatial projection method can effectively identify labeled anomalous samples in medical data, improve data quality, and provide a more reliable foundation for subsequent medical research and model building.

[0112] In one optional implementation, the evolution path of the labeled anomalous samples in the knowledge graph is traced to identify feature conflict points that cause projection deviation. Correction suggestions are generated based on the iteration direction of the conflict points in the dual space. The sample labeling in the initial cohort structure is adjusted according to the correction suggestions to obtain a clinical research cohort, including:

[0113] The evolution path of labeled abnormal samples is extracted from the knowledge graph. The attribute features of each entity node are projected onto the clinical semantic space and the statistical space respectively. Nodes whose projection coordinate differences exceed the difference threshold are marked as candidate conflict nodes. A local subgraph centered on the candidate conflict nodes is constructed and the structural entropy difference is calculated. The candidate conflict node with the largest structural entropy difference is selected as the feature conflict point.

[0114] Starting from the projection coordinates of the feature conflict point in the clinical semantic space, and ending with the projection coordinates of the abnormal sample in the statistical space, the connection vector is calculated and iterated in different directions in the two spaces. When the projection distance of the iteration trajectory converges, the convergence point is mapped to the candidate correction category and the hierarchical label adjustment target.

[0115] The standard evolution path corresponding to the candidate correction category is retrieved in the knowledge graph. The similarity with the abnormal sample path is calculated. The category with the highest similarity is selected as the final correction category and the hierarchical label is updated. After the annotation adjustment is performed on all abnormal samples, the distribution balance and category consistency of the cohort are re-evaluated to obtain a clinical research cohort that meets the requirements.

[0116] In this specific embodiment, tracing the evolution path of annotated abnormal samples in the knowledge graph requires obtaining the complete knowledge graph structure. The knowledge graph contains entity nodes such as diseases, symptoms, diagnoses, and treatments, as well as the edges between them. Each entity node contains multi-dimensional attribute features, which can be projected onto different semantic spaces for analysis.

[0117] The evolution path of extracting labeled anomalous samples from a knowledge graph is achieved by tracing back the connection relationship of the samples in the graph. For each labeled anomalous sample, the corresponding node in the knowledge graph is located by its unique identifier, and its complete evolution path is traced along the incoming and outgoing edges. The evolution path usually contains key nodes of disease progression, such as initial symptoms, intermediate diagnosis and final classification.

[0118] The extracted attribute features of each entity node in the evolution path are projected into the clinical semantic space and the statistical space, respectively. The clinical semantic space is a high-dimensional space constructed based on the knowledge of medical experts, reflecting the professional understanding of disease classification; the statistical space is constructed based on the statistical distribution characteristics of large-scale clinical data. The projection method adopts semantic embedding technology, which transforms the node attribute vectors into the target space through a pre-trained mapping matrix.

[0119] When marking nodes whose projected coordinate differences exceed a threshold as candidate conflict nodes, the Euclidean distance between the projected coordinates of each node in the two spaces is calculated. The difference threshold is usually set as the mean of the projection differences of all nodes plus twice the standard deviation. If the projection difference of a node exceeds this threshold, it is marked as a candidate conflict node. These candidate conflict nodes represent locations where there are inconsistent annotations.

[0120] When constructing a local subgraph centered on a candidate conflict node, the candidate conflict node is selected as the center, and two to three layers of relationships are extended outward to obtain a local subgraph containing related nodes and edges. The structural entropy is calculated for each local subgraph. The structural entropy reflects the complexity and information content of the subgraph. The difference in structural entropy between the local subgraphs in the clinical semantic space and the statistical space is calculated, and the candidate conflict node with the largest difference in structural entropy is selected as the feature conflict point.

[0121] Starting from the projected coordinates of the feature conflict point in the clinical semantic space and ending from the projected coordinates of the labeled abnormal sample in the statistical space, a connection vector is calculated. The connection vector represents the offset direction from clinical understanding to statistical distribution. This offset reveals the essence of the labeled abnormality. Iterative search is performed in the two spaces along different directions. In the clinical semantic space, it iterates along the disease evolution direction defined by expert knowledge, while in the statistical space, it iterates along the direction of data distribution density gradient.

[0122] When the projection distance of the iterative trajectory converges, that is, the projection distance between the search points in the two spaces no longer decreases significantly, the convergence point is mapped to the candidate correction category and the hierarchical label adjustment target. The node category corresponding to the convergence point in the semantic space of the knowledge graph is the candidate correction category, and its hierarchical information is used to determine the adjustment direction of the hierarchical label.

[0123] The standard evolution paths corresponding to candidate corrected categories are retrieved from the knowledge graph. These standard paths are typical disease development trajectories summarized from a large number of correctly classified samples. The similarity between the actual path of the labeled abnormal sample and each standard evolution path is calculated. The similarity calculation uses a weighted combination of path edit distance and node attribute similarity. The category with the highest similarity is selected as the final corrected category, and the hierarchical label is updated accordingly.

[0124] In an application example, an analysis of a group of abnormally labeled lung disease samples revealed that, using the method described above, the evolution path of some samples labeled "pneumonia" in the knowledge graph was closer to the standard path of "interstitial lung disease." The feature conflict point appeared on the imaging feature node. In the clinical semantic space, this node pointed to "ground-glass opacity," while in the statistical space, it leaned towards "honeycomb-like changes." After iterative correction, the category label of these samples was changed from "pneumonia" to "interstitial lung disease," and the hierarchical label was also changed from "acute inflammation" to "chronic fibrosis," making the cohort structure more consistent with clinical reality.

[0125] After adjusting the labels of all abnormal samples, the distribution balance and class consistency of the cohort need to be re-evaluated. Distribution balance is assessed by the variance coefficient of the number of samples in each class, while class consistency is assessed by calculating the clustering metric of the characteristics of samples within the same class. When both of these indicators meet the preset threshold requirements, a clinical research cohort that meets the requirements is considered to have been obtained and can be used for subsequent clinical research and analysis.

[0126] This invention provides a clinical research cohort screening and construction system based on intelligent annotation, comprising:

[0127] The medical data acquisition unit is used to acquire multi-source heterogeneous medical data, including structured medical records and text records;

[0128] The data standardization and fusion unit is used to construct a mapping network based on a multi-source knowledge base to achieve terminology standardization. Through terminology ambiguity constraints and temporal feature verification, the medical data is time-series aligned and feature-fused to obtain a candidate sample set.

[0129] The knowledge graph construction and clustering unit is used to transform the clinical events of the candidate sample set into entity nodes, establish directed relation edges, embed multi-scale temporal constraints and spatial topological constraints to obtain a knowledge graph; based on the knowledge graph, the implicit evolutionary rules and abnormal evolutionary paths are identified through the graph reasoning algorithm, and the candidate sample set is subjected to pattern classification and spectral clustering to obtain multiple sample clusters;

[0130] The queue construction and intelligent annotation unit is used to allocate samples and perform intelligent annotation according to sample clusters and clinical research stratification requirements, so as to obtain the annotation results of each sample and form a preliminary queue structure.

[0131] The dual spatial projection and anomaly detection unit is used to simultaneously project the annotation results of each sample onto a clinical semantic space constructed based on medical ontology relationships and a statistical space constructed based on sample feature distribution, calculate the spatial deviation of each sample and identify annotated abnormal samples.

[0132] The evolution path tracing and queue optimization unit is used to trace the evolution path of the labeled abnormal samples in the knowledge graph, identify the feature conflict points that cause projection deviation, generate correction suggestions based on the iteration direction of the conflict points in the dual space, and adjust the sample labeling in the initial queue structure according to the correction suggestions to obtain the clinical research queue.

[0133] A third aspect of the present invention provides an electronic device, comprising:

[0134] processor;

[0135] Memory used to store processor-executable instructions;

[0136] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0137] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0138] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for clinical study cohort screening and construction based on intelligent labeling, characterized in that, include: Acquire multi-source heterogeneous medical data, including structured medical records and text records; A mapping network is constructed based on a multi-source knowledge base to achieve terminology standardization. Through terminology ambiguity constraints and temporal feature verification, the medical data is temporally aligned and feature fused to obtain a candidate sample set. The clinical events in the candidate sample set are transformed into entity nodes, and directed relation edges are established, embedding multi-scale temporal constraints and spatial topological constraints to obtain a knowledge graph. Based on the knowledge graph, implicit evolutionary patterns and abnormal evolutionary paths are identified through graph reasoning algorithms, and pattern classification and spectral clustering are performed on the candidate sample set to obtain multiple sample clusters, including: The subgraph structure of each candidate sample is extracted from the knowledge graph and multiple rounds of random walks are performed. Starting from the node with zero in-degree, probability transitions are performed according to the edge weights of the directed relation edges. The node sequence and edge sequence of the walk path are recorded as the sampling path. High-frequency combination patterns are extracted by sequence alignment. Patterns that appear more frequently than the pattern recognition threshold are identified as hidden evolutionary patterns. Examine whether there is a dangling starting point with an in-degree greater than zero but no predecessor node, or a dangling ending point with an out-degree greater than zero but no successor node in the evolution path. Search the knowledge graph globally for edges connected to the dangling nodes and determine whether the other end node of the edge is within the sample time span but not included. If it is satisfied, mark the path as an abnormal evolution path with missing nodes. Target evolution nodes are selected based on the topological centrality sorting of nodes in the implicit evolution law, the inclusion of non-abnormal evolution paths with them is statistically analyzed, and the evolution structure features are generated by combining the average time interval of the edges in the path. Spectral clustering is performed on the evolution structure features of all candidate samples to obtain multiple sample clusters. Samples are allocated and intelligently labeled according to sample clusters and clinical research stratification requirements to obtain the labeling results of each sample and form a preliminary cohort structure. The annotation results of each sample are simultaneously projected onto the clinical semantic space constructed based on medical ontology relationships and the statistical space constructed based on sample feature distribution, and the spatial deviation of each sample is calculated and the annotated abnormal samples are identified. By tracing the evolution path of the labeled abnormal samples in the knowledge graph, identifying the feature conflict points that cause projection deviation, generating correction suggestions based on the iteration direction of the conflict points in the dual space, and adjusting the sample labels in the initial queue structure according to the correction suggestions, a clinical research queue is obtained.

2. The method according to claim 1, characterized in that, Terminology standardization is achieved by constructing a mapping network based on a multi-source knowledge base. Through terminology ambiguity constraints and temporal feature verification, the medical data is temporally aligned and feature fused to obtain a candidate sample set, including: Standardization rules for terms are extracted from the medical ontology database, clinical guideline database, and drug knowledge base, respectively. Three independent mapping sub-networks are constructed. The text records in the medical data are transformed in the three sub-networks. When the mapping results are consistent, they are directly adopted. When conflicts occur, the semantic matching confidence of each mapping sub-network is calculated. The result with the highest semantic matching confidence is selected as the standardized term, and the conflict mapping is stored in the term ambiguity database. The standardized terms are input into the semantic understanding model, and a penalty term based on the term ambiguity library is introduced into the model loss function. When the extracted feature vector corresponds to a combination of terms and there is a conflict record in the ambiguity library, the loss weight is corrected. The structured medical records in the medical data are sorted by timestamp, the time transition probability between events is calculated, the time expression of unstructured text is extracted as time anchors, semantic feature association mapping is constructed, consistency verification of the two time relationships is performed, when contradictions are found, the original records are traced back for manual review and marking, after contradictory samples are removed, the remaining features are aligned and fused by time, and the fused features are standardized and mapped to obtain a candidate sample set with unified semantic representation.

3. The method according to claim 1, characterized in that, The clinical events in the candidate sample set are transformed into entity nodes, and directed relation edges are established. Multi-scale temporal constraints and spatial topological constraints are embedded to obtain a knowledge graph, including: The clinical events of each candidate sample in the candidate sample set are divided into multiple levels. Within each level, the clinical events are transformed into entity nodes containing attribute features. Cross-level causal directed edges and intra-level temporal directed edges are established to form a two-dimensional graph topology. Extract the time interval between adjacent entity nodes, divide the time interval into multiple time scales, establish a mapping relationship between time scales and time constraint functions, limit the reasonable range of time interval between the nodes at both ends of the relationship edge, and embed the time constraint function into the edge attributes to form multi-scale time constraints. Calculate the in-degree ratio of each entity node, identify target entity nodes that exceed the centrality threshold, calculate the shortest path length between other entity nodes as the topological distance, establish an inverse mapping from topological distance to constraint strength, embed constraint strength into edge attributes to form spatial topological constraints, combine the relationship edges with embedded double constraints with entity nodes to obtain the knowledge graph.

4. The method of claim 1, wherein, The annotation results of each sample are simultaneously projected onto a clinical semantic space constructed based on medical ontology relationships and a statistical space constructed based on sample feature distribution. The spatial deviation of each sample is calculated and abnormally annotated samples are identified, including: Based on the annotation results of each sample, the corresponding disease category, hierarchical identifier and clinical endpoint event are identified. Based on the disease-symptom-drug ternary relationship in the medical ontology, a semantic relationship graph is constructed. The annotation results of each sample are mapped to coordinate vectors in the semantic relationship graph. The semantic similarity between the annotation results and the standard disease phenotype is calculated, and the projection coordinates of the sample in the clinical semantic space are determined. Principal component analysis is performed on the feature vectors of all samples in the preliminary queue structure to extract the basis vectors of the statistical space composed of multiple principal components. The feature vectors corresponding to the labeling results of each sample are linearly projected onto the basis vectors of the statistical space to obtain the projected coordinates of each sample in the statistical space. The Euclidean distance between the projected coordinates of the same sample in the clinical semantic space and the statistical space is calculated as the spatial deviation. The spatial deviation distribution is constructed and the mean and standard deviation of the deviation are calculated. The deviation threshold is obtained by weighted combination. Samples whose spatial deviation exceeds the deviation threshold are identified as labeled abnormal samples.

5. The method according to claim 1, characterized in that, Tracing the evolution path of the labeled anomalous samples in the knowledge graph, identifying feature conflict points that lead to projection deviation, generating correction suggestions based on the iteration direction of the conflict points in the dual space, and adjusting the sample labeling in the initial cohort structure according to the correction suggestions to obtain the clinical research cohort, including: The evolution path of labeled abnormal samples is extracted from the knowledge graph. The attribute features of each entity node are projected onto the clinical semantic space and the statistical space respectively. Nodes whose projection coordinate differences exceed the difference threshold are marked as candidate conflict nodes. A local subgraph centered on the candidate conflict nodes is constructed and the structural entropy difference is calculated. The candidate conflict node with the largest structural entropy difference is selected as the feature conflict point. Starting from the projection coordinates of the feature conflict point in the clinical semantic space, and ending with the projection coordinates of the abnormal sample in the statistical space, the connection vector is calculated and iterated in different directions in the two spaces. When the projection distance of the iteration trajectory converges, the convergence point is mapped to the candidate correction category and the hierarchical label adjustment target. The standard evolution path corresponding to the candidate correction category is retrieved in the knowledge graph. The similarity with the abnormal sample path is calculated. The category with the highest similarity is selected as the final correction category and the hierarchical label is updated. After the annotation adjustment is performed on all abnormal samples, the distribution balance and category consistency of the cohort are re-evaluated to obtain a clinical research cohort that meets the requirements.

6. A clinical research cohort screening and construction system based on intelligent annotation, used to implement the method as described in any one of claims 1-5, characterized in that, include: The medical data acquisition unit is used to acquire multi-source heterogeneous medical data, including structured medical records and text records; The data standardization and fusion unit is used to construct a mapping network based on a multi-source knowledge base to achieve terminology standardization. Through terminology ambiguity constraints and temporal feature verification, the medical data is time-series aligned and feature-fused to obtain a candidate sample set. The knowledge graph construction and clustering unit is used to transform the clinical events of the candidate sample set into entity nodes, establish directed relation edges, embed multi-scale temporal constraints and spatial topological constraints to obtain a knowledge graph; based on the knowledge graph, the implicit evolutionary rules and abnormal evolutionary paths are identified through the graph reasoning algorithm, and the candidate sample set is subjected to pattern classification and spectral clustering to obtain multiple sample clusters; The queue construction and intelligent annotation unit is used to allocate samples and perform intelligent annotation according to sample clusters and clinical research stratification requirements, obtain the annotation results of each sample, and form a preliminary queue structure. The dual spatial projection and anomaly detection unit is used to simultaneously project the annotation results of each sample onto a clinical semantic space constructed based on medical ontology relationships and a statistical space constructed based on sample feature distribution, calculate the spatial deviation of each sample and identify annotated abnormal samples. The evolution path tracing and queue optimization unit is used to trace the evolution path of the labeled abnormal samples in the knowledge graph, identify the feature conflict points that cause projection deviation, generate correction suggestions based on the iteration direction of the conflict points in the dual space, and adjust the sample labeling in the initial queue structure according to the correction suggestions to obtain the clinical research queue.

7. An electronic device, comprising: include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having stored thereon computer program instructions, wherein, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Electrical drawing and quotation generation method and system based on artificial intelligence

    CN120031620A

  • Medical bill intelligent processing method, system, equipment and medium

    CN121459380A