Management method, device, equipment and storage medium for clinical recruitment of users

Through ontology concept modeling and deep learning analysis combined with knowledge graph construction and encryption processing, the problems of low data utilization efficiency and privacy security in clinical recruitment are solved, and efficient user information management and secure access are achieved.

CN118779920BActive Publication Date: 2025-05-16BEIJING HOUPU PHARM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410988599.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-05-16
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

Traditional clinical recruitment methods face problems such as inefficient recruitment efficiency, insufficient sample representation, and uneven data quality, and it is difficult to effectively utilize multi-source heterogeneous data resources to protect patient privacy and data security.

Method used

By acquiring clinical expert knowledge, ontology concept modeling is carried out, combining bidirectional long and short-term memory neural networks and conditional random fields to conduct deep learning analysis of multi-source heterogeneous user information, build a target knowledge graph, and perform distributed storage and searchable encryption processing to achieve multi-level access control.

Benefits of technology

It improves the security and utilization efficiency of clinical recruitment of user information, enhances data processing capabilities and query efficiency, ensures the security of sensitive information, and refines user permission management, improving the execution efficiency of complex queries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118779920B_ABST
    Figure CN118779920B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of user management, and discloses a management method, device, equipment and storage medium for clinical recruitment users. The method includes: obtaining clinical expert knowledge and performing ontology concept modeling to construct an ontology concept model for clinical recruitment user management; obtaining multi-source heterogeneous user information and performing deep learning analysis to obtain structured user information; calculating the cosine distance and Jaccard coefficient of the structured user information to obtain fused user information; performing graph structuring on the fused user information and the ontology concept model to obtain a target knowledge graph; performing distributed storage of clinical expert knowledge and multi-source heterogeneous user information based on the target knowledge graph to construct a multi-source user information distributed database; performing searchable encryption and multi-level access control to obtain a user information security access control strategy. The present application constructs a distributed database for clinical recruitment user information management to improve the security of user information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of user management, and in particular to a management method, apparatus, device and storage medium for clinical recruitment of users. Background Art

[0002] Traditional clinical recruitment methods face many challenges, such as low recruitment efficiency, insufficient sample representativeness, and uneven data quality. With the advent of the big data era, medical and research institutions have accumulated a large amount of clinical data and patient information, which provides new opportunities for improving clinical recruitment methods.

[0003] Nevertheless, how to effectively utilize these multi-source heterogeneous data resources, extract valuable information from them, and apply them to the clinical recruitment process remains an urgent problem to be solved. Existing data processing methods often have difficulty coping with the heterogeneity and complexity of data, resulting in inefficient information utilization. At the same time, due to the sensitivity of medical data, how to protect patient privacy and data security while making full use of data has also become a problem that needs to be solved. In addition, there are still problems in the clinical recruitment process, such as insufficient knowledge representation, decentralized user information management, and lax data access control. These problems not only affect the efficiency and quality of clinical recruitment, but also increase the risk of data leakage and abuse. Summary of the invention

[0004] The present application provides a management method, apparatus, device and storage medium for clinical recruitment users, which are used to build a distributed database for clinical recruitment user information management and improve the security of user information.

[0005] In a first aspect, the present application provides a method for managing clinical recruitment users, the method for managing clinical recruitment users comprising:

[0006] Acquire clinical expert knowledge and conduct ontology concept modeling to build an ontology concept model for clinical recruitment user management;

[0007] Acquire multi-source heterogeneous user information, and use a bidirectional long short-term memory neural network and a conditional random field to perform deep learning analysis on the multi-source heterogeneous user information to obtain structured user information;

[0008] Calculating the cosine distance and Jaccard coefficient of the structured user information to obtain fused user information;

[0009] Performing graph-structuring processing on the fused user information and the ontology concept model to obtain a target knowledge graph for clinical recruitment user management;

[0010] Based on the target knowledge graph, the clinical expert knowledge and the multi-source heterogeneous user information are distributedly stored to build a multi-source user information distributed database;

[0011] The multi-source user information distributed database is searchably encrypted and multi-level access controlled to obtain a user information security access control strategy.

[0012] In a second aspect, the present application provides a management device for clinical recruitment users, the management device for clinical recruitment users comprising:

[0013] Modeling module, used to acquire clinical expert knowledge and conduct ontology concept modeling, and build an ontology concept model for clinical recruitment user management;

[0014] An analysis module is used to obtain multi-source heterogeneous user information, and use a bidirectional long short-term memory neural network and a conditional random field to perform deep learning analysis on the multi-source heterogeneous user information to obtain structured user information;

[0015] A calculation module, used to calculate the cosine distance and Jaccard coefficient of the structured user information to obtain fused user information;

[0016] A processing module, used for performing graph-structured processing on the fused user information and the ontology concept model to obtain a target knowledge graph for clinical recruitment user management;

[0017] A storage module, used for distributing and storing the clinical expert knowledge and the multi-source heterogeneous user information based on the target knowledge graph, and constructing a multi-source user information distributed database;

[0018] The control module is used to perform searchable encryption and multi-level access control on the multi-source user information distributed database to obtain a user information security access control strategy.

[0019] The third aspect of the present application provides a management device for clinical recruitment users, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the management device for clinical recruitment users executes the above-mentioned management method for clinical recruitment users.

[0020] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the above-mentioned method for managing clinical recruitment users.

[0021] In the technical solution provided by the present application, a structured knowledge system is constructed by ontology concept modeling of clinical expert knowledge, which improves the accuracy and completeness of the representation of clinical recruitment domain knowledge. The bidirectional long short-term memory neural network and conditional random field are used to perform deep learning analysis on multi-source heterogeneous user information, which significantly improves the ability to extract structured information from complex data. Combined with cosine distance and Jaccard coefficient for calculation, multi-dimensional similarity analysis of user information is realized, and the accuracy and comprehensiveness of user information fusion are improved. By graph-structuring the fusion of user information and ontology concept model, a knowledge graph with rich semantics is constructed, and distributed storage is performed based on the target knowledge graph, which improves the data processing capability and query efficiency of the system, and the distributed database is searchable and encrypted, which effectively improves the security of sensitive information while ensuring data availability. By implementing a multi-level access control strategy, user authority management is refined, and a query optimization strategy is designed based on the knowledge graph structure, which significantly improves the execution efficiency of complex queries. By designing a knowledge graph update synchronization strategy, real-time updating of distributed databases is achieved, ensuring the timeliness and consistency of data, thereby improving the security of user information. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0023] Figure 1 A schematic diagram of an embodiment of a method for managing clinical recruitment users in an embodiment of the present application;

[0024] Figure 2 This is a schematic diagram of an embodiment of a management device for clinical recruitment of users in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The present application embodiment provides a management method, device, equipment and storage medium for clinical recruitment of users. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described here can be implemented in an order other than the content illustrated or described here. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 , an embodiment of the management method of clinical recruitment users in the embodiment of the present application includes:

[0027] Step S101, acquiring clinical expert knowledge and performing ontology concept modeling to construct an ontology concept model for clinical recruitment user management;

[0028] It is understandable that the execution subject of the present application may be a management device for clinical recruitment of users, or a terminal or a server, which is not specifically limited here. The present application embodiment is described by taking a server as the execution subject as an example.

[0029] Specifically, clinical expert knowledge is obtained and classified into domains to obtain multiple clinical recruitment related domains. The knowledge of clinical experts covers information on various diseases, treatment methods, patient management, etc. These knowledge are systematically classified and organized to clarify the boundaries and connections between different domains. By classifying clinical recruitment related domains, the core content and key issues of each domain can be more accurately located. Domain key concepts are extracted from multiple clinical recruitment related domains to obtain a domain core concept set, and the domain core concept set is hierarchically organized to obtain a concept hierarchy tree. Natural language processing technology and expert knowledge are applied to extract key concepts in each domain. According to the concept hierarchy tree, the relationship between concepts is defined to obtain a concept relationship network, and the concept relationship network is analyzed for attribute characteristics to obtain a concept attribute set. The concept relationship network is the core part of the ontology model, which defines the relationship between different concepts, such as inclusion relationship, similarity relationship, causal relationship, etc. By defining these relationships, the complex associations between concepts in the domain are revealed. At the same time, the attribute characteristics of the concept relationship network are analyzed to identify the specific attributes of each concept, such as the patient's age, gender, medical history, etc. According to the concept attribute set, the constraint relationship between attributes is defined to obtain the attribute constraint rules. Attribute constraint rules are an important part of the ontology model, which define the constraint relationships between different attributes, such as mutually exclusive relationships, dependent relationships, conditional relationships, etc. These rules help ensure the consistency and accuracy of the model, so that automated reasoning and decision-making can be performed based on these rules in practical applications. The concept hierarchy tree, concept relationship network and attribute constraint rules are integrated to obtain the initial ontology model. The initial ontology model is a comprehensive knowledge representation framework that integrates all key concepts in the field, the relationships between concepts, and the constraint rules between attributes to form a complete knowledge graph. The model consistency is verified according to the initial ontology model to obtain the consistency verification results, and the initial ontology model is iteratively optimized according to the consistency verification results to obtain the ontology concept model for clinical recruitment user management. Logical reasoning and verification techniques are applied to strictly check the consistency of the model to ensure that there are no logical contradictions and inconsistencies in the model. Through multiple iterative optimizations, the accuracy and completeness of the model are gradually improved, and finally a high-quality ontology concept model is obtained.

[0030] Step S102: acquiring multi-source heterogeneous user information, and performing deep learning analysis on the multi-source heterogeneous user information using a bidirectional long short-term memory neural network and a conditional random field to obtain structured user information;

[0031] Specifically, multi-source heterogeneous user information is obtained, and data cleaning is performed on the multi-source heterogeneous user information to remove invalid information, unify the data format, and obtain standardized text data. According to the standardized text data, the words are segmented and part-of-speech tagged to obtain a word sequence, and the word sequence is embedded to obtain a word vector sequence. Word segmentation and part-of-speech tagging are basic steps in natural language processing. Through word segmentation, continuous text can be divided into meaningful word units, and part-of-speech tagging assigns corresponding part-of-speech tags to each word. Word embedding technology is used to convert words into word vectors with semantic information to form a word vector sequence. A bidirectional long short-term memory neural network is used to perform sequence forward LSTM calculation on the word vector sequence to obtain a forward hidden state sequence, and a reverse LSTM calculation is performed on the word vector sequence to obtain a reverse hidden state sequence. Bidirectional LSTM can capture the global information of the context by encoding the sequence information from the forward and backward directions respectively, and improve the model's ability to understand sequence data. The forward hidden state sequence and the reverse hidden state sequence are concatenated to obtain a bidirectional LSTM feature representation, and the bidirectional LSTM feature representation is processed by a fully connected layer to obtain an emission score matrix. Concatenating the forward and reverse hidden state sequences can fuse bidirectional information to form a richer feature representation. Through the fully connected layer processing, the high-dimensional feature representation is mapped to the label space to obtain the emission score corresponding to each time step, forming an emission score matrix. Through the conditional random field, the emission score matrix is ​​calculated for state transition to obtain the optimal labeling sequence. The conditional random field is a probabilistic graph model that performs global optimal labeling on each time step in the sequence according to the context information and the emission score to generate an optimal labeling sequence. Entity recognition is performed on the optimal labeling sequence to obtain a named entity set, and entity relationship extraction is performed on the named entity set to obtain structured user information. Entity recognition forms a named entity set by identifying key entities in the text, such as patients, doctors, symptoms, etc. Entity relationship extraction, on the other hand, identifies the relationship between named entities, such as the relationship between patients and doctors, the relationship between symptoms and treatment methods, and finally obtains structured user information.

[0032] Step S103, calculating the cosine distance and Jaccard coefficient of the structured user information to obtain fused user information;

[0033] Specifically, the structured user information is feature vectorized to obtain a user feature vector set. The feature extraction technology is used to convert the information of each user into a numerical feature vector, so that quantitative analysis can be performed. The cosine similarity between vectors of the user feature vector set is calculated to obtain an inter-user similarity matrix. Cosine similarity evaluates the similarity between users by calculating the cosine value of the angle between vectors. The generated inter-user similarity matrix can reflect the similarity relationship between users. At the same time, the category attributes in the structured user information are one-hot encoded to obtain a category feature binary matrix. One-hot encoding is a common method for processing classified data. By converting the category attributes into binary vectors, it is convenient to calculate the Jaccard coefficient. The Jaccard coefficient is calculated based on the category feature binary matrix to obtain the inter-user Jaccard similarity matrix. The Jaccard coefficient is used to measure the similarity between two sets, that is, the proportion of common features, so as to generate the inter-user Jaccard similarity matrix, which reflects the similarity between users in the classification attributes. The inter-user similarity matrix and the inter-user Jaccard similarity matrix are weighted fused to obtain a comprehensive similarity matrix. Perform user clustering analysis based on the comprehensive similarity matrix to obtain the user group division results. Use a clustering algorithm to group users so that similar users are divided into the same group to achieve preliminary classification of users. Extract representative features for each group in the user group division results to obtain a group feature vector. Extract vectors that can represent the characteristics of each group from each group. These vectors are a summary and conclusion of the group characteristics, which is convenient for subsequent analysis. Calculate the similarity between groups based on the group feature vectors to obtain the inter-group correlation matrix. Perform threshold screening on the inter-group correlation matrix to obtain highly correlated group pairs. By setting a similarity threshold, group pairs with high similarity are screened out. These group pairs have high similarity in certain features. Based on the highly correlated group pairs, perform cross-group fusion of structured user information to obtain fused user information. Merge highly correlated groups and integrate their feature information to form more complete and unified user information.

[0034] Step S104, performing graph-structuring processing on the fused user information and ontology concept model to obtain a target knowledge graph for clinical recruitment user management;

[0035] Specifically, entity recognition is performed on the fused user information to obtain a user entity set. Various key entities, such as patients, doctors, clinical trials, etc., are identified from the fused user information through natural language processing and machine learning technologies. Entity type mapping is performed based on the user entity set and the ontology concept model to obtain a typed entity set. Relationship extraction is performed on the typed entity set to obtain an entity relationship set. By analyzing the association information between entities, the relationship between each entity is extracted, for example, the treatment relationship between a patient and a doctor, the relationship between a patient participating in a clinical trial, etc. Relation semantic alignment is performed based on the entity relationship set and the ontology concept model to obtain a semantic relationship set. The extracted entity relationships are compared and matched with the relationships in the ontology concept model to ensure the semantic consistency and accuracy of the relationships. The typed entity set and the semantic relationship set are converted into graph data structures to obtain an initial knowledge graph. The entities and relationships in the initial knowledge graph are embedded to obtain a graph embedding vector set. The nodes and edges in the graph are converted into low-dimensional vectors to facilitate calculation and analysis. Similarity calculation is performed on the graph embedding vector set to obtain a similarity matrix of entities and relationships. By calculating the similarity between entity and relationship vectors, the degree of association between them is evaluated. According to the similarity matrix, the redundant and contradictory information in the initial knowledge graph is disambiguated to obtain an optimized knowledge graph. Disambiguation refers to identifying and eliminating repeated or contradictory information in the graph through similarity calculation to improve the accuracy and consistency of the knowledge graph. Ontology reasoning is performed on the optimized knowledge graph to obtain a reasoning-expanded knowledge graph. Using the rules and logic in the ontology, the knowledge graph is reasoned and expanded to generate more implicit knowledge. Through ontology reasoning, the content of the knowledge graph is enriched to make it more practical and comprehensive. According to the reasoning-expanded knowledge graph, the graph structure is hierarchically organized, and the entities and relationships in the knowledge graph are hierarchically organized and arranged to give it a clear hierarchical structure, which is easy to manage and query, and the target knowledge graph for clinical recruitment user management is obtained.

[0036] Step S105: Distribute and store clinical expert knowledge and multi-source heterogeneous user information based on the target knowledge graph to build a multi-source user information distributed database;

[0037] Specifically, the target knowledge graph is structurally analyzed to obtain the data model of entity table, relationship table and attribute table. Through comprehensive analysis of the knowledge graph, the nodes and edges in the graph are converted into table structures. According to the data model, data mapping is performed on clinical expert knowledge and multi-source heterogeneous user information to obtain a structured data set. Unstructured or semi-structured data is converted into structured data that conforms to the data model for easy storage and query. A data sharding strategy is designed for the structured data set to obtain a sharding scheme. Large-scale data sets are divided into multiple small fragments for distributed storage and management. Data distributed storage nodes are allocated according to the sharding scheme to obtain the initial data distribution result. By allocating the sharded data to different storage nodes, distributed storage of data is realized, and the scalability and reliability of data storage are improved. A data synchronization mechanism is designed for the initial data distribution result to obtain a knowledge graph update synchronization strategy. According to the knowledge graph update synchronization strategy, the distributed database is updated in real time to obtain a dynamic data storage structure. Through the real-time update mechanism, the consistency and timeliness of the data are ensured, so that the database can accurately reflect the latest knowledge graph content. A query optimizer is designed for the dynamic data storage structure to obtain a query optimization strategy based on the knowledge graph structure. A query optimizer is a component that improves query efficiency by optimizing query paths and execution plans when querying a database. Complex query execution plans are generated based on the query optimization strategy to obtain an optimized query execution plan. Through query optimization, the performance of the database in processing complex queries is improved and the query response time is reduced. A user role and permission management module is designed for the distributed database to obtain a multi-level access control framework. User role and permission management refers to the management of access rights for different users to ensure data security and confidentiality. The multi-level access control framework ensures that only authorized users can access specific data and functions by defining different access levels and permission rules. A corresponding multi-source user information distributed database is generated based on the multi-level access control framework.

[0038] Step S106: Perform searchable encryption and multi-level access control on the multi-source user information distributed database to obtain a user information security access control policy.

[0039] Specifically, authorized users are queried based on a distributed database of multi-source user information, and the authorized users are grouped according to security levels to obtain user security level classification. By performing identity authentication and security level assessment on authorized users, users are classified according to different security requirements to ensure that users of different levels have corresponding access rights. According to the user security level classification, TF-IDF processing is performed on sensitive data in the structured data set to obtain a plaintext keyword weight matrix. TF-IDF processing generates a keyword weight matrix by calculating the importance of each keyword in the document, so that key content in sensitive data can be identified and marked. The plaintext keyword weight matrix is ​​encrypted by AES to obtain an initial encryption key, and round function calculation is performed based on the initial encryption key to obtain a final encryption key. AES encryption is a symmetric encryption algorithm that generates a final encryption key through multiple rounds of function calculation to protect data security. Matrix function encryption is performed on sensitive data to obtain encrypted ciphertext data, and inverse matrix function processing is performed on the encrypted ciphertext data to obtain a decryption matrix. Matrix function encryption implements encryption processing of sensitive data by performing matrix operations on data, and inverse matrix function processing is used in the decryption process to ensure that the data can be correctly restored. At the same time, the encrypted ciphertext data is analyzed to obtain a searchable ciphertext index. Searchable encryption technology allows keyword searches without decrypting the data, thereby achieving efficient data retrieval while protecting data privacy. According to the user security level classification, the user access rights of authorized users are reviewed to obtain access authorization results. By reviewing and verifying the user's access request, it is ensured that only users with corresponding permissions can access specific data. According to the access authorization result, the search request of the authorized user is encrypted to obtain the number of ciphertexts and keyword search requests. By encrypting the search request, the security of the data during transmission is ensured, and the search operation can be performed on the encrypted data. According to the number of ciphertexts and keyword search requests, the encrypted ciphertext data is retrieved and decrypted to obtain the user information security access control policy to ensure that users can access the required information in a secure environment.

[0040] In the embodiment of the present application, a structured knowledge system is constructed by ontology concept modeling of clinical expert knowledge, which improves the accuracy and completeness of the representation of clinical recruitment field knowledge. The bidirectional long short-term memory neural network and conditional random field are used to perform deep learning analysis on multi-source heterogeneous user information, which significantly improves the ability to extract structured information from complex data. Combined with cosine distance and Jaccard coefficient for calculation, multi-dimensional similarity analysis of user information is realized, and the accuracy and comprehensiveness of user information fusion are improved. By graph-structuring the fusion of user information and ontology concept model, a knowledge graph with rich semantics is constructed, and distributed storage is performed based on the target knowledge graph, which improves the data processing ability and query efficiency of the system, and the distributed database is searchable and encrypted. While ensuring data availability, the security of sensitive information is effectively improved. By implementing a multi-level access control strategy, user authority management is refined, and a query optimization strategy is designed based on the knowledge graph structure, which significantly improves the execution efficiency of complex queries. By designing a knowledge graph update synchronization strategy, real-time updating of distributed databases is achieved, ensuring the timeliness and consistency of data, thereby improving the security of user information.

[0041] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0042] (1) Acquire clinical expert knowledge and classify clinical expert knowledge into fields to obtain multiple clinical recruitment related fields;

[0043] (2) Extract key concepts from multiple clinical recruitment-related fields to obtain a set of core concepts, and organize the set of core concepts into a hierarchical structure to obtain a concept hierarchy tree;

[0044] (3) Based on the concept hierarchy tree, define the relationship between concepts to obtain a concept relationship network, and perform attribute feature analysis on the concept relationship network to obtain a concept attribute set;

[0045] (4) Based on the concept attribute set, define the constraint relationship between attributes and obtain the attribute constraint rules;

[0046] (5) Integrate the concept hierarchy tree, concept relationship network and attribute constraint rules to obtain the initial ontology model;

[0047] (6) Perform model consistency verification based on the initial ontology model to obtain consistency verification results, and iteratively optimize the initial ontology model based on the consistency verification results to obtain the ontology conceptual model for clinical recruitment user management.

[0048] Specifically, clinical expert knowledge is obtained, and this knowledge is systematically classified into fields to obtain multiple clinical recruitment-related fields. For example, suppose knowledge is extracted from a database containing a variety of clinical trial information, which may include treatment methods for different diseases, patient screening criteria, clinical trial stages, drug dosages, etc. By classifying this information, several fields are divided, such as cardiovascular disease, oncology, and neurological diseases. When extracting key concepts in multiple clinical recruitment-related fields, natural language processing (NLP) technology and expert knowledge are used. For example, in the field of cardiovascular disease, key concepts such as "coronary heart disease", "hypertension", "myocardial infarction", "blood lipids", and "electrocardiogram" can be extracted. In the process of extracting these concepts, the TF-IDF algorithm can be used to measure the importance of each term. The higher the TF-IDF value, the more important the term is in a specific document. The formula is as follows:

[0049] ;

[0050] Among them, TF Representation term In the documentation The frequency of occurrence, IDF Representation term The inverse document frequency of is defined as:

[0051] ;

[0052] in, is the total number of documents, DF To include the term The number of documents. Through this algorithm, the core concept set in each field is screened out. The core concept set of the field is hierarchically organized to obtain a concept hierarchy tree. Using expert knowledge and field characteristics, related concepts are organized according to their hierarchical relationships. For example, in the field of cardiovascular diseases, "coronary heart disease" and "hypertension" are classified as first-level concepts, while "myocardial infarction" is classified as a second-level concept of "coronary heart disease", and "blood lipids" are classified as a related factor of "hypertension". Through this hierarchical organization, a tree structure is formed. According to the concept hierarchy tree, the relationship between concepts is defined to obtain a concept relationship network. The concept relationship network contains not only hierarchical relationships, but also other types of relationships, such as association relationships, causal relationships, etc. For example, there is a causal relationship between "hypertension" and "myocardial infarction", and there is a diagnostic relationship between "electrocardiogram" and "coronary heart disease". The concept relationship network is analyzed for attribute characteristics to obtain a concept attribute set. The concept attribute set contains the specific characteristics of each concept. For example, the attributes of "hypertension" may include "blood pressure value", "duration", "patient age", etc. According to the concept attribute set, the constraint relationship between attributes is defined to obtain attribute constraint rules. Attribute constraint rules are used to limit the relationship and value range between attributes. For example, "blood pressure value" should be within a certain range, and "patient age" should meet the requirements of a specific test. The concept hierarchy tree, concept relationship network and attribute constraint rules are integrated to obtain the initial ontology model. The initial ontology model is a comprehensive representation of domain knowledge, including all concepts, relationships and attributes, and the integrity and consistency of the model are ensured by rules. According to the initial ontology model, the model consistency verification is performed to ensure that there are no logical contradictions and inconsistencies in the model. Model consistency verification can be performed through logical reasoning and consistency checking tools. For example, verify whether the causal relationship between "hypertension" and "myocardial infarction" is reasonable, and verify whether the value range of "blood pressure value" meets medical standards. Through these verifications, the consistency verification results are obtained. According to the consistency verification results, the initial ontology model is iteratively optimized to obtain the ontology concept model for clinical recruitment user management. For example, if it is found in the consistency verification that the relationship between some concepts is not clearly defined, or the value range of some attributes is unreasonable, adjustments and optimizations can be made according to the verification results until the model meets the consistency and integrity requirements.

[0053] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0054] (1) Obtain multi-source heterogeneous user information and perform data cleaning on the multi-source heterogeneous user information to obtain standardized text data;

[0055] (2) Based on the standardized text data, the words are segmented and POS tagged to obtain a word sequence, and the word sequence is embedded to obtain a word vector sequence;

[0056] (3) Using a bidirectional long short-term memory neural network, a forward LSTM calculation is performed on the word vector sequence to obtain a forward hidden state sequence, and a reverse LSTM calculation is performed on the word vector sequence to obtain a reverse hidden state sequence;

[0057] (4) Concatenate the forward hidden state sequence and the reverse hidden state sequence to obtain a bidirectional LSTM feature representation, and perform a fully connected layer processing on the bidirectional LSTM feature representation to obtain an emission score matrix;

[0058] (5) Through the conditional random field, the state transition calculation of the emission score matrix is ​​performed to obtain the optimal labeling sequence;

[0059] (6) Perform entity recognition on the optimal annotation sequence to obtain a named entity set, and then perform entity relationship extraction on the named entity set to obtain structured user information.

[0060] Specifically, multi-source heterogeneous user information is obtained, including data from different sources and formats, such as electronic health records, social media data, and online questionnaire results. These data usually have different structures and formats. Data cleaning is performed to remove noise, fill missing values, and unify the format to obtain standardized text data. Data cleaning can be achieved through regular expression matching, missing value filling algorithms, and data standardization techniques. According to the standardized text data, the words are segmented and part-of-speech tagged to obtain a word sequence. Word segmentation is to divide the text into independent word units, and part-of-speech tagging assigns a corresponding part-of-speech label to each word. The word sequence is embedded to obtain a word vector sequence. Word embedding is a technology that maps words to a low-dimensional vector space and can capture the semantic information of words. Common word embedding technologies include Word2Vec and GloVe. A bidirectional long short-term memory neural network is used to perform sequence forward LSTM calculation on the word vector sequence to obtain a forward hidden state sequence. LSTM (Long Short-Term Memory) is a neural network that can capture long-range dependencies in a sequence, and bidirectional LSTM combines information in both forward and backward directions. The forward LSTM calculation formula is as follows:

[0061] ;

[0062] in, is the time step The forward hidden state of is the time step The word vector representation of is the time step Similarly, the reverse LSTM calculation is performed on the word vector sequence to obtain the reverse hidden state sequence:

[0063] ;

[0064] in, is the time step The reverse hidden state of is the time step The forward hidden state sequence and the reverse hidden state sequence are concatenated to obtain the bidirectional LSTM feature representation:

[0065] ;

[0066] By concatenating the forward and reverse hidden states, the context information is integrated to form a richer feature representation. The bidirectional LSTM feature representation is processed by the fully connected layer to obtain the emission score matrix. The fully connected layer maps the bidirectional LSTM features to the label space through linear transformation and nonlinear activation function:

[0067] ;

[0068] in, is the weight matrix, is the bias vector, is the time step The emission score. Through the conditional random field, the state transition calculation of the emission score matrix is ​​performed to obtain the optimal labeling sequence. CRF is a sequence labeling model that uses context information to label each element in the sequence. CRF solves the optimal labeling sequence by maximizing the log-likelihood estimate of the sequence:

[0069] ;

[0070] in, For a given input sequence The following annotation sequence The probability of is the normalization factor. Perform entity recognition on the optimal annotation sequence to obtain a named entity set. Identify specific segments in the annotation sequence as entities, extract entity relationships on the named entity set, identify the relationships between entities, and obtain structured user information.

[0071] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0072] (1) Feature vectorization is performed on the structured user information to obtain a user feature vector set, and the cosine similarity between vectors of the user feature vector set is calculated to obtain an inter-user similarity matrix;

[0073] (2) One-hot encode the category attributes in the structured user information to obtain a category feature binary matrix, and calculate the Jaccard coefficient based on the category feature binary matrix to obtain the Jaccard similarity matrix between users;

[0074] (3) Perform weighted fusion on the inter-user similarity matrix and the inter-user Jaccard similarity matrix to obtain a comprehensive similarity matrix;

[0075] (4) Perform user clustering analysis based on the comprehensive similarity matrix to obtain the user group division results, and extract representative features for each group in the user group division results to obtain the group feature vector;

[0076] (5) Calculate the similarity between groups based on the group feature vectors to obtain the inter-group correlation matrix, and perform threshold screening on the inter-group correlation matrix to obtain highly correlated group pairs;

[0077] (6) Based on the highly correlated group pairs, the structured user information is fused across groups to obtain fused user information.

[0078] Specifically, the structured user information is feature vectorized, and each user's information is converted into a feature vector set. Assuming that there is user information such as age, gender, and medical history, this information can be represented as a feature vector. For example, the vector of user A is represented as , where 30 represents age, 1 represents male, and the subsequent 0 and 1 represent disease history, such as whether the user has high blood pressure or diabetes. The cosine similarity between vectors is calculated for the user feature vector set to obtain the user similarity matrix. Cosine similarity measures the similarity between two vectors by calculating the cosine value between them. The formula is as follows:

[0079] ;

[0080] in, Representation vector and The dot product of Represents vectors and By calculating the cosine similarity, we can get the similarity matrix between users. , where each element Indicates user and users The similarity between them. One-hot encode the category attributes in the structured user information to obtain a category feature binary matrix. For example, gender and disease history can be converted into binary form through one-hot encoding, such as male is represented as [1, 0], female is represented as [0, 1], hypertension is represented as [1, 0], and no hypertension is represented as [0, 1]. The Jaccard coefficient is calculated based on the category feature binary matrix to obtain the Jaccard similarity matrix between users. The Jaccard coefficient is used to measure the similarity between two sets. The formula is as follows:

[0081] ;

[0082] in, Representing a collection and collection The size of the intersection of Represents the size of their union. By calculating the Jaccard coefficient, we get the Jaccard similarity matrix between users. , where each element Indicates user and users The inter-user similarity matrix and the inter-user Jaccard similarity matrix are weighted fused to obtain a comprehensive similarity matrix Weighted fusion can be achieved by linearly combining the two, the formula is as follows:

[0083] ;

[0084] in, is the weighting coefficient, which controls the weight ratio of cosine similarity and Jaccard similarity. , balance the contribution of the two similarities to the comprehensive similarity. Perform user cluster analysis based on the comprehensive similarity matrix to obtain the user group division results. Cluster analysis can use algorithms such as K-means or hierarchical clustering. Assume that the K-means algorithm is used to divide users into several groups, each group represents a class of similar users. Through cluster analysis, the user group division results are obtained , where each group contains several users. Perform representative feature extraction on each group in the user group division result to obtain a group feature vector. Representative feature extraction can be achieved by calculating the average value of all user feature vectors in each group, and the formula is as follows:

[0085] ;

[0086] in, For Group The characteristic vector of For Group The number of users in For users The characteristic vector of For Group The set of all users in . The similarity between groups is calculated based on the group feature vector to obtain the inter-group association matrix The similarity between groups can be measured by cosine similarity or other similarity calculation methods, the formula is as follows:

[0087] ;

[0088] in, Indicates group and Groups The similarity between and Group and Groups The feature vector of . Threshold screening is performed on the inter-group correlation matrix to obtain high-correlation group pairs. Threshold screening can be done by setting a similarity threshold To achieve, if , then the group and Groups is a highly correlated group pair. By screening, we can obtain a set of highly correlated group pairs. According to the highly correlated group pairs, the structured user information is fused across groups to obtain fused user information. The user information in the highly correlated groups is integrated to form a larger group to enhance the integrity and consistency of the data. and Groups For a highly correlated group pair, the user information in the two groups can be merged to obtain fused group information.

[0089] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0090] (1) Perform entity recognition on the fused user information to obtain a user entity set, and perform entity type mapping based on the user entity set and the ontology conceptual model to obtain a typed entity set;

[0091] (2) Extract relations from the typed entity set to obtain a set of relations between entities, and align the relations semantically based on the set of relations between entities and the ontology concept model to obtain a semantic relation set;

[0092] (3) Convert the typed entity set and semantic relationship set into a graph data structure to obtain an initial knowledge graph, and embed the entities and relationships in the initial knowledge graph to obtain a graph embedding vector set;

[0093] (4) Calculate the similarity of the graph embedding vector set to obtain the similarity matrix of entities and relationships. Based on the similarity matrix, disambiguate the redundant and contradictory information in the initial knowledge graph to obtain the optimized knowledge graph.

[0094] (5) Perform ontology reasoning on the optimized knowledge graph to obtain a reasoned and extended knowledge graph, and organize the graph structure hierarchically based on the reasoned and extended knowledge graph to obtain the target knowledge graph for clinical recruitment user management.

[0095] Specifically, entity recognition is performed on the fused user information, and important information in the text is extracted as entities to obtain a user entity set. Entity type mapping is performed based on the user entity set and the ontology concept model to obtain a typed entity set. The ontology concept model contains the structure and definition of domain knowledge, and the identified entities are mapped to the corresponding types. Relationship extraction is performed on the typed entity set to identify the associations between entities and obtain a set of inter-entity relationships. Through relationship extraction technology, these relationships are extracted from the text and structured. Relation semantic alignment is performed based on the inter-entity relationship set and the ontology concept model to obtain a semantic relationship set. The extracted relationships are matched and mapped with the relationships in the ontology concept model to ensure the semantic consistency of the relationships. For example, the "suffering from" relationship defined in the ontology model ensures that the relationships extracted from different texts can be uniformly represented and understood. Through semantic alignment, a well-consistent relationship set is obtained to ensure the accuracy and completeness of the data. The typed entity set and the semantic relationship set are converted into a graph data structure to obtain an initial knowledge graph. Entities and relationships are represented as a graph structure, with nodes representing entities and edges representing relationships between entities. The entities and relationships in the initial knowledge graph are embedded to obtain a graph embedding vector set. Graph embedding technology represents the nodes and edges in the graph as low-dimensional vectors for calculation and analysis. Common graph embedding algorithms include DeepWalk and node2vec. Through embedding representation, the entities and relationships in the knowledge graph are converted into vector form, which is convenient for subsequent similarity calculation and disambiguation processing. Similarity calculation is performed on the graph embedding vector set to obtain the similarity matrix of entities and relationships. Similarity calculation can use cosine similarity or other similarity measurement methods. According to the similarity matrix, the redundant and contradictory information in the initial knowledge graph is disambiguated to obtain an optimized knowledge graph. Disambiguation is to identify and eliminate repeated or contradictory information in the graph to ensure the accuracy and consistency of the knowledge graph. For example, if two nodes represent the same entity but there is redundancy, these nodes can be merged through the similarity matrix. Through disambiguation, the knowledge graph is optimized to make it more concise and accurate. Ontology reasoning is performed on the optimized knowledge graph to obtain a reasoning-expanded knowledge graph. The knowledge graph is reasoned and expanded using the rules and logic in the ontology to generate more implicit knowledge. According to the reasoning-expanded knowledge graph, the graph structure is hierarchically organized to obtain the target knowledge graph for clinical recruitment user management. The graph structure hierarchical organization is to hierarchically organize and arrange the entities and relationships in the knowledge graph so that it has a clear hierarchical structure. For example, patients can be used as top-level nodes, diseases as second-level nodes, and treatment methods as third-level nodes. Through hierarchical organization, the knowledge graph is clearer, easier to understand and query.

[0096] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0097] (1) Perform structural analysis on the target knowledge graph to obtain the data model of entity table, relationship table and attribute table;

[0098] (2) Based on the data model, data mapping is performed on clinical expert knowledge and multi-source heterogeneous user information to obtain a structured data set;

[0099] (3) Design a data sharding strategy for the structured data set to obtain a sharding scheme, and allocate data distributed storage nodes according to the sharding scheme to obtain the initial data distribution result;

[0100] (4) Design a data synchronization mechanism for the initial data distribution results to obtain a knowledge graph update synchronization strategy. Based on the knowledge graph update synchronization strategy, the distributed database is updated in real time to obtain a dynamic data storage structure.

[0101] (5) Design a query optimizer for the dynamic data storage structure, obtain a query optimization strategy based on the knowledge graph structure, and generate a complex query execution plan based on the query optimization strategy to obtain an optimized query execution plan;

[0102] (6) Design user role and permission management modules for distributed databases to obtain a multi-level access control framework;

[0103] (7) Generate a corresponding multi-source user information distributed database based on the multi-level access control framework.

[0104] Specifically, the target knowledge graph is structurally analyzed to obtain the data model of entity table, relationship table and attribute table. The knowledge graph consists of entities (nodes), relationships (edges) and attributes (node ​​or edge labels). In the process of structural analysis, each element of the knowledge graph is decomposed into three main parts: entity table, relationship table and attribute table. For example, for a medical knowledge graph, entities may include "patient", "doctor", "disease", "drug", etc., and relationships may include "suffering from", "treatment", "prescription", etc. The attributes are detailed descriptions of these entities and relationships, such as "patient age", "disease symptoms", etc. According to the data model, data mapping is performed on clinical expert knowledge and multi-source heterogeneous user information to obtain a structured data set. The original data is converted into a format that conforms to the target data model. Through data cleaning and conversion, all heterogeneous data are mapped to predefined entity tables, relationship tables and attribute tables. Data sharding strategy design is performed on the structured data set to obtain a sharding scheme. Divide large-scale data into multiple smaller parts for distributed storage and management. For example, sharding is performed by geographic location, time period or data type. The design of the sharding strategy needs to consider factors such as the query frequency, update frequency and data volume of the data. According to the sharding scheme, the data distributed storage nodes are allocated to obtain the initial data distribution results. By allocating the sharded data to different storage nodes, the distributed storage of data is realized, and the scalability and reliability of the storage system are improved. The data synchronization mechanism is designed for the initial data distribution results to obtain the knowledge graph update synchronization strategy. The data synchronization mechanism is the key to ensuring the data consistency on the distributed storage nodes. The knowledge graph update synchronization strategy includes two methods: real-time synchronization and periodic synchronization. For example, when the data on a node is updated, the update is immediately synchronized to other nodes; or the data is synchronized in batches at regular intervals. By designing an effective synchronization mechanism, the data consistency on all nodes is ensured to avoid data inconsistency problems. According to the knowledge graph update synchronization strategy, the distributed database is updated in real time to obtain a dynamic data storage structure. The real-time update mechanism can be implemented through the publish-subscribe mode, the dual-write mechanism, etc. For example, using the publish-subscribe mode, when the data of a node changes, an update message is published, and the subscribing node receives and updates the data. The query optimizer is designed for the dynamic data storage structure to obtain a query optimization strategy based on the knowledge graph structure. The query optimizer selects the optimal query path and execution plan by analyzing the query request. For example, for a complex query request, use indexing, caching and other technologies to improve query efficiency. Query optimization strategies include selecting appropriate index structures, optimizing query statements, and distributed query plans. For example, when querying all the medical records of a patient, the query speed can be improved by pre-establishing indexes. Complex query execution plans are generated according to the query optimization strategy to obtain optimized query execution plans. The query execution plan includes query parsing, logical optimization, physical optimization and other steps.By optimizing the query execution plan, query performance can be significantly improved. For example, when querying a patient's disease history, the relevant records can be directly located through the index, reducing the query time. The user role and permission management module of the distributed database is designed to obtain a multi-level access control framework. Permission management is an important measure to ensure data security. The multi-level access control framework ensures that only authorized users can access specific data by defining access rights for different roles. For example, doctors can access patients' medical records but cannot modify them, and administrators can manage all data but cannot view specific medical records. By designing a reasonable permission management strategy, data security and privacy protection are ensured. The corresponding multi-source user information distributed database is generated according to the multi-level access control framework. The multi-source user information distributed database provides a unified storage and access interface by integrating data from different sources. For example, by integrating data from different hospitals into the same distributed database, data sharing and query across hospitals can be achieved.

[0105] In a specific embodiment, the process of executing step S106 may specifically include the following steps:

[0106] (1) Query authorized users based on a multi-source user information distributed database, and group authorized users by security level to obtain user security level classification;

[0107] (2) According to the user security level classification, the sensitive data in the structured data set is processed by TF-IDF to obtain the plaintext keyword weight matrix;

[0108] (3) Perform AES encryption on the plaintext keyword weight matrix to obtain the initial encryption key, and perform round function calculation based on the initial encryption key to obtain the final encryption key;

[0109] (4) Encrypt the sensitive data using a matrix function to obtain encrypted ciphertext data, and perform inverse matrix function processing on the encrypted ciphertext data to obtain a decryption matrix;

[0110] (5) Analyze the encrypted ciphertext data to obtain a searchable ciphertext index, and review the user access rights of authorized users according to user security level classification to obtain access authorization results;

[0111] (6) Encrypt the search request of the authorized user according to the access authorization result to obtain the number of ciphertexts and keyword search requests, and retrieve and decrypt the encrypted ciphertext data according to the number of ciphertexts and keyword search requests to obtain the user information security access control policy.

[0112] Specifically, the information of authorized users is extracted from the database, and these users are grouped by security level according to the predefined security policy. By grouping, it is ensured that users of different levels can only access the data they are authorized to. According to the user security level classification, the sensitive data in the structured data set is processed by TF-IDF to obtain the plaintext keyword weight matrix. TF-IDF is a method to measure the importance of keywords. The weight of the keyword is obtained by calculating the frequency of each keyword in the document and the inverse frequency in the entire document set. The formula is as follows:

[0113] ;

[0114] Among them, TF Representation term In the documentation The frequency of occurrence, IDF Representation term The inverse document frequency of is defined as:

[0115] ;

[0116] in, is the total number of documents, To include the term The number of documents. Through TF-IDF processing, the keyword weight matrix in the sensitive data is obtained, and the weight of each keyword represents its importance in the entire data set. The plaintext keyword weight matrix is ​​encrypted by AES to obtain the initial encryption key, and the round function calculation is performed based on the initial encryption key to obtain the final encryption key. AES is a symmetric encryption algorithm that generates the final encryption key through multiple rounds of function calculations. Each round function includes four steps: byte substitution, row shift, column confusion, and round key addition. The security of encryption can be enhanced through multiple rounds of functions. Matrix function encryption is performed on sensitive data to obtain encrypted ciphertext data. Matrix function encryption is an advanced encryption technology that implements encryption processing of sensitive data by performing matrix operations on data. Suppose there is a sensitive data matrix D, which is multiplied by the encryption matrix Get the encrypted ciphertext data

[0117] ;

[0118] in, The encryption matrix generated by AES. The decryption matrix is ​​obtained by performing inverse matrix function processing on the encrypted ciphertext data. , used for subsequent decryption operations:

[0119] ;

[0120] Through inverse matrix function processing, data consistency and security are ensured during encryption and decryption. The encrypted ciphertext data is analyzed to obtain a searchable ciphertext index. Searchable encryption technology allows keyword search without decrypting data, thereby achieving efficient data retrieval while protecting data privacy. For example, using inverted index technology, the encrypted keywords and document correspondence are indexed to achieve fast search. According to the user security level classification, the user access rights of authorized users are reviewed to obtain access authorization results. Access rights review is a key step to ensure that only authorized users can access specific data. According to the access authorization results, the search request of the authorized user is encrypted to obtain the ciphertext quantity and keyword search request. For example, the search request submitted by the user contains several keywords, which are converted into encrypted form through encryption technology, and a ciphertext quantity request is generated to query the matching items in the encrypted data. According to the ciphertext quantity and keyword search request, the encrypted ciphertext data is retrieved and decrypted. For example, the user searches for patient records related to "hypertension", matches the relevant records in the encrypted data through the ciphertext search request, and uses the decryption matrix to decrypt the retrieved ciphertext data to obtain the plaintext data.

[0121] The above describes the management method of clinical recruitment users in the embodiment of the present application. The following describes the management device of clinical recruitment users in the embodiment of the present application. Figure 2 In one embodiment of the present application, a management device for clinical recruitment of users includes:

[0122] Modeling module 201, used to acquire clinical expert knowledge and perform ontology concept modeling, and construct an ontology concept model for clinical recruitment user management;

[0123] The analysis module 202 is used to obtain multi-source heterogeneous user information and use a bidirectional long short-term memory neural network and a conditional random field to perform deep learning analysis on the multi-source heterogeneous user information to obtain structured user information;

[0124] A calculation module 203 is used to calculate the cosine distance and Jaccard coefficient of the structured user information to obtain fused user information;

[0125] Processing module 204, used to perform graph-structured processing on the fused user information and ontology concept model to obtain a target knowledge graph for clinical recruitment user management;

[0126] The storage module 205 is used to perform distributed storage of clinical expert knowledge and multi-source heterogeneous user information based on the target knowledge graph, and to construct a multi-source user information distributed database;

[0127] The control module 206 is used to perform searchable encryption and multi-level access control on the multi-source user information distributed database to obtain a user information security access control policy.

[0128] Through the collaboration of the above components, a structured knowledge system was constructed by ontology concept modeling of clinical expert knowledge, which improved the accuracy and completeness of the representation of clinical recruitment domain knowledge. The deep learning analysis of multi-source heterogeneous user information using bidirectional long short-term memory neural network and conditional random field significantly improved the ability to extract structured information from complex data. The multi-dimensional similarity analysis of user information was realized by combining cosine distance and Jaccard coefficient for calculation, which improved the accuracy and comprehensiveness of user information fusion. By graph-structuring the fused user information and ontology concept model, a knowledge graph with rich semantics was constructed. Distributed storage was performed based on the target knowledge graph, which improved the data processing capability and query efficiency of the system. The distributed database was searchable and encrypted, which effectively improved the security of sensitive information while ensuring data availability. By implementing a multi-level access control strategy, user authority management was refined, and query optimization strategies were designed based on the knowledge graph structure, which significantly improved the execution efficiency of complex queries. By designing a knowledge graph update synchronization strategy, real-time updates of distributed databases were achieved, ensuring the timeliness and consistency of data, thereby improving the security of user information.

[0129] The present application also provides a management device for clinical recruitment users, which includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor executes the steps of the clinical recruitment user management method in the above-mentioned embodiments.

[0130] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the method for managing clinical recruitment users.

[0131] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0132] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.

[0133] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for managing clinical recruitment users, characterized in that: The method comprises: Acquire clinical expert knowledge and conduct ontology concept modeling to build an ontology concept model for clinical recruitment user management; Acquire multi-source heterogeneous user information, and use a bidirectional long short-term memory neural network and a conditional random field to perform deep learning analysis on the multi-source heterogeneous user information to obtain structured user information; Calculating the cosine distance and Jaccard coefficient of the structured user information to obtain fused user information; specifically including: performing feature vectorization on the structured user information to obtain a user feature vector set, and performing inter-vector cosine similarity calculation on the user feature vector set to obtain an inter-user similarity matrix; performing one-hot encoding on the category attributes in the structured user information to obtain a category feature binary matrix, and performing Jaccard coefficient calculation based on the category feature binary matrix to obtain an inter-user Jaccard similarity matrix; performing weighted fusion on the inter-user similarity matrix and the inter-user Jaccard similarity matrix to obtain a comprehensive similarity matrix; performing user clustering analysis based on the comprehensive similarity matrix to obtain a user group division result, and performing representative feature extraction on each group in the user group division result to obtain a group feature vector; performing inter-group similarity calculation based on the group feature vector to obtain an inter-group correlation matrix, and performing threshold screening on the inter-group correlation matrix to obtain a high-correlation group pair; performing cross-group fusion on the structured user information based on the high-correlation group pair to obtain fused user information; Performing graph-structuring processing on the fused user information and the ontology concept model to obtain a target knowledge graph for clinical recruitment user management; Based on the target knowledge graph, the clinical expert knowledge and the multi-source heterogeneous user information are distributedly stored to build a multi-source user information distributed database; The multi-source user information distributed database is searchably encrypted and multi-level access controlled to obtain a user information security access control strategy.

2. The method for managing clinical recruitment users according to claim 1, characterized in that: The acquisition of clinical expert knowledge and ontology concept modeling to construct an ontology concept model for clinical recruitment user management includes: Acquiring clinical expert knowledge and classifying the clinical expert knowledge into fields to obtain multiple clinical recruitment related fields; Extracting domain key concepts from the multiple clinical recruitment related fields to obtain a domain core concept set, and hierarchically organizing the domain core concept set to obtain a concept hierarchy tree; According to the concept hierarchy tree, the relationship between concepts is defined to obtain a concept relationship network, and attribute feature analysis is performed on the concept relationship network to obtain a concept attribute set; According to the concept attribute set, the constraint relationship between attributes is defined to obtain attribute constraint rules; Integrating the concept hierarchy tree, the concept relationship network and the attribute constraint rules to obtain an initial ontology model; Model consistency verification is performed according to the initial ontology model to obtain a consistency verification result, and the initial ontology model is iteratively optimized according to the consistency verification result to obtain an ontology conceptual model for clinical recruitment user management.

3. The method for managing clinical recruitment users according to claim 1, characterized in that: The method of acquiring multi-source heterogeneous user information and performing deep learning analysis on the multi-source heterogeneous user information using a bidirectional long short-term memory neural network and a conditional random field to obtain structured user information includes: Acquire multi-source heterogeneous user information, and perform data cleaning on the multi-source heterogeneous user information to obtain standardized text data; According to the standardized text data, words are segmented and part-of-speech tagged to obtain a word sequence, and the word sequence is word embedded to obtain a word vector sequence; Using a bidirectional long short-term memory neural network, performing a sequence forward LSTM calculation on the word vector sequence to obtain a forward hidden state sequence, and performing a reverse LSTM calculation on the word vector sequence to obtain a reverse hidden state sequence; The forward hidden state sequence and the reverse hidden state sequence are concatenated to obtain a bidirectional LSTM feature representation, and the bidirectional LSTM feature representation is fully connected to obtain an emission score matrix; By using a conditional random field, a state transfer calculation is performed on the emission score matrix to obtain an optimal labeling sequence; Entity recognition is performed on the optimal annotation sequence to obtain a named entity set, and entity relationship extraction is performed on the named entity set to obtain structured user information.

4. The method for managing clinical recruitment users according to claim 1, characterized in that: The graph-structured processing of the fused user information and the ontology concept model to obtain a target knowledge graph for clinical recruitment user management includes: Performing entity recognition on the fused user information to obtain a user entity set, and performing entity type mapping based on the user entity set and the ontology conceptual model to obtain a typed entity set; Extracting relationships from the typed entity set to obtain an inter-entity relationship set, and performing relationship semantic alignment based on the inter-entity relationship set and the ontology concept model to obtain a semantic relationship set; Performing graph data structure conversion on the typed entity set and the semantic relationship set to obtain an initial knowledge graph, and embedding entities and relationships in the initial knowledge graph to obtain a graph embedding vector set; Performing similarity calculation on the graph embedding vector set to obtain a similarity matrix of entities and relationships, and disambiguating redundant and contradictory information in the initial knowledge graph based on the similarity matrix to obtain an optimized knowledge graph; The optimized knowledge graph is subjected to ontology reasoning to obtain a reasoned and extended knowledge graph, and the reasoned and extended knowledge graph is hierarchically organized into a graph structure to obtain a target knowledge graph for clinical recruitment user management.

5. The method for managing clinical recruitment users according to claim 4, characterized in that: The method of distributing the clinical expert knowledge and the multi-source heterogeneous user information based on the target knowledge graph to construct a multi-source user information distributed database includes: Performing structural analysis on the target knowledge graph to obtain data models of entity tables, relationship tables, and attribute tables; According to the data model, data mapping is performed on the clinical expert knowledge and the multi-source heterogeneous user information to obtain a structured data set; Designing a data sharding strategy for the structured data set to obtain a sharding scheme, and allocating data distributed storage nodes according to the sharding scheme to obtain an initial data distribution result; Design a data synchronization mechanism for the initial data distribution result to obtain a knowledge graph update synchronization strategy, and update the distributed database in real time according to the knowledge graph update synchronization strategy to obtain a dynamic data storage structure; Designing a query optimizer for the dynamic data storage structure to obtain a query optimization strategy based on the knowledge graph structure, and generating a complex query execution plan based on the query optimization strategy to obtain an optimized query execution plan; Designing a user role and authority management module for the distributed database to obtain a multi-level access control framework; A corresponding multi-source user information distributed database is generated according to the multi-level access control framework.

6. The method for managing clinical recruitment users according to claim 5, characterized in that: The method of performing searchable encryption and multi-level access control on the multi-source user information distributed database to obtain a user information security access control policy includes: Querying authorized users based on the multi-source user information distributed database, and grouping the authorized users by security level to obtain user security level classification; According to the user security level classification, the sensitive data in the structured data set is processed by TF-IDF to obtain a plaintext keyword weight matrix; Performing AES encryption on the plaintext keyword weight matrix to obtain an initial encryption key, and performing round function calculation based on the initial encryption key to obtain a final encryption key; Performing matrix function encryption on the sensitive data to obtain encrypted ciphertext data, and performing inverse matrix function processing on the encrypted ciphertext data to obtain a decryption matrix; Analyze the encrypted ciphertext data to obtain a searchable ciphertext index, and review the user access rights of the authorized user according to the user security level classification to obtain an access authorization result; The search request of the authorized user is encrypted according to the access authorization result to obtain the number of ciphertexts and the keyword search request, and the encrypted ciphertext data is retrieved and decrypted according to the number of ciphertexts and the keyword search request to obtain the user information security access control policy.

7. A management device for clinical recruitment of users, characterized in that: Used to execute the clinical recruitment user management method according to any one of claims 1 to 6, the clinical recruitment user management device comprises: Modeling module, used to acquire clinical expert knowledge and conduct ontology concept modeling, and build an ontology concept model for clinical recruitment user management; An analysis module is used to obtain multi-source heterogeneous user information, and use a bidirectional long short-term memory neural network and a conditional random field to perform deep learning analysis on the multi-source heterogeneous user information to obtain structured user information; A calculation module, used to calculate the cosine distance and Jaccard coefficient of the structured user information to obtain fused user information; A processing module, used for performing graph-structured processing on the fused user information and the ontology concept model to obtain a target knowledge graph for clinical recruitment user management; A storage module, used for distributing and storing the clinical expert knowledge and the multi-source heterogeneous user information based on the target knowledge graph, and constructing a multi-source user information distributed database; The control module is used to perform searchable encryption and multi-level access control on the multi-source user information distributed database to obtain a user information security access control strategy.

8. A management device for clinical recruitment of users, characterized in that: The management device for clinical recruitment of users includes: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory to enable the clinical recruitment user management device to execute the clinical recruitment user management method according to any one of claims 1 to 6.

9. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the method for managing clinical recruitment users according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Non-structured data-oriented domain knowledge extraction method

    CN115510245A

  • Clinical decision-making method, system and equipment based on knowledge graph and natural language processing technology

    CN117316466A