AI-based clinical patient recruitment method and system

By serializing patient information and building a multi-dimensional index structure, combining dynamic adaptability matrix and patient similarity network, the problem of patient status modeling and matching in existing AI recruitment methods is solved, personalized recommendation and efficient recruitment of clinical trials are achieved, and matching accuracy and patient satisfaction are improved.

CN118711731BActive Publication Date: 2025-05-16BEIJING HOUPU PHARM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410937758.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-12
Publication Date
2025-05-16
Estimated Expiration
2044-07-12

AI Technical Summary

Technical Problem

Existing AI recruitment methods are difficult to accurately model patient status, static matching is difficult to adapt to the dynamic nature of clinical trials, and it is difficult to fully utilize patient behavior data for personalized recommendations while protecting patient privacy, neglecting the similarity and group characteristics between patients, making it difficult to achieve accurate matching.

Method used

By serializing the patient information, a patient status vector is formed and stored in a multi-dimensional index structure, probability analysis and abnormal detection training are carried out, a dynamic adaptability matrix and patient similarity network are constructed, and the dynamic matching and personalized recommendations of patient status and clinical trial requirements are achieved.

Benefits of technology

The accuracy and patient satisfaction of patient recruitment matching are improved, and the personalized recommendation of clinical trials is achieved through the dynamic adaptability matrix and patient similarity network, multiple goals of patient allocation are balanced, the overall recruitment efficiency is improved, and the recruitment strategy is dynamically adjusted through abnormal detection models and real-time feedback data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118711731B_ABST
    Figure CN118711731B_ABST
Patent Text Reader

Abstract

The present application relates to the field of information management technology, and discloses a clinical patient recruitment method and system based on AI. The method includes: serializing patient information to obtain a patient state vector and a state domain table; performing probability analysis on the state domain table to obtain a state transition diagram and inputting the decision tree algorithm for anomaly detection training to obtain an anomaly detection model; performing matrix processing to obtain a dynamic adaptability matrix; performing similarity calculation to obtain a comprehensive similarity index of patients and construct a patient similarity network; retrieving similar patients based on the patient similarity network and extracting a label set, and performing weighted calculation on the label set according to the dynamic adaptability matrix to obtain a personalized test recommendation list; performing multi-objective optimization to obtain an initial patient allocation plan, and generating a target patient recruitment strategy based on real-time feedback data through an anomaly detection model. The present application improves the recruitment matching accuracy and patient satisfaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information management technology, and in particular to an AI-based clinical patient recruitment method and system. Background Art

[0002] Clinical trials are a key part of medical research and new drug development, and patient recruitment is one of the most challenging parts. Traditional patient recruitment methods often rely on manual screening and matching, which is inefficient and error-prone. With the digitization of medical data and the rapid development of artificial intelligence technology, AI-based clinical patient recruitment methods have gradually attracted attention and are expected to significantly improve recruitment efficiency and accuracy.

[0003] However, current AI recruitment methods still face many problems. The complexity and heterogeneity of patient data make it a major challenge to accurately model patient status. Secondly, the dynamic nature of clinical trials and the variability of patient status make it difficult for static matching methods to adapt to actual needs. In addition, how to fully utilize patient behavior data for personalized recommendations while protecting patient privacy is also an urgent problem to be solved. Existing recruitment methods often ignore the similarities and group characteristics between patients, making it difficult to achieve accurate matching. Summary of the invention

[0004] The present application provides an AI-based clinical patient recruitment method and system for improving recruitment matching accuracy and patient satisfaction.

[0005] In a first aspect, the present application provides an AI-based clinical patient recruitment method, the AI-based clinical patient recruitment method comprising:

[0006] Serializing the patient information to obtain a patient state vector, and storing the patient state vector into a multidimensional index structure to obtain a state domain table;

[0007] Performing probability analysis on the state domain table to obtain a state transition diagram, and inputting the state transition diagram into a decision tree algorithm for anomaly detection training to obtain an anomaly detection model;

[0008] Matrixing the patient state vector and clinical trial information to obtain a dynamic adaptability matrix;

[0009] Performing similarity calculation on the patient feature matrix and the patient behavior data to obtain a comprehensive patient similarity index, and constructing an index for the comprehensive patient similarity index to obtain a patient similarity network;

[0010] Retrieving similar patients based on the patient similarity network and extracting clinical trial information labels for the similar patients to obtain a target label set, and performing weighted calculation on the target label set according to the dynamic adaptability matrix to obtain a personalized trial recommendation list;

[0011] The personalized trial recommendation list and the patient state vector are subjected to multi-objective optimization to obtain an initial patient allocation plan, and a target patient recruitment strategy is generated according to real-time feedback data through the anomaly detection model.

[0012] In a second aspect, the present application provides an AI-based clinical patient recruitment system, the AI-based clinical patient recruitment system comprising:

[0013] A processing module is used to serialize the patient information to obtain a patient state vector, and store the patient state vector into a multidimensional index structure to obtain a state domain table;

[0014] An analysis module is used to perform probability analysis on the state domain table to obtain a state transition diagram, and input the state transition diagram into a decision tree algorithm for anomaly detection training to obtain an anomaly detection model;

[0015] A matrixing module, used for matrixing the patient state vector and clinical trial information to obtain a dynamic adaptability matrix;

[0016] A calculation module, used for performing similarity calculation on the patient feature matrix and the patient behavior data to obtain a comprehensive patient similarity index, and constructing an index for the comprehensive patient similarity index to obtain a patient similarity network;

[0017] An extraction module, configured to retrieve similar patients based on the patient similarity network and extract clinical trial information labels for the similar patients to obtain a target label set, and to perform weighted calculation on the target label set according to the dynamic adaptability matrix to obtain a personalized trial recommendation list;

[0018] A generation module is used to perform multi-objective optimization on the personalized test recommendation list and the patient state vector to obtain an initial patient allocation plan, and generate a target patient recruitment strategy according to real-time feedback data through the anomaly detection model.

[0019] The third aspect of the present application provides an AI-based clinical patient recruitment device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the AI-based clinical patient recruitment device executes the above-mentioned AI-based clinical patient recruitment method.

[0020] The fourth aspect of the present application provides a computer-readable storage medium, which stores instructions that, when executed on a computer, enable the computer to execute the above-mentioned AI-based clinical patient recruitment method.

[0021] In the technical solution provided by the present application, by storing the patient state vector into a multidimensional index structure to form a state domain table, efficient storage and rapid retrieval of patient information are achieved. By using the state transition diagram and the anomaly detection model, the patient state change can be accurately predicted and abnormal situations can be identified in time. Through the dynamic adaptability matrix, the dynamic matching of the patient state and the clinical trial requirements is achieved. The patient similarity network is constructed by combining the patient characteristics and behavior data. Based on the patient similarity network and the dynamic adaptability matrix, personalized recommendations for clinical trials are achieved, and the patient's acceptance of the recommended trials is improved. Through the multi-objective optimization algorithm, multiple goals of patient allocation, such as maximizing participation and minimizing recruitment time, are balanced, and the overall recruitment efficiency is improved. By using the anomaly detection model and real-time feedback data, the recruitment strategy can be dynamically adjusted, and a time-sensitive decision model is introduced to take into account the changes of patient status and trial requirements over time. Through the reinforcement learning environment, it can continuously learn and improve from the recruitment process. Through the matching evaluation, it is ensured that the final generated recruitment strategy is highly consistent with the patient's state, and the recruitment matching accuracy and patient satisfaction are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0023] Figure 1 This is a schematic diagram of an embodiment of the AI-based clinical patient recruitment method in the embodiments of the present application;

[0024] Figure 2 This is a schematic diagram of an embodiment of an AI-based clinical patient recruitment system in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The embodiments of the present application provide a clinical patient recruitment method and system based on AI. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In the embodiments of the present application, an embodiment of the clinical patient recruitment method based on AI includes:

[0027] Step S101, serializing the patient information to obtain a patient state vector, and storing the patient state vector into a multidimensional index structure to obtain a state domain table;

[0028] It is understandable that the execution subject of the present application can be an AI-based clinical patient recruitment system, or a terminal or a server, which is not limited here. The present application embodiment is described by taking the server as the execution subject as an example.

[0029] Specifically, the electronic health records of the patient are obtained from the patient information. These records contain the patient's medical history, diagnosis information, treatment records and other data. The electronic health records are structured and converted into analyzable structured health status data. According to the structured health status data, the patient's health status is discretized, and the continuous health indicators are segmented or classified to obtain discrete health status. The inclusion and exclusion criteria, trial stages and treatment plans of the clinical trial information are encoded to obtain the trial status code. The inclusion and exclusion criteria refer to the conditions under which the patient can participate in the trial, the trial stages are the different stages of the clinical trial (such as the early, middle and late stages), and the treatment plan involves specific treatment measures. The encoding of this information converts it into numbers or symbols, so that the computer can easily process and analyze it. Based on the discrete health state and trial state coding, a state vector template is constructed. The template is a predefined structure used to guide subsequent data filling and finally form a patient state vector structure. The patient's medical examination results are obtained from the patient information, including blood tests, imaging examinations, pathological examinations, etc. These examination results provide data on the patient's current health status. The examination results and corresponding treatment plans are numerically processed to obtain a health indicator set. According to the health indicator set and the test state encoding, the patient state vector structure is filled according to the predefined state vector template to obtain the initial state vector. The initial state vector is timestamped to record the time when the data was generated to obtain the patient state vector. A multidimensional index structure is created for fast retrieval and operation of multidimensional data. The patient state vector is stored in the multidimensional index structure to form a state domain table. The state domain table is a collection of all patient state vectors, which are organized and managed through a multidimensional index structure so that the required data can be quickly and accurately obtained during subsequent analysis, query and processing.

[0030] Step S102: Perform probability analysis on the state domain table to obtain a state transition diagram, and input the state transition diagram into a decision tree algorithm for anomaly detection training to obtain an anomaly detection model;

[0031] Specifically, the historical data in the state domain table is analyzed in time series, and the state transition frequency matrix is ​​obtained by analyzing the law of patient status changes over time. The frequency matrix records the frequency of transitions from one state to another. According to the state transition frequency matrix, the transition probability between each state is calculated to obtain the Markov transition matrix. Through probability calculation, the frequency information of state transition is converted into a quantifiable probability index, thereby reflecting the dynamic relationship between different states. The Markov transition matrix is ​​graphically processed to generate an initial transition graph. In this graph, nodes represent different health states of patients, and edges represent the transition relationship between states and their probabilities. By analyzing the initial transition graph, node features and edge features are extracted. These features include the node's in-and-out degree, the edge's weight, etc., which constitute the state transition feature set. Cluster analysis is performed on the state transition feature set, and the state combination pattern is obtained by clustering states with similar transition features together. According to the state combination pattern, which state transitions are common and normal and which are abnormal, so as to mark the abnormal state transitions and form the final state transition graph. The state transition graph is sampled to separate the training data set and the validation data set. The training data set is used to train the decision tree algorithm. By learning the rules of state transition, the initial anomaly detector is obtained. The initial anomaly detector is integrated and learned to obtain a more robust initial detection model by combining the judgments of multiple detectors. The validation data set is used to evaluate the performance of the initial detection model, analyze its detection accuracy and false alarm rate, and perform threshold optimization to determine the best detection threshold, and finally obtain the anomaly detection model.

[0032] Step S103, matrix processing is performed on the patient state vector and the clinical trial information to obtain a dynamic adaptability matrix;

[0033] Specifically, the patient state vector is feature extracted and standardized to obtain the patient feature matrix. By analyzing the various health indicators and behavioral data of the patient, it is converted into a set of representative feature values ​​for subsequent analysis. Standardization is to normalize the feature values ​​of different dimensions so that they have the same scale to avoid the influence of different dimensions on the results. At the same time, the clinical trial information is encoded and quantified according to the inclusion and exclusion criteria to obtain the test requirement matrix. By converting the various requirements and conditions of the test into a computable numerical form, the test information can be effectively compared with the patient characteristics. The patient feature matrix and the test requirement matrix are matrix multiplied to obtain the static adaptability matrix. The patient feature matrix is ​​predicted in time series according to the state transition diagram to obtain the predicted state matrix. Time series prediction predicts the future state changes by analyzing the historical change law of the patient's state, and more accurately evaluates the patient's state in the future. The predicted state matrix and the test requirement matrix are matrix multiplied to obtain the target adaptability matrix. By matching the predicted patient state with the test requirements, the adaptability of the future state is evaluated. The time decay function is designed according to the static adaptability matrix and the target adaptability matrix to obtain the time weight vector. The time decay function reflects the change law of adaptability over time. The static adaptability matrix and the target adaptability matrix are weighted and fused through the time weight vector to obtain the first adaptability matrix. The test correlation matrix is ​​constructed according to the historical participation data in the patient information, and the first adaptability matrix is ​​optimized by collaborative filtering to obtain the second adaptability matrix. Collaborative filtering is a recommendation algorithm based on historical data. It optimizes the accuracy of adaptability assessment by analyzing the participation data of similar patients. The second adaptability matrix is ​​decomposed to obtain a low-dimensional representation matrix. Through dimensionality reduction processing, the high-dimensional adaptability data is converted into a representation form in a low-dimensional space to simplify the computational complexity. The matrix is ​​reconstructed and adjusted according to the low-dimensional representation matrix and the output detection information of the anomaly detection model to obtain a dynamic adaptability matrix. The output information of the anomaly detection model is used to identify and correct abnormal data in the adaptability assessment. Through matrix reconstruction and adjustment, the accuracy and reliability of adaptability are improved.

[0034] Step S104: Calculate the similarity of the patient feature matrix and the patient behavior data to obtain a comprehensive similarity index of the patient, and construct an index for the comprehensive similarity index of the patient to obtain a patient similarity network;

[0035] Specifically, the cosine similarity between patients is calculated according to the patient feature matrix to obtain a feature similarity matrix. The similarity is measured by calculating the cosine value of the angle between two vectors. The closer the value is to 1, the more similar the two are. At the same time, the browsing history data of the digital health platform in the patient information is serialized to obtain a behavior sequence vector. By converting the discrete browsing history data into an ordered vector sequence, the data is made more structured and convenient for subsequent analysis. The behavior sequence vector is analyzed for interest features to obtain a patient interest feature matrix. The interest feature analysis analyzes the browsing habits and behavior patterns of patients on the digital health platform, extracts feature values ​​that can represent their interests, and forms an interest feature matrix. The patient interest feature matrix is ​​similarity calculated to obtain an interest similarity matrix. The feature similarity matrix and the interest similarity matrix are weighted and fused to obtain a comprehensive similarity matrix. The comprehensive similarity matrix is ​​threshold filtered to obtain a sparse similarity matrix. By setting a similarity threshold, the similarity values ​​below the threshold are set to zero to obtain a more sparse matrix. The initial similarity network is constructed according to the sparse similarity matrix. In this network, nodes represent patients, and edges represent similar relationships between patients. Only those connections with similarities greater than a threshold are retained to form a more representative similarity network. The initial similarity network is subjected to local sensitive hashing to obtain a hash index structure. Local sensitive hashing is an approximate nearest neighbor search algorithm for high-dimensional data. It significantly improves the efficiency of similarity calculation by mapping similar data points to the same hash bucket. According to the hash index structure and state transition model, the initial similarity network is dynamically updated to obtain the final patient similarity network. The state transition model dynamically adjusts the structure of the similarity network by analyzing the historical changes in the patient's status, so that the network can reflect changes in the patient's status in a timely manner.

[0036] Step S105: Retrieve similar patients based on the patient similarity network and extract clinical trial information labels for the similar patients to obtain a target label set, and perform weighted calculation on the target label set according to the dynamic adaptability matrix to obtain a personalized trial recommendation list;

[0037] Specifically, a graph traversal analysis is performed on the patient similarity network. Through the traversal analysis, the neighbor set of each patient in the similarity network is identified. The neighbor set contains a group of patients who are most similar to the target patient in terms of characteristics and behaviors. According to the neighbor set, the clinical trial information that similar patients have browsed is extracted to form a test information set. The test information set is processed to obtain a test description text corpus. After the corpus is established, a word frequency count is performed to obtain an initial tag set. By counting the frequency of words in the test description text, a set of keywords or tags are identified, which preliminarily reflect the theme and content of the clinical trial. The semantic similarity of the initial tag set is calculated, and a tag similarity matrix is ​​constructed by calculating the semantic similarity between each tag. Hierarchical clustering is performed based on the tag similarity matrix to obtain the target tag set. The hierarchical clustering method aggregates similar tags together to form a more representative tag set, reduce redundancy, and improve the accuracy and representativeness of tags. Collaborative filtering is performed on the target tag set and the patient's implicit behavior data to obtain a patient-tag scoring matrix. Collaborative filtering analyzes patients' behavioral data (such as browsing, clicking, and participation records) to evaluate each patient's preference for different tags and form a scoring matrix. Matrix multiplication is performed on the dynamic adaptability matrix and the patient-tag scoring matrix to obtain a weighted trial score matrix. The dynamic adaptability matrix reflects the degree of match between the patient's status and the trial requirements. Through matrix multiplication, the matching degree is combined with the patient's tag preference to calculate the comprehensive score of each clinical trial. The weighted trial score matrix is ​​sorted by multiple objectives, and multiple factors (such as matching, priority, and historical records) are considered to obtain a sorted trial list. Personalized filtering is performed based on the sorted trial list and the patient's status vector to obtain a personalized trial recommendation list. Personalized filtering selects the most suitable clinical trials by combining the patient's current health status and individual needs.

[0038] Step S106: Perform multi-objective optimization on the personalized trial recommendation list and the patient state vector to obtain an initial patient allocation plan, and generate a target patient recruitment strategy based on real-time feedback data through an anomaly detection model.

[0039] Specifically, the personalized test recommendation list and the patient state vector are feature fused to obtain a multi-objective optimization vector. The optimization objective function is constructed based on the multi-objective optimization vector to obtain the initial optimization problem. The optimization objective function contains multiple objectives, such as maximizing patient matching, minimizing test time and cost, etc. The initial optimization problem is subjected to multi-objective particle swarm optimization to obtain a non-dominated solution set. Multi-objective particle swarm optimization is a multi-objective optimization algorithm based on particle swarm optimization. By simulating the movement and evolution of particle swarms, a set of optimal solutions, namely, a non-dominated solution set, is found. These solutions perform well on multiple objectives. Pareto optimal analysis is performed on the non-dominated solution set to obtain the initial patient allocation plan. Pareto optimal analysis ensures that the overall benefit of the allocation plan is maximized by selecting the optimal solution that is irreplaceable on multiple objectives. Time-sensitive decision-making is performed on the initial patient allocation plan to obtain a sequential allocation strategy. By considering time factors, such as the dynamic changes in patient status and the time requirements of the test, the execution time sequence of the allocation plan is optimized to form a time-sensitive allocation strategy. According to the output detection information of the sequential allocation strategy and the anomaly detection model, a reinforcement learning environment is constructed to obtain a state-action space. The reinforcement learning environment simulates the actual patient recruitment process, forming a state-action space by combining states (patient status and trial requirements) with actions (allocation strategies). The state-action space is processed by Q learning to obtain the optimal action value function. Q learning is a model-free reinforcement learning algorithm that evaluates the value of each state-action pair through continuous iteration and learning, and ultimately finds a strategy that maximizes the total reward. The strategy is updated based on the optimal action value function and real-time feedback data to obtain a dynamic patient recruitment strategy. By incorporating real-time feedback data (such as patient responses and trial progress) into the learning process, the recruitment strategy is dynamically adjusted and optimized to ensure its effectiveness and adaptability in practical applications. The dynamic patient recruitment strategy is searched to obtain the strategy execution sequence, and the matching degree is evaluated based on the strategy execution sequence and the patient state vector to obtain the target patient recruitment strategy. The strategy execution sequence is a specific operation step based on the optimal action value function. By evaluating its matching degree with the patient state vector, it ensures that the final recruitment strategy not only meets the patient's needs but also maximizes the success rate of the clinical trial.

[0040] In the embodiment of the present application, by storing the patient state vector into a multidimensional index structure to form a state domain table, efficient storage and rapid retrieval of patient information are achieved. By using the state transition diagram and the anomaly detection model, the patient state change can be accurately predicted and abnormal situations can be identified in time. Through the dynamic adaptability matrix, the dynamic matching of the patient state and the clinical trial requirements is achieved. The patient similarity network is constructed by combining the patient characteristics and behavior data. Based on the patient similarity network and the dynamic adaptability matrix, personalized recommendations for clinical trials are achieved, and the patient's acceptance of the recommended trial is improved. Through the multi-objective optimization algorithm, multiple goals of patient allocation are balanced, such as maximizing participation, minimizing recruitment time, etc., to improve the overall recruitment efficiency. By using the anomaly detection model and real-time feedback data, the recruitment strategy can be dynamically adjusted, and a time-sensitive decision model is introduced to take into account the changes in patient status and trial requirements over time. Through the reinforcement learning environment, it can continuously learn and improve from the recruitment process. Through the matching evaluation, it is ensured that the final generated recruitment strategy is highly consistent with the patient's state, and the recruitment matching accuracy and patient satisfaction are improved.

[0041] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0042] (1) Obtaining the patient's electronic health record from the patient information, structuring the patient's electronic health record to obtain structured health status data, and discretizing the patient's health status based on the structured health status data to obtain a discrete health status;

[0043] (2) Encode the inclusion criteria, trial phase, and treatment plan of the clinical trial information to obtain the trial status code, and construct a state vector template based on the discrete health state and trial status code to obtain the patient state vector structure;

[0044] (3) Obtain the patient's medical examination results from the patient information, and digitize the patient's medical examination results and treatment plan to obtain a set of health indicators;

[0045] (4) Fill the patient state vector structure according to the health indicator set and the test state code to obtain the initial state vector, and timestamp the initial state vector to obtain the patient state vector;

[0046] (5) Create a multidimensional index structure and store the patient state vector into the multidimensional index structure to obtain a state domain table.

[0047] Specifically, the patient's electronic health record is parsed. Electronic health records are digital archives containing information such as patient medical history, diagnostic information, treatment records, and test results. This information needs to be extracted and organized for subsequent processing. By using natural language processing technology and standardized medical terminology libraries, unstructured text data is converted into structured health status data. The structured data can be represented as a multidimensional array, where each dimension represents a specific health indicator or diagnostic result. Based on the structured health status data, the patient's health status is discretized. Continuous health indicators are converted into discrete health states for more effective analysis and comparison. For example, blood sugar levels can be discretized into categories such as "normal", "high", and "high risk". Set discrete threshold standards, such as setting blood sugar levels The discretization standard is:

[0048] ;

[0049] The inclusion criteria, trial phase and treatment plan of clinical trial information are encoded to obtain the trial status code. The inclusion criteria refer to whether the patient meets the conditions for participating in the trial, the trial phase refers to the different stages of the trial (such as the early, middle and late stages), and the treatment plan is the specific medical intervention measures used in the trial. This information is encoded as numerical values ​​or symbols for calculation and matching. For example, the inclusion criteria can be encoded as a binary variable, where 1 indicates that the conditions are met and 0 indicates that the conditions are not met; the trial phase can be encoded as an integer variable, such as 1 for the early stage, 2 for the middle stage, and 3 for the late stage; the treatment plan can be represented by a string or identifier. Based on the discrete health state and trial state encoding, a state vector template is constructed to guide subsequent data filling to form a patient state vector structure. The patient's medical examination results are obtained from the patient information, and these results and treatment plans are numerically processed to obtain a set of health indicators. The medical examination results include the numerical results of various diagnostic tests, such as blood pressure, heart rate, blood sugar, etc. These data need to be standardized and numerically processed so that they can be compared and analyzed within the same dimension. Numerical processing can be achieved through normalization, standardization and other methods. According to the health indicator set and the test state encoding, fill in the patient state vector structure to obtain the initial state vector. The initial state vector includes all relevant health data and test information to form a comprehensive state representation. The initial state vector of each patient can be represented as a multidimensional vector , where each component represents a specific health indicator or test information. For example, ,in is a discrete state of blood sugar, It is the test state encoding. In order to ensure the temporal traceability of the data, the initial state vector is timestamped, the time when the data was generated is recorded, and the patient state vector is obtained. The timestamp can adopt a standard time format so that each state vector has clear time information, reflecting the dynamic changes in the patient's health status. Create a multidimensional index structure, and store the patient state vector into the multidimensional index structure to form a state domain table. The multidimensional index structure can adopt efficient data structures such as R-tree or KD-tree for fast retrieval and manipulation of multidimensional data. By storing the state vector of each patient in the multidimensional index structure, a comprehensive state domain table is formed, so that the required data can be obtained quickly and accurately during subsequent analysis, query and processing.

[0050] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0051] (1) Perform time series analysis on the historical data in the state domain table to obtain the state transition frequency matrix, and calculate the transition probability between states based on the state transition frequency matrix to obtain the Markov transition matrix;

[0052] (2) Graphically process the Markov transition matrix to obtain an initial transition graph, and extract node features and edge features based on the initial transition graph to obtain a state transition feature set;

[0053] (3) Perform cluster analysis on the state transition feature set to obtain the state combination pattern, and mark the abnormal state transitions according to the state combination pattern to obtain the state transition graph;

[0054] (4) Sampling the state transition diagram to obtain a training data set and a validation data set, and training a decision tree algorithm based on the training data set to obtain an initial anomaly detector;

[0055] (5) Perform ensemble learning on the initial anomaly detector to obtain the initial detection model, and perform performance evaluation and threshold optimization on the initial detection model based on the validation data set to obtain the anomaly detection model.

[0056] Specifically, a time series analysis is performed on the historical data in the state domain table. The state domain table contains the patient's historical health status data and records the state changes at each time point. Through time series analysis, the law of state changes over time is identified and a state transition frequency matrix is ​​constructed. Elements Indicates from the state Transition to state The frequency of the state transition is obtained by counting the number of state transitions. For example, if in the historical data, the state Transition to state The number of times is 10, and the state There are 50 occurrences in total, so

[0057] ;

[0058] According to the state transition frequency matrix , calculate the transition probability between states and obtain the Markov transition matrix . Markov transition matrix Elements Indicates from the state Transition to state The probability of can be obtained by normalizing each row of the frequency matrix. The formula is as follows:

[0059] ;

[0060] in, Indicates from the state The total frequency of transitions to all other states. The Markov transition matrix is ​​obtained by normalizing so that the sum of the transition probabilities in each row is 1. For example, suppose there is a simple system with three health states: healthy, mildly ill, and severely ill. By analyzing historical data, the following state transition frequency matrix is ​​obtained:

[0061] ;

[0062] According to the frequency matrix , calculate the transition probability between states and obtain the Markov transition matrix

[0063] ;

[0064] The Markov transition matrix is ​​processed graphically to obtain an initial transition graph. In the initial transition graph, each node represents a state, and the edge represents the transition relationship between states. The weight of the edge corresponds to the transition probability. By analyzing the initial transition graph, node features and edge features are extracted to obtain a state transition feature set. Node features include the in-degree and out-degree of the node, and edge features include the weight and direction of the edge. Cluster analysis is performed on the state transition feature set to obtain a state combination pattern. Cluster analysis forms several state combination patterns by clustering states with similar transition features together. These patterns reflect the transition rules between different states. For example, the state transition features are clustered using the K-means algorithm or the hierarchical clustering algorithm to identify typical state combination patterns. Based on these patterns, abnormal state transitions are identified and marked to form a state transition graph. The state transition graph is sampled to obtain a training data set and a validation data set. The training data set is used to train the decision tree algorithm. By learning the rules of state transitions, an initial anomaly detector is obtained. The decision tree algorithm is a commonly used classification algorithm. By building a tree structure, classification and judgment are performed according to different state transition features. For example, the transition probability and node features are used as input variables to train a decision tree model to identify normal and abnormal state transitions. In order to improve the performance of the detector, the initial anomaly detector is integrated to obtain an initial detection model. Ensemble learning combines the judgments of multiple detectors to form a more robust model. Common ensemble learning methods include random forest and gradient boosting. For example, multiple decision tree models can be trained to form a final detection model by voting or weighted averaging. The initial detection model is evaluated and the threshold is optimized based on the validation data set to obtain an anomaly detection model. Performance evaluation evaluates the detection effect of the model by calculating indicators such as the accuracy, precision, recall rate and F1 score of the detection model. Threshold optimization adjusts the detection threshold to find the best balance point so that the model achieves the best balance between detection accuracy and false alarm rate. For example, the ROC curve and AUC value are used to evaluate the performance of the model. According to the results of the validation data set, the detection threshold is adjusted to achieve the best detection effect of the model.

[0065] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0066] (1) Feature extraction and standardization of the patient state vector to obtain the patient feature matrix, and encoding and quantifying the clinical trial information according to the inclusion and exclusion criteria to obtain the trial requirement matrix;

[0067] (2) Perform matrix multiplication on the patient feature matrix and the test requirement matrix to obtain a static suitability matrix;

[0068] (3) Perform time series prediction on the patient feature matrix according to the state transition diagram to obtain the predicted state matrix;

[0069] (4) Perform matrix multiplication on the predicted state matrix and the test requirement matrix to obtain the target adaptability matrix;

[0070] (5) Design a time decay function based on the static adaptability matrix and the target adaptability matrix to obtain a time weight vector;

[0071] (6) The static adaptability matrix and the target adaptability matrix are weightedly fused by using the time weight vector to obtain a first adaptability matrix;

[0072] (7) Constructing a trial relevance matrix based on the historical participation data in the patient information, and performing collaborative filtering optimization on the first suitability matrix to obtain a second suitability matrix;

[0073] (8) Performing matrix decomposition on the second adaptability matrix to obtain a low-dimensional representation matrix, and reconstructing and adjusting the matrix according to the low-dimensional representation matrix and the output detection information of the anomaly detection model to obtain a dynamic adaptability matrix.

[0074] Specifically, feature extraction is performed on the patient's state vector. These state vectors include various health indicators and medical history of the patient, such as blood pressure, blood sugar, heart rate, past disease records, etc. The purpose of feature extraction is to convert these multidimensional data into a feature set that can be used for calculation. For example, by using dimensionality reduction methods such as principal component analysis, the data dimension can be reduced and the most representative features can be extracted. These features are standardized to ensure that different features have the same scale and to prevent some features from having an uneven impact on the results due to different dimensions. The clinical trial information is encoded and quantified according to the inclusion and exclusion criteria to obtain the trial requirement matrix. The inclusion and exclusion criteria include whether the patient meets the conditions of the trial, such as age, gender, medical history and other information. By encoding multiple standards, non-numerical information is converted into numerical form. For example, age can be divided into several intervals for quantification, gender can be encoded as 0 and 1, and medical history can be encoded in binary to indicate whether there is a certain disease. The patient feature matrix and the test requirements matrix Perform matrix multiplication to obtain the static adaptability matrix . The matrix multiplication formula is:

[0075] ;

[0076] in, yes Matrix, representing of patients Features, yes Matrix, representing Experimental Requirements, yes Matrix, representing the suitability score of each patient for each test. According to the state transition diagram, the patient feature matrix is ​​predicted in time series to obtain the predicted state matrix The state transition diagram reflects the change of the patient's state over time. Through the time series analysis model (such as ARIMA or LSTM), the patient's future health state is predicted. The predicted future state is combined with the current state to form a predicted state matrix . For the predicted state matrix and the test requirements matrix Perform matrix multiplication to obtain the target adaptability matrix The calculation formula is:

[0077]

[0078] Target Suitability Matrix It reflects the degree of match between the patient's future status and the test requirements. In order to combine the influence of time factors, according to the static adaptability matrix and target suitability matrix Design a time decay function to obtain the time weight vector The time decay function reflects the change of adaptability over time, and usually an exponential decay function can be used:

[0079] ;

[0080] in, is the attenuation coefficient, is time. The time weight vector , for the static adaptability matrix and target suitability matrix Perform weighted fusion to obtain the first adaptability matrix

[0081] ;

[0082] in, and The weights of the current time and the future time are respectively. Construct the trial correlation matrix based on the historical participation data in the patient information , and the first adaptability matrix Perform collaborative filtering optimization to obtain the second adaptability matrix Collaborative filtering optimizes the suitability matrix by analyzing the historical trial participation data of similar patients, thereby improving the accuracy and personalization of recommendations. Perform matrix decomposition to obtain a low-dimensional representation matrix and . Matrix decomposition usually uses the singular value decomposition method to decompose the original matrix into the product of two low-dimensional matrices:

[0083] ;

[0084] in, and are low-dimensional representation matrices for patients and trials, respectively, is a diagonal matrix, representing the eigenvalues. According to the low-dimensional representation matrix and The output detection information of the anomaly detection model is reconstructed and adjusted to obtain a dynamic adaptability matrix The anomaly detection model is used to identify and correct abnormal data in the adaptability matrix. By combining low-dimensional representation and detection information, the matrix is ​​reconstructed to ensure the final dynamic adaptability matrix. More accurate and reliable. For example, suppose there are two patients and three clinical trials. The patient characteristics include age, gender, and medical history. After feature extraction and standardization, the following patient feature matrix is ​​obtained

[0085] ;

[0086] Test Requirements Matrix After encoding:

[0087] ;

[0088] Perform matrix multiplication to obtain the static adaptability matrix

[0089] ;

[0090] Through time series prediction, we get the predicted state matrix

[0091] ;

[0092] Perform matrix multiplication on the predicted state matrix and the test requirement matrix to obtain the target suitability matrix

[0093] ;

[0094] Design the time decay function, assuming , the time weight vector is:

[0095] ;

[0096] Weighted fusion obtains the first adaptability matrix

[0097] ;

[0098] Through collaborative filtering optimization, matrix decomposition and reconstruction are performed, and the output information of the anomaly detection model is combined to finally obtain the dynamic adaptability matrix .

[0099] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0100] (1) Calculate the cosine similarity between patients based on the patient feature matrix to obtain the feature similarity matrix, and serialize the digital health platform browsing history data in the patient information to obtain the behavior sequence vector;

[0101] (2) Perform interest feature analysis on the behavior sequence vector to obtain the patient interest feature matrix, and perform similarity calculation on the patient interest feature matrix to obtain the interest similarity matrix;

[0102] (3) Perform weighted fusion on the feature similarity matrix and the interest similarity matrix to obtain a comprehensive similarity matrix;

[0103] (4) Threshold filtering is performed on the comprehensive similarity matrix to obtain a sparse similarity matrix, and an initial similarity network is constructed based on the sparse similarity matrix;

[0104] (5) Perform local sensitive hashing calculation on the initial similarity network to obtain a hash index structure;

[0105] (6) The initial similarity network is dynamically updated according to the hash index structure and state transition model to obtain the patient similarity network.

[0106] Specifically, the patient characteristics matrix is ​​analyzed. Contains multidimensional feature data for each patient, such as age, gender, medical history, and examination results. In order to calculate the cosine similarity between patients, the feature vectors of each pair of patients are compared. The formula for calculating cosine similarity is:

[0107] ;

[0108] in, and Respectively represent and The feature vector of each patient, · represents the vector inner product, and Represents the modulus of the vector. By calculating the cosine similarity between all patient pairs, the feature similarity matrix is ​​obtained . Serialize the browsing history data of the digital health platform in the patient information to obtain the behavior sequence vector. Convert the patient's browsing behavior on the platform into an ordered vector form. For example, each browsing behavior can be represented as a specific event, arranged in chronological order to form a behavior sequence vector . For the behavior sequence vector Perform interest feature analysis to obtain the patient interest feature matrix Interest feature analysis is to extract patients' interests in different health content by analyzing their browsing behavior. For example, using topic models (such as LDA) or embedding techniques (such as Word2Vec) to convert browsing behavior into interest feature vectors. Perform similarity calculation to obtain the interest similarity matrix The feature similarity matrix and the interest similarity matrix are weighted and fused to obtain a comprehensive similarity matrix. Perform threshold filtering to obtain a sparse similarity matrix Threshold filtering is to set a similarity threshold and set the similarity values ​​below the threshold to zero to reduce the sparsity of the matrix and improve the calculation efficiency. For example, setting the threshold ,if ,but According to the sparse similarity matrix Construct an initial similarity network. In this network, nodes represent patients, edges represent similar relationships between patients, and the weights of edges correspond to non-zero values ​​in the sparse similarity matrix. The initial similarity network reflects the similarity relationship structure between patients. Perform local sensitive hashing calculation on the initial similarity network to obtain a hash index structure. Local sensitive hashing is an approximate nearest neighbor search algorithm for high-dimensional data. By mapping similar data points to the same hash bucket, similar patients can be found quickly. The vectors in the comprehensive similarity matrix are mapped to multiple hash buckets through a hash function to form a hash index structure, thereby accelerating the similarity query process. According to the hash index structure and the state transition model, the initial similarity network is dynamically updated to obtain a patient similarity network. The state transition model reflects the dynamic changes in the patient's health status. By combining the state transition model, the structure of the similarity network is adjusted in real time to make it more accurately reflect the similarity relationship between patients. For example, when the patient's health status changes, the edges and nodes in the network can be dynamically adjusted by updating the hash index and similarity matrix to ensure that the similarity network is always up to date and accurate.

[0109] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0110] (1) Perform graph traversal analysis on the patient similarity network to obtain a neighbor set, and extract clinical trial information browsed by similar patients based on the neighbor set to obtain a trial information set;

[0111] (2) Processing the test information set to obtain a test description text corpus, and performing word frequency statistics based on the test description text corpus to obtain an initial tag set;

[0112] (3) Calculate the semantic similarity of the initial label set to obtain a label similarity matrix, and perform hierarchical clustering based on the label similarity matrix to obtain the target label set;

[0113] (4) Perform collaborative filtering on the target label set and the patient implicit behavior data to obtain the patient-label score matrix, and perform matrix multiplication operation on the dynamic adaptability matrix and the patient-label score matrix to obtain the weighted trial score matrix;

[0114] (5) Perform multi-objective sorting on the weighted test score matrix to obtain a sorted test list, and perform personalized filtering based on the sorted test list and the patient state vector to obtain a personalized test recommendation list.

[0115] Specifically, a graph traversal analysis is performed on the patient similarity network. The patient similarity network is a graph structure in which nodes represent patients, edges represent similarity relationships between patients, and the weight of the edges reflects the similarity. Through graph traversal algorithms such as breadth-first search or depth-first search, other patients similar to the target patient, i.e., the neighbor set, are found. Assuming that breadth-first search is used, starting from the target patient node, the similarity network is traversed to collect patients of all adjacent nodes to form a neighbor set. According to the neighbor set, the clinical trial information browsed by similar patients is extracted to obtain a test information set. The browsing history of each patient on the digital health platform contains the clinical trial information they have browsed. By querying the browsing history of each patient in the neighbor set, the clinical trials that all similar patients have browsed are summarized to form a test information set. The test information set is processed to obtain a test description text corpus. The test description text includes detailed information such as the title, objective, method, and subject conditions of the clinical trial. A structured text corpus is obtained by preprocessing the description text, such as word segmentation, stop word removal, and stem extraction. The word frequency statistics are performed based on the test description text corpus to obtain the initial label set. Word frequency statistics is to identify the most common keywords or tags by calculating the frequency of each word in the corpus. These tags preliminarily reflect the main content and theme of the experiment. The semantic similarity of the initial tag set is calculated to obtain a tag similarity matrix. Semantic similarity calculation evaluates the similarity of different tags by measuring their semantic distance. For example, using pre-trained word vector models such as Word2Vec or GloVe, each tag is represented as a vector, and their semantic similarity is evaluated by calculating the cosine similarity between two vectors. Hierarchical clustering is performed according to the tag similarity matrix to obtain the target tag set. Hierarchical clustering is a clustering method that gradually merges similar tags into clusters. By analyzing the tag similarity matrix, redundant and similar tags in the initial tag set are merged to form a more streamlined and representative target tag set. Assuming that a bottom-up hierarchical clustering method is used, starting from each tag, similar tags are gradually merged until several target tags are formed. Collaborative filtering is performed on the target tag set and the patient's implicit behavior data to obtain a patient-tag score matrix. Implicit behavioral data includes patients' browsing, clicking, searching and other behaviors on the platform. By analyzing these behaviors, each patient's interest and preference for different tags are evaluated. Collaborative filtering predicts the target patient's score for each tag by combining the behavioral data of similar patients. Matrix multiplication is performed on the dynamic adaptability matrix and the patient-tag score matrix to obtain the weighted test score matrix. The dynamic adaptability matrix represents the degree of match between the patient's current health status and the test requirements. Through matrix multiplication, the patient's interest in the tag is combined with the adaptability of the test to calculate the comprehensive score of each test. Multi-objective sorting is performed on the weighted test score matrix to obtain a sorted test list.Multi-objective sorting sorts trials by comprehensively considering multiple factors (such as suitability, interest, and historical participation records) to determine the optimal order of trial recommendations. For example, using the weighted summation method, the scores of each factor are weighted and summarized to obtain the total score of each trial, and then sorted according to the total score. Personalized filtering is performed based on the sorted trial list and the patient state vector to obtain a personalized trial recommendation list. Personalized filtering selects the most suitable clinical trials by combining the patient's current health status and individual needs. For example, the sorted trial list is further filtered based on the patient's specific health indicators, age, gender, and other conditions to ensure that the recommended trials are not only optimal in terms of suitability and interest, but also meet the individual needs of the patient.

[0116] In a specific embodiment, the process of executing step S106 may specifically include the following steps:

[0117] (1) Feature fusion of the personalized trial recommendation list and the patient state vector to obtain a multi-objective optimization vector;

[0118] (2) Constructing the optimization objective function according to the multi-objective optimization vector to obtain the initial optimization problem, and performing multi-objective particle swarm optimization on the initial optimization problem to obtain a non-dominated solution set;

[0119] (3) Perform Pareto optimal analysis based on the non-dominated solution set to obtain the initial patient allocation plan, and make time-sensitive decisions on the initial patient allocation plan to obtain the timing allocation strategy;

[0120] (4) Construct a reinforcement learning environment based on the timing allocation strategy and the output detection information of the anomaly detection model to obtain the state-action space;

[0121] (5) Perform Q learning on the state-action space to obtain the optimal action value function, and update the strategy based on the optimal action value function and real-time feedback data to obtain a dynamic patient recruitment strategy;

[0122] (6) Search for the dynamic patient recruitment strategy to obtain the strategy execution sequence, and perform a matching evaluation based on the strategy execution sequence and the patient state vector to obtain the target patient recruitment strategy.

[0123] Specifically, the personalized test recommendation list is combined with the patient's state vector. The personalized test recommendation list contains the recommendation score of each test, while the patient's state vector includes the patient's current health status and historical data. Assume that the personalized test recommendation list is represented by the vector The patient state vector is represented by the vector Represented by, then the multi-objective optimization vector It can be obtained by weighted summation:

[0124] ;

[0125] in, and are the weight parameters of the personalized trial recommendation list and the patient status vector, indicating their relative importance in the optimization vector. Construct an optimization objective function to obtain the initial optimization problem. The optimization objective function aims to maximize the patient matching while minimizing constraints such as the time and cost of the trial. Assume that the optimization objective function is , then the initial optimization problem can be expressed as:

[0126] Maximize subject to , ;

[0127] in, Indicates Constraints, is the number of constraints. In order to solve the initial optimization problem, multi-objective particle swarm optimization is performed to obtain a non-dominated solution set. Particle swarm optimization is an optimization algorithm that simulates the intelligent behavior of a group. It searches for the optimal solution by continuously updating and iterating the speed and position of particles. Each particle represents a potential solution, its position represents the characteristics of the current solution, and its speed represents the search direction. The formula for updating the position and speed of a particle is as follows:

[0128] ;

[0129] ;

[0130] in, and Represents particles In iteration The speed and position at is the inertia weight, and is the learning factor, and is a random number, It is a particle The best location in history, is the global optimal position. Through multiple iterations, the particle swarm optimization algorithm will eventually converge to a non-dominated solution set, which performs well on multiple objectives but has no obvious superiority or inferiority relationship with each other. Pareto optimal analysis is performed on the non-dominated solution set to obtain the initial patient allocation plan. Pareto optimal analysis ensures that the overall benefit of the allocation plan is maximized by selecting solutions that are irreplaceable on all objectives. By sorting and screening the solutions in the non-dominated solution set, the solution on the Pareto frontier is found. Time-sensitive decisions are made on the initial patient allocation plan to obtain a sequential allocation strategy. Time-sensitive decisions take into account the dynamic changes in the patient's status and the time requirements of the experiment, and optimize the execution time sequence of the allocation plan. For example, according to the expected changes in the patient's health status, the priority and time of the experiment are arranged so that the experiment can be carried out at the most appropriate time. According to the output detection information of the sequential allocation strategy and the anomaly detection model, a reinforcement learning environment is constructed to obtain a state-action space. The reinforcement learning environment simulates the actual patient recruitment process and forms a state-action space by defining states (patient's health status and trial requirements) and actions (allocation strategies). Perform Q learning on the state-action space to obtain the optimal action value function Q-learning is a model-free reinforcement learning algorithm that evaluates the value of each state-action pair through continuous iteration and updating, and ultimately finds a strategy that maximizes the total reward. The update formula is:

[0131] ;

[0132] in, is the learning rate, is the discount factor, It’s an instant reward. It's action After reaching the new state, Is in the new state The value of the optimal action is selected. The strategy is updated according to the optimal action value function and real-time feedback data to obtain a dynamic patient recruitment strategy. Real-time feedback data includes information such as patient responses and the progress of the trial. By incorporating these data into the learning process, the recruitment strategy is dynamically adjusted and optimized to ensure its effectiveness and adaptability in practical applications. The dynamic patient recruitment strategy is searched to obtain a strategy execution sequence. The strategy execution sequence is a specific operation step based on the optimal action value function. By evaluating its matching degree with the patient state vector, it ensures that the final recruitment strategy not only meets the needs of the patient but also maximizes the success rate of the clinical trial. The matching degree is evaluated based on the strategy execution sequence and the patient state vector to obtain the target patient recruitment strategy. The matching degree evaluation ensures that the recommended trial is not only optimal in theory but also can be achieved in actual operation by comparing the patient's current health status with the strategy execution sequence.

[0133] The above describes the clinical patient recruitment method based on AI in the embodiment of the present application. The following describes the clinical patient recruitment system based on AI in the embodiment of the present application. Figure 2 In the embodiments of the present application, an embodiment of the clinical patient recruitment system based on AI includes:

[0134] The processing module 201 is used to serialize the patient information to obtain a patient state vector, and store the patient state vector into a multi-dimensional index structure to obtain a state domain table;

[0135] An analysis module 202 is used to perform probability analysis on the state domain table to obtain a state transition diagram, and input the state transition diagram into a decision tree algorithm for anomaly detection training to obtain an anomaly detection model;

[0136] A matrixing module 203 is used to perform matrixing on the patient state vector and the clinical trial information to obtain a dynamic adaptability matrix;

[0137] The calculation module 204 is used to perform similarity calculation on the patient feature matrix and the patient behavior data to obtain a comprehensive similarity index of the patient, and to construct an index for the comprehensive similarity index of the patient to obtain a patient similarity network;

[0138] Extraction module 205, used to retrieve similar patients based on the patient similarity network and extract clinical trial information labels for similar patients to obtain a target label set, and perform weighted calculation on the target label set according to a dynamic adaptability matrix to obtain a personalized trial recommendation list;

[0139] The generation module 206 is used to perform multi-objective optimization on the personalized trial recommendation list and the patient state vector to obtain an initial patient allocation plan, and generate a target patient recruitment strategy based on real-time feedback data through an anomaly detection model.

[0140] Through the cooperation of the above components, the patient state vector is stored in a multidimensional index structure to form a state domain table, which realizes efficient storage and rapid retrieval of patient information. The state transition diagram and anomaly detection model can accurately predict the change of patient status and timely identify abnormal situations. The dynamic adaptability matrix realizes the dynamic matching of patient status and clinical trial requirements. The patient similarity network is constructed by combining patient characteristics and behavioral data. Based on the patient similarity network and dynamic adaptability matrix, personalized recommendations for clinical trials are realized, which improves patients' acceptance of recommended trials. Through the multi-objective optimization algorithm, multiple goals of patient allocation, such as maximizing participation and minimizing recruitment time, are balanced, which improves the overall recruitment efficiency. By using anomaly detection models and real-time feedback data, the recruitment strategy can be adjusted dynamically, and a time-sensitive decision model is introduced to consider the changes of patient status and trial requirements over time. Through the reinforcement learning environment, it can continuously learn and improve from the recruitment process. Through matching evaluation, it ensures that the final generated recruitment strategy is highly consistent with the patient status, which improves the recruitment matching accuracy and patient satisfaction.

[0141] The present application also provides an AI-based clinical patient recruitment device, which includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes the steps of the AI-based clinical patient recruitment method in the above-mentioned embodiments.

[0142] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein instructions are stored in the computer, and when the instructions are executed on a computer, the computer executes the steps of the AI-based clinical patient recruitment method.

[0143] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0144] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.

[0145] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A clinical patient recruitment method based on AI, characterized in that: The method comprises: Serializing the patient information to obtain a patient state vector, and storing the patient state vector into a multidimensional index structure to obtain a state domain table; Performing probability analysis on the state domain table to obtain a state transition diagram, and inputting the state transition diagram into a decision tree algorithm for anomaly detection training to obtain an anomaly detection model; The patient state vector and clinical trial information are matrixed to obtain a dynamic adaptability matrix; specifically comprising: extracting features and standardizing the patient state vector to obtain a patient feature matrix, and encoding and quantifying the clinical trial information according to the inclusion and exclusion criteria to obtain a test requirement matrix; performing matrix multiplication operations on the patient feature matrix and the test requirement matrix to obtain a static adaptability matrix; performing time series prediction on the patient feature matrix according to the state transition diagram to obtain a predicted state matrix; performing matrix multiplication operations on the predicted state matrix and the test requirement matrix to obtain a target adaptability matrix; and A time attenuation function is designed according to the static adaptability matrix and the target adaptability matrix to obtain a time weight vector; the static adaptability matrix and the target adaptability matrix are weightedly fused by the time weight vector to obtain a first adaptability matrix; a test correlation matrix is ​​constructed according to the historical participation data in the patient information, and the first adaptability matrix is ​​optimized by collaborative filtering to obtain a second adaptability matrix; the second adaptability matrix is ​​matrix decomposed to obtain a low-dimensional representation matrix, and the matrix is ​​reconstructed and adjusted according to the low-dimensional representation matrix and the output detection information of the anomaly detection model to obtain a dynamic adaptability matrix; Performing similarity calculation on the patient feature matrix and the patient behavior data to obtain a comprehensive patient similarity index, and constructing an index for the comprehensive patient similarity index to obtain a patient similarity network; Retrieving similar patients based on the patient similarity network and extracting clinical trial information labels for the similar patients to obtain a target label set, and performing weighted calculation on the target label set according to the dynamic adaptability matrix to obtain a personalized trial recommendation list; The personalized trial recommendation list and the patient state vector are subjected to multi-objective optimization to obtain an initial patient allocation plan, and a target patient recruitment strategy is generated according to real-time feedback data through the anomaly detection model.

2. The AI-based clinical patient recruitment method according to claim 1, characterized in that: The patient information is serialized to obtain a patient state vector, and the patient state vector is stored in a multidimensional index structure to obtain a state domain table, including: Acquire the patient's electronic health record from the patient information, perform structured processing on the patient's electronic health record to obtain structured health status data, and perform discretization processing on the patient's health status according to the structured health status data to obtain a discrete health status; Encoding the inclusion and exclusion criteria, trial phase and treatment plan of the clinical trial information to obtain a trial state code, and constructing a state vector template according to the discrete health state and the trial state code to obtain a patient state vector structure; Obtaining the patient's medical examination results from the patient information, and digitizing the patient's medical examination results and treatment plans to obtain a health indicator set; Filling the patient state vector structure according to the health indicator set and the test state code to obtain an initial state vector, and timestamping the initial state vector to obtain a patient state vector; A multidimensional index structure is created, and the patient state vector is stored in the multidimensional index structure to obtain a state domain table.

3. The AI-based clinical patient recruitment method according to claim 2, characterized in that: The probabilistic analysis of the state domain table is performed to obtain a state transition diagram, and the state transition diagram is input into a decision tree algorithm for anomaly detection training to obtain an anomaly detection model, including: Performing time series analysis on the historical data in the state domain table to obtain a state transition frequency matrix, and calculating the transition probability between states according to the state transition frequency matrix to obtain a Markov transition matrix; Graphically processing the Markov transition matrix to obtain an initial transition graph, and extracting node features and edge features according to the initial transition graph to obtain a state transition feature set; Performing cluster analysis on the state transition feature set to obtain a state combination pattern, and marking abnormal state transitions according to the state combination pattern to obtain a state transition diagram; Sampling the state transition diagram to obtain a training data set and a validation data set, and training a decision tree algorithm according to the training data set to obtain an initial anomaly detector; The initial anomaly detector is subjected to ensemble learning to obtain an initial detection model, and the initial detection model is subjected to performance evaluation and threshold optimization according to the validation data set to obtain an anomaly detection model.

4. The AI-based clinical patient recruitment method according to claim 1, characterized in that: The similarity calculation is performed on the patient feature matrix and the patient behavior data to obtain a patient comprehensive similarity index, and an index is constructed for the patient comprehensive similarity index to obtain a patient similarity network, including: Calculating the cosine similarity between patients according to the patient feature matrix to obtain a feature similarity matrix, and serializing the digital health platform browsing history data in the patient information to obtain a behavior sequence vector; Performing interest feature analysis on the behavior sequence vector to obtain a patient interest feature matrix, and performing similarity calculation on the patient interest feature matrix to obtain an interest similarity matrix; Performing weighted fusion on the feature similarity matrix and the interest similarity matrix to obtain a comprehensive similarity matrix; Performing threshold filtering on the comprehensive similarity matrix to obtain a sparse similarity matrix, and constructing an initial similarity network according to the sparse similarity matrix; Performing local sensitive hashing calculation on the initial similarity network to obtain a hash index structure; The initial similarity network is dynamically updated according to the hash index structure and the state transition model to obtain a patient similarity network.

5. The AI-based clinical patient recruitment method according to claim 1, characterized in that: The retrieving similar patients based on the patient similarity network and extracting clinical trial information labels for the similar patients to obtain a target label set, and performing weighted calculation on the target label set according to the dynamic adaptability matrix to obtain a personalized trial recommendation list, includes: Performing graph traversal analysis on the patient similarity network to obtain a neighbor set, and extracting clinical trial information browsed by similar patients based on the neighbor set to obtain a test information set; Processing the test information set to obtain a test description text corpus, and performing word frequency statistics based on the test description text corpus to obtain an initial tag set; Calculating semantic similarity of the initial tag set to obtain a tag similarity matrix, and performing hierarchical clustering based on the tag similarity matrix to obtain a target tag set; Performing collaborative filtering on the target label set and the patient implicit behavior data to obtain a patient-label score matrix, and performing matrix multiplication operation on the dynamic adaptability matrix and the patient-label score matrix to obtain a weighted trial score matrix; The weighted test score matrix is ​​sorted according to multiple objectives to obtain a sorted test list, and personalized filtering is performed according to the sorted test list and the patient state vector to obtain a personalized test recommendation list.

6. The AI-based clinical patient recruitment method according to claim 1, characterized in that: The performing multi-objective optimization on the personalized test recommendation list and the patient state vector to obtain an initial patient allocation plan, and generating a target patient recruitment strategy according to real-time feedback data through the anomaly detection model, comprises: Performing feature fusion on the personalized test recommendation list and the patient state vector to obtain a multi-objective optimization vector; Constructing an optimization objective function according to the multi-objective optimization vector to obtain an initial optimization problem, and performing multi-objective particle swarm optimization on the initial optimization problem to obtain a non-dominated solution set; Performing a Pareto optimal analysis on the non-dominated solution set to obtain an initial patient allocation plan, and performing a time-sensitive decision on the initial patient allocation plan to obtain a time-series allocation strategy; Constructing a reinforcement learning environment according to the timing allocation strategy and the output detection information of the abnormal detection model to obtain a state-action space; Performing Q learning processing on the state-action space to obtain an optimal action value function, and updating the strategy according to the optimal action value function and real-time feedback data to obtain a dynamic patient recruitment strategy; The dynamic patient recruitment strategy is searched to obtain a strategy execution sequence, and a matching degree is evaluated based on the strategy execution sequence and the patient state vector to obtain a target patient recruitment strategy.

7. An AI-based clinical patient recruitment system, characterized in that: For executing the AI-based clinical patient recruitment method according to any one of claims 1 to 6, the system comprises: A processing module is used to serialize the patient information to obtain a patient state vector, and store the patient state vector into a multidimensional index structure to obtain a state domain table; An analysis module is used to perform probability analysis on the state domain table to obtain a state transition diagram, and input the state transition diagram into a decision tree algorithm for anomaly detection training to obtain an anomaly detection model; The matrixing module is used to perform matrixing on the patient state vector and the clinical trial information to obtain a dynamic adaptability matrix; specifically comprising: extracting and standardizing the patient state vector to obtain a patient feature matrix, encoding and quantifying the clinical trial information according to the inclusion and exclusion criteria to obtain a test requirement matrix; performing matrix multiplication operations on the patient feature matrix and the test requirement matrix to obtain a static adaptability matrix; performing time series prediction on the patient feature matrix according to the state transition diagram to obtain a predicted state matrix; performing matrix multiplication operations on the predicted state matrix and the test requirement matrix to obtain a target adaptability matrix. matrix; designing a time decay function according to the static adaptability matrix and the target adaptability matrix to obtain a time weight vector; performing weighted fusion on the static adaptability matrix and the target adaptability matrix through the time weight vector to obtain a first adaptability matrix; constructing a test correlation matrix according to the historical participation data in the patient information, and performing collaborative filtering optimization on the first adaptability matrix to obtain a second adaptability matrix; performing matrix decomposition on the second adaptability matrix to obtain a low-dimensional representation matrix, and performing matrix reconstruction and adjustment according to the low-dimensional representation matrix and the output detection information of the anomaly detection model to obtain a dynamic adaptability matrix; A calculation module, used for performing similarity calculation on the patient feature matrix and the patient behavior data to obtain a comprehensive patient similarity index, and constructing an index for the comprehensive patient similarity index to obtain a patient similarity network; An extraction module, configured to retrieve similar patients based on the patient similarity network and extract clinical trial information labels for the similar patients to obtain a target label set, and to perform weighted calculation on the target label set according to the dynamic adaptability matrix to obtain a personalized trial recommendation list; A generation module is used to perform multi-objective optimization on the personalized test recommendation list and the patient state vector to obtain an initial patient allocation plan, and generate a target patient recruitment strategy according to real-time feedback data through the anomaly detection model.

8. An AI-based clinical patient recruitment device, characterized in that: The AI-based clinical patient recruitment device includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instructions in the memory to enable the AI-based clinical patient recruitment device to perform the AI-based clinical patient recruitment method as described in any one of claims 1-6.

9. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the AI-based clinical patient recruitment method as described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Patient recruitment method and device in clinical research project

    CN113380353A

  • Probabilistic modeling to match patients to clinical trials

    WO2019079490A1