An intelligent clinical trial plan generation method and related device
Through neural networks and deep learning algorithms, clinical trial knowledge data are processed, and combined with natural language recognition technology, a personalized clinical trial plan for new tumor drugs is generated, solving the inefficiency and high cost problems caused by the complex design of the scheme in the existing technology, and achieving efficient and personalized trial plan generation.
Patent Information
- Application Number
- CN202310029360.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-09
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-01-09
AI Technical Summary
The design of existing new oncology drugs is complex, resulting in low output efficiency and high cost, making it difficult to effectively control.
The neural network algorithm is used to label and clean clinical trial knowledge data, combine natural language recognition technology to obtain user needs, train the experimental data model through deep learning algorithms, generate and improve temporary experimental plans, and finally generate a personalized final experimental plan.
It improves the output rate and personalization of the experimental plan, reduces repetitive work, and realizes efficient generation and personalized customization of the experimental plan.
Smart Images

Figure CN115954072B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent clinical trial plan generation method and related devices. Background Art
[0002] Over the past few decades, cancer research has achieved remarkable success in all aspects of cancer prevention, diagnosis, and treatment. As our understanding of the causes and progression of cancer, as well as the prognostic effects of various treatments, has deepened, clinical trials for new cancer drugs have become increasingly complex. Continuous innovations in modern biomedical methods have further elevated the complexity of clinical trials for new cancer drugs to unprecedented levels.
[0003] The cost of cancer research is extremely high, the project background information is complex, and the relationships between the protocol clauses are more complex than those of general clinical trial protocols. Coupled with the ever-changing frontiers of biotechnology, the slightest carelessness in the design of the trial protocol can lead to the failure of the trial, or the cost-effectiveness of the trial becomes worthless, resulting in low output efficiency of the trial protocol and great difficulty in control. Summary of the Invention
[0004] The purpose of the present invention is to provide an intelligent clinical trial plan generation method and related devices, aiming to solve the problems of low output efficiency and great control difficulty of existing trial plans.
[0005] In a first aspect, an embodiment of the present invention provides an intelligent clinical trial plan generation method, comprising:
[0006] The clinical trial knowledge data is annotated using a neural network algorithm to obtain annotated information, and the annotated information is cleaned, the cleaned annotated information is converted into a unified language content, and the converted annotated information is entered into a knowledge base;
[0007] Obtaining the user's test purpose data, and using natural language recognition technology to identify the test purpose data to obtain purpose text information;
[0008] The user's historical test plan data and usage behavior habit data are trained using a deep learning algorithm to obtain a trained test data model;
[0009] Inputting the text information into the test data model to obtain test results;
[0010] Input the test results into the knowledge base, generate corresponding temporary test plans and knowledge systems and send them to users;
[0011] Obtain the user's improved content of the temporary test plan based on the knowledge system and generate a final test plan.
[0012] In a second aspect, an embodiment of the present invention provides an intelligent clinical trial plan generation device, comprising:
[0013] A knowledge base supplement unit is used to annotate clinical trial knowledge data using a neural network algorithm to obtain annotated information, clean the annotated information, convert the cleaned annotated information into a unified language content, and enter the converted annotated information into the knowledge base;
[0014] A text information acquisition unit is used to acquire the user's test purpose data and identify the test purpose data using natural language recognition technology to obtain the purpose text information;
[0015] A test data model training unit is used to train the user's historical test plan data and usage behavior habit data through a deep learning algorithm to obtain a trained test data model;
[0016] A test result acquisition unit, configured to input the text information into the test data model to obtain a test result;
[0017] A test result input unit, used to input the test results into the knowledge base, generate a corresponding temporary test plan and knowledge system, and send them to the user;
[0018] The final test plan generating unit is used to obtain the user's improved content of the temporary test plan based on the knowledge system and generate a final test plan.
[0019] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the intelligent clinical trial plan generation method described in the first aspect above is implemented.
[0020] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the intelligent clinical trial plan generation method described in the first aspect above is implemented.
[0021] The embodiment of the present invention provides an intelligent clinical trial plan generation method and related devices, the method comprising: labeling clinical trial knowledge data using a neural network algorithm to obtain labeling information, performing data cleaning on the labeling information, converting the labeling information after data cleaning into a unified language content and entering the converted labeling information into a knowledge base; obtaining the user's test purpose data and using natural language recognition technology to identify the test purpose data to obtain the purpose text information; training the user's historical test plan data and usage behavior habit data using a deep learning algorithm to obtain a trained test data model; inputting the text information into the test data model to obtain test results; inputting the test results into the knowledge base to generate a corresponding temporary test plan and knowledge system and send them to the user; obtaining the user's improvement content of the temporary test plan based on the knowledge system and generating a final test plan. Based on the created knowledge base, the present invention provides a goal-oriented multidisciplinary knowledge system, adopts AI algorithm internal learning iteration, provides a professional and personalized clinical plan knowledge map, and supplements it with massive content in the industry to produce clinical drug test plans, fully solving the original repetitive work, improving the output rate of the test plan, and customizing the test plan that meets the user's personality through the test data model. The embodiments of the present invention also provide an intelligent clinical trial plan generation device, a computer-readable storage medium, and a computer device, which have the above-mentioned beneficial effects and are not described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 This is a flow chart of a method for generating an intelligent clinical trial plan according to this embodiment;
[0024] Figure 2 This is a sub-flowchart of an intelligent clinical trial plan generation method of this embodiment;
[0025] Figure 3 This is another sub-flowchart of the method for generating an intelligent clinical trial plan according to this embodiment;
[0026] Figure 4 This is a schematic block diagram of an intelligent clinical trial plan generation device according to this embodiment. DETAILED DESCRIPTION
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0028] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0029] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0030] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0031] See also Figure 1 The present invention provides an intelligent clinical trial plan generation method, comprising:
[0032] S101: labeling clinical trial knowledge data using a neural network algorithm to obtain labeling information, performing data cleaning on the labeling information, converting the cleaned labeling information into a unified language content, and entering the converted labeling information into a knowledge base;
[0033] Among them, the annotated information needs to be reviewed before entering it into the knowledge base. The review can be a system review or a manual review to ensure the authenticity and accuracy of clinical trial knowledge.
[0034] See also Figure 2 , step S101 includes:
[0035] S201: Preprocessing the clinical trial knowledge data to obtain preprocessed data;
[0036] In this embodiment, the clinical trial knowledge data may be preprocessed by removing ID attributes, deleting features containing missing values, and so on.
[0037] S202: Pre-labeling the semi-structured and unstructured data in the pre-processed data according to drug, clinical, and statistical categories and performing label processing on the pre-labeling;
[0038] In this step, the semi-structured and unstructured data in the preprocessed data are pre-labeled according to the categories of drugs, clinical, and statistics. Subsequently, the labeled clinical trial knowledge data is classified and stored according to the categories of drugs, clinical, and statistics, which can improve the retrieval efficiency of clinical trial knowledge.
[0039] S203: Extracting entity information from pre-labeled clinical trial knowledge data using an algorithm combining a bidirectional long short-term memory neural network and a conditional random field;
[0040] Specifically, pre-labeled clinical trial knowledge data is obtained; the pre-labeled clinical trial knowledge data is input into a long short-term memory model to obtain character embedding; the pre-labeled clinical trial knowledge data is input into a language representation pre-training model to obtain word embedding; the character embedding and the word embedding are spliced together, and the spliced character embedding and word embedding are input into a bidirectional long short-term memory network to obtain a processing result; the processing result is input into a conditional random field to obtain entity information.
[0041] S204: Using named entity recognition technology to perform text triple annotation on the entity information to obtain text triples;
[0042] S205: Construct a knowledge graph based on the text triples.
[0043] Furthermore, after constructing the knowledge graph, other knowledge graph data of the same type as the knowledge graph is obtained in the knowledge base, and the knowledge graph with the highest similarity to the knowledge graph in the knowledge base is calculated through the cosine similarity algorithm, and the knowledge graph is associated with the knowledge graph with the highest similarity in the knowledge base and stored in the knowledge base.
[0044] In one embodiment, the method of step S101 can also be used to create a specialized knowledge base, such as a knowledge base for various cancer types such as breast cancer and pancreatic cancer. The creation of a specialized knowledge base can provide users with a more professional and detailed knowledge system.
[0045] In a specific application scenario, a large amount of data was collected and analyzed for non-small cell lung cancer, a currently studied disease. Based on the existing knowledge team, a corresponding knowledge system map was constructed that matched the research and pharmaceutical company proposals. The disassembled paragraphs were annotated accordingly. Common annotations could be broken down into intervention targets, intervention methods, and other content based on the document structure. Clinical annotations could be used to define disease, test indicator, medication, treatment, and other content. Statistical annotations could be used to define error reduction, sample size estimation, grouping recommendations, and other content. By extracting existing disassembly logic and combining a bidirectional long-short-term memory neural network with a conditional random field, a corresponding disassembly model algorithm was generated. Natural language recognition (NLP) technology and the disassembly model algorithm were used to collect, disassemble, and annotate large amounts of data. Finally, after manual review and confirmation, the non-small cell lung cancer knowledge base was generated. An example of an annotation is shown in Table 1.
[0046] Table 1
[0047]
[0048]
[0049] This embodiment realizes automatic supplementation and improvement of the knowledge base through step S101, and constructs a corresponding knowledge graph. When the user needs to query relevant knowledge, the system can provide the user with the knowledge context and knowledge graph with the best similarity, providing professional knowledge support for the user to write the experimental plan.
[0050] S102: Acquire the user's test purpose data, and use natural language recognition technology to identify the test purpose data to obtain purpose text information;
[0051] This step obtains the user's test purpose data through natural language recognition technology, obtains language feature vectors by performing endpoint detection, noise reduction, and feature extraction on the user's test purpose data, and then performs statistical pattern recognition on the language feature vectors to obtain the target text information.
[0052] S103: Training the user's historical test plan data and usage behavior habit data using a deep learning algorithm to obtain a trained test data model;
[0053] See also Figure 3 The specific contents of this step are as follows:
[0054] S301: extracting entities from the user's historical test plan data and usage behavior habit data to obtain entity recognition data;
[0055] In this embodiment, the main information elements extracted are atomic information elements in the text, such as intervention means, intervention objects, prognostic effects, ethical content, and trial plan template formats. Among them, the intervention object is: a certain disease that the researcher wants to solve, or a specific manifestation of the disease, etc.; the intervention means is: the specific method of treating the disease, that is, the drugs, medical devices, and treatment plans being tested in the trial; the prognostic effects are: the test goals achieved by the trial plan, such as effectiveness and safety, etc.; the ethical content is: the theoretical policy points that need to be paid attention to in the necessary review, etc.
[0056] S302: Perform feature selection and feature dimensionality reduction on the entity recognition data to obtain feature data;
[0057] Specifically, a feature subset of the entity recognition data is obtained; the variance of the feature subset is calculated, and 10 features with the smallest variance are selected; a feature compression method is used to reduce the dimension of the 10 features to obtain 5 descriptive features and the descriptive features are used as the feature data of the entity recognition data.
[0058] S303: performing cluster analysis on the feature data to obtain classified data;
[0059] Specifically, the set of feature data is divided into different cluster objects according to feature similarity through the K-means density clustering algorithm, data with similar features are distributed in the same cluster, and data with dissimilar features are distributed outside the cluster; the distribution density of the feature data is calculated, and data analysis is performed on the distribution density of the feature data to obtain classified data.
[0060] S304: Constructing a test data model using the deep learning algorithm;
[0061] S305: Input the classified data into the test data model for training to obtain a trained test data model.
[0062] This embodiment pre-processes the user's historical test plan data and usage behavior habit data through methods such as entity extraction, feature selection, feature dimensionality reduction, and cluster analysis, and then inputs the pre-processed data into the experimental data model to implement training of the experimental data model. The experimental data model can learn the user's behavioral habits and provide users with test plans that suit their habits and personality by combining with the knowledge base, thereby improving the output rate of test plans.
[0063] S104: Inputting the text information into the test data model to obtain test results;
[0064] S105: Input the test results into the knowledge base, generate a corresponding temporary test plan and knowledge system, and send them to the user;
[0065] S106: Obtaining the user's improved content of the temporary test plan based on the knowledge system and generating a final test plan.
[0066] Among them, after generating the final test plan, the system will input the final test plan into the test data model for training and updating, and at the same time convert the final test plan into a knowledge graph through the neural network algorithm and store it in the knowledge base.
[0067] Furthermore, when the final experimental plan is stored in the knowledge base, other users can make relevant modifications to the experimental plan, or make corresponding supplements to the plan based on actual conditions. This embodiment realizes the full cycle control of the experimental plan from the draft of the experimental plan, the version revision process, submission for ethical review, experimental plan execution, and completion of the experimental plan through online collaboration and corresponding version management.
[0068] Furthermore, before generating a temporary experimental plan, an overall constant framework library can be constructed. The specific process is as follows: extract specific features such as diseases, chemical compounds, and gene targets from the existing plan content, merge and compare similar plans, eliminate and blur the parts with variables (such as decreased platelets), retain the shared constant framework part, form the final constant template, annotate the extracted specific features, and construct an overall constant framework library.
[0069] In one embodiment, a deep learning algorithm is used to comprehensively extract and train the patterns of historical search popularity and the templates and corresponding disease and target characteristics that users tend to use in their historical plan development process to form a test data model. When the main researcher or pharmaceutical company-related organization initiates a clinical trial plan based on their own needs, the user inputs their own needs into the system. The system relies on the test data model to capture the best recommended content and multiple related constant templates and knowledge bases from the knowledge base and constant framework library. Finally, the user freely matches and assembles the best recommended content and multiple related constant templates and knowledge bases according to their own needs to obtain a "rough" underlying template of the clinical trial plan. The system also provides users with operations to improve and beautify the underlying template, including but not limited to the use of multiple statistical reference tools, reminders of common historical precautions, reminders of knowledge rules that are strongly correlated or unrelated between conventional knowledge content, references to excellent writing plan templates of the same type, and team online collaboration functions. Finally, users produce a clinical trial plan that is ultimately suitable for the current drug and current disease based on their actual needs.
[0070] In a specific application scenario, content is selected according to the system scope boundary: the final selection is intervention trial + phase III and PD-1 drug application in non-small cell lung cancer. The system searches the knowledge base based on the content finally selected by the user, and recommends the best (highest degree of annotation matching) template and the most commonly used (based on the user's historical writing habits) template. After the user views the corresponding knowledge base and template and makes a selection, the skeleton and constant part of the solution are obtained, that is, the underlying template at the "rough" level is obtained; the user can make necessary key node reminders and fills based on non-small cell lung cancer, PD-1 targets, matters needing attention in phase III or clinical concerns on the underlying template; the user can also Based on the obtained underlying template, the system provides a variety of tools to make relevant modifications to each chapter. For example, according to the sample size estimation tool, the allocation ratio of the corresponding experimental drug group and the control group, the subjects in each group, the dropout rate and the indication drug-related rate are input to calculate the total sample capacity and reversely import the plan; users can annotate and communicate with corresponding annotations through multi-person online collaboration to jointly complete the writing of a chapter or module with a high degree of aggregation; at the same time, multiple excellent industry case templates are provided for reference in the writing process to improve the sense of interaction in the writing process; and finally a complete set of clinical trial plans for PD-1 gene targets for non-small cell lung cancer is generated and put into practice and implementation. The system will dynamically disassemble the content of the trial plan generated this time, and dynamically update the knowledge base after confirmation by the knowledge base review team, and update the trial data model at the same time.
[0071] This embodiment ensures the coordination of work through some online network technologies, and essentially realizes the electronicization of the original offline work content; in addition, starting from the establishment of the experimental project, it realizes the full-process management of the experimental cycle; and constructs an authoritative and professional multidisciplinary knowledge base system, which enables ordinary users to achieve one-stop guidance, creation and generation of experimental plans.
[0072] See also Figure 4 This embodiment provides an intelligent clinical trial plan generation device 400, including:
[0073] The knowledge base supplement unit 401 is used to label the clinical trial knowledge data using a neural network algorithm to obtain labeling information, clean the labeling information, convert the cleaned labeling information into a unified language content, and enter the converted labeling information into the knowledge base;
[0074] The text information acquisition unit 402 is used to acquire the user's test purpose data and identify the test purpose data using natural language recognition technology to obtain purpose text information;
[0075] The test data model training unit 403 is used to train the user's historical test plan data and usage behavior habit data through a deep learning algorithm to obtain a trained test data model;
[0076] A test result acquisition unit 404 is used to input the text information into the test data model to obtain a test result;
[0077] The test result input unit 405 is used to input the test result into the knowledge base, generate a corresponding temporary test plan and knowledge system, and send them to the user;
[0078] The final test plan generating unit 406 is configured to obtain the user's improved content of the temporary test plan based on the knowledge system and generate a final test plan.
[0079] Furthermore, the knowledge base supplement unit 401 includes:
[0080] A preprocessing subunit, configured to preprocess the clinical trial knowledge data to obtain preprocessed data;
[0081] a labeling subunit, configured to pre-label the semi-structured and unstructured data in the pre-processed data according to the categories of medicine, clinical, and statistics, and to perform label processing on the pre-labeled data;
[0082] An entity information extraction subunit, which is used to extract entity information from pre-labeled clinical trial knowledge data using an algorithm that combines a bidirectional long short-term memory neural network and a conditional random field;
[0083] A text triple acquisition subunit is used to use named entity recognition technology to perform text triple annotation on the entity information to obtain text triples;
[0084] The knowledge graph construction subunit is used to construct a knowledge graph based on the text triples.
[0085] Furthermore, the entity information extraction subunit includes:
[0086] A data acquisition subunit, used to acquire pre-labeled clinical trial knowledge data;
[0087] The character embedding acquisition subunit is used to input pre-labeled clinical trial knowledge data into the long short-term memory model to obtain character embeddings;
[0088] The word embedding acquisition subunit is used to input pre-labeled clinical trial knowledge data into the language representation pre-training model to obtain word embeddings;
[0089] a concatenation processing subunit, configured to concatenate the character embedding and the word embedding, and input the concatenated character embedding and word embedding into a bidirectional long short-term memory network to obtain a processing result;
[0090] The input subunit is used to input the processing result into the conditional random field to obtain entity information.
[0091] Furthermore, the test data model training unit 403 includes:
[0092] An entity recognition data acquisition subunit is used to extract entities from the user's historical test plan data and usage behavior habit data to obtain entity recognition data;
[0093] A feature data acquisition subunit, configured to perform feature selection and feature dimension reduction on the entity recognition data to obtain feature data;
[0094] A classification data acquisition subunit, configured to perform cluster analysis on the feature data to obtain classification data;
[0095] A test data model construction subunit, configured to construct a test data model using the deep learning algorithm;
[0096] The classification data input subunit is used to input the classification data into the test data model for training to obtain a trained test data model.
[0097] Furthermore, the feature data acquisition subunit includes:
[0098] A feature subset acquisition subunit, configured to acquire a feature subset of the entity recognition data;
[0099] A feature selection subunit, used to calculate the variance of the feature subset and select the 10 features with the smallest variance;
[0100] The feature dimension reduction subunit is used to reduce the dimension of the 10 features by adopting a feature compression method to obtain 5 descriptive features and use the descriptive features as feature data of the entity recognition data.
[0101] Furthermore, the classification data acquisition subunit includes:
[0102] A feature clustering subunit is used to divide the feature data set into different cluster objects according to feature similarity by using a K-means density clustering algorithm, distributing data with similar features in the same cluster and distributing data with dissimilar features outside the cluster;
[0103] The distribution density calculation subunit is used to calculate the distribution density of the feature data and perform data analysis on the distribution density of the feature data to obtain classification data.
[0104] Furthermore, the final test plan generating unit 406 includes:
[0105] The storage subunit is used to input the final test plan into the test data model for training and updating, and at the same time convert the final test plan into a knowledge graph through a neural network algorithm and store it in a knowledge base.
[0106] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-mentioned devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0107] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When executed, the computer program can implement the methods provided in the above embodiments. The storage medium can include any medium capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0108] The present invention further provides a computer device that may include a memory and a processor. The memory stores a computer program, and the processor, when invoking the computer program in the memory, can implement the method provided in the above embodiment. Of course, the computer device may also include various network interfaces, a power supply, and other components.
[0109] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
[0110] It should also be noted that, in this specification, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprising" or any other variations thereof are intended to cover non-exclusive.
[0111] Inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. A method for generating an intelligent clinical trial plan, characterized in that: include: The clinical trial knowledge data is annotated using a neural network algorithm to obtain annotated information, and the annotated information is cleaned, the cleaned annotated information is converted into a unified language content, and the converted annotated information is entered into a knowledge base; Obtaining the user's test purpose data, and using natural language recognition technology to identify the test purpose data to obtain purpose text information; The user's historical test plan data and usage behavior habit data are trained using a deep learning algorithm to obtain a trained test data model; Inputting the text information into the test data model to obtain test results; Input the test results into the knowledge base, generate corresponding temporary test plans and knowledge systems and send them to users; Obtaining the user's improved content of the temporary test plan based on the knowledge system and generating a final test plan; The labeling of clinical trial knowledge data by a neural network algorithm to obtain labeling information includes: Preprocessing the clinical trial knowledge data to obtain preprocessed data; Pre-labeling the semi-structured and unstructured data in the pre-processed data according to drug, clinical, and statistical categories and performing label processing on the pre-labeling; An algorithm combining bidirectional long short-term memory neural network and conditional random field is used to extract entity information from pre-labeled clinical trial knowledge data; Using named entity recognition technology to perform text triple annotation on the entity information to obtain text triples; Constructing a knowledge graph based on the text triples; The training of the user's historical test plan data and usage behavior habit data by a deep learning algorithm to obtain a trained test data model includes: Perform entity extraction on the user's historical test plan data and usage behavior habit data to obtain entity recognition data; Performing feature selection and feature dimensionality reduction on the entity recognition data to obtain feature data; Performing cluster analysis on the characteristic data to obtain classified data; Constructing a test data model through the deep learning algorithm; The classified data is input into the test data model for training to obtain a trained test data model.
2. The intelligent clinical trial plan generation method according to claim 1, characterized in that: The algorithm combining a bidirectional long short-term memory neural network and a conditional random field to extract entity information from pre-labeled clinical trial knowledge data includes: Access to pre-annotated clinical trial knowledge data; Input the pre-labeled clinical trial knowledge data into the long short-term memory model to obtain character embeddings; Input the pre-labeled clinical trial knowledge data into the language representation pre-training model to obtain word embeddings; Concatenating the character embedding and the word embedding, and inputting the concatenated character embedding and word embedding into a bidirectional long short-term memory network to obtain a processing result; The processing result is input into the conditional random field to obtain entity information.
3. The intelligent clinical trial plan generation method according to claim 1, characterized in that: The feature data obtained by performing feature selection and feature dimensionality reduction on the entity recognition data includes: Obtaining a feature subset of the entity recognition data; Calculate the variance of the feature subset and select the 10 features with the smallest variance; A feature compression method is used to reduce the dimension of the 10 features to obtain 5 descriptive features, and the descriptive features are used as feature data of the entity recognition data.
4. The intelligent clinical trial plan generation method according to claim 1, characterized in that: The cluster analysis of the feature data to obtain the classified data includes: The set of feature data is divided into different cluster objects according to feature similarity by using the K-means density clustering algorithm, data with similar features are distributed in the same cluster, and data with dissimilar features are distributed outside the cluster; The distribution density of the characteristic data is calculated, and data analysis is performed on the distribution density of the characteristic data to obtain classification data.
5. The intelligent clinical trial plan generation method according to claim 1, characterized in that: After obtaining the user's improved content of the temporary test plan and generating the final test plan, the method includes: The final test plan is input into the test data model for training and updating, and at the same time, the final test plan is converted into a knowledge graph through a neural network algorithm and stored in a knowledge base.
6. An intelligent clinical trial plan generation device, characterized in that: include: A knowledge base supplement unit is used to annotate clinical trial knowledge data using a neural network algorithm to obtain annotated information, clean the annotated information, convert the cleaned annotated information into a unified language content, and enter the converted annotated information into the knowledge base; A text information acquisition unit is used to acquire the user's test purpose data and identify the test purpose data using natural language recognition technology to obtain the purpose text information; A test data model training unit is used to train the user's historical test plan data and usage behavior habit data through a deep learning algorithm to obtain a trained test data model; A test result acquisition unit, configured to input the text information into the test data model to obtain a test result; A test result input unit, used to input the test results into the knowledge base, generate a corresponding temporary test plan and knowledge system, and send them to the user; A final test plan generating unit, configured to obtain the user's improved content of the temporary test plan based on the knowledge system and generate a final test plan; The knowledge base supplement unit includes: A preprocessing subunit, configured to preprocess the clinical trial knowledge data to obtain preprocessed data; a labeling subunit, configured to pre-label the semi-structured and unstructured data in the pre-processed data according to the categories of medicine, clinical, and statistics, and to perform label processing on the pre-labeled data; An entity information extraction subunit, which is used to extract entity information from pre-labeled clinical trial knowledge data using an algorithm that combines a bidirectional long short-term memory neural network and a conditional random field; A text triple acquisition subunit is used to use named entity recognition technology to perform text triple annotation on the entity information to obtain text triples; A knowledge graph construction subunit, configured to construct a knowledge graph based on the text triples; The test data model training unit includes: An entity recognition data acquisition subunit is used to extract entities from the user's historical test plan data and usage behavior habit data to obtain entity recognition data; A feature data acquisition subunit, configured to perform feature selection and feature dimension reduction on the entity recognition data to obtain feature data; A classification data acquisition subunit, configured to perform cluster analysis on the feature data to obtain classification data; A test data model construction subunit, configured to construct a test data model using the deep learning algorithm; The classified data input subunit is used to input the classified data into the test data model for training to obtain a trained test data model.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the intelligent clinical trial plan generation method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the intelligent clinical trial plan generation method according to any one of claims 1 to 5.
Citation Information
Patent Citations
A method and system for constructing a health knowledge graph
CN109669994A
Well site test data processing method and device based on knowledge graph
CN113377963A