Drug classification method, device and equipment based on clinical test forbidden drugs
By constructing a Gaussian Bayes algorithm model and using Gaussian kernel function and core principal component analysis, the problem of lack of targetedness and accuracy of existing drug classification methods is solved, and more targeted drug classification is achieved, which improves the accuracy and safety of clinical trials.
Patent Information
- Application Number
- CN202510078047.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The existing drug classification methods are not targeted and cannot guarantee the accuracy of classification. Especially in clinical trials, it is difficult for research doctors to check banned drugs one by one, which affects the accuracy of clinical trial results.
The drug classification method based on banned drugs based on clinical trials is adopted. By receiving the drug information of the drug to be queried, the corresponding disease type and mechanism of action are determined, historical drug information is processed, and the data is mapped to the high-dimensional feature space using Gaussian nuclear function and nuclear principal component analysis, Gaussian Bayesian algorithm model is constructed, drug classification is performed, and alarm prompts are output when a banned drug is detected.
It improves the accuracy of drug classification, makes it more targeted, can be more in line with clinical trial research plans, reduces the work burden of research doctors, reduces risks during clinical trials, and improves the quality of clinical research.
Smart Images

Figure CN120015363A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug classification, and in particular to a drug classification method, device and equipment based on drugs prohibited in clinical trials. Background Art
[0002] During clinical trials, most clinical research plans will restrict the use of combined drugs, and researchers will list these drugs that will interfere with the results of clinical trials as prohibited drugs. Prohibited drugs usually include drugs with similar pharmacological effects to research drugs, drugs that affect the pharmacokinetics of research drugs, and drugs that will enhance or mask adverse reactions. However, there are usually multiple categories of prohibited drugs in each clinical trial, covering a large number of drugs. It is difficult for research doctors to check them one by one when using drugs, which poses a certain risk to the accuracy of clinical trial results.
[0003] Among the current existing drug classification methods, most are classified according to pharmacological effects, structural characteristics, and medical insurance types. This classification method lacks specificity and cannot guarantee the accuracy of the classification.
[0004] The above contents are only used to assist in understanding the technical solution of the present invention and do not constitute an admission that the above contents are prior art. Summary of the invention
[0005] The main purpose of the present invention is to provide a drug classification method, device and equipment based on drugs banned in clinical trials, aiming to solve the technical problem that the current existing drug classification methods are mostly classified according to pharmacological effects, structural characteristics and medical insurance types, etc. Such classification methods lack specificity and cannot guarantee the accuracy of classification.
[0006] To achieve the above object, the present invention provides a drug classification method based on drugs prohibited in clinical trials, and the drug classification method based on drugs prohibited in clinical trials comprises the following steps:
[0007] Receiving input drug information of a drug to be queried;
[0008] Determine the disease type and action mechanism corresponding to the drug to be queried according to the drug information;
[0009] Acquiring historical drug information, and processing the historical drug information to obtain processed drug information;
[0010] Mapping the processed drug information to a high-dimensional feature space through a Gaussian kernel function, and extracting feature information through kernel principal component analysis to obtain a drug information data set after spatial transformation;
[0011] Dividing the spatially transformed drug information dataset into a training dataset and a test dataset;
[0012] Using the Gaussian Bayesian algorithm to perform model training based on the training data set, and using the test data set to evaluate and verify the trained model to complete the construction of the drug classification model;
[0013] Inputting the disease type and the mechanism of action into the drug classification model to determine the drug classification corresponding to the drug to be queried;
[0014] If the drug classification of the drug to be queried is a prohibited drug, a corresponding alarm prompt is output.
[0015] In some embodiments, the processing of the historical drug information to obtain processed drug information includes:
[0016] dividing the historical drug information into a number of initial data sets based on disease type and mechanism of action;
[0017] An initial data set is randomly selected as the current data set, and several samples are selected from the remaining initial data sets to be added to the current data set, until several reference data sets are obtained, and the number of samples between the reference data sets is balanced;
[0018] Standardized automatic scaling is used to scale the features of samples in any reference data so that the features of samples in each reference data are at the same level, and processed drug information is obtained based on several reference data sets that have completed feature scaling.
[0019] In some embodiments, the expression corresponding to scaling the features of samples in any reference data using standardized automatic scaling is: j '=(S j -avg) / d, where S j ' is the eigenvalue after feature j is scaled, S j is the original eigenvalue of feature j before scaling, avg is the average eigenvalue of any reference data set, and d is the variance corresponding to the eigenvalue of any reference data set.
[0020] In some embodiments, the method of using the Gaussian Bayesian algorithm to perform model training based on the training data set, and using the test data set to evaluate and verify the trained model to complete the construction of the drug classification model, includes:
[0021] Determining a priori probabilities for each class in the training set;
[0022] Determine the conditional probability of each attribute under each category;
[0023] Determine the distance between all samples of one category and all samples of another category based on the prior probability of each category and the conditional probability of each attribute under each category;
[0024] Determine the weighted posterior probability using the class-specific attribute weight matrix and class conditional probability, and calculate the corresponding classification accuracy;
[0025] Iterative optimization is performed using the NSGA-II algorithm to determine the spacing and classification accuracy between different categories, and a set of Pareto solutions are recorded;
[0026] When the termination condition of model optimization is reached, the Pareto solution set is output to obtain multiple sets of class-specific attribute weight matrices;
[0027] Using cross-validation evaluation to select a number of optimal class-specific attribute weight matrices from the multiple groups of class-specific attribute weight matrices;
[0028] Using the posterior probability of the test data set based on the several optimal class-specific attribute weight matrices under the test data set;
[0029] Based on the posterior probability, select the category with the highest posterior probability as the predicted category of the test sample, and return the predicted category label;
[0030] The classification accuracy is determined based on the category label, and when the classification accuracy reaches a preset accuracy, the training of the drug classification model is determined to be completed.
[0031] In some embodiments, the expression corresponding to determining the distance between all samples of one category and all samples of another category is: Among them, p(x j |Y0) and p(x j |Y1) are x in the Y0 category j Attributes and Y1 categories under x j The conditional probability of the attribute, P(Y0) and P(Y1) are the prior probabilities of category Y0 and category Y1 respectively, m is the number of samples, and n is the number of features.
[0032] In some embodiments, the classification accuracy is expressed as: Among them, δ(·) is a binary function, which is used to determine the number of samples that the model classifies correctly, y i represents the predicted category label, and y represents the true label.
[0033] In addition, to achieve the above-mentioned purpose, the present invention also proposes a drug classification device based on drugs prohibited for clinical trials, the drug classification device based on drugs prohibited for clinical trials comprising:
[0034] A receiving module, used for receiving input drug information of a drug to be queried;
[0035] An acquisition module, used to determine the disease type and action mechanism corresponding to the drug to be queried according to the drug information;
[0036] A construction module is used to obtain historical drug information, process the historical drug information, and obtain processed drug information;
[0037] The construction module is used to map the processed drug information to a high-dimensional feature space through a Gaussian kernel function, and extract feature information through kernel principal component analysis to obtain a drug information data set after spatial transformation;
[0038] The construction module is used to divide the spatially transformed drug information dataset into a training dataset and a test dataset;
[0039] The construction module is used to perform model training based on the training data set using the Gaussian Bayes algorithm, and to evaluate and verify the trained model using the test data set to complete the construction of the drug classification model;
[0040] A classification module, used for inputting the disease type and the mechanism of action into the drug classification model to determine the drug classification corresponding to the drug to be queried;
[0041] The prompt module is used to output a corresponding alarm prompt if the drug classification of the drug to be queried is a banned drug.
[0042] In some embodiments, the building module is used to divide the historical drug information into several initial data sets based on disease type and mechanism of action;
[0043] An initial data set is randomly selected as the current data set, and several samples are selected from the remaining initial data sets to be added to the current data set, until several reference data sets are obtained, and the number of samples between the reference data sets is balanced;
[0044] Standardized automatic scaling is used to scale the features of samples in any reference data so that the features of samples in each reference data are at the same level, and processed drug information is obtained based on several reference data sets that have completed feature scaling.
[0045] In some embodiments, the expression corresponding to scaling the features of samples in any reference data using standardized automatic scaling is: j '=(S j -avg) / d, where S j ' is the eigenvalue after feature j is scaled, S j is the original eigenvalue of feature j before scaling, avg is the average eigenvalue of any reference data set, and d is the variance corresponding to the eigenvalue of any reference data set.
[0046] In addition, to achieve the above-mentioned purpose, the present invention also proposes a drug classification device based on drugs banned in clinical trials, the drug classification device based on drugs banned in clinical trials comprising: a memory, a processor, and a drug classification program based on drugs banned in clinical trials stored on the memory and executable on the processor, the drug classification program based on drugs banned in clinical trials being configured to implement the steps of the drug classification method based on drugs banned in clinical trials as described above.
[0047] In the present invention, the corresponding disease type and mechanism of action are determined according to the input drug information of the drug to be queried; the drug information after processing the historical drug information is mapped to the high-dimensional feature space through the Gaussian kernel function, and the drug information data set after spatial transformation is obtained through the kernel principal component analysis. The drug classification model is constructed based on the drug information data set and the Gaussian Bayes algorithm is used. Finally, the drug classification corresponding to the drug to be queried is determined through the drug classification model; and the corresponding alarm prompt is output when the banned drug is detected. Through the above method, the drug classification is made more targeted, and the accuracy of the drug classification is improved, so that it is more in line with the clinical trial research plan, which can effectively reduce the workload of research doctors, reduce the risks in the clinical trial process, and improve the quality of clinical research. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 This is a flow chart of a first embodiment of a drug classification method based on banned drugs for clinical trials of the present invention;
[0049] Figure 2 A schematic diagram of a class-specific attribute weight matrix in a drug classification method based on drugs prohibited in clinical trials of the present invention;
[0050] Figure 3 This is a structural block diagram of the first embodiment of the drug classification device based on drugs banned in clinical trials of the present invention.
[0051] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0052] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0053] The embodiment of the present invention provides a drug classification method based on drugs prohibited by clinical trials, referring to Figure 1 , Figure 1 The figure is a flow chart of a first embodiment of a drug classification method based on drugs prohibited in clinical trials according to the present invention.
[0054] In this embodiment, the drug classification method based on the banned drugs in clinical trials includes the following steps:
[0055] Step S10: receiving input drug information of the drug to be queried.
[0056] In this embodiment, the executor of this embodiment is a drug classification device based on drugs banned in clinical trials, wherein the drug classification device based on drugs banned in clinical trials has functions such as data processing, data communication and program running. The drug classification device based on drugs banned in clinical trials can be a computer terminal device or other network device, and of course can also be other devices with similar functions, and this embodiment does not impose any restrictions on this.
[0057] It should be noted that the current existing drug classification methods are mostly based on pharmacological effects, structural characteristics and medical insurance types. This classification method lacks specificity and cannot guarantee the accuracy of the classification.
[0058] In order to solve the above technical problems, in this embodiment, the drug information of the drug to be queried is received as input; the disease type and mechanism of action corresponding to the drug to be queried are determined based on the drug information; a drug classification model is constructed based on historical drug information; the disease type and the mechanism of action are input into the drug classification model to determine the drug classification corresponding to the drug to be queried; if the drug classification of the drug to be queried is a prohibited drug, a corresponding alarm prompt is output. Through the above method, the drug classification is made more targeted, and the accuracy of the drug classification is improved, so that it is more in line with the clinical trial research plan, which can effectively reduce the workload of research doctors, reduce the risks in the clinical trial process, and improve the quality of clinical research. Specifically, it can be achieved in the following way.
[0059] In the specific implementation, this embodiment needs to first receive the input drug information of the drug to be queried. When the doctor performs a drug query, he will enter the relevant drug information. The input drug information includes but is not limited to the name of the drug and its use.
[0060] Step S20: Determine the disease type and action mechanism corresponding to the drug to be queried based on the drug information.
[0061] It should be noted that, unlike the current drug classification, the drug classification in this embodiment is based on the disease type and mechanism of action corresponding to the drug. Based on the drug information obtained above, the disease type and mechanism of action corresponding to the drug to be queried can be determined.
[0062] Step S30: Acquire historical drug information, process the historical drug information, and obtain processed drug information.
[0063] Step S40: Mapping the processed drug information to a high-dimensional feature space through a Gaussian kernel function, and extracting feature information through kernel principal component analysis to obtain a drug information data set after spatial transformation.
[0064] Step S50: Divide the spatially transformed drug information dataset into a training dataset and a test dataset.
[0065] Step S60: Using the Gaussian Bayesian algorithm to perform model training based on the training data set, and using the test data set to evaluate and verify the trained model to complete the construction of the drug classification model.
[0066] It should be noted that the Gaussian kernel function has the following form: K(x,y)=exp(-γ||xy|| 2 ), where x and y are two points in the original space, γ is a parameter, and ||xy|| 2 is the Euclidean distance between x and y. x and y can be obtained from the processed drug information, and the mapping of the high-dimensional feature space is achieved through the above-mentioned functional relationship. The kernel principal component analysis used in this embodiment is an extension of principal component analysis (PCA) and is used to deal with nonlinear data dimensionality reduction problems. Unlike standard PCA, kernel principal component analysis allows the processing of nonlinear data, maps the data to a high-dimensional space in order to better perform dimensionality reduction operations, and can mine nonlinear information in the data.
[0067] Among them, the process of processing the historical drug information to obtain the processed drug information specifically includes: dividing the historical drug information into several initial data sets based on disease type and mechanism of action; arbitrarily selecting an initial data set as the current data set, and selecting several samples from the remaining initial data sets to add to the current data set until several reference data sets are obtained, and the sample quantity balance is achieved between the reference data sets; using standardized automatic scaling to scale the features of samples in any reference data so that the features of samples in each reference data are at the same order of magnitude, and obtaining the processed drug information based on the several reference data sets that have completed feature scaling.
[0068] It should be noted that in order to ensure the balance of the number of samples, in this embodiment, for any data set, several samples are obtained from other data sets and added to the data set. For example, the number of samples in the four data sets 1 to 4 is 100, 66, 110, and 90 respectively. The adjusted data set 1 is all the samples in data set 1 plus 34 samples randomly selected from the remaining data sets 2 to 4. The same method is adopted for each data set to achieve a balance in the number of samples. Among them, the expression corresponding to scaling the features of samples in any reference data using standardized automatic scaling is: S j '=(S j -avg) / d, where S j ' is the eigenvalue after feature j is scaled, S j is the original eigenvalue of feature j before scaling, avg is the average eigenvalue of any reference data set, and d is the variance corresponding to the eigenvalue of any reference data set.
[0069] Furthermore, in this embodiment, the Gaussian Bayesian algorithm is used to train the model based on the training data set, and the trained model is evaluated and verified using the test data set to complete the process of building the drug classification model, specifically determining the prior probability of each category in the training set; determining the conditional probability of each attribute under each category; determining the spacing between all samples of one category and all samples of another category based on the prior probability of each category and the conditional probability of each attribute under each category; determining the weighted posterior probability using the class-specific attribute weight matrix and the class conditional probability, and calculating the corresponding classification accuracy; using the NSGA-II algorithm for iterative optimization to determine different The method comprises the following steps: calculating the spacing between categories and the classification accuracy, and recording a set of Pareto solutions; when the termination condition of the model optimization is reached, outputting the Pareto solution set to obtain multiple groups of class-specific attribute weight matrices; using cross-validation evaluation to screen out several optimal class-specific attribute weight matrices from the multiple groups of class-specific attribute weight matrices; using the posterior probability of the test data set based on the several optimal class-specific attribute weight matrices under the test data set; based on the posterior probability, selecting the category with the highest posterior probability as the predicted category of the test sample, and returning the predicted category label; determining the classification accuracy based on the category label, and determining that the training of the drug classification model is completed when the classification accuracy reaches the preset accuracy.
[0070] It should be noted that the expression corresponding to determining the distance between all samples of one category and all samples of another category is: Among them, p(x j |Y0) and p(x j |Y1) are x in the Y0 category j Attributes and Y1 categories under x jThe conditional probability of the attribute, P(Y0) and P(Y1) are the prior probabilities of the Y0 category and the Y1 category respectively, m is the number of samples, and n is the number of features. The expression of classification accuracy is: Among them, δ(·) is a binary function, which is used to determine the number of samples that the model classifies correctly, y i represents the predicted class label, and y represents the true label. The resulting class-specific attribute weight matrix is as follows: Figure 2 As shown, finally, the class-specific attribute weight matrix is used according to the formula: To calculate the posterior probability of each instance, y represents the category to which the instance belongs, Y represents the set of all categories, m represents the number of all attributes in the data set, and W y,i Represents the weight given to the i-th attribute under category y.
[0071] Step S70: input the disease type and the mechanism of action into the drug classification model to determine the drug classification corresponding to the drug to be queried.
[0072] In a specific implementation, after the construction of the drug classification model is completed, in this embodiment, the acquired disease type and mechanism of action are input into the drug classification model to obtain the specific classification of the drug.
[0073] Step S80: If the drug classification of the drug to be queried is a prohibited drug, a corresponding alarm prompt is output.
[0074] It should be understood that prohibited drugs may affect the results of clinical trials. When the drug classification of the drug to be queried is identified as a prohibited drug, in order to avoid prescribing the wrong drug, a corresponding alarm prompt will be output in this embodiment.
[0075] In addition, in this embodiment, when searching for drugs banned in clinical trials, users can use the following method: Step 1: Enter the drug information into the information database for drug storage according to the above-mentioned drug classification method; Step 2: Obtain drug data from the database, and analyze and preprocess the drug information to generate an XML file; Step 3: Serialize the generated drug information XML file into JSON format data; Step 4: Transmit the serialized JSON format data to the WEB server via the network; Step 5: When the client browser sends a request for drug query, the server responds to the browser's request and returns the drug information in JSON format to the browser; Step 6: The browser parses the returned JSON data; Step 7: Display the query results.
[0076] Furthermore, the drug storage information database is established in the drug management system using the MVC mode, that is, the input, processing and output processes of drug information are separated according to the model layer, view layer and control layer. The drug management system is programmed using Java technology, and the user-friendly and powerful DREAMWEAVER is used as the development tool. The server-side JAVA language is used, the client script is written in JavaScript, the database uses the Oracle 19c database, the server middleware uses WebSphere Application Server 8.5, and the Struts session management, filter and database integration technology are used to build the Web application, and the web page is written using VUE2 technology. Among them, step 3 is specifically as follows: step 3.1: load the drug information XML file into the memory; step 3.2: parse the XML file to form a tree structure; step 3.3: traverse the tree, find the node Node in the tree, and obtain the basic unit Element; step 3.4: parse into JSON data format; step 6 is specifically as follows: step 6.1: the browser receives the JSON data responded by the server; step 6.2: the browser parses the JSON data and obtains the corresponding JSONObject data; step 6.3: go to step 7.
[0077] In this embodiment, the corresponding disease type and mechanism of action are determined based on the input drug information of the drug to be queried; the drug information after processing the historical drug information is mapped to the high-dimensional feature space through the Gaussian kernel function, and the drug information data set after spatial transformation is obtained through kernel principal component analysis. The drug classification model is constructed based on the drug information data set and using the Gaussian Bayes algorithm. Finally, the drug classification corresponding to the drug to be queried is determined through the drug classification model; and the corresponding alarm prompt is output when the banned drug is detected. Through the above method, the drug classification is made more targeted, and the accuracy of the drug classification is improved, so that it is more in line with the clinical trial research plan, which can effectively reduce the workload of research doctors, reduce the risks in the clinical trial process, and improve the quality of clinical research.
[0078] Reference Figure 3 , Figure 3 This is a structural block diagram of the first embodiment of the drug classification device based on drugs banned in clinical trials of the present invention.
[0079] like Figure 3 As shown, the drug classification device based on clinical trial banned drugs proposed in the embodiment of the present invention includes:
[0080] The receiving module 10 is used to receive the input drug information of the drug to be queried;
[0081] An acquisition module 20, for determining the disease type and action mechanism corresponding to the drug to be queried according to the drug information;
[0082] A construction module 30 is used to obtain historical drug information, process the historical drug information, and obtain processed drug information;
[0083] The construction module 30 is used to map the processed drug information to a high-dimensional feature space through a Gaussian kernel function, and extract feature information through kernel principal component analysis to obtain a drug information data set after spatial transformation;
[0084] The construction module 30 is used to divide the spatially transformed drug information dataset into a training dataset and a test dataset;
[0085] The construction module 30 is used to perform model training based on the training data set using the Gaussian Bayes algorithm, and to evaluate and verify the trained model using the test data set to complete the construction of the drug classification model;
[0086] A classification module 40, for inputting the disease type and the mechanism of action into the drug classification model to determine the drug classification corresponding to the drug to be queried;
[0087] The prompt module 50 is used to output a corresponding warning prompt if the drug classification of the drug to be queried is a prohibited drug.
[0088] In this embodiment, the corresponding disease type and mechanism of action are determined based on the input drug information of the drug to be queried; the drug information after processing the historical drug information is mapped to the high-dimensional feature space through the Gaussian kernel function, and the drug information data set after spatial transformation is obtained through kernel principal component analysis. The drug classification model is constructed based on the drug information data set and using the Gaussian Bayes algorithm. Finally, the drug classification corresponding to the drug to be queried is determined through the drug classification model; and the corresponding alarm prompt is output when the banned drug is detected. Through the above method, the drug classification is made more targeted, and the accuracy of the drug classification is improved, so that it is more in line with the clinical trial research plan, which can effectively reduce the workload of research doctors, reduce the risks in the clinical trial process, and improve the quality of clinical research.
[0089] In some embodiments, the construction module 30 is used to divide the historical drug information into several initial data sets based on disease type and mechanism of action;
[0090] An initial data set is randomly selected as the current data set, and several samples are selected from the remaining initial data sets to be added to the current data set, until several reference data sets are obtained, and the number of samples between the reference data sets is balanced;
[0091] Standardized automatic scaling is used to scale the features of samples in any reference data so that the features of samples in each reference data are at the same level, and processed drug information is obtained based on several reference data sets that have completed feature scaling.
[0092] In some embodiments, the expression corresponding to scaling the features of samples in any reference data using standardized automatic scaling is: j '=(S j -avg) / d, where S j ' is the eigenvalue after feature j is scaled, S j is the original eigenvalue of feature j before scaling, avg is the average eigenvalue of any reference data set, and d is the variance corresponding to the eigenvalue of any reference data set.
[0093] In some embodiments, the construction module 30 is used to determine the prior probability of each category in the training set;
[0094] Determine the conditional probability of each attribute under each category;
[0095] Determine the distance between all samples of one category and all samples of another category based on the prior probability of each category and the conditional probability of each attribute under each category;
[0096] Determine the weighted posterior probability using the class-specific attribute weight matrix and class conditional probability, and calculate the corresponding classification accuracy;
[0097] Iterative optimization is performed using the NSGA-II algorithm to determine the spacing and classification accuracy between different categories, and a set of Pareto solutions are recorded;
[0098] When the termination condition of model optimization is reached, the Pareto solution set is output to obtain multiple sets of class-specific attribute weight matrices;
[0099] Using cross-validation evaluation to select a number of optimal class-specific attribute weight matrices from the multiple groups of class-specific attribute weight matrices;
[0100] Using the posterior probability of the test data set based on the several optimal class-specific attribute weight matrices under the test data set;
[0101] Based on the posterior probability, select the category with the highest posterior probability as the predicted category of the test sample, and return the predicted category label;
[0102] The classification accuracy is determined based on the category label, and when the classification accuracy reaches a preset accuracy, the training of the drug classification model is determined to be completed.
[0103] In some embodiments, the expression corresponding to determining the distance between all samples of one category and all samples of another category is: Among them, p(x j |Y0) and p(x j |Y1) are x in the Y0 category j Attributes and Y1 categories under x j The conditional probability of the attribute, P(Y0) and P(Y1) are the prior probabilities of category Y0 and category Y1 respectively, m is the number of samples, and n is the number of features.
[0104] In some embodiments, the classification accuracy is expressed as: Among them, δ(·) is a binary function, which is used to determine the number of samples that the model classifies correctly, y i represents the predicted category label, and y represents the true label.
[0105] An embodiment of the present application also provides a drug classification device based on drugs banned in clinical trials, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus, and the memory is used to store a drug classification program based on drugs banned in clinical trials; the processor is used to implement the above-mentioned drug classification method based on drugs banned in clinical trials when executing the program stored in the memory.
[0106] The communication bus mentioned in the drug classification device based on the banned drugs in clinical trials can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0107] The communication interface is used for communication between the above-mentioned drug classification device based on drugs banned in clinical trials and other devices.
[0108] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0109] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, and discrete hardware components.
[0110] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive Solid State Disk (SSD)), etc.
[0111] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0112] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0113] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0114] It should be understood that the above is only an example and does not constitute any limitation on the technical solution of the present invention. In specific applications, technicians in this field can make settings as needed, and the present invention does not limit this.
[0115] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of the present invention. In practical applications, technicians in this field can select part or all of them according to actual needs to achieve the purpose of the present embodiment, and no limitation is made here.
[0116] In addition, for technical details not fully described in this embodiment, reference can be made to the drug classification method based on banned drugs in clinical trials provided in any embodiment of the present invention, which will not be described in detail here.
[0117] In addition, it should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0118] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0119] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory (ROM) / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0120] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
[0121] It is understandable that the system provided by the embodiment of the present invention corresponds to the method provided by the embodiment of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts in the above method.
Claims
1. A drug classification method based on drugs prohibited for clinical trials, characterized in that: The drug classification method based on prohibited drugs for clinical trials includes: Receiving input drug information of a drug to be queried; Determine the disease type and action mechanism corresponding to the drug to be queried according to the drug information; Acquiring historical drug information, and processing the historical drug information to obtain processed drug information; Mapping the processed drug information to a high-dimensional feature space through a Gaussian kernel function, and extracting feature information through kernel principal component analysis to obtain a drug information data set after spatial transformation; Dividing the spatially transformed drug information dataset into a training dataset and a test dataset; Using the Gaussian Bayesian algorithm to perform model training based on the training data set, and using the test data set to evaluate and verify the trained model to complete the construction of the drug classification model; Inputting the disease type and the mechanism of action into the drug classification model to determine the drug classification corresponding to the drug to be queried; If the drug classification of the drug to be queried is a prohibited drug, a corresponding alarm prompt is output.
2. The drug classification method based on clinical trial banned drugs according to claim 1, characterized in that: The processing of the historical drug information to obtain processed drug information includes: dividing the historical drug information into a number of initial data sets based on disease type and mechanism of action; An initial data set is randomly selected as the current data set, and several samples are selected from the remaining initial data sets to be added to the current data set, until several reference data sets are obtained, and the number of samples between the reference data sets is balanced; Standardized automatic scaling is used to scale the features of samples in any reference data so that the features of samples in each reference data are at the same level, and processed drug information is obtained based on several reference data sets that have completed feature scaling.
3. The drug classification method based on clinical trial prohibited drugs according to claim 2, characterized in that: The expression corresponding to scaling the features of samples in any reference data using standardized automatic scaling is: j '=(S j -avg) / d, where S j ' is the eigenvalue after feature j is scaled, S j is the original eigenvalue of feature j before scaling, avg is the average eigenvalue of any reference data set, and d is the variance corresponding to the eigenvalue of any reference data set.
4. The drug classification method based on clinical trial prohibited drugs according to claim 1, characterized in that: The method of using the Gaussian Bayesian algorithm to perform model training based on the training data set, and using the test data set to evaluate and verify the trained model to complete the construction of the drug classification model, includes: Determining a priori probabilities for each class in the training set; Determine the conditional probability of each attribute under each category; Determine the distance between all samples of one category and all samples of another category based on the prior probability of each category and the conditional probability of each attribute under each category; Determine the weighted posterior probability using the class-specific attribute weight matrix and class conditional probability, and calculate the corresponding classification accuracy; Iterative optimization is performed using the NSGA-II algorithm to determine the spacing and classification accuracy between different categories, and a set of Pareto solutions are recorded; When the termination condition of model optimization is reached, the Pareto solution set is output to obtain multiple sets of class-specific attribute weight matrices; Using cross-validation evaluation to select a number of optimal class-specific attribute weight matrices from the multiple groups of class-specific attribute weight matrices; Using the posterior probability of the test data set based on the several optimal class-specific attribute weight matrices under the test data set; Based on the posterior probability, select the category with the highest posterior probability as the predicted category of the test sample, and return the predicted category label; The classification accuracy is determined based on the category label, and when the classification accuracy reaches a preset accuracy, the training of the drug classification model is determined to be completed.
5. The drug classification method based on clinical trial prohibited drugs according to claim 4, characterized in that: The expression corresponding to determining the distance between all samples of one category and all samples of another category is: Among them, p(x j |Y0) and p(x j |Y1) are x in the Y0 category j Attributes and Y1 categories under x j The conditional probability of the attribute, P(Y0) and P(Y1) are the prior probabilities of category Y0 and category Y1 respectively, m is the number of samples, and n is the number of features.
6. The drug classification method based on clinical trial prohibited drugs according to claim 4, characterized in that: The expression of the classification accuracy is: Among them, δ(·) is a binary function, which is used to determine the number of samples that the model classifies correctly, y i represents the predicted category label, and y represents the true label.
7. A drug classification device based on drugs prohibited in clinical trials, characterized in that: The drug classification device based on the banned drugs in clinical trials comprises: A receiving module, used for receiving input drug information of a drug to be queried; An acquisition module, used to determine the disease type and action mechanism corresponding to the drug to be queried according to the drug information; A construction module is used to obtain historical drug information, process the historical drug information, and obtain processed drug information; The construction module is used to map the processed drug information to a high-dimensional feature space through a Gaussian kernel function, and extract feature information through kernel principal component analysis to obtain a drug information data set after spatial transformation; The construction module is used to divide the spatially transformed drug information dataset into a training dataset and a test dataset; The construction module is used to perform model training based on the training data set using the Gaussian Bayes algorithm, and to evaluate and verify the trained model using the test data set to complete the construction of the drug classification model; A classification module, used for inputting the disease type and the mechanism of action into the drug classification model to determine the drug classification corresponding to the drug to be queried; The prompt module is used to output a corresponding alarm prompt if the drug classification of the drug to be queried is a banned drug.
8. The drug classification device based on clinical trial banned drugs according to claim 7, characterized in that: The building module is used to divide the historical drug information into several initial data sets based on disease type and mechanism of action; An initial data set is randomly selected as the current data set, and several samples are selected from the remaining initial data sets to be added to the current data set, until several reference data sets are obtained, and the number of samples between the reference data sets is balanced; Standardized automatic scaling is used to scale the features of samples in any reference data so that the features of samples in each reference data are at the same level, and processed drug information is obtained based on several reference data sets that have completed feature scaling.
9. The drug classification device based on clinical trial banned drugs according to claim 8, characterized in that: The expression corresponding to scaling the features of samples in any reference data using standardized automatic scaling is: j '=(S j -avg) / d, where S j ' is the eigenvalue after feature j is scaled, S j is the original eigenvalue of feature j before scaling, avg is the average eigenvalue of any reference data set, and d is the variance corresponding to the eigenvalue of any reference data set.
10. A drug classification device based on drugs prohibited in clinical trials, characterized in that: The drug classification device based on drugs banned in clinical trials comprises: a memory, a processor, and a drug classification program based on drugs banned in clinical trials stored in the memory and executable on the processor, wherein the drug classification program based on drugs banned in clinical trials is configured to implement the steps of the drug classification method based on drugs banned in clinical trials as described in any one of claims 1 to 6.
Citation Information
Cited By
Method, apparatus, and device for classifying drug on the basis of drugs prohibited in clinical trial
WO2026152625A1