Drug screening analysis management method and system based on data analysis
By structuring and grouping drug screening data sources, combining clustering algorithms and knowledge graph embedding, and generating resource optimization configuration plans, the problems of multi-dimensional data fusion and resource allocation in traditional drug screening analysis and management are solved, achieving efficient screening and resource optimization, and improving R&D efficiency and success rate.
Patent Information
- Application Number
- CN202510817596.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN120708938A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pharmaceutical management, and in particular to a drug screening analysis management method and system based on data analysis. Background Art
[0002] In the field of pharmaceutical research and development, drug screening and analysis management, as a key link, plays a decisive role in research and development efficiency and success rate.
[0003] While existing drug screening techniques exist within traditional drug screening analysis and management systems, these technologies have significant limitations. For example, some R&D processes prioritize flexible resource allocation and the comprehensive utilization of multidimensional data. They aim to integrate multidimensional characteristics such as structural novelty and reusability during drug screening, while also dynamically adjusting resource allocation based on the real-time priority of candidates. However, traditional methods struggle to achieve multidimensional data integration and dynamic resource allocation, making them unable to fully meet the demands for efficient screening and resource optimization during the R&D process. Summary of the Invention
[0004] Based on this, it is necessary to provide a drug screening analysis management method and system based on data analysis to address the above technical problems, so as to realize multi-dimensional data fusion and dynamic allocation of resources, and achieve the technical effects of efficient screening and resource optimization.
[0005] In a first aspect, the present application provides a drug screening analysis management method based on data analysis, comprising:
[0006] Obtaining and processing drug screening data sources to obtain structured data;
[0007] Based on a preset activity threshold, a clustering algorithm is used to feature-group the structured data to obtain a first data set, where the first data set includes a high activity group, a medium activity group, and a low activity group;
[0008] The first dataset is processed using a clustering algorithm and knowledge graph embedding to obtain a second dataset with candidate labels;
[0009] generating a ranked list based on the second data set, the ranked list including activity values, structural novelty coefficients, and reuse value weights of the drugs;
[0010] The sorted list is matched with the preset resource allocation strategy to generate a resource optimization plan; the resource optimization plan includes the scheduling of laboratory equipment use, the division of labor among R&D personnel, and the budget allocation ratio;
[0011] Input the ranked list into the preset knowledge graph analysis module to obtain candidate data that meets the preset potential conditions;
[0012] Build a dynamic monitoring model based on candidate data and generate R&D stage reports, which include R&D progress and resource consumption efficiency.
[0013] Furthermore, a dynamic monitoring model is built based on candidate data to generate R&D stage reports, including:
[0014] Based on candidate data, a time series model in a dynamic monitoring model is constructed;
[0015] The following formula is used to extract R&D progress indicators from the time series model and combine them with resource consumption data to calculate resource consumption efficiency:
[0016]
[0017]
[0018] Among them, P t represents the R&D progress indicator, R i Indicates the completion degree of the i-th R&D task, T i represents the planned time of the i-th task, D j represents the number of deliverables in the jth phase, E j represents the expected number of results in the jth stage, α and β represent weight coefficients, C e Indicates resource consumption efficiency, M k represents the material cost in the kth time period, L k represents the labor cost in the kth time period, S k represents the equipment cost in the kth time period, and t represents the total time period;
[0019] Use clustering algorithms to classify R&D progress indicators and resource consumption efficiency, and obtain key features and clustering classification results for different R&D stages;
[0020] Generate a primary R&D stage report based on key features and cluster classification results.
[0021] Furthermore, a dynamic monitoring model is built based on candidate data to generate R&D stage reports, including:
[0022] Based on the primary R&D stage report, natural language processing technology is used to perform structured processing to obtain structured processing results;
[0023] Generate R&D stage report based on structured processing results.
[0024] Furthermore, based on the second data set, a sorted list is generated, including:
[0025] The second dataset is processed using a feature extraction algorithm using the following formula to obtain the feature value of each candidate:
[0026]
[0027] Among them, F i represents the eigenvalue of the i-th candidate, M represents the total number of feature dimensions, α k represents the weight coefficient of the kth feature, represents the extraction function of the k-th dimension feature of the i-th candidate;
[0028] Based on the characteristic value of each candidate, a clustering algorithm is used to perform classification processing to obtain the group value corresponding to the candidate;
[0029] Based on the grouping value, generate the classification results corresponding to the candidate;
[0030] Generate a ranked list based on the classification results using the following formula:
[0031]
[0032] Among them, R i represents the ranking score of the i-th candidate, β represents the normalization coefficient, represents the set of candidates in the same group as candidate i, S ij represents the similarity between candidates i and j, Indicates the number of candidates in the same group.
[0033] Furthermore, the sorted list is matched with the preset resource allocation strategy to generate a resource optimization configuration plan, including:
[0034] Get the sorted list and preset resource allocation strategy;
[0035] Use the following formula to process the sorted list and the preset resource allocation strategy using a matching algorithm to obtain the resource optimization configuration plan:
[0036]
[0037]
[0038] Where E represents the penalty value for insufficient resource allocation, n represents the number of resource types, m represents the number of demand objects, γ represents the penalty coefficient, L represents the minimum demand, A represents the actual allocation, S represents the total score of resource allocation, α represents the weight coefficient of resource allocation, R represents the matching degree of resource allocation, and P represents the priority of resource allocation.
[0039] Furthermore, the ranked list is input into a preset knowledge graph analysis module to obtain candidate data that meets the preset potential conditions, including:
[0040] Use graph embedding algorithms to map the sorted list into the knowledge graph space to obtain a structured vector representation;
[0041] Based on the structured vector representation and the knowledge graph nodes in the knowledge graph space, similarity calculation is performed to obtain the similarity value between the structured vector and the knowledge graph nodes;
[0042] Based on the similarity value and the preset matching rules, potential matching nodes are obtained. Potential matching nodes are entities in the candidate that have high semantic relevance to the knowledge graph space.
[0043] Based on the potential matching nodes, entities that meet the preset potential conditions are determined as candidates;
[0044] Use the trained machine learning model to extract features from the candidate objects and obtain a set of feature vectors;
[0045] Use clustering algorithm to group the feature vector set and obtain candidate classification results;
[0046] Based on the candidate classification results and knowledge graph nodes, candidate data that meets the preset potential conditions is generated.
[0047] Furthermore, based on a preset activity threshold, a clustering algorithm is used to feature-group the structured data to obtain a first data set, including:
[0048] Based on the preset activity threshold, the feature quantities in the structured data are processed to obtain a feature set;
[0049] Use clustering algorithms to cluster the feature set and obtain different activity groups;
[0050] Based on different activity groups, a differentiation algorithm was used for optimization to obtain the first data set.
[0051] In a second aspect, the present application further provides a drug screening analysis and management system based on data analysis, the system comprising:
[0052] The data acquisition module is used to obtain the drug screening data source and process the drug screening data source to obtain structured data;
[0053] A first processing module is configured to perform feature grouping on the structured data using a clustering algorithm based on a preset activity threshold to obtain a first data set, where the first data set includes a high activity group, a medium activity group, and a low activity group;
[0054] A second processing module is used to process the first data set using a clustering algorithm and knowledge graph embedding to obtain a second data set with candidate labels;
[0055] a third processing module, configured to generate a ranking list based on the second data set, the ranking list including activity values, structural novelty coefficients, and reuse value weights of the drugs;
[0056] The resource allocation module is used to match the sorted list with the preset resource allocation strategy and generate a resource optimization plan. The resource optimization plan includes the scheduling of laboratory equipment use, the division of labor among R&D personnel, and the budget allocation ratio.
[0057] The fourth processing module is used to input the sorted list into a preset knowledge graph analysis module to obtain candidate data that meets the preset potential conditions;
[0058] The dynamic monitoring module is used to build a dynamic monitoring model based on candidate data and generate R&D stage reports, which include R&D progress and resource consumption efficiency.
[0059] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any method in the first aspect of the present application are implemented.
[0060] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any method in the first aspect of the present application when the computer program is executed by a processor.
[0061] The technical solution provided by the present application includes the following technical effects: by providing a drug screening analysis and management method based on data analysis, the method includes: obtaining a drug screening data source and processing the drug screening data source to obtain structured data; based on a preset activity threshold, using a clustering algorithm to feature group the structured data to obtain a first data set, the first data set including a high activity group, a medium activity group and a low activity group; using a clustering algorithm and knowledge graph embedding to process the first data set to obtain a second data set with candidate labels; based on the second data set, generating a ranked list, the ranked list including the activity value, structural novelty coefficient and reuse value weight of the drug; matching and calculating the ranked list with a preset resource allocation strategy to generate a resource optimization configuration plan; the resource optimization configuration plan includes laboratory equipment usage scheduling, R&D personnel division of labor and budget allocation ratio; inputting the ranked list into a preset knowledge graph analysis module to obtain candidate data that meets the preset potential conditions; constructing a dynamic monitoring model based on the candidate data to generate an R&D stage report, the R&D stage report including R&D progress and resource consumption efficiency, so as to realize multi-dimensional data fusion and dynamic resource allocation, and achieve the technical effects of efficient screening and resource optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0063] Figure 1 This is a flow chart of a drug screening analysis management method based on data analysis in one embodiment of the present invention;
[0064] Figure 2 This is a structural diagram of a drug screening analysis and management system based on data analysis in one embodiment of the present invention. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0066] like Figure 1 As shown, the present application provides a drug screening analysis management method based on data analysis, comprising:
[0067] S101: Obtain a drug screening data source, and process the drug screening data source to obtain structured data.
[0068] Specifically, drug screening data sources include laboratory experimental results, clinical trial data, molecular simulation data, drug chemical properties databases, and related literature. Drug screening data sources are diverse and complex, containing a large amount of noise and irrelevant information, so they need to be processed to obtain structured data.
[0069] S102: Based on a preset activity threshold, a clustering algorithm is used to feature-group the structured data to obtain a first data set, where the first data set includes a high-activity group, a medium-activity group, and a low-activity group.
[0070] Specifically, the preset activity threshold is a standard set based on drug development experience and needs, used to distinguish drugs with different activity levels. A clustering algorithm is an unsupervised learning algorithm that automatically divides data into different groups based on similarities. During processing, the clustering algorithm analyzes features within structured data, such as the chemical structure of the drug molecules and activity test results, and categorizes the data into high-activity, medium-activity, and low-activity groups based on the similarity of these features. Drug molecules in the high-activity group exhibit strong biological activity in activity tests and are potential drug candidates; drug molecules in the medium-activity group have moderate activity and require further optimization and research; and drug molecules in the low-activity group have weak activity and are not worth further development under current screening conditions. By grouping data by feature, drug molecules with different activity levels can be initially screened, providing guidance for subsequent in-depth analysis and research, improving the efficiency and success rate of drug development, and optimizing resource allocation, allowing R&D teams to devote more attention and resources to drug candidates with high potential.
[0071] S103: Process the first data set using a clustering algorithm and knowledge graph embedding to obtain a second data set with candidate labels.
[0072] Specifically, clustering algorithms can group drug molecules based on their activity, structure, and other characteristics, while knowledge graph embedding can convert drug-related knowledge into vector representations, enhancing the relevance and semantic information of the data, thereby making the labels of candidate substances more accurate and comprehensive, and helping to more efficiently screen out potential drug molecules in the future, thereby improving R&D efficiency and success rate.
[0073] S104: Based on the second data set, a ranking list is generated, where the ranking list includes the activity value, structural novelty coefficient, and reuse value weight of the drug.
[0074] Specifically, first, the second data set is processed using a feature extraction algorithm to obtain the characteristic values of each candidate; then, based on the characteristic values of each candidate, a clustering algorithm is used to perform classification processing to obtain the grouping values corresponding to the candidates; finally, based on the grouping values, a ranked list corresponding to the candidates is generated. This list comprehensively considers the drug's activity value, structural novelty coefficient and reuse value weight to determine the priority of the candidates.
[0075] S105: Matching and calculating the sorted list with the preset resource allocation strategy to generate a resource optimization configuration plan; the resource optimization configuration plan includes laboratory equipment usage scheduling, R&D personnel division of labor, and budget allocation ratio.
[0076] Specifically, a matching algorithm is used to match the ranking score of each candidate in the ranking list with the preset resource allocation strategy, and the priority and demand for each candidate in different resource types are calculated; based on the above priority and demand, a resource optimization allocation plan is generated, which specifically includes the scheduling of laboratory equipment use, the division of labor among R&D personnel, and the budget allocation ratio. This step can reasonably arrange the use time of laboratory equipment according to the priority and resource requirements of drug candidates, optimize the task allocation of R&D personnel, and scientifically allocate budgets to ensure that key projects receive sufficient financial support. This resource optimization allocation plan not only helps to shorten the R&D cycle, but also reduces R&D costs and improves overall R&D efficiency.
[0077] S106: Input the sorted list into a preset knowledge graph analysis module to obtain candidate data that meets the preset potential conditions.
[0078] Specifically, the ranked list is input into a preset knowledge graph analysis module, and the ranked list is mapped to the knowledge graph space through a graph embedding algorithm to obtain a structured vector representation. Similarity calculation is performed based on the structured vector representation and the knowledge graph nodes in the knowledge graph space, and finally the candidate data that meets the preset potential conditions is determined. This process can comprehensively consider multi-dimensional characteristics such as the drug's activity value, structural novelty coefficient, and reuse value weight, thereby improving the accuracy and efficiency of drug screening.
[0079] S107: Build a dynamic monitoring model based on the candidate data and generate a research and development phase report. The research and development phase report includes research and development progress and resource consumption efficiency.
[0080] Specifically, a dynamic monitoring model built based on candidate data can track and analyze various indicators during the R&D process in real time, generating reports on R&D progress and resource efficiency. Progress is measured by analyzing task completion and phased deliverables, while resource efficiency is determined by evaluating the use of materials, labor, and equipment costs. This helps R&D teams adjust strategies in a timely manner, optimize resource allocation, and improve R&D efficiency and success rates.
[0081] The embodiment of the present application provides a drug screening analysis and management method based on data analysis, which obtains structured data by obtaining a drug screening data source and processing the drug screening data source; based on a preset activity threshold, a clustering algorithm is used to feature group the structured data to obtain a first data set, the first data set including a high activity group, a medium activity group, and a low activity group; the first data set is processed using a clustering algorithm and knowledge graph embedding to obtain a second data set with candidate labels; based on the second data set, a ranked list is generated, the ranked list including the activity value, structural novelty coefficient, and reuse value weight of the drug; the ranked list is matched and calculated with a preset resource allocation strategy to generate a resource optimization configuration plan; the resource optimization configuration plan includes laboratory equipment usage scheduling, R&D personnel division of labor, and budget allocation ratio; the ranked list is input into a preset knowledge graph analysis module to obtain candidate data that meets preset potential conditions; a dynamic monitoring model is constructed based on the candidate data to generate an R&D stage report, the R&D stage report including R&D progress and resource consumption efficiency, so as to realize multi-dimensional data fusion and dynamic resource allocation, and achieve the technical effects of efficient screening and resource optimization.
[0082] Furthermore, a dynamic monitoring model is built based on candidate data to generate R&D stage reports, including:
[0083] Based on candidate data, a time series model in a dynamic monitoring model is constructed;
[0084] The following formula is used to extract R&D progress indicators from the time series model and combine them with resource consumption data to calculate resource consumption efficiency:
[0085]
[0086] Among them, P t represents the R&D progress indicator, R i Indicates the completion degree of the i-th R&D task, T i represents the planned time of the i-th task, D j represents the number of deliverables in the jth phase, E j represents the expected number of results in the jth stage, α and β represent weight coefficients, C e Indicates resource consumption efficiency, M k represents the material cost in the kth time period, L k represents the labor cost in the kth time period, S k represents the equipment cost in the kth time period, and t represents the total time period;
[0087] Use clustering algorithms to classify R&D progress indicators and resource consumption efficiency, and obtain key features and clustering classification results for different R&D stages;
[0088] Generate a primary R&D stage report based on key features and cluster classification results.
[0089] Specifically, first, a time series model is constructed within the dynamic monitoring model based on candidate data. By analyzing the performance of candidates at different time points, trends and patterns within the R&D process are captured. Next, a formula is used to extract R&D progress indicators from the time series model and, combined with resource consumption data, calculate resource consumption efficiency. R&D progress indicators are measured by analyzing task completion and phased deliverables, while resource consumption efficiency is determined by evaluating the use of materials, manpower, and equipment costs. Next, a clustering algorithm is used to classify R&D progress indicators and resource consumption efficiency, resulting in key characteristics and cluster classification results for different R&D phases. Finally, based on these key characteristics and cluster classification results, a preliminary R&D phase report is generated. This report helps R&D teams stay informed of project progress, identify potential issues, and optimize resource allocation, thereby improving R&D efficiency and success rates.
[0090] Furthermore, a dynamic monitoring model is built based on candidate data to generate R&D stage reports, including:
[0091] Based on the primary R&D stage report, natural language processing technology is used to perform structured processing to obtain structured processing results;
[0092] Generate R&D stage report based on structured processing results.
[0093] Specifically, based on the initial R&D phase report, natural language processing technology is used to perform structured processing to obtain structured processing results. Natural language processing technology can convert unstructured text data into structured data, facilitating storage, query, and analysis, thereby enhancing the value of the data. The structured processing results can more intuitively display key information in the R&D process, such as R&D progress and resource consumption, to support subsequent decision-making. Finally, based on the structured processing results, an R&D phase report is generated. This report can help the R&D team to promptly understand project progress, identify potential problems, optimize resource allocation, and improve R&D efficiency and success rate.
[0094] Furthermore, based on the second data set, a sorted list is generated, including:
[0095] The second dataset is processed using a feature extraction algorithm using the following formula to obtain the feature value of each candidate:
[0096]
[0097] Among them, F i represents the eigenvalue of the i-th candidate, M represents the total number of feature dimensions, α k represents the weight coefficient of the kth feature, represents the extraction function of the k-th dimension feature of the i-th candidate;
[0098] Based on the characteristic value of each candidate, a clustering algorithm is used to perform classification processing to obtain the group value corresponding to the candidate;
[0099] Based on the grouping value, generate the classification results corresponding to the candidate;
[0100] Generate a ranked list based on the classification results using the following formula:
[0101]
[0102] Among them, R i represents the ranking score of the i-th candidate, β represents the normalization coefficient, represents the set of candidates in the same group as candidate i, S ij represents the similarity between candidates i and j, Indicates the number of candidates in the same group.
[0103] Specifically, the second data set is processed using a feature extraction algorithm to obtain the characteristic values of each candidate. This process can extract representative information from the original data and improve the efficiency and effectiveness of data analysis. Then, based on the characteristic values of each candidate, a clustering algorithm is used for classification processing to obtain the grouping values corresponding to the candidate. Clustering algorithms such as K-Means can assign data points to different clusters based on the characteristics. Then, based on the grouping values, the classification results corresponding to the candidates are generated. Finally, a formula is used to generate a ranked list based on the classification results. The ranked list can comprehensively consider multi-dimensional features such as the drug's activity value, structural novelty coefficient, and reuse value weight, thereby improving the accuracy and efficiency of drug screening.
[0104] Furthermore, the sorted list is matched with the preset resource allocation strategy to generate a resource optimization configuration plan, including:
[0105] Get the sorted list and preset resource allocation strategy;
[0106] Use the following formula to process the sorted list and the preset resource allocation strategy using a matching algorithm to obtain the resource optimization configuration plan:
[0107]
[0108] Where E represents the penalty value for insufficient resource allocation, n represents the number of resource types, m represents the number of demand objects, γ represents the penalty coefficient, L represents the minimum demand, A represents the actual allocation, S represents the total score of resource allocation, α represents the weight coefficient of resource allocation, R represents the matching degree of resource allocation, and P represents the priority of resource allocation.
[0109] Specifically, a ranking list and a preset resource allocation strategy are obtained. The ranking list includes information such as the drug's activity value, structural novelty coefficient, and reuse value weight. The preset resource allocation strategy sets the rules and priorities for resource allocation based on R&D needs and resource constraints. Then, a matching algorithm is used to process the ranking list and the preset resource allocation strategy. By calculating parameters such as the penalty value for insufficient resource allocation, the number of resource types, the number of demand objects, the penalty coefficient, the minimum demand, the actual allocation, the total score of resource allocation, the weight coefficient, the matching degree, and the priority, a resource optimization configuration plan is obtained. This plan covers the scheduling of laboratory equipment use, the division of labor among R&D personnel, and the budget allocation ratio. It aims to improve R&D efficiency and resource utilization, ensure that key projects receive sufficient resource support, and thus improve overall R&D efficiency and success rate.
[0110] Furthermore, the ranked list is input into a preset knowledge graph analysis module to obtain candidate data that meets the preset potential conditions, including:
[0111] Use graph embedding algorithms to map the sorted list into the knowledge graph space to obtain a structured vector representation;
[0112] Based on the structured vector representation and the knowledge graph nodes in the knowledge graph space, similarity calculation is performed to obtain the similarity value between the structured vector and the knowledge graph nodes;
[0113] Based on the similarity value and the preset matching rules, potential matching nodes are obtained. Potential matching nodes are entities in the candidate that have high semantic relevance to the knowledge graph space.
[0114] Based on the potential matching nodes, entities that meet the preset potential conditions are determined as candidates;
[0115] Use the trained machine learning model to extract features from the candidate objects and obtain a set of feature vectors;
[0116] Use clustering algorithm to group the feature vector set and obtain candidate classification results;
[0117] Based on the candidate classification results and knowledge graph nodes, candidate data that meets the preset potential conditions is generated.
[0118] Specifically, a graph embedding algorithm is used to map the sorted list to the knowledge graph space to obtain a structured vector representation. This process can convert complex graph data into low-dimensional vectors, making the graph data more intuitive and easy to understand, and retaining the structural information and node attributes of the graph. Then, based on the structured vector representation and the knowledge graph nodes in the knowledge graph space, a similarity calculation is performed to obtain the similarity value between the structured vector and the knowledge graph node. This step can discover the potential relationship between the candidate and the entity in the knowledge graph. Then, based on the similarity value and the preset matching rules, potential matching nodes are obtained. The above-mentioned potential matching nodes are entities in the candidate that have a high semantic association with the knowledge graph space. Subsequently, based on the potential matching nodes, entities that meet the preset potential conditions are determined as candidates. Then, the trained machine learning model is used to extract features from the candidate to obtain a set of feature vectors. This step can extract representative information from the candidate and improve the efficiency and effectiveness of data analysis. Finally, a clustering algorithm is used to group the feature vector set to obtain the candidate classification results. Based on the candidate classification results and the knowledge graph nodes, candidate data that meets the preset potential conditions is generated. The above process can comprehensively consider multi-dimensional characteristics such as the drug's activity value, structural novelty coefficient, and reuse value weight, thereby improving the accuracy and efficiency of drug screening.
[0119] Furthermore, based on a preset activity threshold, a clustering algorithm is used to feature-group the structured data to obtain a first data set, including:
[0120] Based on the preset activity threshold, the feature quantities in the structured data are processed to obtain a feature set;
[0121] Use clustering algorithms to cluster the feature set and obtain different activity groups;
[0122] Based on different activity groups, a differentiation algorithm was used for optimization to obtain the first data set.
[0123] Specifically, first, based on the preset activity threshold, the feature quantities in the structured data are processed to obtain a feature set. This step can extract features related to drug activity from a large amount of data, providing a basis for subsequent analysis. Then, a clustering algorithm is used to cluster and group the feature set to obtain different activity groups. Clustering algorithms such as K-Means can assign data points to different clusters based on features, thereby distinguishing high-activity, medium-activity, and low-activity groups. Finally, based on different activity groups, a differentiation algorithm is used for optimization to obtain the first data set. The differentiation algorithm can further distinguish and optimize the features of different activity groups, improving the accuracy and availability of the data. This process can effectively conduct preliminary screening of drug candidates and provide a scientific basis for subsequent in-depth analysis and resource allocation.
[0124] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0125] In one embodiment, Figure 2 As shown, the present application also provides a drug screening analysis and management system 200 based on data analysis, including:
[0126] The data acquisition module 201 is used to acquire a drug screening data source and process the drug screening data source to obtain structured data;
[0127] A first processing module 202 is configured to perform feature grouping on the structured data using a clustering algorithm based on a preset activity threshold to obtain a first data set, where the first data set includes a high activity group, a medium activity group, and a low activity group;
[0128] A second processing module 203 is configured to process the first data set using a clustering algorithm and knowledge graph embedding to obtain a second data set with candidate labels;
[0129] A third processing module 204 is configured to generate a ranking list based on the second data set, the ranking list including the activity value, structural novelty coefficient, and reuse value weight of the drug;
[0130] Resource allocation module 205 is used to match the sorted list with the preset resource allocation strategy and generate a resource optimization plan. The resource optimization plan includes the laboratory equipment usage schedule, R&D personnel division of labor and budget allocation ratio;
[0131] The fourth processing module 206 is used to input the sorted list into a preset knowledge graph analysis module to obtain candidate data that meets the preset potential conditions;
[0132] The dynamic monitoring module 207 is used to build a dynamic monitoring model based on the candidate data and generate a research and development phase report, which includes the research and development progress and resource consumption efficiency.
[0133] Specifically, the data acquisition module 201 acquires and processes the drug screening data source to obtain structured data; the first processing module 202 performs feature grouping on the structured data based on a preset activity threshold and a clustering algorithm to obtain a first data set; the second processing module 203 uses a clustering algorithm and knowledge graph embedding to process the first data set to obtain a second data set with candidate labels; the third processing module 204 generates a ranked list containing drug activity values, structural novelty coefficients, and reuse value weights based on the second data set; the resource allocation module 205 matches the ranked list with the preset resource allocation strategy to calculate and generate a resource optimization allocation plan; the fourth processing module 206 inputs the ranked list into the knowledge graph analysis module to obtain candidate data that meets the preset potential conditions; the dynamic monitoring module 207 constructs a dynamic monitoring model based on the candidate data to generate a research and development stage report. The above system covers all aspects of drug screening analysis management, realizes full process management from data acquisition to resource optimization allocation and research and development process monitoring, effectively improves the efficiency and success rate of drug research and development, and provides a systematic solution for drug research and development.
[0134] In one embodiment, the dynamic monitoring module 207 is further configured to:
[0135] Based on candidate data, a time series model in a dynamic monitoring model is constructed;
[0136] The following formula is used to extract R&D progress indicators from the time series model and combine them with resource consumption data to calculate resource consumption efficiency:
[0137]
[0138] Among them, P t represents the R&D progress indicator, R i Indicates the completion degree of the i-th R&D task, T i represents the planned time of the i-th task, D j represents the number of deliverables in the jth phase, E j represents the expected number of results in the jth stage, α and β represent weight coefficients, C e Indicates resource consumption efficiency, M k represents the material cost in the kth time period, L k represents the labor cost in the kth time period, S k represents the equipment cost in the kth time period, and t represents the total time period;
[0139] Use clustering algorithms to classify R&D progress indicators and resource consumption efficiency, and obtain key features and clustering classification results for different R&D stages;
[0140] Generate a primary R&D stage report based on key features and cluster classification results.
[0141] In one embodiment, the dynamic monitoring module 207 is further configured to:
[0142] Based on the primary R&D stage report, natural language processing technology is used to perform structured processing to obtain structured processing results;
[0143] Generate R&D stage report based on structured processing results.
[0144] In one embodiment, the third processing module 204 is further configured to:
[0145] The second dataset is processed using a feature extraction algorithm using the following formula to obtain the feature value of each candidate:
[0146]
[0147] Among them, F i represents the eigenvalue of the i-th candidate, M represents the total number of feature dimensions, α k represents the weight coefficient of the kth feature, represents the extraction function of the k-th dimension feature of the i-th candidate;
[0148] Based on the characteristic value of each candidate, a clustering algorithm is used to perform classification processing to obtain the group value corresponding to the candidate;
[0149] Based on the grouping value, generate the classification results corresponding to the candidate;
[0150] Generate a ranked list based on the classification results using the following formula:
[0151]
[0152] Among them, R i represents the ranking score of the i-th candidate, β represents the normalization coefficient, represents the set of candidates in the same group as candidate i, S ij represents the similarity between candidates i and j, Indicates the number of candidates in the same group.
[0153] In one embodiment, the resource configuration module 205 is further configured to:
[0154] Get the sorted list and preset resource allocation strategy;
[0155] Use the following formula to process the sorted list and the preset resource allocation strategy using a matching algorithm to obtain the resource optimization configuration plan:
[0156]
[0157] Where E represents the penalty value for insufficient resource allocation, n represents the number of resource types, m represents the number of demand objects, γ represents the penalty coefficient, L represents the minimum demand, A represents the actual allocation, S represents the total score of resource allocation, α represents the weight coefficient of resource allocation, R represents the matching degree of resource allocation, and P represents the priority of resource allocation.
[0158] In one embodiment, the fourth processing module 206 is further configured to:
[0159] Use graph embedding algorithms to map the sorted list into the knowledge graph space to obtain a structured vector representation;
[0160] Based on the structured vector representation and the knowledge graph nodes in the knowledge graph space, similarity calculation is performed to obtain the similarity value between the structured vector and the knowledge graph nodes;
[0161] Based on the similarity value and the preset matching rules, potential matching nodes are obtained. Potential matching nodes are entities in the candidate that have high semantic relevance to the knowledge graph space.
[0162] Based on the potential matching nodes, entities that meet the preset potential conditions are determined as candidates;
[0163] Use the trained machine learning model to extract features from the candidate objects and obtain a set of feature vectors;
[0164] Use clustering algorithm to group the feature vector set and obtain candidate classification results;
[0165] Based on the candidate classification results and knowledge graph nodes, candidate data that meets the preset potential conditions is generated.
[0166] In one embodiment, the first processing module 202 is further configured to:
[0167] Based on the preset activity threshold, the feature quantities in the structured data are processed to obtain a feature set;
[0168] Use clustering algorithms to cluster the feature set and obtain different activity groups;
[0169] Based on different activity groups, a differentiation algorithm was used for optimization to obtain the first data set.
[0170] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0171] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0172] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0173] The above-described embodiments merely represent several implementation methods of the embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the concept of the embodiments of the present application, and these modifications and improvements fall within the scope of protection of the embodiments of the present application.
Claims
1. A drug screening analysis and management method based on data analysis, characterized in that: The method comprises: Obtaining a drug screening data source and processing the drug screening data source to obtain structured data; Based on a preset activity threshold, using a clustering algorithm to feature group the structured data to obtain a first data set, the first data set including a high activity group, a medium activity group, and a low activity group; Processing the first dataset using a clustering algorithm and knowledge graph embedding to obtain a second dataset with candidate labels; generating a ranked list based on the second data set, the ranked list including activity values, structural novelty coefficients, and reuse value weights of the drugs; Matching and calculating the sorted list with the preset resource allocation strategy to generate a resource optimization solution; the resource optimization solution includes the scheduling of laboratory equipment use, the division of labor among R&D personnel, and the budget allocation ratio; Inputting the ranked list into a preset knowledge graph analysis module to obtain candidate data that meets preset potential conditions; A dynamic monitoring model is constructed based on the candidate data to generate a research and development phase report, wherein the research and development phase report includes research and development progress and resource consumption efficiency.
2. A drug screening analysis and management method based on data analysis according to claim 1, characterized in that: The step of constructing a dynamic monitoring model based on the candidate data and generating a research and development phase report includes: Based on the candidate data, construct a time series model in the dynamic monitoring model; The following formula is used to extract the R&D progress indicator from the time series model and combine it with the resource consumption data to calculate the resource consumption efficiency: Among them, P t represents the R&D progress indicator, R i Indicates the completion degree of the i-th R&D task, T i represents the planned time of the i-th task, D j represents the number of deliverables in the jth phase, E j represents the expected number of results in the jth stage, α and β represent weight coefficients, C e Indicates resource consumption efficiency, M k represents the material cost in the kth time period, L k represents the labor cost in the kth time period, S k represents the equipment cost in the kth time period, and t represents the total time period; Using a clustering algorithm to classify the R&D progress indicators and the resource consumption efficiency to obtain key features and clustering classification results of different R&D stages; Based on the key features and the cluster classification results, a primary R&D stage report is generated.
3. A drug screening analysis and management method based on data analysis according to claim 2, characterized in that: The method of constructing a dynamic monitoring model based on the candidate data and generating a research and development stage report further includes: Based on the primary R&D stage report, natural language processing technology is used to perform structured processing to obtain a structured processing result; Based on the structured processing results, the R&D stage report is generated.
4. The drug screening analysis and management method based on data analysis according to claim 1, characterized in that: Generating a sorted list based on the second data set includes: The second data set is processed using a feature extraction algorithm using the following formula to obtain the feature value of each candidate: Among them, F i represents the eigenvalue of the i-th candidate, M represents the total number of feature dimensions, α k represents the weight coefficient of the kth feature, represents the extraction function of the k-th dimension feature of the i-th candidate; Based on the characteristic value of each candidate object, a clustering algorithm is used to perform classification processing to obtain a group value corresponding to the candidate object; Based on the grouping value, generating a classification result corresponding to the candidate; The sorted list is generated based on the classification results using the following formula: Among them, R i represents the ranking score of the i-th candidate, β represents the normalization coefficient, represents the set of candidates in the same group as candidate i, S ij represents the similarity between candidates i and j, Indicates the number of candidates in the same group.
5. The drug screening analysis and management method based on data analysis according to claim 1, characterized in that: The matching calculation of the sorted list with the preset resource allocation strategy to generate a resource optimization configuration plan includes: Obtaining the sorted list and the preset resource allocation strategy; The resource optimization configuration solution is obtained by processing the sorted list and the preset resource allocation strategy using a matching algorithm using the following formula: Where E represents the penalty value for insufficient resource allocation, n represents the number of resource types, m represents the number of demand objects, γ represents the penalty coefficient, L represents the minimum demand, A represents the actual allocation, S represents the total score of resource allocation, α represents the weight coefficient of resource allocation, R represents the matching degree of resource allocation, and P represents the priority of resource allocation.
6. The drug screening analysis and management method based on data analysis according to claim 1, characterized in that: The step of inputting the ranked list into a preset knowledge graph analysis module to obtain candidate data that meets preset potential conditions includes: Mapping the ranked list into the knowledge graph space using a graph embedding algorithm to obtain a structured vector representation; Performing similarity calculation based on the structured vector representation and the knowledge graph nodes in the knowledge graph space to obtain a similarity value between the structured vector and the knowledge graph nodes; Based on the similarity value and the preset matching rules, a potential matching node is obtained, where the potential matching node is an entity in the candidate that has a high semantic association with the knowledge graph space; Based on the potential matching nodes, determining the entities that meet the preset potential conditions as the candidates; Using a trained machine learning model to extract features from the candidate to obtain a set of feature vectors; Use clustering algorithm to group the feature vector set and obtain candidate classification results; Based on the candidate classification results and the knowledge graph nodes, the candidate data that meets the preset potential conditions is generated.
7. The drug screening analysis and management method based on data analysis according to claim 1, characterized in that: The method of performing feature grouping on the structured data using a clustering algorithm based on a preset activity threshold to obtain a first data set includes: Based on the preset activity threshold, processing the feature quantity in the structured data to obtain a feature set; Using a clustering algorithm to cluster and group the feature set to obtain different activity groups; Based on the different activity groups, a differentiation algorithm is used for optimization to obtain the first data set.
8. A drug screening analysis and management system based on data analysis, characterized in that: The system comprises: A data acquisition module is used to acquire a drug screening data source and process the drug screening data source to obtain structured data; A first processing module is configured to perform feature grouping on the structured data using a clustering algorithm based on a preset activity threshold to obtain a first data set, wherein the first data set includes a high activity group, a medium activity group, and a low activity group; A second processing module is configured to process the first data set using a clustering algorithm and knowledge graph embedding to obtain a second data set with candidate labels; a third processing module, configured to generate a ranking list based on the second data set, the ranking list including activity values, structural novelty coefficients, and reuse value weights of the drugs; A resource allocation module is used to match the sorted list with a preset resource allocation strategy to generate a resource optimization solution, which includes the scheduling of laboratory equipment use, the division of labor among R&D personnel, and the budget allocation ratio; A fourth processing module is configured to input the ranked list into a preset knowledge graph analysis module to obtain candidate data that meets preset potential conditions; A dynamic monitoring module is used to build a dynamic monitoring model based on the candidate data and generate a research and development stage report, wherein the research and development stage report includes research and development progress and resource consumption efficiency.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a drug screening analysis management method based on data analysis according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a drug screening analysis management method based on data analysis according to any one of claims 1 to 7 are implemented.