Drug research and development investment risk control value analysis method, device, equipment and medium

By constructing a risk analysis method for multi-source clinical data, the problem of inaccurate risk assessment in new drug research and development has been solved, efficient and accurate analysis of risk control values ​​has been achieved, and the efficiency and success rate of drug research and development have been improved.

CN120832550APending Publication Date: 2025-10-24CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510989269.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing technologies have problems in the new drug research and development process, such as inaccurate risk assessment, mismatch between risk and investment, low capital utilization, and lack of resource integration, which have hindered the continuous investment of R&D resources and restricted the development of innovative drugs.

Method used

By obtaining the R&D data of the target drug, extracting and matching similar R&D data of similar drugs, constructing multi-source clinical data, performing multi-dimensional feature encoding and feature fusion, and using the preset risk analysis model to perform probabilistic fusion analysis, risk control values ​​are generated.

Benefits of technology

It achieves efficient and accurate analysis of risk control values ​​in drug development, improves the scientificity and accuracy of risk assessment, provides a reliable basis for decision-making, and improves the efficiency and success rate of drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832550A_ABST
    Figure CN120832550A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a drug research and development input risk control value analysis method, device and equipment and a medium, the method comprises the following steps: obtaining research and development data of a target drug, and extracting stage research and development data of different research and development stages from the research and development data; similar research and development data of similar drugs are matched; collecting the two kinds of data into multi-source clinical data, and extracting structured and unstructured data of different clinical stages from the multi-source clinical data; performing data coding enhancement on the structured data and the non-structured data to obtain structural coding features and non-structural entity features, and performing feature fusion on the structural coding features and the non-structural entity features to obtain a target data feature space; constructing a risk factor index according to the target data feature space and determining a target risk dimension probability; and carrying out probability fusion on the target risk dimension probability to obtain a risk control value. The accuracy and objectivity of drug research and development risk control value analysis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data analysis, and in particular to a drug research and development investment risk control value analysis method, device, equipment and medium. BACKGROUND

[0002] Each research and development stage of new drug clinical has the characteristics of high risk, high investment and long cycle, and the existing investment risk control scheme suitable for different stages of new drug clinical research and development has the problems of inaccurate risk assessment, risk investment mismatch, low fund utilization rate and lack of resource integration, thereby hindering the continuous investment of research and development resources and limiting the research and development of innovative drugs.

[0003] For example, in the medical and health field, before an insurance institution plans to invest in a new anti-cancer drug research and development project, it needs to evaluate the risk of the research and development project according to the research and development progress. The traditional risk assessment scheme is: mainly relying on experts to analyze a small amount of existing preclinical data and making artificial judgments based on similar historical research and development projects. This scheme is difficult to quantify the dynamic risks of target verification failure, incorrect research and development path selection and other factors in research and development, resulting in less objective and accurate conclusions, which is easy to misjudge the investment of research and development project resources.

[0004] For another example, in the financial technology business, before a financial technology company helps an insurance company evaluate new drug research and development project investment. The traditional risk assessment scheme is: using DCF (Discounted: Cashflow Model, cash flow discount model) and other financial models for model analysis. This scheme cannot integrate biological clinical medical data due to the limitation of the model itself, and the analysis conclusion is usually large deviation and low credibility.

[0005] Therefore, how to effectively combine biological clinical data with model analysis to improve the accuracy and objectivity of drug research and development risk control value analysis has become a technical problem to be solved. SUMMARY

[0006] The present application provides a drug research and development investment risk control value analysis method, device, equipment and medium, which mainly aims to solve the problem of effectively combining biological clinical data with model analysis and improving the accuracy and objectivity of drug research and development risk control value analysis.

[0007] In a first aspect, to achieve the above-mentioned purpose, the present application provides a drug research and development investment risk control value analysis method, comprising: obtaining research and development data of a target drug, extracting stage research and development data of the target drug at different research and development stages from the research and development data, and matching similar research and development data of similar drugs according to the research and development data; The phase research and development data and the similar research and development data are collected as multi-source clinical data, and structured clinical data and unstructured clinical data of the multi-source clinical data at different clinical stages are extracted according to a preset phase division standard; The structured clinical data is subjected to multi-dimensional feature coding to obtain structure coding features, and the unstructured clinical data is subjected to entity recognition and relationship extraction to obtain unstructured entity features; The structure coding features and the unstructured entity features are subjected to multi-head attention feature fusion to obtain a target data feature space; A risk factor index of the multi-source clinical data is constructed according to the target data feature space; A plurality of target risk dimension probabilities of the multi-source clinical data are determined according to the risk factor index; The target risk dimension probabilities are subjected to probability fusion analysis by using a preset risk analysis model to obtain a risk control value of the multi-source clinical data.

[0008] In a second aspect, the present application further provides a drug research and development investment risk control value analysis device, comprising: A research and development data acquisition module is configured to acquire research and development data of a target drug, extract phase research and development data of the target drug at different research and development stages from the research and development data, and match similar research and development data of similar drugs according to the research and development data; A structured data extraction module is configured to collect the phase research and development data and the similar research and development data as multi-source clinical data, and extract structured clinical data and unstructured clinical data of the multi-source clinical data at different clinical stages according to a preset phase division standard; A data coding enhancement module is configured to perform multi-dimensional feature coding on the structured clinical data to obtain structure coding features, and perform entity recognition and relationship extraction on the unstructured clinical data to obtain unstructured entity features; A feature data fusion module is configured to perform multi-head attention feature fusion on the structure coding features and the unstructured entity features to obtain a target data feature space; A risk index determination module is configured to construct a risk factor index of the multi-source clinical data according to the target data feature space; A risk probability calculation module is configured to determine a plurality of target risk dimension probabilities of the multi-source clinical data according to the risk factor index; A risk probability fusion module is configured to perform probability fusion analysis on the target risk dimension probabilities by using a preset risk analysis model to obtain a risk control value of the multi-source clinical data.

[0009] In a third aspect, the present application further provides an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the above-mentioned drug research and development investment risk control value analysis method.

[0010] In a fourth aspect, the present application further provides a computer readable storage medium, wherein the computer readable storage medium stores at least one computer program, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned drug research and development investment risk control value analysis method.

[0011] In the embodiments of the present application, the target drug research and development data can be quickly and comprehensively collected by efficient data acquisition tools, breaking down information barriers, and accurately extracting stage research and development data at different research and development stages, ensuring data accuracy and consistency, providing reliable basis for research and development process monitoring, and improving the efficiency and success rate of drug research and development; through stage analysis, the changes of data at each stage can be captured in real time, potential risk patterns can be mined by using algorithm models, and semantic ambiguity between heterogeneous systems can be eliminated, while the clinical context information is retained, and the data analysis capability is improved; through multi-dimensional feature coding technology, the clinical key threshold information is retained, and the dimension difference is eliminated, the category type feature is converted into dense representation through embedding vector technology, and the dimension disaster problem caused by high base number category is solved; through multi-head attention feature fusion technology, the feature representation capability is significantly improved, the interference of modal difference on subsequent calculation is eliminated, and the information loss problem in heterogeneous data fusion is effectively solved; the efficiency and accuracy of multi-source clinical data risk analysis are realized, the contribution of each feature to the risk is automatically quantified, the subjective bias of manual screening is avoided, and the efficiency of investment research and development based on clinical data is improved; the probability weight coefficient is determined by the preset risk analysis model, the scientificity of risk assessment is improved, and the data comparability is enhanced; the quantifiable risk control value is converted into the comprehensive risk probability according to the preset standard, the digitization and standardization of risk assessment are realized, the computer can quickly process and analyze multi-source clinical data, and clear and accurate decision-making basis can be provided for subsequent risk control, which greatly improves the efficiency and accuracy of medical risk assessment. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creating labor intensity.

[0013] Figure 1 An application environment schematic diagram of a drug research and development investment risk control value analysis method according to an embodiment of the present application; Figure 2 A flowchart of a drug research and development investment risk control value analysis method according to an embodiment of the present application; Figure 3 A flowchart of determining a plurality of target risk dimension probabilities of multi-source clinical data according to risk factor indexes according to an embodiment of the present application; Figure 4 A module schematic diagram of a drug research and development investment risk control value analysis device according to an embodiment of the present application; Figure 5 A structure schematic diagram of an electronic device for implementing a drug research and development investment risk control value analysis method according to an embodiment of the present application; Figure 6 Another structure schematic diagram of an electronic device for implementing a drug research and development investment risk control value analysis method according to an embodiment of the present application.

[0014] The object, function characteristics and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0015] In order to make the person in the art better understand the technical solutions of the present disclosure, and understand the implementation process of how to apply technical means to solve the technical problems and achieve the corresponding technical effects, and to implement the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, not all the embodiments. The embodiments of the present disclosure and each feature in the embodiments can be combined with each other without conflict, and the formed technical solutions are all within the protection scope of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the present disclosure.

[0016] It should be noted that the terms "first", "second", and the like in the description and in the claims of the present disclosure and above-described accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular chronological or sequential order. It should be understood that the data thus used can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, device, product, or apparatus including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products, or apparatuses.

[0017] The embodiment of the present application provides a drug research and development investment risk control value analysis method, and the execution subject of the drug research and development investment risk control value analysis method includes but is not limited to at least one of electronic devices such as a server, a terminal and the like which can be configured to execute the device provided by the embodiment of the present application. In other words, the drug research and development investment risk control value analysis method can be executed by software or hardware installed in a terminal device or a server device. The server includes but is not limited to a single server, a server cluster, a cloud server or a cloud server cluster and the like. The server can be a stand-alone server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and big data and artificial intelligence platforms and the like basic cloud computing services.

[0018] The drug research and development investment risk control value analysis method can be applied to, for example Figure 1The application environment of the application is a drug research and development investment risk control value analysis method. The client communicates with the server through the network. The server can obtain the research and development data of the target drug through the client, break through the information barrier, accurately extract the stage research and development data of different research and development stages, ensure the accuracy and consistency of the data, provide reliable basis for the research and development process monitoring, improve the efficiency and success rate of drug research and development, capture the data changes of each stage in real time through stage analysis, use algorithm model to mine potential risk patterns, eliminate the semantic ambiguity between heterogeneous systems, retain the clinical context information while improving the data analysis ability, retain the clinical key threshold information through multi-dimensional feature coding technology, eliminate the dimension difference, convert the category type feature into dense representation through embedding vector technology, solve the dimension disaster problem caused by high base number category, significantly improve the feature representation ability through multi-head attention feature fusion technology, eliminate the interference of modal difference on subsequent calculation, effectively solve the information loss problem in heterogeneous data fusion, realize the efficiency and accuracy of multi-source clinical data risk analysis, automatically quantify the contribution of each feature to the risk, avoid the subjective bias of manual screening, improve the efficiency of investment research and development based on clinical data, determine the probability weight coefficient through the preset risk analysis model, improve the scientificity of risk assessment, enhance the data comparability, convert the comprehensive risk probability into quantifiable risk control value according to the preset standard, realize the digitization and standardization of risk assessment, facilitate the computer to quickly process and analyze multi-source clinical data, and provide clear and accurate decision basis for subsequent risk control, greatly improve the efficiency and accuracy of medical risk assessment, and finally output the risk control value feedback to the client. The client can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be realized by an independent server or a server cluster composed of multiple servers. The application will be described in detail through specific embodiments.

[0019] Referring to Figure 2 Fig. 1 is a flowchart of a drug research and development investment risk control value analysis method provided by an embodiment of the application. In this embodiment, the drug research and development investment risk control value analysis method comprises the following steps: S1, obtaining research and development data of a target drug, extracting stage research and development data of the target drug at different research and development stages from the research and development data, and matching similar research and development data of similar drugs according to the research and development data.

[0020] In the embodiment of the application, the research and development data of the target drug refers to collecting, sorting and various related information generated in the research and development process of a specific drug. These data run through the entire stage from early exploration to final marketing of drug research and development, cover multiple aspects of data such as basic characteristics, pharmacological action, toxicological reaction and clinical trial effect of the drug.

[0021] Among them, the research and development data of the target drug can be obtained according to the research and development department records. The drug research and development team will record various experimental data in detail during the experiment; for example, in the drug discovery stage, researchers will screen a large number of compounds, record the activity, selectivity and other indicators of each compound, and usually save them in the laboratory's electronic database or paper experimental report. Through Python, JAVA and other programming languages, the HTTP protocol is connected to the preset electronic database, and the HTTP request is sent to obtain the research and development data in the electronic database.

[0022] In detail, drug research and development is a long and complex process, which can be generally divided into different stages such as drug discovery, preclinical research, clinical trials (I, II, III), post-marketing research (IV), etc. Precise separation of data corresponding to each specific stage from the overall research and development data, which covers experimental design, operation process, observation results, analysis conclusions, etc. in this stage, can clearly present the specific situation and progress of the target drug in each research and development link.

[0023] In the field of drug research and development, drugs with similar structures, mechanisms of action, indications, etc. to the target drug are found, and these similar drugs are screened from a large amount of drug research and development data through drug feature matching, and their data corresponding to the target drug in the research and development process are obtained. In order to carry out comparative analysis and provide reference and reference for the research and development of the target drug.

[0024] Specifically, the extraction of stage research and development data of the target drug in different research and development stages includes data cleaning, data standardization, stage identification and division, etc. The data cleaning refers to that the research and development data may contain incorrect, missing or repeated information, for example, there may be incorrect data values in the experimental records, or some data is missing due to equipment failure; through data cleaning, incorrect data can be corrected, missing values can be filled (such as using mean, median filling or model-based prediction filling), and repeated data can be deleted, to ensure the accuracy and integrity of the data.

[0025] Among them, research and development data from different sources may have different formats and units; for example, some drug concentration data is expressed in moles per liter (mol / L), and some is expressed in milligrams per milliliter (mg / mL). Convert the data to a standard format and unit for subsequent processing and analysis, and standardize the text data, such as unified terminology, and unify "heart disease" and "heart disease" into one expression.

[0026] Further, according to the general process and industry standards of drug research and development, the starting and ending marks of each research and development stage are determined, the drug discovery stage usually starts from the screening of a large number of compounds and ends with the determination of a lead compound with potential activity; the preclinical research stage starts from the optimization of the lead compound and ends with the completion of a series of animal experiments and the acquisition of sufficient safety data; after data cleaning and standardization, the corresponding stage identifier is added to each research and development data in the research and development data set; the research and development stage field can be added in the database, or the stage prefix can be added to the data file to realize the stage research and development data.

[0027] If the research and development data is stored in the database, the structured query language (SQL) can be used to query and extract according to the stage identifier; for the research and development data stored in the file system, a script can be written by using a programming language (such as Python) to filter the files to obtain the stage research and development data of different research and development stages.

[0028] Specifically, the similar research and development data of the matched similar drugs is screened based on the similarity of the drug structure, the drug molecule is represented as a molecular fingerprint, and the molecular fingerprint is a method of encoding molecular structure information into a binary vector, wherein the commonly used molecular fingerprints include MACCS fingerprint, ECFP fingerprint and the like.

[0029] Further, by calculating the similarity (such as Tanimoto coefficient) between the molecular fingerprints, drugs with similar structures to the target drug can be screened, or drugs with the same or similar target points as the target drug can be found, the target point information of the drug is obtained through a drug database (such as DrugBank, ChEMBL), and then matched and screened, for the research and development data of the extracted similar drugs, correlation analysis is performed to find similar parts of the target drug research and development data, and a data mining technology (such as association rule mining) is used to discover potential relationships between data, so as to obtain similar research and development data.

[0030] In the embodiment of the application, the target drug research and development data can be quickly and comprehensively collected through efficient data acquisition tools, the information barriers can be broken, and the stage research and development data of different research and development stages can be accurately extracted to ensure the accuracy and consistency of the data, and provide reliable basis for the research and development process monitoring; at the same time, the machine learning algorithm is used for similar drug screening and research and development data matching, potential correlations can be mined, successful experience can be learned from, the research and development process can be accelerated, the research and development cost and risk can be reduced, and the efficiency and success rate of drug research and development can be improved.

[0031] S2, the stage research and development data and the similar research and development data are collected as multi-source clinical data, and the structured clinical data and the unstructured clinical data of the multi-source clinical data in different clinical stages are extracted according to a preset stage division standard.

[0032] In the embodiment of the present application, the stage research and development data and the similar research and development data of different sources, formats and structures can be integrated together by establishing a unified data warehouse to solve the problems of data dispersion and heterogeneity.

[0033] Specifically, a rule-based fusion algorithm is adopted to merge the data of different data sources according to preset business rules; clustering analysis technology in machine learning can be used to classify data with similar features into one category to realize deep fusion of data; in addition, data correlation technology can mine the potential relationship between different data and construct the correlation network between data, thereby forming complete, accurate and valuable multi-source clinical data to provide strong support for drug research and development decision-making.

[0034] Further, the multi-source clinical data includes structured data (such as laboratory examination result values) and unstructured data (such as doctor's handwritten medical record text and pathological section images); the data sources can be determined through system research, and standardized interfaces can be developed for different data sources to obtain related clinical data.

[0035] For example, in the medical health scene, the medical insurance service quality improvement project aims to break the data island, help medical institutions, pharmacies, clinics and other institutions improve service quality and improve medical efficiency, and provide more convenient and efficient medical insurance services; big data technology and machine learning algorithms are used to analyze and mine data in multiple dimensions, and multiple core capabilities such as heat distribution, site recommendation, price comparison, and false medical identification are built to improve the accuracy of medical insurance policy propaganda and the satisfaction of community residents with medical insurance services.

[0036] For example, in the financial technology scene, multi-source acquisition and risk assessment of financial data, financial institutions need to acquire multi-source data to assess the credit risk of customers and provide personalized financial services, which can be achieved by government-issued economic data, various financial websites and databases, API interfaces (such as Yahoo Finance, Alpha Vantage, etc.), network crawler information, customer data center, etc.; the feature data of the customer is input into the model to evaluate the credit risk level, and according to the risk assessment result, the customer is provided with a personalized financial service scheme such as loan amount, interest rate, etc.

[0037] In the embodiment of the present application, the structured clinical data and unstructured clinical data of the multi-source clinical data in different clinical stages are extracted according to the preset stage division standard, which includes: determining the stage clinical data of different clinical stages according to the preset stage division standard; performing field standardization processing on the structured modality data in the stage clinical data to obtain structured clinical data; Semantic vectorization embedding is performed on the unstructured modal data in the stage clinical data to generate unstructured clinical data.

[0038] In detail, according to a preset stage division standard, different clinical stage data are obtained by data analysis technology, such as text format of analyzed electronic medical records and special format of medical images, so as to understand the overall data; then, by means of data classification algorithm, the data are preliminarily classified into corresponding clinical stages according to key features in the preset standard, such as time nodes and treatment events.

[0039] Among them, for structured data, database query technology is used to accurately filter and extract from the database according to stage characteristics; for unstructured data, natural language processing technology is used to identify text information related to each clinical stage through keyword extraction and semantic analysis; at the same time, according to data checking technology, the integrity and accuracy of the extracted data are checked to ensure that the stage clinical data are true and reliable, and to provide a solid data foundation for subsequent drug research and development analysis.

[0040] Further, for structured data, it is checked whether the data contain fixed fields, separators (such as commas and tabs) or predefined data patterns (such as date formats and numerical ranges); for unstructured data, natural language processing technology is used to detect text continuity, semantic complexity or binary features of images / audio.

[0041] Among them, for composite data containing both structured fields and unstructured text (such as tables and free text in electronic medical records), a rule engine (such as regular expression matching table boundaries) or a machine learning model (such as BERT text classification) is used to split the data into independent modules and label them respectively.

[0042] Further, the same type of fields from different sources are mapped to a unified term set, and a rule engine (such as Drools) is configured to verify whether key fields (such as patient ID and examination date) are missing, and to fill in the missing fields using historical data or statistical values (such as median and mode) of the same type of patients.

[0043] Specifically, a BERT variant dedicated to the medical field is used to extract text semantic features, and a medical ontology library (such as SNOMED CT) is used to identify key entities (such as disease names and drug doses) in the text and encode them as structured labels, thereby obtaining unstructured clinical data.

[0044] In the embodiments of the present application, efficient integration and deep utilization of clinical data are realized at the computer level, which enables rapid integration of massive multi-source data, breaks down data silos, and through phased analysis, can capture real-time changes in data at each stage and use algorithm models to mine potential risk patterns.

[0045] For example, by comparing the fluctuations of structured data indicators with unstructured text feedback at different stages of drug development, emerging risk factors can be identified in a timely manner. This dynamic and phased approach enables the computer system to adapt to changes in the development process and provide early warning of risks, thereby providing more reliable safety assurance for drug development.

[0046] In the embodiments of the present application, for structured data, field standardization processing eliminates semantic ambiguity between heterogeneous systems through unified term coding, unit conversion and format specification; for unstructured data, semantic vectorization embedding technology uses a pre-trained model in the medical field to convert text, image and other data into high-dimensional semantic vectors, while preserving clinical context information and improving data analysis capabilities.

[0047] S3, multi-dimensional feature coding of the structured clinical data is performed to obtain structure coding features, and entity recognition and relationship extraction are performed on the unstructured clinical data to obtain non-structured entity features.

[0048] In the embodiments of the present application, the numerical, categorical and time-series features of the structured clinical data are discretized, vector embedded and hierarchically convoluted for feature extraction, respectively, and finally spliced and fused to generate structure coding features.

[0049] In the embodiments of the present application, the multi-dimensional feature coding of the structured clinical data to obtain structure coding features comprises: Data feature classification is performed on the structured clinical data to obtain numerical, categorical and time-series features; The numerical features are discretized and the discretized numerical features are standardized and coded to obtain numerical coding features; The categorical features are category vector embedded to obtain category coding features; A convolution kernel of a convolutional neural network is obtained, and the convolution kernel is used to perform hierarchical feature extraction on the time-series features to obtain time-series coding features; The numerical coding features, the category coding features and the time-series coding features are spliced to obtain structure coding features.

[0050] In detail, by analyzing the field attributes of structured data, the structured data is divided into three basic feature types: numerical feature recognition, which detects whether the field contains continuous or discrete numerical values, and distinguishes between integer type (such as patient ID) and measurement type numerical values through data distribution analysis (such as whether there is a decimal point, numerical range); categorical feature recognition, which determines whether the field is a set of limited discrete values, and assists in judgment through statistical unique value quantity and data proportion threshold; time-series feature recognition, which identifies fields with time dependence, matches date format or time interval description through regular expression, and confirms the time sequence relationship combined with business logic.

[0051] Further, the continuous numerical values are converted into discrete intervals and unified scales, the skewness of numerical distribution is solved, the numerical range is divided into intervals with fixed width, it is ensured that each interval contains the same number of samples, the numerical values are linearly mapped to the interval [0, 1], and the data is converted based on the mean and standard deviation.

[0052] Specifically, the discrete categories are converted into continuous vectors, the dimension disaster problem of high base number categories is solved, a binary vector is created for each category, the category label is replaced with a numerical statistic corresponding to the category, and the category is mapped to a low-dimensional dense vector by using a word vector model trained based on a medical knowledge graph or a clinical text corpus.

[0053] In detail, the local and global patterns of time series data are captured by a convolutional neural network (CNN), local features are extracted by sliding on the time axis, and different length convolution kernels (such as 3, 5, and 7 days) are used in parallel to capture patterns of different time spans, so that the numerical, category, and time series encoding features are spliced along the feature dimension to form a comprehensive feature representation, so as to ensure that the lengths of the feature vectors are consistent (for example, the numerical encoding is 10-dimensional, the category encoding is 32-dimensional, and the time series encoding is 64-dimensional, and after splicing, it is 106-dimensional), wherein the weights can be manually assigned according to the feature importance.

[0054] In the embodiment of the application, the multi-dimensional feature encoding technology not only retains the clinical key threshold information, but also eliminates the dimensional difference, and the category type feature is converted into a dense representation by embedding vector technology, solving the dimension disaster problem caused by high base number categories.

[0055] S4, multi-head attention feature fusion is performed on the structure coding features and the non-structure entity features to obtain a target data feature space.

[0056] In the embodiment of the application, the structure coding features and the non-structure entity features are fused by linear transformation, attention matrix construction, and multi-head attention mechanism, and finally spliced to generate a target data feature space.

[0057] In the embodiment of the application, the multi-head attention feature fusion on the structure coding features and the non-structure entity features to obtain a target data feature space comprises: The structure coding features and the non-structure entity features are respectively linearly transformed to obtain a first feature vector group and a second feature vector group; The first attention matrix and the second attention matrix of the first feature vector group and the second feature vector group are respectively constructed; According to the first attention matrix and the second attention matrix, the first attention features and the second attention features of different attention heads are calculated by using a pre-constructed multi-head attention mechanism; The first attention feature and the second attention feature are spliced to generate a target data feature space.

[0058] In detail, the structured and unstructured features are spatially converted by a fully connected neural network to make the features of different modalities comparable. The feature vector of the structured data after multi-dimensional encoding (such as the spliced results of numerical, categorical, and time sequence features) is input into a fully connected layer, and is linearly combined by a weight matrix and a bias term. For example, if the original structure feature dimension is 128, it can be projected to a 256-dimensional space by linear transformation to generate a first feature vector group.

[0059] Among them, the vector representation of unstructured data (such as text entities and image region features) is processed similarly; for example, the disease entity vector (such as a 300-dimensional word vector of “diabetes”) extracted from the text is projected to the same 256-dimensional space as the structure feature by a fully connected layer to generate a second feature vector group.

[0060] Further, for any two vectors (such as vector A and vector B) in the first feature vector group, a similarity score is calculated to finally generate a square matrix (such as 256x256), where each element represents the association strength of the corresponding feature pair.

[0061] Specifically, the first / second attention matrix is divided into multiple subspaces (such as 8 heads) along the feature dimension, and the attention weight is calculated independently for each head; for example, the 256-dimensional feature vector group is divided into 8 32-dimensional sub-vector groups corresponding to 8 attention heads; the Softmax function is applied to the attention matrix of each subspace to normalize the score to a probability distribution (sum = 1) representing the attention degree of each feature pair under the head, and the original feature vector is weighted and summed according to the attention weight. The aggregation results of all heads are spliced along the feature dimension (such as 8 32-dimensional vectors spliced into a 256-dimensional vector) to obtain the first attention feature and the second attention feature.

[0062] In detail, the attention representations of structured and unstructured features are combined to generate the final target feature space, ensuring that the first attention feature and the second attention feature have consistent dimensions (such as both being 256-dimensional), and if they are inconsistent, they are adjusted by a fully connected layer. The two attention feature vectors are directly spliced along the feature dimension (such as 256-dimensional + 256-dimensional = 512-dimensional) to form the fused target feature space.

[0063] In the embodiment of the present application, through the multi-head attention feature fusion technology, efficient integration of structured and unstructured clinical data is realized at the computer level, the feature representation capability is significantly improved, and the interference of modal difference on subsequent calculation is eliminated; the correlation strength between features is dynamically modeled by using the attention matrix, compared with the traditional fixed weight fusion method, the clinical semantic correlation can be adaptively captured; the multi-head attention mechanism parallelly mines the interaction mode from multiple subspaces, effectively solving the information loss problem in heterogeneous data fusion.

[0064] S5, constructing a risk factor index of the multi-source clinical data according to the target data feature space.

[0065] In the embodiment of the present application, the key risk features are screened through feature importance analysis, and the risk factor index is generated from the target data feature space by combining time series smoothing and threshold mapping.

[0066] In the embodiment of the present application, the risk factor index of the multi-source clinical data according to the target data feature space comprises: performing feature importance analysis on the target data feature space to obtain a weight coefficient of each target data feature in the target data feature space; screening a key risk feature subset greater than a preset weight threshold according to the weight coefficient; determining an initial risk score according to the key risk feature subset, and performing time series smoothing processing on the initial risk score to obtain a target risk score; performing index mapping on the target risk score according to a preset risk threshold to obtain a risk factor index.

[0067] In detail, a random forest or a gradient boosting tree model is used to determine the importance of each feature by calculating the total decrease in impurity (such as Gini coefficient) brought by each feature when the decision tree node is split; the mutual information value of each feature and the risk label (such as whether a complication occurs) is calculated, and the greater the value, the stronger the correlation; a threshold is set according to domain experience or statistical distribution, for example, features with a weight higher than 1.5 times of the median of all features (such as weight>0.5) are defined as key features, or the top 20% of features are selected according to the weight.

[0068] Further, the value of each key feature is multiplied by its weight and then added, the feature value is divided into different risk levels according to the clinical guidelines and is assigned a value, the average value of the last N scores (such as the last 3 examinations) is calculated to eliminate the influence of single abnormal value, and the smoothing window size needs to be adjusted according to the data update frequency to avoid excessive smoothing that masks the real risk changes.

[0069] Specifically, continuous risk scores are converted into discrete risk levels or labels to facilitate clinical decision-making. A threshold that balances sensitivity and specificity is selected through ROC curve analysis. In addition to the overall risk level, it can also be mapped to specific risk factor indicators.

[0070] In an embodiment of the present invention, high efficiency and accuracy of risk analysis of multi-source clinical data are achieved. Feature importance analysis uses the built-in interpretation mechanism of machine learning models (such as random forests or neural networks) to automatically quantify the contribution of each feature to risk, avoiding the subjective bias of manual screening; secondly, key feature screening based on weight thresholds ensures the representativeness of feature subsets through dynamic optimization algorithms (such as cross-validation parameter adjustment), effectively eliminates data noise and captures risk trends, realizes full-process automation of risk analysis, and improves the efficiency of investment and research and development based on clinical data.

[0071] S6. Determine the probabilities of multiple target risk dimensions of the multi-source clinical data based on the risk factor indicators.

[0072] In the embodiment of the present invention, risk factor indicators are standardized, dimensionally mapped, and probability modeled, and finally, the probabilities of multiple target risk dimensions with weighted distribution are calculated.

[0073] like Figure 3 As shown, in an embodiment of the present invention, determining the probabilities of multiple target risk dimensions of the multi-source clinical data according to the risk factor indicators includes: Standardizing the risk factor indicators to obtain standardized risk indicators; Obtaining a risk dimension mapping matrix, and mapping the standardized risk indicators to a plurality of preset risk dimension spaces according to the risk dimension mapping matrix; Conduct probability distribution modeling on multiple mapped standardized risk indicators and construct probability density functions for multiple risk dimensions; Calculate the cumulative distribution probability value of each risk dimension according to the probability density function; The cumulative distribution probability values ​​are weighted to obtain multiple target risk dimension probabilities.

[0074] In detail, data normalization is used to eliminate the dimensional differences between different risk indicators to ensure the accuracy of subsequent mapping and modeling. The value of each risk indicator is linearly mapped to a fixed interval, and the transformation is performed based on the mean and standard deviation of the indicator so that the data mean is 0 and the standard deviation is 1.

[0075] Furthermore, risk dimensions are divided according to clinical domain knowledge, and the strength of association between each standardized indicator and the risk dimension is determined through expert experience or data-driven methods. The value of each standardized indicator is multiplied by its association strength and then added to the corresponding risk dimension. The mapping matrix needs to be updated regularly to reflect the latest clinical guidelines, and it is necessary to ensure that the sum of the association strengths of all indicators is 1 (if multi-dimensional mapping is allowed) or equal to 1 (if single-dimensional mapping).

[0076] Among them, the appropriate probability distribution model is selected according to the data characteristics, and the distribution parameters are determined by maximum likelihood estimation or moment estimation method. For small sample data, non-parametric methods such as kernel density estimation can be used to avoid distribution assumption bias; for multimodal distribution, multiple distribution models (such as Gaussian mixture model) can be mixed to improve fitting accuracy.

[0077] Specifically, the probability that a risk dimension exceeds a certain threshold is calculated by integrating the probability density function, the possibility of extreme risk is quantified, and the fitted probability density function is integrated to obtain the cumulative probability from negative infinity to the current value; and the final target risk dimension probability is generated by weighted integration of the cumulative probabilities of each dimension, reflecting the contribution of different dimensions to the overall risk.

[0078] In the embodiment of the present invention, efficient quantitative analysis of the probability of clinical risk dimensions is achieved, which significantly improves the intelligent level of risk assessment. The standardized processing uses an adaptive algorithm to automatically match the data distribution type (such as normal and lognormal), eliminating dimensional differences while retaining data characteristics.

[0079] S7. Use a preset risk analysis model to perform a probability fusion analysis on the target risk dimension probability to obtain the risk control value of the multi-source clinical data.

[0080] In an embodiment of the present invention, the probabilities of each risk dimension are weighted, fused, and normalized through a risk analysis model, and the comprehensive risk probability is converted into a quantifiable risk control value based on preset standards to achieve risk assessment of multi-source clinical data.

[0081] In an embodiment of the present invention, the method of performing a probability fusion analysis on the target risk dimension probability using a preset risk analysis model to obtain the risk control value of the multi-source clinical data includes: Determine the probability weight coefficient of the target risk dimension probability corresponding to each risk dimension using a preset risk analysis model; Performing probability normalization processing according to the probability weight coefficient and the target risk dimension probability to obtain a comprehensive risk probability; According to the preset risk level classification standard and the corresponding risk control value correspondence, the comprehensive risk probability is converted into a control value to obtain a risk control value.

[0082] In detail, a risk analysis model is constructed depending on an advanced machine learning algorithm, such as a neural network algorithm, by collecting a large amount of historical clinical data covering the performance of various risk dimensions in different cases and the final risk result, inputting the data into the neural network model for training, and continuously adjusting the internal parameters of the model to learn the complex relationship between each risk dimension and the final risk. After sufficient training, the model can automatically assign a reasonable probability weight coefficient to each risk dimension according to the input target risk dimension probability, and the coefficient reflects the importance of the risk dimension in the overall risk.

[0083] Further, the probability weight coefficient and the target risk dimension probability are weighted and summed and normalized. The weighted sum is obtained by multiplying each target risk dimension probability by its corresponding probability weight coefficient and then adding all the products to quantitatively summarize the contribution of different risk dimensions to the overall risk. The normalization process is to convert the result of the weighted sum into a unified probability range, usually between 0 and 1, to make the comprehensive risk probability comparable and interpretable, and to intuitively reflect the size of the overall risk.

[0084] Specifically, the risk level division standard, for example, divides the comprehensive risk probability into low, medium and high levels, and clearly defines the range of each level, and sets corresponding risk control values for each risk level, which can be determined according to actual business needs and risk tolerance. After obtaining the comprehensive risk probability, it is compared with the risk level division standard to determine its risk level, and then the corresponding risk control value is found according to the corresponding relationship, which is similar to a mapping operation, to convert the abstract comprehensive risk probability into a specific and operable risk control value, providing a clear basis for subsequent risk management and decision-making.

[0085] In the embodiments of the present application, the probability weight coefficient is determined by the preset risk analysis model, the correlation between each risk dimension and the overall risk is deeply mined by means of machine learning algorithm, the weight is accurately assigned, and the scientific nature of risk assessment is improved. The probability normalization process uses an efficient algorithm to normalize the weighted result to a unified range, enhancing data comparability. The comprehensive risk probability is converted into a quantifiable risk control value according to the preset standard, realizing the digitization and standardization of risk assessment, facilitating the rapid processing and analysis of multi-source clinical data by computers, and providing clear and accurate decision-making basis for subsequent risk management and control, greatly improving the efficiency and accuracy of medical risk assessment.

[0086] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0087] As Figure 4 shown, it is an embodiment of the present application to provide a kind of functional module diagram of drug research and development investment risk control value analysis device.

[0088] In the embodiment of the present disclosure, a drug research and development investment risk control value analysis device is provided, which corresponds to the above-mentioned drug research and development investment risk control value analysis method. As Figure 4 shown, the drug research and development investment risk control value analysis device 100 can be installed in an electronic device, according to the function realized, the drug research and development investment risk control value analysis device 100 includes research and development data acquisition module 101, structure data extraction module 102, data coding enhancement module 103, feature data fusion module 104, risk index determination module 105, risk probability calculation module 106 and risk probability fusion module 107. The detailed description of each functional module is as follows: The research and development data acquisition module is used for acquiring the research and development data of the target drug, extracting the stage research and development data of the target drug at different research and development stages from the research and development data, and matching the similar research and development data of similar drugs according to the research and development data; The structure data extraction module is used for collecting the stage research and development data and the similar research and development data into multi-source clinical data, and extracting structured clinical data and unstructured clinical data of the multi-source clinical data at different clinical stages according to a preset stage division standard; The data coding enhancement module is used for multi-dimensional feature coding of the structured clinical data to obtain structure coding features, and entity recognition and relationship extraction of the unstructured clinical data to obtain non-structure entity features; The feature data fusion module is used for multi-head attention feature fusion of the structure coding features and the non-structure entity features to obtain a target data feature space; The risk index determination module is used for constructing risk factor indexes of the multi-source clinical data according to the target data feature space; The risk probability calculation module is used for determining a plurality of target risk dimension probabilities of the multi-source clinical data according to the risk factor indexes; The risk probability fusion module is used for probability fusion analysis of the target risk dimension probabilities by using a preset risk analysis model to obtain a risk control value of the multi-source clinical data.

[0089] In an embodiment, the structure data extraction module 102, when extracting the structured clinical data and the unstructured clinical data of the multi-source clinical data at different clinical stages according to the preset stage division standard, is used for: determine stage clinical data of different clinical stages according to a preset stage division standard; perform field standardization processing on structured modality data in the stage clinical data to obtain structured clinical data; perform semantic vectorization embedding on unstructured modality data in the stage clinical data to generate unstructured clinical data.

[0090] In an embodiment, the data coding enhancement module 103, when performing multi-dimensional feature coding on the structured clinical data to obtain structure coding features, is configured to: perform data feature classification on the structured clinical data to obtain numerical features, categorical features, and time-series features; perform discretization processing on the numerical features, and perform standardization coding on the discretized numerical features to obtain numerical coding features; perform category vector embedding on the categorical features to obtain category coding features; obtain a convolution kernel of a convolutional neural network, and perform hierarchical feature extraction on the time-series features according to the convolution kernel to obtain time-series coding features; perform feature concatenation on the numerical coding features, the category coding features, and the time-series coding features to obtain structure coding features.

[0091] In an embodiment, the feature data fusion module 104, when performing multi-head attention feature fusion on the structure coding features and the non-structured entity features to obtain a target data feature space, is configured to: perform linear transformation on the structure coding features and the non-structured entity features respectively to obtain a first feature vector group and a second feature vector group; construct a first attention matrix and a second attention matrix of the first feature vector group and the second feature vector group respectively; according to the first attention matrix and the second attention matrix, calculate first attention features and second attention features of different attention heads by using a pre-constructed multi-head attention mechanism; perform feature concatenation on the first attention features and the second attention features to generate a target data feature space.

[0092] In an embodiment, the risk indicator determination module 105, when constructing a risk factor indicator of the multi-source clinical data according to the target data feature space, is configured to: perform feature importance analysis on the target data feature space to obtain a weight coefficient of each target data feature in the target data feature space; filter out a key risk feature subset greater than a preset weight threshold according to the weight coefficient; determine an initial risk score according to the key risk feature subset, perform time series smoothing on the initial risk score to obtain a target risk score; perform index mapping on the target risk score according to a preset risk threshold to obtain a risk factor index.

[0093] In an embodiment, the risk probability calculation module 106, when performing determination of a plurality of target risk dimension probabilities of the multi-source clinical data according to the risk factor index, is configured to: perform standardization processing on the risk factor index to obtain a standardized risk index; obtain a risk dimension mapping matrix, and map the standardized risk index into a plurality of preset risk dimension spaces according to the risk dimension mapping matrix; perform probability distribution modeling on the mapped plurality of standardized risk indexes to construct a plurality of risk dimension probability density functions; calculate a cumulative distribution probability value of each risk dimension according to the probability density function; perform weight distribution on the cumulative distribution probability value to obtain a plurality of target risk dimension probabilities.

[0094] In an embodiment, the risk probability fusion module 107, when performing probability fusion analysis on the target risk dimension probability by using a preset risk analysis model to obtain a risk control value of the multi-source clinical data, is configured to: determine a probability weight coefficient of each risk dimension corresponding to the target risk dimension probability by using a preset risk analysis model; perform probability normalization processing on the probability weight coefficient and the target risk dimension probability to obtain a comprehensive risk probability; perform control value conversion on the comprehensive risk probability according to a preset risk level division standard and a corresponding risk control value corresponding relationship to obtain a risk control value.

[0095] In the present application, the specific limitations of the drug research investment risk control value analysis device can be referred to the limitations of the drug research investment risk control value analysis method in the above, which will not be repeated here. Each module in the above drug research investment risk control value analysis device can be realized by software, hardware and their combinations. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to call and execute the operations corresponding to each module by the processor.

[0096] In one embodiment, a computer device is provided, which can be a server, and its internal structure diagram can be as shown in Figure 5As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the service side of a drug R&D investment risk control value analysis method.

[0097] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a method for analyzing the risk control value of drug R&D investment.

[0098] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed: Acquiring R&D data of a target drug, extracting R&D data of the target drug at different R&D stages from the R&D data, and matching similar R&D data of similar drugs based on the R&D data; Aggregating the stage R&D data and the similar R&D data into multi-source clinical data, and extracting structured clinical data and unstructured clinical data at different clinical stages from the multi-source clinical data according to preset stage division criteria; Performing multi-dimensional feature encoding on the structured clinical data to obtain structural encoding features, and performing entity recognition and relationship extraction on the unstructured clinical data to obtain unstructured entity features; Performing multi-head attention feature fusion on the structural coding features and the non-structural entity features to obtain a target data feature space; constructing risk factor indicators for the multi-source clinical data according to the target data feature space; determine a plurality of target risk dimension probabilities of the multi-source clinical data according to the risk factor indicators; perform probability fusion analysis on the target risk dimension probabilities by using a preset risk analysis model to obtain a risk control value of the multi-source clinical data.

[0099] In several embodiments provided in the present application, it should be understood that the disclosed devices and apparatuses can be implemented in other manners. For example, the above-described system embodiments are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, another division manner can be used.

[0100] In addition, each function module in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be realized in the form of hardware, or in the form of hardware plus software function modules.

[0101] Therefore, no matter from which point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present application is defined by the appended claims rather than the above description, and therefore all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be regarded as limiting the claims to which they relate.

[0102] In some embodiments of the present embodiment, a computer readable storage medium is provided, and a computer program is stored on the computer readable storage medium. The computer program is characterized in that when executed by a processor, the computer program implements the steps of the method described in the above embodiments.

[0103] The readable storage medium of the present application stores a computer program, and the computer program can realize the following when executed by a processor of an electronic device: obtain research and development data of a target drug, extract stage research and development data of the target drug at different research and development stages from the research and development data, and match similar research and development data of similar drugs according to the research and development data; aggregate the stage research and development data and the similar research and development data into multi-source clinical data, extract structured clinical data and unstructured clinical data of the multi-source clinical data at different clinical stages according to a preset stage division standard; perform multi-dimensional feature coding on the structured clinical data to obtain structured coding features, and perform entity recognition and relationship extraction on the unstructured clinical data to obtain unstructured entity features; perform multi-head attention feature fusion on the structured coding features and the unstructured entity features to obtain a target data feature space; construct a risk factor index of the multi-source clinical data according to the target data feature space; determine a plurality of target risk dimension probabilities of the multi-source clinical data according to the risk factor index; perform probability fusion analysis on the target risk dimension probabilities by using a preset risk analysis model to obtain a risk control value of the multi-source clinical data.

[0104] It should be noted that the functions or steps described above in relation to the computer-readable storage medium or the computer device can correspond to the relevant descriptions of the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0105] The computer-readable storage medium can also store at least one computer executable program / instruction, such as computer readable instructions. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory, etc. The computer-readable storage medium may, for example, include read-only memory (ROM), hard disk, flash memory, etc. For example, the non-transitory computer-readable storage medium can be connected to a computing device such as a computer, and then when the computing device runs the computer readable instructions stored on the computer readable storage medium, the various methods described above can be performed.

[0106] In addition, the computer device can also include (but not limited to) a data bus, an input / output (I / O) bus, a display, and an input / output device (for example, a keyboard, a mouse, a speaker, etc.), etc.

[0107] In one embodiment, the at least one computer executable instruction can also be compiled or composed into a software product / computer program product, wherein one or more computer executable instructions are executed by the processor to perform the steps of the various functions and / or methods described in the embodiments of the present technology.

[0108] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium and can include the processes of the above-mentioned embodiments when executed. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory.

[0109] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit, module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units or modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0110] In the embodiments provided by the present disclosure, it should be understood that the disclosed apparatus and method can also be implemented by other manners. The apparatus embodiments described above are merely illustrative, for example, the flowcharts and block diagrams in the drawings show the possible implementation architecture, function and operation of the apparatus, method and computer program product according to the embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, program segment or part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the block can occur in different order from that noted in the drawings. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for implementing the specified function or action, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0111] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

[0112] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for example introduction, and do not represent actual use.

Claims

1. A method of analyzing a drug development investment risk control value, characterized by, The method comprises: obtaining research and development data of a target drug, extracting stage research and development data of the target drug at different research and development stages from the research and development data, and matching similar research and development data of similar drugs according to the research and development data; pooled into multi-source clinical data, and structured clinical data and unstructured clinical data of the multi-source clinical data at different clinical stages are extracted according to a preset stage division standard; the structured clinical data is subjected to multi-dimensional feature coding to obtain structure coding features, and entity recognition and relationship extraction are performed on the unstructured clinical data to obtain unstructured entity features; the structure coding features and the unstructured entity features are subjected to multi-head attention feature fusion to obtain a target data feature space; risk factor indexes of the multi-source clinical data are constructed according to the target data feature space; a plurality of target risk dimension probabilities of the multi-source clinical data are determined according to the risk factor indexes; a risk control value of the multi-source clinical data is obtained by performing probability fusion analysis on the target risk dimension probabilities by using a preset risk analysis model.

2. The method of claim 1, wherein the drug development investment risk control value analysis is performed by a computer system. The multi-dimensional feature coding of the structured clinical data to obtain structure coding features comprises: data feature classification is performed on the structured clinical data to obtain numerical features, categorical features and time series features; the numerical features are subjected to discretization processing, and the discretized numerical features are subjected to standardization coding to obtain numerical coding features; category vector embedding is performed on the categorical features to obtain category coding features; a convolution kernel of a convolutional neural network is obtained, and hierarchical feature extraction is performed on the time series features according to the convolution kernel to obtain time series coding features; the numerical coding features, the category coding features and the time series coding features are spliced to obtain the structure coding features.

3. The method of claim 1, wherein the drug development investment risk control value analysis is performed by a computer system. The multi-head attention feature fusion of the structure coding features and the unstructured entity features to obtain a target data feature space comprises: linear transformation is performed on the structure coding features and the unstructured entity features respectively to obtain a first feature vector group and a second feature vector group; a first attention matrix and a second attention matrix of the first feature vector group and the second feature vector group are constructed respectively; first attention features and second attention features of different attention heads are calculated by using a pre-constructed multi-head attention mechanism according to the first attention matrix and the second attention matrix; the first attention features and the second attention features are spliced to generate a target data feature space.

4. The drug R&D investment risk control value analysis method according to claim 1, characterized in that: The construction of risk factor indexes of the multi-source clinical data according to the target data feature space comprises: feature importance analysis is performed on the target data feature space to obtain a weight coefficient of each target data feature in the target data feature space; a key risk feature subset greater than a preset weight threshold is screened according to the weight coefficient; an initial risk score is determined according to the key risk feature subset, and a target risk score is obtained by performing time series smoothing processing on the initial risk score; The target risk score is index-mapped according to a preset risk threshold to obtain a risk factor index.

5. The method of claim 1, wherein the drug development investment risk control value analysis is performed by a computer system. The determining of the plurality of target risk dimension probabilities of the multi-source clinical data according to the risk factor index comprises: ​ The risk factor index is standardized to obtain a standardized risk index; A risk dimension mapping matrix is obtained, and the standardized risk index is mapped into a plurality of preset risk dimension spaces according to the risk dimension mapping matrix; The plurality of standardized risk indexes after mapping are subjected to probability distribution modeling to construct a probability density function of the plurality of risk dimensions; The cumulative distribution probability value of each risk dimension is calculated according to the probability density function; The cumulative distribution probability value is subjected to weight distribution to obtain a plurality of target risk dimension probabilities.

6. The method of claim 1, wherein the drug development investment risk control value analysis is performed by a computer system. The plurality of target risk dimension probabilities are subjected to probability fusion analysis by using a preset risk analysis model to obtain a risk control value of the multi-source clinical data, comprising: The probability weight coefficient of the target risk dimension probability corresponding to each risk dimension is determined by using a preset risk analysis model; The comprehensive risk probability is obtained by performing probability normalization processing according to the probability weight coefficient and the target risk dimension probability; The comprehensive risk probability is subjected to control value conversion according to a preset risk level division standard and a corresponding risk control value corresponding relationship to obtain a risk control value.

7. The method of claim 1, wherein the drug development investment risk control value analysis is performed by a computer system. The structured clinical data and unstructured clinical data of the multi-source clinical data at different clinical stages are extracted according to a preset stage division standard, comprising: ​ The stage clinical data of different clinical stages are determined according to a preset stage division standard; The structured clinical data is obtained by performing field standardization processing on the structured modality data in the stage clinical data; The unstructured clinical data is generated by performing semantic vectorization embedding on the unstructured modality data in the stage clinical data.

8. A pharmaceutical research and development investment risk control value analysis device characterized by comprising: The device comprises: A research and development data acquisition module is configured to acquire research and development data of a target drug, extract stage research and development data of the target drug at different research and development stages from the research and development data, and match similar research and development data of similar drugs according to the research and development data; A structured data extraction module is configured to collect the stage research and development data and the similar research and development data into multi-source clinical data, and extract structured clinical data and unstructured clinical data of the multi-source clinical data at different clinical stages according to a preset stage division standard; A data coding enhancement module is configured to perform multi-dimensional feature coding on the structured clinical data to obtain structured coding features, and perform entity recognition and relationship extraction on the unstructured clinical data to obtain unstructured entity features; A feature data fusion module is configured to perform multi-head attention feature fusion on the structured coding features and the unstructured entity features to obtain a target data feature space; A risk index determination module is configured to construct a risk factor index of the multi-source clinical data according to the target data feature space; A risk probability calculation module is configured to determine a plurality of target risk dimension probabilities of the multi-source clinical data according to the risk factor index. A risk probability fusion module is configured to perform probability fusion analysis on the target risk dimension probability by using a preset risk analysis model to obtain a risk control value of the multi-source clinical data.

9. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the method for analyzing a risk control value of a drug research and development investment according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: The computer program is executed by the processor to implement the method for analyzing a risk control value of a drug research and development investment according to any one of claims 1 to 7.