Deep learning-based medicine compliance risk early warning and decision suggestion method and system
By adopting deep learning and a variety of machine learning algorithms in the medical compliance risk warning system, the shortcomings of the existing system in risk identification, assessment accuracy and early warning functions are solved, and more efficient and accurate medical compliance risk warning is achieved.
Patent Information
- Application Number
- CN202411275387.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing medical compliance risk warning system has shortcomings in risk identification, assessment accuracy, early warning functions and system optimization, and it is difficult to cope with complex compliance requirements and rapidly changing market environment.
A deep learning-based approach is adopted, combining natural language processing, computer vision and machine learning algorithms to realize multi-dimensional medical compliance risk scanning, supplier risk assessment and qualification assessment, enhance the forward-looking and dynamic nature of early warning functions, and optimize the system through feedback loops and incremental learning.
It has achieved more comprehensive risk identification, more accurate quantitative assessment and more forward-looking risk prediction, which has improved the efficiency and accuracy of medical compliance risk warning, and helped enterprises better cope with complex compliance environments.
Smart Images

Figure CN119940900A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence and big data analysis, and specifically relates to a pharmaceutical compliance risk warning and decision-making recommendation method and system based on deep learning. Background Art
[0002] In the pharmaceutical industry, compliance risk warning is an important part of corporate operations, involving multiple links such as drug sales, marketing, clinical trials, and supply chain management. Traditional compliance risk warning methods mainly rely on manual inspections and post-reviews, which are inefficient and inaccurate, and are difficult to cope with increasingly complex compliance requirements and rapidly changing market environments.
[0003] In recent years, with the development of artificial intelligence (AI) and big data technology, more and more companies have begun to try to use these technologies to improve the efficiency and accuracy of compliance risk warning. For example, natural language processing (NLP) technology can be used to analyze large amounts of text data and identify potential risk points; machine learning algorithms can quantify and score these risk points. However, the application of these technologies is still in its early stages, and there are many areas that need further improvement and optimization.
[0004] Although artificial intelligence and big data technologies have shown great potential in compliance risk early warning, existing solutions still have obvious shortcomings in the following aspects:
[0005] 1. Insufficient risk identification: The risk identification of existing systems mostly relies on predefined rules and models, which cannot fully cover all potential risk points, especially when facing emerging risks and complex scenarios.
[0006] 2. Poor evaluation accuracy: Although many systems have introduced machine learning algorithms, the accuracy and reliability of the models still need to be improved in practical applications. In particular, when processing multi-dimensional and unstructured data, the performance of existing algorithms is not satisfactory.
[0007] 3. Limited early warning function: Most systems focus on identifying current risks and lack the ability to predict and warn of potential future risks, making it impossible to provide companies with forward-looking risk prevention measures.
[0008] 4. Insufficient system optimization: The existing system lacks an effective feedback mechanism and dynamic learning capabilities, making it difficult to continuously optimize and upgrade based on actual usage and emerging risk types. Summary of the invention
[0009] In view of the deficiencies in the above-mentioned background technology, the present invention proposes a pharmaceutical compliance risk warning and decision-making recommendation method and system based on deep learning. By introducing advanced artificial intelligence and big data analysis technologies, it aims to achieve comprehensive risk identification, accurate quantitative assessment, and forward-looking risk prediction and warning, helping pharmaceutical companies to achieve pharmaceutical compliance risk warning more efficiently and accurately.
[0010] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0011] The present invention provides a pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning, which mainly includes the following steps:
[0012] S1 Medical compliance risk identification and assessment;
[0013] S1.1 Multi-dimensional medical compliance risk scanning;
[0014] Use natural language processing technology to analyze evidence chains and data records to identify potential medical compliance risk points; then use random forests and gradient boosting trees to quantify and score the identified medical compliance risk points;
[0015] S1.2 Generate a heat map of medical compliance risks;
[0016] S1.3 Compliance check of sensitive business data;
[0017] Use computer vision and natural language processing techniques to analyze sensitive business data and identify non-compliant content, while using anomaly detection algorithms to identify suspicious visit patterns and frequencies;
[0018] S2Supplier risk assessment;
[0019] Integrate internal and external data to build a comprehensive supplier risk scoring model, and use time series analysis to predict future risk trends of suppliers;
[0020] S3 supplier qualification assessment;
[0021] S3.1 Data integration and validation;
[0022] Obtain supplier information from third-party authoritative data sources in real time through API interfaces, and use data comparison algorithms to automatically verify the consistency between the qualification information provided by suppliers and official data;
[0023] S3.2 Dynamic monitoring and early warning;
[0024] Develop scheduled tasks to regularly obtain the latest supplier qualification information from third-party interfaces, and establish an early warning mechanism based on a rule engine, which will automatically trigger the early warning mechanism when it is detected that the supplier qualifications have changed or are about to expire.
[0025] Furthermore, the specific implementation process of step S1.1 is as follows:
[0026] S1.1.1 Data preprocessing;
[0027] Assume there are n features m samples x1,x2,...,x m , the input matrix in Target vector The superscript T indicates matrix transposition; missing values are processed by scaling the obtained eigenvalues to 0-1 or standard normal distribution, and category features are encoded using one-hot encoding or label encoding;
[0028] S1.1.2 Random Forest Algorithm;
[0029] The random forest consists of multiple decision trees. The specific training process of each decision tree is as follows:
[0030] a) Randomly select m samples with replacement from m samples;
[0031] b) Randomly select k features from n features;
[0032] c) Use these k features to build a decision tree;
[0033] d) Repeat steps ac to construct T decision trees;
[0034] Use information gain or Gini coefficient to divide decision tree nodes;
[0035] S1.1.3 Gradient boosting tree algorithm;
[0036] The gradient boosting tree constructs a set of weak learners in an iterative manner, initializing F0(x) = \arg\min{γ}\sum_{i=1}^m L(y i ,γ), F0(x) represents the initial model, arg represents the extreme value operation on γ, i.e., the minimum value min, γ represents the constant prediction value, y i represents the target vector, L(y i ,γ) represents the loss function, which measures the difference between the predicted γ and the true target y i The gap between them; for the number of iterations m = 1toM, we have:
[0037] a) Calculate negative gradient
[0038] i=1,...,n;F(x i ) represents the input x i The predicted value of Indicates that in the mth iteration, the model is the previous iteration Result;
[0039] b) Fit a regression tree h m (x) to the target negative gradient
[0040] c) Calculate the multiplier γ m :
[0041] Indicates that in the m-1th iteration, the input x i The predicted value, h m (x i ) indicates the effect of the newly added weak learner (such as regression tree) on x in the mth round. i The predicted value of .
[0042] d) Update the model:
[0043] F m (x) represents the final model after the mth iteration, represents the model at the m-1th iteration, h m (x) represents the regression tree, η represents the learning rate;
[0044] The prediction formula of the gradient boosting tree is: GBT\Score=FM(x)=F0(x)+\sum{i=1}^Mηγ i h i (x), M represents the total number of iterations, γ i represents the coefficient of the regression tree learned in the i-th iteration, h i (x) represents the prediction result of the i-th decision tree;
[0045] S1.1.4 The final risk score is the weighted average of the random forest prediction result RF_Score and the gradient boosting tree prediction result GBT_Score;
[0046] S1.1.5 Feature importance calculation;
[0047] For random forests, the importance of feature j can be calculated by the average impurity reduction, and its specific calculation formula is as follows: Imp(j) = (1 / T)\sum_{i=1}^T\sum_{k\in splitnodes(i,j)}ΔI(S k ,j),ΔI(S k ,j) indicates that the impurity of feature j on node k is reduced, S krepresents the data set of the kth node, splitnodes(i,j) represents the set of nodes in the i-th tree that are split by feature j; for gradient boosted trees, feature importance can be measured by the number of times a feature is selected as a split feature in all trees;
[0048] S1.1.6 Use cross-validation to evaluate model performance;
[0049] S1.1.7 Use grid search or random search to optimize model hyperparameters;
[0050] S1.1.8 Define the risk level based on the final risk score Risk_Score:
[0051] Risk_Score ranges from 0-20: low risk;
[0052] Risk_Score ranges from 21-50: medium risk;
[0053] Risk_Score ranges from 51-80: high risk;
[0054] Risk_Score ranges from 81-100: Very high risk.
[0055] Furthermore, the specific implementation process of step S1.3 is as follows:
[0056] S1.3.1 Collect sales representatives’ visit records, visit photos and visit reports;
[0057] S1.3.2 Feature extraction;
[0058] a) Visit pattern characteristics: visit frequency, visit time distribution, visit duration and visit location;
[0059] b) Image features: Use a pre-trained convolutional neural network to extract feature vectors of visit photos and identify key objects in visit photos;
[0060] c) Text features: Use pre-trained BERT or other pre-trained language models to extract text features of visit reports and identify keywords and phrases in visit reports;
[0061] d) Merge all features;
[0062] S1.3.3 Anomaly detection algorithm;
[0063] Initialize the anomaly detector, build a prediction model and perform model training. The model includes input layer, encoding layer, decoding layer and output layer. Set the dynamic risk threshold, and then start to identify suspicious access patterns and frequencies of sensitive business data to determine whether there are anomalies. If the model prediction score is greater than the dynamic risk threshold, it is determined that an anomaly exists and a corresponding report is generated.
[0064] Furthermore, the specific implementation process of step S2 is as follows:
[0065] S2.1 Establish the prediction model formula;
[0066] First, define the following features:
[0067] X_1: supplier's historical compliance record score;
[0068] X_2: the supplier’s financial health score;
[0069] X_3: Quality rating of products provided by suppliers;
[0070] X_4: the market reputation score of the supplier;
[0071] Then define the supplier's risk score Y as a linear combination of the above characteristics:
[0072] Y=\beta_0+\beta_1X_1+\beta_2X_2+\beta_3X_3+\beta_4X_4+\epsilon
[0073] Where \beta_0 represents the intercept term, \beta_1, \beta_2, \beta_3, \beta_4 represent the weight coefficients of each feature, and \epsilon represents the error term;
[0074] S2.2 Model prediction process;
[0075] S2.2.1 Data collection: Collect relevant data about suppliers from internal and external channels, including historical compliance records, financial health, product quality and market reputation;
[0076] S2.2.2 Data preprocessing: Clean and standardize the collected data, handle missing values and outliers, and merge data from different sources;
[0077] S2.2.3 Feature extraction: Extract feature values X_1, X_2, X_3, X_4 from the processed data;
[0078] S2.2.4 Model training: Use the training set data to fit the linear regression model and estimate the model parameters \beta_0,\beta_1,\beta_2,\beta_3,\beta_4 by the least squares method;
[0079] S2.2.5 Model evaluation: Use the test set data to evaluate the model performance and calculate the mean square error and coefficient of determination to measure the prediction accuracy of the model;
[0080] S2.2.6 Model application: Apply the trained model to the data of new suppliers to generate risk score Y;
[0081] S2.3 Obtain the risk score of the new supplier through the above process;
[0082] S2.4 Risk classification: Based on the above risk scores, new suppliers are classified into low risk, medium risk and high risk levels.
[0083] Furthermore, the specific implementation process of step S3.1 is as follows:
[0084] S3.1.1 Data acquisition;
[0085] Obtain data D_s provided by suppliers in real time through the API interface; obtain data D_o from third-party authoritative data sources in real time through the API interface;
[0086] S3.1.2 Data preprocessing;
[0087] Clean and standardize the acquired data to ensure the data format is consistent;
[0088] S3.1.3 Similarity calculation;
[0089] For each field i, calculate the similarity J(d_i,o_i) between the data d_i provided by the supplier and the official data o_i;
[0090] S3.1.4Consistency determination;
[0091] Set a similarity threshold \theta. For each field i, if the similarity J(d_i,o_i) is greater than the threshold \theta, d_i and o_i are considered consistent and marked as consistent; otherwise, d_i and o_i are considered inconsistent and marked as inconsistent.
[0092] Furthermore, the specific implementation process of step S3.2 is as follows:
[0093] S3.2.1 Scheduled tasks;
[0094] Regularly obtain the latest supplier qualification information from third-party interfaces;
[0095] S3.2.2 Establish an early warning mechanism based on a rule engine;
[0096] (1) Rule 1: Notification of change of qualifications;
[0097] Rule description: If the supplier's qualification information changes, a change notification will be triggered;
[0098] Application scenario: When any field in the latest supplier qualification information obtained by the scheduled task is inconsistent with the existing data in the system;
[0099] Trigger condition: compare the latest supplier qualification information with the existing data in the system, and trigger if any inconsistency is found;
[0100] Notification content: indicate the specific fields that have been changed and their old and new values, and notify relevant personnel to confirm and process;
[0101] (2) Rule 2: Qualification expiration warning notice;
[0102] Rule description: If the supplier's qualification is about to expire, an expiration warning notification will be triggered;
[0103] Application scenario: In the latest supplier qualification information obtained by the scheduled task, the qualification expiration date field shows that it will expire within the next 30 days;
[0104] Trigger condition: The difference between the current date and the qualification expiration date is less than or equal to 30 days;
[0105] Notification content: indicate the supplier name, qualification expiration date and remaining days, and remind relevant personnel to renew qualifications or other processing;
[0106] S3.2.3 When it is detected that the supplier's qualifications have changed or are about to expire, the early warning mechanism will be automatically triggered;
[0107] (1) Data comparison;
[0108] The system compares the latest supplier qualification information with the existing data in the system;
[0109] (2) Application of rules;
[0110] Rule 1: Qualification change notification: Check the latest supplier qualification information with each field in the existing data. If any inconsistency is found, a change notification is triggered;
[0111] Rule 2: Qualification expiration warning notification: Check the qualification expiration date field in the latest supplier qualification information. If the expiration date is within the next 30 days, trigger the expiration warning notification;
[0112] (3) Notification sending;
[0113] Based on the output of the rule engine, the system notifies relevant personnel via email, SMS or system message.
[0114] Furthermore, it also includes step S4 risk identification and prediction in work tasks: based on historical task data, market data and risk data, identify the main risks that sales representatives may face in the process of completing tasks; at the same time, use the prediction model to quantitatively evaluate the risks that may arise in future work tasks, and reserve response measures in the plan.
[0115] Furthermore, the specific implementation process of step S4 is as follows:
[0116] S4.1 Data collection and preprocessing;
[0117] Collect historical data, market data and risk data, process missing values and outliers in the data, and ensure data quality;
[0118] S4.2 Feature extraction and selection;
[0119] Extract features related to task completion, and then use feature selection methods to screen important features;
[0120] S4.3 Model training and evaluation;
[0121] Divide the data set into a training set and a test set, train the linear regression prediction model, calculate the mean square error and determination coefficient, minimize the loss function, use the test set to evaluate the performance of the linear regression prediction model, and adjust the linear regression prediction model parameters to improve the prediction accuracy;
[0122] The formula of the linear regression prediction model is as follows:
[0123] y=\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_nx_n
[0124] Among them, y represents the predicted risk value, x_1, x_2, \dots, x_n represent the factors affecting the risk, and \beta_0, \beta_1, \dots, \beta_n represent the model parameters.
[0125] The present invention provides a pharmaceutical compliance risk early warning and decision-making suggestion system based on deep learning, which is used to implement the pharmaceutical compliance risk early warning and decision-making suggestion method based on deep learning. The system includes:
[0126] Pharmaceutical compliance risk identification and assessment module, used to implement multi-dimensional pharmaceutical compliance risk scanning, generate pharmaceutical compliance risk heat maps, and implement compliance checks on sensitive business data;
[0127] A supplier risk assessment module, which is used to establish a supplier risk assessment prediction model, to implement the prediction process of the supplier risk assessment prediction model, to output the prediction results of the supplier risk assessment prediction model, and to analyze and apply the prediction results of the supplier risk assessment prediction model;
[0128] The supplier qualification evaluation module is used to realize data integration and verification in the supplier qualification evaluation process, as well as to realize dynamic monitoring and early warning of supplier qualifications;
[0129] The risk identification and prediction module in work tasks is used to realize data collection and preprocessing in work tasks, and to realize the extraction and selection of data features in work tasks, and to establish a risk identification and prediction model in work tasks, and to realize the training and evaluation process of the risk identification and prediction model in work tasks.
[0130] The present invention provides a deep learning-based pharmaceutical compliance risk warning and decision-making recommendation device, which mainly includes: a memory and a processor; the memory stores executable instructions, and the processor is configured to execute the executable instructions in the memory to implement the steps of the deep learning-based pharmaceutical compliance risk warning and decision-making recommendation method.
[0131] The beneficial effects of the present invention are:
[0132] The present invention is a deep learning-based pharmaceutical compliance risk warning and decision-making recommendation method and system, which uses AI technology to achieve comprehensive risk scanning and identify explicit and implicit risks in all directions; accurately quantifies and scores risks through machine learning algorithms to achieve intelligent quantitative evaluation; it can not only identify current risks, but also predict possible problems in the future, and continuously improve the accuracy and practicality of the system through feedback loops and incremental learning. Compared with the prior art, the present invention has the following advantages:
[0133] 1. Improve the comprehensiveness and accuracy of pharmaceutical compliance risk identification;
[0134] Existing systems have limited capabilities in identifying emerging risks and complex scenarios, and are prone to missing potential risk points. This invention will introduce advanced deep learning and natural language processing technologies to improve the processing capabilities of multi-dimensional, unstructured data and comprehensively identify potential compliance risks.
[0135] 2. Improve the accuracy and reliability of risk assessment;
[0136] Due to the lack of diversity and quality of training data, the performance of existing systems in risk assessment is not ideal. This invention will improve the training effect of the model through multi-source data fusion and high-quality data set construction, and adopt complex algorithms and models to improve the accuracy and reliability of risk quantitative assessment.
[0137] 3. Enhance the foresight and dynamism of early warning function;
[0138] The warning function of the existing system relies on historical data, lacks foresight, and the warning mechanism is fixed and difficult to adjust dynamically. The present invention will introduce a prediction model and a dynamic adjustment mechanism to enhance the foresight and accuracy of the warning function and timely discover and warn of potential risks.
[0139] 4. Optimize the system’s feedback and learning mechanisms
[0140] The existing system lacks effective feedback mechanism and incremental learning ability, and it is difficult to optimize according to actual usage. The present invention will build a perfect feedback loop and incremental learning mechanism to ensure that the system can be continuously optimized and updated, and improve adaptability and response speed.
[0141] 5. The present invention significantly improves the efficiency and accuracy of pharmaceutical compliance risk warning, helping enterprises to better cope with complex compliance environments and ensure the legality and security of operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0142] Figure 1 This is a flow chart of a deep learning-based pharmaceutical compliance risk warning and decision-making recommendation method of the present invention.
[0143] Figure 2 This is a heat map of pharmaceutical compliance risks.
[0144] Figure 3 This is a structural block diagram of a deep learning-based pharmaceutical compliance risk warning and decision-making recommendation system of the present invention.
[0145] Figure 4 This is a block diagram of the structure of a deep learning-based pharmaceutical compliance risk warning and decision-making recommendation device of the present invention. DETAILED DESCRIPTION
[0146] In a first aspect, the present invention provides a pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning.
[0147] The present invention provides a pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning, such as Figure 1 As shown, the specific implementation steps are as follows:
[0148] S1 Medical compliance risk identification and assessment;
[0149] S1.1 Multi-dimensional medical compliance risk scanning;
[0150] Use natural language processing (NLP) technology to analyze evidence chains and data records to identify potential medical compliance risk points; then use machine learning algorithms such as random forests and gradient boosting trees to quantify and score the identified medical compliance risk points. The specific implementation process is as follows:
[0151] S1.1.1 Data preprocessing;
[0152] First, the raw data needs to be preprocessed. The raw data is mainly the front-line business data recorded by sales staff in real time. The system provides a convenient data entry interface for pharmaceutical sales staff. They can record the entire process of the visit in real time, including visit records (time, location, object), visit reports (key content), etc. These data are directly derived from the daily work of sales staff, truly reflect the development of front-line business, and are the key data basis for risk identification and assessment.
[0153] Assume there are n features m samples x1,x2,...,x m , the input matrix in Target vector The superscript T indicates matrix transposition. First, missing values are processed using methods such as mean filling and interpolation. Then, the obtained eigenvalues are scaled to 0-1 or standard normal distribution to achieve data standardization. Finally, one-hot encoding or label encoding is used to encode category features.
[0154] S1.1.2 Random Forest algorithm;
[0155] The random forest consists of multiple decision trees. The specific training process of each decision tree is as follows:
[0156] a) Randomly select m samples with replacement from m samples (bootstrap sampling);
[0157] b) Randomly select k features from n features;
[0158] c) Use these k features to build a decision tree;
[0159] d) Repeat steps ac to construct T decision trees.
[0160] Use information gain or Gini coefficient to divide decision tree nodes, where information gain IG(T,a) = H(T)-H(T|a), H(T) = -Σp(i|t)log2(p(i|t)); Gini coefficient: Gini(T) = 1-Σ(p(i|t)) 2 . Among them, i represents the category, t represents the node, a represents the characteristic attribute used to segment the data set at node t, H(T) represents the entropy value of the data set T, reflecting the degree of disorder of the data set, H(T|a) represents the conditional entropy of the data set T under the condition of feature a, and p(i|t) represents the conditional probability that the sample belongs to category i in node t.
[0161] The prediction formula of random forest is: Among them, h i (x) represents the prediction result of the i-th decision tree, and x represents the feature vector of the sample.
[0162] S1.1.3 Gradient Boosting Tree algorithm;
[0163] The gradient boosting tree constructs a set of weak learners (usually decision trees) in an iterative manner, initializing F0(x) = \arg\min{γ}\sum_{i=1}^m L(y i ,γ), γ represents a constant, its value needs to be determined in the initialization phase, y i represents the target vector, L(y i ,γ) represents the loss function, which measures the gap between the predicted value and the true value; for the number of iterations m = 1toM, we have:
[0164] a) Calculate negative gradient
[0165] F(x i ) means that in sample x i , the predicted value of the current model.
[0166] b) Fit a regression tree h m (x) to the target negative gradient
[0167] c) Calculate the multiplier γ m :
[0168] Indicates that in the last iteration, in sample x i The predicted value at h m (x i ) represents the regression tree generated in the mth iteration.
[0169] d) Update the model:
[0170] F m (x) represents the predicted value of the model at sample x after the mth iteration. Indicates the predicted value at sample x in the previous iteration (round m-1), h m (x) represents the regression tree, η represents the learning rate (usually 0.1 or 0.01)
[0171] The prediction formula of the gradient boosting tree is: GBT\Score=FM(x)=F0(x)+\sum{i=1}^Mηγ i h i (x), M represents the total number of iterations, γ i represents the reduction coefficient determined in the i-th iteration, h i (x) represents the prediction result of the i-th decision tree.
[0172] S1.1.4 Integrated scoring;
[0173] The final risk score is the weighted average of the random forest prediction result RF_Score and the gradient boosting tree prediction result GBT_Score. The specific calculation formula is as follows: Risk_Score = w1*RF_Score+w2*GBT_Score, where w1+w2=1, w1 and w2 represent the weight of the random forest prediction result and the weight of the gradient boosting tree prediction result, respectively. The weight can be determined based on the cross-validation results or business needs.
[0174] S1.1.5 Feature importance calculation;
[0175] For random forests, the importance of feature j can be calculated by the average impurity reduction, and its specific calculation formula is as follows: Imp(j) = (1 / T)\sum_{i=1}^T\sum_{k\in splitnodes(i,j)}ΔI(S k ,j), where ΔI(S k ,j) indicates the reduction of impurity of feature j on node k, k indicates the node number of the decision tree, S k represents the sample set reaching node K, splitnodes(i,j) represents the set of nodes that use feature j as the split feature in the i-th decision tree.
[0176] For gradient boosted trees, feature importance can be measured by the number of times the feature was selected as a split feature across all trees.
[0177] S1.1.6 Model evaluation;
[0178] Cross-validation is used to evaluate model performance. Commonly used indicators include:
[0179] Mean Squared Error (MSE): y i represents the true target value of the i-th sample, Represents the predicted value of the i-th sample (i.e., model output);
[0180] R 2 Fraction: Represents the mean of the true target values of all samples.
[0181] S1.1.7 Hyperparameter optimization;
[0182] Use grid search or random search to optimize model hyperparameters, such as tree depth, minimum number of leaf node samples, feature sampling ratio, etc.
[0183] S1.1.8 Risk level classification;
[0184] Based on the final risk score Risk_Score, the risk level can be defined:
[0185] Risk_Score ranges from 0-20: low risk;
[0186] Risk_Score ranges from 21-50: medium risk;
[0187] Risk_Score ranges from 51-80: high risk;
[0188] Risk_Score ranges from 81-100: Very high risk.
[0189] S1.2 Generation of medical compliance risk heat map:
[0190] Based on risk scores and impact range, a visual heat map of medical compliance risks is automatically generated, such as Figure 2 As shown: the horizontal axis represents different business link dimensions, and each small square represents a specific risk point; the vertical axis represents the impact of different risk types; different color gradients represent different risk levels, red represents extremely high risk, orange represents high risk, yellow represents medium risk, and green represents low risk. The darker the color, the higher the risk. Figure 2 It can be seen that in certain areas / links there is a large concentration of red / orange risk points, reflecting a systemic high-risk area.
[0191] S1.3 Compliance check of sensitive business data;
[0192] Apply computer vision technology and natural language processing (NLP) technology to analyze sensitive business data and identify non-compliant content, while using anomaly detection algorithms to identify suspicious visit patterns and frequencies. The specific implementation process is as follows:
[0193] S1.3.1 Data collection and preprocessing;
[0194] The following data were collected: sales representatives’ visit records (including time, location, duration, etc.), visit photos, and visit reports;
[0195] S1.3.2 Feature extraction;
[0196] a) Visit pattern characteristics: visit frequency (daily / weekly / monthly), visit time distribution, visit duration and visit location;
[0197] b) Image features (using computer vision technology): Use a pre-trained convolutional neural network to extract feature vectors of visit photos and identify key objects in visit photos, such as variety name, applicable department, etc.;
[0198] c) Text features (using natural language processing (NLP) technology): Use pre-trained BERT or other pre-trained language models to extract text features of visit reports and identify keywords and phrases in visit reports, such as drug names, hospital names, etc.
[0199] d) Merge all features.
[0200] S1.3.3 Anomaly detection algorithm;
[0201] S1.3.3.1 Data preparation;
[0202] Collect relevant data such as sales representatives’ visit records, visit photos, and visit reports, and perform necessary preprocessing, such as supplementing missing values, removing outliers, extracting features, etc., to organize the data into a format suitable for model input.
[0203] S1.3.3.2Build anomaly detection model;
[0204] Based on the preprocessed data, a machine learning model suitable for anomaly detection tasks is constructed.
[0205] S1.3.3.3 Model training;
[0206] The prepared data is divided into training set, validation set and test set. The model is trained with the training set so that the model can learn the difference between normal behavior patterns and abnormal behavior patterns.
[0207] S1.3.3.4 Determine anomaly score threshold;
[0208] The model performance corresponding to different anomaly score thresholds is tested on the validation set, and the optimal threshold that can achieve satisfactory precision and recall is selected. The anomaly score here is the model output, reflecting the probability that the sample is abnormal.
[0209] S1.3.3.5 Model evaluation and tuning;
[0210] The generalization ability of the model is fully evaluated on the reserved test set, and the model is tuned as necessary according to the evaluation indicators (such as F1 score, etc.) until the performance is satisfactory.
[0211] S1.3.3.6 Model application and report generation;
[0212] Apply the trained model to actual sales data, mark samples with anomaly scores higher than the set threshold as anomalies, and generate easy-to-read reports for these anomaly cases, including:
[0213] (1) Overview of the abnormality: specific manifestations of the abnormality, time of occurrence of the abnormality, etc.;
[0214] (2) Anomaly score: the anomaly score output by the model;
[0215] (3) Analysis of abnormal causes: main characteristics of the abnormality and ranking of their importance;
[0216] (4) Case comparison: the normal case that is most similar to this abnormal case;
[0217] (5) Other auxiliary information: such as the place of visit, duration, etc.
[0218] S2Supplier risk assessment;
[0219] Integrate internal and external data to build a comprehensive supplier risk scoring model, and use time series analysis to predict the supplier's future risk trends. The specific implementation process is as follows:
[0220] S2.1 Establish the prediction model formula;
[0221] First, define the following features:
[0222] X_1: supplier's historical compliance record score;
[0223] X_2: the supplier’s financial health score;
[0224] X_3: Quality rating of products provided by suppliers;
[0225] X_4: the market reputation score of the supplier;
[0226] Then define the supplier's risk score Y as a linear combination of the above characteristics, as follows:
[0227] Y=\beta_0+\beta_1X_1+\beta_2X_2+\beta_3X_3+\beta_4X_4+\epsilon
[0228] Among them, \beta_0 represents the intercept term, \beta_1, \beta_2, \beta_3, \beta_4 represent the weight coefficients of each feature respectively, and \epsilon represents the error term.
[0229] S2.2 Model prediction process;
[0230] S2.2.1 Data Collection: Obtain supplier information from third-party authoritative data sources in real time through API interfaces, such as data sources such as Tianyancha and the Industrial and Commercial Bureau that collect supplier-related data, including historical compliance records, financial health, product quality, and market reputation. Data can be obtained in real time through API interfaces or extracted from historical records.
[0231] S2.2.2 Data preprocessing: Clean and standardize the collected data, handle missing values and outliers, and merge data from different sources.
[0232] S2.2.3 Feature extraction: Extract feature values X_1, X_2, X_3, X_4 from the processed data. These features can be original data or secondary indicators obtained through certain calculations, thus forming the training set data and test set data.
[0233] S2.2.4 Model training: Use the training set data to fit the linear regression model and estimate the model parameters \beta_0, \beta_1, \beta_2, \beta_3, \beta_4 by the least squares method (OLS).
[0234] S2.2.5 Model evaluation: Use the test set data to evaluate the model performance and calculate the mean square error (MSE) and coefficient of determination (R 2 ) to measure the prediction accuracy of the model.
[0235] S2.2.6 Model application: Apply the trained model to the data of new suppliers to generate a risk score Y.
[0236] S2.3 prediction results;
[0237] Through the above process, the risk score of the supplier can be obtained. Assume that there is the following new supplier data:
[0238] X_1 = 85 (historical compliance record score);
[0239] X_2 = 90 (financial health score);
[0240] X_3 = 80 (product quality score);
[0241] X_4 = 75 (market reputation score);
[0242] Assume that after model training, the parameters obtained are: \beta_0 = 5, \beta_1 = 0.3, \beta_2 = 0.25, \beta_3 = 0.2, \beta_4 = 0.15, then the risk score Y of the supplier is:
[0243] Y=5+0.3\times 85+0.25\times 90+0.2\times 80+0.15\times 75
[0244] The specific calculation result is: Y=5+25.5+22.5+16+11.25=80.25.
[0245] The vendor's final risk score was 80.25, indicating a relatively medium risk level.
[0246] S2.4 Result analysis and application;
[0247] Risk classification: Based on the above risk scores, suppliers are divided into low risk, medium risk and high risk levels. For example, a score below 60 is low risk, 60-80 is medium risk, and above 80 is high risk.
[0248] S3 supplier qualification assessment;
[0249] S3.1 Data integration and validation;
[0250] The API interface is used to obtain supplier information from third-party authoritative data sources such as Tianyancha and the Industrial and Commercial Bureau in real time. At the same time, a data comparison algorithm is designed to automatically verify the consistency between the qualification information provided by the supplier and the official data. The specific implementation process is as follows:
[0251] S3.1.1 Data acquisition;
[0252] Obtain data D_s provided by suppliers in real time through the API interface; obtain data D_o from third-party authoritative data sources such as Tianyancha and Industrial and Commercial Bureau in real time through the API interface.
[0253] The obtained data structure is as follows:
[0254] Data provided by the supplier: D_s = \{d_1,d_2,\ldots,d_n\}, d_1,d_2,\ldots,d_n\ represents the data sequence provided by the supplier.
[0255] Data from official data sources: D_o = \{o_1,o_2,\ldots,o_n\}, o_1,o_2,\ldots,o_n\ represents the data sequence of official data sources.
[0256] S3.1.2 Data preprocessing;
[0257] The acquired data is cleaned and standardized to ensure consistent data formats. For text fields, they are converted to lowercase and special characters and spaces are removed; for numeric fields, the units and precision are standardized. For example, the company name, license number, and registration date do not need to be processed, and the address field is preprocessed, that is, "Street" is unified into "St.".
[0258] S3.1.3 Similarity calculation;
[0259] For each field i, the similarity J(d_i,o_i) between the data d_i provided by the supplier and the official data o_i is calculated. The present invention uses a simple Jaccard similarity coefficient to calculate the similarity of each field. For each field i, there is the following relationship: J(d_i,o_i) = \frac{|d_i\cap o_i|}{|d_i\cup o_i|}, where |d_i\cap o_i| represents the size of the intersection of d_i and o_i, and |d_i\cup o_i| represents the size of the union of d_i and o_i.
[0260] S3.1.4Consistency determination;
[0261] Set a similarity threshold \theta (for example, 0.8). For each field i, if the similarity J(d_i,o_i) is greater than the threshold \theta, d_i and o_i are considered consistent and marked as consistent; otherwise, d_i and o_i are considered inconsistent and marked as inconsistent.
[0262] S3.1.5 Summary of results;
[0263] Summarize the consistency determination results of all fields and generate a final consistency report.
[0264] Through the above method, enterprises can automatically verify the consistency between the qualification information provided by suppliers and official data to ensure the accuracy and reliability of the data.
[0265] S3.2 Dynamic monitoring and early warning;
[0266] Develop scheduled tasks to regularly obtain the latest supplier qualification information from third-party interfaces, and establish an early warning mechanism based on the rule engine. When it is detected that the supplier qualification has changed or is about to expire, the early warning mechanism will be automatically triggered. The specific implementation process is as follows:
[0267] S3.2.1 Scheduled tasks;
[0268] Regularly obtain the latest supplier qualification information from third-party interfaces.
[0269] S3.2.2 Establish an early warning mechanism based on a rule engine;
[0270] (1) Rule 1: Notification of change of qualifications;
[0271] Rule description: If the supplier's qualification information changes, a change notification will be triggered.
[0272] Application scenario: When any field (such as company name, license number, address, registration date, etc.) in the latest supplier qualification information obtained by the scheduled task is inconsistent with the existing data in the system.
[0273] Trigger condition: The latest supplier qualification information is compared with the existing data in the system, and the trigger is triggered if any inconsistency is found.
[0274] Notification content: Indicate the specific fields that have been changed and their old and new values, and notify relevant personnel for confirmation and processing.
[0275] (2) Rule 2: Qualification expiration warning notice;
[0276] Rule description: If the supplier's qualifications are about to expire (for example, within the next 30 days), an expiration warning notification is triggered.
[0277] Application scenario: In the latest supplier qualification information obtained by the scheduled task, the qualification expiration date field shows that it will expire within the next 30 days.
[0278] Trigger condition: The difference between the current date and the qualification expiration date is less than or equal to 30 days.
[0279] Notification content: Indicate the supplier name, qualification expiration date and remaining days, and remind relevant personnel to renew qualifications or other processing.
[0280] S3.2.3 When it is detected that the supplier's qualifications have changed or are about to expire, the early warning mechanism will be automatically triggered;
[0281] (1) Data comparison;
[0282] The system compares the latest supplier qualification information with the existing data in the system.
[0283] (2) Application of rules;
[0284] Rule 1: Qualification change notification: Check the latest supplier qualification information with each field in the existing data. If any inconsistency is found, a change notification is triggered.
[0285] Rule 2: Qualification expiration warning notification: Check the qualification expiration date field in the latest supplier qualification information. If the expiration date is within the next 30 days, an expiration warning notification is triggered.
[0286] (3) Notification sending;
[0287] Based on the output of the rule engine, the system notifies relevant personnel via email, SMS or system messages.
[0288] The following are some example scenarios for dynamic monitoring and early warning applications:
[0289] Scenario 1: Qualification change;
[0290] The supplier address recorded in the system is "123Pharma Street", and the address field in the newly acquired supplier qualification information is "456PharmaAvenue", triggering Rule 1 to send a change notification indicating the old and new values of the address field.
[0291] Scenario 2: Qualification expiration warning;
[0292] The supplier qualification expiration date recorded in the system is "2024-08-15", and the current date is "2024-07-16", triggering rule 2 to send an expiration warning notification to remind relevant personnel that their qualifications will expire within 30 days.
[0293] S4 Risk identification and prediction in work tasks;
[0294] Based on historical task data, market data, and risk data, identify the main risks that sales representatives may face in the process of completing tasks. For example, customer acquisition risks, price risks caused by flow, etc. At the same time, use the prediction model to quantitatively evaluate the risks that may arise in future work tasks and reserve countermeasures in the plan. The specific implementation process is as follows:
[0295] S4.1 Data collection and preprocessing;
[0296] Collect historical data, market data and risk data, including task completion status, market environment and risk events, etc., process missing values and outliers in the data, and ensure data quality.
[0297] S4.2 Feature extraction and selection;
[0298] Extract features related to task completion, such as sales, number of customers, market demand, etc., and then use feature selection methods (such as correlation analysis, principal component analysis) to screen important features.
[0299] S4.3 Model training and evaluation;
[0300] The data set is divided into a training set and a test set (e.g. 80% for training and 20% for testing), a linear regression prediction model is trained, and the mean square error (MSE) and coefficient of determination (R 2 ), minimize the loss function (such as mean square error MSE), use the test set to evaluate the performance of the linear regression prediction model, and adjust the parameters of the linear regression prediction model to improve the prediction accuracy.
[0301] Among them, the formula of the linear regression prediction model is as follows:
[0302] y=\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_nx_n
[0303] Among them, y represents the predicted risk value, x_1, x_2, \dots, x_n represent factors affecting the risk (such as market demand, customer visit rate, etc.), \beta_0, \beta_1, \dots, \beta_n represent model parameters.
[0304] The calculation formula of mean square error (MSE) is as follows:
[0305] \text{MSE}=\frac{1}{n}\sum_{i=1}^{n}(y_i-\hat{y}_i)^2
[0306] Among them, y_i represents the actual value, \hat{y}_i represents the predicted value, and n represents the number of samples.
[0307] Among them, the coefficient of determination (R 2 ) is calculated as follows:
[0308] R^2=1-\frac{\sum_{i=1}^{n}(y_i-\hat{y}_i)^2}{\sum_{i=1}^{n}(y_i-\bar{y})^2}
[0309] Among them, \bar{y} represents the average of the actual values.
[0310] In a second aspect, the present invention provides a pharmaceutical compliance risk warning and decision-making recommendation system based on deep learning, which is mainly used to implement a pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning provided in the first aspect.
[0311] A pharmaceutical compliance risk warning and decision-making recommendation system based on deep learning of the present invention, such as Figure 3 As shown, it mainly includes the following modules:
[0312] The pharmaceutical compliance risk identification and assessment module is mainly used to implement multi-dimensional pharmaceutical compliance risk scanning, to generate pharmaceutical compliance risk heat maps, and to implement compliance checks on sensitive business data.
[0313] The supplier risk assessment module is mainly used to establish a supplier risk assessment prediction model, to implement the prediction process of the supplier risk assessment prediction model, to output the prediction results of the supplier risk assessment prediction model, and to analyze and apply the prediction results of the supplier risk assessment prediction model.
[0314] The supplier qualification evaluation module is mainly used to realize data integration and verification in the supplier qualification evaluation process, as well as to realize dynamic monitoring and early warning of supplier qualifications.
[0315] The risk identification and prediction module in work tasks is mainly used to realize data collection and preprocessing in work tasks, to realize the extraction and selection of data features in work tasks, to establish risk identification and prediction models in work tasks, and to realize the training and evaluation process of risk identification and prediction models in work tasks.
[0316] In the third aspect, the present invention provides a pharmaceutical compliance risk warning and decision-making recommendation device based on deep learning, which is mainly used to run a pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning provided in the first aspect.
[0317] like Figure 4 As shown, a pharmaceutical compliance risk warning and decision-making recommendation device based on deep learning of the present invention mainly includes: a processor, a memory, an input device and an output device; wherein the number of processors can be one or more, Figure 4 Only one processor is used as an example; the processor, memory, input device and output device are all connected through a bus or other means. Figure 4 In the example, the connection through a bus is taken as an example; the memory, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules; the processor can execute various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory, that is, to realize a pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning provided in the first aspect; the input device can be used to receive input digital or character information, and generate key signal input related to the user settings and function control of the device; the output device may include display devices such as display screens.
[0318] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning, characterized in that: The following steps are involved: S1 Medical compliance risk identification and assessment; S1.1 Multi-dimensional medical compliance risk scanning; Use natural language processing technology to analyze evidence chains and data records to identify potential medical compliance risk points; Then, random forest and gradient boosting trees were used to quantitatively score the identified medical compliance risk points; S1.2 Generate a heat map of medical compliance risks; S1.3 Compliance check of sensitive business data; Use computer vision and natural language processing techniques to analyze sensitive business data and identify non-compliant content, while using anomaly detection algorithms to identify suspicious visit patterns and frequencies; S2Supplier risk assessment; Integrate internal and external data to build a comprehensive supplier risk scoring model, and use time series analysis to predict future risk trends of suppliers; S3 supplier qualification assessment; S3.1 Data integration and validation; Obtain supplier information from third-party authoritative data sources in real time through API interfaces, and use data comparison algorithms to automatically verify the consistency between the qualification information provided by suppliers and official data; S3.2 Dynamic monitoring and early warning; Develop scheduled tasks to regularly obtain the latest supplier qualification information from third-party interfaces, and establish an early warning mechanism based on a rule engine, which will automatically trigger the early warning mechanism when it is detected that the supplier qualifications have changed or are about to expire.
2. The pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning according to claim 1 is characterized in that: The specific implementation process of step S1.1 is as follows: S1.1.1 Data preprocessing; Assume there are n features m samples x1,x2,...,x m , input matrix X = [x1,x2,...,x m ] T ,in Target vector Y = [y1, y2, ..., y m ] T , the superscript T indicates the matrix transpose; Use missing value processing, scale the obtained feature values to 0-1 or standard normal distribution, and use one-hot encoding or label encoding to encode category features; S1.1.2 Random Forest Algorithm; The random forest consists of multiple decision trees. The specific training process of each decision tree is as follows: a) Randomly select m samples with replacement from m samples; b) Randomly select k features from n features; c) Use these k features to build a decision tree; d) Repeat steps ac to construct T decision trees; Use information gain or Gini coefficient to divide decision tree nodes; S1.1.3 Gradient boosting tree algorithm; The gradient boosting tree constructs a set of weak learners in an iterative manner, initializing F0(x) = \arg\min{γ}\sum_{i=1}^m L(yi,γ), where F0(x) represents the initial model, arg represents the extreme value operation on γ, that is, the minimum value min, γ represents the constant prediction value, yi represents the target vector, and L(yi,γ) represents the loss function, which measures the gap between the predicted γ and the true target yi; for the number of iterations m=1toM, we have: a) Calculate negative gradient i=1,...,n; F(xi) represents the predicted value of input xi, Indicates that in the mth iteration, the model is the previous iteration Results ; b) Fit a regression tree h m (x) to the target negative gradient c) Calculate the multiplier γ m : Indicates the input x at the m-1th iteration i The predicted value, h m (x i ) represents the predicted value of xi by the newly added weak learner (such as regression tree) in the mth round; d) Update the model: F m (x) represents the final model after the mth iteration, represents the model at the m-1th iteration, h m (x) represents the regression tree, η represents the learning rate; The prediction formula of the gradient boosting tree is: GBT\Score=FM(x)=F0(x)+\sum{i=1}^Mηγ i h i (x), M represents the total number of iterations, γ i represents the coefficient of the regression tree learned in the i-th iteration, h i (x) represents the prediction result of the i-th decision tree; S1.1.4 The final risk score is the weighted average of the random forest prediction result RF_Score and the gradient boosting tree prediction result GBT_Score; S1.1.5 Feature importance calculation; For random forests, the importance of feature j can be calculated by the average impurity reduction, and its specific calculation formula is as follows: Imp(j) = (1 / T)\sum_{i=1}^T\sum_{k\in splitnodes(i,j)}ΔI(S k ,j),ΔI(S k ,j) indicates that the impurity of feature j on node k is reduced, S k represents the data set of the kth node, splitnodes(i,j) represents the set of nodes in the i-th tree that are split by feature j; for gradient boosted trees, feature importance can be measured by the number of times a feature is selected as a split feature in all trees; S1.1.6 Use cross-validation to evaluate model performance; S1.1.7 Use grid search or random search to optimize model hyperparameters; S1.1.8 Define the risk level based on the final risk score Risk_Score: Risk_Score ranges from 0-20: low risk; Risk_Score ranges from 21-50: medium risk; Risk_Score ranges from 51-80: high risk; Risk_Score ranges from 81-100: Very high risk.
3. The pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning according to claim 1 is characterized in that: The specific implementation process of step S1.3 is as follows: S1.3.1 Collect sales representatives’ visit records, visit photos and visit reports; S1.3.2 Feature extraction; a) Visit pattern characteristics: visit frequency, visit time distribution, visit duration and visit location; b) Image features: Use a pre-trained convolutional neural network to extract feature vectors of visit photos and identify key objects in visit photos; c) Text features: Use pre-trained BERT or other pre-trained language models to extract text features of visit reports and identify keywords and phrases in visit reports; d) Merge all features; S1.3.3 Anomaly detection algorithm; Initialize the anomaly detector, build a prediction model and perform model training. The model includes input layer, encoding layer, decoding layer and output layer. Set the dynamic risk threshold, and then start to identify suspicious access patterns and frequencies of sensitive business data to determine whether there are anomalies. If the model prediction score is greater than the dynamic risk threshold, it is determined that an anomaly exists and a corresponding report is generated.
4. The pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning according to claim 1, characterized in that: The specific implementation process of step S2 is as follows: S2.1 Establish the prediction model formula; First, define the following features: X_1: supplier's historical compliance record score; X_2: the supplier’s financial health score; X_3: Quality rating of products provided by suppliers; X_4: the market reputation score of the supplier; Then define the supplier's risk score Y as a linear combination of the above characteristics: Y=\beta_0+\beta_1X_1+\beta_2X_2+\beta_3X_3+\beta_4X_4+\epsilon Where \beta_0 represents the intercept term, \beta_1, \beta_2, \beta_3, \beta_4 represent the weight coefficients of each feature, and \epsilon represents the error term; S2.2 Model prediction process; S2.2.1 Data collection: Collect relevant data about suppliers from internal and external channels, including historical compliance records, financial health, product quality and market reputation; S2.2.2 Data preprocessing: Clean and standardize the collected data, handle missing values and outliers, and merge data from different sources; S2.2.3 Feature extraction: Extract feature values X_1, X_2, X_3, X_4 from the processed data; S2.2.4 Model training: Use the training set data to fit the linear regression model and estimate the model parameters \beta_0,\beta_1,\beta_2,\beta_3,\beta_4 by the least squares method; S2.2.5 Model evaluation: Use the test set data to evaluate the model performance and calculate the mean square error and coefficient of determination to measure the prediction accuracy of the model; S2.2.6 Model application: Apply the trained model to the data of new suppliers to generate risk score Y; S2.3 Obtain the risk score of the new supplier through the above process; S2.4 Risk classification: Based on the above risk scores, new suppliers are classified into low risk, medium risk and high risk levels.
5. The pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning according to claim 1, characterized in that: The specific implementation process of step S3.1 is as follows: S3.1.1 Data acquisition; Obtain data D_s provided by suppliers in real time through the API interface; obtain data D_o from third-party authoritative data sources in real time through the API interface; S3.1.2 Data preprocessing; Clean and standardize the acquired data to ensure the data format is consistent; S3.1.3 Similarity calculation; For each field i, calculate the similarity J(d_i,o_i) between the data d_i provided by the supplier and the official data o_i; S3.1.4Consistency determination; Set a similarity threshold \theta. For each field i, if the similarity J(d_i,o_i) is greater than the threshold \theta, d_i and o_i are considered consistent and marked as consistent; otherwise, d_i and o_i are considered inconsistent and marked as inconsistent.
6. The pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning according to claim 1 is characterized in that: The specific implementation process of step S3.2 is as follows: S3.2.1 Scheduled tasks; Regularly obtain the latest supplier qualification information from third-party interfaces; S3.2.2 Establish an early warning mechanism based on a rule engine; (1) Rule 1: Notification of changes in qualifications; Rule description: If the supplier's qualification information changes, a change notification will be triggered; Application scenario: When any field in the latest supplier qualification information obtained by the scheduled task is inconsistent with the existing data in the system; Trigger condition: compare the latest supplier qualification information with the existing data in the system, and trigger if any inconsistency is found; Notification content: indicate the specific fields that have been changed and their old and new values, and notify relevant personnel to confirm and process; (2) Rule 2: Qualification expiration warning notice; Rule description: If the supplier's qualification is about to expire, an expiration warning notification will be triggered; Application scenario: In the latest supplier qualification information obtained by the scheduled task, the qualification expiration date field shows that it will expire within the next 30 days; Trigger condition: The difference between the current date and the qualification expiration date is less than or equal to 30 days; Notification content: indicate the supplier name, qualification expiration date and remaining days, and remind relevant personnel to renew qualifications or other processing; S3.2.3 When it is detected that the supplier's qualifications have changed or are about to expire, the early warning mechanism will be automatically triggered; (1) Data comparison; The system compares the latest supplier qualification information with the existing data in the system; (2) Application of rules; Rule 1: Qualification change notification: Check the latest supplier qualification information with each field in the existing data. If any inconsistency is found, a change notification is triggered; Rule 2: Qualification expiration warning notification: Check the qualification expiration date field in the latest supplier qualification information. If the expiration date is within the next 30 days, trigger the expiration warning notification; (3) Notification sending; Based on the output of the rule engine, the system notifies relevant personnel via email, SMS or system message.
7. The pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning according to claim 1 is characterized in that: It also includes step S4, risk identification and prediction in work tasks: based on historical task data, market data and risk data, identify the main risks that sales representatives may face in the process of completing tasks; at the same time, use predictive models to quantitatively evaluate the risks that may arise in future work tasks, and reserve response measures in the plan.
8. The pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning according to claim 7 is characterized in that: The specific implementation process of step S4 is as follows: S4.1 Data collection and preprocessing; Collect historical data, market data and risk data, process missing values and outliers in the data, and ensure data quality; S4.2 Feature extraction and selection; Extract features related to task completion, and then use feature selection methods to screen important features; S4.3 Model training and evaluation; Divide the data set into a training set and a test set, train the linear regression prediction model, calculate the mean square error and determination coefficient, minimize the loss function, use the test set to evaluate the performance of the linear regression prediction model, and adjust the linear regression prediction model parameters to improve the prediction accuracy; The formula of the linear regression prediction model is as follows: y=\beta_0+\beta_1x_1+\beta_2x_2+\dots+\beta_nx_n Among them, y represents the predicted risk value, x_1, x_2, \dots, x_n represent the factors affecting the risk, and \beta_0, \beta_1, \dots, \beta_n represent the model parameters.
9. A pharmaceutical compliance risk warning and decision-making recommendation system based on deep learning, characterized by: The system for implementing the pharmaceutical compliance risk warning and decision-making recommendation method based on deep learning as described in any one of claims 1 to 8 comprises: Pharmaceutical compliance risk identification and assessment module, used to implement multi-dimensional pharmaceutical compliance risk scanning, generate pharmaceutical compliance risk heat maps, and implement compliance checks on sensitive business data; A supplier risk assessment module, which is used to establish a supplier risk assessment prediction model, to implement the prediction process of the supplier risk assessment prediction model, to output the prediction results of the supplier risk assessment prediction model, and to analyze and apply the prediction results of the supplier risk assessment prediction model; The supplier qualification evaluation module is used to realize data integration and verification in the supplier qualification evaluation process, as well as to realize dynamic monitoring and early warning of supplier qualifications; The risk identification and prediction module in work tasks is used to realize data collection and preprocessing in work tasks, and to realize the extraction and selection of data features in work tasks, and to establish a risk identification and prediction model in work tasks, and to realize the training and evaluation process of the risk identification and prediction model in work tasks.
10. A medical compliance risk warning and decision-making recommendation device based on deep learning, characterized in that: include: Memory and processor; The memory stores executable instructions, and the processor is configured to execute the executable instructions in the memory to implement the steps of the deep learning-based pharmaceutical compliance risk warning and decision-making recommendation method as described in any one of claims 1 to 8.
Citation Information
Cited By
Data source access test method, system and device of credit platform and medium
CN120856623A
Machine learning-based clinical mass spectrum risk prediction method and equipment
CN122084815A
Machine Learning-Based Risk Prediction Methods and Equipment for Clinical Mass Spectrometry
CN122084815B