Drug sales management method based on artificial intelligence
By constructing an AI-based drug sales management method, integrating multi-source heterogeneous data, establishing predictive models and a blockchain traceability system, and optimizing inventory allocation, the problems of data integration and prediction accuracy in traditional drug sales management have been solved, achieving automation and transparency in drug sales management.
Patent Information
- Application Number
- CN202511445874.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-11
AI Technical Summary
Traditional drug sales management methods struggle to integrate multi-source heterogeneous data, fail to deeply explore the relationship between sales patterns and external factors, cannot accurately capture the changing characteristics of drug market demand, and lack strong adaptability.
An AI-based drug sales management approach is adopted. By collecting heterogeneous data from multiple sources, cleaning and standardizing it, a drug sales demand prediction model is constructed, a real-time transaction monitoring mechanism and a blockchain traceability system are established, an intelligent drug classification system is built, inventory allocation and control are carried out, and the prediction model is optimized through multi-level anomaly detection and causal relationship analysis.
It has improved the automation level of drug sales management, reduced the forecast cycle from weekly to daily, improved forecast accuracy, accelerated inventory turnover, shortened supply chain response time, reduced inventory costs, improved the accuracy of abnormal transaction identification, met data traceability requirements, and achieved optimization of the entire chain circulation efficiency and transparent supervision.
Smart Images

Figure CN120931329A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data management technology, and more specifically, to an artificial intelligence-based method for managing drug sales. Background Technology
[0002] With the rapid development of the modern pharmaceutical distribution system, especially in areas such as digital transformation, supply chain optimization, and intelligent management, the need for accurate prediction of drug sales demand, real-time monitoring of transaction anomalies, and end-to-end traceability is becoming increasingly urgent. While traditional drug sales management methods have made some progress in basic data collection, inventory management, and sales statistics, they have not systematically solved core issues such as the standardized integration of multi-source heterogeneous data, the analysis of discrepancies between predictive models and actual sales data, and the collaborative management of drug categories across regions. These shortcomings make it difficult to meet the refined management needs of drugs from different regions, through different channels, and for different product categories.
[0003] Therefore, how to integrate multi-source heterogeneous data, deeply explore the relationship between sales patterns and external factors, accurately capture the changing characteristics of drug market demand, and develop a highly adaptable intelligent management method for drug sales has become an urgent problem to be solved. Summary of the Invention
[0004] This invention provides an artificial intelligence-based drug sales management method that solves the technical problems of difficulty in integrating multi-source heterogeneous data, deeply mining the correlation between sales patterns and external factors, accurately capturing the characteristics of changes in drug market demand, and lacking strong adaptability.
[0005] This invention provides a drug sales management method based on artificial intelligence, comprising the following steps: Collect drug sales data from multiple heterogeneous sources, and clean, standardize, and store the drug sales data to obtain a drug sales dataset. Based on drug sales datasets and factors influencing drug sales, a predictive model for drug sales demand is trained and constructed. The prediction model forecasts the demand for drug sales in a future preset period to obtain the prediction results. Based on the forecast results, a real-time transaction monitoring mechanism will be established to detect and warn of abnormalities in drug transaction behavior. Based on the prediction results and real-time transaction monitoring mechanism, a blockchain traceability system is constructed to establish an immutable chain of drug sales records; Based on the forecast results and the blockchain traceability system, an intelligent drug classification system is constructed to obtain a refined inventory allocation and management strategy. The allocation and control strategy divides the various categories of drugs into zones, resulting in multiple sub-regional drug categories. A survey was conducted based on sub-regional drug category data and forecast results to obtain drug sales survey data. Based on pharmaceutical sales survey data, deviation analysis, correlation analysis, and causal relationship analysis were conducted on the actual sales data and forecast results of pharmaceutical categories in sub-regions to obtain the analysis results. Based on the analysis results, we will conduct in-depth analysis and screen out interfering factors that are emerging factors affecting sales data, so as to obtain accurate emerging factors. Accurate emerging influencing factors are fed back into the prediction model for parameter optimization and model iteration.
[0006] As a preferred embodiment of the present invention, a real-time transaction monitoring mechanism is established, including: Multiple risk level assessment threshold ranges are set based on the prediction results; Risk scores are generated by comparing the degree of deviation between actual sales transaction data and forecast data; The current risk level is determined by comparing the scoring results with multiple risk level assessment threshold ranges. It also includes constructing a multi-level anomaly detection architecture based on multiple risk level assessment threshold ranges to perform multi-level anomaly and violation detection on the prediction results, thereby obtaining detection results; the detection results indicate that sales data contains falsified or untrue information due to improper means or illegal sales behavior; (excessive drug procurement, abnormal price fluctuations, transactions at abnormal times). The first layer uses a rule engine to match multiple preset abnormal sales transaction rules; it then uses these rules to instantly filter out abnormal sales behaviors to identify abnormal sales activities. The second layer performs abnormality measurement risk scoring on abnormal sales behavior. If there are at least two abnormal sales behaviors, the risk score is assessed according to multiple risk level assessment threshold ranges in sequence to obtain the assessment results. The corresponding level of early warning mechanism is triggered in sequence according to the assessment results to obtain a graded early warning. The third layer is based on hierarchical early warning feedback, which dynamically adjusts the range of early warning thresholds to accurately define the risk level boundaries for multi-level abnormal violation detection.
[0007] As a preferred embodiment of the present invention, a blockchain traceability system is constructed, comprising: Based on the prediction results and real-time transaction monitoring mechanism, drug sales data is linked in an immutable chain with production batch information, logistics trajectory, and quality inspection reports; A distributed ledger architecture is established based on chain associations, storing the full lifecycle information of each drug transaction in the form of blocks; It also includes building a smart contract mechanism to trigger traceability and liability determination procedures when abnormal sales behavior is detected; The system enables rapid traceability and accountability for drug sales records throughout their entire lifecycle through traceability and accountability procedures.
[0008] As a preferred embodiment of the present invention, an intelligent drug classification system is constructed, comprising: Based on the prediction results and multiple influencing factors, the various categories of drugs are intelligently partitioned and categorized to form multiple sub-regional drug categories with different demand characteristics; the multiple influencing factors include, but are not limited to, sales geographical location, sales channels, drug categories, etc. The sub-regional drug categories are dynamically divided based on drug sales trend data to obtain drug categories with multiple trend characteristics. Specifically, by analyzing the historical sales data time series of sub-regional drug categories, and through trend analysis, drug categories whose sales data in the current period show an upward trend, a downward trend, a stable trend, and a fluctuating trend are identified, and drug categories with different trend characteristics are assigned to different trend partitions. For pharmaceutical categories in areas with an upward trend, develop an aggressive expansionary inventory strategy, increase safety stock levels and replenishment frequency, and prioritize the allocation of prime shelf locations and promotional resources; For pharmaceutical categories in the downward trend zone, a conservative and contractionary inventory strategy should be developed to reduce inventory levels and order quantities, and to strengthen clearance sales and recommend alternative products. For pharmaceutical categories with stable trends, develop a balanced maintenance inventory strategy to maintain existing inventory levels and replenishment pace. For pharmaceutical categories with fluctuating trends, develop flexible and adaptive inventory strategies, and set up dynamic safety stock and flexible replenishment mechanisms; the fluctuating trend zones refer to the frequent and obvious fluctuations in sales data, which are unstable trends. Based on multiple trend characteristics of drug categories, correlation analysis is conducted on drug categories in sub-regions according to drug efficacy and effects in order to establish a mechanism for coordinating efficacy-related drug categories. The efficacy-related category combination mechanism constructs a drug efficacy-related knowledge graph to identify drug category combinations that include synergistic therapeutic effects, complementary functional effects, and advantages of combined drug use, forming an efficacy-related category matrix. The efficacy-related category matrix includes: In the upward trend zone, when the sales of core efficacy drugs increase, the zone configuration and sales data of efficacy-related categories are promoted simultaneously, so as to drive the sales data of related categories to grow in tandem through the main category; In the declining trend zone, when the sales of core efficacy drugs decline, the configuration structure of efficacy-related categories should be adjusted in a timely manner, and sales demand should be shifted through the conversion of substitute categories and efficacy upgrades. Based on the efficacy-related product category matrix, a cross-regional efficacy-related product category linkage mechanism is established. When the sales data of core efficacy drugs in a certain region shows abnormal fluctuations, it triggers a sales warning for related efficacy categories in other regions, thereby achieving coordinated operation of the sales supply chain for main and auxiliary drugs.
[0009] As a preferred embodiment of the present invention, a survey is conducted based on sub-regional drug category data and prediction results, including: Basic data on drug categories in each sub-region are collected to obtain key sales indicators; these key sales indicators include quantitative indicators such as actual sales volume, inventory level, turnover rate, customer group characteristics, sales linkage coefficient, delay in the spread of abnormal events, scope of risk diffusion, and intensity of cross-regional demand transmission. The actual sales key indicators collected are compared and matched with the forecast results to identify sub-regions with abnormal forecast deviations and the drug categories in the current sub-region. For drug categories with abnormal forecast deviations, conduct in-depth research on the sales environment, market competition, and factors influencing changes in consumer behavior in the current sub-regions of the drug categories. Integrate all sales-influencing factors from the survey information to form structured drug sales survey data; By mining sales correlation patterns, abnormal sales behavior propagation paths, and risk diffusion patterns among sub-regions based on drug sales survey data, we can identify emerging factors and potential risks affecting drug sales. The emerging factors affecting drug sales form an interconnected network of influences. The influence network identifies, quantifies, and predicts complex relationships among multiple factors to analyze and provide forward-looking warnings about changes in the drug sales environment.
[0010] The influencing network specifically includes: The trend of electronic prescriptions driven by the popularization of digital healthcare and the development of telemedicine services have a synergistic effect, jointly promoting the digital transformation of drug sales channels and impacting the traditional pharmacy sales model. This digital transformation process directly catalyzes the in-depth application of artificial intelligence-assisted diagnostic technology. With the improvement of digital medical infrastructure, the deep integration of artificial intelligence-assisted diagnostic technology and the trend of electronic prescriptions has driven the rise of precision medicine demand and personalized drug sales models. The popularization of precision medicine models, in turn, has accelerated the improvement of health management awareness throughout society. The increased demand for preventative medications driven by heightened health management awareness is amplified exponentially through social media health information dissemination channels. The two factors mutually promote each other and reshape the structure of drug categories, forming a new drug consumption ecosystem oriented towards preventative healthcare. Social media health information dissemination guides and influences consumers' drug selection behavior, forming a positive feedback loop with the improvement of health management awareness. This loop effect further strengthens the market acceptance and user stickiness of telemedicine services. The development of telemedicine services has increased the demand for timeliness and convenience in drug delivery. This is interdependent with multiple factors such as the popularization of digital healthcare, the trend of electronic prescriptions, and the application of artificial intelligence-assisted diagnostic technology. Ultimately, this has led to the construction of an integrated online and offline drug sales ecosystem characterized by technology-driven, demand-oriented, and service-integrated features. This ecosystem achieves intelligent collaborative optimization of the entire drug sales chain through a multi-factor linkage mechanism.
[0011] As a preferred embodiment of the present invention, deviation analysis, correlation analysis, and causal relationship analysis are performed on the actual sales data and predicted results of drug categories in sub-regions to obtain the analysis results, including: The absolute deviation, relative deviation, and deviation distribution characteristics of the actual sales data and forecast results for each sub-region are obtained to identify the drug categories and time points with abnormal deviations, and to obtain the deviation analysis results. Based on the deviation analysis results, the correlation coefficients of sales data between different sub-regions and different drug categories are obtained to obtain the correlation analysis results of sales data between drug categories. Based on the correlation analysis results, the root causes affecting prediction bias are identified through causal inference algorithms, including external environmental factors, model parameter settings, and causal factors affecting data quality issues, so as to obtain the causal relationship analysis results. By integrating the results of the three analyses, a comprehensive analysis result is formed, which includes a description of deviation characteristics, a correlation map, and a causal chain of influence. Based on the comprehensive analysis results and drug sales survey data, we will conduct an in-depth analysis of emerging influencing factors affecting drug sales and screen out interfering factors of emerging influencing factors in order to obtain accurate emerging influencing factors.
[0012] As a preferred embodiment of the present invention, the emerging influencing factors affecting drug sales are analyzed in depth and interfering factors are screened out, including: Establish a multi-dimensional factor analysis framework to quantitatively evaluate emerging influencing factors through three dimensions: impact intensity, effect duration, and duration, and form an influencing factor matrix; An interference factor identification system is constructed based on the influencing factor matrix to identify and screen out interference factors among emerging influencing factors; Among them, outlier detection is performed on emerging influencing factors to obtain outlier detection results; The outlier detection results identify abnormal fluctuation points in policy and environmental factor data; When the data on the frequency of medical insurance catalog adjustments shows a preset outlier exceeding the standard deviation of the historical mean, it is marked as a potential interference factor and reviewed and confirmed. The confirmed outlier data is processed using the median replacement method, and the processing results are fed back to the influencing factor matrix for weight adjustment. Based on the outlier detection results, redundant variables in socioeconomic factors were analyzed and eliminated. When the correlation coefficient between changes in residents' income level and health awareness in socioeconomic factors exceeds a preset threshold, it is identified as a highly correlated redundant variable. Combining the influence strength in the influence factor matrix, the main factors of the redundant variables with strong influence strength are retained, while the secondary factors of the redundant variables are removed to avoid multicollinearity interference. The screening results are then updated to the influence factor matrix.
[0013] As a preferred embodiment of the present invention, it further includes: By identifying low-variance interference terms among emerging influencing factors, when the variance of new drug development progress data within a preset time range is less than a preset value of the overall variance, a low-information-content interference factor is obtained. Low-information interference factors are removed from the model input variables to obtain the removal results; The rejection results are transmitted to a dynamic monitoring mechanism for recording; Based on the elimination results, non-stationary interference series in the competitive landscape factors are screened out by time series stationarity test, and ADF test is performed on market concentration change data. When the P value is greater than the preset value, it is determined to be a non-stationary interference factor. The interference factors of non-stationary sequences are eliminated by difference transformation or detrending processing to obtain the processing result; The processing results are fed back to the interference factor processing verification program to obtain accurate interference factor processing results; As a preferred embodiment of the present invention, it further includes: Establish a dynamic monitoring mechanism for interference factors. Based on the identification results of various interference factors, set the detection threshold for interference factors to be a dual standard of influence weight being lower than a preset value and predicted contribution being lower than a preset value. Among them, when a certain interference factor meets the interference factor standard for a continuous preset time, it is automatically removed from the influence factor matrix. The removal decision triggers the redistribution of the weights of the influence factor matrix. A verification procedure for handling interference factors was constructed, and backtesting was performed on the set of influencing factors after filtering out interference factors. The emerging influencing factors before and after filtering out interference factors are input into the prediction model to obtain two output results. The root mean square error improvement is obtained based on the two output results. When the improvement exceeds the preset improvement value, the interference factor removal is confirmed to be effective. The set of emerging influencing factors after interference factor removal is then input into the multi-dimensional factor analysis framework for the next round of evaluation. Otherwise, the feedback is sent to the dynamic monitoring mechanism to re-evaluate the removal criteria, forming a closed-loop optimization mechanism to ensure the continuous accuracy of influencing factor identification and interference factor removal.
[0014] Secondly, a computer-readable storage medium is provided for storing computer-readable instructions that, when read by a computer, enable the execution of the aforementioned artificial intelligence-based drug sales management method.
[0015] The beneficial effects of this invention are as follows: By automatically collecting, cleaning, and standardizing multi-source heterogeneous data, the data integration time is significantly shortened and efficiency is significantly improved; the intelligent demand forecasting based on the LSTM-XGBoost fusion forecasting model reduces the forecasting cycle from weekly to daily, significantly improves forecasting accuracy, and greatly reduces the amount of manual intervention, effectively improving the automation level of drug sales management; through the intelligent drug classification system and the cross-trend regional inventory coordination and linkage mechanism, the drug circulation turnover speed is greatly improved, the supply chain response time is significantly shortened, and a systematic optimization of the entire chain circulation efficiency is achieved.
[0016] By constructing multi-dimensional feature vectors, we can deeply explore sales influencing factors such as time characteristics, policy characteristics, and correlation characteristics, and accurately identify the characteristics of changes in drug market demand. By using Pearson correlation coefficient and Spearman correlation coefficient to analyze the correlation strength and direction of sales data between different sub-regions, the accuracy of correlation identification is significantly improved. By using Granger causality test and cointegration test methods to analyze the causal relationship and long-term equilibrium relationship of sales data between regions, we can successfully identify the action path and intensity of external factors such as the lag effect of policy changes, the periodic influence of seasonal factors, and the impact of sudden events. We can also establish a dynamic influencing factor knowledge graph to achieve a comprehensive insight and in-depth analysis of drug sales patterns.
[0017] Based on the vector autoregression model, the dynamic impact path and transmission mechanism of external shocks on sales in each sub-region are identified, significantly improving the accuracy of demand change early warning. By analyzing the root causes of prediction inaccuracies through causal inference algorithms, the mutual influence relationship of sales data between sub-regions and the regional demand transmission mechanism are explored, and cross-regional synergistic effects are identified, greatly improving the accuracy of demand fluctuation prediction. A factor influence quantification model is established, assigning weight coefficients and impact delay parameters to each key influencing factor, enabling real-time monitoring and accurate prediction of market demand change characteristics.
[0018] Through an intelligent drug classification system and a multi-dimensional zoning management strategy, the proportion of slow-moving drug inventory has been significantly reduced, and the capital turnover rate has been greatly improved. Based on the ABC-XYZ classification method combined with trend zoning characteristics, a multi-dimensional inventory allocation and management strategy is implemented. For upward trend zoning, an aggressive expansion strategy is implemented; for downward trend zoning, a conservative contraction strategy is implemented; for stable trend zoning, a balanced maintenance strategy is adopted; and for volatile trend zoning, a flexible adaptation strategy is constructed. As a result, inventory holding costs have been significantly reduced and stockout rates have been significantly controlled. A cross-trend zoning inventory coordination and linkage mechanism and an inventory substitution scheme based on efficacy association have been established. Overall inventory turnover efficiency has been significantly improved, and significant optimization of inventory costs and maximization of economic benefits have been achieved.
[0019] A two-tiered abnormal transaction monitoring mechanism combining rules and algorithms has been established, significantly improving the accuracy of abnormal transaction identification and greatly reducing the cost of investigation. Blockchain technology has been used to create an immutable chain of drug sales records, linking sales data with production batch information, logistics trajectories, and quality inspection reports, fully meeting the data traceability requirements of the "Regulations on the Supervision and Administration of Drug Circulation," and significantly shortening the response time for quality issues. A smart contract mechanism has been established, automatically triggering traceability queries and responsibility identification when abnormal transactions are detected or prediction deviations exceed thresholds, significantly improving the efficiency of full-process supervision and achieving transparent supervision and compliance assurance throughout the entire drug lifecycle.
[0020] By inputting the survey results of key influencing factors into the prediction model in the form of structured data, and using incremental learning technology to achieve online adaptive updates of the model, the prediction accuracy has been continuously and significantly improved. A model performance monitoring mechanism has been established to regularly evaluate the improvement in prediction accuracy, and to dynamically adjust the model architecture and intelligently optimize the algorithm based on the evaluation results, ensuring that the system can adapt to changes in the market environment and maintain a long-term stable high-precision prediction capability. An inventory allocation decision support system has been built, integrating three core algorithm modules: trend prediction, cost optimization, and risk assessment, to generate personalized inventory parameter suggestions for each drug category, achieving precise and intelligent inventory allocation management. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of a drug sales management method based on artificial intelligence provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the allocation and control strategy method provided in the embodiments of the present invention. Detailed Implementation
[0022] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.
[0023] like Figure 1 As shown, an artificial intelligence-based drug sales management method includes the following steps: Step 1: Collect drug sales data from multiple heterogeneous sources, and clean, standardize, and store the drug sales data to obtain a drug sales dataset; Step 2: Based on the drug sales dataset and factors influencing drug sales, train and construct a predictive model for drug sales demand; the predictive model predicts drug sales demand for a preset future period to obtain the prediction results; Step 3: Based on the prediction results, establish a real-time transaction monitoring mechanism to detect and issue early warnings for abnormal drug transaction behavior; Step 4: Based on the prediction results and real-time transaction monitoring mechanism, construct a blockchain traceability system to establish an immutable chain of drug sales records; Step 5: Based on the prediction results and the blockchain traceability system, construct an intelligent drug classification system to obtain a refined inventory allocation and management strategy.
[0024] like Figure 2 As shown, the allocation control strategy further includes: Step 51: The allocation and control strategy divides the drug categories into zones and categorizes them to obtain multiple sub-regional drug categories; Step 52: Conduct a survey based on the sub-regional drug category data and the prediction results to obtain drug sales survey data; Step 53: Based on the drug sales survey data, conduct deviation analysis, correlation analysis, and causal relationship analysis on the actual sales data and forecast results of drug categories in sub-regions to obtain the analysis results; Step 54: Based on the analysis results, further analyze and screen out interfering factors that are emerging factors affecting sales data to obtain accurate emerging factors; and feed the accurate emerging factors back into the prediction model for parameter optimization and model iteration.
[0025] Furthermore, the prediction model is a fusion model, which uses a weighted average method. The sub-model output integrates a Long Short-Term Memory (LSTM) network, Extreme Gradient Boosting (XGBoost), and an Autoregressive Integral Moving Average (ARIMA) model; the weight coefficients of each sub-model are dynamically determined based on its root mean square error (RMSE) on the validation set, and the specific calculation formula is as follows: Where i and j represent LSTM, XGBoost, and ARIMA models, respectively; the weights are recalculated each time the model is retrained, and satisfy the following conditions: .
[0026] The construction of the blockchain traceability system includes: adopting the Hyperledger Fabric consortium blockchain framework, with network nodes including drug regulatory authorities, manufacturers, and distributors; using the Byzantine Fault Tolerance (PBFT) algorithm for the consensus mechanism; and using Go language for the smart contracts, where the preset exception triggering condition is: when the risk score output by the real-time transaction monitoring mechanism is greater than 55 points or the prediction deviation continues to exceed 25% for a specified period of time, the traceability query contract is automatically executed; the traceability query contract traces the entire life cycle record of the drug by querying the Merkle tree structure of the on-chain blocks, and calls the responsibility location algorithm to automatically generate a responsibility report based on the transaction timestamp, digital signature, and process metadata.
[0027] The anomaly detection rule base in the real-time transaction monitoring mechanism is generated and updated in the following way: the initial rule set is defined by business experts based on historical violation cases; subsequently, frequent anomaly patterns are discovered by analyzing historical transaction data and using association rule learning (such as the Apriori algorithm), and added to the rule base after being confirmed by experts; each rule is configured with confidence and support thresholds and is assigned a dynamic weight, which is periodically adjusted according to the historical trigger accuracy of the rule.
[0028] The trend analysis in constructing the intelligent drug classification system includes: using statistical significance tests (such as the Mann-Kendall trend test) to determine the trend of the sales data time series; for each drug category, calculating the Sen's slope of its sales for n consecutive months (n≥3), and determining an upward trend when the slope is greater than 0 and the p-value is less than 0.05; determining a downward trend when the slope is less than 0 and the p-value is less than 0.05; determining a stable trend when the p-value is greater than or equal to 0.05; and determining a fluctuating trend when the unit root test (ADF test) indicates that the series is non-stationary and the trend cannot be clearly defined by trend tests.
[0029] It also includes screening out interfering factors from emerging factors affecting drug sales: using variance inflation factor (VIF) to detect and screen out redundant variables with multicollinearity higher than a preset threshold (VIF>10); using low variance filtering to remove low-information features with variances lower than 1% of the sum of variances of all features; and performing unit root tests (ADF tests) on time series influencing factors to screen out non-stationary series and standardize stationary series.
[0030] The smart contract in the blockchain traceability system contains the core TraceabilityQuery function. The execution logic of this function is as follows: Listen to the message queue output by the real-time transaction monitoring module. When the risk score is greater than 55 or the prediction deviation rate of a certain category exceeds 25-28% for t consecutive hours, the contract will be automatically invoked.
[0031] Based on the drug batch number in the triggering event, the blockchain ledger is traversed, and by comparing the drug information stored in the block body (such as production batch and logistics tracking number), the Merkle tree structure is used to quickly locate and retrieve all relevant transaction blocks.
[0032] The contract's embedded ResponsibilityLocator analyzes the retrieved on-chain data and automatically determines the most likely stage and responsible party for any anomalies based on transaction type, process (production, logistics, sales), digital signature, and timestamp anomalies. It then generates a structured responsibility report, stores it on the blockchain, and notifies relevant regulatory nodes.
[0033] The trend judgment in the intelligent drug classification system employs statistical methods to enhance objectivity: The Mann-Kendall nonparametric trend test was used to analyze the historical sales time series of pharmaceutical products. The Z-statistic and p-value were calculated, and a significant trend was considered to exist when the p-value was less than the significance level α (usually set to 0.05).
[0034] The magnitude of the trend is calculated using Sen's slope estimation method. Sen's slope is a robust slope estimation method that is insensitive to outliers in the time series.
[0035] The trend is determined by combining the significance test results and Sen's slope value: when p < 0.05 and Sen's slope > 0, it is determined to be a significant upward trend; when p < 0.05 and Sen's slope < 0, it is determined to be a significant downward trend; when p ≥ 0.05, it is determined to be a stable trend (no significant trend); when the sequence is determined to be non-stationary by the ADF test, but the Mann-Kendall test does not show a significant monotonic trend, it is determined to be a fluctuating trend.
[0036] The data preprocessing procedure for the emerging influencing factors affecting drug sales, aimed at screening out confounding factors, includes the following steps: Obtain the variance inflation factor (VIF) for all continuous variables. Variables with VIF values greater than 10-15 are considered to have severe multicollinearity. Based on business knowledge, retain the more representative variables and remove the remaining redundant variables.
[0037] Obtain the variance of all features and remove features whose variance is less than 1-2% of the sum of the variances of all features. These features have too small a variation and are considered low-information noise.
[0038] For time series factors such as "new drug R&D progress index" and "market concentration", the ADF unit root test is performed. If the test result shows that the series is non-stationary (p value greater than 0.05), the series is subjected to first-order differencing or detrending until it passes the stationarity test, so as to eliminate the interference of non-stationary series as input to the prediction model.
[0039] Specifically, the AI-based drug sales management system constructed in this embodiment of the invention adopts a distributed architecture design, with the following specific configuration: The server uses 2-4 Intel Xeon Gold 6248 (16-24 cores) processors and 64-256GB of memory as the main computing nodes, equipped with high-performance SSDs to ensure data read and write efficiency; the database cluster adopts a 3-5 node HBase cluster architecture, with each node having a storage capacity of 1-5TB, and uses RAID5 configuration to ensure data security and reliability; it is equipped with NVIDIA Tesla V100 or A100 GPU accelerator cards for AI model training and inference calculations, with GPU memory of 16-80GB, supporting parallel computing to improve model training efficiency.
[0040] The operating system uses CentOS 7.6 or later, providing a stable Linux operating environment; the database system uses HBase 2.4.9 or later, supporting distributed storage and real-time querying of large-scale data; the AI computing framework uses TensorFlow 2.6-2.10 and XGBoost 1.5-1.7, used for building deep learning models and gradient boosting models respectively; the stream processing engine uses Apache Flink 1.12-1.16, supporting real-time data stream processing and low-latency computing; the blockchain platform uses Hyperledger Fabric 2.2-2.5, building a consortium blockchain environment to ensure data immutability.
[0041] The system is built on a client / server (C / S) architecture, comprising a four-layer software system: a basic support layer (database, middleware, network communication), a platform service layer (data service, algorithm service, blockchain service), a core business layer (prediction unit, monitoring unit, traceability unit, classification unit), and an application layer (user interface, API interface, reporting system). Each layer interacts with the others through standardized interfaces, ensuring the system's modularity and scalability.
[0042] First, deploy all software modules on the server, configure API interface permissions and data encryption certificates, and establish a secure data transmission channel. Then, import historical sales data from the past 2-5 years (approximately 20-80 million records), totaling about 300-800GB, using batch import to complete the data migration. Complete the pre-training process of the prediction model, which takes approximately 24-72 hours, generating initial model parameters. Configure an abnormal transaction rule base, pre-setting judgment rules and threshold parameters for 10-30 types of violation patterns, specifically including: Large transaction anomaly pattern: Transactions whose single transaction amount exceeds 2-5 times the historical average standard deviation, with the threshold set at the historical average + 2σ to 3σ. Frequent abnormal transaction pattern: The same customer makes more than 20-50 transactions within 24 hours, or the cumulative transaction amount within 7 days exceeds 1.5-3 times the monthly average; Abnormal pricing patterns: Transactions where drug sales prices deviate from the average market price by more than ±20%-30%, or where the price difference of the same drug in different channels exceeds 10%-20%; Inventory anomaly pattern: Anomalies where inventory data does not match sales data, with a discrepancy rate exceeding 3%-8%; Abnormal Time Pattern: Large transactions outside of business hours (9 PM to 7 AM), with a single transaction amount exceeding 3,000-8,000 yuan; Geographical anomaly: Cross-regional transactions where the customer's purchase location is more than 300-800 kilometers away from the registered address; Abnormal drug combination pattern: purchasing a large number of drugs with the same effect in a short period of time, exceeding the normal dosage by 5-15 times; Abnormal payment patterns: using multiple payment methods to split a single large transaction, or cash transactions accounting for more than 60%-90% of total transactions; Abnormal customer behavior patterns: Newly registered customers making their first transaction exceeding 5,000-15,000 yuan, or dormant customers suddenly activating and making large transactions; Prescription drug violation patterns: selling prescription drugs without a prescription, or selling drugs that do not match the prescription; Data integrity anomaly pattern: Abnormal records with a critical field missing rate exceeding 5%-15%, or data formats that do not conform to standard specifications; Supplier anomaly mode: expired supplier qualifications, abnormal price fluctuations exceeding 15%-25%, or unauthorized supplier transactions; Abnormal return patterns: Abnormal situations where the return rate exceeds 10%-20% or the return amount accounts for more than 8%-15% of the sales amount; Abnormal system operation mode: A large number of data modification operations in a short period of time, with a single user modifying more than 500-2000 records within 1 hour; Compliance anomaly mode: Sales activities that violate drug management regulations, such as selling expired drugs, counterfeit drugs, or unapproved imported drugs. Various anomaly thresholds can be dynamically adjusted according to actual business conditions.
[0043] Example 1 In this embodiment of the invention, a complete data acquisition and processing system was constructed based on the actual needs of drug sales management. Specifically: First, the data collection targets are clearly defined, including multi-dimensional data such as drug sales transaction records, inventory change data, supplier information, customer purchasing behavior, price fluctuation information, and regulatory compliance records; the collection scope is determined to include heterogeneous data sources such as medical institutions at all levels, drug retail enterprises, drug wholesale enterprises, and e-commerce platforms; and data interface technology and web crawling technology are used to establish connection channels with various data sources.
[0044] Real-time acquisition of pharmaceutical sales transaction data is achieved through API interfaces; data from the inventory management system is obtained through database connections; warehouse environment data is collected using sensor networks; market price information is obtained through third-party data services; automated data collection equipment periodically captures publicly available data from the internet; and sales data from special channels are manually entered to ensure the comprehensiveness and accuracy of data collection.
[0045] Each data source is validated and verified. This data collection method ensures data diversity and integrity, reflecting real-time drug sales and adapting to the data needs of different business scenarios. During the collection process, factors such as data timeliness, accuracy, and compliance are considered. Multi-source data collection allows predictive models to more accurately analyze drug sales trends, especially in responding to market changes, enabling the rapid acquisition of key information. Simultaneously, this collection method reduces data silos, facilitating subsequent data integration.
[0046] Based on the results of multi-source heterogeneous data collection, each data source is assigned a corresponding data cleaning task. Specific tasks are allocated according to the data quality and characteristics of that data source to ensure comprehensive data processing. Various technical methods are used to clean and standardize the collected pharmaceutical sales data: Sales transaction data includes core fields such as transaction time, drug code, quantity, amount, buyer information, and sales channel. Inventory data includes drug inventory levels, safety stock thresholds, replenishment records, and expired drug disposal records. Supplier data includes information such as supplier qualifications, supply capacity, pricing structure, and cooperation history; Customer data includes characteristics such as customer type, purchasing preferences, credit status, and geographical distribution; Regulatory data includes drug approval numbers, quality inspection reports, adverse reaction reports, recall records, etc.
[0047] Based on the five data types collected, information from each business scenario is integrated into an electronic pharmaceutical sales data archive. Each archive will include data from different data sources, formats, and types, ensuring that data for each pharmaceutical category can be stored, updated, and managed uniformly; this electronic archive records the detailed sales trajectory of each pharmaceutical category. The included pharmaceutical sales dataset includes basic pharmaceutical information, sales history, market performance, etc., and also records special management information for special pharmaceuticals (such as controlled substances, high-risk drugs, etc.).
[0048] Data cleaning removes duplicate records, corrects errors, and fills in missing data to ensure data quality. The deduplication process compares unique identifiers such as drug codes, transaction times, and customer IDs to identify and merge duplicate transaction records. Error correction utilizes a standardized database of basic drug information to automatically identify and correct spelling errors in drug names, specifications, and manufacturers. Outlier detection uses statistical analysis to identify extreme data points such as price and sales anomalies, and determines the appropriate handling method based on business rules. Missing value imputation employs time series interpolation and regression prediction to fill in missing sales data.
[0049] Furthermore, a data collection program is launched daily at a set time (e.g., 7:00-9:00 AM) to automatically acquire the previous day's sales data from various pharmacies, hospitals, and distributors via a pre-defined API interface. The data collection volume is approximately 500,000 to 2 million records per day, with a single API call response time controlled within 100-500ms. Simultaneously, a web crawler is launched to collect sales data from partner e-commerce platforms. The crawler uses a distributed architecture and supports concurrent processing.
[0050] Data standardization processing employs multimodal AI technology based on a BERT pre-trained model to intelligently identify and standardize information such as drug names. The system calls the BERT model to perform synonym mapping for drug names (e.g., acetaminophen tablets and paracetamol tablets); it uses regular expressions to extract and format key fields such as specifications, quantities, and prices, improving the accuracy of field extraction; and it supports automatic identification and processing of structured data (e.g., Excel spreadsheets) and unstructured data (e.g., PDF sales orders).
[0051] The cleaned and transformed standardized data is stored in an HBase distributed database, with a single node supporting 500-2000 data entries per second and data latency controlled within 30-100ms. The entire data processing flow is completed within 10-30 minutes. A data quality monitoring mechanism is established to monitor data integrity, accuracy, and consistency indicators in real time, thereby improving the accuracy of data cleaning.
[0052] Example 2 Based on the cleaned pharmaceutical sales dataset, an intelligent demand forecasting model is constructed. Specifically: In the feature engineering phase, key feature variables are extracted from the raw data, including time features (seasonality, periodicity, trend), drug features (category, price, specifications), market features (competitive intensity, policy impact), and customer features (purchase frequency, preference analysis). Correlation analysis, principal component analysis, and other methods are used to screen out the feature combinations that have the greatest impact on sales forecasting.
[0053] The model construction employs an ensemble learning approach, combining time series analysis, machine learning, and deep learning techniques. Specifically, the ARIMA model captures trends and seasonal characteristics of time series data; the random forest model handles nonlinear relationships and feature interactions; the LSTM neural network learns long-term dependencies; and the Prophet model handles the impact of holidays and special events. A weighted fusion method is used to integrate the prediction results from multiple models, improving prediction accuracy and stability.
[0054] The model training process employs a rolling time window approach, using historical sales data as the training set to progressively predict and update model parameters. Cross-validation and grid search are used to optimize model hyperparameters, ensuring the model's generalization ability. A model evaluation system is established, using metrics such as root mean square error, mean absolute percentage error, and prediction accuracy to evaluate model performance.
[0055] Dataset partitioning and model selection: Partition by time (to avoid future data leakage), such as using data from 2021-2022 as the training set (60-70%), January-June 2023 as the validation set (15-20%), and July-December 2023 as the test set (15-25%).
[0056] Choose the appropriate model based on the characteristics of the data: Time series models: suitable for univariate or strongly time-dependent data, such as ARIMA (handling linear trends) and Prophet (automatically identifying trends and seasonal effects, suitable for business users to use quickly). Machine learning models: suitable for scenarios with multiple factors, such as random forest (handling non-linear relationships and resisting overfitting) and XGBoost (high accuracy, suitable for data with high feature dimensions). Deep learning models are suitable for massive amounts of data (such as tens of millions of records), such as LSTM (which captures long and short-term time dependencies and is suitable for high-frequency data in online pharmacies).
[0057] By defining the objective function, with the goal of "minimizing the error between predicted sales and actual sales", common loss functions include mean squared error (MSE) and mean absolute error (MAE). Tune parameters (such as the number of trees in a random forest or the hidden layer dimension of an LSTM) using grid search or Bayesian optimization, and evaluate the optimal parameters using a validation set. Quantitative indicators: RMSE (Root Mean Square Error, reflecting the magnitude of the error), MAE (Mean Absolute Error, reflecting the average deviation), R² (Coefficient of Determination, the closer to 1, the better the fit). Qualitative analysis: Observe the consistency of trends between the predicted curve and the actual curve (e.g., whether it accurately captures the sales peak before the Spring Festival).
[0058] Predicting future drug sales demand over a predetermined period: Based on a trained model, input data on future influencing factors, output prediction results, and adapt them to business needs.
[0059] Preset time period: Set according to business needs (such as the next 7 days, 30 days, 90 days), short-term for replenishment, long-term for production planning.
[0060] Input the influencing factors data for the forecast period: Known factors include: future holiday arrangements (such as the date of the Spring Festival in 2024), planned promotional activities (such as the "Double 11" discount plan), and fixed policies (such as no adjustments to the medical insurance catalog in 2024). Predictive factors include: estimating future influenza incidence rates based on historical incidence rate trends (using time series models for extrapolation), and assuming competitor prices remain at current levels (if no available plans).
[0061] Prediction logic: The model outputs sales based on input features (such as weekday characteristics for the next 7 days, promotional plans, and estimated incidence rates) combined with historical patterns. For example, the Prophet model decomposes the trend (long-term growth / decline), seasonal effects (monthly / quarterly cycles), and holiday effects, and outputs the predicted value after superimposing them; For example, LSTM models predict future sales by remembering past sales patterns (such as sales being higher on Fridays than on Thursdays) and combining them with future features.
[0062] Output format: Output by time period granularity (e.g., daily / weekly sales) with confidence intervals (e.g., 80-95% confidence interval, reflecting forecast uncertainty).
[0063] If the sales forecast is negative (unreasonable), adjust it to 0; if it is known that a hospital is about to add a new procurement plan, appropriately increase the forecast value for the corresponding region. Forecast results are summarized by region / drug category (e.g., "Total demand for cold medicine in all pharmacies in Beijing in the next 30 days").
[0064] Regularly verify the prediction error using actual sales data. If the error exceeds the threshold (e.g., RMSE exceeds 20%), retrain the model (add new data, supplement unconsidered influencing factors, such as sudden outbreaks of epidemics). Dynamically update features (such as adding "online consultation volume" as a predictive feature for related drugs).
[0065] Specifically, this includes: In the feature engineering phase, constructing 64-256 dimensional feature vectors, including time features (quarters, holidays, days of the week, etc.), policy features (adjustments to the medical insurance catalog, changes in drug approval policies, etc.), correlation features (the sales correlation between cold medicine and thermometers, etc.), and economic indicator features (residents' income levels, health awareness, etc.). Feature selection is then performed using methods such as Pearson correlation coefficient and mutual information to ultimately determine the feature combinations most influential on prediction.
[0066] The LSTM neural network design employs a 2-4 layer structure, with 64-128 neurons in the hidden layers. It uses the Adam optimizer, with a learning rate of 0.1-0.3, a batch size of 16-64, and 50-80 training epochs. The network can capture long-term dependencies and trend changes in pharmaceutical sales data, improving the accuracy of identifying long-term trends such as annual sales cycles.
[0067] XGBoost model parameter configuration: This model mainly handles nonlinear relationships and the impact of sudden factors, such as the impact of epidemics, policy adjustments, and competitor product launches on sales, and can effectively predict the impact of sudden events.
[0068] The ARIMA time series model employs an automatic parameter selection algorithm, determining the optimal (p, d, q) parameter combination through the AIC criterion. It is primarily used to capture the seasonal and periodic characteristics of sales data. The model ensemble uses a weighted fusion strategy, with LSTM weights of 0.3-0.5, XGBoost weights of 0.3-0.5, and ARIMA weights of 0.1-0.3, achieving a final prediction accuracy of over 88%-95%.
[0069] The model training employs incremental learning techniques, automatically fine-tuning the parameters of the fusion model daily using the latest sales data, with the fine-tuning time controlled within 20-60 minutes. A model performance monitoring mechanism is established, automatically triggering a model retraining process when the prediction accuracy drops by more than 3%-8%.
[0070] Example 3 Based on the output of the prediction model, a three-layer real-time transaction monitoring mechanism is constructed. The specific implementation process is as follows: The first-layer rule engine performs real-time matching using a pre-set database of violation rules. This database contains 15-25 violation patterns: excessive drug procurement (single purchase quantity exceeding historical averages by 2-5 times), abnormal price fluctuations (prices deviating from the market average by more than 25%-35%), abnormal time transactions (large transactions late at night or on holidays), frequent returns (more than 3-5 returns within 7 days), and cross-regional abnormal transactions (transactions exceeding normal delivery range), etc. Each rule has clearly defined trigger conditions and weighting coefficients. The system monitors transaction flows in real time and immediately marks any matched violation pattern.
[0071] The second-layer anomaly scoring system quantitatively assesses the marked abnormal behavior. The scoring algorithm comprehensively considers factors such as the weight of the violation rule, the degree of prediction bias, and historical risk records to calculate a comprehensive risk score. The risk score ranges from 0 to 100 points, with 0-25 points indicating low risk, 26-55 points indicating medium risk, and 56-100 points indicating high risk. Different levels of alerts are triggered based on the risk level: low-risk alerts display a yellow warning on the monitoring panel; medium-risk alerts send email notifications to relevant personnel and generate work orders in the system; high-risk alerts immediately notify management personnel by phone and automatically freeze relevant transactions.
[0072] The third-layer dynamic optimization module continuously optimizes system parameters based on early warning feedback and detection accuracy. By statistically analyzing historical early warning accuracy, it identifies false alarms and missed alarms, and dynamically adjusts rule weights and risk thresholds. A feedback learning mechanism is established, using manually confirmed early warning results as training samples to continuously optimize the performance of the anomaly detection algorithm.
[0073] Specifically, the system employs the Apache Flink real-time computing engine as its streaming processing platform, supporting real-time processing capabilities of 5,000-15,000 transactions per second, with processing latency controlled within 50-200ms. The rules engine utilizes the Drools rules engine, supporting dynamic rule configuration and hot updates.
[0074] The anomaly detection algorithm employs the Isolation Forest algorithm for unsupervised anomaly detection. Algorithm parameters are set as follows: number of trees 50-200, subsample size 128-512, and anomaly ratio 3%-8%. The algorithm can identify multi-dimensional anomaly patterns, including single-dimensional anomalies and multi-dimensional combined anomalies. The risk scoring algorithm comprehensively considers factors such as violation rule weights (weight range 0.05-1.0), prediction bias (bias threshold 10%-35%), and historical risk records (risk decay coefficient 0.8-0.95) to calculate a comprehensive risk score.
[0075] The early warning grading mechanism is designed as follows: 0-25 points indicate low risk (green), and a prompt message is displayed on the monitoring panel; 26-55 points indicate medium risk (yellow), and relevant personnel are notified by email and a work order is generated in the system; 56-100 points indicate high risk (red), and management personnel are immediately notified by phone and relevant transactions are automatically frozen to prevent further losses.
[0076] The reinforcement learning optimization employed the Q-learning algorithm, adjusting the warning threshold parameters based on feedback from manual review. The learning rate was set at 0.05-0.2, the discount factor at 0.85-0.95, and the exploration rate gradually decreased to 0.05-0.15. Through continuous learning, the system's false alarm rate decreased, and the warning accuracy significantly improved.
[0077] Example 4 To ensure the immutability and traceability of drug sales data, a traceability system based on blockchain technology will be constructed. Specific implementation includes: The distributed ledger architecture is designed as a consortium blockchain, consisting of a network of key nodes including drug regulatory authorities, manufacturers, distributors, and medical institutions. Each node maintains a complete copy of the ledger, and data consistency is ensured through a consensus mechanism. The block structure includes a block header (timestamp, previous block hash, Merkle root) and a block body (transaction data, digital signature, smart contract execution result).
[0078] The smart contract mechanism is equipped with automated traceability and anomaly handling rules. When the system detects an abnormal transaction, the smart contract automatically triggers a traceability procedure to query the complete supply chain information of the drug from production to sales. The contract includes a pre-set liability location algorithm that automatically identifies the responsible party based on the anomaly type and the stage at which it occurred, and generates a liability report.
[0079] The data upload process uses a batch submission method, packaging verified transaction data onto the blockchain every hour to reduce network load while ensuring data timeliness. Sensitive information is encrypted before being uploaded to the blockchain to protect business privacy. A multi-factor verification mechanism is established, requiring at least three nodes to confirm each transaction before it is written to the blockchain.
[0080] The visual query interface provides multi-dimensional traceability query functions. Users can quickly locate the entire life cycle record of a target drug by criteria such as batch number, production date, and sales channel. The interface displays key information such as drug production information, quality inspection reports, distribution trajectory, sales records, predictive matching degree, and abnormal detection results, forming a complete traceability chain diagram.
[0081] Specifically, this includes using Hyperledger Fabric 2.2-2.5 as the underlying blockchain platform to build a consortium blockchain network composed of key nodes such as drug regulatory authorities, manufacturers, distributors, and medical institutions. Each node maintains a complete copy of the ledger, and the PBFT (Byzantine Fault Tolerance) consensus mechanism is used to ensure data consistency, with consensus time controlled within 2-5 seconds.
[0082] The block structure design consists of two parts: a block header and a block body. The block header contains metadata such as the timestamp, the hash value of the previous block (SHA-256 algorithm), the Merkle root hash, and the block height. The block body contains key traceability elements such as drug transaction data, digital signatures, smart contract execution results, prediction matching degree, anomaly detection results, and regulatory compliance status. Each block size is controlled within 0.5-2MB to ensure network transmission efficiency.
[0083] The smart contract mechanism is developed using the Go programming language and includes automated traceability and anomaly handling rules. The contract pre-defines responsibility identification algorithms, including: an anomaly type identification algorithm (based on rule matching and machine learning), a responsibility link identification algorithm (tracking through timestamps and data flow), and a responsibility entity confirmation algorithm (combining digital signatures and access control). When the system detects an abnormal transaction or a prediction deviation exceeding the 25%-35% threshold, the smart contract automatically triggers a traceability procedure to query the complete supply chain information of the drug from production to sales, generating a responsibility report within 8-15 seconds.
[0084] The data upload optimization employs a batch submission strategy, packaging verified transaction data onto the blockchain every 30 minutes to 2 hours, processing 500-8000 transaction records per batch, effectively reducing network load. Before upload, commercially sensitive information is encrypted using the AES-256 encryption algorithm to protect business privacy. A multi-factor verification mechanism is established, requiring at least 2-5 nodes to confirm each transaction before it is written to the blockchain, ensuring data reliability.
[0085] Query performance optimization utilizes indexing mechanisms and caching strategies to support fast queries based on prediction accuracy, anomaly type, and time range, with query response times controlled within 3-5 seconds. The query interface supports visual displays, including timeline charts, flowcharts, and relationship diagrams.
[0086] Example 5 Based on the prediction results and multi-dimensional analysis, an intelligent drug classification management system is constructed. The specific implementation plan is as follows: The multi-dimensional classification first categorizes drugs based on dimensions such as geographical location, sales channels, drug category, and prediction accuracy. The geographical dimension is divided according to the three-level administrative divisions of province, city, and county; the channel dimension distinguishes sales channels such as hospitals, pharmacies, e-commerce, and clinics; the category dimension is classified according to therapeutic area and pharmacological action; and the accuracy dimension is graded based on historical prediction deviation rate.
[0087] Trend analysis uses time series decomposition technology to identify the sales trend characteristics of each drug category. Upward trend identification: Sales growth rate is positive for 3-6 consecutive months and the growth is accelerating; Downward trend identification: Sales decline for 2-4 consecutive months and the decline gradually widens; Stable trend identification: Sales fluctuation is within ±6% and there is no obvious trend; Fluctuating trend identification: Sales changes exceed 15%-20% and there is no discernible pattern.
[0088] Differentiated inventory strategies are developed based on different trend characteristics. For drugs with an upward trend, an aggressive expansion strategy is adopted: safety stock is increased by 20%-40%, replenishment frequency is changed from monthly to weekly, and prime shelf locations are prioritized. For drugs with a downward trend, a conservative contraction strategy is adopted: replenishment cycles are extended, and promotional clearance sales are strengthened. For drugs with a stable trend, a balanced maintenance strategy is adopted: existing inventory parameters are kept unchanged, and adjustments are made regularly. For drugs with a volatile trend, a flexible adaptation strategy is adopted: dynamic safety stock is set, and a rapid replenishment mechanism is established.
[0089] Efficacy association analysis constructs relationships between drug categories using a drug knowledge graph. The knowledge graph contains information such as drug indications, contraindications, drug interactions, and combination therapy regimens. Graph neural network algorithms are used to uncover hidden association patterns and identify drug combinations with synergistic therapeutic effects. An association coefficient matrix is then established to quantify the strength of associations between different drugs.
[0090] The cross-regional linkage mechanism establishes an automated early warning and coordination system. When the sales of core drugs in a certain region experience abnormal fluctuations, the system automatically analyzes the sales of related products in other regions and predicts potential chain reactions. Triggering mechanisms include inventory alerts, replenishment recommendations, and adjustments to promotional strategies. The coordination mechanism supports cross-regional inventory transfers and resource reallocation.
[0091] The refined inventory allocation and control strategy also includes: based on the ABC-XYZ classification method, further integrating trend zoning features to achieve multi-dimensional and precise inventory control. The classification algorithm combines fuzzy theory with traditional classification methods, uses entropy weighting to determine classification weights, and subdivides pharmaceuticals into 6-12 control categories, with each category configured with differentiated inventory parameters.
[0092] For pharmaceutical categories exhibiting an upward trend, an aggressive expansion-oriented inventory allocation and control strategy is adopted: the safety stock coefficient is set at 1.5-2.0 times the standard demand, replenishment frequency is increased to 2-3 times per week, a priority shelf location allocation mechanism is established, and a 20%-30% demand growth buffer inventory is set. When forecasts indicate a continued upward trend, an automatic expansion replenishment mechanism is activated, with the replenishment trigger point set at 70%-90% of the safety stock, and the replenishment quantity being 1.8-2.0 times the economic order quantity. Simultaneously, a coordinated inventory increase plan for efficacy-related categories is established; when the inventory of core pharmaceutical products increases, the inventory of efficacy-related categories increases synchronously by a coefficient of 0.4-0.8.
[0093] For pharmaceutical categories experiencing a downward trend, a conservative, contractionary inventory allocation and control strategy will be implemented: the safety stock coefficient will be adjusted to 0.5-0.8 times the standard demand, the replenishment cycle will be extended to once every 2-3 weeks, and an intelligent clearance and promotion trigger mechanism will be established. When the inventory turnover days exceed 60 days, a price optimization algorithm will be automatically activated, with promotional discounts ranging from 5% to 15%. Simultaneously, a substitute recommendation algorithm will be activated to recommend pharmaceuticals with similar efficacy and an upward trend. A tiered inventory clearance plan will be constructed, prioritizing sales based on the proximity of the expiration date, prioritizing the clearance of near-expiration pharmaceuticals to maximize and preserve inventory value.
[0094] For pharmaceutical categories with stable trends, a balanced maintenance inventory allocation and control strategy will be implemented: maintaining a safety stock of 1.0-1.2 times the standard demand and a fixed replenishment cycle of 10-14 days. A precise demand forecasting-inventory matching mechanism will be established, using moving averages to smooth out short-term fluctuations and ensure a high degree of alignment between inventory levels and actual demand. Simultaneously, an inventory efficiency monitoring system will be established to maintain inventory turnover at an industry-leading level, with a target turnover rate of 8-12 times per year.
[0095] For pharmaceutical categories with fluctuating trends, a flexible and adaptive inventory allocation and control strategy is constructed: a dynamic safety stock range of 0.8-1.8 times the standard demand is set, and a two-tiered elastic replenishment mechanism is established, including basic replenishment (based on fixed periods) and emergency replenishment (triggered by demand fluctuations). When demand fluctuations exceed a preset threshold ±25%, the emergency replenishment procedure is automatically triggered, and the emergency replenishment quantity is dynamically calculated based on the fluctuation amplitude. A multi-scenario inventory buffer pool is constructed, with corresponding inventory adjustment plans preset for different fluctuation modes such as seasonal fluctuations, sudden event impacts, and policy influences.
[0096] When abnormal inventory pressure occurs in a certain trend zone, intelligent algorithms identify the inventory redundancy space in other zones and calculate the optimal inventory transfer plan. The inventory transfer cost model comprehensively considers factors such as transportation costs, time costs, and opportunity costs to ensure the economic efficiency of the transfer decision. Simultaneously, it establishes inventory substitution plans based on efficacy correlation; when a certain drug is out of stock, it automatically recommends several alternative drugs with similar efficacy as preferred options.
[0097] An inventory allocation decision support system is constructed, integrating three core algorithm modules: trend prediction, cost optimization, and risk assessment. The trend prediction module outputs trend prediction results based on an LSTM-XGBoost fusion model; the cost optimization module uses an integer programming algorithm to calculate the optimal inventory allocation; and the risk assessment module uses Monte Carlo simulation to evaluate inventory risk. The system generates personalized inventory parameter suggestions for each drug category, including key indicators such as optimal reorder point, economic order quantity, and maximum inventory limit.
[0098] Example 6 The prediction results are evaluated and optimized using multi-dimensional analysis methods, and a closed-loop feedback mechanism is established to continuously improve the system's prediction accuracy and business value. The specific implementation process includes: The deviation analysis and deep diagnostic mechanism adopts a multi-level deviation analysis framework to obtain the predictive deviation index of drug categories in each sub-region.
[0099] The absolute deviation (MAE) reflects the direct difference between the predicted and actual values. ,in, For the first One actual observation value, For the first One predicted value, The total number of samples; Relative bias (MAPE) eliminates the influence of dimensions, facilitating comparisons between different product categories. Specifically: ,in, For the first One actual observation value, For the first One predicted value, The total number of samples; The root mean square error (RMSE) is more sensitive to large deviations, and the formula is as follows: ,in, For the first One actual observation value, For the first One predicted value, The total number of samples.
[0100] Deviation distribution analysis uses kernel density estimation to identify the distribution patterns and outliers of deviations, and the 3σ criterion to identify outliers. A multi-level deviation early warning mechanism is established: a green safe zone (deviation rate 1-10%), a yellow watch zone (deviation rate 10%-15%), an orange warning zone (deviation rate 15%-30%), and a red danger zone (deviation rate exceeding 30%). Corresponding handling strategies are developed for different warning levels: yellow warnings trigger parameter fine-tuning; orange warnings initiate model retraining; and red warnings involve comprehensive model diagnosis and reconstruction.
[0101] In-depth deviation diagnosis includes time-dimensional deviation analysis (identifying temporal patterns in prediction accuracy), spatial-dimensional deviation analysis (discovering the impact of geographical factors on prediction accuracy), and category-dimensional deviation analysis (analyzing differences in prediction difficulty among different drug types). A deviation attribution algorithm is established to automatically identify the main factors causing prediction deviations, including data quality issues, model fit issues, and changes in the external environment.
[0102] The correlation analysis and association mining system employs multiple statistical methods to deeply explore the relationships between variables. Pearson correlation coefficient analyzes linear relationships and is suitable for normally distributed data; Spearman correlation coefficient analyzes monotonic relationships and is suitable for non-normally distributed data; Kendall's τ correlation coefficient analyzes ordinal correlation and is more robust to outliers. A correlation heatmap visually displays the strength of associations between different sub-regions and drug categories. A correlation coefficient with an absolute value greater than 0.7 is defined as a strong correlation, 0.4-0.7 as a moderate correlation, and less than 0.4 as a weak correlation.
[0103] Construct a dynamic correlation monitoring system, employing a sliding window technique (window size 15-45 days) to continuously monitor correlation changes and promptly identify evolving trends in relationships. Identify highly correlated product category combinations to provide data support for joint promotions and inventory coordination. Establish a correlation prediction model to forecast future trends in relationships, providing forward-looking guidance for strategy adjustments.
[0104] Multi-dimensional correlation analysis is implemented, specifically including: geographic correlation analysis to identify sales linkages between different regions; temporal correlation analysis to discover lagged correlations between time series; category correlation analysis to uncover substitution and complementarity relationships between drugs; and channel correlation analysis to explore synergistic effects between different sales channels. The strength of these multi-dimensional correlations is quantified using a composite correlation index, providing a basis for precision marketing and inventory optimization.
[0105] Causal relationship analysis and prediction mechanisms utilize various causal inference methods to identify the direction and strength of causal relationships between variables. Granger causality tests identify causal relationships between time series, rejecting the null hypothesis when the test statistic F is greater than the critical value; cointegration tests (Johansen method) analyze long-term equilibrium relationships between variables, identifying common trends and adjustment speeds; vector autoregression (VAR) models characterize the dynamic relationships of multivariate systems, capturing the mutual influences between variables.
[0106] A structured causal graph is constructed, using a directed acyclic graph (DAG) to represent the causal relationship network between variables. Causal discovery algorithms (such as PC and FCI algorithms) are used to automatically learn the causal structure from observational data. Causal inference analysis is performed to distinguish between genuine causal relationships and spurious correlations, providing reliable causal evidence for decision-making.
[0107] Impulse response function analysis quantifies the dynamic impact of external shocks on the system: the duration and decay rate of policy change shocks (such as adjustments to medical insurance policies); the lagged impact of seasonal factors (such as the flu season) on the sales of related drugs; and the short-term and long-term impacts of sudden events (such as epidemic outbreaks). An impact transmission path diagram is established to identify key nodes and amplification mechanisms in the spread of the impact.
[0108] The model iterative optimization and adaptive learning mechanism establishes a continuous model optimization framework based on the analysis results. Parameter optimization employs a variety of advanced algorithms: grid search for exhaustive parameter tuning; random search to improve search efficiency; Bayesian optimization to guide the search direction based on prior knowledge; genetic algorithm to simulate the evolutionary process to find the optimal solution; and particle swarm optimization (PSO) to search for the global optimum through swarm intelligence.
[0109] Feature optimization employs a multi-level feature selection strategy: filtering methods (such as chi-square test and mutual information) quickly screen relevant features; wrapping methods (such as recursive feature elimination) evaluate the predictive performance of feature subsets; and embedded methods (such as LASSO regression and RandomForest feature importance) select features during model training. A dynamic feature importance evaluation mechanism is established to periodically update feature weights, eliminate redundant features, and introduce new effective features.
[0110] Structural optimization adjusts the model architecture based on prediction performance and business needs: neural network depth optimization (determining the optimal number of layers through validation set error); ensemble learning strategy optimization (adjusting the number and weights of base learners); hybrid model architecture design (combining the advantages of different model types). A mechanism for balancing model complexity and performance is established to optimize model efficiency while ensuring prediction accuracy.
[0111] Adaptive learning mechanisms enable continuous model evolution: online learning algorithms support real-time model updates and handle concept drift in data streams; incremental learning methods incorporate new data without retraining the entire model; transfer learning techniques transfer the knowledge of trained models to new scenarios; and multi-task learning simultaneously optimizes multiple related prediction tasks.
[0112] Establish an A / B testing mechanism and model evaluation system, and design rigorous controlled trials to compare the effects of different optimization schemes. The A / B testing framework includes: experimental design (determining test indicators, sample allocation, and test period); statistical analysis (hypothesis testing, confidence intervals, and effect size calculation); and result interpretation (statistical significance, actual business value, and potential risks). Establish a multi-dimensional evaluation indicator system: prediction accuracy indicators (RMSE, MAPE, SMAPE); business value indicators (inventory turnover improvement, sales growth rate, and cost savings); and system performance indicators (response time, concurrent processing capability, and stability).
[0113] Build a model performance monitoring and early warning system to monitor key indicators such as prediction accuracy, feature distribution, and prediction bias in real time. When model performance deteriorates (accuracy decreases by more than 3%-8% or bias exceeds the threshold for 5-10 consecutive days), automatically trigger the model retraining process. Establish a model version management mechanism to support model rollback and A / B comparison, ensuring the safety and controllability of model updates.
[0114] The intelligent optimization decision support system integrates all the above-mentioned analysis and optimization functions, providing managers with visualized optimization suggestions and decision support. System outputs include: model performance diagnostic reports, optimization strategy recommendations, risk assessment results, and implementation path planning. An optimization effect evaluation mechanism is established to track the actual effects of optimization measures, forming a closed-loop continuous improvement system.
[0115] Through the specific implementation methods described above, this invention achieves intelligent management of pharmaceutical sales data, establishing a complete system from data collection, predictive analysis, real-time monitoring to traceability management. The entire system possesses self-learning, self-optimization, and self-adaptive capabilities, continuously improving prediction accuracy and business value, providing a scientific, efficient, and intelligent technical solution for pharmaceutical sales management. The system's prediction accuracy, inventory turnover rate, and anomaly detection accuracy are all improved, providing crucial technical support for the digital transformation and intelligent upgrading of the pharmaceutical industry.
[0116] A computer-readable storage medium, characterized in that it is used to store computer-readable instructions, which, when read by a computer, enable the execution of an artificial intelligence-based drug sales management method: Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media.
[0117] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.
Claims
1. A drug sales management method based on artificial intelligence, characterized in that, Includes the following steps: Collect drug sales data from multiple heterogeneous sources, and clean, standardize, and store the drug sales data to obtain a drug sales dataset. Based on drug sales datasets and factors influencing drug sales, a predictive model for drug sales demand is trained and constructed. The prediction model forecasts the demand for drug sales in a future preset period to obtain the prediction results. Based on the forecast results, a real-time transaction monitoring mechanism will be established to detect and warn of abnormalities in drug transaction behavior. Based on the prediction results and real-time transaction monitoring mechanism, a blockchain traceability system is constructed to establish an immutable chain of drug sales records; Based on the forecast results and the blockchain traceability system, an intelligent drug classification system is constructed to obtain a refined inventory allocation and management strategy.
2. The drug sales management method based on artificial intelligence according to claim 1, characterized in that, Establish a real-time transaction monitoring mechanism, including: Multiple risk level assessment threshold ranges are set based on the prediction results; Risk scores are generated by comparing the degree of deviation between actual sales transaction data and forecast data; The current risk level is determined by comparing the scoring results with multiple risk level assessment threshold ranges. It also includes constructing a multi-level anomaly detection architecture based on multiple risk level assessment threshold ranges to perform multi-level anomaly violation detection on the prediction results and obtain the detection results; The first layer uses a rule engine to match multiple preset abnormal sales transaction rules; it then uses these rules to instantly filter out abnormal sales behaviors to identify abnormal sales activities. The second layer performs abnormality measurement risk scoring on abnormal sales behavior. If there are at least two abnormal sales behaviors, the risk score is assessed according to multiple risk level assessment threshold ranges in sequence to obtain the assessment results. The corresponding level of early warning mechanism is triggered in sequence according to the assessment results to obtain a graded early warning. The third layer is based on hierarchical early warning feedback, which dynamically adjusts the range of early warning thresholds to accurately define the risk level boundaries for multi-level abnormal violation detection.
3. The drug sales management method based on artificial intelligence according to claim 1, characterized in that, Building a blockchain traceability system includes: Based on the prediction results and real-time transaction monitoring mechanism, drug sales data is linked in an immutable chain with production batch information, logistics trajectory, and quality inspection reports; A distributed ledger architecture is established based on chain associations, storing the full lifecycle information of each drug transaction in the form of blocks; It also includes building a smart contract mechanism to trigger traceability and liability determination when abnormal sales behavior is detected; Rapid traceability and accountability for drug sales records throughout their entire lifecycle are achieved through traceability queries and responsibility identification.
4. The drug sales management method based on artificial intelligence according to claim 1, characterized in that, Constructing an intelligent drug classification system includes: Based on the prediction results and multiple influencing factors, the various categories of drugs are intelligently partitioned and categorized to form multiple sub-regional drug categories with different demand characteristics. The sub-regional drug categories are dynamically divided based on drug sales trend data to obtain drug categories with multiple trend characteristics. Specifically, by analyzing the historical sales data time series of sub-regional drug categories, and through trend analysis, drug categories whose sales data in the current period show an upward trend, a downward trend, a stable trend, and a fluctuating trend are identified, and drug categories with different trend characteristics are assigned to different trend partitions. Based on multiple trend characteristics of drug categories, correlation analysis is conducted on drug categories in sub-regions according to drug efficacy and effects in order to establish a mechanism for coordinating efficacy-related drug categories. The efficacy-related category combination mechanism constructs a drug efficacy-related knowledge graph to identify drug category combinations that include synergistic therapeutic effects, complementary functional effects, and advantages of combined drug use, forming an efficacy-related category matrix. Based on the efficacy-related product category matrix, a cross-regional efficacy-related product category linkage mechanism is established. When the sales data of core efficacy drugs in a certain region shows abnormal fluctuations, it triggers a sales warning for related efficacy categories in other regions, thereby achieving coordinated operation of the sales supply chain for main and auxiliary drugs.
5. The drug sales management method based on artificial intelligence according to claim 1, characterized in that, A survey was conducted based on sub-regional drug category data and forecast results, including: Collect basic data on drug categories in each sub-region to obtain key sales indicators; compare and match the collected actual key sales indicators with the prediction results to identify sub-regions with abnormal prediction deviations and drug categories in the current sub-region; For drug categories with abnormal forecast deviations, conduct in-depth research on the sales environment, market competition, and factors influencing changes in consumer behavior in the current sub-regions of the drug categories. Integrate all sales-influencing factors from the survey information to form structured drug sales survey data; By mining sales correlation patterns, abnormal sales behavior propagation paths, and risk diffusion patterns among sub-regions based on drug sales survey data, we can identify emerging factors and potential risks affecting drug sales. The emerging factors affecting drug sales form an interconnected network of influences. The influence network identifies, quantifies, and predicts complex relationships among multiple factors to analyze and provide forward-looking warnings about changes in the drug sales environment.
6. The drug sales management method based on artificial intelligence according to claim 1, characterized in that, Deviation analysis, correlation analysis, and causal relationship analysis were performed on the actual sales data and forecast results of drug categories in sub-regions to obtain the following analysis results: The absolute deviation, relative deviation, and deviation distribution characteristics of the actual sales data and forecast results for each sub-region are obtained to identify the drug categories and time points with abnormal deviations, and to obtain the deviation analysis results. Based on the deviation analysis results, the correlation coefficients of sales data between different sub-regions and different drug categories are obtained to obtain the correlation analysis results of sales data between drug categories. Based on the correlation analysis results, the root causes affecting prediction bias are identified through causal inference algorithms, including external environmental factors, model parameter settings, and causal factors affecting data quality issues, so as to obtain the causal relationship analysis results. By integrating the results of the three analyses, a comprehensive analysis result is formed, which includes a description of deviation characteristics, a correlation map, and a causal chain of influence. Based on the comprehensive analysis results and drug sales survey data, we will conduct an in-depth analysis of emerging influencing factors affecting drug sales and screen out interfering factors of emerging influencing factors in order to obtain accurate emerging influencing factors.
7. The drug sales management method based on artificial intelligence according to claim 6, characterized in that, This involves in-depth analysis of emerging factors affecting drug sales and screening out confounding factors, including: Establish a multi-dimensional factor analysis framework to quantitatively evaluate emerging influencing factors through three dimensions: impact intensity, effect duration, and duration, and form an influencing factor matrix; An interference factor identification system is constructed based on the influencing factor matrix to identify and screen out interference factors among emerging influencing factors; Among them, outlier detection is performed on emerging influencing factors to obtain outlier detection results; The outlier detection results identify abnormal fluctuation points in policy and environmental factor data; When the data on the frequency of medical insurance catalog adjustments shows a preset outlier exceeding the standard deviation of the historical mean, it is marked as a potential interference factor and reviewed and confirmed. The confirmed outlier data is processed using the median replacement method, and the processing results are fed back to the influencing factor matrix for weight adjustment. Based on the outlier detection results, redundant variables in socioeconomic factors were analyzed and eliminated. When the correlation coefficient between changes in residents' income level and health awareness in socioeconomic factors exceeds a preset threshold, it is identified as a highly correlated redundant variable. Combining the influence strength in the influence factor matrix, the main interference factor of the redundant variable with the strong influence strength is retained, and the secondary interference factor of the redundant variable is removed. The screening results are then updated to the influence factor matrix.
8. The drug sales management method based on artificial intelligence according to claim 7, characterized in that, Also includes: By identifying low-variance interference terms among emerging influencing factors, when the variance of new drug development progress data within a preset time range is less than a preset value of the overall variance, a low-information-content interference factor is obtained. Low-information interference factors are removed from the model input variables to obtain the removal results; The rejection results are transmitted to a dynamic monitoring mechanism for recording; Based on the elimination results, non-stationary interference series in the competitive landscape factors are screened out by time series stationarity test, and ADF test is performed on market concentration change data. When the P value is greater than the preset value, it is determined to be a non-stationary interference factor. The interference factors of non-stationary sequences are eliminated by difference transformation or detrending processing to obtain the processing result; The processing results are fed back to the interference factor processing verification to obtain accurate interference factor processing results.
9. A drug sales management method based on artificial intelligence according to claim 8, characterized in that, Also includes: Establish a dynamic monitoring mechanism for interference factors. Based on the identification results of various interference factors, set the detection threshold for interference factors to be a dual standard of influence weight being lower than a preset value and predicted contribution being lower than a preset value. Among them, when a certain interference factor meets the interference factor standard for a continuous preset time, it is automatically removed from the influence factor matrix. The removal decision triggers the redistribution of the weights of the influence factor matrix. A verification of interference factor treatment was constructed, and the set of influencing factors after screening out interference factors was backtested for verification. The emerging influencing factors before and after filtering out interference factors are input into the prediction model to obtain two output results. The degree of improvement of the root mean square error is obtained based on the two output results. When the degree of improvement exceeds the preset improvement value, the interference factor is confirmed to be effective. The set of emerging influencing factors after the interference factors are removed is then input into the multi-dimensional factor analysis framework for the next round of evaluation. Otherwise, the feedback is sent to the dynamic monitoring mechanism to re-evaluate the removal criteria.
10. A computer-readable storage medium, characterized in that, Used to store computer-readable instructions, which, when read by a computer, enable the execution of an artificial intelligence-based drug sales management method as described in any one of claims 1-9.
Citation Information
Patent Citations
Data analysis method and system for drug sales
CN119379326A
Drug sales inventory management system based on artificial intelligence
CN119919059A
Internet-based medicine wholesale and retail comprehensive service method and system
CN120525162A
Cited By
Commodity display image intelligent scoring method and system based on dynamic rule configuration
CN122289261A