Method and system for analyzing market share of drugs based on multi-source heterogeneous data
By using a multi-source heterogeneous data analysis method, we collect and standardize drug sales data, decompose demand signals using wavelet transform, and calculate dynamic market share by combining regional event characteristics. This solves the problem of insufficient identification of market demand fluctuations in traditional methods and achieves more accurate market share analysis.
Patent Information
- Application Number
- CN202511319586.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2026-03-24
- Estimated Expiration
- 2045-09-16
AI Technical Summary
Existing pharmaceutical market analysis techniques struggle to accurately identify fluctuations in market demand signals and lack sufficient analysis of regional market differences, resulting in unreliable analysis results.
A drug market share analysis method using multi-source heterogeneous data is adopted. By collecting drug sales data, performing time alignment and unit standardization, decomposing demand signals using wavelet transform technology, and combining the occurrence time and propagation sequence of regional events, a weighted composite dynamic market share is calculated, and historical data is used for correction.
It improves the accuracy and reliability of market share analysis, enabling accurate identification of regional panic demand and actual medical needs during public health emergencies, and generating more precise market share forecasts.
Smart Images

Figure CN121094862B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of big data analysis and processing, in particular to a drug market share analysis method and system based on multi-source heterogeneous data. BACKGROUND
[0002] With the rapid development of the pharmaceutical market, accurate analysis and prediction of drug market share are of great significance to market decision-making and drug supply chain management of pharmaceutical enterprises. Under the influence of public health emergencies, drug market demand shows significant volatility and regional differences, which poses new challenges to traditional market analysis methods.
[0003] Existing drug market analysis techniques are mainly based on time series analysis methods, which analyze and predict historical sales data by establishing mathematical models. These methods use signal processing techniques such as Fourier transform to decompose market demand signals into periodic fluctuations and trend changes, thereby identifying market change patterns.
[0004] However, the existing technology is difficult to accurately identify the fluctuation component of the market demand signal, and lacks analysis of regional market differences, resulting in insufficient reliability of the analysis results; this situation needs to be further improved. SUMMARY
[0005] In order to solve the problem that the existing drug market analysis technology is difficult to accurately identify the fluctuation component of the market demand signal, and lacks analysis of regional market differences, resulting in insufficient reliability of the analysis results, the present application provides a drug market share analysis method and system based on multi-source heterogeneous data, which adopts the following technical solutions:
[0006] In a first aspect, the present application provides a drug market share analysis method based on multi-source heterogeneous data, comprising the following steps:
[0007] Collecting drug sales data to obtain a drug market sales original data set in different regions;
[0008] According to the original data set, the data is time-aligned and unit-standardized to obtain a regional standardized data set;
[0009] Based on the standardized data set and the daily average sales surge rate index, the occurrence time and propagation order of public health emergencies in each region are determined;
[0010] According to the standardized data set, the daily drug sales data of each region is taken as the overall demand signal, and the occurrence time and propagation order of the region are combined to decompose the overall demand signal using wavelet transform to obtain regional panic demand components and actual medical demand baseline;
[0011] calculate a dynamic market share of each region based on the regional panic demand component and the actual medical demand baseline;
[0012] calculate a regional corrected market share prediction result according to the dynamic market share of each region and the historical data baseline.
[0013] By adopting the above technical solution, with the rapid development of the pharmaceutical market, the importance of market share analysis for enterprise decision-making and supply chain management is increasingly prominent; due to the sudden public health event, the demand of the pharmaceutical market will fluctuate sharply, and the traditional single data source analysis method is difficult to accurately grasp the market change rule; for example, in a certain emergency, the same category of drugs in different regions showed significant demand differences, some regions appeared short-term panic buying, while some regions maintained normal demand levels. The complex market dynamic characteristics make it difficult for traditional analysis methods to cope; the present application first collects drug sales data to build a complete market sales data set; through time alignment and unit standardization processing, the heterogeneity problem of different source data is solved; then, based on the standardized data set, the daily average sales surge rate index is set, the occurrence time and propagation order of the emergency in each region are identified; the daily drug sales of each region is taken as the total demand signal, and the wavelet transform technology is used to accurately decompose the demand signal in combination with the identified occurrence time and propagation order of each region, to obtain the regional panic demand component and the actual medical demand baseline; the dynamic market share of each region is calculated by weighted synthesis; finally, the dynamic market share is corrected in combination with the historical data baseline to generate the final prediction result; through multi-source data fusion, the information integrity is improved, the demand component is accurately identified by using wavelet transform, and the regional propagation characteristics are considered, so that the market share analysis is more accurate and reliable.
[0014] Optionally, based on the standardized data set and the daily average sales surge rate index, the occurrence time and propagation order of the sudden public health event in each region are determined, specifically including the following steps:
[0015] Based on the historical data set, a sales fluctuation feature model is constructed, a regional weight coefficient is set in combination with the regional market size, and a daily average sales surge rate index is determined;
[0016] Calculate the growth rate of the daily drug sales of each region relative to the average sales of the previous N days to obtain a daily average sales surge rate sequence of each region, 3≤N≤7;
[0017] According to the daily average sales surge rate sequence and the daily average sales surge rate index, the time point when each region first exceeds the index is selected as the regional event occurrence time;
[0018] Sort the regions in chronological order based on the regional event occurrence time to obtain a regional propagation order of the public health emergency.
[0019] Under the influence of the public health emergency, accurately identifying the occurrence time and propagation order of the event in different regions is crucial for analyzing the changes in the drug market; due to the differences in market size and sensitivity to the event in each region, the traditional fixed threshold detection method cannot accurately reflect the regional characteristics; the present application first constructs a sales fluctuation feature model based on historical data, sets corresponding weight coefficients in combination with the market size of each region, and forms a targeted daily sales surge rate index; then, 3 to 7 days are selected as the reference time window, and the growth rate of daily drug sales of each region relative to the average sales of the previous N days is calculated to generate a daily sales surge rate sequence reflecting market changes; by comparing these sequences with the set surge rate index, the time point at which each region first exceeds the index is determined, which is marked as the event occurrence time of the region; finally, based on the identified event occurrence time of each region, the regions are sorted in chronological order to obtain the complete regional propagation order; by considering the regional market characteristics to set a dynamic index, accurate identification of the event occurrence time is achieved, and the propagation rule of the event is revealed through time series analysis.
[0020] Optionally, the wavelet transform adopts a db4 wavelet function, the total demand signal is decomposed by 3 layers, and the total demand signal is decomposed by the wavelet transform to obtain a regional panic demand component and an actual medical demand baseline, specifically including the following steps:
[0021] The db4 wavelet function is used to perform 3-layer wavelet decomposition on the total demand signal to obtain high-frequency components and low-frequency components;
[0022] According to the regional event occurrence time, the regional panic demand component is extracted from the high-frequency components;
[0023] According to the regional propagation order, the low-frequency components are reconstructed to obtain an actual medical demand baseline.
[0024] By adopting the above technical solution, accurately distinguishing between panic demand and actual medical demand is a key challenge in pharmaceutical market demand analysis. Because these two types of demand differ significantly in their temporal characteristics and fluctuations, traditional time series analysis methods struggle to achieve effective separation. For example, during a sudden event, some regions experience short-term concentrated panic buying; this sudden high-frequency fluctuation mixes with stable medical demand, resulting in a complex superposition of market demand signals. This application first uses the db4 wavelet function to perform a three-level wavelet decomposition on the overall demand signal. Through this three-level decomposition, the demand signal is separated into high-frequency components reflecting rapid changes and low-frequency components reflecting long-term trends. Based on the previously identified regional event occurrence times, the fluctuations related to the event are located and extracted from the high-frequency components; these fluctuations represent the regional panic demand components. Considering the continuity and regional propagation characteristics of actual medical demand, the method reconstructs the low-frequency components according to the obtained regional propagation sequence, obtaining a baseline reflecting the true medical demand. Wavelet transform enables multi-scale analysis of the demand signal, and combined with the spatiotemporal characteristics of the event, targeted component extraction and reconstruction provide a more reliable data foundation for market share analysis.
[0025] Optionally, based on the occurrence time of the regional event, the regional panic demand component is extracted from the high-frequency component, specifically including the following steps:
[0026] In the high-frequency components, the corresponding signal analysis intervals are divided according to the time position corresponding to the occurrence time of the regional event;
[0027] Based on the fluctuation characteristics of historical data and the size of regional markets, an energy threshold is determined, and signal segments exceeding the energy threshold are extracted within the signal analysis interval as candidate intervals for panic demand.
[0028] The candidate intervals for panic demand are subjected to morphological feature analysis, and the morphological features include amplitude mutation rate, duration and decay rate.
[0029] The regional panic demand component is obtained by superimposing the intervals that conform to the morphological characteristics.
[0030] By adopting the technical scheme, firstly, the application determines the corresponding analysis interval in the high-frequency component signal according to the identified regional event occurrence time; the pertinence of the analysis range is ensured; then, by analyzing the fluctuation characteristics in the historical data and combining the regional market size to set a suitable energy threshold, the signal segment exceeding the energy threshold is preliminarily screened out in the delimited analysis interval as a candidate interval of panic demand; in-depth morphological feature analysis is performed on the candidate intervals, and the amplitude mutation rate, the duration and the decay rate are mainly investigated; the amplitude mutation rate reflects the intensity of demand growth, the duration reflects the maintenance time of the panic state, and the decay rate represents the demand decline trend; finally, the intervals that meet the morphological feature requirements are superimposed to obtain the complete regional panic demand component; the limitations of a single threshold are avoided, the typical characteristics of panic demand are fully considered, and the accuracy of panic demand identification is improved.
[0031] Optionally, according to the regional propagation order, the low-frequency component is reconstructed to obtain an actual medical demand baseline, and the method comprises the following steps:
[0032] Based on the regional propagation order, time series correlation analysis is performed on the low-frequency components of each region.
[0033] According to the time series correlation analysis result, the low-frequency component is trend reconstructed to generate a regional benchmark demand curve.
[0034] According to the overall demand signal and the regional panic demand component, an initial medical demand is calculated.
[0035] The initial medical demand is corrected by using the regional benchmark demand curve to obtain an actual medical demand baseline.
[0036] By adopting the technical scheme, since the medical demand has a propagation correlation between different regions, it is difficult to reflect the mutual influence between the regions by simply analyzing the low-frequency signal of a single region; firstly, according to the obtained regional propagation order, the time series correlation analysis is performed on the low-frequency components of each region to reveal the conduction law and influence strength of the demand change between the regions; based on the time series correlation analysis result, the low-frequency component is trend reconstructed to generate a benchmark demand curve reflecting the conduction characteristics between the regions; then, by stripping the identified regional panic demand component from the overall demand signal, a preliminary medical demand estimation is obtained; the initial medical demand is corrected by using the regional benchmark demand curve generated in the early stage to eliminate the abnormal fluctuations in the regional propagation process, and finally an actual characteristic medical demand baseline is obtained; the dynamic reconstruction of the medical demand baseline is realized.
[0037] Optionally, based on the regional panic demand component and the actual medical demand baseline, a dynamic market share of each region is calculated, specifically including the following steps:
[0038] The peak ratio and duration ratio of the regional panic demand component relative to the overall demand signal are calculated to determine the panic demand intensity coefficient of each region;
[0039] The growth rate and volatility rate of the actual medical demand baseline are extracted to calculate the medical demand change index of each region;
[0040] The panic demand share and medical demand share of each region are calculated based on the panic demand intensity coefficient and the medical demand change index, respectively;
[0041] The panic demand share and medical demand share of each region are weighted and synthesized to obtain the dynamic market share of each region.
[0042] By using the above technical solution, since panic demand is short-term and sudden and medical demand is continuous and stable, simply adding the two demands with equal weight cannot reflect the real structure of the market. The application first calculates the peak ratio and duration ratio of the regional panic demand component to quantify the intensity characteristics of panic demand. The peak ratio reflects the intensity of demand surge, and the duration ratio represents the duration of panic state. By analyzing the growth rate and volatility rate of the actual medical demand baseline, a medical demand change index reflecting the stability of demand is calculated. Based on these characteristic indexes, the market share of each region in the dimensions of panic demand and medical demand is calculated. The panic demand share and medical demand share calculated are weighted and synthesized to obtain a dynamic market share that can accurately reflect the market structure. The characteristics of different types of demand are considered, and the accuracy of market share calculation is improved.
[0043] In a second aspect, the application provides a drug market share analysis system based on multi-source heterogeneous data, comprising:
[0044] A data acquisition module is configured to acquire drug sales data and obtain a drug market sales original data set in different regions;
[0045] A standardization processing module is configured to perform time alignment and unit standardization processing on the data based on the original data set to obtain a regional standardized data set;
[0046] An event detection module is configured to determine the occurrence time and propagation order of a sudden public health event in each region based on the standardized data set and the daily sales surge rate index;
[0047] a demand decomposition module configured to, according to the standardized data set, take daily drug sales data of each region as an overall demand signal, combine the regional occurrence time and the propagation order, and decompose the overall demand signal by using wavelet transform to obtain a regional panic demand component and an actual medical demand baseline;
[0048] a market share calculation module configured to calculate a weighted and synthesized dynamic market share of each region based on the regional panic demand component and the actual medical demand baseline;
[0049] a prediction analysis module configured to calculate a regionally corrected market share prediction result according to the dynamic market share of each region and a historical data baseline.
[0050] In a third aspect, the present application provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the above-mentioned drug market share analysis method based on multi-source heterogeneous data when executing the computer program.
[0051] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the above-mentioned drug market share analysis method based on multi-source heterogeneous data when executed by a processor.
[0052] In summary, the present application has at least one of the following beneficial technical effects:
[0053] The present application first collects drug sales data, processes the data by time alignment and unit standardization to construct a unified data set, identifies the regional occurrence time and the propagation order of the sudden event based on the daily sales surge rate index, takes the regional drug sales as an overall demand signal, uses wavelet transform technology to realize demand signal decomposition, obtains a regional panic demand component and an actual medical demand baseline, calculates a regional dynamic market share by weighted synthesis, and corrects the final prediction result in combination with historical data; through multi-source data fusion and wavelet transform analysis, the accuracy of market share analysis is improved;
[0054] The application firstly adopts db4 wavelet function to perform 3-layer wavelet decomposition on the overall demand signal; through 3-layer decomposition, the demand signal is separated into high-frequency components reflecting rapid changes and low-frequency components embodying long-term trends; based on the identified event occurrence time of the region in advance, the fluctuation part related to the event is located and extracted in the high-frequency components, and these fluctuations represent the regional panic demand component; considering the continuity and regional propagation characteristics of actual medical demand, the method reconstructs the low-frequency components according to the obtained regional propagation sequence to obtain the baseline reflecting the real medical demand; through wavelet transform, multi-scale analysis of the demand signal is realized, targeted component extraction and reconstruction are performed in combination with the spatiotemporal characteristics of the event, and a more reliable data basis is provided for market share analysis;
[0055] The application firstly quantifies the intensity characteristics of panic demand by calculating the peak ratio and duration ratio of the regional panic demand component. The peak ratio reflects the degree of sudden increase in demand, and the duration ratio represents the maintenance time length of the panic state. The medical demand change index reflecting the demand stability is calculated by analyzing the growth rate and fluctuation rate of the actual medical demand baseline; based on these characteristic indexes, the market share of each region in the panic demand and medical demand dimensions is calculated respectively; the panic demand share and the medical demand share calculated are weighted and synthesized to obtain the dynamic market share which can accurately reflect the market structure; considering the characteristic differences of different types of demand, the accuracy of market share calculation is improved. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 is a flowchart of a drug market share analysis method based on multi-source heterogeneous data according to an embodiment of the application;
[0057] Figure 2 is a flowchart of step 300 in a drug market share analysis method based on multi-source heterogeneous data according to an embodiment of the application;
[0058] Figure 3 is a flowchart of step 400 in a drug market share analysis method based on multi-source heterogeneous data according to an embodiment of the application;
[0059] Figure 4 is a flowchart of step 420 in a drug market share analysis method based on multi-source heterogeneous data according to an embodiment of the application;
[0060] Figure 5 is a flowchart of step 430 in a drug market share analysis method based on multi-source heterogeneous data according to an embodiment of the application;
[0061] Figure 6is a flowchart of step 500 in a method for analyzing drug market share based on multi-source heterogeneous data according to an embodiment of the present application.
[0062] Figure 7 is a module schematic diagram of a system for analyzing drug market share based on multi-source heterogeneous data according to an embodiment of the present application.
[0063] Figure 8 is an internal structure diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0064] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification and the appended claims of the present application, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or," as used in the present application, refers to any or all possible combinations of one or more of the listed items.
[0065] Hereinafter, the terms "first" and "second" are only for the purpose of description, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, the meaning of "a plurality of" is two or more, unless otherwise specified.
[0066] The embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0067] In a first aspect, the present application provides a method for analyzing drug market share based on multi-source heterogeneous data, referring to Figure 1 , comprising the following steps:
[0068] S100, collecting drug sales data to obtain raw data sets of drug market sales in different regions.
[0069] In this embodiment, market sales data is collected through a standardized data authorization and management process. The collected data mainly includes desensitized terminal sales data summary information, authorized market research data, and publicly released industry statistical data, etc.
[0070] S200, according to the raw data set, time alignment and unit standardization processing is performed on the data to obtain a regional standardized data set.
[0071] In this embodiment, time alignment refers to unifying sales data of different sources and different recording times to daily granularity; unit standardization processing includes specification unification, quantity standardization, and amount normalization processing, which converts data of different measurement units into a comparable standard form.
[0072] Specifically, a drug specification conversion mapping table is established to record the conversion relationship between different packaging specifications and standard units; a regional price index database is constructed, including the price level coefficients of each region; and time granularity conversion rules are designed to aggregate scattered time point data according to unified rules.
[0073] S300, based on the standardized data set and the daily sales surge rate index, the occurrence time and propagation order of the public health emergency in each region are determined.
[0074] In this embodiment, the daily sales surge rate index refers to the growth rate of daily drug sales relative to the previous average level; the regional occurrence time refers to the time point when the sales significantly increase for the first time in each region; and the propagation order refers to the propagation sequence of the emergency between different regions.
[0075] Specifically, a regional reference sales database is established to record the past sales reference values of each region; a sales fluctuation early warning rule library is constructed, including the upper limit of normal fluctuation range, the surge early warning threshold, and the duration requirement. When the daily sales surge rate of a region continuously exceeds the early warning threshold, the time point is marked as the regional occurrence time. By recording the occurrence time of each region, the propagation path is determined according to the chronological order to form the regional propagation sequence.
[0076] S400, according to the standardized data set, the daily drug sales data of each region is taken as the overall demand signal, combined with the occurrence time and propagation order of the region, and the wavelet transform is used to decompose the overall demand signal to obtain the regional panic demand component and the actual medical demand baseline.
[0077] In this embodiment, the overall demand signal refers to the standardized regional daily sales data sequence; the panic demand component refers to the short-term abnormal demand caused by the emergency; and the actual medical demand baseline refers to the real clinical drug demand after excluding panic factors.
[0078] Specifically, a demand signal feature library is established, including periodic features, seasonal features, and trend features of normal medical demand; a regional correlation model is constructed to describe the demand transmission relationship between adjacent regions. By decomposing the overall demand signal, the high-frequency part that meets the panic feature is extracted, and the remaining part is analyzed for regional correlation to obtain the medical demand baseline. First, the correlation between regions is defined, including geographical adjacency relationship, traffic connection relationship and population flow relationship, among which, the direct adjacent value is 1, the indirect adjacent value is 0.5, the value based on main traffic trunk connectivity is 0-1, and the flow index is calculated based on permanent population migration data. Then, the demand transmission function is established, and the weighted calculation method is adopted: the transmission coefficient = w1 x geographical adjacency coefficient + w2 x traffic intensity coefficient + w3 x population flow index, wherein w1, w2, w3 are weights and the sum is 1. The time delay of demand change in adjacent regions is calculated, and the time lag function Δt = a x distance / transmission coefficient + b is fitted, wherein a and b are to-be-determined coefficients. On this basis, the demand correlation equation Y(t) = α x X(t-Δt) + β x Y(t-1) is established, wherein Y(t) is the demand of the target region at time t, X(t-Δt) is the demand of the source region after a delay of Δt, α is the transmission coefficient, and β is the autocorrelation coefficient. Finally, the accuracy of the model is verified using historical data, the prediction error is calculated and the parameters are adjusted, and a correction mechanism is introduced to handle outliers, thereby accurately describing the demand transmission relationship between regions. The regional correlation analysis is based on the above model, first, the theoretical medical demand value of each region is calculated using the demand correlation equation, then the actual residual demand is compared with the theoretical value, when the deviation exceeds the preset threshold, the weighted average method is used for correction, wherein the actual value weight is 0.6 and the theoretical value weight is 0.4; finally, the corrected data is smoothed, and the 5-day moving average method is used to eliminate short-term fluctuations to obtain the final medical demand baseline.
[0079] S500, based on the regional panic demand component and the actual medical demand baseline, the weighted and synthesized dynamic market share of each region is calculated.
[0080] In this embodiment, the panic demand intensity refers to the proportion of panic demand to total demand; the medical demand change refers to the growth trend of the baseline demand; and the dynamic market share refers to the market share index considering the influence of double demand.
[0081] Specifically, a demand feature quantification index system is established, including peak intensity index, duration cycle index and decay rate index; a market share calculation rule is designed, and the weight coefficient is determined according to the characteristics of different demand types. By weighting the panic demand share and the medical demand share, the market share index reflecting the dynamic change of the market is obtained.
[0082] S600, calculate the corrected market share prediction result of each region according to the dynamic market share and the historical data baseline.
[0083] In this embodiment, the historical data baseline refers to the market share level of each region under normal circumstances, and the corrected prediction result refers to the market share estimate considering the historical regularity.
[0084] Specifically, a market share correction model is established, which includes a historical baseline deviation term, a market size adjustment term and a regional competition factor; a prediction result evaluation index library is constructed to test the rationality of the prediction result. Through baseline correction and rationality test of the dynamic market share, the final output is the prediction result that meets the actual market. The historical baseline deviation term is obtained by calculating the standard deviation of the market share in the past 12 months, reflecting the degree of historical fluctuation; the market size adjustment term is calculated based on the GDP growth rate, population growth rate and medical resource input increment of each region, used to correct the development potential; the regional competition factor considers the number of main competitors, the market share of new entrants and the brand concentration index, used to evaluate the competition situation, and the final correction model adopts the form of weighted summation and is smoothed by a piecewise function.
[0085] In one embodiment, with reference to Figure 2 In step S300, based on the standardized data set and the daily sales surge rate index, the occurrence time and propagation order of the public health emergency in each region are determined, which specifically includes the following steps:
[0086] S310, based on the historical data set, a sales fluctuation feature model is constructed, and a regional weight coefficient is set according to the market size of the region to determine the daily sales surge rate index.
[0087] In this embodiment, the sales fluctuation feature model refers to a mathematical model describing the normal fluctuation of drug sales, including periodic fluctuation, random fluctuation and trend change; the regional weight coefficient is a correction parameter reflecting the difference in market size of different regions, which is determined by population, medical resource density and economic development level; the daily sales surge rate index is a quantitative standard for judging abnormal growth of sales.
[0088] Specifically, first, the monthly sales average and standard deviation of each region in the past two years are calculated by simple statistical methods as the basic fluctuation reference value. On this basis, each region is divided into large, medium and small regions according to the population size, and is assigned a basic weight of 1.2, 1.0 and 0.8 respectively. For the density of medical resources, the number of medical institution beds per 10,000 people is used for correction: when the number of beds exceeds 50, the weight is increased by 0.1; when it is less than 30, the weight is decreased by 0.1. The final surge rate index value is 2 times the basic fluctuation standard deviation multiplied by the corrected regional weight. Among them, for periodic fluctuations, the main periodic components in the sales sequence are extracted using Fourier transform, usually including 7 days, 30 days and 365 days three basic periods; for random fluctuations, the residual sequence is fitted using ARMA model, in which the autoregressive order and moving average order are determined by AIC criterion; for trend changes, the long-term trend is fitted using polynomial regression method, and the order is selected based on the goodness of fit R² value. The fluctuation characteristics of the three dimensions are combined by an additive model: sales = trend item + periodic item + random item, each component is equipped with a corresponding confidence interval to judge whether the actual sales significantly deviates from the normal range.
[0089] S320, calculate the growth rate of the daily sales of each region relative to the average sales of the previous N days, and obtain the daily average sales surge rate sequence of each region.
[0090] In this embodiment, the average sales of the previous N days refers to the arithmetic mean of the sales of the previous N natural days of a specific date, and the value of N ranges from 3 to 7 days; the daily average sales surge rate sequence refers to a time sequence composed of the growth rate of daily sales relative to the average value of the previous N days.
[0091] S330, according to the daily average sales surge rate sequence and the daily average sales surge rate index, screen the time point when each region first exceeds the index as the region event occurrence time.
[0092] In this embodiment, the first time exceeding the index refers to the first time when the surge rate is higher than the preset index appears within the observation period; the region event occurrence time refers to the time node when the influence of the emergency is confirmed to appear, which needs to meet the requirements of continuity and significance. Single-day anomaly is not enough to confirm the occurrence of the event, and it needs to exceed the index for consecutive days and meet a specific growth pattern.
[0093] Specifically, for the surge rate sequence of each region, first identify the time point when the surge rate exceeds the indicator value, and the exceeding amplitude should not be less than 10%; second, observe the change trend of the surge rate in the next three days, if the surge rate of at least two days in the three days is maintained at more than 75% of the indicator value, then it is confirmed as an effective surge. In order to improve accuracy, the data of the first day and the last day of each month are cross-verified with the last day of the previous month and the first day of the next month respectively, to avoid false surges caused by monthly closing. For regions with frequent fluctuations in surge rate, smoothing processing is introduced: use three-day moving average instead of single-day value for judgment.
[0094] S340, based on the region event occurrence time, the regions are sorted in chronological order to obtain the region propagation order of the public health emergency.
[0095] In this embodiment, the chronological order refers to the time order of confirming the occurrence of events in each region; the region propagation order refers to the diffusion path of the emergency event in geographical space, which needs to consider the geographical position relationship between regions. The propagation order not only reflects the time sequence, but also needs to meet the continuity characteristics in space.
[0096] Specifically, first, sort the regions according to the event occurrence time to form an initial propagation sequence. Then verify the spatial continuity, the time interval between adjacent regions should not exceed 5 days, otherwise the data of the region need to be reviewed; for cross-provincial regions, according to the connection situation of traffic trunk lines, the time interval is allowed to be extended to 7 days. In the metropolitan area where the population flows frequently, even non-adjacent regions may also appear synchronous or near-synchronous surges, in which case these regions are classified into the same propagation level. Finally, a hierarchical propagation path diagram is generated: the first level is the first region, and then every 3 days a propagation level is divided until all regions with surges are covered.
[0097] In one embodiment, the wavelet transform uses db4 wavelet function to decompose the overall demand signal for 3 layers, referring to Figure 3 In step S400, the overall demand signal is decomposed by wavelet transform to obtain the regional panic demand component and the actual medical demand baseline, which specifically includes the following steps:
[0098] S410, using db4 wavelet function to perform 3-layer wavelet decomposition on the overall demand signal to obtain high-frequency components and low-frequency components.
[0099] In this embodiment, the db4 wavelet function is an orthogonal wavelet function, and the 3-layer wavelet decomposition means that the signal is decomposed into different frequency components in turn, and each layer of decomposition can obtain an approximation component and a detail component; the high-frequency component corresponds to the short-term rapid fluctuation part, and the low-frequency component corresponds to the long-term change trend.
[0100] Specifically, the daily sales data sequence is first normalized to compress the value range to 0-1. The first layer decomposition obtains high-frequency signals with a frequency greater than 1 / 4 and low-frequency signals with a frequency less than 1 / 4; the second layer decomposition further decomposes the low-frequency signals of the first layer into medium-frequency signals with a frequency of 1 / 8-1 / 4 and low-frequency signals with a frequency less than 1 / 8; and the third layer decomposition obtains three final frequency bands: short-term fluctuation signals (with a frequency greater than 1 / 4), medium-term fluctuation signals (with a frequency between 1 / 8 and 1 / 4), and long-term trend signals (with a frequency less than 1 / 8). The short-term and medium-term fluctuation signals are combined as high-frequency components, and the long-term trend signals are taken as low-frequency components.
[0101] S420, according to the regional event occurrence time, extracting the regional panic demand component in the high-frequency component.
[0102] In this embodiment, the regional panic demand component refers to short-term abnormal demand fluctuation caused by a sudden event; the high-frequency component extraction needs to position and filter the signal in the time domain, and the signal analysis window is determined in combination with the event occurrence time.
[0103] Specifically, first, the event occurrence time is taken as the center, and 5 days are extended forward and 15 days are extended backward to form a 20-day analysis interval. The peak value point of the high-frequency component is searched for in the interval, and the peak value amplitude is required to be more than 2 times the standard deviation of the original signal before decomposition. The found peak value interval is subjected to morphological analysis: the rising slope should be greater than the falling slope, and the peak value duration should not exceed 7 days. The high-frequency signal segment meeting these characteristics is extracted as the panic demand component, and other high-frequency fluctuations are filtered as random fluctuations.
[0104] S430, reconstructing the low-frequency component according to the regional propagation order to obtain an actual medical demand baseline.
[0105] In this embodiment, the actual medical demand baseline reflects the real clinical demand level after excluding the panic factor; the low-frequency component reconstruction needs to consider the propagation relationship between regions to ensure the continuity and rationality of the baseline change.
[0106] Specifically, first, the low-frequency component is subjected to trend analysis to identify the inflection point position and change rate. The regions of adjacent propagation levels should present similar baseline change characteristics, and the change time difference should match the propagation time. When it is found that the low-frequency trend of a region is seriously inconsistent with that of the surrounding regions, a linear interpolation method is used for correction: taking the change trend of the adjacent regions as a reference, the baseline slope of the abnormal region is adjusted. Finally, the corrected low-frequency component and the normal range high-frequency fluctuation are superimposed to obtain the final medical demand baseline.
[0107] In one embodiment, with reference to Figure 4In step S420, according to the regional event occurrence time, the regional panic demand component is extracted from the high-frequency component, including the following steps:
[0108] S421, in the high-frequency component, according to the time position corresponding to the regional event occurrence time, the corresponding signal analysis interval is divided.
[0109] In this embodiment, the signal analysis interval refers to the time window containing the panic demand characteristics, which needs to cover the premonitory period before the event and the panic subsidence period after the event; the time position refers to the corresponding point of the event occurrence time in the signal sequence.
[0110] Specifically, taking the event occurrence time as the reference point, 5 days before the event as the premonitory observation period, and 15 days after the event as the panic subsidence period, an analysis interval with a total length of 20 days is formed. If the analysis interval overlaps with the interval of the adjacent event, the midpoint of the two event occurrence times is taken as the boundary to divide them into different analysis intervals.
[0111] S422, according to the fluctuation characteristics of historical data and the regional market size, the energy threshold is determined, and the signal segment exceeding the energy threshold in the signal analysis interval is extracted as the panic demand candidate interval.
[0112] In this embodiment, the energy threshold refers to a quantitative standard for judging the degree of signal anomaly; the panic demand candidate interval refers to a continuous time period that may contain panic demand characteristics; the fluctuation characteristics include the amplitude and frequency characteristics of the signal.
[0113] Specifically, the root mean square value of the signal in the previous 30 days of the analysis interval is calculated as the reference energy level, and the reference value of large regions is multiplied by 1.2 and the reference value of small regions is multiplied by 0.8 for size correction. Set the sliding window width to 3 days, and calculate the signal energy point by point in the window. When the window energy exceeds 2 times the corrected reference value, mark the window as the start point of the candidate interval; when the energy falls below the reference value, mark it as the end point. If the interval between two candidate intervals is less than 2 days, they are combined into a continuous candidate interval.
[0114] S423, the morphological feature analysis is performed on the panic demand candidate interval, and the morphological features include the amplitude mutation rate, the duration and the decay rate.
[0115] In this embodiment, the amplitude mutation rate describes the rising speed of the signal; the duration indicates the maintenance time of the abnormal state; and the decay rate reflects the demand decline characteristics. These characteristics are used to distinguish panic demand and other types of demand fluctuations.
[0116] Specifically, three characteristic values of each candidate interval are calculated, the amplitude mutation rate is the maximum daily growth rate of the signal from the reference level to the peak value, the duration is the number of days that the signal amplitude is maintained at more than 75% of the peak value, and the decay rate is the average daily amplitude of the signal from the peak value to the reference level. The candidate interval needs to meet the following conditions at the same time: the amplitude mutation rate is greater than 50%, the duration is between 2 and 7 days, and the decay rate is less than half of the mutation rate.
[0117] S424, superimpose the intervals meeting the morphological characteristics to obtain the regional panic demand component.
[0118] In this embodiment, the regional panic demand component refers to the finally confirmed abnormal demand signal; the superposition process needs to maintain the continuity of the signal and eliminate the interference of overlapping areas.
[0119] Specifically, first, arrange the intervals screened through the morphological characteristics in chronological order, and check the overlap of adjacent intervals. For overlapping intervals, keep the interval with a larger amplitude mutation rate; for intervals with a spacing of less than 3 days, realize smooth connection through the cubic spline interpolation method. Finally, integrate the signals of all intervals to form a complete panic demand component sequence.
[0120] In one embodiment, referring to Figure 5 , in step S430, reconstruct the low-frequency component according to the regional propagation order to obtain the actual medical demand baseline, specifically including the following steps:
[0121] S431, based on the regional propagation order, perform time series correlation analysis on the low-frequency components of each region.
[0122] In this embodiment, time series correlation analysis refers to the study of the sequence and mutual influence of low-frequency demand changes in different regions; the low-frequency component reflects the long-term change trend of demand; and the regional propagation order provides a space-time framework for demand conduction.
[0123] Specifically, form analysis pairs by combining adjacent regions in the propagation sequence, and calculate the cross-correlation coefficient of the low-frequency components of each pair of regions. For directly adjacent regions, the cross-correlation coefficient is required to be not less than 0.7; for regions connected by traffic trunk lines, the cross-correlation coefficient should be not less than 0.6. Calculate the time delay of the low-frequency trend of adjacent regions to verify whether it is consistent with the propagation time sequence: generally, the delay of adjacent regions should be within the range of 2-5 days, and regions exceeding this range need to be marked as abnormal region pairs for subsequent correction.
[0124] S432, according to the time series correlation analysis result, reconstruct the trend of the low-frequency component to generate a regional baseline demand curve.
[0125] In this embodiment, trend reconstruction refers to correcting abnormal low-frequency changes based on regional correlation; the regional benchmark demand curve represents reasonable demand change trends; the reconstruction process needs to maintain the continuity of trend changes between regions.
[0126] Specifically, first, determine the reference region: select the region with the highest cross-correlation coefficient and located at the front end of the propagation as the main reference region. Extract the trend feature points of this region, including the inflection point position, change slope, and peak level. According to the propagation timing, these feature parameters are transmitted backward to generate the trend template of other regions. For the previously marked abnormal region pairs, use linear interpolation method to align their trend features, ensuring smooth transition of trend changes between regions.
[0127] S433, calculate the initial medical demand according to the overall demand signal and the regional panic demand component.
[0128] In this embodiment, the initial medical demand refers to the medical drug demand preliminarily separated from the total demand; the calculation process needs to consider the superposition characteristics of demand composition; the overall demand signal contains two main components: medical demand and panic demand.
[0129] Specifically, a simple signal subtraction method is used: subtract the identified panic demand component from the overall demand signal to obtain the preliminarily separated medical demand. To avoid negative values, set the demand lower limit to 30% of the overall demand. Smooth the separation result: use a 5-day moving average window to eliminate short-term fluctuations, but preserve the main trend changes.
[0130] S434, modify the initial medical demand using the regional benchmark demand curve to obtain the actual medical demand baseline.
[0131] In this embodiment, the actual medical demand baseline refers to the final determined real clinical drug demand level; the modification process needs to balance the deviation between the initial calculation result and the benchmark trend.
[0132] Specifically, compare the initial medical demand with the regional benchmark demand curve to calculate the daily deviation rate. When the deviation rate exceeds 20%, use the weighted average method for modification: give the initial demand a weight of 0.6 and the benchmark curve a weight of 0.4. Perform a reasonableness test on the modified result: the daily change should not exceed 15%, and the monthly total amount should match the medical resource supply capacity. Finally, realize the smooth transition of the medical demand baseline in each region through cubic spline interpolation.
[0133] In one embodiment, referring to Figure 6 , in step S500, based on the regional panic demand component and the actual medical demand baseline, calculate the weighted and synthesized dynamic market share of each region, specifically including the following steps:
[0134] S510, calculate the peak ratio and duration ratio of the regional panic demand component relative to the overall demand signal, and determine the panic demand intensity coefficient of each region.
[0135] In this embodiment, the peak ratio refers to the ratio of the panic demand peak value to the total demand peak value; the duration ratio refers to the ratio of the panic demand duration to the total duration of the observation period; and the panic demand intensity coefficient reflects the degree of regional panic.
[0136] Specifically, the panic demand of each region is characterized: first, calculate the peak ratio P = panic demand maximum value / total demand maximum value, which reflects the demand intensity; then calculate the duration ratio D = number of days when the panic demand exceeds the baseline value / total number of days in the observation period (take 30 days), which reflects the persistence. The panic demand intensity coefficient K is calculated by weighting: K = 0.6 × P + 0.4 × D. When the P value exceeds 0.5, it indicates that the panic demand dominates, and the K value increases by 0.1; when the D value exceeds 0.3, it indicates that the duration is long, and the K value increases by 0.05. Finally, the K values of each region are normalized so that the sum is 1.
[0137] S520, extract the growth rate and volatility rate of the actual medical demand baseline, and calculate the medical demand change index of each region.
[0138] In this embodiment, the growth rate reflects the upward trend of demand; the volatility rate represents the stability of demand; and the medical demand change index comprehensively reflects the dynamic characteristics of regional medical demand.
[0139] Specifically, first calculate the growth rate G = (observation period final value - observation period initial value) / observation period initial value, which reflects the overall trend; then calculate the volatility rate V = standard deviation of daily average change rate / average of daily average change rate, which reflects the stability. The medical demand change index M = 0.7 × G + 0.3 × (1-V). For regions with negative growth rate, the G value is recorded as 0; for regions with volatility rate higher than 0.5, the V value is limited within 0.5. It can also be adjusted according to the administrative level, for example, the M value of the provincial capital city is increased by 10%, the M value of the prefecture-level city remains unchanged, and the M value of the county-level city is decreased by 10%.
[0140] S530, calculate the panic demand share and medical demand share of each region according to the panic demand intensity coefficient and the medical demand change index respectively.
[0141] In this embodiment, the panic demand share reflects the market allocation under the influence of the emergency; the medical demand share reflects the market structure under normal conditions; and the calculation of the two shares needs to introduce a regional characteristic correction mechanism. The demand share calculation needs to consider the influence of objective factors such as population density and medical resources.
[0142] Specifically, the panic demand share S1 = K x (1 + population density correction coefficient), wherein the correction coefficient of the area with a population density higher than the average value by 50% is 0.2, the correction coefficient of the area within the average value ± 50% is 0, and the correction coefficient of the area with a population density lower than the average value by 50% is -0.1. The medical demand share S2 = M x (1 + medical resource correction coefficient), wherein the correction coefficient of the area with a number of medical beds per 10,000 people higher than the average value is 0.15, and the correction coefficient of the area with a number of medical beds per 10,000 people lower than the average value is -0.1. The two shares are normalized respectively to ensure that the sum of each is 1.
[0143] S540, the panic demand share and the medical demand share of each area are weighted and synthesized to obtain a dynamic market share of each area.
[0144] In this embodiment, the dynamic market share needs to balance the influence of panic demand and medical demand; the weighted synthesis adopts a time-varying weight scheme to reflect the demand characteristics in different stages. The dynamic adjustment of the market share needs to be adapted to the development stage of the event.
[0145] Specifically, different weight configurations are adopted according to the time process, for example, a weight ratio of 0.7:0.3 is adopted in the initial stage of the event, the dynamic market share = 0.7 x S1 + 0.3 x S2; a weight ratio of 0.5:0.5 is adopted in the development stage of the event, the dynamic market share = 0.5 x S1 + 0.5 x S2; a weight ratio of 0.3:0.7 is adopted in the subsiding stage of the event, the dynamic market share = 0.3 x S1 + 0.7 x S2. The calculation result is normalized to ensure that the sum is 100%, and the market configuration is dynamically adjusted once a week.
[0146] The method for analyzing the market share of a drug based on multi-source heterogeneous data provided in the application realizes the unified collection and standardized processing of different terminal data by establishing a multi-source data collection network, uses wavelet transform technology to decompose the overall demand signal into a panic demand component and a medical demand baseline, and solves the technical problem that the market demand composition is complex and difficult to quantify during a public health emergency. By introducing a dynamic calculation scheme of time-varying weights, the market structure characteristics in different periods are accurately reflected, and the defects that the traditional static analysis method is easily disturbed by short-term fluctuations are avoided. The implementation of the application can provide more accurate market share analysis results for pharmaceutical enterprises, help the enterprises to make accurate market judgments and reasonable resource allocation decisions during an emergency, and improve the emergency response capability of the pharmaceutical supply chain.
[0147] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the application.
[0148] In a second aspect, the application provides a drug market share analysis system based on multi-source heterogeneous data. The drug market share analysis system based on multi-source heterogeneous data is described below in combination with the drug market share analysis method based on multi-source heterogeneous data.
[0149] With reference to Figure 7 A drug market share analysis system based on multi-source heterogeneous data comprises:
[0150] A data collection module is configured to collect drug sales data and obtain drug market sales original data sets in different regions.
[0151] A standardization processing module is configured to perform time alignment and unit standardization processing on the data according to the original data sets, and obtain regional standardized data sets.
[0152] An event detection module is configured to determine the occurrence time and propagation order of a public health emergency in each region based on the standardized data sets and the daily sales surge rate index.
[0153] A demand decomposition module is configured to use the daily drug sales data in each region as the overall demand signal according to the standardized data sets, combine the occurrence time and propagation order of the region, and use wavelet transform to decompose the overall demand signal to obtain a regional panic demand component and an actual medical demand baseline.
[0154] A market share calculation module is configured to calculate the dynamic market share of each region after weighted synthesis based on the regional panic demand component and the actual medical demand baseline.
[0155] A prediction analysis module is configured to calculate the regional corrected market share prediction result based on the dynamic market share of each region and the historical data baseline.
[0156] In one embodiment, the application provides an electronic device, which can be a server, and the internal structure diagram of the electronic device can be as shown in Figure 8 The electronic device comprises a processor, a memory and a network interface connected through a system bus. The processor of the electronic device is configured to provide computing and control capabilities. The memory of the electronic device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The database of the electronic device is configured to store data. The network interface of the electronic device is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement a drug market share analysis method based on multi-source heterogeneous data.
[0157] Those skilled in the art can understand that Figure 8The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0158] In one embodiment, an electronic device is also provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above-mentioned method embodiments when executing the computer program.
[0159] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing relevant hardware. The above-mentioned computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not as a limitation, RAM can be in various forms such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0160] The above are preferred embodiments of the present application, which do not limit the protection scope of the present application, therefore: any equivalent changes made on the structure, shape, principle of the present application shall be covered within the protection scope of the present application.
Claims
1. A method for analyzing drug market share based on multi-source heterogeneous data, characterized in that, Includes the following steps: Collect drug sales data to obtain raw datasets of drug market sales in different regions; Based on the original dataset, time alignment and unit standardization are performed on the data to obtain a standardized dataset divided into regions. Based on the standardized dataset and the daily average sales surge rate index, the occurrence time and spread sequence of public health emergencies in each region are determined, and the regional event occurrence time and regional spread sequence are obtained. Based on the standardized dataset, the daily drug sales data of each region are used as the overall demand signal. Combining the occurrence time and propagation order of the region, wavelet transform is used to decompose the overall demand signal to obtain the regional panic demand component and the actual medical demand baseline. Based on the regional panic demand components and the actual medical demand baseline, calculate the dynamic market share of each region after weighted synthesis. Based on the dynamic market share and historical data baseline of each region, the regional adjusted market share forecast results are calculated. The wavelet transform employs the db4 wavelet function to perform a three-level decomposition of the overall demand signal. This wavelet transform decomposes the overall demand signal to obtain regional panic demand components and the baseline of actual medical demand. Specifically, the steps include: The overall demand signal is decomposed into high-frequency and low-frequency components using the db4 wavelet function. Based on the time of occurrence of the regional events, regional panic demand components are extracted from the high-frequency components; Based on the propagation order of the region, the low-frequency component is reconstructed to obtain the baseline of actual medical needs; Specifically, based on the occurrence time of the regional event, the regional panic demand component is extracted from the high-frequency components, including the following steps: In the high-frequency components, the corresponding signal analysis intervals are divided according to the time position corresponding to the occurrence time of the regional event; Based on the fluctuation characteristics of historical data and the size of regional markets, an energy threshold is determined, and signal segments exceeding the energy threshold are extracted within the signal analysis interval as candidate intervals for panic demand. The candidate intervals for panic demand are subjected to morphological feature analysis, and the morphological features include amplitude mutation rate, duration and decay rate. The regional panic demand components are obtained by superimposing the intervals that conform to the morphological characteristics described. The process of reconstructing the low-frequency components based on the regional propagation order to obtain the baseline of actual medical needs includes the following steps: Based on the propagation order of the regions, a time-series correlation analysis is performed on the low-frequency components of each region. Based on the time-series correlation analysis results, the low-frequency components are reconstructed to generate a regional baseline demand curve; Calculate the initial medical demand based on the overall demand signal and the regional panic demand components; The initial medical demand is corrected using the regional baseline demand curve to obtain the actual medical demand baseline.
2. The method for analyzing drug market share based on multi-source heterogeneous data according to claim 1, characterized in that, Based on the standardized dataset and the daily average sales surge rate indicator, the timing and sequence of transmission of public health emergencies in various regions are determined, specifically including the following steps: A sales fluctuation feature model is constructed based on historical datasets, and regional weight coefficients are set in combination with regional market size to determine the daily average sales surge rate indicator. Calculate the growth rate of daily drug sales in each region relative to the average sales of the previous N days to obtain the daily average sales surge rate sequence for each region, where 3≤N≤7. Based on the daily average sales surge rate sequence and the daily average sales surge rate index, the time point when each region first exceeds the index is selected as the regional event occurrence time. Based on the time of occurrence of the events in the regions, the regions are sorted in chronological order to obtain the regional transmission sequence of the public health emergency.
3. The method for analyzing drug market share based on multi-source heterogeneous data according to claim 1, characterized in that, Based on the aforementioned regional panic demand components and the baseline of actual medical demand, the weighted composite dynamic market share for each region is calculated, specifically including the following steps: Calculate the peak ratio and duration ratio of the regional panic demand component relative to the overall demand signal to determine the panic demand intensity coefficient for each region; Extract the growth rate and volatility of the actual medical demand baseline, and calculate the medical demand change index for each region; The panic demand share and medical demand share of each region are calculated based on the panic demand intensity coefficient and the medical demand change index, respectively. The dynamic market share of each region is obtained by weighting and combining the share of panic demand and medical demand in each region.
4. A drug market share analysis system based on multi-source heterogeneous data, characterized in that, The method for analyzing drug market share based on multi-source heterogeneous data according to any one of claims 1-3 includes: The data acquisition module is used to collect drug sales data and obtain raw datasets of drug market sales in different regions. The standardization processing module is used to perform time alignment and unit standardization on the original dataset to obtain a standardized dataset by region. The event detection module is used to determine the occurrence time and spread sequence of public health emergencies in various regions based on the standardized dataset and the daily average sales surge rate index. The demand decomposition module is used to decompose the overall demand signal by taking the daily drug sales data of each region as the overall demand signal based on the standardized dataset, and combining the occurrence time and propagation order of the region with wavelet transform to obtain the regional panic demand component and the actual medical demand baseline. The market share calculation module is used to calculate the dynamic market share of each region after weighted synthesis based on the regional panic demand component and the actual medical demand baseline. The predictive analysis module is used to calculate the regionally corrected market share prediction results based on the dynamic market share of each region and historical data baseline.
5. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the pharmaceutical market share analysis method based on multi-source heterogeneous data as described in any one of claims 1-3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the drug market share analysis method based on multi-source heterogeneous data as described in any one of claims 1-3.
Citation Information
Patent Citations
Medical sales prediction system and method based on IPSO-LSTM model
CN115760210A
Inventory management method and system for drug demand quantity prediction based on big data analysis
CN118380120A