Potential customer prediction system, potential customer prediction method, potential customer prediction program

The potential customer prediction system addresses the challenge of predicting early adopters for new drugs by stratifying customers based on attribute and activity information, determining early adopters, and estimating adoption probabilities, thereby optimizing sales and marketing efforts and recovering development costs efficiently.

JP7798423B1Active Publication Date: 2026-01-14TCROSS INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2025567549
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-10-29
Publication Date
2026-01-14
Estimated Expiration
2045-10-29

AI Technical Summary

Technical Problem

Pharmaceutical companies face challenges in predicting potential customers for new drugs before or shortly after launch, as there is a lack of accumulated market data for new drugs, making it impossible to use conventional analytical models, and they must rely on experience and intuition.

Method used

A potential customer prediction system that includes data acquisition, stratification of customers based on attribute and activity information, determination of early adopters using factor analysis and binary classification algorithms, and estimation of adoption probabilities.

Benefits of technology

Enables highly efficient allocation of sales and marketing resources to high-probability customer groups, facilitating rapid recovery of research and development costs and improving profitability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798423000003
    Figure 0007798423000003
  • Figure 0007798423000004
    Figure 0007798423000004
  • Figure 0007798423000005
    Figure 0007798423000005
Patent Text Reader

Abstract

To provide a potential customer prediction system capable of predicting potential customers who are likely to become early adopters of new drugs. The potential customer prediction system 1 includes a calculation unit 6 that controls the potential customer prediction system 1, and a storage unit 7. The calculation unit 6 includes a data acquisition unit 2 that acquires a plurality of pieces of attribute information J1 and a plurality of pieces of activity information J2 related to each of a plurality of doctors collected before the launch of a new drug, a stratification unit 3 that stratifies the plurality of doctors based on the plurality of pieces of attribute information J1 and the plurality of pieces of activity information J2, an early adopter determination unit 4 that determines whether each of the stratified doctors is an early adopter EA, and an adoption probability estimation unit 5 that estimates an adoption probability P that indicates the probability that each of the plurality of doctors is an early adopter using a binary classification algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a system for predicting potential customers likely to adopt a new drug before or during its early stages of launch. [Background technology]

[0002] Conventionally, there are known technologies that use past market data to perform qualitative and quantitative analysis of pharmaceutical products with a proven sales track record and use the analysis results to optimize sales and marketing activities. For example, Patent Document 1 introduces a technology that predicts the probability of a target customer adopting a pharmaceutical product using the attributes of the target customer and the history of the pharmaceutical company's past sales and marketing activities. Furthermore, Patent Documents 2 and 3 introduce technologies that stratify customers based on their attributes and behavioral history, and identify customer segments that are likely to be interested in a particular pharmaceutical product. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 7636838 [Patent Document 2] Japanese Patent Application Laid-Open No. 2024-129535 [Patent Document 3] International Publication No. 2024 / 190678 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, pharmaceutical companies have contributed to improving the prognosis and quality of life (QOL) of many patients by developing and launching new drugs one after another. However, the research and development of new drugs requires a huge amount of time and cost. Furthermore, due to the increasing complexity of drug discovery targets, the expansion of clinical trial scale, and the sophistication of regulatory requirements, the research and development costs of new drugs have been increasing year by year. For this reason, pharmaceutical companies have been seeking to streamline sales promotion in order to recover research and development costs as quickly as possible and enable sustainable research and development activities.

[0005] The target customers for new drugs are so-called early adopters, those who will use or prescribe the new drug early. However, for new drugs before or immediately after their launch, market data such as sales and marketing activities, adoption records, and prescription records has not been accumulated. For this reason, it has been impossible to use conventional analytical models that use market data to estimate target customers, and companies have had to rely on experience and intuition.

[0006] Therefore, an object of the present disclosure is to provide a potential customer prediction system that can predict potential customers who are likely to become early adopters. [Means for solving the problem]

[0007] The potential customer estimation system disclosed herein comprises a calculation means, the calculation means comprising: a data acquisition means for acquiring a plurality of attribute information and a plurality of activity information regarding each of a plurality of customers collected before the new drug is launched; a stratification means for stratifying the plurality of customers based on the plurality of attribute information and the plurality of activity information; an early adopter determination means for determining whether each of the stratified customers is an early adopter; and an adoption probability estimation means for estimating an adoption probability indicating the probability that each of the plurality of customers is an early adopter using a binary classification algorithm, wherein the stratification means includes a factor analysis means for quantifying the attribute information and activity information to generate observed variables, extracting common factors that commonly contribute to the observed variables and idiosyncratic factors that uniquely contribute to the observed variables, and calculating the contributions of the common factors and idiosyncratic factors to the observed variables as factor scores, and the early adopter determination means calculates a total value of the factor scores for each customer, generates a factor score distribution that indicates the distribution of the number of customers for each total value, and determines early adopters using the factor score distribution. [Effects of the Invention]

[0008] The potential customer estimation system disclosed herein stratifies multiple customers based on their attribute information and activity information, determines whether each stratified customer is an early adopter, and calculates the probability of them adopting a new drug. This allows for highly efficient investment of limited sales and marketing resources in customer groups with a high probability of becoming early adopters. This allows for rapid recovery of enormous research and development costs in the early stages of a new drug's life cycle, leading to healthy cash flow and improved profitability. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram of a potential customer prediction system according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a schematic diagram showing the configuration of information used in a potential customer prediction system. [Figure 3] FIG. 1 is a schematic diagram of a factor analysis model. [Figure 4]FIG. 10 is a table showing a factor score list. [Figure 5] FIG. 1 is a schematic diagram showing factor score distribution and early adopter criteria. [Figure 6] FIG. 1 is a conceptual diagram of a binomial logistic regression analysis model. [Figure 7] FIG. 10 is a table showing a potential customer list. [Figure 8] FIG. 1 is a schematic diagram showing the distribution of early adopters and estimated probabilities. [Figure 9] 1 is a flowchart showing the flow of processing of a potential customer prediction system. DETAILED DESCRIPTION OF THE INVENTION

[0010] An embodiment of a potential customer prediction system embodying the present disclosure will be described below with reference to the accompanying drawings. In the following description, a "doctor" is used as an example of a "customer," and a "new drug" is used as an example of a "new product." Furthermore, a customer may be not only a "doctor," but also a "medical institution."

[0011] As shown in FIGS. 1 and 2, the potential customer prediction system 1 includes a calculation unit 6 and a storage unit 7.

[0012] The calculation unit 6 includes a data acquisition unit 2 that acquires multiple pieces of attribute information J1 and multiple pieces of activity information J2 regarding each of multiple doctors collected before the new drug is launched, a stratification unit 3 that stratifies the multiple doctors based on the multiple pieces of attribute information J1 and the multiple pieces of activity information J2, an early adopter determination unit 4 that determines whether each of the stratified doctors is an early adopter EA, and an adoption probability estimation unit 5 that estimates an adoption probability P that indicates the probability that each of the multiple doctors is an early adopter EA using a binary classification algorithm.

[0013] The data acquisition unit 2 acquires attribute information J1 and activity information J2 related to multiple doctors who are sales and marketing targets. The attribute information J1 and activity information J2 are information that can be extracted from an existing customer list created based on sales data from the sale of existing pharmaceuticals. The attribute information J1 includes, for example, ID, name, gender, age, prefecture, facility name, medical department, affiliated facility, job title, etc. The activity information J2 includes, for example, conference attendance history, paper publication history, lecture history, etc.

[0014] The stratification unit 3 stratifies multiple doctors using attribute information J1 and activity information J2. "Stratification" refers to classifying doctors with the same characteristics and properties and treating them as a homogeneous group. One method for stratifying multiple doctors is factor analysis (maximum likelihood method). The stratification unit 3 performs factor analysis using a factor analysis model M1. The stratification unit 3 can extract groups of doctors with potentially high interest from multivariate data such as doctor attribute and activity information using not only factor analysis but also a method that can extract similarities and latent structures among doctors from a customer list. The stratification unit 3 can include models for implementing principal component analysis, cluster analysis, self-organizing maps (SOM), latent Dirichlet allocation (LDA), and other statistical or machine learning methods, such as cluster analysis models and principal component analysis models, or analytical models that combine these.

[0015] As shown in Figure 3, the factor analysis model M1 consists of an observed variable V1, common factors (F1, F2) that contribute to multiple observed variables V1 in common, and an intrinsic factor (error term E) that contributes uniquely to each observed variable V1. In the example of Figure 3, the error term E that contributes uniquely to each of the observed variables V1-1 to V1-5 is shown as error terms E1 to E5. The common factor and intrinsic factor are latent factors underlying the observed variable V1 and explain the variance structure of the observed data. In the factor analysis model M1, the relationship between each observed variable and the common factor is calculated as a factor loading, and factor scores FS (expertise factor score FS1, activity factor score FS2, error term score ES) for each physician for each common factor are calculated based on the factor loadings. The factor scores FS are statistically quantified latent characteristics (expertise, activity, etc.) estimated for the observed variables V1-1 to V1-5.

[0016] The observed variable V1 is a quantification of the attribute information J1 and activity information J2, and serves as basic data for determining early adopter EA. The observed variable V1 includes, for example, academic positions, domestic and international publications, conference presentations, the number of clinical cases, and other professional activities. While the observed variable V1 can be modified as needed, here we will explain it using the observed variable V1-1 as "academic positions," the observed variable V1-2 as "number of domestic and international publications," the observed variable V1-3 as "conference presentations," the observed variable V1-4 as "number of clinical cases," and the observed variable V1-5 as "other activities." The academic positions are the experience of serving as an officer or secretary in academic societies related to the target disease. The number of paper presentations is the number of papers presented as a lead author or co-author in Japan and overseas. The conference presentations are the experience of presenting at conferences in Japan and overseas. In quantifying the attribute information J1 and activity information J2, for example, "male" can be set to "0" and "female" to "1." The quantification can be performed manually by the user or automatically by the system using preset values, etc.

[0017] The common factors include, for example, an Expertise Factor F1 and an Activity Factor F2. The arrow pointing from the common factor to the observed variable V1 represents the structure of the contribution from the common factor to the observed variable V1. Specifically, the Expertise Factor F1 mainly contributes to academic position history, number of domestic and international paper publications, and academic conference presentation history, while the Activity Factor F2 mainly contributes to academic conference presentation history, number of facility cases, and other activity achievements. Because the contributions of multiple common factors intersect for academic conference presentation history, cross-loadings are added.

[0018] The error term E is the unique variance of each observed variable V1. The arrow pointing from the error term E to the observed variable V1 represents the structure of the contribution from the error term E to the observed variable V1.

[0019] The factor analysis model M1 is expressed by Equation 1.

[0020]

number

[0021] x is a p×1 vector of observed variables V1, Λ is a p×m factor loading matrix, f is an m×1 vector of common factors, and ε is a p×1 vector of unique factors. The number of factors m is set based on the analysis conditions. The stratification unit 3 extracts common factors and unique factors and then calculates the factor score FS.

[0022] As shown in FIG. 4, the stratification unit 3 generates a factor score list L1 in which attribute information J1 and activity information J2 are linked to factor scores FS.

[0023] 5, the stratification unit 3 calculates the sum of the factor scores FS for each doctor and generates a factor score distribution G that represents the number of doctors for each sum. The factor score distribution G roughly follows a normal distribution.

[0024] The early adopter determination unit 4 determines doctors who fall within the top threshold % of the factor score distribution G as early adopters (EA), and determines other doctors as non-early adopters (NEA). The threshold % can be set to 10-20%, preferably 16%. In this example, factor scores of 0.82 or higher correspond to the top 16%.

[0025] Here, the "top 16%" is based on the "Innovator Theory" of American sociologist Professor Everett M. Rogers. In this disclosure, the "top 16%" is based on the 16% calculated by combining "Innovators (not shown) = 2.5%" and "Early Adopters (EA) = 13.5%" in the "Innovator Theory."

[0026] The adoption probability estimation unit 5 uses a binary classification algorithm to estimate the probability (adoption probability P) that each of the multiple doctors is an early adopter EA. An example of a binary classification algorithm is the binomial logistic regression analysis model M2. The adoption probability estimation unit 5 is only required to find a group that is highly similar to the extracted group of doctors and estimate the adoption probability P, and the specific algorithm can be freely selected as long as it can realize this technical concept. The adoption probability estimation unit 5 can include not only the binomial logistic regression analysis model M2, but also models for implementing binary classification algorithms such as eXtreme Gradient Boosting (XGBoost), random forests, support vector machines (SVMs), neural networks, etc., or models that combine these.

[0027] The binomial logistic regression analysis model M2 estimates the adoption probability based on the objective variable V3 and the explanatory variable V2. The objective variable V3 stores the determination result (0: non-early adopter NEA, 1: early adopter EA) by the early adopter determination unit 4. The explanatory variable V2 stores the attribute variable V2.

[0028] The binomial logistic regression analysis model M2 is expressed by Equation 2.

[0029]

number

[0030] Equation 2 is a logit function that represents the probability pi (adoption probability P) that doctor i is an early adopter EA. β0 is the intercept, β j (j=1, 2, ..., k) is the regression coefficient of each explanatory variable V2. The hiring probability estimation unit 5 calculates the hiring probability pi of doctor i as a percentage based on the estimation formula obtained here.

[0031] Figure 6 shows a graph that illustrates the concept of the binomial logistic regression analysis model M2. The horizontal axis of the graph represents the explanatory variable V2, and the vertical axis represents the adoption probability P.

[0032] The graph of the binomial logistic regression analysis model M2 shows that the adoption probability increases nonlinearly as the value of the explanatory variable V2 increases. The dashed line on the graph of the binomial logistic regression analysis model M2 represents an adoption probability of P=0.5. The adoption probability estimation unit 5 can classify products into Class 0 (non-early adopters NEA) and Class 1 (early adopters EA) using the adoption probability P=0.5 as the boundary.

[0033] 7, the hiring probability estimation unit 5 generates a potential customer list L2 that links attribute information J1 with high-probability doctors (objective variable V3), hiring probability P, and explanatory variable V2. The potential customer list L2 can also include observed variables V1 and latent variables (not shown) that are statistical features extracted by factor analysis or the like.

[0034] The calculation unit 6 is configured to be able to output the factor score list L1 and the potential customer list L2. For example, the calculation unit 6 can output the hiring probability P to a file in Excel (registered trademark) format or other spreadsheet format.

[0035] As shown in Figure 8, the calculation unit 6 not only outputs the hiring probability P as a numerical value, but also combines it with attribute information J1 by region, facility size, medical department, etc. to generate graphs and tables, or generate mapping data superimposed on a map, etc.

[0036] The calculation unit 6 can be configured to plot the distribution of the number of early adopter EAs and the distribution of the adoption probability P as a choropleth map using geospatial data in GeoJSON (Geographic JavaScript Object Notation) format. A choropleth map is a map that displays statistical values ​​for each region in different colors to visually distinguish them. The calculation unit 6 can obtain the locations included in the doctor attribute information J1 as geospatial data and generate visualized data using GeoJSON format, a standard format for geographic information systems (GIS). The GeoJSON format is a general-purpose data structure that describes the boundaries of regions and facilities as coordinate information, and has the advantage of being easy to draw using a web browser or visualization tool.

[0037] The calculation unit 6 uses geospatial data in GeoJSON format to plot a choropleth map that indicates the adoption probability P by region using different shades of color. In the present disclosure, doctors whose adoption probability P indicates a certain percentage (for example, 50% or more) are highlighted as a high-probability group. This allows users to intuitively grasp the concentration of early adopter EAs in a specific region and clarify the regions where sales and marketing resources should be prioritized.

[0038] Furthermore, the calculation unit 6 can be configured to retrain the binary classification algorithm using actual prescription data and sales data (hereinafter, prescription history information J3) collected after the new drug is launched as additional training data. Specifically, after the initial adoption probability P is estimated, the data acquisition unit 2 acquires the prescription history information J3, and the adoption probability estimation unit 5 generates a new explanatory variable V2 based on the prescription history information J3, reconfigures the binary classification algorithm, and re-estimates the adoption probability P. The prescription history information J3 can include the number of sales visits, the number of online symposium participations, etc. This configuration makes it possible to build a re-training model that reflects the latest physician behavior, and is expected to improve the time-series accuracy of the binary classification algorithm.

[0039] In other words, before a new drug is launched or in the early stages of its launch, physicians in the high-probability group are identified and approached with priority. After a certain period of time has passed since the initial launch and prescription performance information J3 has been accumulated, it will be possible to integrate sales and marketing activity data in addition to the estimated adoption probability P. Therefore, in the early stages of launch, targets will be selected using analysis based on this disclosure alone, and after launch, the data obtained through this disclosure can be combined with sales and marketing activity data to create an expanded model that includes new activity information J2. This expanded model will serve as a foundation for verifying the effectiveness and optimal allocation of activities, further enhancing the usefulness of sales and marketing.

[0040] For example, one year after a drug is launched, it will be possible to grasp both the activity history and sales data for each doctor, which will enable development from early adopter assessments at the initial stage to a sophisticated prescription probability model based on actual activity effects and sales performance. This will enable consistent strategic use from the prediction model at the early stage of a drug's launch to the performance model at its mature stage.

[0041] This re-learning model can be used in combination with the "optimal sales activity promotion model" currently under patent application. This makes it possible to input the adoption probability P updated through re-learning and automatically recommend optimal sales activity content (frequency of visits, means of providing information, invitations to lectures, etc.) for each doctor. This configuration makes it possible to improve the efficiency of sales and marketing activities and quickly recover research and development costs throughout the entire period from the early stage of market launch to the mature stage.

[0042] This disclosure features high connectivity with other analytical models and sales support systems, allowing for easy integration into existing corporate systems. For example, by combining the potential customer prediction system 1 of this disclosure with the probabilistic model of Patent No. 7636838 (a model that estimates the adoption probability P for each physician or facility), the results of the early adopter determination unit 4 can be input into the probabilistic model as initial parameters, further improving the accuracy of estimating the adoption probability after market launch. This also enables dynamic updates according to data while maintaining consistent prediction logic from the early stage of new drug introduction to the mature stage.

[0043] The functions of the present disclosure are realized, for example, by an information processing device that includes a central processing unit (CPU), a main memory (RAM), a non-volatile memory (ROM, flash memory, hard disk, etc.), an input / output interface, a display device, an input device, and a network interface.

[0044] The CPU of the information processing device functions as the calculation unit 6 of the present disclosure. The CPU reads and executes programs and data stored in the storage device, and performs various processes related to the present disclosure, namely, a data acquisition step by the data acquisition unit 2, a stratification step such as factor analysis by the stratification unit 3, an early adopter determination step by the early adopter determination unit 4, an adoption probability estimation step by the adoption probability estimation unit 5, and a drawing step and an output step by the calculation unit 6. In other words, the CPU is configured to read and execute a potential customer prediction program stored in the storage device. It is also possible to implement a potential customer prediction program that functions as the calculation unit 6 in the CPU.

[0045] The storage device of the information processing device functions as the storage unit 7 of the present disclosure. It stores program code, attribute information J1, activity information J2, observed variables V1, factor scores FS, early adopter EA, objective variable V3, explanatory variables V2, parameters of the factor analysis model M1, and parameters of the binomial logistic regression analysis model M2. The input / output interface enables data exchange with external storage and other information processing devices, and the display device displays analysis results and visualized graphs. The input device consists of a keyboard, mouse, touch panel, etc., and is used by users to set analysis conditions and thresholds. The network interface is used to connect to medical institutions and external databases and for processing in a cloud environment.

[0046] The present disclosure may be executed in a standalone operating environment such as a stand-alone personal computer or tablet terminal, or in a client-server system in which multiple information processing devices cooperate via a network, or in a cloud computing environment. For example, the program of the present disclosure may be implemented on an in-house server of a pharmaceutical company, and accessed from a terminal in the company's sales and marketing department to perform the early adopter determination process and the adoption probability estimation process.

[0047] This disclosure may also be provided as a cloud service. In this case, physician attribute information J1 and activity information J2 are transmitted to a cloud server via a secure network, and the stratification process, early adopter determination process, adoption probability estimation process, rendering process, and output process are executed sequentially on the cloud. Users can access the results using a browser or dedicated application and change analysis conditions such as filtering and sorting as needed.

[0048] Furthermore, the present disclosure may be implemented as modular software and integrated with an existing sales force automation (SFA) or customer relationship management (CRM) system. In this case, the algorithm of the present disclosure processes attribute information J1 and activity information J2 provided by the existing system as input, and displays a potential customer list L2 on a dashboard, providing information that can be immediately used by field sales representatives.

[0049] The model disclosed herein can be flexibly configured according to the execution environment and the user's operational style. For example, in a standalone environment, it can be linked to an on-premise database and complete processing without external communication. On the other hand, in a cloud environment, data can be input and results can be shared in real time from multiple locations, facilitating nationwide market analysis and strategy formulation.

[0050] The present disclosure may be provided as a program. The program includes a set of instructions for causing a computer to execute a series of processes according to the present disclosure, i.e., a data acquisition process, a stratification process, an early adopter determination process, an adoption probability estimation process, a rendering process, and an output process, on an information processing device. The program may be provided in a form recorded on a tangible recording medium (such as a CD-ROM, DVD, Blu-ray (registered trademark) Disc, a magnetic disk, or a semiconductor memory) or in a form downloadable via a network such as the Internet. Furthermore, the program may be provided as a service in a cloud environment, and may be configured to be accessed by a user via a browser or a dedicated client.

[0051] Next, the processing flow of the potential customer prediction system 1 configured as above will be described with reference to FIG.

[0052] The data acquisition unit 2 acquires attribute information J1 and activity information J2 of multiple doctors (S1). The stratification unit 3 quantifies the attribute information J1 and activity information J2 to generate an observation variable V1 (S2). The stratification unit 3 extracts an expertise factor F1, an activity factor F2, and an error term E, and calculates the contribution of the expertise factor F1, the activity factor F2, and the error term E to the observation variable V1 as a factor score FS (S3). The calculation unit 6 generates and outputs a factor score list L1 (S4).

[0053] The early adopter determination unit 4 generates a factor score distribution G and determines doctors who are in the top 16% of the factor score distribution G as early adopters EA (S5). The adoption probability estimation unit 5 uses whether multiple doctors are early adopters as a dependent variable V3 and an observed variable V1 as an explanatory variable V2, and estimates an adoption probability P based on the dependent variable V3 and the explanatory variable V2 (S6). The calculation unit 6 generates and outputs a potential customer list L2 (S7).

[0054] Therefore, according to the potential customer prediction system 1 of this embodiment, the factor score FS is used to determine early adopter EA, so that it is possible to statistically determine the group of doctors who are highly interested in a new drug before or in the early stages of its market launch without using past market data.

[0055] Next, three embodiments of a chronic disease, a rare disease, and an acute disease will be described as application examples of the present disclosure.

[0056] An example of a case involving a chronic disease will be explained. Company X, a pharmaceutical company, is developing a new drug for the treatment of diabetes. The new drug has a new mechanism of action different from existing drugs, and good efficacy has been reported in domestic clinical studies. Before the drug was released to the market, Company X faced the issue of "which doctors at which facilities should be given priority in providing information and having them introduce the new drug?"

[0057] The target doctors were those in the diabetes department, endocrinology and metabolism department, and clinics. The company had a customer list that it had accumulated through past sales activities, but for products with no past track record, such as new drugs, there was no data or model to quantitatively determine which doctors would be early adopters.

[0058] Therefore, Company A decided to predict promising potential customers using the potential customer prediction system 1 disclosed herein. The data acquisition unit 2 of the potential customer prediction system 1 acquired attribute information J1 and activity information J2 from a customer list (a list of multiple doctors related to the diabetes field) held by Company A. A factor analysis was performed using factor analysis model M1 to determine early adopters EA.

[0059] The adoption probability estimation unit 5 calculated the adoption probability P using the binomial logistic regression analysis model M2 with early adopters EA as the objective variable V3, and found that the adoption probability P for the top 50 people was 0.75 or higher. Based on the potential customer list L2, Company A created a list of doctors that should be approached intensively in the early stages of market launch and formulated a visiting plan for sales representatives. In fact, doctors estimated to be in the high probability group adopted the new drug early, and market penetration accelerated through case reports and exchanges of opinions both inside and outside the facility.

[0060] An example of implementation targeting rare diseases will be described. Company B, a pharmaceutical company, is developing a new drug to treat rare diseases. The company's sales background was that it had not previously been able to fully grasp the nationwide doctor network or prescription trends in the rare disease field, making it difficult to determine where to focus its existing sales resources. Therefore, the company decided to introduce the potential customer prediction system 1 disclosed herein in order to carry out effective targeting before the drug is launched.

[0061] First, we constructed a list of doctors specializing in rare disease A as target customers. This list included attribute information J1 and activity information J2, such as past clinical research participation, conference presentations on related diseases, number of domestic and international papers written, number of cases, attendance at rare disease treatment facilities, and KOL activity history. The data acquisition unit 2 of the potential customer prediction system 1 acquired attribute information J1 and activity information J2 from Company A's customer list (a list of multiple doctors related to the diabetes field). The stratification unit 3 quantified attribute information J1 and activity information J2 to generate observed variable V1, and then performed factor analysis by applying observed variable V1 to factor analysis model M1. Specifically, factor analysis model M1 extracted a factor representing "disease expertise" (expertise factor F1) and a factor representing "academic activity" (activity factor F2), and calculated a factor score FS for each of the multiple doctors. The early adopter determination unit 4 determined early adopters EA based on the factor score distribution G.

[0062] The adoption probability estimation unit 5 calculated the adoption probability P using the binomial logistic regression analysis model M2 with early adopter EA as the objective variable V3, and found that the adoption probability P for the top 50 people was 0.80 or higher.

[0063] Based on the analysis results, Company B formulated a sales plan that prioritized visiting physicians in the high-risk group. After the drug was launched, these physicians quickly adopted the new drug, and information spread throughout the rare disease treatment network through case sharing and lectures. As a result, efficient product distribution and improved patient access were achieved even in a limited market environment.

[0064] An example of an acute disease will be described. Company C, a pharmaceutical company, was developing a new drug for acute disease B, which had a new mechanism of action that quickly suppressed severe attacks. In the acute disease field, the speed at which treatment can be initiated is directly related to the prognosis, so it was extremely important to deliver the product to appropriate doctors and medical institutions quickly from the early stages of its launch.

[0065] The sales background was that medical treatment for acute diseases is concentrated in specific medical institutions and facilities with emergency response capabilities, and it was thought that doctors affiliated with these facilities would play an important role in the initial spread of product adoption. However, Hei Pharmaceutical did not have a detailed understanding of prescription trends at emergency response facilities nationwide, and faced challenges in optimally allocating sales resources. Therefore, the company decided to introduce the potential customer prediction system 1 disclosed herein and select priority targets before launching the product.

[0066] First, a list of doctors at hospitals with emergency response facilities or intensive care units (ICUs) that can treat target disease B was created. The data acquisition unit 2 of the potential customer prediction system 1 acquired the doctor's attribute information J1 as well as activity information J2, such as the number of past cases of acute disease B or related diseases, presentation history at academic conferences in the emergency department or intensive care department, number of papers in related fields, history of participation in clinical trials, and history of involvement in creating in-hospital treatment protocols. The stratification unit 3 quantified the attribute information J1 and activity information J2 and performed factor analysis by applying the observed variable V1 to the factor analysis model M1.

[0067] The factor analysis model M1 extracted the expertise factor F1, which represents "experience in acute care," and the activity factor F2, which represents "academic communication ability," and calculated a factor score FS for each of multiple doctors. The early adopter determination unit 4 determined early adopters EA based on the factor score distribution G.

[0068] The adoption probability estimation unit 5 calculated the adoption probability P using the binomial logistic regression analysis model M2 with early adopter EA as the objective variable V3, and showed that doctors belonging to a specific emergency response facility are likely to initially adopt the new drug with an adoption probability of P = 85% or more.

[0069] Based on the results of this analysis, Company C developed a priority visit plan for doctors in the high-risk group and provided information about the new drug and support for its introduction from the early stages of its launch. As a result, early access to patients with acute disease B was ensured, evaluation of the product in clinical settings was rapidly accumulated, and its spread to the nationwide emergency medical network was accelerated.

[0070] In this way, by utilizing the potential customer prediction system 1 disclosed herein, pharmaceutical companies can identify promising customers with a statistically high adoption probability P without relying on past performance in the early stages of a new drug's launch, dramatically improving the accuracy of their sales and marketing activities. While traditionally, sales representatives often decided who to visit based on their experience and intuition, selecting targets based on objective and reproducible assessment results provided by this disclosure reduces waste in sales activities. Furthermore, by prioritizing visits to doctors in the high-probability group, efficient market penetration can be achieved from the early stages of a drug's launch, making it possible to make the most of limited human and time resources.

[0071] The present disclosure is not limited to the above embodiments, and the type of input data, schema, output format, user interface, system configuration, etc. can be appropriately modified depending on the purpose of use and environment, as long as it does not deviate from the spirit of the present disclosure. [Explanation of symbols]

[0072] 1. Potential customer prediction system 2 Data Acquisition Section 3 Stratification Department 4 Early Adopter Judgment Department 5. Recruitment probability estimation section 6 Arithmetic section 7 Memory section J1 attribute information J2 activity information J3 Prescription performance information V1 Observed variables V2 explanatory variables V3 Response variable F1 Expertise Factor F2 active factor E error term M1 factor analysis model M2 binomial logistic regression model FS Factor Score G-factor score distribution EA Early Adopters NEA Non-Early Adopters L1 factor score list L2 Potential Customer List

Claims

1. A potential customer prediction system including a calculation means, the calculation means includes: a data acquisition means for acquiring a plurality of pieces of attribute information and a plurality of pieces of activity information relating to each of a plurality of customers collected before the new product is launched; a stratification means for stratifying the plurality of customers based on the plurality of pieces of attribute information and the plurality of pieces of activity information; an early adopter determination means for determining whether each of the stratified customers is an early adopter; and an adoption probability estimation means for estimating an adoption probability indicating the probability that each of the plurality of customers is an early adopter using a binary classification algorithm; the stratification means includes factor analysis means that quantifies the attribute information and the activity information to generate observed variables, extracts common factors that commonly contribute to the observed variables and unique factors that uniquely contribute to the observed variables, and calculates the contributions of the common factors and the unique factors to the observed variables as factor scores; The potential customer prediction system is characterized in that the early adopter determination means calculates a total value of the factor scores for each customer, generates a factor score distribution that indicates the distribution of the number of customers for each of the total values, and determines the early adopters using the factor score distribution.

2. The potential customer prediction system according to claim 1 , wherein the stratification means includes a factor analysis model, a cluster analysis model, a principal component analysis model, or an analysis model that is a combination of these.

3. The potential customer prediction system of claim 1 , wherein the binary classification algorithm comprises a binary logistic regression model, gradient boosting, random forest, support vector machine, neural network, or a model that is a combination thereof.

4. 2. The potential customer prediction system according to claim 1, wherein the early adopter determination means determines customers in the top 10 to 20% of the factor score distribution as early adopters.

5. The potential customer prediction system according to claim 1 , wherein the calculation means plots the distribution of the number of early adopters as a choropleth map using geospatial data in GeoJSON format.

6. The potential customer prediction system according to claim 1 , wherein the calculation means plots the distribution of the hiring probability as a choropleth map using geospatial data in GeoJSON format.

7. the data acquisition means acquires a plurality of pieces of prescription record information for each of a plurality of customers collected after the new product is launched; The potential customer prediction system according to claim 1 , wherein the adoption probability estimation means reconfigures the binary classification algorithm based on the prescription record information.

8. 2. The potential customer prediction system according to claim 1, wherein the calculation means outputs the adoption probability to a file in Excel (registered trademark) format or other spreadsheet format.

9. A potential customer prediction method for predicting potential customers using a calculation means, a data acquisition step in which the calculation means acquires a plurality of pieces of attribute information and a plurality of pieces of activity information relating to each of a plurality of customers; a stratification step in which the calculation means stratifies a plurality of customers; an early adopter determination step in which the calculation means determines whether each of the stratified customers is an early adopter; an adoption probability estimation step in which the calculation means estimates an adoption probability indicating the probability that each of the plurality of customers is an early adopter using a binary classification algorithm; In the stratification step, the calculation means digitizes the attribute information and the activity information to generate observed variables, extracts common factors that commonly contribute to the observed variables and intrinsic factors that uniquely contribute to the observed variables, and calculates the contributions of the common factors and the intrinsic factors to the observed variables as factor scores; A potential customer prediction method, wherein in the early adopter determination step, the calculation means calculates a total value of the factor scores for each customer, generates a factor score distribution indicating a distribution of the number of customers for each of the total values, and determines the early adopters using the factor score distribution.

10. 10. The potential customer prediction method according to claim 9, wherein in the stratification step, the calculation means calculates the factor scores using a factor analysis model, a cluster analysis model, a principal component analysis model, or a model that is a combination of these.

11. The method for predicting potential customers according to claim 9 , wherein the binary classification algorithm comprises a binary logistic regression model, gradient boosting, random forest, support vector machine, neural network, or a model that is a combination thereof.

12. 10. The potential customer prediction method according to claim 9, wherein in the early adopter determination step, the calculation means determines customers in the top 10 to 20% of the factor score distribution as early adopters.

13. The potential customer prediction method according to claim 9, further comprising a drawing step in which the calculation means draws the distribution of the number of early adopters for each location as a choropleth map using geospatial data in GeoJSON format.

14. The potential customer prediction method according to claim 9, further comprising a drawing step in which the calculation means draws the distribution of the hiring probability for each location as a choropleth map using geospatial data in GeoJSON format.

15. In the data acquisition step, the calculation means acquires a plurality of pieces of prescription record information for each of a plurality of customers collected after the new product is launched, The potential customer prediction method according to claim 9 , wherein in the adoption probability estimation step, the calculation means reconstructs the binary classification algorithm based on the prescription record information.

16. 10. The potential customer forecasting method according to claim 9, further comprising an output step in which the calculation means outputs the adoption probability to a file in Excel (registered trademark) format or other spreadsheet format.

17. A potential customer prediction program that, when executed on a computer, causes the computer to function as the potential customer prediction system according to claim 1.

Citation Information

Patent Citations

  • Information providing method, information providing device, and information providing program

    JP2010198493A

  • Segmentation device, segmentation program and segmentation method

    JP2015069502A

  • Information processing method, information processing device, and program

    JP2020057221A

  • Information processing device and information processing method

    JP7636838B1

  • System and method for rebate marketing

    WO2005031624A1