Medical marketing system-oriented business data analysis method and system

By constructing a linked data model and a large language model for pharmaceutical marketing operations, the problems of data fragmentation, single analysis dimensions, and weak decision support in pharmaceutical marketing data analysis have been solved. This has enabled comprehensive production diagnosis and optimization suggestions, improving management efficiency and the scientific nature of decision-making.

CN121998263APending Publication Date: 2026-05-08CHIA TAI TIANQING PHARMA GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHIA TAI TIANQING PHARMA GRP CO LTD
Filing Date
2026-04-09
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing pharmaceutical marketing analysis technologies suffer from insufficient data integration dimensions, lack of customized analysis models, weak anomaly identification and case extraction capabilities, and weak decision support capabilities, leading to resource waste and low management efficiency.

Method used

This paper designs a business data analysis method for the pharmaceutical marketing system. By collecting multi-source data, constructing a linked data model, applying a quantitative evaluation algorithm for production health, and combining it with a large language model for intelligent diagnosis, the method outputs anomaly identification, excellent case extraction, and optimization suggestions to achieve full-dimensional production diagnosis.

Benefits of technology

It has achieved full transparency and visualization of the core links of pharmaceutical marketing business data, improved management efficiency and scientific decision-making, and promoted the transformation of marketing management from extensive to intensive.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure SMS_7
    Figure SMS_7
  • Figure SMS_26
    Figure SMS_26
Patent Text Reader

Abstract

The invention relates to a business data analysis method and system oriented to a medicine marketing system, and the method comprises the steps: collecting the multi-source data of four subjects, namely, corresponding organizations, personnel, hospitals and products, in a medicine marketing business, building a linkage data model related to the multi-source data of the four subjects, getting through the four dimensions, namely performance, cost, customer distribution and interactive behaviors, and building a linkage data model; a clear, comprehensive and accurate unified data system is established, online core data is realized, a core analysis layer and a diagnosis calibration layer are established based on a large language model, series analysis of four-dimensional data under multiple granularities of organizations, personnel and clients is realized, hierarchical and drillable reports are generated and displayed in combination with an output display layer, and the report data is displayed. Analysis of each subject of organizations, personnel, hospitals and products is visually presented, traditional manual number reading and manual analysis modes are replaced, available business optimization suggestions are rapidly output, business is assisted to objectively measure the operating health degree, marketing management is promoted to be transformed from extensive mode to refined mode, and management efficiency and decision scientificity are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a business data analysis method and system for the pharmaceutical marketing system, belonging to the field of intelligent analysis technology for pharmaceutical marketing data. Background Technology

[0002] The core competitiveness of pharmaceutical marketing lies in "precise resource allocation" and "efficiency of input and output". With stricter regulations and intensified market competition in the pharmaceutical industry, companies urgently need to optimize the allocation of marketing resources through data-driven approaches to improve the return on investment.

[0003] Pharmaceutical marketing involves multiple levels of entities: at the organizational level (representative offices), it is necessary to control the balance between overall performance and expenses; at the personnel level (sales representatives), it is necessary to improve the efficiency of individual resource utilization; at the hospital level (partner clients), it is necessary to achieve in-depth operation of high-value clients; and at the product level, it is necessary to monitor the performance of different products and the rationality of expenses. The performance, expenses, allocation, and interaction data generated by each entity are the core basis for evaluating the efficiency of investment.

[0004] Traditional pharmaceutical marketing analysis models rely on manual processing of scattered data, only providing a basic overview of "performance and expenses," failing to answer core questions such as "Is the cost investment reasonable?", "Are the interactive behaviors effective?", and "Why do some personnel / hospitals have outstanding productivity?" For example, they cannot quickly identify abnormal personnel or hospitals with "high costs and low performance," nor can they replicate the successful experience of "low costs and high growth," leading to resource waste and low management efficiency.

[0005] With the widespread application of big data technology in the pharmaceutical industry, enterprises have an increasingly urgent need for "comprehensive, precise, and implementable" production analysis. Existing technologies, such as the sales data statistics system and general marketing analysis tools commonly used in the pharmaceutical industry, have core solutions including the pharmaceutical sales data statistics system and general marketing analysis tools. Among them, the pharmaceutical sales data statistics system mainly connects to the sales management system and expense reimbursement system, collects basic data on organizational and personnel performance and expenses, as well as customer allocation data, to provide basic statistical reports such as performance achievement rate, expense ratio, and regional ranking. It supports single-dimensional data filtering and export, but lacks the ability to perform multi-subject data correlation analysis. Moreover, the output format is mainly fixed-format tables, which only present data results and do not include anomaly identification, root cause analysis, and optimization suggestions.

[0006] While general marketing analytics tools support the uploading and integration of structured data, they lack dedicated data interfaces designed for the multidimensional data architecture of pharmaceutical marketing, resulting in poor adaptability. These tools provide general planned expense ratio calculations and trend analysis functions, but lack customized models for pharmaceutical scenarios, such as "correlation between interactive behavior and performance" and "tiered customer return on investment analysis." Furthermore, they can only generate basic visualization charts (bar charts, line charts) and cannot perform multi-level drill-downs (e.g., from organization to personnel, hospitals), leading to weak decision support capabilities. Therefore, while these technical solutions provide basic data statistics tools for pharmaceutical marketing, their limitations in data integration depth and the universality of analytical models prevent them from achieving "full-dimensional return on investment diagnosis" and "implementable decision support," making it difficult to meet the refined marketing management needs of pharmaceutical companies.

[0007] Therefore, the existing technology currently has the following shortcomings: 1. Insufficient Data Integration Dimensions: Existing technologies only cover partial data, failing to achieve full-dimensional correlation among organizations, personnel, hospitals, and products. Causal Chain: The effectiveness of pharmaceutical marketing investment is influenced by the interaction of multiple stakeholders (e.g., personnel interaction directly affects hospital performance, and product characteristics determine hospital suitability). Data fragmentation prevents the establishment of a complete "input-behavior-output" logic, leading to biased analysis conclusions and an inability to accurately pinpoint the root causes of investment problems. This results in a lack of data support for resource optimization, leading to wasted investment.

[0008] 2. Lack of Pharmaceutical Customization in Analytical Models: Existing technologies use general analytical models that are not adapted to the multi-entity characteristics of pharmaceutical marketing. Causal Chain: In pharmaceutical marketing, there is a strong binding relationship between "personnel-hospitals-products" (e.g., a specific representative is responsible for a specific product at a specific hospital). General models cannot identify anomalies in production and investment under such relationships → cannot accurately calculate the planned expense ratios of multiple entities, and have difficulty distinguishing between effective and ineffective investments → marketing management decisions are highly unpredictable and cannot be optimized in a targeted manner.

[0009] 3. Lack of anomaly identification and case extraction capabilities: Existing technologies only present data results and lack intelligent diagnostic mechanisms. Causal chain: The pharmaceutical marketing data volume is large, and it is difficult for manual screening to quickly identify key scenarios such as "high cost and low performance" and "low cost and high growth" → anomalies cannot be warned in a timely manner, and good experiences cannot be quickly replicated → management efficiency is low, and it is difficult to continuously improve return on investment.

[0010] 4. Weak Decision Support Capabilities: Existing technology outputs are limited in form and lack tiered, actionable optimization suggestions. Cause-and-effect chain: Marketing managers need to quickly obtain clear directions on "how to optimize the overall picture and how to adjust specific areas." The presentation of raw data requires significant time for interpretation → low decision-making efficiency and susceptibility to misjudgments due to interpretation biases → impacting the effectiveness of marketing strategies and slowing ROI improvement. Summary of the Invention

[0011] The technical problem to be solved by this invention is to provide a business data analysis method for the pharmaceutical marketing system, which addresses the core pain points in existing pharmaceutical marketing investment and production management, such as data fragmentation, single analysis dimensions, inaccurate investment and production correlation, weak decision support, and inefficiency of manual data analysis. It realizes transparency and visualization of the entire core link of performance, expenses, customer allocation, and customer interaction, helps the business objectively measure the health of operations, promotes the transformation of marketing management from extensive to intensive, and improves management efficiency and scientific decision-making.

[0012] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: The present invention designs a business data analysis method for the pharmaceutical marketing system, comprising the following steps: Step A. Based on the pre-set systems in the pharmaceutical marketing business, collect multi-source data corresponding to the four main entities of organization, personnel, hospital and product, and the multi-source data includes performance data, expense data, allocation data and interaction data, and then proceed to step B; Step B. Clean the collected multi-source data and build a linkage data model to link the multi-source data of the four subjects (organization, personnel, hospital, and product) based on the four-level unique association keys. Then proceed to step C. Step C. Based on the multi-source associated data of the four subjects (organization, personnel, hospital, and product) under the linkage data model, apply the preset quantitative assessment algorithm for production health for each of the four subjects (organization, personnel, hospital, and product) to quantify and score them according to the dimensions of performance data, cost data, allocation data, and interaction data. Perform quantitative analysis of production health for each subject to obtain a comprehensive production health score for the subject, and then proceed to step D. Step D. Input the comprehensive health scores of the four main entities—organization, personnel, hospital, and product—as well as the multi-source correlation data of the four entities under the linkage data model into the target large language model. Combined with the structured prompts under the preset medical scenario application, drive the target large language model to perform intelligent diagnosis and output a preliminary diagnostic conclusion including anomaly identification, excellent case extraction, comprehensive diagnosis, and optimization suggestions. Then, verify the preliminary diagnostic conclusion through a preset multi-dimensional cross-validation mechanism to obtain the final full-dimensional production diagnosis conclusion.

[0013] As a preferred technical solution of the present invention: In step A, a single sales representative, a single doctor, and a single product are taken as the data collection objects. Real-time data collection is performed for multi-source data of dynamic data types, and batch data collection is performed for multi-source data of static basic data types. In accordance with the marketing compliance requirements of the pharmaceutical industry, compliance verification is performed on the collected multi-source data.

[0014] As a preferred technical solution of the present invention: In step A, the indicators of the multi-source data corresponding to the organization include office identification O-ID, performance target and actual value, total cost, SL plan cost, cost ratio, total number and level of customers, doctor interaction rate, regional market capacity, and month-on-month and year-on-year performance data. The various indicators of the multi-source data corresponding to personnel include sales representative P-ID, office identification O-ID, individual performance, SL plan costs and actual expenditures, number and level of assigned doctors, frequency, type and target of interaction, coverage rate of target level doctors, compliance of visits, and assessment cycle data. The indicators of the multi-source data corresponding to the hospital include the hospital H-ID, the office identifier O-ID, the hospital level, the cooperative product Pr-ID, the performance contribution, the SL plan expense ratio, the number of assigned doctors, the number and form of interaction, the participation coverage rate, and the departmental breakdown data. The various indicators of the multi-source data corresponding to the product include product Pr-ID, office identification O-ID, performance target and actual value, achievement rate, SL plan cost, cost ratio, matching hospital level and department, interactive conversion data, and category attributes.

[0015] As a preferred embodiment of the present invention: step B includes performing the following steps B1 to B3; Step B1. Perform abnormal data rule-based removal, data standardization and unification, and missing value medical business association completion on the collected multi-source data in sequence to clean and update the collected multi-source data, and then proceed to step B2; Step B2. Using the O-ID indicator corresponding to the organization's office, the P-ID indicator corresponding to the personnel, the H-ID indicator corresponding to the hospital, and the Pr-ID indicator corresponding to the product, construct a pre-defined four-level unique association key for the four entities: organization, personnel, hospital, and product, and establish P-ID. O-ID, H-ID O-ID, Pr-ID The inheritance relationship of O-ID and P-ID H-ID, P-ID Pr-ID, H-ID The association relationship of Pr-ID is then determined, and then proceed to step B3; Step B3. Connect the preset multi-source data collection time dimension t, and establish a linked data model according to the following formula; M_{t,(O,P,H,Pr),k}=V_{t,(O,P,H,Pr),k}+Tag_{(O-ID,P-ID,H-ID,Pr-ID)}; In the formula, O represents organization, P represents personnel, H represents hospital, Pr represents product, (O,P,H,Pr) represents the four-subject association combination of organization, personnel, hospital, and product, k represents the k-th indicator under the four-subject association combination of organization, personnel, hospital, and product, Tag_{(O-ID,P-ID,H-ID,Pr-ID) represents the association tag of the four-subject unique association key of organization, personnel, hospital, and product, V_{t,(O,P,H,Pr),k} represents the collected data value of the k-th indicator under the four-subject association combination of organization, personnel, hospital, and product corresponding to the time dimension t, and M_{t,(O,P,H,Pr),k} represents the linkage data matrix of the k-th indicator under the four-subject association combination of organization, personnel, hospital, and product.

[0016] As a preferred technical solution of the present invention: Step C is based on multi-source associated data of four subjects—organization, personnel, hospital, and product—under the linkage data model, and is performed according to the following formula for each of the four subjects: organization, personnel, hospital, and product.

[0017] The overall health score of the main entity's production status was calculated; among which, Indicates the first A comprehensive score for the overall health of the production of each entity. , , , The numbers represent the order of the numbers. Preset weights for performance data, expense data, allocation data, and interaction data under each entity. , , , The numbers represent the order of the numbers. Quantitative scoring of performance data, expense data, allocation data, and interaction data under each entity.

[0018] As a preferred technical solution of the present invention: Step C further includes a comprehensive score of the production health of each subject, including the organization, personnel, hospital and product, as well as a quantitative score of performance data, cost data, allocation data and interaction data of each subject. According to the following anomaly identification rules and excellent case selection rules, the anomaly points are accurately located and the excellent cases are reproducible and extracted. The rules for identifying anomalies and selecting outstanding cases include: the preset threshold for the overall health score of the abnormal entity, the preset threshold for each core indicator under multi-source correlated data, and the preset threshold for the overall health score of the abnormal entity and the threshold for the target core indicator under multi-source correlated data.

[0019] As a preferred technical solution of the present invention: in step D, the preliminary diagnostic conclusion is verified by the following preset multi-dimensional cross-validation mechanism to obtain the final full-dimensional production diagnostic conclusion. A quantitative indicator verification mechanism is used to verify whether the abnormal or excellent cases identified in the preliminary diagnostic conclusion meet the preset threshold quantitative judgment rules. If they do, the abnormal or excellent cases identified in the preliminary diagnostic conclusion are retained; otherwise, the corresponding abnormal or excellent cases identified in the preliminary diagnostic conclusion are rejected. The subject association verification mechanism, based on the linkage data model, verifies whether the abnormal root causes analyzed in the preliminary diagnostic conclusion are consistent with the standard quantitative data of the associated subject. If they are, the abnormal root causes analyzed in the preliminary diagnostic conclusion are retained; otherwise, the abnormal root causes analyzed in the preliminary diagnostic conclusion are rejected. The pharmaceutical business logic verification mechanism, based on pharmaceutical marketing business rules and industry compliance requirements, verifies whether the optimization suggestions in the preliminary diagnosis conclusion are consistent with the actual business situation. If they are, the optimization suggestions in the preliminary diagnosis conclusion are retained; otherwise, the optimization suggestions in the preliminary diagnosis conclusion are rejected.

[0020] As a preferred technical solution of the present invention, it further includes step E as follows: after step D is completed, step E is performed; Step E. Based on the aforementioned linked data model and the final comprehensive production diagnosis conclusion, generate and display hierarchical, drill-down reports. The reports support drilling down from the organizational level to the personnel level, hospital level, and product level for details. The reports also visualize the comprehensive production health scores of each entity (organization, personnel, hospital, and product), as well as the core indicators, identified anomalies, and optimization suggestions under multi-source related data.

[0021] Corresponding to the above, the technical problem to be solved by the present invention is to provide a system for implementing a business data analysis method for the pharmaceutical marketing system. The system is divided into modules based on functional applications, which efficiently implement each step of the business data analysis method, realizes the transparency and visualization of the entire core link of performance, expenses, customer allocation, and customer interaction, helps the business to objectively measure the health of its operations, promotes the transformation of marketing management from extensive to intensive, and improves management efficiency and the scientific nature of decision-making.

[0022] To solve the aforementioned technical problems, this invention adopts the following technical solution: This invention designs a system for implementing a business data analysis method for the pharmaceutical marketing system, comprising a data acquisition layer, a data processing layer, a core analysis layer, and a diagnostic calibration layer. The data acquisition layer executes step A, collecting multi-source data from four entities—organization, personnel, hospital, and product—corresponding to preset systems in the pharmaceutical marketing business. The data processing layer executes step B, constructing a linkage data model to link the multi-source data of the four entities. The core analysis layer executes step C, obtaining a comprehensive health score for the operational status of the four entities (organization, personnel, hospital, and product). The diagnostic calibration layer executes step D, applying a target large-scale language model to perform intelligent diagnosis, outputting preliminary diagnostic conclusions, and verifying them through a preset multi-dimensional cross-validation mechanism to obtain a final, comprehensive operational diagnostic conclusion.

[0023] As a preferred technical solution of the present invention, it also includes an output display layer, which is used to generate and display hierarchical and drill-down reports based on the linked data model and the final full-dimensional production diagnosis conclusion. The reports support drilling down from the organizational level to the personnel level, hospital level, and product level for details, and visually present the comprehensive production health score of each subject of the organization, personnel, hospital, and product, as well as the core indicators, identified anomalies, and optimization suggestions under multi-source related data.

[0024] The business data analysis method and system for the pharmaceutical marketing system described in this invention, compared with existing technologies, has the following technical advantages: This invention designs a business data analysis method and system for the pharmaceutical marketing system. It collects multi-source data from four main entities in pharmaceutical marketing operations: organizations, personnel, hospitals, and products. A linked data model is constructed to connect the multi-source data of these four entities, integrating performance, expenses, customer allocation, and interactive behavior across four dimensions. This establishes a clear, comprehensive, and accurate unified data system adapted to the marketing operations of pharmaceutical companies, enabling the online processing of core data. Based on a large language model, a business analysis system including a core analysis layer and a diagnostic calibration layer is built, achieving interconnected analysis of four-dimensional data at multiple granularities (organization, personnel, and customers). Combined with an output presentation layer, hierarchical and drill-down reports are generated and displayed, visually presenting the analysis of each entity (organization, personnel, hospitals, and products). This replaces the traditional manual data reading and analysis mode, quickly outputting actionable business optimization suggestions, helping businesses objectively measure operational health, promoting the transformation of marketing management from extensive to intensive, and improving management efficiency and the scientific nature of decision-making. Attached Figure Description

[0025] Figure 1 This is a schematic diagram illustrating the working framework of an embodiment of the business data analysis method for the pharmaceutical marketing system designed according to the present invention. Detailed Implementation

[0026] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0027] The design concept of this invention revolves around building a complete solution across the entire process of "data-analysis-application," with its core reliance on data integration and AI empowerment to achieve automated analysis, as detailed below: 1. Reconstruct underlying data: Connect key links such as customer allocation, interaction behavior, performance results, and expense expenditures to realize the online processing of core data, build a clear, comprehensive, and accurate unified data foundation, lay a solid foundation for subsequent analysis and AI product development, and reduce the cost of manual data collection and processing from the source.

[0028] 2. Build a customized analysis system: Based on AI capabilities, build a business analysis system to realize the serial analysis of "customer allocation-interaction-performance-expense" at multiple granularities of organization, personnel and customers, and realize the visualization of analysis results through the dashboard.

[0029] 3. Implement automated decision support: Automatically generate operational analysis reports such as organizational health status based on the needs of management at all levels. Combined with the operational dashboard system, it completely replaces the traditional manual data reading and analysis mode, and quickly outputs actionable business optimization suggestions.

[0030] Furthermore, the design takes the following points into consideration for the core business: 1. Kanban-enabled management: Through production kanban, the entire chain of organization, personnel and customers can be analyzed in a multi-dimensional way, breaking down the barriers of extensive management, making the entire marketing business process quantifiable, traceable and reviewable, and providing intuitive support for refined operation.

[0031] 2. Improved efficiency through automated reporting: Automatically generates organizational health and productivity analysis reports for all levels of management, efficiently addressing the pain points of cumbersome data retrieval and inefficient data interpretation on the business side, helping to quickly uncover business issues behind the data and drive precise optimization.

[0032] 3. Building a solid foundation of underlying capabilities: Constructing a clear, comprehensive, and accurate underlying data system not only improves the overall efficiency of data product development and data analysis, but also lays a solid data foundation for the research and development and implementation of subsequent AI products.

[0033] Based on the above design, this invention specifically designs a business data analysis method and system for the pharmaceutical marketing system. In practical applications, the system design includes a data acquisition layer, a data processing layer, a core analysis layer, a diagnostic calibration layer, and an output display layer. The business data analysis method is specifically implemented through the following steps A to E.

[0034] Step A. The data acquisition layer, based on the pharmaceutical enterprise sales management system, expense reimbursement system, customer management system, and product management system in pharmaceutical marketing operations, collects multi-source data on individual sales representatives, individual doctors, and individual products at a time granularity down to the day. This data covers four main entities: organization, personnel, hospital, and product, encompassing performance, expenses, allocation, and interaction. In accordance with pharmaceutical industry marketing compliance requirements, the collected multi-source data undergoes compliance verification to ensure it complies with these requirements. Specifically, for dynamic data types, real-time data acquisition is performed; for static, basic data types, batch data acquisition is performed. Then, proceed to Step B.

[0035] The application of the data acquisition layer connects the four core business systems of pharmaceutical companies: sales management system, expense reimbursement system, customer management system, and product management system. It enables automatic data collection from the source, avoiding errors and inefficiencies caused by manual data entry. All collected data is bound to a unique business identifier ID, providing a foundation for subsequent four-level association. It achieves multi-source, high-granularity, and compliant data collection from four main entities in pharmaceutical marketing: organization, personnel, hospitals, and products. This builds a scenario-based, structured, and complete basic data pool for subsequent association analysis. Unlike the indiscriminate data access of general systems, this layer is designed with a dedicated collection dimension and dual-mode access mechanism for pharmaceutical marketing business processes and compliance requirements, ensuring the accuracy and timeliness of the data source.

[0036] In practical applications, the collection of multi-source data involves the following specific metrics: The indicators corresponding to the organization's multi-source data include office identification (O-ID), performance targets and actual values, total costs, SL plan costs, cost ratio, total number and level of customers, doctor interaction rate, regional market size, and performance data compared to the previous and current periods. The various indicators of the multi-source data corresponding to personnel include sales representative P-ID, office identification O-ID, individual performance, SL plan costs and actual expenditures, number and level of assigned doctors (A / B / C level), interaction frequency, type and target, coverage rate of target level A and B doctors, visit compliance, and assessment cycle data. The indicators of the multi-source data corresponding to the hospital include the hospital H-ID, the office identifier O-ID, the hospital level, the cooperative product Pr-ID, the performance contribution, the SL plan expense ratio, the number of assigned doctors, the number and form of interaction, the participation coverage rate, and the departmental breakdown data. The various indicators of the multi-source data corresponding to the product include product Pr-ID, office identification O-ID, performance target and actual value, achievement rate, SL plan cost, cost ratio, matching hospital level and department, interactive conversion data, and category attributes.

[0037] Step B. The data processing layer cleans the collected multi-source data and constructs a linkage data model to link the multi-source data of the four subjects (organization, personnel, hospital, and product) based on the four-level unique association keys. Then, it proceeds to step C.

[0038] In practical applications, the data processing layer is specifically designed to execute the following steps B1 to B3.

[0039] Step B1. Perform abnormal data rule-based removal, data standardization and unification, and missing value medical business association completion on the collected multi-source data in sequence to clean and update the collected multi-source data, and then proceed to step B2.

[0040] In practical applications, the design and execution of step B1 above removes invalid data, standardizes data, and completes missing data to ensure data accuracy and consistency. All cleaning rules are deeply adapted to the characteristics of the pharmaceutical industry, as detailed below: 1. Abnormal data rule-based removal: Construct a pharmaceutical business rule verification library to automatically remove invalid data through preset rules. Core verification rules include: negative performance value, SL plan expenses exceeding the industry average for the product / region by more than 3 times, interaction frequency contradicting actual working hours, hospital level significantly inconsistent with performance contribution, and the number of doctors allocated exceeding the reasonable range (0-200), etc. 2. Data Standardization and Unification: All collected data undergoes standardization processing in terms of format, units, and representation. The core unified standards are as follows, ensuring the comparability of data from different sources: Uniform format: Date format is "YYYY-MM-DD", numerical values ​​are retained to 2 decimal places, and ID is a fixed-length character code; Units are standardized: Amounts are expressed in "yuan", quantities are expressed in "person / time / box", and percentages are expressed in "%". Standardized terminology: Hospital levels are standardized as "Grade III Class A / Grade III Class B / Grade II / Grade I", doctor classifications are standardized as "Grade A / Grade B / Grade C", and interaction types are standardized as "compliant visits / academic communication / document delivery / conference participation"; 3. Missing Value Imputation Based on Medical Business Relationships: This method replaces the common mean / median imputation with a business relationship imputation approach to avoid data distortion. The core imputation rules include: Key data (performance, SL plan expenses, number of assigned doctors): Completed by back-inferring through multi-level ID association, such as back-inferring the performance contribution of a single product or a single hospital through the association of "P-ID-H-ID-Pr-ID"; Non-critical data (interaction type, meeting coverage): marked as "not collected" and not included in quantitative calculations to avoid analytical bias caused by invalid data completion.

[0041] Step B2. Using the O-ID indicator corresponding to the organization's office, the P-ID indicator corresponding to the personnel, the H-ID indicator corresponding to the hospital, and the Pr-ID indicator corresponding to the product, construct a pre-defined four-level unique association key for the four entities: organization, personnel, hospital, and product, and establish P-ID. O-ID, H-ID O-ID, Pr-ID The inheritance relationship of O-ID and P-ID H-ID, P-ID Pr-ID, H-ID The association relationship of Pr-ID is established, and then step B3 is performed.

[0042] Regarding the four-level unique association key, the office identification O-ID indicator constitutes the top-level association key, which manages the data of its subordinate personnel, hospitals, and products; the sales representative P-ID indicator is associated with the assigned hospital and the product data it is responsible for; the hospital H-ID indicator is associated with the assigned personnel and the cooperative product data; and the product Pr-ID indicator is associated with the sales personnel and the cooperative hospital data.

[0043] Step B3. Connect the preset multi-source data collection time dimension t, and establish a linked data model according to the following formula; M_{t,(O,P,H,Pr),k}=V_{t,(O,P,H,Pr),k}+Tag_{(O-ID,P-ID,H-ID,Pr-ID)}; In the formula, O represents organization, P represents personnel, H represents hospital, and Pr represents product. (O,P,H,Pr) represents the four-subject association combination of organization, personnel, hospital, and product. k represents the k-th indicator under the four-subject association combination of organization, personnel, hospital, and product. Tag_{(O-ID,P-ID,H-ID,Pr-ID)} represents the association tag of the four-level unique association key of organization, personnel, hospital, and product, realizing full-link traceability of data from top level to smallest granularity. V_{t,(O,P,H,Pr),k} represents the collected data value of the k-th indicator under the four-subject association combination of organization, personnel, hospital, and product corresponding to the time dimension t. M_{t,(O,P,H,Pr),k} represents the linkage data matrix of the k-th indicator under the four-subject association combination of organization, personnel, hospital, and product.

[0044] The data processing layer performs pharmaceutical-specific cleaning and precise construction of a four-level linkage data model on the collected multi-source data. Unlike traditional general data cleaning and simple association, this layer designs a triple cleaning mechanism and a precise association algorithm with four unique association keys for pharmaceutical marketing business logic. It solves the data fragmentation problem from the bottom up and builds a unique, drill-down, and traceable hierarchical data chain of "organization → personnel → hospital → product", providing a precise data foundation for subsequent analysis.

[0045] In practical applications, based on the above-mentioned interconnected data model, a top-down hierarchical drill-down algorithm is designed. The drill-down logic is: O-ID → P-ID / H-ID / Pr-ID → P-ID+H-ID+Pr-ID (smallest granularity), ensuring that each data result can be accurately traced to a specific organization, person, hospital, or product portfolio. Moreover, the drill-down process does not require manual intervention and is automatically implemented by the system based on the ID association relationship.

[0046] Regarding the design of a linked data model to achieve the association of multi-source data from four main entities, in practical applications, knowledge graph technology can also be used to construct a multi-entity association network, which can more intuitively present the complex relationship between "personnel-hospital-products" and is suitable for the massive data scenarios of large pharmaceutical companies.

[0047] Step C. The core analysis layer, based on the multi-source correlated data of four entities—organization, personnel, hospital, and product—under the linked data model, applies a pre-set operational health quantification and evaluation algorithm to each of the four entities. The algorithm quantifies and scores data across performance, expense, allocation, and interaction dimensions using the following formula:

[0048] The overall health score of the main entity's commissioning is calculated to achieve a quantitative analysis of the entity's commissioning health, and then proceeds to step D; where, Indicates the first A comprehensive score for the overall health of the production of each entity. , , , The numbers represent the order of the numbers. Preset weights for performance data, expense data, allocation data, and interaction data under each entity. , , , The numbers represent the order of the numbers. Quantitative scoring of performance data, expense data, allocation data, and interaction data under each entity.

[0049] In the calculation of the overall health score of the above-mentioned entities in production, regarding the preset weights , , , In practical applications, differentiated weighting coefficients are set based on the core business needs of various entities in pharmaceutical marketing to ensure that the scoring results align with actual business requirements. Specifically, the weighting for organizational entities is 0.4 for performance, 0.3 for expenses, 0.15 for allocation, and 0.15 for interaction. This weighting design focuses on achieving organizational performance goals and overall cost control. For personnel entities, the weighting is 0.35 for performance, 0.25 for expenses, 0.2 for allocation, and 0.2 for interaction. The weighting is based on the individual performance of personnel, the rationality of customer allocation, and the effectiveness of interaction and conversion. For hospitals, the weighting is 0.3 for performance, 0.25 for cost, 0.2 for allocation, and 0.25 for interaction. This weighting is designed with the hospital's performance contribution, interaction frequency, and matching cost investment as the core. For products, the weighting is 0.45 for performance, 0.3 for cost, 0.1 for allocation, and 0.15 for interaction. This weighting is designed with the product's performance achievement and the rationality of cost investment as the core.

[0050] And about , , , In practical applications, the calculation uses a universal dimensional quantitative scoring formula designed for four main entities: organizations, personnel, hospitals, and products. All indicators are adapted to the pharmaceutical marketing scenario, and industry averages, team averages, and benchmark values ​​are introduced as references to ensure the objectivity and accuracy of the scoring, as detailed below.

[0051] Quantitative scoring of performance dimensions (0-100): Taking into account both performance achievement rate and month-on-month growth rate, the core formula is: =min ((PA / PA_std)×60 + (PG / PG_std)×40, 100 ), where: PA is the performance achievement rate (PA=actual performance / target performance×100%), PA_std is the benchmark value of the performance achievement rate (default 100%), PG is the performance month-on-month growth rate (PG =(current period performance-previous period performance) / previous period performance×100%), PG_std is the benchmark value of the performance month-on-month growth rate (default 0%).

[0052] Cost dimension quantitative scoring (0-100): The score is based on the deviation of the SL plan expense ratio from the industry / company benchmark. The smaller the deviation, the higher the score. Core formula: = 100 - min (|FER / FER_std - 1|×100, 80), where FER is the SL plan expense rate (FER=SL plan expense / actual performance×100%), and FER_std is the benchmark value of the SL plan expense rate (take the average compliance value of the same product / region in the pharmaceutical industry).

[0053] Quantitative scoring of allocation dimensions (0-100): Taking into account both customer allocation suitability and the proportion of high-value customers, the core formula is: = min ((CA / CA_std)×50 + (CR / CR_std)×50, 100), where CA is the customer allocation fit, CR is the allocation ratio of high-value customers (A / B level), and CA_std and CR_std are the corresponding baseline values ​​(default 100% and 60%).

[0054] Interaction dimension quantitative scoring (0-100): Taking into account both effective interaction rate and high-value customer interaction coverage, the core formula is: = min ((IR / IR_std)×60 + (HC / HC_std)×40, 100), where IR is the effective interaction rate, HC is the interaction coverage rate of high-value customers (A / B level), and IR_std and HC_std are the corresponding baseline values ​​(default 80% and 90%).

[0055] In practical applications, after achieving quantitative analysis of the main body's operational health in step C, further design is made to comprehensively score the operational health of the main body based on the organization, personnel, hospital, and product, as well as to quantitatively score the performance data, cost data, allocation data, and interaction data of each main body. According to the following rules for identifying outliers and selecting excellent cases, the accurate location of outliers and the replicable extraction of excellent cases are achieved. The rules for identifying anomalies and selecting outstanding cases include: the preset threshold for the overall health score of the abnormal entity, the preset threshold for each core indicator under multi-source correlated data, and the preset threshold for the overall health score of the abnormal entity and the threshold for the target core indicator under multi-source correlated data.

[0056] Regarding anomaly identification and best practice extraction, in practical applications, a comprehensive score based on the main production health is used. and linked data models The system conducts detailed analysis on four main entities: organization, personnel, hospital, and product. It also designs specific quantitative rules for identifying anomalies and selecting best practices to achieve accurate anomaly localization and replicable extraction of best practices. The rules are all based on health scores and quantitative values ​​of core indicators to avoid the subjectivity of human judgment, as detailed below.

[0057] 1. Organizational Entity Analysis Key metrics calculations: performance achievement rate, per capita efficiency (performance per employee = total organizational performance / number of sales personnel), total expense ratio, SL plan expense ratio, doctor interaction rate, and customer allocation rate; Trend Analysis: Based on Linked Data Model Calculate the month-on-month growth rate (PG) of performance, expenses, and interaction rate to identify overall growth / decline trends; Comprehensive diagnosis: Combining the overall health score of the main production unit Based on trend analysis results, assess the balance between "performance-expenses-interactions" and output results such as " The core conclusion is: "=85 (excellent) but PG=-27.7% (performance declined sequentially)". Anomaly / Excellent Judgment: <40 indicates an organizational abnormality. Organizations with a score ≥80 and a month-on-month growth rate of core indicators higher than the industry average are considered outstanding cases.

[0058] 2. Personnel Analysis Health assessment: The overall health score of the main personnel in production is calculated using the above formula. ,position For individuals with scores below 40, we simultaneously analyze individual scores across four dimensions to accurately pinpoint the root cause of the anomaly (e.g., ...). A cost ratio of less than 30 is considered abnormal. <30 indicates an abnormal interaction rate); Anomaly identification quantification rules: ① <40; ② Performance achievement rate PA<50% and SL plan expense ratio FER>20% (low performance, high expense); ③ Performance month-on-month growth rate PG<-30% (significant performance decline); ④ Interaction coverage rate of A-level and B-level doctors HC<50% (insufficient interaction with high-value clients); Quantitative rules for extracting excellent cases: ① ≥85; ② The month-on-month growth rate of performance (PG) is greater than 50% and the expense ratio (FER) of the SL plan is less than the team average; ③ The interaction coverage rate (HC) of A-level and B-level doctors is 100% and the interaction conversion rate is greater than the team average; Case Feature Extraction: Extract allocation characteristics (A / B level customer ratio), interaction characteristics (interaction frequency / type), and cost characteristics (SL plan cost investment direction) of outstanding personnel to form a replicable personnel operation strategy; 3. Hospital Main Body Analysis Health assessment: The overall health score of the hospital's main structure is calculated using the formula above. ,position For hospitals with an abnormal score of <40, analyze the root cause of the abnormality in conjunction with the hospital's level (e.g., tertiary hospitals). <30 represents high-level, low-performance); Anomaly identification quantification rules: ① <40; ②SL plan expense ratio FER>30% and performance growth rate PG<0 (high cost, low growth); ③ranks in the bottom 10% of hospitals of the same level in terms of performance contribution and interaction rate IR<60%; Quantitative rules for extracting excellent cases: ① ≥85; ② Month-on-month growth rate of performance (PG) > 20% and SL plan expense ratio (FER) < average of hospitals of the same level; ③ Month-on-month growth of interaction frequency > 50% and performance growth > 15% (high interaction, high conversion); Case Feature Extraction: Extract interaction features (interaction format / frequency), cost features (cost input structure), and product features (high-performing product types) from outstanding hospitals to form replicable hospital operation strategies.

[0059] 4. Product Main Body Analysis Health assessment: The overall health score of the main product's production launch is calculated using the formula above. ,position For products with an abnormal score of <40, analyze the individual scores across four dimensions to pinpoint the root cause of the abnormality (e.g., ...). A cost ratio of less than 30 is considered abnormal. <30% have no performance); Key metrics calculations: performance achievement rate, month-on-month growth rate, SL plan expense ratio, and month-on-month change in expenses; Anomaly identification quantification rules: ① <40; ②SL plan expense ratio FER>50% and performance achievement rate PA<30% (high expense, low performance); ③Performance PA<20% for two consecutive assessment cycles (no performance contribution); Quantitative rules for extracting excellent cases: ① ≥85; ② Performance achievement rate PA>150% and SL plan expense ratio FER<average of similar products; ③ Month-on-month growth rate PG>30% and stable expense ratio (FER fluctuation <5%); Adaptability analysis: combining linked data models We analyze the appropriate hospital levels / departments and target sales groups for outstanding products to provide accurate data for product resource allocation.

[0060] The core analysis layer executes step C above by designing a comprehensive health assessment model for the four main entities: organization, personnel, hospital, and product. It introduces dynamic weighting coefficients for the pharmaceutical scenario and transforms qualitative indicators of four dimensions—performance, cost, allocation, and interaction—into quantitative scores of 0-100. Through weighted calculation, a comprehensive health score is obtained, enabling a quantitative evaluation of the production effect and solving the problem of traditional analysis being "primarily qualitative with no quantitative standards."

[0061] In practical applications, machine learning algorithms (such as decision trees and cluster analysis) are introduced into the design to replace rule-based identification, which improves the accuracy of identifying outliers in complex scenarios and is suitable for scenarios with more data dimensions.

[0062] Step D. The diagnostic calibration layer inputs the comprehensive health scores of the four main entities—organization, personnel, hospital, and product—as well as the multi-source correlation data of the four entities under the linkage data model into the target large language model. Combined with the structured prompts under the preset medical scenario application, the target large language model performs intelligent diagnosis and outputs a preliminary diagnostic conclusion that includes anomaly identification, excellent case extraction, comprehensive diagnosis, and optimization suggestions. The preliminary diagnostic conclusion is then verified through a preset multi-dimensional cross-validation mechanism to obtain the final full-dimensional production diagnosis conclusion.

[0063] In practical applications, the diagnostic calibration layer introduces large language models (LLM, such as Qwen and GPT) for intelligent diagnosis. However, unlike the traditional simple application of "data + prompts", it designs a customized structured prompt framework for medical scenarios and a cross-validation mechanism for multi-dimensional analysis results to solve the problems of poor scenario adaptability and result uncertainty in LLM analysis, and ensure the accuracy and reliability of analysis conclusions. The specific implementation is as follows.

[0064] Customized Structured Prompt Framework for Medical Scenarios: Based on Linked Data Model Based on the quantitative health score results, standardized and structured prompts are designed to avoid analytical bias caused by vague prompts. The prompt framework includes five core elements: analysis object, data range, quantitative indicators, analysis objectives, and output format, ensuring the relevance and accuracy of LLM analysis. An example is as follows: Analysis object: 8 sales representatives (P-ID: P001-P008) of the Beijing XX office (O-ID: O001), 11 partner hospitals (H-ID: H001-H011), and 6 key products (Pr-ID: Pr001-Pr006); Data range: Linked data model for Month X, 202X. The data includes full data, with core indicators across four dimensions: performance, expenses, allocation, and interaction; quantitative indicators: a comprehensive score of the operational health of each entity. And scores for each dimension S1-S4, performance achievement rate PA, SL plan expense ratio FER, interaction rate IR, etc.; Analysis objectives: identify abnormal entities, extract excellent cases, analyze the root causes of abnormalities, and propose targeted optimization suggestions; Output format: abnormal entity (ID+HS+root cause of abnormality), excellent entity (ID+HS+core characteristics), optimization suggestions (customized for each entity, tailored to pharmaceutical marketing business).

[0065] Regarding the execution of LLM intelligent diagnostics, structured prompts and linked data models. The quantitative data is input into the pre-trained LLM, and the LLM outputs preliminary analysis conclusions, including four major modules: outlier identification, excellent case extraction, comprehensive diagnosis, and optimization suggestions.

[0066] The further diagnostic calibration layer verifies the preliminary diagnostic conclusions through the following preset multi-dimensional cross-validation mechanism to obtain the final full-dimensional production diagnostic conclusions.

[0067] A quantitative indicator verification mechanism is used to verify whether the abnormal or excellent cases identified in the preliminary diagnostic conclusion meet the preset threshold quantitative judgment rules. If they do, the abnormal or excellent cases identified in the preliminary diagnostic conclusion are retained; otherwise, the corresponding abnormal or excellent cases identified in the preliminary diagnostic conclusion are rejected. The subject association verification mechanism, based on the linkage data model, verifies whether the abnormal root causes analyzed in the preliminary diagnostic conclusion are consistent with the standard quantitative data of the associated subject. If they are, the abnormal root causes analyzed in the preliminary diagnostic conclusion are retained; otherwise, the abnormal root causes analyzed in the preliminary diagnostic conclusion are rejected. The pharmaceutical business logic verification mechanism, based on pharmaceutical marketing business rules and industry compliance requirements, verifies whether the optimization suggestions in the preliminary diagnosis conclusion are consistent with the actual business situation. If they are, the optimization suggestions in the preliminary diagnosis conclusion are retained; otherwise, the optimization suggestions in the preliminary diagnosis conclusion are rejected.

[0068] Regarding optimization suggestions, in practical applications, large language models (such as Qwen and GPT) can be integrated to generate more scenario-based optimization suggestions based on the analysis results, which are suitable for scenarios that require in-depth business interpretation.

[0069] The multi-dimensional cross-validation mechanism is one of the core innovations of this invention. It uses triple validation to calibrate the initial LLM analysis conclusions, eliminating contradictory results and correcting biased conclusions, ensuring the consistency, rationality, and business adaptability of the final analysis results. Conclusions that fail validation are rejected and re-entered into the LLM analysis until validation is successful. Finally, the LLM analysis conclusions that pass triple cross-validation are deeply integrated with the multi-entity operational health metric scoring results to generate a final comprehensive operational diagnostic conclusion, providing accurate and reliable basis for subsequent decision-making.

[0070] Step E. The output display layer generates and displays hierarchical, drill-down reports based on the linked data model and the final full-dimensional production diagnosis conclusion. The reports support drilling down from the organizational level to the personnel level, hospital level, and product level for details. The report also visualizes the comprehensive production health score of each entity (organization, personnel, hospital, and product), as well as the core indicators, identified anomalies, and optimization suggestions under multi-source related data.

[0071] The output presentation layer is designed to meet the needs of managers at different levels in pharmaceutical companies, enabling collaboration between "top-down control and bottom-up execution." Unlike traditional fixed-format reports, this layer supports multi-dimensional drill-down and personalized customization. Furthermore, the alerts and suggestions are based on quantitative analysis results, making them highly practical. Specific applications are as follows.

[0072] 1. Hierarchical drill-down report display: based on a four-level linked data model The design incorporates a four-tier report structure, from overall to detailed, with one-click drill-down between reports. Each report carries a four-level association key label, enabling end-to-end data traceability, as detailed below: Overall Overview Report (Organizational Level): Presents the overall performance, expenses, key interactive indicators, comprehensive score of the main production health, and comprehensive diagnostic conclusions of the branch office, and supports drill-down to personnel / hospital / product details; Personnel Details Report (Personnel Level): Displays the overall health score of each representative's main production and the scores of each dimension, the quantitative values ​​of core indicators, outliers / excellent features, and supports drill-down to the representative's assigned hospital and responsible product details; Hospital Detailed Report (Hospital Level): Displays the overall health score of each hospital's main production and the scores of each dimension, the quantitative values ​​of core indicators, the comparison with the average of hospitals of the same level, outliers / excellent features, and supports drill-down to the hospital's allocated personnel and cooperative product details; Product Details Report (Product Layer): Displays the overall health score of each product's main production status, scores for each dimension, performance achievement, cost input, month-on-month changes, and adaptability analysis. It also supports drill-down to details of the product's sales personnel and partner hospitals.

[0073] In practical applications, BI tools (such as Power BI and Tableau) can be used to design and develop customized reports to enhance interactivity and visualization, and meet personalized reporting needs.

[0074] 2. Tiered Anomaly Early Warning Module: A red-yellow-green three-color tiered early warning system is designed. Based on the comprehensive health score of the main entity and the quantitative rules for anomaly identification, abnormal entities are tiered and marked, and an anomaly list is automatically generated. The list includes five core elements: abnormal entity ID, anomaly type, anomaly root cause, impact degree, and related data. The specific early warning levels are as follows: Red Alert: If the score is less than 40 or meets the core anomaly identification rules (such as low performance and high expenses), it is considered a severe anomaly and needs to be dealt with immediately. Yellow alert: 40≤ A value less than 60 indicates a mild anomaly, requiring monitoring and optimization within a specified timeframe. Green is normal: A value of ≥60 indicates a normal condition; continuous monitoring is required.

[0075] 3. Customized Decision Recommendation Module: Based on the final analysis conclusions, this module pushes hierarchical, customized, and actionable optimization suggestions. These suggestions are strongly linked to anomalies / best practice cases and are adapted to the business characteristics of different stakeholders, as detailed below: At the organizational level: focus on balancing and optimizing overall performance, costs, and interactions, such as "optimizing the cost structure of the SL plan and reducing cost investment in high-cost, low-performance areas" and "identifying key personnel / hospitals with declining performance and improving interaction conversion in a targeted manner." At the personnel level: Focus on targeted optimization based on the root causes of individual abnormalities, such as "strengthening A-level and B-level customer interaction and conversion training for probationary representatives and replicating the interaction strategies of outstanding representatives" and "providing guidance on cost investment for low-performing and high-cost representatives to reduce the SL plan cost rate." At the hospital level: Focus on optimizing the alignment between hospital level and performance, costs, and interactions, such as "increasing the frequency of interactions for high-level hospitals with low performance and adjusting the allocation of personnel" and "optimizing the cost structure of hospitals with high costs and low growth and reducing ineffective cost investment." At the product level: focus on optimizing the fit between product performance and expenses, such as "verifying the rationality of investment in high-expense products with no performance, and considering reducing or suspending expense investment", and "replicating the hospital / personnel matching strategy of excellent products to expand the market coverage of high-performance products".

[0076] The business data analysis method and system designed for the pharmaceutical marketing system in this invention revolves around a technical system built around pharmaceutical scenario-specific data linkage, AI-driven full-process, and multi-dimensional result cross-validation. Addressing the strong business binding characteristics of the four main entities in pharmaceutical marketing—organization, personnel, hospitals, and products—it designs a four-layer collaborative architecture: data acquisition layer, data processing layer, core analysis layer, diagnostic calibration layer, and output display layer. This architecture follows a closed-loop time sequence: multi-source data access → pharmaceutical scenario-specific cleaning and association → AI quantitative analysis → LLM intelligent diagnosis → cross-validation calibration → layered visualization output.

[0077] The core innovation of this invention lies in the design of a precise construction algorithm for a four-level linkage data model in pharmaceutical marketing, a dynamic weighting algorithm for quantitative assessment of the health of multiple entities' production and operation, and a multi-dimensional cross-validation mechanism for LLM analysis results. This overcomes the pain points of traditional analysis systems, such as coarse data association, general analysis models, and superficial AI applications. It achieves precise association, intelligent diagnosis, and efficient decision-making of production and operation data across the entire pharmaceutical marketing chain. The overall architecture has three core characteristics: deep adaptability to pharmaceutical scenarios, high accuracy of analysis results, and strong feasibility of decision recommendations.

[0078] The following is a summary of the meanings of the terms involved in the business data analysis method and system for the pharmaceutical marketing system designed in this invention.

[0079] Artificial Intelligence (AI) refers to systems or machines that simulate human intelligence and are capable of cognitive activities such as learning, reasoning, and decision-making.

[0080] Large Language Model (LLM): refers to a deep learning model (such as GPT, Qwen, Llama, etc.) trained on massive amounts of text data, which has powerful natural language understanding and generation capabilities.

[0081] Natural Language Processing (NLP): A branch of artificial intelligence that primarily studies the interaction between computers and human language, including the understanding, analysis, and generation of natural language.

[0082] Structured data refers to data that has been organized and formatted so that it can be stored and processed in a defined structure (such as database tables or Excel spreadsheets).

[0083] Performance: In the context of pharmaceutical sales, performance specifically refers to the actual sales revenue generated for drugs or products, and is a core indicator for measuring operating results.

[0084] Sales & Marketing Plan Expense (SL Plan Expense): In the pharmaceutical sales context, this specifically refers to the dedicated expenses set aside to increase drug sales (i.e., "volume growth"). This expense is primarily used for marketing activities such as customer maintenance, academic promotion, and conference organization, and is a core component of "planned expenses."

[0085] Planned Expense Rate: This refers to the ratio of planned expenses to related performance revenue. In the pharmaceutical industry, this indicator is primarily used to measure the rationality and control over marketing expenditures.

[0086] Performance Achievement Rate: This refers to the ratio of actual sales revenue to target sales revenue, and is mainly used to evaluate the target achievement of sales teams or sales personnel.

[0087] Customer Allocation Rate: This refers to the percentage of customers allocated to sales personnel out of the total number of customers. These customers include core clients in pharmaceutical sales, such as doctors and hospitals.

[0088] Visit Coverage Rate: This refers to the percentage of doctors actually visited by a salesperson within the assessment period, relative to the total number of doctors assigned to them. Here, "visit" specifically refers to effective on-site communication actions that comply with industry compliance requirements.

[0089] Interaction Rate: This refers to the percentage of effective interactions between sales personnel and doctors out of total interaction opportunities. Effective interactions include compliance visits, academic communication, and document delivery—actions that conform to industry standards.

[0090] Month-on-month growth rate: This refers to the increase in data for the current statistical period compared to the data for the immediately preceding statistical period. The statistical period can be monthly, quarterly, or yearly, and is suitable for multi-period performance evaluation scenarios in pharmaceutical sales.

[0091] The design technology solution of the present invention is applied to specific embodiments, according to... Figure 1 Execution, for example, Beijing XX Office A uses this system for comprehensive production analysis, the specific application process is as follows: 1. The data collection layer accesses the office's organizational data (target performance of RMB 999,000, actual performance of RMB 1,784,000), the individual performance / expenses / allocation / interaction data of 8 representatives, the performance / expenses / interaction data of 11 cooperating hospitals, and the performance / expense data of 6 key products.

[0092] 2. Data processing layer cleans the data (removing one invalid fee record), establishes associations through "Office ID", "Personnel ID", "Hospital ID", and "Product ID", and constructs a hierarchical data chain.

[0093] 3. The core analysis layer and diagnostic calibration layer complete the full-dimensional diagnosis: At the organizational level: the performance achievement rate was 178.2%, but it declined by 27.7% compared to the previous period. The SL plan expense ratio increased slightly by 12.6%, and the interaction rate increased by 13.1%. The diagnosis was "performance exceeded expectations but growth is under pressure, expenses have been reduced but efficiency needs to be optimized". At the personnel level: Identify anomalies such as "0 RMB in sales but 9,600 RMB in expenses" and "46 interactions (lowest in the team)" for Li XX (probationary period), and extract the excellent case of XX with "134.04% sales growth and expense ratio lower than the team average"; At the hospital level: Identify anomalies such as Beijing XX Hospital A's "expense ratio of 35.56%, but performance decreased by 43.68% quarter-on-quarter," and extract the excellent case of Beijing XX Hospital B's "interaction frequency increased by 62.13%, driving performance growth of 15.35%"; At the product level: Identify the anomaly of the Tianyun product's "expense ratio surge of 50.25%" and the excellent performance of the Ganping product's "stable revenue growth and low expense ratio".

[0094] 4. Output presentation layer generates hierarchical reports, marking anomalies in red and pushing decision-making suggestions: strengthen interactive conversion training for probationary representatives; optimize the cost structure of Beijing XX Hospital A; replicate the high-interaction strategy of Beijing XX Hospital B; verify the reasonableness of Tianyun product costs. After the office implemented the suggestions, the overall month-on-month decline narrowed to 10% the following month, and the performance of all three abnormal representatives achieved breakthroughs.

[0095] The business data analysis method and system designed for the pharmaceutical marketing system in this invention exhibit the following four beneficial effects in practical applications.

[0096] 1. Comprehensive data correlation for precise production diagnosis: This invention integrates data from four main entities—organization, personnel, hospitals, and products—to construct an "input-behavior-output" linkage system. Direct technical effects: Breaking down data silos and enabling full-link data traceability from the overall picture to details; Ultimately: It can accurately pinpoint the root causes of production problems (e.g., whether declining performance is due to insufficient personnel interaction or uneven allocation of hospital resources), avoiding the one-sidedness of traditional analysis and providing data support for targeted optimization.

[0097] 2. Customized Pharmaceutical Analysis Model for Enhanced Diagnostic Adaptability: This invention designs a proprietary return on investment analysis algorithm tailored to the multi-entity characteristics of pharmaceutical marketing. Direct Technical Effects: It can accurately identify anomalies and best practices for different entities, and calculate multi-dimensional planned expense ratios; Ultimate Effects: It solves the problem that general models cannot adapt to pharmaceutical scenarios, making analysis conclusions more aligned with actual business needs, helping companies quickly replicate successful experiences and avoid resource waste.

[0098] 3. Anomaly Warning and Suggestion Push for Improved Management Efficiency: This invention reduces the cost of manual data interpretation by intelligently marking anomalies and pushing implementation suggestions. Direct technical effect: Managers can quickly focus on core issues and optimization directions without being immersed in massive amounts of data; Ultimate effect: Significantly improves marketing management efficiency, achieving a closed loop of "problem discovery - cause analysis - optimization implementation," and driving continuous improvement in return on investment.

[0099] 4. Hierarchical Reporting to Adapt to Diverse Management Needs: This invention features hierarchical reports from overall data to detailed information, meeting the needs of managers at different levels. Direct Technical Effects: Senior managers can grasp the overall production status, while lower-level managers can obtain specific optimization actions; Ultimate Effects: Achieving synergy between "top-down control and bottom-up execution," ensuring the effective implementation of resource optimization strategies and improving the overall return on investment of the marketing system.

[0100] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A business data analysis method for the pharmaceutical marketing system, characterized in that, Includes the following steps: Step A. Based on the pre-set systems in the pharmaceutical marketing business, collect multi-source data corresponding to the four main entities of organization, personnel, hospital and product, and the multi-source data includes performance data, expense data, allocation data and interaction data, and then proceed to step B; Step B. Clean the collected multi-source data and build a linkage data model to link the multi-source data of the four subjects (organization, personnel, hospital, and product) based on the four-level unique association keys. Then proceed to step C. Step C. Based on the multi-source associated data of the four subjects (organization, personnel, hospital, and product) under the linkage data model, apply the preset quantitative assessment algorithm for production health for each of the four subjects (organization, personnel, hospital, and product) to quantify and score them according to the dimensions of performance data, cost data, allocation data, and interaction data. Perform quantitative analysis of production health for each subject to obtain a comprehensive production health score for the subject, and then proceed to step D. Step D. Input the comprehensive health scores of the four main entities—organization, personnel, hospital, and product—as well as the multi-source correlation data of the four entities under the linkage data model into the target large language model. Combined with the structured prompts under the preset medical scenario application, drive the target large language model to perform intelligent diagnosis and output a preliminary diagnostic conclusion including anomaly identification, excellent case extraction, comprehensive diagnosis, and optimization suggestions. Then, verify the preliminary diagnostic conclusion through a preset multi-dimensional cross-validation mechanism to obtain the final full-dimensional production diagnosis conclusion.

2. The business data analysis method for the pharmaceutical marketing system according to claim 1, characterized in that: In step A, data collection is performed on a single sales representative, a single doctor, or a single product. Real-time data collection is performed on multi-source data of dynamic data types, and batch data collection is performed on multi-source data of static basic data types. In accordance with the marketing compliance requirements of the pharmaceutical industry, compliance verification is performed on the collected multi-source data.

3. The business data analysis method for the pharmaceutical marketing system according to claim 1, characterized in that: In step A, the indicators of the multi-source data corresponding to the organization include office identification O-ID, performance target and actual value, total cost, SL plan cost, cost ratio, total number and level of customers, doctor interaction rate, regional market capacity, and month-on-month and year-on-year performance data. The various indicators of the multi-source data corresponding to personnel include sales representative P-ID, office identification O-ID, individual performance, SL plan costs and actual expenditures, number and level of assigned doctors, frequency, type and target of interaction, coverage rate of target level doctors, compliance of visits, and assessment cycle data. The indicators of the multi-source data corresponding to the hospital include the hospital H-ID, the office identifier O-ID, the hospital level, the cooperative product Pr-ID, the performance contribution, the SL plan expense ratio, the number of assigned doctors, the number and form of interaction, the participation coverage rate, and the departmental breakdown data. The various indicators of the multi-source data corresponding to the product include product Pr-ID, office identification O-ID, performance target and actual value, achievement rate, SL plan cost, cost ratio, matching hospital level and department, interactive conversion data, and category attributes.

4. The business data analysis method for the pharmaceutical marketing system according to claim 1, characterized in that: Step B includes performing steps B1 to B3 as follows; Step B1. Perform abnormal data rule-based removal, data standardization and unification, and missing value medical business association completion on the collected multi-source data in sequence to clean and update the collected multi-source data, and then proceed to step B2; Step B2. Using the O-ID indicator corresponding to the organization's office, the P-ID indicator corresponding to the personnel, the H-ID indicator corresponding to the hospital, and the Pr-ID indicator corresponding to the product, construct a pre-defined four-level unique association key for the four entities: organization, personnel, hospital, and product, and establish P-ID. O-ID, H-ID O-ID, Pr-ID The inheritance relationship of O-ID and P-ID H-ID, P-ID Pr-ID, H-ID The association relationship of Pr-ID is then determined, and then proceed to step B3; Step B3. Connect the preset multi-source data collection time dimension t, and establish a linked data model according to the following formula; M_{t,(O,P,H,Pr),k}=V_{t,(O,P,H,Pr),k}+Tag_{(O-ID,P-ID,H-ID,Pr-ID)}; In the formula, O represents organization, P represents personnel, H represents hospital, Pr represents product, (O,P,H,Pr) represents the four-subject association combination of organization, personnel, hospital, and product, k represents the k-th indicator under the four-subject association combination of organization, personnel, hospital, and product, Tag_{(O-ID,P-ID,H-ID,Pr-ID) represents the association tag of the four-subject unique association key of organization, personnel, hospital, and product, V_{t,(O,P,H,Pr),k} represents the collected data value of the k-th indicator under the four-subject association combination of organization, personnel, hospital, and product corresponding to the time dimension t, and M_{t,(O,P,H,Pr),k} represents the linkage data matrix of the k-th indicator under the four-subject association combination of organization, personnel, hospital, and product.

5. The business data analysis method for the pharmaceutical marketing system according to claim 1, characterized in that: Step C is based on multi-source associated data of four subjects—organization, personnel, hospital, and product—under the linkage data model. For each of the four subjects (organization, personnel, hospital, and product), the following formula is used: The overall health score of the main entity's production status was calculated; among which, Indicates the first A comprehensive score for the overall health of the production of each entity. , , , The numbers represent the order of the numbers. Preset weights for performance data, expense data, allocation data, and interaction data under each entity. , , , The numbers represent the order of the numbers. Quantitative scoring of performance data, expense data, allocation data, and interaction data under each entity.

6. The business data analysis method for the pharmaceutical marketing system according to claim 5, characterized in that: Step C also includes a comprehensive score of the operational health of the main entities, including the organization, personnel, hospital, and product, as well as quantitative scores of performance data, cost data, allocation data, and interaction data under each entity. The following anomaly identification rules and excellent case selection rules are used to achieve accurate positioning of anomalies and replicable extraction of excellent cases. The rules for identifying anomalies and selecting outstanding cases include: the preset threshold for the overall health score of the abnormal entity, the preset threshold for each core indicator under multi-source correlated data, and the preset threshold for the overall health score of the abnormal entity and the threshold for the target core indicator under multi-source correlated data.

7. The business data analysis method for the pharmaceutical marketing system according to claim 1, characterized in that: In step D, the preliminary diagnostic conclusion is verified through the following preset multi-dimensional cross-validation mechanism to obtain the final full-dimensional production diagnostic conclusion. A quantitative indicator verification mechanism is used to verify whether the abnormal or excellent cases identified in the preliminary diagnostic conclusion meet the preset threshold quantitative judgment rules. If they do, the abnormal or excellent cases identified in the preliminary diagnostic conclusion are retained; otherwise, the corresponding abnormal or excellent cases identified in the preliminary diagnostic conclusion are rejected. The subject association verification mechanism, based on the linkage data model, verifies whether the abnormal root causes analyzed in the preliminary diagnostic conclusion are consistent with the standard quantitative data of the associated subject. If they are, the abnormal root causes analyzed in the preliminary diagnostic conclusion are retained; otherwise, the abnormal root causes analyzed in the preliminary diagnostic conclusion are rejected. The pharmaceutical business logic verification mechanism, based on pharmaceutical marketing business rules and industry compliance requirements, verifies whether the optimization suggestions in the preliminary diagnosis conclusion are consistent with the actual business situation. If they are, the optimization suggestions in the preliminary diagnosis conclusion are retained; otherwise, the optimization suggestions in the preliminary diagnosis conclusion are rejected.

8. The business data analysis method for the pharmaceutical marketing system according to claim 1, characterized in that: It also includes step E as follows: after step D is completed, proceed to step E; Step E. Based on the aforementioned linked data model and the final comprehensive production diagnosis conclusion, generate and display hierarchical, drill-down reports. The reports support drilling down from the organizational level to the personnel level, hospital level, and product level for details. The reports also visualize the comprehensive production health scores of each entity (organization, personnel, hospital, and product), as well as the core indicators, identified anomalies, and optimization suggestions under multi-source related data.

9. A system for implementing the business data analysis method for the pharmaceutical marketing system as described in any one of claims 1 to 8, characterized in that: The system comprises a data acquisition layer, a data processing layer, a core analysis layer, and a diagnostic calibration layer. The data acquisition layer executes step A, collecting multi-source data from four entities—organization, personnel, hospitals, and products—corresponding to pre-defined systems in the pharmaceutical marketing business. The data processing layer executes step B, constructing a linked data model to connect the multi-source data of the four entities. The core analysis layer executes step C, obtaining a comprehensive health score for the operational status of each of the four entities (organization, personnel, hospitals, and products). The diagnostic calibration layer executes step D, applying a target large-scale language model to perform intelligent diagnosis, outputting preliminary diagnostic conclusions, and verifying them through a pre-defined multi-dimensional cross-validation mechanism to obtain a final, comprehensive operational diagnostic conclusion.

10. The system for implementing a business data analysis method for a pharmaceutical marketing system according to claim 9, characterized in that: It also includes an output display layer, which is used to generate and display hierarchical and drill-down reports based on the linked data model and the final full-dimensional production diagnosis conclusion. The reports support drilling down from the organizational level to the personnel level, hospital level, and product level for details, and visually present the comprehensive production health score of each subject of the organization, personnel, hospital, and product, as well as the core indicators, identified anomalies, and optimization suggestions under multi-source related data.

Citation Information

Patent Citations

  • Intelligent enterprise finance and accounting data analysis system and method based on artificial intelligence

    CN120876128A

  • Multi-level-oriented insurance operation index dynamic monitoring and intelligent decision-making system

    CN121684842A

  • Hospital employee performance dynamic accounting and assessment system based on multi-dimensional medical business data fusion

    CN121724508A