Data-driven business intelligence supervision and behavior tracing system and method

By constructing a multidimensional correlation graph and generating a golden diagnosis and treatment path template, the problem of the medical insurance business supervision system's difficulty in identifying abnormal diagnosis and treatment paths has been solved, realizing the rational use of medical insurance funds and the intelligent upgrade of business supervision.

CN121304354BActive Publication Date: 2026-05-05ZHEJIANG HONGZHEN YISHEN DATA TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG HONGZHEN YISHEN DATA TECHNOLOGY CO LTD
Filing Date
2025-10-16
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

The existing medical insurance business supervision system is unable to identify abnormal behaviors in the diverse treatment paths of patients with the same diagnosed disease in real time, resulting in the waste of medical insurance funds. The traditional rule base cannot adapt to the dynamic changes in medical insurance business models and lacks data-driven business intelligence analysis capabilities.

Method used

A multidimensional correlation graph is constructed to generate a golden treatment path template. By comparing the patient's treatment path in real time, abnormal treatment behaviors can be accurately identified. This includes obtaining basic data on the patient's medical services, constructing a multidimensional correlation graph, generating a disease treatment path pool, constructing a treatment trajectory coordinate system, generating a golden treatment path template, and comparing the treatment path of the target patient in real time.

Benefits of technology

It enables real-time monitoring of the treatment pathways of patients with the same diagnosed disease, accurately identifies abnormal behaviors such as excessive examinations and duplicate treatments, improves the efficiency and accuracy of medical insurance business supervision, ensures the rational use of medical insurance funds, and promotes the transformation of medical insurance business from traditional passive response to intelligent prediction and proactive intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121304354B_ABST
    Figure CN121304354B_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing technology and discloses a data-driven business intelligent supervision and behavior tracing system and method. The method includes acquiring basic patient medical service data to construct a multi-dimensional correlation graph; extracting patients with the same diagnosis from the graph to generate a disease treatment path pool; constructing a point set of the treatment trajectory coordinate system to generate a golden treatment path template; matching and aligning the real-time treatment path of the target patient with the template to determine whether it is an abnormal path, and if so, conducting source tracing analysis. This invention achieves accurate identification of abnormal treatment paths in "different treatments for the same disease", effectively reducing the waste of medical insurance funds and improving the efficiency of medical insurance supervision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to a data-driven business intelligent supervision and behavior tracing system and method. Background Technology

[0002] Supervision of medical insurance operations plays a crucial role in maintaining order in the medical service sector and ensuring the security of medical insurance funds. With the increasing volume of medical data and the growing complexity of medical practices, traditional supervisory models face numerous challenges, necessitating intelligent methods to enhance supervisory efficiency and achieve refined management of medical insurance operations. The rational use of medical insurance funds directly impacts the sustainable development of medical insurance operations and the long-term stability of the medical insurance system; therefore, the level of intelligence in the medical insurance supervisory system has become a key indicator for measuring the modernization level of medical insurance operation management.

[0003] Chinese patent CN119048189A discloses an intelligent supervision system, method, and electronic device based on medical insurance data. Through data storage and filtering mechanisms, it provides regulatory support for pharmacies purchasing drugs covered by medical insurance. This system primarily focuses on compliance checks of drug formulation, but lacks in-depth analysis capabilities of the entire medical treatment process, fails to integrate medical insurance supervision with intelligent business decision-making, and cannot provide forward-looking guidance for medical insurance operations from a data-driven perspective.

[0004] Chinese patent CN119579332A discloses a medical insurance outpatient supervision system and method based on fraud anomaly profiling technology. It utilizes modules such as medical transaction data collection, behavior analysis, and risk assessment to achieve accurate identification and prevention of fraudulent behavior. While the system focuses on detecting abnormal behavior, its analysis methods rely on a pre-set rule base, making it difficult to adapt to the dynamic changes in medical insurance business models. It lacks data-driven business intelligence analysis capabilities and cannot achieve real-time intervention in medical insurance operations.

[0005] However, in actual medical insurance operations, patients with the same diagnosis may exhibit diverse treatment paths due to individual differences, the complexity of their conditions, and varying physician judgments. Some of these paths may deviate from reasonable limits, involving excessive testing, duplicate treatments, or treatment evasion, leading to a waste of medical insurance funds. Traditional rule bases, due to their static preset characteristics, struggle to dynamically identify abnormal paths hidden within reasonable differences in treatment. For example, in treatment paths, some unnecessary examinations may be disguised as routine procedures. Lacking real-time dynamic comparison and analysis methods, the monitoring system cannot promptly detect such anomalies, resulting in the loss of medical insurance funds and impacting the sustainable development of medical insurance operations. Medical insurance business supervision urgently needs to shift from "post-event accountability" to "in-process intervention," and upgrade from "static rules" to "dynamic intelligence" to achieve intelligent management of medical insurance business processes. Summary of the Invention

[0006] To overcome the aforementioned deficiencies of existing technologies, this invention provides a data-driven business intelligent supervision and behavior tracing system and method. By constructing a multi-dimensional correlation graph, generating a golden treatment path template, and comparing patient treatment paths in real time, it accurately identifies abnormal treatment behaviors, effectively solving the problem that existing technologies cannot detect abnormal treatment paths under the condition of "different treatments for the same disease" in real time, ensuring the rational use of medical insurance funds, and maintaining the fairness and sustainability of the medical insurance system.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] Data-driven business intelligence supervision and behavior tracing methods include:

[0009] Acquire basic patient medical service data and construct a multi-dimensional correlation graph that comprehensively displays the diagnosis and treatment process;

[0010] Extract patients with the same diagnosis from the multidimensional association graph to generate a disease diagnosis and treatment pathway pool;

[0011] Based on the disease diagnosis and treatment pathway pool, a diagnosis and treatment trajectory coordinate system is constructed to obtain the diagnosis and treatment trajectory point set;

[0012] Generate a golden treatment path template based on the set of treatment trajectory points;

[0013] Obtain the real-time diagnosis and treatment path of the target patient, match and align the real-time diagnosis and treatment path with the golden diagnosis and treatment path template, and determine whether the real-time diagnosis and treatment path of the target patient is an abnormal diagnosis and treatment path;

[0014] If an abnormal treatment pathway is identified, a source tracing analysis will be conducted on the abnormal treatment pathway.

[0015] Furthermore, the patient's basic medical service data includes at least the patient's diagnosis code information, the timeline of treatment items, and medical insurance settlement information;

[0016] The method for constructing the multidimensional association graph includes: constructing a preliminary diagnostic association graph based on patient diagnostic code information; and constructing a multidimensional association graph based on the time sequence records of medical treatment items, medical insurance settlement information, and the constructed preliminary diagnostic association graph.

[0017] Furthermore, the method for constructing a preliminary diagnostic association graph based on patient diagnostic code information includes: extracting diagnostic codes as nodes from patient diagnostic code information; if two diagnostic codes appear simultaneously in more than Y patients and have a sequential order, then adding a directed edge to connect the two diagnostic code nodes, where Y is a preset threshold parameter.

[0018] Furthermore, the method for constructing a multidimensional association graph based on the time-series records of medical treatment items, medical insurance settlement information, and the constructed preliminary diagnostic association graph includes:

[0019] In the preliminary diagnostic association graph, the diagnostic code is the core node. Around the core node, diagnostic item nodes are added based on the time sequence records of diagnostic items and medical insurance settlement nodes are added based on medical insurance settlement information. The diagnostic code core node is connected to the diagnostic item nodes and medical insurance settlement nodes respectively, and the medical insurance settlement nodes are connected to the corresponding diagnostic item nodes.

[0020] Furthermore, the method for extracting co-diagnostic patient groups from a multidimensional association graph includes:

[0021] Determine the target diagnostic code, and construct a diagnostic nearest neighbor set in the multidimensional association graph with the target diagnostic code as the center and set a diagnostic similarity radius r. The diagnostic nearest neighbor set contains all diagnostic codes whose distance from the target diagnostic code does not exceed r.

[0022] Collect diagnostic records corresponding to all diagnostic codes in the diagnostic nearest neighbor set to form the original diagnostic group; based on the multidimensional association graph, screen the original diagnostic group for complications and comorbidities to obtain the pure diagnostic group;

[0023] Calculate the diagnostic purity index of the pure diagnostic group;

[0024] If the diagnostic purity index is greater than the preset purity threshold, the pure diagnostic group will be identified as the same diagnostic patient group.

[0025] Furthermore, the method for calculating the diagnostic purity index of the pure diagnostic group includes:

[0026] The number of diagnostic records corresponding to all diagnostic codes in the nearest neighbor set of diagnoses is counted and marked as the total number of diagnostic records. The number of diagnostic records of complications and comorbidities that are screened out is counted. The diagnostic purity index is equal to the difference between the total number of diagnostic records and the number of diagnostic records of complications and comorbidities that are screened out, and then divided by the total number of diagnostic records.

[0027] Furthermore, the method for generating the disease diagnosis and treatment path pool is as follows: based on the same diagnosis patient group, extract the time sequence record of each patient's diagnosis and treatment items to generate the disease diagnosis and treatment path pool.

[0028] Furthermore, the method for constructing a treatment trajectory coordinate system based on the disease treatment pathway pool to obtain the treatment trajectory point set includes:

[0029] Based on the time sequence records of patients' treatment items in the disease treatment pathway pool, a standardized treatment trajectory vector V' is constructed.

[0030] The time point of the first diagnosis and treatment is obtained from the time sequence record of the diagnosis and treatment project. The time point of the first diagnosis and treatment is used as the origin O to construct the coordinate system of the diagnosis and treatment trajectory. The standardized diagnosis and treatment trajectory vector V' is mapped to the coordinate system of the diagnosis and treatment trajectory to obtain the set of diagnosis and treatment trajectory points.

[0031] Furthermore, the timeline of the diagnosis and treatment items includes at least the diagnosis and treatment item code, the cost of the diagnosis and treatment item, and the specific implementation time;

[0032] The method for constructing a standardized treatment trajectory vector V' based on the time sequence records of patients' treatment items in the disease treatment pathway pool includes:

[0033] Using the initial treatment time as the baseline t0, calculate the time interval sequence Δt between the specific implementation time of each subsequent treatment item and t0;

[0034] Map the diagnostic and treatment item codes to standardized diagnostic and treatment item codes to generate a standardized diagnostic and treatment item sequence P;

[0035] Calculate the ratio of the cost of each treatment item to the average cost of the same item in the same region during the same period, and generate a series of standardized cost coefficients C;

[0036] The time interval sequence Δt, the standardized treatment item sequence P, and the cost standardization coefficient sequence C are combined into a three-dimensional treatment trajectory vector V=[Δt,P,C]. The three-dimensional treatment trajectory vector V is then subjected to dimension normalization to obtain the standardized treatment trajectory vector V'.

[0037] A data-driven business intelligence supervision and behavior tracing system is used to implement the aforementioned data-driven business intelligence supervision and behavior tracing method. The system includes:

[0038] The graph construction module is used to acquire basic data on patients' medical services and construct a multi-dimensional relational graph that comprehensively displays the diagnosis and treatment process.

[0039] Disease diagnosis and treatment pathway pool generation module: used to extract patients with the same diagnosis from the multidimensional association graph and generate a disease diagnosis and treatment pathway pool;

[0040] Treatment trajectory generation module: used to construct a treatment trajectory coordinate system based on the disease treatment path pool, and obtain the treatment trajectory point set;

[0041] Path template generation module: Generates a golden treatment path template based on the set of treatment trajectory points;

[0042] Abnormal Path Judgment Module: Used to obtain the real-time diagnosis and treatment path of the target patient, match and align the real-time diagnosis and treatment path with the golden diagnosis and treatment path template, and determine whether the real-time diagnosis and treatment path of the target patient is an abnormal diagnosis and treatment path;

[0043] Source tracing analysis module: used to conduct source tracing analysis on abnormal diagnosis and treatment pathways.

[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0045] This invention comprehensively integrates patient medical service data by constructing a multi-dimensional correlation graph, providing a complete data foundation for medical insurance business supervision. Based on this, a disease diagnosis and treatment path pool and a golden diagnosis and treatment path template are generated, accurately extracting mainstream diagnosis and treatment business models and forming a dynamic, data-driven intelligent business supervision benchmark. Through real-time matching and alignment of diagnosis and treatment paths with templates, real-time monitoring of the diagnosis and treatment paths of patients with the same diagnosed disease is achieved. In particular, it can accurately identify abnormal business behaviors such as excessive examinations and duplicate treatments hidden in reasonable treatment differences, avoiding waste of medical insurance funds due to abnormal treatment paths. This invention not only breaks through the static limitations of traditional rule bases, achieving real-time intervention "in progress," but also generates automatic "explanation suggestions" for the responsible physician's path, truly promoting a "quasi-real-time self-discipline + external discipline dual-wheel" approach to medical insurance business. From a data-driven business intelligence perspective, this invention transforms medical insurance supervision from a traditional passive response model to a business model of intelligent prediction and proactive intervention, significantly improving the efficiency and accuracy of medical insurance business supervision and ensuring the rational use of medical insurance funds. This invention provides decision support for medical insurance regulatory departments through data-driven business intelligence analysis, enabling medical insurance supervision to shift from "experience-driven" to "data-driven", from "outcome-based" to "process-based", and from "single-dimensional" to "multi-dimensional correlation". It provides data-driven intelligent decision support for the sustainable development of medical insurance business and realizes the intelligent upgrade of medical insurance business supervision. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a flowchart illustrating the principle of the data-driven intelligent business supervision and behavior tracing method in this invention.

[0048] Figure 2 This is a schematic diagram of the preliminary diagnostic correlation diagram constructed in an embodiment of the present invention;

[0049] Figure 3 This is a flowchart illustrating the principle of obtaining a disease diagnosis and treatment pathway pool based on a pure diagnostic group, as described in this invention.

[0050] Figure 4 This is a flowchart illustrating the principle of the present invention for determining whether a target patient's real-time treatment path is an abnormal treatment path;

[0051] Figure 5 This is a functional block diagram of the data-driven business intelligent supervision and behavior tracing system of the present invention. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] Example 1

[0054] Please see Figure 1 As shown, this embodiment provides a data-driven business intelligent supervision and behavior tracing method, including:

[0055] Step S10: Obtain basic patient medical service data and construct a multi-dimensional correlation graph that comprehensively displays the diagnosis and treatment process;

[0056] The basic data for patient medical services includes the patient's unique identifier, patient's diagnosis code, timeline of treatment items, and medical insurance settlement information; the timeline of treatment items includes the treatment item code, specific implementation time, treatment item cost, and information on the implementing department and physician;

[0057] Patient medical service basic data is obtained through interfaces such as the Health Information System (HIS) and the medical insurance settlement system. This includes patient unique identification information, patient diagnostic codes, treatment item timeline records, and medical insurance settlement information. Among these, patient unique identification information, such as the personal identification code linked to the medical insurance card number, ensures the uniqueness of patient identity and avoids data confusion. Patient diagnostic codes are based on the International Classification of Diseases (ICD) standard, providing a unified framework for diagnostic categories. Treatment item timeline records include treatment item codes, specific implementation times (when the item started or occurred), treatment item costs, implementing hospital, department, and physician information, recording the time sequence of treatment behavior and the responsible party, comprehensively reflecting the patient's treatment process. The collection of physician information allows subsequent anomaly tracing to directly locate the specific implementing entity, achieving precise "behavior-personnel" correlation and overcoming the limitations of traditional supervision that can only track institutional-level issues. Medical insurance settlement information covers reimbursement amounts, co-payment ratios, and pooled fund payment amounts, linking treatment behavior with the flow of medical insurance funds.

[0058] The collected basic patient medical service data is cleaned and formatted. For example, the custom codes for diagnosis and treatment items used by different medical institutions, such as "A123" for a blood routine test at a certain hospital, are mapped to the national standard code, such as "B456". This achieves standardized integration of cross-institutional medical data, ensuring consistency in the codes for diagnosis and treatment items across different medical institutions and eliminating data confusion and analytical obstacles caused by coding differences. Disorganized time records, such as "2024 / 1 / 19:00", are converted to the standard format "YYYY-MM-DDHH:MM:SS" to ensure the accuracy of time-series analysis. Duplicate records are filtered, such as multiple registrations of the same patient for the same item on the same day, and logically contradictory data, such as treatment time earlier than diagnosis time, improving data quality. Traditional medical insurance supervision data is scattered and formatted inconsistently. Differences in coding systems among different institutions prevent data fusion, directly affecting the accuracy of subsequent analysis. Multi-source data collection covers all elements of the diagnosis and treatment process, providing a complete dataset for full-chain supervision; standardized processing eliminates data silos, making cross-institutional data comparable.

[0059] The specific process of constructing a multidimensional association graph that comprehensively displays the diagnosis and treatment process includes: constructing a preliminary diagnosis association graph based on the patient's diagnosis code information in the patient's basic medical service data; and constructing a multidimensional association graph that comprehensively displays the diagnosis and treatment process based on the time sequence records of diagnosis and treatment items, medical insurance settlement information, and the constructed preliminary diagnosis association graph.

[0060] The method for constructing a preliminary diagnostic association graph based on patient diagnostic code information includes: extracting diagnostic codes as nodes from patient diagnostic code information; if two diagnostic codes appear simultaneously in more than Y patients and have a sequential order, then adding a directed edge to connect the two diagnostic code nodes, where Y is a preset threshold parameter.

[0061] Specifically, the patient diagnostic code information includes diagnostic codes based on the ICD-10 / ICD-11 standards, such as I10 and I25. These diagnostic codes are used as graph nodes to represent disease diagnostic entities. The co-occurrence frequency and order of any two diagnostic codes in the patient population are counted (e.g., the number of patients diagnosed with A before B). When the number of co-occurring patients exceeds a preset threshold parameter Y, a directed edge from A to B is added, representing the disease progression relationship. Y is a preset threshold parameter used to filter accidental associations and control the sparsity of the preliminary diagnostic association graph. A larger Y value requires more patients to have both diagnostic codes simultaneously to form an association, resulting in a sparser graph and more reliable associations. A smaller Y value requires fewer patients to co-occur and form an association, resulting in a denser graph but potentially containing more accidental associations. Y is an empirically set integer threshold, preferably 30, to ensure the significance of co-occurrence relationships. The value of Y can be adjusted according to the number of patients with the target disease. Generally, if the sample size exceeds 500 people, Y can be set to 30; if the sample size is small, Y can be appropriately reduced to 10-20 to ensure that there are still enough samples to form a reliable relationship.

[0062] For example, please refer to Figure 2 As shown, taking the ICD-10 coding system as an example, I10 represents primary hypertension, I25 represents coronary heart disease, and I50 represents heart failure. With I10, I25, and I50 as three nodes, among 1000 patients, 320 patients were first diagnosed with hypertension (I10) and then developed coronary heart disease (I25), 180 patients progressed from hypertension (I10) to heart failure (I50), and 90 patients deteriorated from coronary heart disease (I25) to heart failure (I50). If the number of these associated patients exceeds the preset threshold parameter Y (e.g., set to 30), three directional disease progression paths (directed edges) are generated: from I10 to I25 (hypertension leading to coronary heart disease), from I10 to I50 (hypertension leading to heart failure), and from I25 to I50 (coronary heart disease secondary to heart failure), forming a preliminary diagnostic association graph.

[0063] Traditional rule bases can only identify pre-defined fixed "diagnosis-diagnosis" associations (e.g., hypertension is always accompanied by diabetes), failing to dynamically discover data-driven patterns of real disease progression. The preliminary diagnostic association graph, through data-driven association modeling, automatically mines temporal associations between diagnoses, supplementing potential associations not explicitly defined in traditional rule bases and improving the adaptability of regulatory models. The directionality of directed edges accurately reflects the logic of disease progression, avoiding confusion of causal relationships, such as misclassifying complications as underlying diseases. A dynamically adjustable Y-threshold mechanism allows the preliminary diagnostic association graph to adapt to the regulatory needs of different disease domains—for rare diseases (small patient base), the Y-value is lowered to retain necessary associations; for common diseases (e.g., hypertension), the Y-value is increased to filter noise. This dynamic adjustment flexibility is difficult for traditional fixed rule bases to achieve in terms of association filtering elasticity. The preliminary diagnostic association graph provides the first layer of filtering rules for anomaly identification such as "different treatments for the same disease"—only paths conforming to the logic of real disease progression are included in the normal range, significantly reducing the search space for anomaly detection and improving the efficiency and accuracy of real-time supervision.

[0064] The specific process of constructing a multidimensional association graph that comprehensively displays the diagnosis and treatment process based on the time sequence records of diagnosis and treatment items, medical insurance settlement information, and the constructed preliminary diagnosis association graph includes: based on the time sequence records of diagnosis and treatment items and medical insurance settlement information, in the constructed preliminary diagnosis association graph, taking the diagnosis code as the core node, adding diagnosis and treatment item nodes and medical insurance settlement nodes around the core node to construct a multidimensional association graph that comprehensively displays the diagnosis and treatment process; the diagnosis and treatment item nodes are established based on the time sequence records of diagnosis and treatment items, and the medical insurance settlement nodes are established based on the medical insurance settlement information; calculating the association weights between each node in the multidimensional association graph.

[0065] In the multidimensional association graph, the diagnostic code core node is connected to the treatment item node and the medical insurance settlement node, respectively. The medical insurance settlement node is connected to the corresponding treatment item node. Treatment sequence edges are added according to the order of implementation of the treatment items. The order of implementation of the treatment items is determined according to the specific implementation time in the treatment item sequence.

[0066] Multidimensional correlation graphs are used to achieve structured correlation and integration of medical behavior data and medical insurance fund payment data. This breaks through the technical bottleneck of existing technologies where medical behavior trajectories and fund flow data are independent and cannot be cross-dimensionally correlated. It enables the medical insurance supervision system to conduct linked analysis on the rationality of diagnosis, the necessity of treatment items, and the compliance of medical insurance payments based on a unified full-chain data model, thereby improving the ability to identify abnormal behaviors across links such as excessive examinations and illegal reimbursements.

[0067] Information such as treatment item codes and specific implementation times are extracted from the time-series records of treatment items. Each independent treatment item is abstracted as a treatment item node in the graph. For example, "blood pressure monitoring" and "drug therapy" are treatment item nodes. The attributes of each treatment item node include the treatment item code, specific implementation time, treatment item cost, implementing hospital, department, and physician information. Through treatment item nodes, the multidimensional association graph can intuitively display the treatment measures for a specific diagnosis (such as hypertension), and reflect the logical sequence of treatment through time-series edges (such as "blood pressure monitoring → drug therapy"). Each diagnostic code core node is connected to the corresponding treatment item node. For example, the hypertension diagnostic code core node is connected to treatment item nodes such as blood pressure monitoring and drug therapy, indicating that these are treatment measures for this diagnosis.

[0068] Data such as reimbursement amounts and payment ratios are extracted from medical insurance settlement information, and each medical insurance settlement record related to medical treatment is abstracted into a medical insurance settlement node in a graph. For example, a medical insurance payment record for a blood pressure monitoring session is an independent medical insurance settlement node, and the attributes of the medical insurance settlement node include payment amount, co-payment ratio, etc. Medical insurance settlement nodes extend from the core node of the diagnostic code, reflecting the direct link between diagnosis and medical insurance payment. For example, the core node of the hypertension diagnostic code connects to the node recording its medical insurance payment information. Medical insurance settlement nodes are connected to corresponding treatment item nodes, clearly indicating the correspondence between treatment items and medical insurance payments. For example, the treatment item node for a blood pressure monitoring session corresponds to the settlement node recording the medical insurance payment for that session. Simultaneously, based on the order of implementation of treatment items, the order is determined according to the specific time in the treatment item timeline, and treatment timeline edges are added. For example, if blood pressure monitoring occurs before medication, there is a timeline edge pointing from the blood pressure monitoring node to the medication treatment node, reflecting the time sequence and logical order of treatment. Through such connections, diagnostic codes, treatment items, and medical insurance settlements are closely linked, comprehensively showcasing the treatment process and allowing medical insurance regulatory departments to clearly understand the relationship between medical behavior and the flow of medical insurance funds.

[0069] The association weights between nodes in the multidimensional association graph include the association weights between the diagnostic code node and the treatment item node, the association weights between the diagnostic code node and the medical insurance settlement node, and the weights between the treatment item node and the medical insurance settlement node.

[0070] The association weight between diagnostic code nodes and treatment item nodes is used to determine the relevance of treatment items to a specific diagnosis. It is calculated using two factors: First, the usage frequency factor. This involves counting the total number of patients diagnosed with the diagnostic code (marked as the number of patients with treatment codes) and the number of patients who actually received the treatment item (marked as the number of patients with treatment items). The ratio of the number of patients with treatment items to the number of patients with diagnostic codes is the usage frequency factor. A higher ratio indicates that the treatment item is more common in the treatment of the disease. For example, if most hypertensive patients receive blood pressure monitoring, a high frequency factor indicates a strong association between the two. Second, the clinical relevance factor. This involves reviewing authoritative medical clinical pathways, treatment guidelines, and literature to assess the recommended value of treatment items in treating or diagnosing the disease. For example, a value of 1 is assigned to first-line recommendations, 0.5 to second-line recommendations, and 0 to non-recommendations. The usage frequency factor and the clinical relevance factor are then weighted and summed according to a preset weight ratio (e.g., 50% each) to obtain the association weight. The association weight between diagnostic code nodes and treatment item nodes, combined with usage frequency and clinical guidelines, can accurately determine the relevance between treatment items and diagnoses, avoid treating occasional treatment items as routine, and identify treatment items that do not comply with guidelines. For example, if a treatment item is used infrequently and is not recommended by clinical guidelines, its weight will be low, which may indicate that the item is abnormal. This helps to accurately identify violations such as over-examination and duplicate treatment.

[0071] The association weight between the diagnostic code node and the medical insurance settlement node reflects the strength of the relationship between diagnosis and medical insurance payment. It is considered from two perspectives: first, the medical insurance payment ratio, which is the proportion of the total cost paid by medical insurance for all treatment items directly related to the diagnostic code; a higher ratio indicates more comprehensive medical insurance coverage and a closer association. Second, the stability of medical insurance payment; this is calculated by first determining the standard deviation of the medical insurance payment amount related to the diagnosis, then dividing it by the average to obtain the relative volatility value. Subtracting this relative volatility value from 1 represents stability; lower volatility indicates higher stability. The weighted average of these two factors yields the association weight between the diagnostic code node and the medical insurance settlement node. The association weight between the diagnostic code node and the medical insurance settlement node, considering both payment ratio and stability, reflects the degree of medical insurance support for diagnosis and the standardization of fees. A high and stable payment ratio indicates strong medical insurance support and standardized medical services; a low payment ratio or poor stability may indicate irregularities in charging, helping regulatory authorities to promptly identify anomalies in medical insurance fund payments and avoid fund waste.

[0072] The weighting of the treatment item node and the medical insurance settlement node focuses on the rationality of the treatment item and its medical insurance payment. This involves two aspects: first, the proportion of medical insurance payment for the item, calculated by dividing the amount paid by medical insurance for the treatment item by the total cost of the item; a higher proportion indicates greater medical insurance support. Second, cost variability, which involves collecting cost data for the treatment item from multiple similar patients, calculating the standard deviation and mean, and obtaining the cost variability (the ratio of standard deviation to mean). Subtracting this cost variability from 1 represents cost stability; a higher variability indicates lower cost stability and potential anomalies. The weighting of these two factors yields the weight between the treatment item node and the medical insurance settlement node. By using the payment proportion and cost variability, the rationality of individual treatment item costs can be assessed. For example, a low payment proportion and high cost variability may indicate unreasonable charges, supplementing the shortcomings of traditional regulatory oversight of individual treatment item costs.

[0073] The correlation weights between each node are calculated, quantifying the strength of the correlation between diagnosis and treatment items, diagnosis and medical insurance payment, and treatment items and medical insurance payment. By utilizing indicators such as medical insurance payment stability and cost variability, anomalies can be deeply explored from a cost perspective, revealing abnormal behaviors that are difficult to detect using traditional methods, such as situations where cost fluctuations are abnormal but treatment items appear reasonable.

[0074] Step S20: Extract the same diagnosis patient group from the multidimensional association graph and generate a disease diagnosis and treatment path pool;

[0075] The specific process of extracting patients with the same diagnosis from the multidimensional association graph and generating a disease diagnosis and treatment path pool includes: determining the target diagnostic code; constructing a diagnostic nearest neighbor set in the multidimensional association graph with the target diagnostic code as the center and setting a diagnostic similarity radius r; the diagnostic nearest neighbor set contains all diagnostic codes that are no more than r away from the target diagnostic code; collecting the diagnostic records corresponding to all diagnostic codes in the diagnostic nearest neighbor set to form the original diagnostic group; counting the number of diagnostic records corresponding to all diagnostic codes in the diagnostic nearest neighbor set and marking it as the total number of diagnostic records; based on the multidimensional association graph, screening the original diagnostic group for complications and comorbidities to obtain a pure diagnostic group; counting the number of diagnostic records for complications and comorbidities that were screened out; and obtaining the disease diagnosis and treatment path pool based on the pure diagnostic group.

[0076] After determining the target diagnostic code, a diagnostic similarity radius *r* is set based on the hierarchical structure and semantic similarity of ICD encoding. A diagnostic nearest neighbor set is then constructed, containing all diagnostic codes whose distance to the target diagnostic code does not exceed *r*. *r* measures the coding proximity between two diagnostic codes, specifically determined by whether their ICD encoding prefixes are identical. For example, if the first two characters of two diagnostic codes are the same (e.g., I10 and I11), the distance is considered to be 1; if the first character is the same but the second character is different (e.g., I10 and I25), the distance is considered to be 2.

[0077] For example, setting the diagnostic similarity radius r=1 and the target diagnostic code to I10 (primary hypertension), based on the ICD-10 cardiovascular system classification hierarchy, I11 (hypertensive heart disease), I15 (secondary hypertension), etc., whose coding distance is within the range of r, are included in the nearest neighbor set. By expanding the scope of similar diagnoses, a candidate patient group covering similar pathological features is formed. This process quantifies the semantic distance between diagnoses through hierarchical matching of coding structures (such as major categories, subcategories, and item levels), expanding the scope of clinical relevance analysis of patient groups and improving the rationality of dividing patient groups with the same diagnosis. For example, the diagnosis of hypertension combined with cardiac complications can be included in the same analysis category, improving the clinical relevance of patient groups.

[0078] Diagnostic records from all diagnostic codes in the nearest neighbor set are collected to form an initial diagnostic group. Complications and comorbidities are then screened out from this initial group based on a multidimensional association graph. This is achieved through two dimensions: First, the association weight between diagnostic code nodes and treatment item nodes is used to determine complications. For example, if the weight of the diagnostic code "diabetic nephropathy" with the treatment item node "renal function test" is higher than a preset threshold (e.g., 0.7), and the renal function test is performed after the diagnosis of diabetes in the treatment sequence, then, based on clinical guidelines, it is inferred that "diabetic nephropathy" is a complication of diabetes, and this diagnostic record is retained. Second, irrelevant comorbidities are eliminated based on the co-occurrence frequency between diagnostic code nodes. For example, if the co-occurrence frequency of the diagnosis of "cold" in hypertensive patients is lower than the frequency threshold Y' (e.g., 5%), it is considered to have no direct association with the core diagnosis and is screened out. Through the above screening, a pure diagnostic group is obtained, excluding interference from secondary diagnoses. The number of screened records is counted, providing a data basis for purity calculation.

[0079] Please see Figure 3 As shown, the specific process of obtaining the disease treatment pathway pool based on the pure diagnostic group includes: calculating the diagnostic purity index of the pure diagnostic group; if the diagnostic purity index is greater than the preset purity threshold, the pure diagnostic group is identified as a patient group with the same diagnosis; based on the patient group with the same diagnosis, the time sequence records of the treatment items of each patient are extracted to generate the disease treatment pathway pool.

[0080] The diagnostic purity index is used to assess the similarity and relevance between the pure diagnostic group and the target diagnostic code of the center. The calculation formula is: Diagnostic Purity = (Total number of diagnostic records - Number of comorbidity and comorbidity diagnostic records removed) / Total number of diagnostic records. For example, if the total number of diagnostic records is 100 and the number of comorbidity and comorbidity diagnostic records removed is 20, then the diagnostic purity index is (100-20) / 100 = 0.8. If this index is greater than the preset purity threshold (e.g., 0.75), it indicates that the pure diagnostic group has a high similarity and relevance to the center's diagnostic code, the data quality is good, and it can accurately reflect the diagnosis and treatment of the disease. In this case, the pure diagnostic group is identified as a patient group with the same diagnosis. This ensures that the diagnostic information of the selected patient group has sufficient purity, providing a reliable data foundation for subsequent diagnosis and treatment pathway analysis. If the preset purity threshold is not reached, the diagnostic similarity radius r needs to be adjusted or the screening rules need to be re-evaluated until the purity requirements are met. Based on the time-series records of treatment items for patients with the same diagnosis, a disease treatment path pool is generated, which includes information such as treatment item code, specific implementation time, treatment item cost, implementing department, and physician, forming the basic dataset for subsequent path analysis.

[0081] By setting the diagnostic similarity radius *r* and matching ICD coding levels, this approach overcomes the limitations of traditional medical insurance supervision that relies solely on a single diagnostic code to segment patient groups. It allows for the inclusion of clinically relevant similar diagnoses within the analysis scope; for example, hypertension and its complications can be treated as a whole. This avoids the fragmentation of similar patients due to subtle differences in diagnostic codes, improving the clinical rationality of patient group segmentation and resulting in a larger effective sample size. Utilizing the weights and co-occurrence frequencies of multidimensional association graphs for comorbidity and comorbidity screening, compared to traditional rule-based methods with fixed exclusion lists, dynamically identifies secondary diagnoses with genuine clinical relevance to the target diagnosis. For example, it judges the rationality of complications based on treatment item weights, rather than simply excluding all accompanying diagnoses. It effectively filters out incidentally irrelevant diagnoses, such as occasional colds in hypertensive patients, while retaining pathologically relevant comorbidity diagnoses, such as hypertensive heart disease. This allows the final patient group with the same diagnosis to focus more on the core treatment pathway of the "same disease," reducing path analysis bias caused by mixed diagnoses and improving the homogeneity of the patient group.

[0082] The introduction of diagnostic purity indices provides a quantitative evaluation standard for the effectiveness of patient populations. By setting a purity threshold (e.g., 0.75), it ensures that the proportion of core diagnoses in the included patient data reaches a preset level, preventing the subsequent path pool from failing to reflect mainstream treatment patterns due to excessive extraneous diagnoses. For example, a diagnostic purity index of 0.8 indicates that 80% of diagnostic records belong to the target diagnosis and its reasonable complications, while 20% are irrelevant diagnoses that have been screened out. This quantitative evaluation mechanism provides a traceable technical basis for data quality control. By combining the semantic similarity of ICD codes with the weights of treatment items, not only is the population at the diagnostic level expanded, but the inherent correlation between diagnosis and treatment behavior is also implicitly explored. For example, by associating a certain diagnostic code with a high weight of a specific treatment item, the clinical rationality of that diagnosis as a complication can be verified in reverse. This cross-dimensional data correlation analysis provides a new perspective for medical insurance supervision.

[0083] Without screening and purity assessment of the patient population, the pathway pool generated directly from the original diagnostic data may contain a large number of irrelevant treatment behaviors. This can lead to the inclusion of abnormal pathway features in the subsequently constructed optimal treatment pathway template, distorting the matching results between the real-time treatment pathway and the template. Screening and purity assessment ensure that the treatment trajectories in the pathway pool primarily reflect the reasonable treatment patterns for the target diagnosis. This allows the optimal treatment pathway template to accurately capture the time series, project combinations, and cost characteristics of mainstream pathways, providing a reliable benchmark for identifying abnormal pathways and improving the accuracy of the optimal treatment pathway template.

[0084] Step S30: Based on the disease diagnosis and treatment path pool, construct a diagnosis and treatment trajectory coordinate system to obtain the diagnosis and treatment trajectory point set;

[0085] The specific process of constructing a treatment trajectory coordinate system based on the disease treatment pathway pool and obtaining the treatment trajectory point set includes: constructing a standardized treatment trajectory vector V' based on the time sequence records of the patient's treatment items in the disease treatment pathway pool; constructing a treatment trajectory coordinate system with the first treatment time point as the origin O; mapping the standardized treatment trajectory vector V' onto the treatment trajectory coordinate system to obtain the treatment trajectory point set T.

[0086] The specific process of constructing a standardized treatment trajectory vector V' based on the time sequence records of patients' treatment items in the disease treatment pathway pool includes: obtaining the initial treatment time point from the treatment item time sequence records; using the initial treatment time point as the reference point t0, calculating the time interval sequence Δt={Δt1,Δt2,...,Δt...} of the specific implementation time of each subsequent treatment item relative to t0. n}, where n represents the total number of medical procedures performed by the patient during the diagnosis and treatment process, and Δt nThis represents the time interval between the nth treatment item and the initial treatment time point t0; the treatment item codes are mapped to standardized treatment item codes, generating a standardized treatment item sequence P={p1,p2,...,p...} n}, p n The standardized treatment item code represents the nth treatment item; calculate the ratio of the cost of each treatment item to the average cost of the same item in the same region during the same period, and generate a cost standardization coefficient sequence C={c1,c2,...,c n}, c n It is the cost standardization coefficient of the nth treatment item; the time interval sequence Δt, the standardized treatment item sequence P, and the cost standardization coefficient sequence C are combined into a three-dimensional treatment trajectory vector V=[Δt,P,C]. The three-dimensional treatment trajectory vector V is subjected to dimension normalization to eliminate the dimensional differences between different dimensions, and the standardized treatment trajectory vector V' is obtained.

[0087] Using the patient's first treatment time t0 as the baseline, calculate the time interval sequence Δt={Δt1,Δt2,...,Δt} for each subsequent treatment item. n}, where n represents the total number of medical procedures performed by the patient during the diagnosis and treatment process, and Δt n This represents the time interval between the nth treatment item and the initial treatment time t0. For example, if a patient's initial treatment time is 9:00 AM on March 1, 2025 (t0), followed by blood tests (10:30 AM on March 1, Δt1 = 1.5 hours) and a CT scan (2:00 PM on March 2, Δt2 = 29 hours), then the Δt sequence is {1.5, 29, ...}. By quantifying the time difference, the time dimension of the treatment process is transformed into a numerical sequence, realizing the numerical quantification of the treatment time sequence and solving the problem of the difficulty in accurately quantifying and comparing the time sequence in traditional monitoring. By converting the treatment time sequence into a calculable time interval through the time interval sequence Δt, the "sequence of treatment steps" and the "time consumption of each step" can be accurately captured. For example, if a patient with hypertension only begins their first drug treatment 72 hours after diagnosis, their Δt is much greater than the average for similar patients. This time deviation can be visually identified through the x-axis coordinate in the coordinate system, enabling real-time quantitative monitoring of the timeliness of diagnosis and treatment and solving the problem that traditional manual review is difficult to quantify time differences.

[0088] The process maps medical service item codes to national standard codes. For example, mapping a hospital's "electrocardiogram" code "C001" to "M123" in the medical insurance catalog generates a standardized medical service item sequence P. This process eliminates differences in coding systems among different medical institutions. For instance, different hospitals may use different internal codes for "blood routine tests," but after mapping, they are unified as "B456," ensuring consistency of cross-institutional data. This establishes a unified cross-institutional medical service item identification system, ensuring data consistency across institutions and avoiding missed or misjudged items due to coding differences.

[0089] "Contemporary period" refers to the time period within the same statistical cycle as the target patient's treatment behavior, usually based on a natural time range, such as "first quarter of 2025" or "same year," or a medical regulatory cycle, such as the medical insurance settlement year, to ensure the time dimension comparability of cost data. "Same region" refers to the medical institution where the target patient sought treatment belonging to the same medical insurance pooling area or geographical administrative region, ensuring that cost data are affected by the same economic level and medical insurance policies. "Same item" refers to the treatment item with the same standardized code as the target treatment item, such as both being "B456 blood routine," ensuring consistency across institutions. C reflects the ratio of the actual cost of this treatment item to the average cost of the same item in the same period and region. The purpose of calculating C is to eliminate the influence of cost differences between different regions and different medical institutions, making the cost data of treatment items comparable. For example, if the average cost of the same item in the same period in a certain region is 100 yuan, and the actual cost of a patient's nth treatment item is 120 yuan, then c... n That is 1.2; if the other patient's cost is 80 yuan, then c n It's 0.8. This standardized coefficient provides a more intuitive understanding of the relative cost of each patient's medical services, making it easier to identify abnormal spending behavior.

[0090] Combine Δt, P, and C into a three-dimensional treatment trajectory vector V=[Δt,P,C], and then perform normalization processing, such as min-max standardization or Z-score standardization, to eliminate the dimensional differences between time, project code, and cost coefficient, thus obtaining a standardized treatment trajectory vector V'.

[0091] The method for mapping the standardized treatment trajectory vector V' to the treatment trajectory coordinate system is as follows: The time interval sequence Δt in the standardized treatment trajectory vector V' is mapped to the x-axis of the treatment trajectory coordinate system, with the unit being a standardized time unit. For example, days can be used as the baseline. Normalization is applied to map time differences of different magnitudes to the [0,1] interval or retain the original numerical units. For instance, if a patient's first treatment time is t0 = June 1, 2025, 09:00, and the Δt values ​​for subsequent treatments are 1.5 hours and 29 hours respectively, then the x-axis coordinates correspond to Δt1 = 0.0625 days and Δt2 = 1.2083 days. The standardized treatment item sequence P is mapped to the y-axis, and the original treatment item codes are mapped to unique numerical identifiers, such as using sequence number coding or unique hot coding. For example, "blood pressure monitoring" corresponds to code B456, and "drug therapy" corresponds to code M123, represented as discrete numerical points in the coordinate system. For example, B456 is mapped to 5, and M123 is mapped to 10. The cost standardization coefficient sequence C is mapped to the z-axis and is a dimensionless value, reflecting the ratio of actual cost to the average cost in the same period and region. For example, if the actual cost of a certain medical procedure is 1.2 times the average cost, then the z-axis coordinate is 1.2.

[0092] The diagnostic and treatment trajectory coordinate system transforms the abstract time-series records of diagnostic and treatment items into a computable three-dimensional spatial trajectory, enabling the differences in "different treatments for the same disease" to be quantified through spatial distance, breaking through the limitations of traditional rule bases in capturing complex path deviations.

[0093] Step S40: Generate a golden treatment path template based on the treatment trajectory point set T;

[0094] The method for generating a golden treatment path template based on the treatment trajectory point set T includes: calculating the local density ρ of each point in the treatment trajectory point set T, and identifying the point T with the highest density. max ; Calculate the distance δ between each point in the treatment trajectory point set T and the point with the highest density; Identify core trajectory points based on local density ρ and distance δ; Connect the core trajectory points in the treatment trajectory coordinate system to form the main path; Incorporate the treatment trajectory points around each core trajectory point whose local density is greater than the preset density threshold into the corresponding branch path to form the golden treatment path template.

[0095] Local density ρ reflects the degree of clustering of similar points around each treatment trajectory point. The method for calculating local density is as follows: for each trajectory point T in the treatment trajectory point set T... i Count the number of neighboring points within a preset radius r1; this number is T. i Local density ρ i1 ≤ i ≤ n. The radius r1 is set based on the spatial resolution of the treatment trajectory coordinate system, usually taking the mean or standard deviation of the Euclidean distance of all trajectory points, for example, the composite distance of 0.5 standardized time units (e.g., 0.5 days) and 1 treatment item coding interval. Areas with high local density mean that many patients underwent similar treatments at similar time points, and the cost levels were also relatively similar, indicating that this may be a common and reasonable treatment pattern. After obtaining the local density of each trajectory point, find the point T with the highest density. max Then, calculate T for each of the other points. i With T max Euclidean distance δ in the treatment trajectory coordinate system i Distance δ reflects the degree to which each treatment trajectory point deviates from the most common treatment pattern.

[0096] The method for identifying core trajectory points based on local density ρ and distance δ is as follows: Plot each trajectory point on a decision map with ρ as the x-axis and δ as the y-axis. Points meeting the decision criteria are designated as core trajectory points. The decision criteria are: local density ρ is greater than k1 times the average density of all points, and distance δ is greater than k2 times the average distance of all points. k1 is a set density threshold multiple; for example, k1 set to 1.5 means the local density of the core trajectory point must be at least 1.5 times the average density. k2 is a set distance threshold multiple; for example, k2 set to 2.0 means the distance between the core trajectory point and the point with the highest density must be at least twice the average distance. Core trajectory points meeting the decision criteria are typically located in the upper right corner of the decision map. These core trajectory points represent the most common and mainstream treatment patterns, are similar to the treatment paths of most patients, and exhibit significant differences between each other. Traditional clustering algorithms are susceptible to noise interference. This invention adopts a core trajectory point recognition mechanism based on dual thresholds of local density and distance, which effectively filters noise points and retains high-frequency and differentiated diagnosis and treatment pattern nodes, thereby refining the diagnosis and treatment trajectory point set and improving the accuracy of mainstream diagnosis and treatment pattern extraction.

[0097] Following the chronological order of the core trajectory points on the treatment trajectory coordinate system, these points are connected sequentially to form the main path of the golden treatment pathway template, representing the most common treatment sequence. For example, the main path for a hypertension patient is "initial treatment - blood pressure monitoring - medication." For each core trajectory point, treatment trajectory points with higher local density (greater than a preset density threshold, such as setting the density threshold to the average density) within its neighborhood are included in the corresponding path branches. For example, after the "medication" core point, branches such as "electrocardiogram" and "echocardiography" are added to correspond to examinations for different complications. In this way, a tree structure consisting of a main trunk and multiple branches forms the complete golden treatment pathway template. This template encompasses the actual treatment pathways of most patients, reflects the most common and reasonable treatment model, and can serve as a standard for evaluation and reference.

[0098] Step S40 transforms the patient population's treatment trajectory into an interpretable tree-like template, preserving the stability of mainstream patterns while accommodating reasonable differences in treatment. Breaking through the static limitations of traditional rule bases, it enables the automatic generation and updating of "data-driven" regulatory benchmarks, providing a reference standard that combines accuracy and flexibility for real-time anomaly monitoring. This not only solves the key problem of "how to define normal treatment pathways," but also, through the dynamic evolution of the template, allows medical insurance supervision to shift from "post-event verification" to "intelligent guidance during the process," significantly improving regulatory efficiency.

[0099] Step S50: Obtain the real-time diagnosis and treatment path of the target patient, match and align the real-time diagnosis and treatment path with the golden diagnosis and treatment path template, and determine whether the real-time diagnosis and treatment path of the target patient is an abnormal diagnosis and treatment path.

[0100] Please see Figure 4 As shown, the method for determining whether a target patient's real-time treatment path is an abnormal treatment path includes: representing the real-time treatment path as a real-time treatment trajectory vector V'. t Mapping these coordinates to the treatment trajectory coordinate system yields the real-time treatment trajectory point set T. t ; The real-time diagnosis and treatment trajectory point set T t Match and align with the main path and branch paths of the golden treatment pathway template, and calculate the minimum matching distance d. min Determine the minimum matching distance d min Does it exceed the preset deviation threshold d0? If so, mark the real-time diagnosis and treatment path of the target patient as an abnormal diagnosis and treatment path and enter it into the abnormal diagnosis and treatment path database; otherwise, it is regarded as a normal diagnosis and treatment path.

[0101] After obtaining the real-time treatment path of the target patient, which includes the time sequence records of treatment items, the patient's treatment item codes are mapped to standardized treatment item codes under the national standard coding system, generating the corresponding standardized treatment trajectory vector V'. t V't Mapping these coordinates onto the constructed diagnostic and treatment trajectory coordinate system yields the patient's real-time trajectory point set T. t T t It reflects the actual path of the target patient during the diagnosis and treatment process.

[0102] Using the Dynamic Time Warping (DTW) algorithm, the real-time diagnosis and treatment trajectory point set T of the target patient is... t The DTW algorithm matches the main path and each branch path of the golden treatment path template separately. By flexibly adjusting along the time axis, it allows for optimal alignment of time series of different lengths, solving the problem that traditional Euclidean distance cannot handle phase differences in time series and improving the matching accuracy of treatment trajectories of different lengths. The specific process is as follows: T... t Each treatment trajectory point in the model is dynamically matched with points in the golden treatment path template. The cumulative distance matrix is ​​calculated, and the path with the minimum cumulative distance is found by backtracking. The distance corresponding to this path is T. t The minimum matching distance d is calculated by iterating through all template paths in the golden treatment path template and taking the minimum value. min d min It reflects the degree of difference between the actual treatment pathway of the target patient and the most similar standard pathway.

[0103] Set a deviation threshold d0 as the standard for judging whether the treatment path is abnormal. The minimum matching distance d... min Compare with d0, if d min If the actual treatment path exceeds d0, it indicates a significant difference between the target patient's actual treatment path and the standard path, deviating from the conventional treatment model. In this case, it is marked as an abnormal treatment path and added to the abnormal treatment path database for subsequent analysis and intervention; if d... min If the distance does not exceed d0, the patient's treatment path is considered to be basically in line with the conventional pattern and is a normal treatment path, requiring no special attention or source analysis. The mean plus two times the standard deviation is taken as d0 by calculating the matching distance distribution of all normal trajectory points in the golden treatment path template to cover 95% of the normal fluctuation range. For example, if the average matching distance of normal trajectories in the golden treatment path template is 0.8 and the standard deviation is 0.2, then d0 = 0.8 + 2 × 0.2 = 1.2. When d... min When the deviation is greater than 1.2, it is determined to be an abnormal treatment path. This application establishes a real-time comparison mechanism based on three-dimensional spatial distance calculation, realizing real-time quantitative assessment of treatment path deviation, transforming abstract differences in treatment behavior into measurable numerical indicators, and significantly improving the efficiency and accuracy of anomaly identification.

[0104] Transforming abstract differences in medical practices into measurable numerical indicators improves verification efficiency. Combined with dynamic time warping technology, it can identify hidden anomalies caused by adjustments in the order of treatment, such as behaviors that evade supervision by delaying key examinations—scenarios that traditional rule bases cannot cover. Through comprehensive analysis across three dimensions, it can simultaneously pinpoint anomaly types; for example, deviations in the time dimension indicate abnormal treatment cycles, and deviations in the cost dimension point to issues with the reasonableness of charges, providing a clear direction for subsequent tracing and offering higher accuracy than single-dimensional analysis. By incorporating cost standardization coefficients into the matching dimension, it can discover violations that appear to conform to treatment guidelines but exhibit abnormal cost clusters. For example, multiple patients incurring significantly higher costs than the regional average for the same treatment items may reveal a new violation pattern where medical institutions decompose items and inflate costs to embezzle medical insurance funds. Furthermore, the construction of a treatment trajectory coordinate system provides a unified framework for cross-institutional treatment behavior analysis, enabling the identification of "cross-hospital violations" such as repeated examinations and overtreatment of the same patient in different hospitals—a blind spot in traditional single-institutional data supervision.

[0105] Step S60: If the abnormal diagnosis and treatment path is determined, a source tracing analysis is carried out on the abnormal diagnosis and treatment path.

[0106] Specifically, after determining that the real-time treatment path of the target patient is an abnormal treatment path, the deviation degree dev is calculated for each abnormal treatment path from three dimensions: time, treatment items, and cost. t dev p dev c .

[0107] Time dimension deviation dev t The calculation method is as follows: compare the actual time of each treatment item in the abnormal path with the time range of the corresponding item in the golden treatment path template, and calculate the proportion of the sum of the absolute values ​​of the time differences to the total treatment time. For example, in the golden treatment path template, "drug therapy" should be performed within 1-3 days after the first treatment. If a patient receives it on the 7th day, the time difference is 4 days. If the total treatment time is 10 days, then dev t =4 / 10=0.4.

[0108] Deviation of diagnostic and treatment items p The calculation method is to count the number of newly added, missing, or replaced treatment items in the abnormal treatment pathway compared to the golden treatment pathway template, and divide this number by the total number of treatment items in the template. For example, if the golden treatment pathway template includes three items: "monitoring-medication-follow-up examination," and the abnormal treatment pathway adds "CT examination" but lacks "follow-up examination," then dev... p =2 / 3≈0.67.

[0109] Cost dimension deviation dev cThe calculation method is as follows: Calculate the sum of the absolute values ​​of the differences between the standardized cost coefficients of each item in the abnormal treatment pathway and the mean of the golden treatment pathway template, and divide by the number of items. For example, the mean standardized cost coefficient in the golden treatment pathway template is 1.0, and the standardized cost coefficients of the three items in the abnormal treatment pathway are 1.5, 0.8, and 1.3, respectively. The sum of the differences is 0.5 + 0.2 + 0.3 = 1.0, dev c =1.0 / 3≈0.33.

[0110] will dev t dev p dev c Each dimension is assigned a preset weight, such as time (30%), project (40%), and cost (30%), and then weighted and summed to obtain the overall deviation score (Dev). A higher Dev score indicates a more severe anomaly. Subsequently, the abnormal diagnosis paths are ranked according to the Dev scores, and the violation type is inferred based on the deviation of each dimension: if the time dimension deviation score is higher... t At its highest, this could lead to problems such as prolonged hospital stays and excessive hospitalization; if the deviation of the treatment item dimension is dev... p At its highest, there may be issues such as over-examination, over-treatment, and duplicate charges; if the cost dimension deviates by a certain degree... c At its highest, there may be issues such as unreasonable charges or charges exceeding the standard.

[0111] In the source tracing analysis, the diagnostic and treatment item nodes corresponding to abnormal treatment paths are associated with the executing hospital, department, and physician information. For example, in the abnormal treatment path of a hypertensive patient, dev p The results were high, and the newly added "genetic testing" item showed that the usage frequency of this item among patients with the same diagnosis was only 5%. Furthermore, the number of times the physician ordered this item in the past three months was significantly higher than the department average. Combined with the abnormally low out-of-pocket ratio of this item in the medical insurance settlement process, it can be concluded that the physician is suspected of illegally ordering unnecessary high-value tests.

[0112] In existing technologies, medical insurance supervision often stops at anomaly identification, lacking systematic means to trace the source of violations, resulting in the inability to accurately locate the responsible parties and violation patterns. This application achieves precise tracing from abnormal treatment paths to specific violation clues through multi-dimensional quantitative analysis and graph association mining, enabling precise classification of violations and transforming vague "abnormalities" into verifiable issues such as "excessive examinations" and "double billing." Through the analysis of the main behavioral chain of the association graph, violations across treatment stages can be uncovered. For example, if multiple departments in a hospital abnormally reduce the weight of treatment items for the same diagnosed patient, it reveals a possible systematic avoidance of necessary examinations, a deep-seated problem that cannot be discovered by single-patient analysis. By using cross-analysis of cost and item dimensions, violations of "abnormal charging for compliant items" can be identified. For example, if a treatment item falls within the guideline recommendation range, but the cost standardization coefficient is consistently higher than 1.5 without reasonable clinical basis, it suggests possible overcharging.

[0113] Example 2

[0114] This embodiment, based on Embodiment 1, provides a data-driven business intelligent supervision and behavior tracing system, such as... Figure 5 As shown, it includes:

[0115] The graph construction module is used to acquire basic data on patients' medical services and construct a multi-dimensional relational graph that comprehensively displays the diagnosis and treatment process.

[0116] Disease diagnosis and treatment pathway pool generation module: used to extract patients with the same diagnosis from the multidimensional association graph and generate a disease diagnosis and treatment pathway pool;

[0117] Treatment trajectory generation module: used to construct a treatment trajectory coordinate system based on the disease treatment path pool, and obtain the treatment trajectory point set;

[0118] Path template generation module: Generates a golden treatment path template based on the set of treatment trajectory points;

[0119] Abnormal Path Judgment Module: Used to obtain the real-time diagnosis and treatment path of the target patient, match and align the real-time diagnosis and treatment path with the golden diagnosis and treatment path template, and determine whether the real-time diagnosis and treatment path of the target patient is an abnormal diagnosis and treatment path;

[0120] Source tracing analysis module: used to conduct source tracing analysis on abnormal diagnosis and treatment pathways.

[0121] In the graph construction module, the specific process of constructing a multidimensional association graph that comprehensively displays the diagnosis and treatment process includes: constructing a preliminary diagnosis association graph based on the patient's diagnosis code information in the patient's basic medical service data; and constructing a multidimensional association graph that comprehensively displays the diagnosis and treatment process based on the time sequence records of diagnosis and treatment items, medical insurance settlement information, and the constructed preliminary diagnosis association graph.

[0122] In the disease diagnosis and treatment pathway pool generation module, the specific process of extracting patients with the same diagnosis from the multidimensional association graph and generating the disease diagnosis and treatment pathway pool includes: determining the target diagnostic code; setting a diagnostic similarity radius r with the target diagnostic code as the center in the multidimensional association graph to construct a diagnostic nearest neighbor set; the diagnostic nearest neighbor set contains all diagnostic codes that are no more than r away from the target diagnostic code; collecting the diagnostic records corresponding to all diagnostic codes in the diagnostic nearest neighbor set to form an original diagnostic group; counting the number of diagnostic records corresponding to all diagnostic codes in the diagnostic nearest neighbor set and marking it as the total number of diagnostic records; based on the multidimensional association graph, screening the original diagnostic group for complications and comorbidities to obtain a pure diagnostic group; counting the number of diagnostic records for complications and comorbidities that were screened out; and obtaining the disease diagnosis and treatment pathway pool based on the pure diagnostic group.

[0123] The methods and systems of this application may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the method is for illustrative purposes only, and the steps of the method of this application are not limited to the order specifically described above, unless otherwise specifically stated.

[0124] In addition, the parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of the corresponding technical solutions in the prior art have not been described in detail, so as to avoid excessive elaboration.

[0125] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A data-driven business intelligent supervision and behavior tracing method, characterized in that, The method includes: Acquire basic patient medical service data, which includes at least patient diagnostic code information, treatment item time sequence records, and medical insurance settlement information; construct a preliminary diagnostic association graph based on patient diagnostic code information; extract diagnostic codes as nodes from patient diagnostic code information; in the constructed preliminary diagnostic association graph, with the diagnostic code as the core node, add treatment item nodes based on treatment item time sequence records and medical insurance settlement nodes based on medical insurance settlement information around the core node; the diagnostic code core node is connected to the treatment item nodes and medical insurance settlement nodes respectively, and the medical insurance settlement nodes are connected to the corresponding treatment item nodes, thus constructing a multi-dimensional association graph that comprehensively displays the treatment process; Extract patients with the same diagnosis from the multidimensional association graph, and extract the time sequence records of each patient's diagnosis and treatment items based on the patients with the same diagnosis to generate a disease diagnosis and treatment path pool; Based on the time sequence records of patients' treatment items in the disease treatment pathway pool, a standardized treatment trajectory vector V' is constructed; the first treatment time point is obtained from the time sequence records of treatment items, and a treatment trajectory coordinate system is constructed with the first treatment time point as the origin O. The standardized treatment trajectory vector V' is mapped to the treatment trajectory coordinate system to obtain the treatment trajectory point set. Calculate the local density of each point in the treatment trajectory point set and identify the point with the highest density; calculate the distance between each point in the treatment trajectory point set and the point with the highest density; identify the core trajectory points based on the local density and the distance; connect the core trajectory points in the treatment trajectory coordinate system to form the main path; include the treatment trajectory points around each core trajectory point whose local density is greater than a preset density threshold into the corresponding branch path to form a golden treatment path template. Obtain the real-time diagnosis and treatment path of the target patient, match and align the real-time diagnosis and treatment path with the golden diagnosis and treatment path template, and determine whether the real-time diagnosis and treatment path of the target patient is an abnormal diagnosis and treatment path; If an abnormal treatment pathway is identified, a source tracing analysis will be conducted on the abnormal treatment pathway.

2. The data-driven intelligent business supervision and behavior tracing method according to claim 1, characterized in that, The method for constructing a preliminary diagnostic association graph based on patient diagnostic code information includes: if two diagnostic codes appear simultaneously in more than Y patients and there is a sequential order, then a directed edge is added to connect the two diagnostic code nodes, where Y is a preset threshold parameter.

3. The data-driven intelligent business supervision and behavior tracing method according to claim 2, characterized in that, The method for extracting patients with the same diagnosis from a multidimensional association graph includes: Determine the target diagnostic code, and construct a diagnostic nearest neighbor set in the multidimensional association graph with the target diagnostic code as the center and set a diagnostic similarity radius r. The diagnostic nearest neighbor set contains all diagnostic codes whose distance from the target diagnostic code does not exceed r. Collect diagnostic records corresponding to all diagnostic codes in the diagnostic nearest neighbor set to form the original diagnostic group; based on the multidimensional association graph, screen the original diagnostic group for complications and comorbidities to obtain the pure diagnostic group; Calculate the diagnostic purity index of the pure diagnostic group; If the diagnostic purity index is greater than the preset purity threshold, the pure diagnostic group will be identified as the same diagnostic patient group.

4. The data-driven business intelligent supervision and behavior tracing method according to claim 3, characterized in that, The method for calculating the diagnostic purity index of the pure diagnostic group includes: The number of diagnostic records corresponding to all diagnostic codes in the nearest neighbor set of diagnoses is counted and marked as the total number of diagnostic records. The number of diagnostic records of complications and comorbidities that are screened out is counted. The diagnostic purity index is equal to the difference between the total number of diagnostic records and the number of diagnostic records of complications and comorbidities that are screened out, and then divided by the total number of diagnostic records.

5. The data-driven intelligent business supervision and behavior tracing method according to claim 4, characterized in that, The timeline of the diagnosis and treatment items shall include at least the diagnosis and treatment item code, the cost of the diagnosis and treatment item, and the specific implementation time; The method for constructing a standardized treatment trajectory vector V' based on the time sequence records of patients' treatment items in the disease treatment pathway pool includes: Using the initial treatment time as the baseline t0, calculate the time interval sequence Δt between the specific implementation time of each subsequent treatment item and t0; Map the diagnostic and treatment item codes to standardized diagnostic and treatment item codes to generate a standardized diagnostic and treatment item sequence P; Calculate the ratio of the cost of each treatment item to the average cost of the same item in the same region during the same period, and generate a series of standardized cost coefficients C; The time interval sequence Δt, the standardized treatment item sequence P, and the cost standardization coefficient sequence C are combined into a three-dimensional treatment trajectory vector V=[Δt,P,C]. The three-dimensional treatment trajectory vector V is then subjected to dimension normalization to obtain the standardized treatment trajectory vector V'.

6. A data-driven business intelligent supervision and behavior tracing system, used to implement the data-driven business intelligent supervision and behavior tracing method according to any one of claims 1-5, characterized in that, The system includes: The graph construction module is used to acquire basic data on patients' medical services and construct a multi-dimensional relational graph that comprehensively displays the diagnosis and treatment process. Disease diagnosis and treatment pathway pool generation module: used to extract patients with the same diagnosis from the multidimensional association graph and generate a disease diagnosis and treatment pathway pool; Treatment trajectory generation module: used to construct a treatment trajectory coordinate system based on the disease treatment path pool, and obtain the treatment trajectory point set; Path template generation module: Generates a golden treatment path template based on the set of treatment trajectory points; Abnormal Path Judgment Module: Used to obtain the real-time diagnosis and treatment path of the target patient, match and align the real-time diagnosis and treatment path with the golden diagnosis and treatment path template, and determine whether the real-time diagnosis and treatment path of the target patient is an abnormal diagnosis and treatment path; Source tracing analysis module: used to conduct source tracing analysis on abnormal diagnosis and treatment pathways.

Citation Information

Patent Citations

  • Intelligent supervision system and method based on medical insurance data and electronic equipment

    CN119048189A

  • Medical insurance common outpatient service supervision system and method based on fraud abnormal portrait technology

    CN119579332A

  • Medical insurance data checking method and system based on big data, electronic equipment and medium

    CN119722110A