Doctor drug dosage supernormal grading early warning method and system based on machine learning

By applying machine learning methods in drug dosage monitoring, combining feature reconstruction of autoencoder and attention mechanism, and building a drug dosage abnormal detection model, it solves the problems of inefficient and insufficient recognition capabilities of drug dosage monitoring in the existing technology, real-time and comprehensive monitoring and timely early warning are achieved.

CN120108758AActive Publication Date: 2025-06-06DALIAN UNIV OF TECH

Patent Information

Application Number
CN202510300342.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-06
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The prior art relies on manual review and post-spot checks in drug dosage monitoring, which is inefficient and has high subjectivity and lag. It is difficult to achieve real-time and comprehensive monitoring, and it is difficult to identify potential dosage abnormalities or prescription behaviors that deviate from the norm.

Method used

A machine learning-based method is adopted, combining feature reconstruction autoencoder and attention mechanism to construct an abnormal detection model for drug dosage. By constructing the sample set, sample features are obtained, local outlier factors are calculated using the LOF algorithm, and the features are reconstructed through local cross attention (LCA) optimization, and finally a drug dosage abnormality detection model is obtained.

Benefits of technology

Real-time and comprehensive monitoring of drug dosage is achieved, accurately identifying dosage abnormalities, providing timely warnings and dosage guidance, and reducing safety hazards in drug use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108758A_ABST
    Figure CN120108758A_ABST
Patent Text Reader

Abstract

The invention discloses a machine learning-based doctor drug dosage supernormal grading early warning method and system, and the system comprises a data layer which is mainly a data center, collects risk analysis data of drug consumable usage, departments, suppliers and the like from a hospital, carries out the data cleaning, and carries out the management of the data, including personnel management and query statistics; the business layer is mainly used for realizing key personnel monitoring and data intelligent research and judgment through an algorithm model, including clue mining, associated factor analysis and data visualization analysis; and the application layer comprises an intelligent early warning and forecasting function module, realizes red, yellow and green three-level risk early warning through big data analysis of the business layer according to threshold setting, and performs visual display. According to the method, a visual grading early warning mechanism condition is adopted, the drug dosage of a doctor is detected by using an anomaly detection algorithm model, the detection result is more accurate, the warning effect is more obvious, and meanwhile, the detection result provides a reference basis for the abnormal dosage condition for subsequent related departments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence medical technology, and specifically relates to a method and system for early warning of abnormal classification of doctor's drug usage based on machine learning. Background Art

[0002] In modern medical practice, rational use of drugs is an important part of ensuring patient safety and treatment effectiveness. When prescribing, doctors need to comprehensively judge the appropriate type and dosage of drugs based on the patient's specific condition, physiological characteristics and other factors. However, due to various reasons, including but not limited to human error, information asymmetry or lack of experience, sometimes the dosage of drugs exceeds the normal range, which may not only lead to poor drug efficacy, but also cause serious adverse reactions and even endanger the patient's life. At present, most hospitals and clinics mainly rely on the professional judgment and experience of doctors to avoid such problems. Although some medical institutions have introduced electronic prescription systems, which can provide certain drug interaction checking functions, they are limited in their ability to identify and warn of abnormal drug dosage. In addition, traditional electronic prescription systems often lack the ability to deeply analyze historical data and cannot effectively discover potential risk patterns or trends.

[0003] As the standardization of the medical system gradually improves, digital tools are being taken seriously in various places. Using digital means to monitor clinical behavior has become a trend in the management of medical institutions in the future, and the medical field has entered the "big data era". In the context of the era of big data, it is very necessary to deeply integrate data mining and analysis with hospital work. However, there is currently no method for detecting abnormal behavior of doctors in medication in the existing technology. At present, the monitoring of drug dosage mainly relies on manual review and post-inspection. This method is not only inefficient, but also has high subjectivity and lag, and it is difficult to achieve real-time and comprehensive monitoring. At the same time, since manual review is usually based on doctor experience or general dosage standards, it is difficult to identify potential dosage anomalies or prescription behaviors that deviate from the norm. In addition, the existing methods also have obvious deficiencies in timely feedback, making it difficult to provide doctors with timely dosage guidance, resulting in the failure to timely discover and deal with the safety hazards of some drug use. The present invention is based on an anomaly detection model that integrates feature reconstruction autoencoder (Autoencoder) and attention mechanism (Attention Mechanism). Based on a feature reconstruction model, the core architecture is an autoencoder. By combining the attention mechanism (LCA and MLKA), the model's ability to capture normal patterns is enhanced, and the hierarchical warning function is realized by coordinating the reconstruction error and threshold. The basic information of patients, doctors, and medications are historical data recorded in the hospital medical records. Summary of the invention

[0004] The purpose of the present invention is to solve the problem that the monitoring of drug dosage in the prior art mainly relies on manual review and post-inspection, which is not only inefficient, but also has high subjectivity and lag, making it difficult to achieve real-time and comprehensive monitoring. At the same time, since manual review is usually based on doctor's experience or general dosage standards, it is difficult to identify potential dosage anomalies or prescription behaviors that deviate from the norm. In addition, the existing methods also have obvious deficiencies in timely feedback, making it difficult to provide doctors with timely dosage guidance, resulting in the problem that some safety hazards in the use of drugs have not been discovered and dealt with in a timely manner.

[0005] In order to solve the above problems, the present invention provides a method for early warning of abnormal doctor drug usage based on machine learning, comprising: S1: construct sample set; The sample set includes: basic patient information, doctor information, medication information, and medical insurance data; S2: Obtain sample characteristics; The ratio of patient age to drug dosage is used as a new sample, and the X-means algorithm is used to iteratively cluster the new samples to obtain the sample set feature F org ; S3: Use the LOF algorithm to calculate the local outlier factor of each sample point; Using sample set feature F org The sample points are obtained to obtain the minimum reachable distance, and the local reachable density of each sample point is calculated according to the minimum reachable distance. The local outlier factor of each sample point is obtained by the LOF algorithm. S4: Reconstructed features through local cross attention LCA optimization; Obtain reference features through local cross attention LCA and pass weights , value vector Generate the final reconstruction features; S5: Use the features reconstructed in step S4 to train the sample set and obtain a drug dosage anomaly detection model.

[0006] In a preferred embodiment, the sample set in step S1 includes: basic patient information, doctor information, medication information, and medical insurance data; Patient basic information includes: recording month, patient age, patient gender, and length of hospital stay; doctor information includes: subspecialty, doctor name, and doctor title; medication information includes: disease diagnosis, drug dosage, drug name, drug category, drug code, drug manufacturer, drug unit price, and drug unit; medical insurance data includes: drg group number, drg group number, and drg group name; Step S2: The specific steps of obtaining sample features are as follows: The ratio of patient age to drug dosage is used as the new sample, and the formula is: in, represents a new sample, Indicates The patient's age, Indicates The amount of medicine required for each patient, =1, 2, ..., N; Use X-means algorithm to analyze new samples Perform iterative clustering, and take the initial number of clusters K as 2 or 3. For each initial cluster , using the binary K-means algorithm Split into two subclusters and , and calculate the center after the split, evaluate the results before and after the split, and calculate the BIC value, that is, the Bayesian information criterion, to determine whether the split is reasonable. The formula is: in, represents the log-likelihood value of the split model, The number of parameters representing the split model includes the number of centers and variance parameters, Indicates the number of samples in the cluster. If the BIC value after classification is higher than the original value, the split is accepted. If the splitting condition is met, the number of clusters K is updated: K=K+1, and the partitioning test is repeated until all clusters cannot be split any further; after clustering is completed, the centroid of each cluster is obtained and used together with the sample set obtained in step S1 as the sample set feature F org ; Step S3 uses the LOF algorithm to calculate the local outlier factor of each sample point. The specific steps are: For each sample set feature F org Sample points , define its k-distance, that is, the distance from the point to its kth nearest neighbor, the formula is: In the formula, yes The Nearest neighbor; The LOF algorithm introduces the reachable distance, which is defined as the point And its first Neighbors The distance between them is: Even if and The distance is less than The k-distance of , still uses the k-distance as the minimum reachable distance; Each sample point The local reachable density formula is: The LOF algorithm compares sample points The local density of the point is measured by the density of other points in its neighborhood Whether it is abnormal, the formula is: In the formula, if ,but The density of its neighborhood is similar and it is not an outlier. but If the density is lower than the neighborhood density, it is an outlier; according to the local outlier factor of the sample The distribution sets the threshold, and in practical applications, the threshold [0.9, 0.95] is selected; Step S4 optimizes the reconstructed features through local cross attention LCA. The specific steps are: LCA adds a local perception mask to restrict the query feature to match the reference feature in the neighborhood within a local range. The formula is as follows: In the formula, is the sample set feature F org The sample features in is a learnable normal reference representation, is the local mask matrix; By weight Sum value vector Generate the final reconstruction features, the formula is: In the formula, Represents the reconstruction features and weights of a certain layer of LCA module Represents the weighting of the reference, the value vector A mapping form that represents a learnable reference representation; At the same time, the output of LCA is combined with the output of the mask learning key attention module and the hyperparameter Weighting ensures that the model pays more attention to the LCA output. The formula is: In the formula, Represents the final output reconstruction feature, represents the reconstruction feature of the LCA module, represents the mask learning key attention matrix; Step S5 uses step S4 to reconstruct the features The training sample set is used to obtain the drug dosage anomaly detection model. The specific steps are as follows: The mean square error and cosine similarity loss function are used to measure the difference between the reconstructed features and the original input features. The loss function is defined as follows: In the formula, and Represents the height and width of the feature map; by minimizing the reconstruction loss, the model learns the distribution of normal usage features and ignores abnormal patterns. The reconstructed feature set is input into the model for training optimization iteration to obtain the optimal method.

[0007] A machine learning-based doctor's abnormal drug usage classification warning system, including: a data module, a business module, and an application module; M1: The working mode of the data module is reflected in the use of data modules to build sample sets, which include: basic patient information, doctor information, medication information, and medical insurance data; Patient basic information includes: recording month, patient age, patient gender, and length of hospital stay; doctor information includes: subspecialty, doctor name, and doctor title; medication information includes: disease diagnosis, drug dosage, drug name, drug category, drug code, drug manufacturer, drug unit price, and drug unit; medical insurance data includes: drg group number, drg group number, and drg group name; M2: The working mode of the application module is reflected in obtaining sample characteristics; The ratio of patient age to drug dosage is used as the new sample, and the formula is: in, represents a new sample, Indicates The patient's age, Indicates The amount of medicine required for each patient, =1, 2, ..., N; Use X-means algorithm to analyze new samples Perform iterative clustering, and take the initial number of clusters K as 2 or 3. For each initial cluster , using the binary K-means algorithm Split into two subclusters and , and calculate the center after the split, evaluate the results before and after the split, and calculate the BIC value, that is, the Bayesian information criterion, to determine whether the split is reasonable. The formula is: in, represents the log-likelihood value of the split model, The number of parameters representing the split model includes the number of centers and variance parameters, Indicates the number of samples in the cluster. If the BIC value after classification is higher than the original value, the split is accepted. If the splitting condition is met, the number of clusters K is updated: K=K+1, and the partitioning test is repeated until all clusters cannot be split any further; after clustering is completed, the centroid of each cluster is obtained and used together with the sample set obtained in step S1 as the sample set feature F org ; M3: Use LOF algorithm to calculate the local outlier factor of each sample point; For each sample set feature F org Sample points , define its k-distance, that is, the distance from the point to its kth nearest neighbor, the formula is: In the formula, yes The Nearest neighbor; The LOF algorithm introduces the reachable distance, which is defined as the point And its first Neighbors The distance between them is: Even if and The distance is less than The k-distance of , still uses the k-distance as the minimum reachable distance; Each sample point The local reachable density formula is: The LOF algorithm compares sample points The local density of the point is measured by the density of other points in its neighborhood Whether it is abnormal, the formula is: In the formula, if ,but The density of its neighborhood is similar and it is not an outlier. but If the density is lower than the neighborhood density, it is an outlier; according to the local outlier factor of the sample The distribution sets the threshold, and in practical applications, the threshold [0.9, 0.95] is selected; M4: Optimizing feature reconstruction through local cross attention LCA; LCA adds a local perception mask to restrict the query feature to match the reference feature in the neighborhood within a local range. The formula is as follows: In the formula, is the sample set feature F org The sample features in is a learnable normal reference representation, is the local mask matrix; By weight Sum value vector Generate the final reconstruction features, the formula is: In the formula, Represents the reconstruction features and weights of a certain layer of LCA module Represents the weighting of the reference, the value vector A mapping form that represents a learnable reference representation; At the same time, the output of LCA is combined with the output of the mask learning key attention module and the hyperparameter Weighting ensures that the model pays more attention to the LCA output. The formula is: In the formula, Represents the final output reconstruction feature, represents the reconstruction feature of the LCA module, represents the mask learning key attention matrix; M5: Reconstruct the feature using step M4 The sample set is trained to obtain a drug dosage anomaly detection model; The mean square error and cosine similarity loss function are used to measure the difference between the reconstructed features and the original input features. The loss function is defined as follows: In the formula, and Represents the height and width of the feature map; by minimizing the reconstruction loss, the model learns the distribution of normal usage features and ignores abnormal patterns. The reconstructed feature set is input into the model for training optimization iteration. Finally, the business module uses the trained model algorithm for anomaly detection and stores the detection results in the database.

[0008] The beneficial effects of the present invention are as follows: a visual graded early warning mechanism is adopted, and an abnormal detection algorithm model is used to detect the doctor's drug usage. The detection result is more accurate and the warning effect is more obvious. At the same time, the detection result provides a reference for abnormal usage for subsequent relevant departments. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1A system block diagram of the machine learning-based doctor drug dosage abnormal classification early warning system of the present invention; Figure 2 A schematic diagram of the modeling process of the abnormal drug dosage detection model for doctors of the present invention; Figure 3 The present invention is a schematic flow chart of the steps of using an abnormal amount of consumables. DETAILED DESCRIPTION

[0010] Example 1 A machine learning-based method for early warning of abnormal doctor drug usage, comprising the following steps: S1: construct sample set; The sample set includes: basic patient information, doctor information, medication information, and medical insurance data; S2: Obtain sample characteristics; The ratio of patient age to drug dosage is used as a new sample, and the X-means algorithm is used to iteratively cluster the new samples to obtain the sample set feature F org ; S3: Use the LOF algorithm to calculate the local outlier factor of each sample point; Using sample set feature F org The sample points are obtained to obtain the minimum reachable distance, and the local reachable density of each sample point is calculated according to the minimum reachable distance. The local outlier factor of each sample point is obtained by the LOF algorithm. S4: Reconstructed features through local cross attention LCA optimization; Obtain reference features through local cross attention LCA and pass weights , value vector Generate the final reconstruction features; S5: Use the features reconstructed in step S4 to train the sample set and obtain a drug dosage anomaly detection model.

[0011] The sample set in step S1 includes: basic patient information, doctor information, medication information, and medical insurance data; Patient basic information includes: recording month, patient age, patient gender, and length of hospital stay; doctor information includes: subspecialty, doctor name, and doctor title; medication information includes: disease diagnosis, drug dosage, drug name, drug category, drug code, drug manufacturer, drug unit price, and drug unit; medical insurance data includes: drg group number, drg group number, and drg group name; Step S2: The specific steps of obtaining sample features are as follows: The ratio of patient age to drug dosage is used as the new sample, and the formula is: in, represents a new sample, Indicates The patient's age, Indicates The amount of medicine required for each patient, =1, 2, ..., N; Use X-means algorithm to analyze new samples Perform iterative clustering, and take the initial number of clusters K as 2 or 3. For each initial cluster , using the binary K-means algorithm Split into two subclusters and , and calculate the center after the split, evaluate the results before and after the split, and calculate the BIC value, that is, the Bayesian information criterion, to determine whether the split is reasonable. The formula is: in, represents the log-likelihood value of the split model, The number of parameters representing the split model includes the number of centers and variance parameters, Indicates the number of samples in the cluster. If the BIC value after classification is higher than the original value, the split is accepted. If the splitting condition is met, the number of clusters K is updated: K=K+1, and the partitioning test is repeated until all clusters cannot be split any further; after clustering is completed, the centroid of each cluster is obtained and used together with the sample set obtained in step S1 as the sample set feature F org ; Step S3 uses the LOF algorithm to calculate the local outlier factor of each sample point. The specific steps are: For each sample set feature F org Sample points , define its k-distance, that is, the distance from the point to its kth nearest neighbor, the formula is: In the formula, yes The Nearest neighbor; The LOF algorithm introduces the reachable distance, which is defined as the point And its first Neighbors The distance between them is: Even if and The distance is less than The k-distance of , still uses the k-distance as the minimum reachable distance; Each sample point The local reachable density formula is: The LOF algorithm compares sample points The local density of the point is measured by the density of other points in its neighborhood Whether it is abnormal, the formula is: In the formula, if ,but The density of its neighborhood is similar and it is not an outlier. but If the density is lower than the neighborhood density, it is an outlier; according to the local outlier factor of the sample The distribution sets the threshold, and in practical applications, the threshold [0.9, 0.95] is selected; Step S4 optimizes the reconstructed features through local cross attention LCA. The specific steps are: LCA adds a local perception mask to restrict the query feature to match the reference feature in the neighborhood within a local range. The formula is as follows: In the formula, is the sample set feature F org The sample features in is a learnable normal reference representation, is the local mask matrix; By weight Sum value vector Generate the final reconstruction features, the formula is: In the formula, Represents the reconstruction features and weights of a certain layer of LCA module Represents the weighting of the reference, the value vector A mapping form that represents a learnable reference representation; At the same time, the output of LCA is combined with the output of the mask learning key attention module and the hyperparameter Weighting ensures that the model pays more attention to the LCA output. The formula is: In the formula, Represents the final output reconstruction feature, represents the reconstruction feature of the LCA module, represents the mask learning key attention matrix; Step S5 uses step S4 to reconstruct the features The training sample set is used to obtain the drug dosage anomaly detection model. The specific steps are as follows: The mean square error and cosine similarity loss function are used to measure the difference between the reconstructed features and the original input features. The loss function is defined as follows: In the formula, and Represents the height and width of the feature map; by minimizing the reconstruction loss, the model learns the distribution of normal usage features and ignores abnormal patterns. The reconstructed feature set is input into the model for training optimization iteration to obtain the optimal method.

[0012] A machine learning-based doctor's abnormal drug usage classification warning system, including: a data module, a business module, and an application module; M1: The working mode of the data module is reflected in the use of data modules to build sample sets, which include: basic patient information, doctor information, medication information, and medical insurance data; Patient basic information includes: recording month, patient age, patient gender, and length of hospital stay; doctor information includes: subspecialty, doctor name, and doctor title; medication information includes: disease diagnosis, drug dosage, drug name, drug category, drug code, drug manufacturer, drug unit price, and drug unit; medical insurance data includes: drg group number, drg group number, and drg group name; M2: The working mode of the application module is reflected in obtaining sample characteristics; The ratio of patient age to drug dosage is used as the new sample, and the formula is: in, represents a new sample, Indicates The patient's age, Indicates The amount of medicine required for each patient, =1, 2, ..., N; Use X-means algorithm to analyze new samples Perform iterative clustering, and take the initial number of clusters K as 2 or 3. For each initial cluster , using the binary K-means algorithm Split into two subclusters and , and calculate the center after the split, evaluate the results before and after the split, and calculate the BIC value, that is, the Bayesian information criterion, to determine whether the split is reasonable. The formula is: in, represents the log-likelihood value of the split model, The number of parameters representing the split model includes the number of centers and variance parameters, Indicates the number of samples in the cluster. If the BIC value after classification is higher than the original value, the split is accepted. If the splitting condition is met, the number of clusters K is updated: K=K+1, and the partitioning test is repeated until all clusters cannot be split any further; after clustering is completed, the centroid of each cluster is obtained and used together with the sample set obtained in step S1 as the sample set feature F org ; M3: Use LOF algorithm to calculate the local outlier factor of each sample point; For each sample set feature F org Sample points , define its k-distance, that is, the distance from the point to its kth nearest neighbor, the formula is: In the formula, yes The Nearest neighbor; The LOF algorithm introduces the reachable distance, which is defined as the point And its first Neighbors The distance between them is: Even if and The distance is less than The k-distance of , still uses the k-distance as the minimum reachable distance; Each sample point The local reachable density formula is: The LOF algorithm compares sample points The local density of the point is measured by the density of other points in its neighborhood Whether it is abnormal, the formula is: In the formula, if ,but The density of its neighborhood is similar and it is not an outlier. but If the density is lower than the neighborhood density, it is an outlier; according to the local outlier factor of the sample The distribution sets the threshold, and in practical applications, the threshold [0.9, 0.95] is selected; M4: Optimizing feature reconstruction through local cross attention LCA; LCA adds a local perception mask to restrict the query feature to match the reference feature in the neighborhood within a local range. The formula is as follows: In the formula, is the sample set feature F org The sample features in is a learnable normal reference representation, is the local mask matrix; By weight Sum value vector Generate the final reconstruction features, the formula is: In the formula, Represents the reconstruction features and weights of a certain layer of LCA module Represents the weighting of the reference, the value vector A mapping form that represents a learnable reference representation; At the same time, the output of LCA is combined with the output of the mask learning key attention module and the hyperparameter Weighting ensures that the model pays more attention to the LCA output. The formula is: In the formula, Represents the final output reconstruction feature, represents the reconstruction feature of the LCA module, Mask learning key attention matrix; M5: Reconstruct the feature using step M4 The sample set is trained to obtain a drug dosage anomaly detection model; The mean square error and cosine similarity loss function are used to measure the difference between the reconstructed features and the original input features. The loss function is defined as follows: In the formula, and Represents the height and width of the feature map; by minimizing the reconstruction loss, the model learns the distribution of normal usage features and ignores abnormal patterns. The reconstructed feature set is input into the model for training optimization iteration. Finally, the business module uses the trained model algorithm for anomaly detection and stores the detection results in the database.

[0013] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A machine learning-based doctor's drug dosage abnormal graded warning method, characterized in that: include: S1: construct sample set; The sample set includes: basic patient information, doctor information, medication information, and medical insurance data; S2: Obtain sample characteristics; The ratio of patient age to drug dosage is used as a new sample, and the X-means algorithm is used to iteratively cluster the new samples to obtain the sample set feature F org ; S3: Use the LOF algorithm to calculate the local outlier factor of each sample point; Using sample set feature F org The sample points are obtained to obtain the minimum reachable distance, and the local reachable density of each sample point is calculated according to the minimum reachable distance. The local outlier factor of each sample point is obtained by the LOF algorithm. S4: Reconstructed features through local cross attention LCA optimization; Obtain reference features through local cross attention LCA and pass weights , value vector Generate the final reconstruction features; S5: Use the features reconstructed in step S4 to train the sample set and obtain a drug dosage anomaly detection model.

2. The machine learning-based doctor's drug usage abnormal graded warning method according to claim 1 is characterized in that: Step S1: The sample set includes: basic patient information, doctor information, medication information, and medical insurance data; Patient basic information includes: recording month, patient age, patient gender, and length of hospital stay; doctor information includes: subspecialty, doctor name, and doctor title; medication information includes: disease diagnosis, drug dosage, drug name, drug category, drug code, drug manufacturer, drug unit price, and drug unit; medical insurance data includes: drg group number, drg group number, and drg group name; Step S2: The specific steps of obtaining sample features are as follows: The ratio of patient age to drug dosage is used as the new sample, and the formula is: in, represents a new sample, Indicates The patient's age, Indicates The amount of medicine required for each patient, =1, 2, ..., N; Use X-means algorithm to analyze new samples Perform iterative clustering, and take the initial number of clusters K as 2 or 3. For each initial cluster , using the binary K-means algorithm Split into two subclusters and , and calculate the center after the split, evaluate the results before and after the split, and calculate the BIC value, that is, the Bayesian information criterion, to determine whether the split is reasonable. The formula is: in, represents the log-likelihood value of the split model, The number of parameters representing the split model includes the number of centers and variance parameters, Indicates the number of samples in the cluster. If the BIC value after classification is higher than the original value, the split is accepted. If the splitting condition is met, the number of clusters K is updated: K=K+1, and the partitioning test is repeated until all clusters cannot be split any further; after clustering is completed, the centroid of each cluster is obtained and used together with the sample set obtained in step S1 as the sample set feature F org ; Step S3 uses the LOF algorithm to calculate the local outlier factor of each sample point. The specific steps are: For each sample set feature F org Sample points , define its k-distance, that is, the distance from the point to its kth nearest neighbor, the formula is: In the formula, yes The first Nearest neighbor; The LOF algorithm introduces the reachable distance, which is defined as the point And its first Neighbors The distance between them is: Even if and The distance is less than The k-distance of , still uses the k-distance as the minimum reachable distance; Each sample point The local reachable density formula is: The LOF algorithm compares sample points The local density of the point is measured by the density of other points in its neighborhood Whether it is abnormal, the formula is: In the formula, if ,but The density of its neighborhood is similar and it is not an outlier. but If the density is lower than the neighborhood density, it is an outlier; according to the local outlier factor of the sample The distribution sets the threshold, and in practical applications, the threshold [0.9, 0.95] is selected; Step S4 optimizes the reconstructed features through local cross attention LCA. The specific steps are: LCA adds a local perception mask to restrict the query feature to match the reference feature in the neighborhood within a local range. The formula is as follows: In the formula, is the sample set feature F org The sample features in is a learnable normal reference representation, is the local mask matrix; By weight Sum value vector Generate the final reconstruction features, the formula is: In the formula, Represents the reconstruction features and weights of a certain layer of LCA module Represents the weighting of the reference, the value vector A mapping form that represents a learnable reference representation; At the same time, the output of LCA is combined with the output of the mask learning key attention module and the hyperparameter Weighting ensures that the model pays more attention to the LCA output. The formula is: In the formula, Represents the final output reconstruction feature, represents the reconstruction features of the LCA module, represents the mask learning key attention matrix; Step S5 uses step S4 to reconstruct the features The training sample set is used to obtain the drug dosage anomaly detection model. The specific steps are as follows: The mean square error and cosine similarity loss function are used to measure the difference between the reconstructed features and the original input features. The loss function is defined as follows: In the formula, and Represents the height and width of the feature map; by minimizing the reconstruction loss, the model learns the distribution of normal usage features and ignores abnormal patterns. The reconstructed feature set is input into the model for training optimization iteration to obtain the optimal method.

3. A machine learning-based doctor's abnormal drug usage classification warning system, characterized in that: include: Data module, business module, application module; M1: The working mode of the data module is reflected in the use of data modules to build sample sets, which include: basic patient information, doctor information, medication information, and medical insurance data; Patient basic information includes: recording month, patient age, patient gender, and length of hospital stay; doctor information includes: subspecialty, doctor name, and doctor title; medication information includes: disease diagnosis, drug dosage, drug name, drug category, drug code, drug manufacturer, drug unit price, and drug unit; medical insurance data includes: drg group number, drg group number, and drg group name; M2: The working mode of the application module is reflected in obtaining sample characteristics; The ratio of patient age to drug dosage is used as the new sample, and the formula is: in, represents a new sample, Indicates The patient's age, Indicates The amount of medicine required for each patient, =1, 2, ..., N; Use X-means algorithm to analyze new samples Perform iterative clustering, and take the initial number of clusters K as 2 or 3. For each initial cluster , using the binary K-means algorithm Split into two subclusters and , and calculate the center after the split, evaluate the results before and after the split, and calculate the BIC value, that is, the Bayesian information criterion, to determine whether the split is reasonable. The formula is: in, represents the log-likelihood value of the split model, The number of parameters representing the split model includes the number of centers and variance parameters, Indicates the number of samples in the cluster. If the BIC value after classification is higher than the original value, the split is accepted. If the splitting condition is met, the number of clusters K is updated: K=K+1, and the partitioning test is repeated until all clusters cannot be split any further; after clustering is completed, the centroid of each cluster is obtained and used together with the sample set obtained in step S1 as the sample set feature F org ; M3: Use LOF algorithm to calculate the local outlier factor of each sample point; For each sample set feature F org Sample points , define its k-distance, that is, the distance from the point to its kth nearest neighbor, the formula is: In the formula, yes The first Nearest neighbor; The LOF algorithm introduces the reachable distance, which is defined as the point And its first Neighbors The distance between them is: Even if and The distance is less than The k-distance of , still uses the k-distance as the minimum reachable distance; Each sample point The local reachable density formula is: The LOF algorithm compares sample points The local density of the point is measured by the density of other points in its neighborhood Whether it is abnormal, the formula is: In the formula, if ,but The density of its neighborhood is similar and it is not an outlier. but If the density is lower than the neighborhood density, it is an outlier; according to the local outlier factor of the sample The distribution sets the threshold, and in practical applications, the threshold [0.9, 0.95] is selected; M4: Optimizing feature reconstruction through local cross attention LCA; LCA adds a local perception mask to restrict the query feature to match the reference feature in the neighborhood within a local range. The formula is as follows: In the formula, is the sample set feature F org The sample features in is a learnable normal reference representation, is the local mask matrix; By weight Sum value vector Generate the final reconstruction features, the formula is: In the formula, Represents the reconstruction features and weights of a certain layer of LCA module Represents the weighting of the reference, the value vector A mapping form that represents a learnable reference representation; At the same time, the output of LCA is combined with the output of the mask learning key attention module and the hyperparameter Weighting ensures that the model pays more attention to the LCA output. The formula is: In the formula, Represents the final output reconstruction feature, represents the reconstruction features of the LCA module, represents the mask learning key attention matrix; M5: Reconstruct the feature using step M4 The sample set is trained to obtain a drug dosage anomaly detection model; The mean square error and cosine similarity loss function are used to measure the difference between the reconstructed features and the original input features. The loss function is defined as follows: In the formula, and Represents the height and width of the feature map; by minimizing the reconstruction loss, the model learns the distribution of normal usage features and ignores abnormal patterns. The reconstructed feature set is input into the model for training optimization iteration. Finally, the business module uses the trained model algorithm for anomaly detection and stores the detection results in the database.

Citation Information

Patent Citations

  • Medical abnormity violation big data risk early warning method based on unsupervised machine learning and integrated learning

    CN117764741A

  • Medicine project reasonable use method and system based on diagnosis and medicine project reasoning

    CN118507077A

  • Smart multidosing

    US20200245925A1

Cited By

  • Doctor drug dosage supernormal grading early warning method and system based on machine learning

    CN121306602A