Medical treatment risk detection method, device and equipment and computer storage medium

By obtaining user medical communication behavior data, using pre-trained target expectation value tables and clustering analysis, combined with pre-trained classifiers, the user behavior patterns are automatically identified, which solves the problem of low accuracy of medical risk detection in the existing technology, and achieves efficient and dynamic risk identification.

CN120299641APending Publication Date: 2025-07-11CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510311844.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing medical risk detection model has low detection accuracy, making it difficult to adapt to changing medical treatment behaviors, and relies on a large number of labeled sample data and fixed data characteristics, resulting in insufficient recognition rate.

Method used

By obtaining user medical communication behavior data, using pre-trained target expectation value tables and clustering analysis, combined with pre-trained classifiers, it automatically recognizes user behavior patterns, dynamically adjusts and optimizes tags, and improves risk identification accuracy.

Benefits of technology

It realizes efficient identification of medical treatment risks under a small number of marked sample data, improves detection accuracy, adapts to changes in different medical treatment behaviors, and reduces misjudgment and misjudgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299641A_ABST
    Figure CN120299641A_ABST
Patent Text Reader

Abstract

The invention discloses a medical treatment risk detection method, device and equipment and a computer storage medium, and relates to the technical field of data processing. The method comprises the following steps: acquiring medical-seeing communication behavior data of a user; inputting the medical-seeing communication behavior data of the user into a medical-seeing risk detection model, and determining a medical-seeing risk classification result corresponding to the medical-seeing communication behavior data of the user according to a pre-trained target expected value table; the target expected value table comprises a corresponding relation between the doctor-seeing communication behavior data and the doctor-seeing risk classification result, the target expected value table is obtained by training the doctor-seeing communication behavior data and a target user tag, and the target user tag is determined according to a relative distance among a user clustering tag, a user prediction tag and a data point; the user clustering label is obtained by clustering according to the local density and the minimum distance of the user data points, and the user prediction label is obtained through a pre-trained classifier. According to the embodiment of the invention, the optimal decision is made according to the target expected value table, and the risk identification accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data processing, and in particular, relates to a method, device, equipment, and computer storage medium for detecting medical treatment risks. Background Art

[0002] Currently, there are more and more ways for users to defraud insurance when seeking medical treatment. Users may use others' medical insurance cards for medical reimbursement, or lend their medical insurance cards to others to defraud medical insurance funds; or patients do not actually receive inpatient treatment after going through the inpatient formalities and declare expenses to the medical insurance department.

[0003] When identifying medical treatment risks in the prior art, preset rules or statistical methods are often used for judgment, and abnormal behaviors are identified through statistical analysis results. However, due to the small number of medical treatment risk samples and the rapid evolution of insurance fraud means, the recognition rate of fraud behaviors by the judgment methods obtained from medical treatment risk samples is low, and it is difficult to adapt to the changing medical treatment behaviors, resulting in a low accuracy rate of existing medical treatment risk detection. Summary of the Invention

[0004] Embodiments of this application provide a method, device, equipment, and computer storage medium for detecting medical treatment risks to solve the problem of low detection accuracy rate of existing medical treatment risk detection models.

[0005] In a first aspect, embodiments of this application provide a method for detecting medical treatment risks, and the method includes:

[0006] Obtain user medical treatment communication behavior data;

[0007] Input the user medical treatment communication behavior data into a medical treatment risk detection model, and determine the medical treatment risk classification result corresponding to the user medical treatment communication behavior data according to a pre-trained target expected value table; the target expected value table includes the corresponding relationship between medical treatment communication behavior data and medical treatment risk classification results, the target expected value table is obtained by training medical treatment communication behavior data and target user labels, the target user labels are determined according to user clustering labels, user prediction labels, and the relative distance of user data points, the user clustering labels are obtained by clustering according to the local density and minimum distance of user data points, and the user prediction labels are obtained through a pre-trained classifier.

[0008] In a second aspect, embodiments of this application provide a device for detecting medical treatment risks, and the device includes:

[0009] An obtaining module, configured to obtain user medical treatment communication behavior data;

[0010] A determination module is configured to input user medical treatment communication behavior data into a medical treatment risk detection model, and determine a medical treatment risk classification result corresponding to the user medical treatment communication behavior data according to a pre-trained target expected value table. The target expected value table includes the corresponding relationship between the medical treatment communication behavior data and the medical treatment risk classification result. The target expected value table is obtained by training the medical treatment communication behavior data and the target user labels. The target user labels are determined according to the user clustering labels, the user prediction labels, and the relative distance of the user data points. The user clustering labels are obtained by clustering according to the local density and the minimum distance of the user data points. The user prediction labels are obtained by a pre-trained classifier.

[0011] In a third aspect, an embodiment of the present application provides a terminal device, which includes a processor and a memory storing computer program instructions. When the processor executes the computer program instructions, the method for detecting medical treatment risks as described in the first aspect is implemented.

[0012] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method for detecting medical treatment risks as described in the first aspect is implemented.

[0013] In a fifth aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute the method for detecting medical treatment risks as described in the first aspect.

[0014] An embodiment of the present application provides a method, apparatus, device, and computer storage medium for detecting medical treatment risks. The method first obtains user medical treatment communication behavior data; inputs the user medical treatment communication behavior data into a medical treatment risk detection model, and determines a medical treatment risk classification result corresponding to the user medical treatment communication behavior data according to a pre-trained target expected value table. The target expected value table includes the corresponding relationship between the medical treatment communication behavior data and the medical treatment risk classification result. The target expected value table is obtained by training the medical treatment communication behavior data and the target user labels. The target user labels are determined according to the user clustering labels, the user prediction labels, and the relative distance of the user data points. The user clustering labels are obtained by clustering according to the local density and the minimum distance of the user data points. The user prediction labels are obtained by a pre-trained classifier. By clustering, the behavior patterns of users can be automatically identified, users with similar behaviors can be grouped, and classification is performed based on the distribution characteristics of the data itself without relying on a large number of labeled sample data. Determining the target user labels and adjusting and optimizing the initial clustering results can obtain the potential connections between data points. Training the medical treatment risk detection model according to the training medical treatment communication behavior data and the target user labels to obtain the target expected value table can make the optimal decision based on the current state and historical experience, and improve the accuracy of risk identification. Description of the Drawings

[0015] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0016] Figure 1 is a schematic flowchart of a method for detecting medical treatment risks provided by an embodiment of the present application;

[0017] Figure 2 is a schematic flowchart of a method for training a risk detection model provided by an embodiment of the present application;

[0018] Figure 3 is a schematic flowchart of a method for determining a user's abnormal score provided by an embodiment of the present application;

[0019] Figure 4 is a schematic flowchart of a method for determining a user's clustering label provided by an embodiment of the present application;

[0020] Figure 5 is a schematic flowchart of a method for determining a target user label provided by an embodiment of the present application;

[0021] Figure 6 is a schematic structural diagram of a device for detecting medical treatment risks provided by an embodiment of the present application;

[0022] Figure 7 is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0023] The following will describe in detail the features and exemplary embodiments of various aspects of the present application. To make the objectives, technical solutions and advantages of the present application clearer, the following further describes the present application in detail with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.

[0024] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0025] The technical solutions of the prior art usually use rules or machine learning methods to identify insurance fraud behaviors. Machine learning technology characterizes insurance fraud behaviors through a series of operation steps such as data preprocessing, feature engineering, model construction, model training, and model evaluation. However, the samples of insurance fraud behaviors are often a minority compared to normal samples, and there is a problem of class imbalance in the data. Moreover, the prior art solutions have strict requirements for data selection, such as comprehensive, accurate, and high-quality transaction records and details, etc. The cost of manual data annotation is high, and the utilization of unannotated data is insufficient. The commonly used ensemble learning strategy in the solutions constructs a fusion model, which highly depends on the characteristics of the data itself and is prone to insufficient recognition of medical fraud samples. And with the development of technology, fraudsters continuously adopt new technologies and methods, such as using cutting-edge technologies like big data and artificial intelligence, and combining complex means such as making money through part-time jobs for fraud. The commonly used ensemble learning strategy in the prior art solutions constructs a fusion model, which highly depends on the structure of the model itself and fixed data characteristics, and requires a long time to retrain and deploy. It cannot dynamically optimize the medical fraud recognition strategy to adapt to the changing medical behaviors, resulting in a low detection accuracy of the existing medical risk detection models.

[0026] To solve the problems of the prior art, embodiments of the present application provide a method, apparatus, device, and computer storage medium for detecting medical treatment risks. The method first obtains user medical treatment communication behavior data; inputs the user medical treatment communication behavior data into a medical treatment risk detection model, and determines a medical treatment risk classification result corresponding to the user medical treatment communication behavior data according to a pre-trained target expectation value table; the target expectation value table includes the corresponding relationship between the medical treatment communication behavior data and the medical treatment risk classification result, and the target expectation value table is obtained by training the medical treatment communication behavior data and the target user labels. The target user labels are determined according to the user clustering labels, user prediction labels, and relative distances of user data points. The user clustering labels are obtained by clustering according to the local density and minimum distance of the user data points, and the user prediction labels are obtained through a pre-trained classifier. By clustering, the behavior patterns of users can be automatically identified, users with similar behaviors can be grouped, and classification can be performed based on the distribution characteristics of the data itself without relying on a large number of labeled sample data. Determining the target user labels and adjusting and optimizing the initial clustering results can obtain the potential connections between data points. Training the medical treatment risk detection model according to the training medical treatment communication behavior data and the target user labels to obtain the target expectation value table can make optimal decisions based on the current state and historical experience, and improve the accuracy of risk identification.

[0027] The following introduces the method for detecting medical treatment risks provided by the embodiments of the present application with reference to the accompanying drawings.

[0028] Figure 1 The flowchart of the method for detecting medical treatment risks provided by an embodiment of the present application is shown. As Figure 1 shown, the method may include the following steps: S101 and S102.

[0029] S101, obtain user medical treatment communication behavior data.

[0030] Among them, the user medical treatment communication behavior data is behavior data generated by the user during the medical treatment process through communication means, such as telephone, text message, online platform, etc.

[0031] In some embodiments, the user medical treatment communication behavior data may include user basic information, user behavior data of browsing medical software, medical insurance data, medical treatment time, medical treatment hospital, medical industry communication activities, user search keywords, and other data.

[0032] The user basic information may include demographic information such as user province and city codes, age, and gender, which are used as the primary key to associate subsequent other behavior data, and are used to analyze the medical treatment behavior characteristics of different user groups.

[0033] The user behavior data of browsing medical software refers to the medical browsing behaviors of users on different platforms. When obtaining medical insurance data, medical treatment time, and the hospitals visited, extract the longitude and latitude coordinates of the base stations and hospitals where the users are resident, calculate whether the shortest distance from the base stations passed by the users to the hospitals is within a certain range, and record the duration of the users' stay at the hospital base stations within the specified time period, so as to judge the frequency of the users approaching medical institutions.

[0034] Medical industry communication activities refer to the call duration and number of calls made / answered by users to / from the medical industry / key insurance companies. Identify medical industry communication activities through phone numbers and text messages and calculate relevant communication metrics.

[0035] User search keywords are obtained from the user search engine log files and search application data on the device. Record the user search keywords, their types, frequencies, and the number of days, as well as the upstream, downstream, and total traffic of the search behavior. Analyze the users' medical needs through the user keywords to judge whether the users have actually undergone medical treatment.

[0036] Analyze the users' behavior patterns during the medical treatment process by obtaining the users' medical treatment communication behavior data, so as to provide basic data support for subsequent medical treatment risk assessment.

[0037] S102, input the users' medical treatment communication behavior data into the medical treatment risk detection model, and determine the medical treatment risk classification result corresponding to the users' medical treatment communication behavior data according to the pre-trained target expected value table; the target expected value table includes the corresponding relationship between the medical treatment communication behavior data and the medical treatment risk classification result, and the target expected value table is obtained by training the medical treatment communication behavior data and the target user labels. The target user labels are determined according to the user clustering labels, user prediction labels, and the relative distances of the user data points. The user clustering labels are obtained by clustering the local densities and minimum distances of the user data points, and the user prediction labels are obtained through a pre-trained classifier.

[0038] Among them, the target expected value table is the expected value table for the medical treatment risk classification corresponding to different users' medical treatment communication behavior data, and the medical treatment risk classification result is the risk result of users' medical insurance fraud.

[0039] In some embodiments, the medical treatment risk classification of users may include low risk, medium risk, and high risk, or may also be classified, such as no risk, the risk of using others' medical insurance cards, the risk of staying in the hospital without actual treatment, etc.

[0040] In some embodiments, determining the medical treatment risk classification result corresponding to the users' medical treatment communication behavior data according to the pre-trained target expected value table may include:

[0041] Determine the expected values of different medical treatment risk classifications corresponding to different users' medical treatment communication behavior data according to the target expected value table;

[0042] Determine the medical treatment risk classification result according to the expected value of all medical treatment risk classifications.

[0043] In some embodiments, when determining the medical treatment risk classification result according to the expected value of all medical treatment risk classifications, the multiple expected values of each type of medical treatment risk classification can be fused into one expected value. Specifically, it can be done by taking the average value, weighted summation, or selecting the maximum value, so that each type of medical treatment risk classification corresponds to a fused expected value, and the medical treatment risk classification corresponding to the maximum fused expected value is selected as the medical treatment risk classification result.

[0044] In some embodiments, it is also possible to determine the medical treatment risk classification of the maximum expected value corresponding to the medical treatment communication behavior data of different types of users, calculate the number of occurrences of each medical treatment risk classification, and select the medical treatment risk classification with the largest number as the medical treatment risk classification result.

[0045] Through clustering, the behavior patterns of users can be automatically identified, users with similar behaviors can be grouped, and classification can be carried out based on the distribution characteristics of the data itself without relying on a large number of labeled sample data. Determine the target user labels, adjust and optimize the initial clustering results, and potential connections between user data points can be obtained. Train the medical treatment risk detection model according to the training medical treatment communication behavior data and target user labels to obtain the target expected value table, which can make the optimal decision based on the current state and historical experience and improve the accuracy of risk identification.

[0046] In some embodiments, before inputting the user's medical treatment communication behavior data into the medical treatment risk detection model, desensitize all identifiable information to ensure that it cannot be traced back to an individual; adopt a secure data storage solution, clearly define the purpose and scope of data use, and strictly control data access rights to ensure that the data is only used for medical treatment risk detection and is fully protected during transmission and storage to prevent unauthorized access; regularly conduct compliance reviews on the stored data.

[0047] In some embodiments, before inputting the user's medical treatment communication behavior data into the medical treatment risk detection model, it is also possible to clean and preprocess the user's medical treatment communication behavior data to remove irrelevant information such as noise, stop words, and punctuation marks. Combine the user behavior data, the location base station data near the medical institution, the communication duration and frequency with the medical institution, the user's search keywords, and the medical Internet behavior data, etc., and perform corresponding feature derivation. Construct features such as "access frequency to different medical institutions" based on the location data and call records to form a complete user feature set. And perform normalization processing on the features to ensure that features with different dimensions can be compared and analyzed on the same scale. Finally, input the normalized user feature set into the medical treatment risk detection model.

[0048] In some embodiments, such as Figure 2 shown, before inputting the user's medical treatment communication behavior data into the medical treatment risk detection model and determining the medical treatment risk classification result corresponding to the user's medical treatment communication behavior data according to the pre-trained target expectation value table, the method may further include: S201 to S205.

[0049] S201, Obtain the training medical treatment communication behavior data.

[0050] Among them, the training medical treatment communication behavior data is the medical treatment communication behavior data used to train the model, and the training medical treatment communication behavior data includes a small amount of data with marked user medical treatment risk classification results and a large amount of data without marked user medical treatment risk classification results.

[0051] S202, Calculate the local density and minimum distance of the user data points, determine the cluster center according to the local density and minimum distance and perform clustering to obtain the user cluster label, and the user data points include the training medical treatment communication behavior data.

[0052] Among them, the user data point is a multi-dimensional feature vector composed of the training medical treatment communication behavior data, the local density is the number of neighbors around the user data point, which is used to measure the density of the user data point, the minimum distance is the distance from the user data point to its nearest neighbor, which is used to measure the isolation degree of the user data point, and the cluster center is the core point of each cluster in the clustering algorithm.

[0053] In some embodiments, the training medical treatment communication behavior data set can be D = {x1, x2,..., xn}, where xi represents the i-th user data point. Each user data point xi is multi-dimensional data, that is, the dimensions included in the user feature set. For example, xi = {xi1, xi2,..., xim+1}, which are {medical treatment frequency, medical information contact degree,..., potential insurance fraud abnormal behavior score} of the i-th user data point respectively.

[0054] S203, Input the training medical treatment communication behavior data into the pre-trained classifier, and determine the user prediction label according to the training medical treatment communication behavior data and the preset weight.

[0055] Among them, the pre-trained classifier is a classification model that has been trained with training data, and can be a decision tree classification model, a support vector machine model or a neural network model.

[0056] S204, Calculate the relative distance of the user data points, and determine the target user label according to the user cluster label, the user prediction label and the relative distance of the user data points.

[0057] Among them, the relative distance is the ratio of the distance from the user data point to the cluster center to the average distance within the cluster, which is used to measure the representativeness of the user data point in the clustering.

[0058] Comprehensively considering the clustering results and the prediction results of the classifier, dynamically adjust the labels according to the relative distances of user data points, determine the final target user labels for each user data point, and improve the accuracy of the labels.

[0059] S205, train the medical treatment risk detection model based on the training medical treatment communication behavior data and the target user labels to obtain a target expected value table, and update the risk detection model according to the target expected value table.

[0060] Use the training data and the target user labels to train the model to update the expected value table, so that the target expected value table can better reflect the relationship between user behavior and risk, thereby improving the model's prediction ability for new data.

[0061] Combining multiple information such as clustering analysis, the prediction results of the pre-trained classifier, and the relative distances of user data points, clustering analysis can discover the natural groupings in the data, the pre-trained classifier provides preliminary predictions, and the relative distances are used to adjust the labels. Analyzing user behavior from different perspectives can better capture the complexity of user behavior and improve the accuracy and reliability of risk classification.

[0062] In some embodiments, the training medical treatment communication behavior data includes user search keywords. Before inputting the training medical treatment communication behavior data into the pre-trained classifier, the method may further include:

[0063] Determine the user anomaly score according to the proportion of preset keywords in the search keywords;

[0064] Input the user anomaly score into the pre-trained classifier.

[0065] Wherein, the search keywords are high-risk words for medical treatment risks, other words related to medical insurance reimbursement, and their synonyms and near-synonyms. The high-risk words are keywords frequently associated with insurance fraud behaviors. Other words refer to medical consultations related to medical insurance reimbursement, disease common sense, etc., such as related disease descriptions, physical condition descriptions, etc. The preset keywords are high-risk words and their synonyms and near-synonyms.

[0066] According to the proportion of preset keywords in the user search keywords, users with abnormal behaviors can be analyzed in advance, and the user anomaly score is input into the classifier, providing more information for the classifier and reducing the amount of data that needs to be further analyzed, improving the accuracy and efficiency of the classification process.

[0067] In some embodiments, named entity recognition is performed on the keywords searched by the user based on the Han Language Processing (HANLP) library to identify the entities in the search record, which are then matched with the keyword database, so as to obtain information such as the search times, days, upstream and downstream traffic, types, and timestamps of high-risk words and other words searched by each user within a certain period, thereby obtaining the proportion of preset keywords in the search keywords.

[0068] In some embodiments, the semantic analysis ability of the Natural Language Processing (NLP) algorithm can also be utilized to further analyze the context in which the keywords appear. For example, if a user frequently searches a large number of abnormal keywords related to medical insurance reimbursement without searching other normal medical-related words, it is determined that the user's behavior is abnormal. According to the context, it is judged whether there are other searches for insurance fraud behaviors, such as asking how to exaggerate the condition, fabricate treatment records, and how to seek medical treatment with another person's medical insurance card, etc., so as to improve the accuracy of identification.

[0069] In some embodiments, the medical risk detection model can be a Q-Learning algorithm model, and the target expected value table can be a Q function table.

[0070] In the scenario of medical fraud identification, Q-learning makes decisions by learning an action value function to evaluate the value of examining specific input information. For example, when the system needs to examine the user's geographical location information to identify potential insurance fraud behaviors, Q-learning will calculate the expected cumulative reward for examining the geographical location and select the action with the highest Q value to execute, helping the system make decisions.

[0071]

[0072] Among them, Q(s,a;θ) is the value of the Q function, representing the expected return of taking action a in state s. s represents the constructed user feature set, such as medical insurance usage records and medical communication records, etc. a represents examining specific user input information, such as querying medical records and analyzing geographical locations, etc. θ is the neural network parameter, E S′~P is the expectation for all possible new states s′. s′ represents the new state after taking the action. P is the distribution of the new state. r represents the positive feedback for successfully identifying insurance fraud behaviors or the negative feedback for misjudgment. γ is the discount factor, which helps the system balance the relationship between current and future rewards. a′ is the action that can be selected in the new state s′, and θ' is the target network parameter, which is used to stabilize the learning process.

[0073] By continuously interacting with the environment, the Q value is updated according to the obtained positive reward or negative feedback, and gradually the optimal strategy is selected.

[0074] The improved Q target network no longer requires the Q network to pass parameters to it, which helps reduce the noise in the parameter passing process. The state s of the environment only affects the Q network, and the improved Q target network only accepts good samples as the training set, which can more effectively identify potential fraudsters. The two networks are trained separately. The loss functions for training the two networks are set as follows:

[0075]

[0076] Where Y is the label value, m is the set threshold, samples greater than m are considered bad samples, Q' is the value of the improved Q network, M is the total number of samples, L is the overall loss function, and L1 is the loss function for a single sample, which is used to measure the difference between the improved Q network and the original Q network.

[0077] At the same time, borrowing the idea of unsupervised learning, it is judged whether the samples in Q belong to the Q' samples, that is, it is judged whether the samples detected by Q are good samples. At the same time, experience replay is used to stabilize the training process, storing cases of successfully identifying fraud behaviors and misjudgments, storing the interaction experience of the model in the environment, and randomly sampling these experiences for learning during the training process, which is beneficial to breaking the temporal correlation between data and improving the generalization ability of the model.

[0078] The neural network automatically learns the feature representation of the data, and regards the medical fraud identification process as a sequential decision-making problem, where each decision corresponds to different dimensions of user data, such as location information, search keywords, etc. In the initial stage, the model selects actions through a random policy and receives feedback according to the reward function. As the learning progresses, the model gradually adjusts its policy to maximize the cumulative reward.

[0079] In some embodiments, as Figure 3 shown, determining the user anomaly score according to the proportion of the preset keywords in the search keywords may include: S301 to S303.

[0080] S301, dividing the number of the first preset keywords by the number of the first search keywords to obtain the keyword ratio, where the first search keywords are the search keywords within a preset period, and the first preset keywords are the preset keywords in the first search keywords;

[0081] S302, dividing the number of the first preset keywords minus the second preset keywords by the second search keywords to obtain the keyword growth rate, where the second search keywords are the search keywords in the previous period of the preset period, and the second preset keywords are the preset keywords in the second search keywords;

[0082] S303, multiplying the keyword growth rate plus one by the keyword ratio to obtain the user anomaly score.

[0083] In some embodiments, the formula for calculating the user anomaly score is: User anomaly score = (First preset keyword / First search keyword) * (1 + Keyword growth rate).

[0084] If the number of times a user searches for high-risk words within a short period of time significantly exceeds the set threshold, it is marked as a potential abnormal behavior. For users marked as abnormal, comprehensive evaluation is carried out in combination with other behavior data to reduce the false alarm rate. Other behavior data refers to user location base station data, medical communication behavior data, etc. By analyzing the user's mobile phone base station positioning data, the user's daily activity trajectory is drawn. Check whether these trajectories highly coincide with medical institutions, especially within a certain period of time after searching for high-risk words. At the same time, check the call and text message records of the user with medical institutions, insurance companies, etc., and trace back the historical account period. At the same time, trace back data such as the usage traffic and number of times of the user browsing medical institution APPs, mini-programs, and official accounts.

[0085] Based on the analysis results of the above multi-dimensional data, a comprehensive evaluation model is established to calculate a risk score for each abnormal user. Risk score = a × Anomaly score + b × Coincidence degree score of base station positioning and medical insurance medical institutions + c × Communication behavior score + d × Medical Internet behavior score. Where a refers to the weight of the high-risk word search frequency, b refers to the weight of the coincidence degree of base station positioning and medical insurance medical institutions, c refers to the communication behavior weight, d refers to the weight of the medical institution Internet behavior, the communication behavior score refers to the proportion of relevant call and text message behaviors in the total communication behaviors, and the medical Internet behavior score refers to the sum of the usage times and time of the user browsing medical institution-related APPs, mini-programs, and official accounts accounting for the proportion of the total Internet behavior times and time.

[0086] According to different data intervals of different dimensions, different thresholds are divided, and the final score is divided into high, medium, and low-risk insurance fraudsters. Based on expert experience, the score in the interval [0, 0.3) is low risk, [0.3, 0.6) is medium risk, and [0.6, 1] is high risk. Through the above comprehensive evaluation process, the false alarm rate caused by single-dimensional data can be reduced.

[0087] In some embodiments, according to the actual application effect and historical data, the threshold of the anomaly score is dynamically adjusted to improve the accuracy. Based on expert experience, if the false alarm rate is higher than 5%, the threshold in the high-risk interval is increased by 10% to reduce false alarms; if the missed alarm rate exceeds 5%, the threshold in the low-risk interval is reduced by 10% to increase the number of detected anomalies. At the same time, the high-risk word list in the keyword search database is adjusted according to the actual situation to adapt to the medical insurance policies and insurance fraud behavior characteristics of different regions. During the process of dynamically adjusting the threshold, the effects of the new thresholds are continuously tested and verified to reduce misjudgment and missed judgment.

[0088] In some embodiments, before determining the user anomaly score based on the proportion of the preset keyword in the search keywords, the method further includes:

[0089] Respectively determine the preset keyword semantic vector and the search keyword semantic vector corresponding to the preset keyword and the search keyword;

[0090] Calculate the similarity between the preset keyword semantic vector and the search keyword semantic vector, and use the search keywords with a similarity higher than the set threshold as associated keywords;

[0091] Determine the user anomaly score according to the proportions of the preset keyword and the associated keywords in the search keywords.

[0092] By calculating the semantic vector similarity between the preset keyword and the search keyword, it is possible to more accurately identify the search keywords semantically related to the preset keyword. This method not only considers the surface matching of keywords, but also considers the semantic relevance, reducing false alarms caused by surface keyword matching, and thus more accurately evaluating the anomaly degree of user behavior.

[0093] In some embodiments, as Figure 4 shown, calculating the local density and the minimum distance of user data points, determining the cluster center according to the local density and the minimum distance and performing clustering to obtain user cluster labels, may include: S401 to S404.

[0094] S401, calculate the average value and the standard deviation of the distances between user data points;

[0095] S402, add the product of the standard deviation and the preset adjustable parameter to the average value to obtain the truncation distance;

[0096] S403, determine the local density of each user data point according to the truncation distance and the distances between user data points, and obtain the target distance from each user data point to the target data point, where the target data point is the neighbor data point with the highest local density of the user data point;

[0097] S404, select the user data points with the truncation distance and the target distance greater than the set threshold as the cluster center and perform clustering to obtain user cluster labels.

[0098] User data points whose local density and minimum distance meet the set threshold are selected as cluster centers, representing different user behavior patterns; for each user data point that is not a cluster center, it is assigned to the cluster with the nearest cluster center, thereby assisting in user behavior grouping identification and more accurately understanding user medical behavior. For example, for medical insurance data, if a user visits different medical institutions multiple times in a short period of time and the reimbursement amount is abnormally high, or for operator data, the user's geographic location data does not match the location of the visit, based on the above data, if a user's local density is low, his behavior is significantly different from that of most users, and the minimum distance is high, far away from other high-density areas, then the user may be a potential fraudster.

[0099] In some embodiments, calculating the local density and minimum distance of the user data points may include:

[0100] Calculate the local density ρ i , for each user data point x i , and its local density calculation formula is:

[0101]

[0102] d c =μ+k·σ

[0103] Among them, d ij is the user data point x i and x j The distance between c is the cutoff distance, indicating the range of proximity, and μ is the user data point x i is the mean of the distances between them, σ is the standard deviation, and k is an adjustable parameter.

[0104] Since the fixed cut-off distance cannot adapt to the inherent structure and changes of operators and medical insurance data, the model may lack generalization ability in new medical groups, affecting the clustering stability. The cut-off distance is adaptively adjusted and dynamically adjusted based on the data distribution characteristics of the user feature set. The adaptability is improved according to the corresponding data characteristics, thereby improving the accuracy of medical risk detection. For outliers, the isolation forest anomaly detection algorithm is used for verification.

[0105] Calculate the minimum distance δ i , for each user data point x i , find the minimum distance to the point with the highest local density among all its neighbors:

[0106]

[0107] Among them, δ i is the user data point x iThe minimum distance to the highest local density among all its neighbors, ρ j For the user data point x i The local density, ρ i For the user data point x j The local density, for the point with the highest local density, δ i Is defined as the distance between the second nearest points to it.

[0108] In some embodiments, as Figure 5 Shown, determining the target user label according to the user clustering label, the user prediction label and the relative distance of the user data point, and obtaining the updated clustering result may include: S501 to S503.

[0109] S501, calculating the confidence of the user clustering label, and using the user clustering label with the confidence greater than the set threshold as the target classification label of the user data point;

[0110] S502, calculating the target relative distance between the user data point and the target data point, where the target data point is the data point with the target classification label that is closest to the user data point;

[0111] S503, when the user clustering label and the user prediction label of the user data point are different and the target relative distance is less than the set threshold, updating the user clustering label according to the user prediction label to obtain the target user label, and updating the clustering result according to the target user label.

[0112] By calculating the target relative distance between the user data point and the target data point, and updating the user clustering label according to the user prediction label when specific conditions are met, the clustering result can be dynamically adjusted, which not only depends on the clustering result, but also combines the confidence evaluation, thereby improving the reliability of the clustering result.

[0113] In some embodiments, inputting the training medical communication behavior data into a pre-trained classifier, and determining the user prediction label according to the training medical communication behavior data and the preset weight may include:

[0114] Gradually expanding the labeled data set in an iterative manner. First, use the labeled data D l Train an initial classifier f θ , and the training objective is to minimize the loss function, using the cross-entropy loss function:

[0115]

[0116] Where, L CE Is the cross-entropy loss function value, used to measure the difference between the model prediction value and the true label, x i Is the user data point, y i Is xi The true label, D l is the labeled dataset, θ is the parameter of the classifier, f θ (x i ) is the output of the classifier, n l is the number of data points in the labeled dataset.

[0117] Using the initial classifier f θ to predict the unlabeled dataset D u and assign pseudo-labels to the prediction results with high confidence The calculation method is:

[0118]

[0119] Among them, is the pseudo-label, that is, the predicted class of the unlabeled data point, c is the index of the class, f θ (x j ′) c is the output vector of the model for x j ′, and x j ′ is the user data point in the unlabeled dataset.

[0120]

[0121] Among them, D u is the unlabeled dataset, x j ′ is the user data point in the unlabeled dataset, n u is the number of data points in the unlabeled dataset.

[0122] When adjusting the pseudo-labels, it is carried out according to the relative distance between the prediction result of the classifier and the user data point. For the unlabeled user data point x u , the assigned pseudo-label is The prediction of the classifier f θ for x u is The nearest labeled point is x v .

[0123] The relative distance is:

[0124]

[0125] Among them, r uv is the relative distance between the data points x u and x v , d(x u ,x v ) is the distance between x u and x v , and d avg(xv) is the distance from x vAverage distance between adjacent data points.

[0126] If the predicted label is inconsistent with the pseudo-label and the relative distance is small, it indicates that x u may be in the region of the classification boundary. The pseudo-label is adjusted as follows:

[0127]

[0128] where θ is a preset threshold for determining whether to adjust the pseudo-label.

[0129] Repeat the above process. For each iteration, use the updated training set, that is, the new training set D formed by merging the pseudo-labeled data and the labeled data (k) , to increase the number of the training set of the model. Use the new training set D (k) to update the model parameters by minimizing the loss function. Integrate operator data and medical insurance data, and identify potential fraud behaviors by analyzing the unreasonable user abnormal medical treatment frequency, geographical location, etc., and gradually improve the ability to identify fraud behaviors.

[0130]

[0131] D (k) = D L ∪ D′ U

[0132] where θ (k) is the model parameter after the k-th iteration, D (k) is the training set of the k-th iteration, τ is the loss function, f(x; θ) is the output of the classifier, representing the prediction of the model for the input data x, y is the label, which can be the true label of the labeled data or the pseudo-label of the unlabeled data, D L is the labeled data set, and D′ U is the adjusted unlabeled data set.

[0133] When the change in the model parameters is less than a certain threshold, or the model iteration reaches the preset number of times k, stop the iteration. Use the model parameter θ (k) obtained in the last iteration as the final result. The iteration formula is:

[0134] ||θ (k) - θ (k-1) || < ε

[0135] where θ (k-1) is the model parameter after the (k - 1)-th iteration, and ε is the convergence threshold.

[0136] Based on an improved self-training algorithm, the relative distance is introduced to compare the pseudo-labels and the predicted labels. Under limited samples, it ensures that each point is in the appropriate cluster, improving the accuracy of medical fraud identification.

[0137] In some embodiments, after inputting the user's medical communication behavior data into the medical risk detection model and determining the medical risk classification result corresponding to the user's medical communication behavior data according to the pre-trained target expectation value table, the method further includes:

[0138] Obtain the true medical risk classification result of the user's medical communication behavior data;

[0139] Update the target expectation value table according to the user's medical communication behavior data and the true medical risk classification result.

[0140] By regularly updating the target expectation value table, it supports the continuous improvement of the model. By continuously updating the target expectation value table, the model can continuously learn new data patterns, thus better adapting to the changing medical environment, reducing the misclassification rate of the model, and being able to more accurately identify and classify the medical risks of users.

[0141] Among them, when the model successfully identifies a true medical behavior or effectively identifies an insurance fraud behavior, positive feedback is given. A larger positive reward is provided when an insurance fraud behavior is found to encourage the model to better identify insurance fraud behaviors; if the model fails to effectively identify an insurance fraud behavior or mislabels a normal medical treatment as an insurance fraud behavior, the model receives negative feedback.

[0142] Figure 6 Fig. 600 shows a device for medical risk detection provided by an embodiment of the present application. As shown in the figure, the device may include:

[0143] An acquisition module 601, configured to acquire the user's medical communication behavior data;

[0144] A determination module 602, configured to input the user's medical communication behavior data into the medical risk detection model, and determine the medical risk classification result corresponding to the user's medical communication behavior data according to the pre-trained target expectation value table; the target expectation value table includes the corresponding relationship between the medical communication behavior data and the medical risk classification result, and the target expectation value table is obtained by training the medical communication behavior data and the target user labels. The target user labels are determined according to the user clustering label, the user prediction label, and the relative distance of the user data points. The user clustering label is obtained by clustering according to the local density and the minimum distance of the user data points, and the user prediction label is obtained through a pre-trained classifier.

[0145] In some embodiments, the device 600 for medical risk detection may further include:

[0146] The acquisition module 601 is further configured to acquire training medical communication behavior data;

[0147] The clustering module is configured to calculate the local density and minimum distance of user data points, determine the clustering center according to the local density and minimum distance, and perform clustering to obtain user clustering labels, where the user data points include training medical communication behavior data;

[0148] The classification module is configured to input the training medical communication behavior data into a pre-trained classifier, and determine user prediction labels according to the training medical communication behavior data and preset weights;

[0149] The determination module 602 is further configured to calculate the relative distance of user data points, and determine target user labels according to the user clustering labels, user prediction labels, and the relative distance of user data points;

[0150] The training module is configured to train the medical risk detection model according to the training medical communication behavior data and target user labels to obtain a target expected value table, and update the risk detection model according to the target expected value table.

[0151] In some embodiments, the determination module 602 is further configured to determine a user anomaly score according to the proportion of preset keywords in the search keywords;

[0152] The classification module is further configured to input the user anomaly score into a pre-trained classifier.

[0153] In some embodiments, the medical risk detection device 600 may further include:

[0154] The calculation module is configured to divide the number of first preset keywords by the number of first search keywords to obtain a keyword ratio, where the first search keywords are the search keywords within a preset period, and the first preset keywords are the preset keywords in the first search keywords;

[0155] The calculation module is further configured to divide the number of first preset keywords minus the second preset keywords by the second search keywords to obtain a keyword growth rate, where the second search keywords are the search keywords in the previous period of the preset period, and the second preset keywords are the preset keywords in the second search keywords;

[0156] The calculation module is further configured to multiply the keyword growth rate plus one by the keyword ratio to obtain a user anomaly score.

[0157] In some embodiments, the determination module 602 is further configured to respectively determine a preset keyword semantic vector and a search keyword semantic vector corresponding to the preset keywords and search keywords;

[0158] The calculation module is further configured to calculate the similarity between the semantic vectors of the preset keywords and the search keywords, and use the search keywords with a similarity higher than the set threshold as associated keywords;

[0159] The determination module 602 is further configured to determine the user anomaly score according to the proportions of the preset keywords and the associated keywords in the search keywords.

[0160] In some embodiments, the calculation module is further configured to calculate the average value and the standard deviation of the distances between user data points;

[0161] The calculation module is further configured to add the product of the standard deviation and a preset adjustable parameter to the average value to obtain a truncation distance;

[0162] The determination module 602 is further configured to determine the local density of each user data point according to the truncation distance and the distance between user data points, and obtain the target distance from each user data point to the target data point, where the target data point is the neighbor data point with the highest local density of the user data point;

[0163] The clustering module is further configured to select the user data points with truncation distances and target distances greater than the set threshold as clustering centers and perform clustering to obtain user clustering labels.

[0164] In some embodiments, the device for detecting medical treatment risks may further include:

[0165] The calculation module is further configured to calculate the confidence level of the user clustering label, and use the user clustering label with a confidence level greater than the set threshold as the target classification label of the user data point;

[0166] The calculation module is further configured to calculate the target relative distance between the user data point and the target data point, where the target data point is the data point with the target classification label that is closest to the user data point;

[0167] The update module is configured to, when the user clustering label and the user prediction label of the user data point are different and the target relative distance is less than the set threshold, update the user clustering label according to the user prediction label to obtain the target user label, and update the clustering result according to the target user label.

[0168] In some embodiments, the acquisition module is further configured to acquire the true medical treatment risk classification result of the user's medical treatment communication behavior data;

[0169] The update module is further configured to update the target expected value table according to the user's medical treatment communication behavior data and the true medical treatment risk classification result.

[0170] Figure 6 Each module in the shown device can implement Figure 1 each step in, and achieve the corresponding technical effects, which will not be elaborated here for the sake of brevity.

[0171] Figure 7 The figure shows a schematic diagram of the hardware structure of the terminal device provided by the embodiments of the present application.

[0172] The terminal device may include a processor 701 and a memory 702 storing computer program instructions.

[0173] Specifically, the above-mentioned processor 701 may include a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0174] The memory 702 may include a mass storage for data or instructions. By way of example and not limitation, the memory 702 may include a Hard Disk Drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. In one example, the memory 702 may include removable or non-removable (or fixed) media, or the memory 702 is a non-volatile solid-state memory. The memory 702 may be inside or outside the integrated gateway disaster recovery device.

[0175] In one example, the memory 702 may include a Read Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method of detecting medical treatment risks according to the present disclosure.

[0176] The processor 701 reads and executes the computer program instructions stored in the memory 702 to implement Figure 1 the method for detecting medical treatment risks in the illustrated embodiments.

[0177] In one example, the terminal device may further include a communication interface 703 and a bus 704. Among them, as Figure 7 shown, the processor 701, the memory 702, and the communication interface 703 are connected through the bus 704 to complete mutual communication.

[0178] The communication interface 703 is mainly used to implement communication between various modules, devices, units, and / or equipment in the embodiments of the present application.

[0179] The bus 704 includes hardware, software, or both, and couples the components of the terminal device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. In suitable cases, the bus 704 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0180] In addition, in combination with the method for detecting medical treatment risks in the above embodiments, the embodiments of the present application may provide a computer storage medium to implement. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, any one of the methods for detecting medical treatment risks in the above embodiments is implemented.

[0181] The embodiments of the present application also provide a computer program product, including a computer program, and when the computer program is executed by a processor, any one of the methods for detecting medical treatment risks in the above embodiments is implemented.

[0182] It should be clear that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0183] The functional blocks shown in the above-described structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or text segments used to perform the required tasks. The program or text segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memory (ROM), flash memory, erasable read-only memory (EROM), floppy disks, compact disc read-only memory (CD-ROM), optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The text segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0184] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0185] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It should also be understood that each block in the block diagrams and / or flowcharts, and the combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0186] As described above, this is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

Claims

1. A method for detecting medical treatment risks, characterized in that, Including: Obtain the user's medical treatment communication behavior data; Input the user's medical treatment communication behavior data into a medical treatment risk detection model, and determine the medical treatment risk classification result corresponding to the user's medical treatment communication behavior data according to a pre-trained target expected value table; The target expected value table includes the corresponding relationship between medical treatment communication behavior data and medical treatment risk classification results. The target expected value table is obtained by training medical treatment communication behavior data and target user labels. The target user labels are determined according to user clustering labels, user prediction labels, and the relative distance of user data points. The user clustering labels are obtained by clustering based on the local density and minimum distance of user data points. The user prediction labels are obtained through a pre-trained classifier.

2. The method for detecting medical risks according to claim 1, wherein Before inputting the user's medical treatment communication behavior data into the medical treatment risk detection model and determining the medical treatment risk classification result corresponding to the user's medical treatment communication behavior data according to the pre-trained target expected value table, the method further includes: Obtain training medical treatment communication behavior data; Calculate the local density and minimum distance of user data points, determine the clustering center based on the local density and minimum distance, and perform clustering to obtain user clustering labels. The user data points include training medical treatment communication behavior data; Input the training medical treatment communication behavior data into a pre-trained classifier, and determine user prediction labels according to the training medical treatment communication behavior data and preset weights; Calculate the relative distance of user data points, and determine target user labels according to user clustering labels, user prediction labels, and the relative distance of user data points; Train the medical treatment risk detection model according to the training medical treatment communication behavior data and target user labels to obtain a target expected value table, and update the risk detection model according to the target expected value table.

3. The method for detecting medical treatment risks according to claim 2, wherein, The training medical treatment communication behavior data includes user search keywords. Before inputting the training medical treatment communication behavior data into the pre-trained classifier, the method further includes: Determine the user anomaly score according to the proportion of preset keywords in the search keywords; Input the user anomaly score into the pre-trained classifier.

4. The method for detecting medical treatment risks according to claim 3, wherein The determination of the user anomaly score according to the proportion of preset keywords in the search keywords includes: Divide the number of first preset keywords by the number of first search keywords to obtain a keyword ratio. The first search keywords are the search keywords within a preset period, and the first preset keywords are the preset keywords in the first search keywords; Divide the number of first preset keywords minus the number of second preset keywords by the second search keywords to obtain a keyword growth rate. The second search keywords are the search keywords in the previous period of the preset period, and the second preset keywords are the preset keywords in the second search keywords; Multiply the keyword growth rate plus one by the keyword ratio to obtain the user anomaly score.

5. The method for detecting medical treatment risks according to claim 3, wherein Before determining the user anomaly score according to the proportion of preset keywords in the search keywords, the method further includes: Respectively determine the preset keyword semantic vector and the search keyword semantic vector corresponding to the preset keywords and search keywords; Calculate the similarity between the semantic vectors of the preset keyword and the search keyword, and use the search keyword with a similarity higher than the set threshold as the associated keyword; Determine the user anomaly score according to the proportion of the preset keyword and the associated keyword in the search keyword.

6. The method for detecting medical treatment risks according to claim 2, wherein, The calculating the local density and the minimum distance of the user data points, determining the clustering center according to the local density and the minimum distance, and performing clustering to obtain the user clustering label, including: Calculate the average value and the standard deviation of the distances between user data points; Add the product of the standard deviation and the preset adjustable parameter to the average value to obtain the truncation distance; Determine the local density of each user data point according to the truncation distance and the distance between user data points, and obtain the target distance from each user data point to the target data point, where the target data point is the neighbor data point with the highest local density of the user data point; Select the user data points with the truncation distance and the target distance greater than the set threshold as the clustering center and perform clustering to obtain the user clustering label.

7. The method for detecting medical treatment risks according to claim 2, wherein The determining the target user label according to the user clustering label, the user prediction label, and the relative distance of the user data point, and obtaining the updated clustering result, including: Calculate the confidence of the user clustering label, and use the user clustering label with the confidence greater than the set threshold as the target classification label of the data point; Calculate the target relative distance between the user data point and the target data point, where the target data point is the data point with the target classification label that is closest to the user data point; When the user clustering label and the user prediction label of the data point are different and the target relative distance is less than the set threshold, update the user clustering label according to the user prediction label to obtain the target user label, and update the clustering result according to the target user label.

8. The method for detecting medical treatment risks according to claim 1, wherein After inputting the user's medical treatment communication behavior data into the medical treatment risk detection model and determining the medical treatment risk classification result corresponding to the user's medical treatment communication behavior data according to the pre-trained target expectation value table, the method further includes: Obtain the true medical treatment risk classification result of the user's medical treatment communication behavior data; Update the target expectation value table according to the user's medical treatment communication behavior data and the true medical treatment risk classification result.

9. A device for detecting medical treatment risks, characterized in that, Including: An acquisition module for acquiring the user's medical treatment communication behavior data; A determination module for inputting the user's medical treatment communication behavior data into the medical treatment risk detection model and determining the medical treatment risk classification result corresponding to the user's medical treatment communication behavior data according to the pre-trained target expectation value table; The target expectation value table includes the corresponding relationship between the medical treatment communication behavior data and the medical treatment risk classification result. The target expectation value table is obtained by training the medical treatment communication behavior data and the target user label. The target user label is determined according to the user clustering label, the user prediction label, and the relative distance of the user data point. The user clustering label is obtained by clustering according to the local density and the minimum distance of the user data point. The user prediction label is obtained by a pre-trained classifier.

10. A terminal device, characterized in that, The device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the method for detecting medical treatment risks as described in any one of claims 1-8 is implemented.

11. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the method for detecting medical treatment risks described in any one of claims 1-8 is implemented.

12. A computer program product, characterized in that, When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute the method for detecting medical treatment risks described in any one of claims 1-8.

Citation Information

Cited By

  • Sparse positive sample risk discrimination method and system based on clustering analysis

    CN122091258A