A method and device for preventing medical insurance fraud

By collecting medical data in real time and combining it with machine learning and deep learning technologies, fraud risks are dynamically assessed, solving the problems of high false negative and high false alarm rates in existing medical insurance fraud prevention methods, and achieving accurate identification and efficient processing of medical insurance fraud.

CN122134469APending Publication Date: 2026-06-02PICC INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PICC INFORMATION TECH CO LTD
Filing Date
2026-01-07
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing methods for preventing medical insurance fraud rely on manual review and static rule matching, resulting in high false negative or false positive rates. They also lack the ability to fuse multimodal features, making it difficult to cope with new fraud methods and causing losses to the medical insurance fund.

Method used

Medical data is collected in real time through API interfaces, and a normal medical consumption behavior model is established by combining machine learning algorithms. Fraud analysis is carried out using deep learning image recognition and natural language processing technologies to dynamically assess fraud risks and trigger early warning mechanisms and handling procedures.

Benefits of technology

It enables real-time and accurate identification of medical insurance fraud, improves the accuracy and efficiency of fraud detection and reduces the economic and reputational risks for insurance companies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134469A_ABST
    Figure CN122134469A_ABST
Patent Text Reader

Abstract

This invention proposes a method and device for preventing medical insurance fraud, belonging to the field of insurance regulatory technology, aiming to solve the problems of low efficiency and poor accuracy in existing manual review of claims materials. The system includes a data acquisition layer, a data processing layer, an analysis and identification layer, an assessment and early warning layer, and a decision-making layer. It connects to systems of medical institutions, medical insurance departments, and pharmacies through API interfaces to collect and clean multi-source data; it establishes a normal medical consumption behavior model based on machine learning algorithms, combining deep learning image recognition and natural language processing technologies to accurately identify the authenticity of invoices and the consistency of medical behavior logic; it classifies risk levels through risk assessment algorithms, triggering early warnings and suspending the claims process for medium- and high-risk applications, and taking corresponding measures after verification. This system achieves real-time monitoring and accurate identification of fraudulent behavior, improves claims efficiency, reduces insurance company risks, and maintains the fairness and sustainability of the medical insurance system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of insurance regulatory technology, and in particular to a method and device for preventing medical insurance fraud. Background Technology

[0002] Medical insurance fraud detection, as a core technology in the insurance regulatory field, is widely used in the compliance management of the medical security system. With the continuous expansion of medical insurance coverage and the intelligent evolution of fraud methods, traditional manual review models are no longer sufficient to meet the demands of massive data processing. Existing technological systems mainly rely on human experience and judgment, conducting preliminary screening through paper document review and rule matching. Their technical architecture covers key aspects such as data collection, manual review, and result determination. Specifically, in this technological ecosystem, medical institutions, pharmacies, and medical insurance departments operate independently, with data interaction relying on manual entry and paper-based transmission, resulting in particularly prominent data silos. Therefore, when dealing with new fraud patterns such as forged invoices and fabricated illnesses, existing systems often need to combine single-dimensional features such as medical coding rules and cost threshold judgments, lacking the ability for multimodal collaborative analysis of medical behavior patterns, textual semantic logic, and image features.

[0003] However, existing methods for preventing medical insurance fraud rely heavily on manual review and static rule matching, lacking a dynamic baseline model of medical behavior. This can lead to a double dilemma of high false negative or high false positive rates. Specifically, traditional manual review suffers from low data processing efficiency, typically handling fewer than 5,000 applications per day, and current systems fail to semantically link these applications to medical knowledge graphs. Furthermore, the recognition rate for printed features of forged invoices is less than 65%, and the accuracy rate for textual logic verification of fabricated medical conditions is less than 72%. These technical deficiencies create significant blind spots in fraud detection. Consequently, existing technologies often employ single-feature analysis, which suffers from insufficient multimodal feature fusion and weak adversarial fraud pattern recognition capabilities, thus impacting the security of the medical insurance fund and the system's credibility. According to WHO statistics, these technical deficiencies cause hundreds of billions of yuan in losses to the medical insurance fund annually, and with the emergence of AI-generated forged data, traditional rule engines are struggling to cope with the continuous iteration of new fraud methods. Summary of the Invention

[0004] The present invention aims to at least partially solve one of the technical problems in the related art.

[0005] Therefore, the first objective of this invention is to propose a method for preventing medical insurance fraud.

[0006] Another objective of this invention is to provide a medical insurance fraud prevention device.

[0007] The third objective of this invention is to provide a computer device.

[0008] A fourth objective of this invention is to provide a non-transitory computer-readable storage medium.

[0009] To achieve the above objectives, a first aspect of the present invention provides a method for preventing medical insurance fraud, comprising: S1 collects medical data from medical institutions, pharmacies, and medical insurance departments in real time through API interfaces and database connections; S2 establishes a normal medical consumption behavior model based on machine learning algorithms, and uses deep learning image recognition algorithms to verify the authenticity of invoices based on their printing features. At the same time, it uses natural language processing technology to analyze the semantic and logical consistency between diagnostic descriptions and medication and treatment methods. S3. Based on the normal medical consumption behavior model, the results of invoice authenticity identification and semantic logic analysis, extract risk features and use a dynamic risk assessment model to calculate the probability value of fraud risk, and classify the risk into different levels. S4 triggers an early warning mechanism and suspends the claims process for claims with medium to high risk levels. At the same time, it sends an early warning signal to the insurance company's review personnel for investigation and verification. Based on the investigation results, it either adds the company to the fraud blacklist or restores the normal claims process.

[0010] In one embodiment of the present invention, S1 includes: S11 uses SSL / TLS encryption protocol to encrypt medical data during transmission, ensuring data security during transmission; S12 interacts with medical institution systems in real time via a RESTful API interface, encapsulates medical data in JSON format, and performs integrity verification.

[0011] In one embodiment of the present invention, S2 includes: S21 uses the LeNet-5 improved convolutional neural network structure, where the convolutional layer C1 contains 6 5×5 convolutional kernels and the pooling layer S2 uses a 2×2 max pooling window. S22: Construct a knowledge graph containing over 1.2 million medical entity relationships, and verify the logical consistency between diagnostic descriptions and medication and treatment methods through semantic matching algorithms.

[0012] In one embodiment of the present invention, S4 includes: S41 verifies the identity of the policyholder through two methods: telephone interviews and on-site investigations. The telephone interviews use a preset script template for voice recognition interaction. S42 stipulates that multi-dimensional measures will be taken against policyholders confirmed to have committed fraud, including adding the policyholder's information to the fraud blacklist, freezing their medical insurance accounts, and simultaneously submitting electronic evidence packages to the regulatory authorities.

[0013] To achieve the above objectives, a second aspect of the present invention provides a medical insurance fraud prevention device, comprising: The multi-source data real-time acquisition module is used to collect medical data from medical institutions, pharmacies, and medical insurance departments in real time through API interfaces and database connections; The multi-algorithm behavior modeling and feature analysis module is used to build a normal medical consumption behavior model based on machine learning algorithms, and to verify the authenticity of invoices by combining deep learning image recognition algorithms with printing features. At the same time, natural language processing technology is used to analyze the semantic and logical consistency between diagnostic descriptions and medication and treatment methods. The dynamic risk assessment and classification module is used to extract risk features and calculate the fraud risk probability value using the dynamic risk assessment model based on the normal medical consumption behavior model, the invoice authenticity identification results and the semantic logic analysis results, and classify the risk into different levels. The early warning triggering and handling module is used to trigger the early warning mechanism and suspend the claim process for claims with medium to high risk levels. At the same time, it sends an early warning signal to the insurance company's review personnel for investigation and verification. Based on the investigation results, it executes the fraud blacklist entry or resumes the normal claims process.

[0014] The present invention provides a method and apparatus for preventing medical insurance fraud, which can identify medical insurance fraud in real time and accurately, effectively improve the accuracy and efficiency of fraud detection, and reduce the economic and reputational risks of insurance companies.

[0015] To achieve the above objectives, a third aspect of this application provides a computer device comprising a processor and a memory; wherein the processor runs a program corresponding to the executable program code stored in the memory, for implementing a medical insurance fraud prevention method as described in the first aspect embodiment.

[0016] To achieve the above objectives, a fourth aspect of this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a medical insurance fraud prevention method as described in the first aspect.

[0017] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0018] Figure 1 This is a flowchart of a method for analyzing the flow rate, efficiency, and water consumption rate of a hydropower plant generator unit according to an embodiment of the present invention; Figure 2This is a structural diagram of an analysis device for the flow rate, efficiency, and water consumption rate of a hydropower plant generator set according to an embodiment of the present invention. Figure 3 It is a computer device according to an embodiment of the present invention. Detailed Implementation

[0019] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0021] The following description, with reference to the accompanying drawings, describes a method and apparatus for preventing medical insurance fraud according to an embodiment of the present invention.

[0022] Example 1 Figure 1 This is a flowchart of a medical insurance fraud prevention method according to an embodiment of the present invention, such as... Figure 1 As shown, it includes: S1 collects medical data from medical institutions, pharmacies, and medical insurance departments in real time through API interfaces and database connections.

[0023] Specifically, the present invention collects medical data from medical institutions, pharmacies, and medical insurance departments in real time through API interfaces and database connections. The technical implementation principle of this step is based on the integration of a distributed system architecture and heterogeneous data sources, employing standardized data interface protocols to achieve cross-institutional and cross-system data synchronization and transmission. The specific operation methods include: deploying a data collection agent program at the medical institution level to establish a secure communication link with the system's main server via HTTPS protocol; periodically extracting drug sales records at the pharmacy level through database connections; and achieving batch data synchronization at the medical insurance department through a data exchange platform. The collected data types cover diagnostic records, drug prescriptions, treatment cost details, electronic invoice information, and basic information of the insured. Invoice information must meet national tax standard formats, such as invoice code, invoice number, invoice date, and amount.

[0024] Furthermore, the system supports multi-threaded concurrent data acquisition, with a maximum configurable concurrency level. To adapt to the responsiveness of different data sources, the data collection frequency can be set to real-time (<1 second latency) or near real-time (5-30 second latency) according to business needs, and supports breakpoint resumption and data verification mechanisms, such as MD5 hash verification and data integrity checks. Furthermore, the system employs authentication mechanisms such as OAuth2.0 or JWT to ensure the legitimacy and security of data access.

[0025] Furthermore, the quality of the data collected in this step directly affects the training effect of subsequent machine learning models and the accuracy of fraud detection. By integrating multi-source data, the system can construct a complete profile of medical behavior, providing reliable data support for abnormal behavior detection, image authenticity recognition, and natural language logic analysis. At the same time, the efficiency and security of this step ensure the stable operation of the system in high-concurrency, multi-institutional access environments, laying a solid foundation for achieving intelligent and automated medical insurance fraud prevention.

[0026] Furthermore, S1 includes: S11 uses SSL / TLS encryption protocol to encrypt medical data during transmission, ensuring data security during transmission.

[0027] Specifically, in the process of medical data transmission, the SSL / TLS encryption protocol is used to encrypt the data end-to-end. SSL and its evolved version TLS are security protocols widely used in Internet communication. The principle is to achieve identity authentication, data encryption and integrity verification of the communicating parties through a combination of asymmetric encryption, symmetric encryption and message authentication code (MAC).

[0028] Furthermore, in the medical insurance fraud prevention system of this invention, the data acquisition module exchanges data with external systems such as medical institutions, pharmacies, and medical insurance departments via the HTTPS protocol. HTTPS is essentially a secure encapsulation of the HTTP protocol on top of the SSL / TLS protocol. Specifically, before establishing a connection, the system first performs a handshake protocol, ensuring the legitimacy of both communicating parties by exchanging public keys, generating session keys, and verifying digital certificates. After the handshake is completed, the system uses a symmetric encryption algorithm to encrypt the transmitted data and simultaneously employs the HMAC algorithm to verify data integrity, preventing data tampering during transmission.

[0029] Furthermore, the system supports TLS 1.2 and above, with TLS 1.3 recommended for improved performance and security. The preferred encryption suites include modern security combinations such as ECDHE, RSA, AES256, GCM, and SHA384, ensuring forward confidentiality and resistance to man-in-the-middle attacks. For certificate management, the system uses X.509 standard certificates issued by a trusted CA and supports OCSP real-time certificate status verification to enhance the real-time nature and reliability of authentication.

[0030] Furthermore, this encryption mechanism is widely used in API communication, database connections, and data synchronization processes between the system and medical institutions, as well as in data exchange platforms. For example, when the system retrieves diagnostic records from a hospital server via a RESTful API, all data is transmitted through a TLS encrypted channel, ensuring that the data is not eavesdropped on or tampered with, whether on a public or private network. In addition, the system also supports encrypted transmission methods based on protocols such as IPSec or MQTT over TLS to adapt to the data interaction needs of different network environments.

[0031] Furthermore, this step, on the one hand, provides a secure channel for the real-time acquisition of multi-source data, ensuring the confidentiality and integrity of the data and laying a reliable foundation for subsequent data integration and analysis; on the other hand, it complies with the data transmission security requirements in the "Basic Requirements for Information System Security Level Protection," enhancing the system's compliance and credibility. Through the deployment of SSL / TLS encryption protocols, the system effectively prevents data leakage and man-in-the-middle attacks, thereby ensuring the secure flow of medical insurance data at the source.

[0032] S12 interacts with medical institution systems in real time via a RESTful API interface, encapsulates medical data in JSON format, and performs integrity verification.

[0033] Specifically, the step of "real-time data interaction with medical institution systems through RESTful API interfaces, encapsulating medical data in JSON format and performing integrity verification" in this invention is based on modern Web service architecture. It achieves real-time data synchronization and verification across systems and platforms through standardized interfaces, ensuring that the system can obtain structured and parsable medical data, providing high-quality input for subsequent fraud identification.

[0034] Furthermore, the system adopts RESTful API as the core communication method with medical institution systems. Based on the HTTP protocol, RESTful API enables real-time acquisition and updating of medical data by defining unified resource paths and operation methods. Medical data is encapsulated in JSON format before transmission, a format with good readability and cross-language compatibility, facilitating system parsing and processing.

[0035] Furthermore, to ensure data integrity and consistency, the system performs a multi-dimensional verification mechanism upon receiving JSON data. First, it performs format verification, checking whether the JSON conforms to the preset schema specifications, such as the existence of fields, data type matching, and the completeness of required fields. Second, it performs logical verification, such as verifying whether the diagnosis date is within a reasonable range and whether the medication and diagnosis have a medical relevance. In addition, the system employs a hash verification mechanism, calculating hash values ​​for key fields (such as invoice codes, amounts, and diagnosis codes) and comparing them with hash values ​​provided by the source system to ensure that the data has not been tampered with during transmission.

[0036] Furthermore, the system supports custom validation rules, such as setting field length limits (e.g., `invoice_number` length is 12 digits) and amount range limits (e.g., the amount of a single prescription does not exceed a certain limit). ), time interval threshold (e.g., the interval between two visits must not be less than 10 ... The system also supports asynchronous data retrieval and polling mechanisms to ensure the continuity and timeliness of data collection even when medical institutions experience system response delays.

[0037] Furthermore, this step is widely deployed in practical applications for system integration with healthcare entities such as hospitals, clinics, and pharmacies, and is particularly suitable for scenarios requiring high-frequency, low-latency data updates. Through standardized interfaces and structured data encapsulation, the system can achieve seamless integration with heterogeneous systems, providing a reliable data foundation for subsequent machine learning modeling, behavioral pattern analysis, and risk assessment.

[0038] Furthermore, this step significantly improves the efficiency and security of data collection, ensuring the integrity and consistency of medical data and providing high-quality input data for the fraud detection model. Simultaneously, through a real-time data interaction mechanism, the system can quickly respond to new claims, gaining a time window for subsequent intelligent analysis and risk assessment, thereby improving the overall accuracy and response speed of fraud detection.

[0039] S2 establishes a normal medical consumption behavior model based on machine learning algorithms, and uses deep learning image recognition algorithms to verify the authenticity of invoices based on their printing features. At the same time, it uses natural language processing technology to analyze the semantic and logical consistency between diagnostic descriptions and medication and treatment methods.

[0040] Specifically, this step involves constructing a normal medical consumption behavior model based on machine learning algorithms, and combining deep learning image recognition and natural language processing technologies to achieve a comprehensive analysis of the authenticity of medical invoices and the consistency between diagnosis and treatment logic. This step integrates behavioral modeling, image feature extraction, and semantic reasoning, exhibiting a high degree of intelligence and automation.

[0041] Furthermore, the system first models historical normal medical consumption data using machine learning algorithms (such as decision trees, random forests, and support vector machines). Specifically, the system extracts key features from structured data collected from medical institutions, pharmacies, and medical insurance departments, including frequency of visits, types and dosages of drugs, and distribution of treatment methods and costs. Through feature engineering, the raw data is transformed into vector representations that can be used for model training, for example, using TF-IDF, One-Hot encoding, or embedded feature extraction methods. During model training, a cross-entropy loss function is employed. Error calculations are performed between the predicted results and the true labels, and the model parameters are optimized using the backpropagation algorithm. After training, the model can perform behavioral pattern matching on new claim applications to determine whether they deviate from normal medical consumption behavior.

[0042] Furthermore, the system employs an image recognition algorithm based on Convolutional Neural Networks (CNNs) to authenticate invoices. CNNs extract printing features from invoice images, such as font format, QR code structure, and anti-counterfeiting watermarks, through multiple convolutional and pooling operations. The output feature map of the convolutional layers is calculated using a formula... Perform calculations, where and These represent the height and width of the convolutional kernel, respectively. The pooling layer uses the max pooling formula. To reduce feature dimensionality and improve model robustness, the system uses a large dataset of labeled invoice images during training and employs data augmentation techniques to enhance the model's generalization ability.

[0043] Furthermore, the system analyzes the semantic and logical consistency between diagnostic descriptions and medication / treatment methods. After word segmentation and embedding, the text data is input into a pre-trained semantic analysis model for contextual modeling. Combined with a medical knowledge graph, the system can reason about the semantic relationships between diagnostic terms and treatment plans, identifying logical contradictions such as "prescribing expensive anti-cancer drugs despite a diagnosis of a mild cold." This process is completed collaboratively by a semantic matching algorithm and a logical rule engine, ensuring the analysis results possess medical rationality and legal compliance.

[0044] Furthermore, this step is widely applied in the automated review process of insurance companies for claims applications. The system can process electronic invoices and text data from different medical institutions in real time, identifying forged invoices and fabricated medical conditions, thereby effectively reducing the burden of manual review and improving the efficiency of fraud detection. In terms of technical effectiveness, this step significantly improves the accuracy and recall rate of fraud detection, reduces the false positive rate, and provides a reliable basis for subsequent risk assessment and handling.

[0045] Furthermore, S2 includes: S21 uses the LeNet-5 improved convolutional neural network structure, where the convolutional layer C1 contains six 5×5 convolutional kernels and the pooling layer S2 uses a 2×2 max pooling window.

[0046] Specifically, in the intelligent recognition module of this invention, an improved LeNet-5 convolutional neural network structure is used for image recognition to verify the authenticity of medical invoices. This structure is adaptively optimized based on the traditional LeNet-5 to improve the recognition accuracy and robustness of medical invoice images. Specifically, the convolutional layer C1 contains six 5×5 convolutional kernels, each used to extract local features from the image, such as anti-counterfeiting watermarks, QR codes, font styles, and other key visual information in the invoice. The convolution operation follows the standard two-dimensional convolution formula:

[0047] in, For the input image, For convolution kernel, and These are the height and width of the convolution kernel, respectively. To output feature map at position The response value. In this embodiment, The number of convolutional kernels is 6, and the output feature map size is 28×28 (assuming the input image is 32×32).

[0048] Furthermore, in pooling layer S2, a 2×2 max-pooling window with a stride of 2 is used to downsample the six 28×28 feature maps output from layer C1 to reduce the spatial dimensionality and enhance the model's adaptability to translation invariance. The mathematical expression of the max-pooling operation is:

[0049] in, To determine the pooling window size, Step size, For the input feature map, This is the output feature map after pooling. The size of the pooled feature map is 14×14, which retains key features while reducing the computational complexity of subsequent operations.

[0050] Furthermore, this step, through a combination of convolutional and pooling layers, enables the system to effectively extract structured features from invoice images and perform classification and judgment through subsequent fully connected layers. In practical applications, this model can be deployed on cloud servers or edge computing devices to process batches of uploaded invoice images in real time, identify features of counterfeit invoices, and thus assist in the determination of fraudulent activities. Its technical value lies in significantly improving the accuracy and processing efficiency of image recognition, providing reliable technical support for the automated detection of medical insurance fraud.

[0051] S22: Construct a knowledge graph containing over 1.2 million medical entity relationships, and verify the logical consistency between diagnostic descriptions and medication and treatment methods through semantic matching algorithms.

[0052] Specifically, in the medical insurance fraud prevention system, a knowledge graph containing more than 1.2 million medical entity relationships is constructed, and the logical consistency between the diagnostic description and the medication and treatment methods is verified through semantic matching algorithms. This step is based on natural language processing (NLP) and medical ontology modeling technology, aiming to extract medical entities and their relationships from unstructured or semi-structured medical texts, and use the knowledge graph for semantic reasoning to determine whether medical behavior conforms to medical logic.

[0053] Furthermore, the system first extracts entities such as diseases, medications, treatments, symptoms, and examinations from diagnostic records, prescription information, and treatment instructions using a medical entity recognition (NER) model. Subsequently, it employs dependency parsing and relation extraction models (such as BERT-based sequence labeling models) to identify semantic relationships between entities, such as "disease-medication," "disease-treatment," and "examination-diagnosis." The constructed knowledge graph is stored in the form of triples, i.e. ,in and For entities, It represents the relationships between entities. The map contains over 1.2 million entries, covering authoritative medical standards such as ICD-10 disease codes, ATC drug classifications, and treatment guidelines.

[0054] Furthermore, the system employs a semantic similarity calculation method based on pre-trained language models (such as DeepSeek, Qwen, and GLM) to match diagnostic descriptions with actual medications and treatments. Specifically, the diagnostic text and the drug / treatment text are encoded into vector representations respectively. and Logical consistency is assessed using cosine similarity. If the similarity is below a preset threshold, further fraud risk assessment is triggered.

[0055] Furthermore, this step is widely used to review claims submitted by policyholders, especially in cases involving complex diseases, multiple drug combinations, or unconventional treatments, effectively identifying fabricated medical conditions or irrational medication use. Through the semantic reasoning capabilities of knowledge graphs, the system can identify logical contradictions such as "a diabetic patient prescribing anti-allergy medication," thereby improving the intelligence level of fraud detection.

[0056] Furthermore, this step significantly enhances the system's ability to judge the rationality of medical behaviors, compensates for the shortcomings of traditional rule engines in semantic understanding, improves the accuracy and recall rate of fraud identification, and provides a solid data foundation for subsequent risk assessment and early warning mechanisms.

[0057] S3. Based on the normal medical consumption behavior model, the results of invoice authenticity identification, and the results of semantic logic analysis, risk features are extracted and a dynamic risk assessment model is used to calculate the probability value of fraud risk, and the risk is divided into different levels.

[0058] Specifically, in the risk assessment module, the system extracts key risk features based on a normal medical consumption behavior model, invoice authenticity verification results, and semantic logic analysis results. It then uses a dynamic risk assessment model to calculate the probability value of fraud risk, ultimately classifying the risk into different levels. This step technically achieves the fusion analysis of multi-source heterogeneous data and the dynamic optimization of machine learning models.

[0059] Furthermore, the system first obtains cleaned and modeled medical behavior data from the data analysis module, including visit frequency, correlation between medications and diagnoses, and distribution of treatment costs. Simultaneously, the intelligent recognition module provides invoice image recognition results (such as authenticity verification and anti-counterfeiting feature matching) and semantic logic analysis results output by the natural language processing module (such as semantic matching degree between diagnosis and medication, and logical consistency score). The system maps this multi-dimensional data into risk feature vectors. Each of them A numerical representation of a specific risk characteristic, for example This indicates the abnormal frequency of medical treatment locations. The credibility score of the invoice image is indicated. This indicates the semantic matching degree between diagnosis and medication.

[0060] Furthermore, the dynamic risk assessment model employs probabilistic models such as logistic regression or Naive Bayes to weight the risk feature vector and output a fraud risk probability value. The calculation formula is as follows:

[0061] in, For the model weight vector, The bias term is used for training and optimization using historical labeled data. Model parameters are typically updated iteratively using the Adam optimizer, with a learning rate set to... Regularization term This is to prevent overfitting.

[0062] Furthermore, the system sets tiered thresholds based on the probability value of fraud risk. For example, when... When, it is judged as low risk; when When, it is judged as medium risk; when When a risk level is reached, it is classified as high-risk. This risk classification mechanism complies with the standard requirements for risk classification in the ISO / IEC 27001 information security management system.

[0063] Furthermore, this step is widely used for initial risk screening in the claims review process. For example, when processing claims in batches, the system can automatically mark high-risk claims and trigger an early warning mechanism, diverting medium- and low-risk claims to automated processing channels, thereby significantly improving review efficiency. In addition, this module can be combined with medical knowledge graphs to weight the semantic logic analysis results, further enhancing the accuracy of risk assessment.

[0064] Furthermore, this step, by quantifying risk characteristics and introducing a dynamic assessment model, enables real-time and accurate identification of fraudulent behavior, providing a scientific basis for subsequent automated processing and manual review, and effectively reducing the insurance company's payout risk and operating costs.

[0065] S4 triggers an early warning mechanism and suspends the claims process for claims with medium to high risk levels. At the same time, it sends an early warning signal to the insurance company's review personnel for investigation and verification. Based on the investigation results, it either adds the company to the fraud blacklist or restores the normal claims process.

[0066] Specifically, in some implementations, the early warning and handling module of this invention triggers an early warning mechanism and suspends the claims process for medium- to high-risk claims based on multi-dimensional risk assessment results. Simultaneously, it sends an early warning signal to the insurance company's review personnel for investigation and verification, and, based on the investigation results, either adds the claim to a fraud blacklist or resumes the normal claims process. This step is technically implemented based on risk threshold settings, automated process control, and a human-machine collaboration mechanism.

[0067] Furthermore, the system first considers the risk probability value output by the risk assessment module. Claims are categorized into three risk levels: low, medium, and high. The criteria for determining medium and high risk are as follows: if... If so, a medium-risk warning will be triggered; if This triggers a high-risk warning. Threshold and It can be dynamically adjusted based on the distribution of historical fraud data, and is usually set to... , This is to ensure that the system achieves a balance between false alarm rate and false negative rate.

[0068] Furthermore, when the system determines that a claim application is of medium to high risk, it will automatically activate the early warning mechanism, sending an early warning signal to the reviewers via email, SMS, or in-system push notifications. The early warning signal contains key information such as the policyholder's basic information, description of abnormal behavior, risk score, and suspicious data fragments (such as invoice image recognition results, inconsistencies between diagnosis and medication logic), facilitating the reviewers to quickly locate the problem.

[0069] Furthermore, the system suspends the claims process through a workflow control module. Specifically, this involves marking the claims application status as "pending investigation" in the database and locking the relevant payment interfaces to prevent erroneous payments before confirmation. After the reviewers complete the investigation, the system takes further action based on the investigation results: if fraud is confirmed, the policyholder's information is added to a fraud blacklist, claims are refused, and regulatory reporting can be triggered; if fraud is not confirmed, the system automatically restores the application status to "normal" and expedites the subsequent review process, improving processing efficiency.

[0070] Furthermore, this step is widely applicable to high-concurrency medical insurance claims review scenarios, especially in complex situations involving out-of-town medical treatment, high drug costs, and frequent medical visits, effectively preventing the spread of fraud. Through automated early warning and process control, the system significantly reduces the delay of manual intervention and improves the timeliness of fraud response.

[0071] Furthermore, this step enables rapid interception and precise handling of potential fraudulent activities, effectively reducing insurance companies' payout losses and operational risks. Simultaneously, the establishment of a blacklist mechanism provides data support for subsequent fraud identification, further enhancing the system's intelligence and continuous optimization capabilities.

[0072] Furthermore, S4 includes: S41 verifies the policyholder's identity through two methods: telephone interviews and on-site investigations. The telephone interviews use a pre-set script template for voice recognition interaction.

[0073] Specifically, in some implementations, this invention verifies the policyholder's identity through both telephone interviews and on-site investigations. The telephone interviews utilize pre-set script templates for voice recognition interaction, enabling automated verification of the policyholder's identity information and behavioral consistency analysis. This step is achieved through the collaborative work of voice recognition, natural language processing (NLP), and a pre-set logical verification mechanism.

[0074] Furthermore, the telephone access module guides the policyholder to respond via voice using preset script templates. The system employs a deep learning-based automatic speech recognition model to convert speech signals into text information. The input to the speech recognition model is audio data in 16kHz sampling rate, single-channel, 16-bit PCM format, and the output is a text sequence. The recognition accuracy (WER) can reach over 92% on a standard corpus.

[0075] Furthermore, the dialogue templates are dynamically generated based on the insured's historical medical records to ensure that the questions are targeted. For example, if the insured recently received treatment at a certain hospital, the system will prioritize asking about the details of that visit. The identification results are semantically parsed using an NLP module to extract key entities (such as hospital name, visit time, and medication name) and compared with the medical records stored in the system. If a significant discrepancy is identified between the insured's answer and the historical records (such as conflicting visit times or inconsistent hospital names), a further investigation process is triggered.

[0076] Furthermore, the on-site investigation involves auditors visiting the medical institution or pharmacy mentioned by the policyholder, as prompted by the system, to verify the authenticity of their medical records, prescription information, and invoices. This step, combined with the voice recognition results from the telephone interview, forms a dual verification mechanism, significantly improving the accuracy of identity verification and the reliability of fraud detection.

[0077] Furthermore, this technical solution can be deployed in insurance company claims review systems as an auxiliary investigation tool for medium- to high-risk cases. Its effectiveness lies in achieving efficient and automated verification of policyholder identity information through the combination of structured dialogue and voice recognition technology. This reduces manual investigation costs, improves the response speed and accuracy of fraud detection, and effectively curbs common insurance fraud behaviors such as fabricated medical visits and forged invoices.

[0078] S42 stipulates that multi-dimensional measures will be taken against policyholders confirmed to have committed fraud, including adding the policyholder's information to the fraud blacklist, freezing their medical insurance accounts, and simultaneously submitting electronic evidence packages to the regulatory authorities.

[0079] Specifically, after confirming that the policyholder has committed fraud, the system will take multi-dimensional measures, including adding the policyholder's information to the fraud blacklist, freezing their medical insurance account, and simultaneously submitting an electronic evidence package to the regulatory authorities.

[0080] Furthermore, the entry of fraud blacklists is based on a unified identity verification system, typically using a combination of fields such as the insured's ID number, medical insurance card number, and name for uniqueness verification to ensure the accuracy and immutability of the blacklist data. The blacklist database can be deployed in a distributed system, supporting high-concurrency writes and real-time queries, and complies with the ISO / IEC 27001 information security management system standard. Freezing medical insurance accounts is achieved through API interfaces or database connections with the medical insurance department. The system sends freeze commands to the medical insurance account management platform, typically using HTTPS for encrypted communication and adhering to the HL7FHIR standard for data formatting to ensure the compliance and traceability of the commands.

[0081] Furthermore, the generation and submission of electronic evidence packages involve three key stages: data packaging, encryption, and transmission. Evidence packages typically contain structured and unstructured data such as policyholder identity information, images of forged invoices, abnormal medical records, and system identification logs, packaged in ZIP or PDF / A format. To ensure data integrity, the system uses the SHA-256 algorithm to perform hash verification on the evidence package and signs the data using asymmetric encryption (such as RSA-2048) to ensure the legal validity of the evidence during transmission. Evidence packages are submitted through a pre-defined regulatory interface, and the transmission protocol supports HTTPS or SFTP, meeting the compliance requirements of the Cybersecurity Law and the Data Security Law regarding the transmission of sensitive data.

[0082] Furthermore, this step is typically triggered automatically by the system backend. After confirming fraud, auditors execute the final disposition through an access control interface. The system supports multi-level approval processes to ensure the legality and auditability of the disposition. This disposition mechanism not only effectively curbs the recurrence of fraudulent activities but also provides regulatory authorities with structured and verifiable electronic evidence, improving regulatory efficiency and the credibility of enforcement evidence.

[0083] The medical insurance fraud detection method of this invention can identify medical insurance fraud in real time and accurately, improve the accuracy of fraud detection and the efficiency of claims processing, and reduce the economic and reputational risks of insurance companies.

[0084] Example 2 This invention proposes a medical insurance fraud prevention system that can monitor and identify medical insurance fraud in real time and accurately, effectively preventing fraudulent activities and maintaining the normal order of the insurance market and the legitimate rights and interests of all parties.

[0085] In one embodiment of the present invention, the data acquisition module includes: interfacing with relevant systems such as medical institutions, medical insurance departments, and pharmacies via API interfaces, database connections, or data exchange platforms. To ensure the stability and security of the interfacing systems, security measures such as encrypted transmission and access control are employed to protect privacy and integrity during data transmission. Data acquisition content includes: obtaining medical information such as the insured's diagnostic records, examination reports, prescriptions, and treatment cost details from medical institutions; obtaining drug purchase records from pharmacies, including drug names, quantities, and prices; and obtaining the insured's medical insurance payment records and reimbursement records from medical insurance departments. Simultaneously, electronic invoice information, such as invoice codes, numbers, issuance dates, and amounts, is collected to ensure the accuracy and traceability of invoice information. Data integration involves integrating the collected medical information, invoice information, basic information of the insured, and insurance contract terms. A unified data format and standard are established to ensure data consistency and comparability. The integrated data is stored in a database to provide a foundation for subsequent data analysis and intelligent identification.

[0086] Furthermore, the data analysis module includes: employing data cleaning algorithms to remove invalid and erroneous data, such as duplicate records, missing values, and outliers. The cleaned data is then validated and verified to ensure its accuracy and reliability. Model building: Based on machine learning algorithms, such as decision trees, random forests, and support vector machines, a model of normal medical consumption behavior is established. Historical claims data is used to train the model, and its parameters are adjusted to accurately reflect the characteristics of normal medical consumption behavior. The model is then validated and tested to evaluate its accuracy and generalization ability. Behavioral pattern analysis: Clustering analysis algorithms are used to identify abnormally frequent medical visits and time intervals, such as multiple visits to different locations within a short period. Association rule mining algorithms are used to discover unreasonable combinations of drugs and diagnoses, such as a drug having no or low association with a certain disease. The analysis results are compared with the normal medical consumption behavior model to determine whether the insured's medical behavior is abnormal.

[0087] Furthermore, the intelligent recognition module includes: employing deep learning-based image recognition algorithms, such as convolutional neural networks (CNNs), to authenticate medical invoices. It trains an image recognition model to accurately identify printing features, font formats, QR codes, and anti-counterfeiting marks on invoices. Invoices are processed in batches to improve recognition efficiency. Natural Language Processing (NLP): NLP techniques, such as word embedding and semantic analysis, are used to analyze the logical consistency between diagnostic descriptions and medication / treatment methods. A medical knowledge graph is established, containing medical concepts such as diseases, drugs, and treatments, and their relationships. The knowledge graph is used to perform semantic matching and logical reasoning on diagnostic descriptions and medication / treatment methods to determine whether there are fabricated medical conditions or unreasonable medical practices.

[0088] Furthermore, the risk assessment module includes: extracting risk characteristics based on data analysis and intelligent identification results, such as the degree of abnormality in medical behavior, the authenticity of invoices, and the rationality of the illness and treatment. Risk assessment algorithms, such as logistic regression and Naive Bayes, are used to calculate the fraud risk probability value for each claim. Risk level classification: Based on the risk probability value, risks are classified into low, medium, and high levels. Risk thresholds are set, and claims exceeding the thresholds are alerted and processed.

[0089] Furthermore, the early warning and handling module includes: Early Warning Mechanism: An early warning system is established to automatically issue early warning signals for medium- to high-risk claims. These signals can be sent to insurance company reviewers via email, SMS, system notifications, etc. Suspension of Claims Process: After an early warning is issued, the claims process is automatically suspended to prevent potential fraud. This suspension can be controlled through the system, such as locking the claim application or stopping payment. Investigation and Verification: Based on the early warning information, insurance company reviewers conduct further investigations and verifications of the claims application. These investigations and verifications can be conducted through telephone interviews, on-site investigations, and document review. Handling Measures: If fraud is confirmed, the policyholder's information is added to a fraud blacklist, claims are refused, and legal action is taken. For serious cases of fraud, the matter is reported to regulatory authorities to assist in the investigation. For low-risk claims, the normal review process is expedited to improve claims efficiency.

[0090] In one embodiment of the present invention, the deep learning-based image recognition algorithm includes: In a medical insurance fraud prevention system, the deep learning-based image recognition algorithm is mainly used for the authentication of medical invoices. Here, a convolutional neural network (CNN) is used as the core algorithm. CNN is a deep learning model specifically designed for processing data with a grid structure (such as images). Convolutional layer formula: The convolutional layer is the core component of a CNN, responsible for extracting image features. For an input image I and a convolutional kernel K, each element Oi,j of the output feature map O of the convolution operation can be calculated using the following formula:

[0091] Where h and w are the height and width of the convolution kernel, respectively, Ii+m,j+n is the pixel value of the input image at position (i+m,j+n), and Km,n is the weight value of the convolution kernel at position (m,n).

[0092] Activation function: A convolutional layer is usually followed by an activation function, such as ReLU (Rectified Linear Unit), whose formula is:

[0093] The ReLU function sets negative values ​​to zero and positive values ​​to remain unchanged, which helps solve the vanishing gradient problem. Pooling layer formula: Pooling layers are used to reduce the spatial resolution of feature maps and reduce computational cost. Commonly used pooling operations include max pooling and average pooling. The formula for max pooling is:

[0094] Where k is the size of the pooling window, s is the stride, and O is the input feature map.

[0095] In one embodiment of the present invention, training the image recognition model includes: Model architecture: adopting a classic CNN architecture, such as LeNet-5, AlexNet, or the more advanced ResNet. Taking LeNet-5 as an example: Input layer: receiving image data of medical invoices, usually grayscale or RGB images. Convolutional layer C1: 6 x 5 convolutional kernels, outputting 6 x 28 feature maps. Pooling layer S2: 2 x 2 max pooling window, outputting 6 x 14 feature maps. Convolutional layer C3: 16 x 5 convolutional kernels, outputting 16 x 10 feature maps. Pooling layer S4: 2 x 2 max pooling window, outputting 16 x 5 feature maps. Fully connected layer F5: 120 neurons. Fully connected layer F6: 84 neurons. Output layer: setting the appropriate number of neurons according to the number of categories of authenticity of the invoice (e.g., 2 categories: true, false), using the Softmax activation function. Training Process: Data Preparation: Collect and label a large number of real and fake medical invoice images. Data Augmentation: Perform rotation, scaling, and translation operations on the images to increase data diversity. Model Initialization: Randomly initialize model parameters. Forward Propagation: Input image data into the model and calculate the output of each layer. Loss Calculation: Calculate the difference between the predicted result and the true label using the cross-entropy loss function. Backpropagation: Calculate the gradient based on the loss function and update the model parameters using optimization algorithms (such as SGD, Adam). Iterative Training: Repeat the forward propagation, loss calculation, and backpropagation processes until the model converges or reaches the preset number of iterations.

[0096] In one embodiment of the present invention, the data processing process using natural language processing technology includes: Data processing flow: Text preprocessing: Cleaning text data such as diagnostic descriptions and medication records, removing irrelevant characters, punctuation marks, etc. Word segmentation and word embedding: Using a word segmentation tool (such as jieba) to segment the text into words or phrases, and mapping each word to a low-dimensional vector space (word embedding). Semantic analysis: Using a pre-trained language model (DeepSeek, Qwen, GLM) or a custom semantic analysis algorithm, analyzing the logical consistency between the diagnostic description and medication / treatment methods. Medical knowledge graph construction: Constructing a graph containing medical concepts such as diseases, drugs, and treatment methods and their relationships to assist semantic matching and logical reasoning. Logical reasoning and consistency check: Combining the medical knowledge graph, performing semantic matching and logical reasoning on the diagnostic description and medication / treatment methods to determine whether there is a fabricated illness or unreasonable medical behavior.

[0097] In one embodiment of the present invention, system integration and data transmission include: System integration method: API interface: Integration with systems such as medical institutions, medical insurance departments, and pharmacies via RESTful API to achieve real-time data transmission. Database connection: Direct connection to the partner's database to obtain the required data through SQL queries. Data exchange platform: Data extraction, transformation, and loading (ETL) are performed using middleware or a data exchange platform (such as an ETL tool). Data transmission process: Data encryption: Before data transmission, the data is encrypted using SSL / TLS encryption protocol to ensure data security during transmission. Authentication and authorization: Strict authentication and authorization are performed when integrating with the system to ensure that only legitimate systems can access the data. Data transmission: The encrypted data is transmitted to the target system via HTTP / HTTPS protocol. For large data volumes, chunked transmission or streaming transmission can be used. Data reception and decryption: After receiving the data, the target system performs decryption to restore the original data. Data verification: The received data undergoes integrity verification and consistency checks to ensure that the data is not lost or tampered with during transmission.

[0098] In one embodiment of the present invention, the system is trained using historical claims data. The specific process of model training includes: Data preparation: Data collection: Collecting historical claims data, including the insured's medical information, invoice information, reimbursement records, etc. Data cleaning: Removing invalid and erroneous data, such as duplicate records, missing values, outliers, etc. Verifying and validating the cleaned data to ensure its accuracy and reliability. Data labeling: Labeling the historical claims data to distinguish between normal claims and fraudulent claims. Model training process: Feature extraction: Extracting meaningful features from the historical claims data, such as medical frequency, drug costs, diagnosis type, etc. Model selection: Selecting a suitable machine learning or deep learning model according to task requirements, such as decision tree, random forest, support vector machine, or CNN, etc. Model initialization: Randomly initializing model parameters.

[0099] Forward Propagation: Input feature data into the model and calculate the prediction result. Loss Calculation: Calculate the difference between the prediction result and the true label using an appropriate loss function (e.g., cross-entropy loss, mean squared error loss). Backpropagation: Calculate the gradient based on the loss function and update the model parameters using optimization algorithms (e.g., SGD, Adam). Iterative Training: Repeat the forward propagation, loss calculation, and backpropagation process, using a validation set to monitor model performance and prevent overfitting. Stop training when the model's performance on the validation set no longer improves. Model Evaluation and Tuning: Evaluate the model's final performance using a test set and tune the model based on the evaluation results, such as adjusting model parameters, adding features, or changing the model structure.

[0100] In one embodiment of the present invention, system development and integration includes: completing the development of each module to ensure smooth data interaction and process connection between modules; conducting system integration testing to ensure the stability and reliability of the overall system function; system docking and data transmission, docking with partners such as medical institutions and pharmacies to ensure real-time data transmission and updates; establishing a data security mechanism to protect the security and privacy of data transmission. System training and optimization includes: training the system using historical claims data, continuously adjusting model parameters and rule thresholds to improve the accuracy of fraud identification; regularly evaluating and optimizing the system's performance to ensure it can adapt to constantly changing fraud methods. New fraud case data collection and updating includes: continuously collecting new fraud case data to continuously update the system's identification capabilities; establishing a fraud case database to provide data support for the continuous improvement and optimization of the system. Fraud prevention report generation and decision support includes: regularly generating fraud prevention reports for insurance company management and regulatory authorities to summarize fraud trends and prevention effectiveness; providing decision support to insurance companies to help them develop more effective fraud prevention strategies. System maintenance and upgrade includes: establishing a system maintenance mechanism to ensure the normal operation and stability of the system. The system will be upgraded and updated regularly based on business development needs and technological advancements.

[0101] The embodiments of this invention also have the following technical effects: System integration and automated processing: The system connects with multiple systems such as medical institutions, medical insurance departments, and pharmacies through API interfaces and database connections, achieving real-time data collection and integration. This cross-system data integration capability provides a comprehensive and accurate data foundation for subsequent fraud identification. Multi-dimensional data analysis and intelligent identification: The system uses machine learning algorithms (such as decision trees and random forests) to establish a normal medical consumption behavior model, and combines cluster analysis, association rule mining, and other algorithms to conduct multi-dimensional analysis of the insured's medical behavior. Simultaneously, it employs deep learning-based image recognition algorithms and natural language processing technology to intelligently identify the authenticity of medical invoices and the logical consistency between diagnostic descriptions and medication / treatment methods. This multi-dimensional and multi-level data analysis and intelligent identification method improves the accuracy and efficiency of fraud identification. Risk assessment and early warning mechanism: Based on the data analysis and intelligent identification results, the system extracts risk characteristics and uses risk assessment algorithms to calculate the fraud risk probability value of each claim application. Based on the risk probability value, the system classifies risks into different levels and automatically issues early warning signals for medium- and high-risk claim applications. This risk assessment and early warning mechanism enables insurance companies to promptly detect and address potential fraudulent activities. Continuous optimization and decision support: The system also possesses continuous optimization and decision support capabilities. It continuously updates its identification capabilities by collecting new fraud case data; it regularly generates fraud prevention reports for insurance company management and regulatory authorities, summarizing fraud trends and prevention effectiveness; and it provides decision support to insurance companies, helping them develop more effective fraud prevention strategies.

[0102] This invention improves the efficiency and security of data collection, ensuring data accuracy and integrity. It enables precise identification and early warning of medical insurance fraud, improving the accuracy of fraud detection. Through image recognition and natural language processing technologies, it further enhances the intelligence level of fraud detection. The risk assessment module provides insurance companies with a scientific basis for risk assessment, helping to reduce their risks. The early warning and handling module enables rapid response and handling of fraudulent activities, effectively preventing their occurrence.

[0103] Example 3 To achieve the above embodiments, such as Figure 2 As shown, this embodiment also provides a medical insurance fraud prevention device 10, including: The multi-source data real-time acquisition module 100 is used to collect medical data from medical institutions, pharmacies and medical insurance departments in real time through API interface and database connection; The multi-algorithm behavior modeling and feature analysis module 200 is used to establish a normal medical consumption behavior model based on machine learning algorithms, and to verify the authenticity of invoices by combining deep learning image recognition algorithms with printing features. At the same time, it uses natural language processing technology to analyze the semantic and logical consistency between diagnostic descriptions and medication and treatment methods. The dynamic risk assessment and classification module 300 is used to extract risk features and calculate the fraud risk probability value using the dynamic risk assessment model based on the normal medical consumption behavior model, the invoice authenticity identification results and the semantic logic analysis results, and classify the risk into different levels. The early warning triggering and handling execution module 400 is used to trigger the early warning mechanism and suspend the claim process for claims with medium to high risk levels. At the same time, it sends an early warning signal to the insurance company's review personnel for investigation and verification, and executes the fraud blacklist entry or resumes the normal claim process based on the investigation results.

[0104] Furthermore, the multi-source data real-time acquisition module 100 is also used for: SSL / TLS encryption protocols are used to encrypt medical data during transmission, ensuring data security during the transmission process; Real-time data interaction with medical institution systems is achieved through RESTful API interfaces, and medical data is encapsulated in JSON format and its integrity is verified.

[0105] Furthermore, the multi-algorithm behavior modeling and feature analysis module 200 is also used for: The LeNet-5 improved convolutional neural network architecture is used, where the convolutional layer C1 contains six 5×5 convolutional kernels and the pooling layer S2 uses a 2×2 max pooling window. We constructed a knowledge graph containing over 1.2 million medical entity relationships and verified the logical consistency between diagnostic descriptions and medication and treatment methods through semantic matching algorithms.

[0106] Furthermore, the early warning triggering and handling execution module 400 is also used for: The policyholder's identity is verified through two methods: telephone interviews and on-site surveys. The telephone interviews use a pre-set script template for voice recognition interaction. For policyholders confirmed to have committed fraud, multi-dimensional measures will be taken, including adding the policyholder's information to the fraud blacklist, freezing their medical insurance accounts, and simultaneously submitting electronic evidence packages to the regulatory authorities.

[0107] This invention discloses a medical insurance fraud prevention device that can identify medical insurance fraud in real time and accurately, effectively improving the accuracy and efficiency of fraud detection and reducing the economic and reputational risks of insurance companies.

[0108] Example 4 To implement the methods of the above embodiments, the present invention also provides a computer device, such as... Figure 3 As shown, the computer device 600 includes a memory 601 and a processor 602; wherein, the processor 602 reads executable program code stored in the memory 601 to run a program corresponding to the executable program code, so as to implement the various steps of the medical insurance fraud prevention method described above.

[0109] Example 5 To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a medical insurance fraud prevention method as described in the foregoing embodiments.

[0110] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0111] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A method for preventing medical insurance fraud, characterized in that, include: S1 collects medical data from medical institutions, pharmacies, and medical insurance departments in real time through API interfaces and database connections; S2 establishes a normal medical consumption behavior model based on machine learning algorithms, and uses deep learning image recognition algorithms to verify the authenticity of invoices based on their printing features. At the same time, it uses natural language processing technology to analyze the semantic and logical consistency between diagnostic descriptions and medication and treatment methods. S3. Based on the normal medical consumption behavior model, the results of invoice authenticity identification and semantic logic analysis, extract risk features and use a dynamic risk assessment model to calculate the probability value of fraud risk, and classify the risk into different levels. S4 triggers an early warning mechanism and suspends the claims process for claims with medium to high risk levels. At the same time, it sends an early warning signal to the insurance company's review personnel for investigation and verification. Based on the investigation results, it either adds the company to the fraud blacklist or restores the normal claims process.

2. The method as described in claim 1, characterized in that, S1 includes: S11 uses SSL / TLS encryption protocol to encrypt medical data during transmission, ensuring data security during transmission; S12 interacts with medical institution systems in real time via a RESTful API interface, encapsulates medical data in JSON format, and performs integrity verification.

3. The method as described in claim 1, characterized in that, S2 includes: S21 uses the LeNet-5 improved convolutional neural network structure, where the convolutional layer C1 contains 6 5×5 convolutional kernels and the pooling layer S2 uses a 2×2 max pooling window. S22: Construct a knowledge graph containing over 1.2 million medical entity relationships, and verify the logical consistency between diagnostic descriptions and medication and treatment methods through semantic matching algorithms.

4. The method as described in claim 1, characterized in that, The S4 includes: S41 verifies the identity of the policyholder through two methods: telephone interviews and on-site investigations. The telephone interviews use a preset script template for voice recognition interaction. S42 stipulates that multi-dimensional measures will be taken against policyholders confirmed to have committed fraud, including adding the policyholder's information to the fraud blacklist, freezing their medical insurance accounts, and simultaneously submitting electronic evidence packages to the regulatory authorities.

5. A medical insurance fraud prevention device, characterized in that, include: The multi-source data real-time acquisition module is used to collect medical data from medical institutions, pharmacies, and medical insurance departments in real time through API interfaces and database connections; The multi-algorithm behavior modeling and feature analysis module is used to build a normal medical consumption behavior model based on machine learning algorithms, and to verify the authenticity of invoices by combining deep learning image recognition algorithms with printing features. At the same time, natural language processing technology is used to analyze the semantic and logical consistency between diagnostic descriptions and medication and treatment methods. The dynamic risk assessment and classification module is used to extract risk features and calculate the fraud risk probability value using the dynamic risk assessment model based on the normal medical consumption behavior model, the invoice authenticity identification results and the semantic logic analysis results, and classify the risk into different levels. The early warning triggering and handling module is used to trigger the early warning mechanism and suspend the claim process for claims with medium to high risk levels. At the same time, it sends an early warning signal to the insurance company's review personnel for investigation and verification. Based on the investigation results, it executes the fraud blacklist entry or resumes the normal claims process.

6. The apparatus as claimed in claim 5, characterized in that, The multi-source data real-time acquisition module is also used for: SSL / TLS encryption protocols are used to encrypt medical data during transmission, ensuring data security during the transmission process; Real-time data interaction with medical institution systems is achieved through RESTful API interfaces, and medical data is encapsulated in JSON format and its integrity is verified.

7. The apparatus as claimed in claim 5, characterized in that, The multi-algorithm behavior modeling and feature analysis module is also used for: The LeNet-5 improved convolutional neural network architecture is used, where the convolutional layer C1 contains six 5×5 convolutional kernels and the pooling layer S2 uses a 2×2 max pooling window. We constructed a knowledge graph containing over 1.2 million medical entity relationships and verified the logical consistency between diagnostic descriptions and medication and treatment methods through semantic matching algorithms.

8. The apparatus as claimed in claim 5, characterized in that, The early warning triggering and handling execution module is also used for: The policyholder's identity is verified through two methods: telephone interviews and on-site surveys. The telephone interviews use a pre-set script template for voice recognition interaction. For policyholders confirmed to have committed fraud, multi-dimensional measures will be taken, including adding the policyholder's information to the fraud blacklist, freezing their medical insurance accounts, and simultaneously submitting electronic evidence packages to the regulatory authorities.

9. A computer device, characterized in that, Including processor and memory; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement a medical insurance fraud prevention method as described in any one of claims 1-4.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a medical insurance fraud prevention method as described in any one of claims 1-4.