Accident claim settlement method and device, computer equipment and storage medium

By automatically extracting medical diagnostic information from claims and reimbursement documents through image recognition and text embedding models, the problem of inefficiency and inconsistent judgment in traditional insurance claims has been solved, achieving efficient and accurate determination of the cause of the accident.

CN120894147APending Publication Date: 2025-11-04CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510828557.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

In the traditional insurance claims process, document processing relies on manual operation, which leads to inefficiency, inconsistent judgments, and a high risk of errors.

Method used

Image recognition technology is used to automatically extract medical diagnosis information from claims and reimbursement documents, and semantic similarity retrieval is performed through text embedding model and diagnosis information knowledge base to determine the cause of the accident.

Benefits of technology

It improves the efficiency and accuracy of claims processing, reduces human intervention, and ensures consistency and accuracy in judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894147A_ABST
    Figure CN120894147A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence and information security, and relates to an accident claim settlement method, comprising: acquiring a claim reimbursement document uploaded by a user, the claim reimbursement document being image data and including medical diagnosis information; performing image recognition on the claim settlement reimbursement document, and extracting medical diagnosis information; according to a pre-constructed diagnosis information knowledge base, semantic similarity retrieval is carried out on the medical diagnosis information, a plurality of pieces of target historical diagnosis information are obtained, and the diagnosis information knowledge base is constructed by adopting a text embedding model and the historical diagnosis information; and determining a target accident reason according to the plurality of pieces of target historical diagnosis information, and executing a claim settlement processing flow based on the target accident reason. The invention further provides an accident claim settlement device, computer equipment and a storage medium. The method can be applied to a business management program system of financial science and technology, medical diagnosis information in the claim settlement reimbursement document can be automatically extracted, an accident reason can be intelligently judged, and the claim settlement efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and is applied to online financial technology business scenarios, particularly to a method, device, computer equipment, and storage medium for claims settlement. Background Technology

[0002] In the traditional insurance claims process, users need to submit various claim documents, and the insurance company's claims personnel need to manually review these documents, extract key information, and determine whether the cause of the accident meets the compensation requirements of the insurance terms.

[0003] Currently, document processing in the insurance claims process relies primarily on manual operation. Claims personnel need to review paper or electronic documents submitted by users, manually enter various information, and determine the cause of the loss based on experience. This method is not only time-consuming and labor-intensive, but also prone to inefficiency and inconsistent judgments when dealing with a large number of cases.

[0004] Therefore, there is an urgent need for a method that can automatically extract medical diagnostic information from claims documents and intelligently determine the cause of the accident in order to improve the efficiency and accuracy of claims processing and enhance the user experience. Summary of the Invention

[0005] The purpose of this application is to provide a method, apparatus, computer equipment, and storage medium for claims settlement, which aims to automatically extract medical diagnostic information from claims reimbursement documents and intelligently determine the cause of the accident, thereby improving the efficiency and accuracy of claims settlement.

[0006] Firstly, a method for handling claims is provided, which employs the following technical solution:

[0007] Obtain the claim reimbursement documents uploaded by the user. The claim reimbursement documents are image data and contain medical diagnosis information.

[0008] Image recognition is performed on the claim reimbursement documents to extract the medical diagnosis information;

[0009] Based on a pre-built diagnostic information knowledge base, semantic similarity retrieval is performed with the medical diagnostic information to obtain several target historical diagnostic information. The diagnostic information knowledge base is constructed using a text embedding model and historical diagnostic information.

[0010] The cause of the target incident is determined based on the aforementioned historical diagnostic information of several targets;

[0011] The corresponding claims processing procedure will be executed based on the cause of the incident.

[0012] Secondly, an accident claims processing device is provided, which adopts the following technical solution:

[0013] The document acquisition module is used to acquire the claim reimbursement documents uploaded by the user. The claim reimbursement documents are image data and contain medical diagnosis information.

[0014] The image recognition module is used to perform image recognition on the claim reimbursement documents and extract the medical diagnosis information;

[0015] The information retrieval module is used to perform semantic similarity retrieval with the medical diagnostic information based on a pre-built diagnostic information knowledge base, and obtain several target historical diagnostic information. The diagnostic information knowledge base is constructed using a text embedding model and historical diagnostic information.

[0016] The incident determination module is used to determine the cause of the incident based on the historical diagnostic information of the aforementioned targets;

[0017] The claims processing module is used to execute the corresponding claims processing procedures based on the target cause of the accident.

[0018] Thirdly, a computer device is provided, including a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the above-described claims settlement method.

[0019] Fourthly, a computer-readable storage medium is provided, on which computer-readable instructions are stored, which, when executed by a processor, implement the steps of the above-described claims settlement method.

[0020] Compared with the prior art, the embodiments of this application have the following main advantages:

[0021] This application employs image recognition technology to automatically extract medical diagnostic information from claim reimbursement documents, avoiding the tediousness and errors of manual data entry and improving the efficiency and accuracy of information extraction. By utilizing a diagnostic information knowledge base constructed using a text embedding model and historical diagnostic information, semantic similarity retrieval is performed with the extracted medical diagnostic information to quickly find target historical diagnostic information similar to the current medical diagnostic information. Furthermore, semantic similarity retrieval allows for a more accurate understanding of the semantics of medical diagnostic information, reducing errors in determining the cause of the accident due to misunderstandings. By determining the target cause of the accident based on several retrieved target historical diagnostic information, the subjectivity and uncertainty of manual judgment are reduced, improving the consistency of claims processing. Based on this application's solution, automated technology is used to perform image recognition, information extraction, and semantic similarity retrieval on user-uploaded claim reimbursement documents, ultimately determining the target cause of the accident, significantly reducing manual intervention and improving claims processing efficiency and accuracy. Attached Figure Description

[0022] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;

[0024] Figure 2 This is a flowchart illustrating one embodiment of the claims settlement method according to this application;

[0025] Figure 3 This is a flowchart of an embodiment prior to step S201 of this application;

[0026] Figure 4 This is a flowchart illustrating an embodiment of step S203 of this application;

[0027] Figure 5 This is a schematic diagram of the structure of one embodiment of the claims settlement device according to this application;

[0028] Figure 6 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0030] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0031] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0032] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables.

[0033] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0034] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers.

[0035] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0036] It should be noted that the claims settlement method provided in this application is generally executed by a server / terminal device, and correspondingly, the claims settlement device is generally set in the server / terminal device.

[0037] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0038] Continue to refer to Figure 2 A flowchart illustrating an embodiment of the claims settlement method according to this application is shown. The claims settlement method includes the following steps:

[0039] Step S201: Obtain the claim reimbursement documents uploaded by the user. The claim reimbursement documents are image data and contain medical diagnosis information.

[0040] In this embodiment, the claims settlement method operates on electronic devices (e.g., Figure 1 The server / terminal device shown can obtain user-uploaded claim documents via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, Wi-Fi connections, Bluetooth connections, Wi-MAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future-developed wireless connection methods.

[0041] In this embodiment, the user-uploaded claim reimbursement documents are obtained. Users can upload claim reimbursement documents through the insurance company's mobile application, website, or other electronic channels. These documents may include, but are not limited to, medical expense invoices, outpatient medical records, inpatient medical records, examination reports, discharge summaries, and diagnostic certificates, which contain medical diagnostic information. Medical diagnostic information refers to the professional judgment made on a patient's disease, symptoms, or health status through medical examination, testing, or clinical evaluation, and typically includes the disease name, symptoms, surgical or treatment methods, pathological results, etc. During the claims process, medical diagnostic information can be used to clarify whether the cause of the accident meets the insurance liability. Claim reimbursement documents are generally paper documents or electronic documents. To facilitate users in uploading these documents, they are usually converted into image data through methods such as taking photos, scanning, screenshotting, and exporting.

[0042] Step S202: Perform image recognition on the claim reimbursement documents to extract the medical diagnosis information;

[0043] In this embodiment, image recognition technology is used to perform image recognition on the claim reimbursement documents and extract medical diagnosis information from the recognized content. For example, a preset image recognition technology, such as optical character recognition technology, is used to identify and extract text information and image layout information from the claim reimbursement documents; the text information is cleaned and structured to extract key fields (such as diagnosis conclusion, surgery name, drug name, etc.); the information is then combined with the image layout information and key fields to obtain the final medical diagnosis information.

[0044] Step S203: Based on the pre-built diagnostic information knowledge base, perform semantic similarity retrieval with the medical diagnostic information to obtain several target historical diagnostic information. The diagnostic information knowledge base is constructed using a text embedding model and historical diagnostic information.

[0045] In this embodiment, historical diagnostic information can be obtained from the insurance company's historical claims data. This historical diagnostic information includes diagnostic descriptions of various diseases and corresponding classifications of the causes of claims. For example, historical diagnostic information may include diagnoses of diseases such as "acute appendicitis," "type 2 diabetes," and "hypertension," as well as their corresponding causes of claims such as "disease" and "accident."

[0046] Text embedding models can employ natural language processing models such as BERT, Word2Vec, BGE, or GloVe. Taking the BGE model as an example, this model uses deep learning techniques to analyze the semantics and contextual relationships of words in text. When processing historical diagnostic information, the BGE model considers the specific meaning of each word in the medical diagnostic context, as well as the relationships between words.

[0047] For example, for the historical diagnosis information "acute appendicitis," the text embedding model analyzes the semantics of the words "acute" and "appendicitis," as well as the overall meaning formed by their combination. The model converts this semantic information and contextual relationships into a point in a vector space, i.e., a vector representation. This vector representation contains the semantic features of the historical diagnosis information, making semantically similar historical diagnosis information closer together in the vector space.

[0048] In this embodiment, a diagnostic information knowledge base is pre-constructed based on a text embedding model and historical diagnostic information. The construction process can involve cleaning and standardizing the collected historical diagnostic information (e.g., the same terms "myocardial infarction" and "heart attack"); encoding each piece of historical diagnostic information into a vector using a text embedding model and storing it in a vector database to obtain the diagnostic information knowledge base. This vector database can employ a database system that supports storing and retrieving high-dimensional vectors, such as Faiss, Milvus, or Pinecone. These vector databases support efficient semantic similarity retrieval, enabling them to quickly find the vector most similar to the query vector.

[0049] Specifically, the extracted medical diagnostic information is input into a text embedding model, which generates a query vector from the medical diagnostic information. In a pre-built diagnostic information knowledge base, several historical diagnostic information (such as Top-5) with high semantic similarity to the query vector are retrieved as target historical diagnostic information.

[0050] Step S204: Determine the cause of the target incident based on the aforementioned historical diagnostic information of several targets;

[0051] In this embodiment, the cause of the target accident is determined based on several historical diagnostic information of the target. For example, if multiple historical diagnostic information of the target points to the same cause of the accident (such as 80% or 100% of the historical diagnostic information of the target being "traffic accident injury"), then the same cause of the accident is determined as the cause of the target accident.

[0052] Step S205: Execute the corresponding claims processing procedure based on the cause of the target accident.

[0053] In this embodiment, a corresponding claims processing procedure is matched to the target cause of the accident and executed. This claims processing procedure can be pre-set and may include claims amount calculation, automatic or manual review, payment, and user notification processes.

[0054] In this embodiment, image recognition technology is used to automatically extract medical diagnostic information from claim reimbursement documents, avoiding the tediousness and errors of manual data entry and improving the efficiency and accuracy of information extraction. By utilizing a diagnostic information knowledge base constructed using a text embedding model and historical diagnostic information, semantic similarity retrieval is performed with the extracted medical diagnostic information to quickly find target historical diagnostic information similar to the current medical diagnostic information. Furthermore, semantic similarity retrieval allows for a more accurate understanding of the semantics of the medical diagnostic information, reducing errors in determining the cause of the accident due to misunderstandings. By determining the target cause of the accident based on several retrieved target historical diagnostic information, the subjectivity and uncertainty of manual judgment are reduced, improving the consistency of claims processing. Based on this application, by using automated technology to perform image recognition, information extraction, and semantic similarity retrieval on user-uploaded claim reimbursement documents, the target cause of the accident is ultimately determined, significantly reducing manual intervention and improving the efficiency and accuracy of claims processing.

[0055] In some optional implementations of this embodiment, refer to Figure 3 Before step S201, that is, before obtaining the claim reimbursement documents uploaded by the user, the aforementioned electronic device may also perform the following steps:

[0056] Step S101: Obtain several pre-collected historical diagnostic information records;

[0057] Specifically, this involves acquiring several pre-collected historical diagnostic information entries. These entries can be extracted from the insurance company's historical claims database, including diagnostic information from all processed claims over the past few years. This historical diagnostic information covers various types of illnesses, accidental injuries, and other medical conditions, with each entry linked to a specific cause of the claim.

[0058] Step S102: Using a text embedding model, analyze the first semantics of each word in the historical diagnostic information and the first contextual information between words, map the first semantics and the first contextual information to a vector space, and convert them into a first vector representation;

[0059] Specifically, the text embedding model employs the BGE (BAAI General Embedding) model, which is based on the Transformer architecture and can map text into low-dimensional dense vectors. It supports bilingual (Chinese and English) and mixed text, performing well across more than 100 languages. This model features multi-functionality, multilingual support, and multi-granularity, enabling dense retrieval, multi-vector retrieval, and sparse retrieval, supporting inputs of varying granularities from short sentences to long documents. Through pre-training on a large amount of medical text, this BGE model can better understand the semantics of medical terminology and diagnostic descriptions.

[0060] For each historical diagnosis record, the BGE model analyzes the semantics (first semantics) of the medical terms, symptom descriptions, disease names, and other words, as well as the logical relationships and contextual information (first contextual information) between these words. For example, for the historical diagnosis record "right femoral neck fracture caused by a fall from a height," the model understands that "fall from a height" is the cause and "right femoral neck fracture" is the result, and that there is a causal relationship between the two.

[0061] The BGE model maps these initial semantic and contextual information into a multi-dimensional vector space, generating a first vector representation that comprehensively represents the semantic features of the historical diagnostic information. In the vector space, historical diagnostic information with similar semantics (such as different types of fractures) will cluster in nearby regions, while historical diagnostic information with significant semantic differences (such as fractures and colds) will be distributed in more distant regions.

[0062] Step S103: Store the first vector representation in the vector database to construct a diagnostic information knowledge base.

[0063] Specifically, the vector database can be the Milvus vector database. The generated first vector representation is stored in the Milvus vector database to construct a diagnostic information knowledge base, providing a foundation for subsequent semantic similarity retrieval.

[0064] In this embodiment, by employing a text embedding model, the first semantic meaning of each word in historical diagnostic information and the first contextual information between words can be analyzed in depth. This not only captures the meaning of the words themselves but also understands the relationships between words in specific contexts, thus representing the content of the diagnostic information more comprehensively. By mapping the first semantic meaning and the first contextual information to a vector space, a first vector representation is obtained and stored in the knowledge base in vector form. This makes the diagnostic information easier to process and compare, providing convenience for subsequent tasks such as information retrieval, classification, and recommendation.

[0065] In some optional implementations of this embodiment, refer to Figure 4 Step S203 above, which involves performing a semantic similarity search between the medical diagnostic information and the pre-built diagnostic information knowledge base to obtain several target historical diagnostic information entries, may include the following steps:

[0066] Step S2031: Assign weights to the medical diagnostic information according to preset document priorities;

[0067] Specifically, before performing semantic similarity retrieval, the extracted medical diagnostic information is weighted according to a preset document priority. Different types of claim reimbursement documents have different levels of importance in the claim process; for example, a formal hospital diagnosis certificate may be more authoritative than a regular outpatient medical record. Therefore, corresponding weights are assigned to the medical diagnostic information extracted from different claim reimbursement documents based on the document type. For example, a diagnosis certificate might have a weight of 1.5, while a regular outpatient medical record might have a weight of 1.0.

[0068] Step S2032: Using a text embedding model, analyze the second semantics of each word and the second contextual information between words in the weighted medical diagnosis information, map the second semantics and the second contextual information to a vector space, and convert them into a second vector representation;

[0069] Specifically, after weight allocation, a text embedding model is used to process the medical diagnostic information. This text embedding model is the same as the one used when constructing the diagnostic information knowledge base. Using this text embedding model, the semantics (second semantics) of each word in the weighted medical diagnostic information and the contextual relationships between words (second contextual information) are analyzed. Then, these second semantics and second contextual information are mapped into a vector space to obtain a second vector representation.

[0070] For example, for the extracted medical diagnosis information "type 2 diabetes mellitus complicated with diabetic nephropathy (early stage)," the model will analyze the second semantics of words such as "type 2," "diabetes," "complication," "diabetic nephropathy," and "early stage" and the second contextual relationship between them, and generate a second vector representation that can represent the semantic features of the medical diagnosis information.

[0071] Step S2033: Obtain the first vector representation in the diagnostic information knowledge base, and calculate the similarity between the first vector representation and the second vector representation;

[0072] Specifically, the first vector representation of all historical diagnostic information is obtained from the diagnostic information knowledge base, and the similarity between the first vector representation and the second vector representation of the current medical diagnostic information is calculated. Similarity calculation can employ cosine similarity, Euclidean distance, or other vector similarity measures.

[0073] Step S2034: If the similarity is greater than a preset threshold, then the historical diagnostic information corresponding to the first vector representation is determined as the target historical diagnostic information.

[0074] Specifically, if the similarity between the first vector representation of historical diagnostic information and the second vector representation of current medical diagnostic information is greater than a preset threshold (e.g., 0.85), then the historical diagnostic information is identified as the target historical diagnostic information. In this way, the system can find multiple target historical diagnostic information that are semantically similar to the current medical diagnostic information.

[0075] In this embodiment, by assigning weights to medical diagnostic information based on preset document priorities, it is ensured that important or critical diagnostic information receives more attention and consideration in subsequent processing. This weighting method helps focus on the most diagnostically valuable information, improving the efficiency and accuracy of information processing. Through precise weighting and information focusing, in-depth semantic and contextual analysis, efficient vector representation and similarity calculation, and accurate determination of target historical diagnostic information, the intelligence and efficiency of medical diagnosis are enhanced.

[0076] In some optional implementations of this embodiment, step S204, namely determining the cause of the target incident based on the aforementioned historical diagnostic information, may include the following steps:

[0077] Extract the historical causes of accidents from the historical diagnostic information of the several targets respectively, and determine whether the several historical causes of accidents are the same;

[0078] When several historical causes of accidents are the same, the historical cause of accident is determined as the target cause of accident.

[0079] Specifically, when determining the cause of a target claim, the corresponding historical claim cause is first extracted from the historical diagnostic information of each target. For example, if three historical diagnostic records are retrieved, and their corresponding historical claim causes might be "disease," "disease," and "disease," then it is checked whether these historical claim causes are the same. If all historical claim causes are the same (such as "disease" in the example above), then this consistent historical claim cause is directly determined as the target claim cause for the current case.

[0080] In some optional implementations of this embodiment, after the above-described step of determining whether the causes of several historical incidents are the same, the following steps may also be included:

[0081] When several historical causes of accidents are different, a preset first prompt word template is obtained. The first prompt template is used to instruct the language model to deduce the cause of the medical diagnosis based on the correlation between the diagnostic information.

[0082] In this embodiment, if the historical causes of accidents are different, for example, if the historical causes of accidents corresponding to the three retrieved target historical diagnostic information are "disease", "accident" and "disease" respectively, further analysis is needed to determine the most appropriate cause of accident.

[0083] Specifically, a preset first prompt word template is obtained. This first prompt word template is pre-designed and is used to instruct the language model to deduce the most appropriate cause of the accident based on the correlation between diagnostic information. The first prompt word template may be in the form of: "Please analyze the most likely cause of the accident based on the following current medical diagnostic information and historical similar cases. Current diagnosis: {current medical diagnostic information}; Historical case 1: {historical diagnostic information 1}, cause of accident: {historical cause of accident 1}; Historical case 2: {historical diagnostic information 2}, cause of accident: {historical cause of accident 2}; ...".

[0084] By combining the medical diagnostic information, the target historical diagnostic information, and the first prompt word template, a first prompt text is generated;

[0085] The first prompt text is input into the language model, and the cause of the target accident is generated based on the medical diagnosis information and the target's historical diagnosis information.

[0086] Specifically, the current medical diagnosis information, the target historical diagnosis information and the corresponding historical cause of the accident are filled into the first prompt word template to generate a complete first prompt text, and then the first prompt text is input into the language model.

[0087] Specifically, the language model analyzes the similarities and differences between the current diagnosis and various historical cases, taking into account medical knowledge and insurance claim rules, to deduce the most appropriate cause of the claim. For example, if the current diagnosis is "fracture caused by a fall," and historical cases include "fracture caused by a traffic accident" (caused by "accident") and "fracture caused by osteoporosis" (caused by "illness"), the language model will infer, based on the "fall" factor in the current diagnosis, that the cause of the claim is more likely to be "accident" rather than "illness."

[0088] Specifically, the first prompt text is input into the language model, and the language model infers the cause of the accident based on the instructions in the first prompt text, medical diagnosis information, and the target's historical diagnosis information.

[0089] In this embodiment, when several historical accident causes in target historical diagnostic information are the same, these causes are directly identified as the target accident causes, resulting in a concise and efficient outcome. When several historical accident causes are different, a preset first prompt word template is obtained, and the language model is instructed to deduce the accident cause of the medical diagnostic information based on the correlation between the diagnostic information. This approach is more flexible and intelligent, capable of handling complex and changing situations. The application of the language model can handle complex natural language understanding and generation tasks, thereby deriving more accurate and reasonable target accident causes.

[0090] In some optional implementations of this embodiment, step S202, namely, performing image recognition on the claim reimbursement documents to extract the medical diagnosis information, may include the following steps:

[0091] The claim reimbursement document is image-recognized using optical character recognition (OCR) to obtain text information and image layout information.

[0092] Specifically, in the image recognition process, optical character recognition (OCR) technology is first used to convert the text in the claim reimbursement documents (image data) into editable text. OCR technology can be achieved using tools such as Tesseract or ABBYY FineReader. The OCR tool is then used to perform image recognition on the claim reimbursement documents, extracting text information and recognizing the text's position, size, font, and other image layout information within the image.

[0093] For example, when given an image of a medical expense invoice, OCR can identify textual information such as the hospital name, patient information, diagnosis name, treatment items, and cost on the invoice, while also recording the positional relationship of this textual information in the image.

[0094] The text information is concatenated based on the image layout information to obtain concatenated text data;

[0095] Specifically, based on the identified image layout information, the text information is reasonably spliced ​​together. For example, text information in the same area is merged into a paragraph, or text information in the same table is organized according to row and column relationships, thereby obtaining structured spliced ​​text data.

[0096] Obtain a preset second prompt word template, and combine the concatenated text data with the second prompt word template to generate a second prompt text. The second prompt word template is used to instruct the language model to extract key information from the input data.

[0097] Specifically, a pre-designed second prompt template is obtained. This template instructs the language model to extract key medical diagnostic information from the text. The prompt template might look like this: "Please extract the patient's main diagnostic information from the following medical document content, including the disease name, symptom description, and doctor's diagnosis. Medical document content: {concatenated text data}". Then, the concatenated text data is filled into the second prompt template to generate the second prompt text.

[0098] The second prompt text is input into the language model, and the language model is used to extract key information from the concatenated text data as medical diagnostic information.

[0099] Specifically, the language model can be GPT-4, Claude, or other large-scale language models capable of understanding medical terminology and document structure, thereby accurately extracting key medical diagnostic information. The second prompt text is input into the language model, which then extracts key information from the concatenated text data based on the instructions of the second prompt text, serving as the medical diagnostic information.

[0100] For example, from a complex medical record, a language model may extract medical diagnostic information such as "primary diagnosis: type 2 diabetes; complication: diabetic nephropathy (early stage); symptoms: polydipsia, polyuria, weight loss".

[0101] In this embodiment, optical character recognition (OCR) is used to perform image recognition on claim reimbursement documents, enabling rapid and accurate extraction of text information from the image while simultaneously acquiring image layout information. By concatenating the extracted text information according to the image layout information, concatenated text data is obtained. This effectively combines the position and relationship of the text within the original image, helping to maintain the integrity and coherence of the text. By inputting the second prompt text into a language model, the powerful natural language processing capabilities of the language model are utilized to extract key information from the concatenated text data. This information serves as medical diagnostic information, capable of handling complex text structures and semantic relationships, accurately identifying key information related to medical diagnosis, and achieving intelligent extraction of key information, further enhancing the system's intelligence level.

[0102] Further reference Figure 5 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of an accident claims processing device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0103] like Figure 5 As shown, the claims processing device 400 described in this embodiment includes: a document acquisition module 401, an image recognition module 402, an information retrieval module 403, an accident determination module 404, and a claims processing module 405. Wherein:

[0104] The document acquisition module 401 is used to acquire the claim reimbursement documents uploaded by the user. The claim reimbursement documents are image data and contain medical diagnosis information.

[0105] Image recognition module 402 is used to perform image recognition on the claim reimbursement documents and extract the medical diagnosis information;

[0106] The information retrieval module 403 is used to perform semantic similarity retrieval with the medical diagnostic information based on a pre-built diagnostic information knowledge base to obtain several target historical diagnostic information. The diagnostic information knowledge base is constructed using a text embedding model and historical diagnostic information.

[0107] The incident determination module 404 is used to determine the cause of the incident based on the aforementioned historical diagnostic information of several targets;

[0108] The claims processing module 405 is used to execute the corresponding claims processing procedure based on the target cause of the accident.

[0109] In this embodiment, the user-uploaded claim reimbursement documents are obtained. Users can upload claim reimbursement documents through the insurance company's mobile application, website, or other electronic channels. These documents may include, but are not limited to, medical expense invoices, outpatient medical records, inpatient medical records, examination reports, discharge summaries, and diagnostic certificates, which contain medical diagnostic information. Medical diagnostic information refers to the professional judgment made on a patient's disease, symptoms, or health status through medical examination, testing, or clinical evaluation, and typically includes the disease name, symptoms, surgical or treatment methods, pathological results, etc. During the claims process, medical diagnostic information can be used to clarify whether the cause of the accident meets the insurance liability. Claim reimbursement documents are generally paper documents or electronic documents. To facilitate users in uploading these documents, they are usually converted into image data through methods such as taking photos, scanning, screenshotting, and exporting.

[0110] Specifically, image recognition technology is used to perform image recognition on claim reimbursement documents and extract medical diagnosis information from the recognized content. For example, using preset image recognition technology, such as optical character recognition (OCR), text information and image layout information in claim reimbursement documents are identified and extracted; the text information is cleaned and structured to extract key fields (such as diagnosis conclusion, surgery name, drug name, etc.); and the information is combined with image layout information and key fields to obtain the final medical diagnosis information.

[0111] Specifically, historical diagnostic information can be obtained from the insurance company's historical claims data. This information includes descriptions of various diseases and corresponding classifications of the causes of claims. For example, historical diagnostic information may include diagnoses of diseases such as "acute appendicitis," "type 2 diabetes," and "hypertension," as well as their corresponding causes of claims such as "disease" and "accident."

[0112] Text embedding models can employ natural language processing models such as BERT, Word2Vec, BGE, or GloVe. Taking the BGE model as an example, this model uses deep learning techniques to analyze the semantics and contextual relationships of words in text. When processing historical diagnostic information, the BGE model considers the specific meaning of each word in the medical diagnostic context, as well as the relationships between words.

[0113] For example, for the historical diagnosis information "acute appendicitis," the text embedding model analyzes the semantics of the words "acute" and "appendicitis," as well as the overall meaning formed by their combination. The model converts this semantic information and contextual relationships into a point in a vector space, i.e., a vector representation. This vector representation contains the semantic features of the historical diagnosis information, making semantically similar historical diagnosis information closer together in the vector space.

[0114] In this embodiment, a diagnostic information knowledge base is pre-constructed based on a text embedding model and historical diagnostic information. The construction process can involve cleaning and standardizing the collected historical diagnostic information (e.g., the same terms "myocardial infarction" and "heart attack"); encoding each piece of historical diagnostic information into a vector using a text embedding model and storing it in a vector database to obtain the diagnostic information knowledge base. This vector database can employ a database system that supports storing and retrieving high-dimensional vectors, such as Faiss, Milvus, or Pinecone. These vector databases support efficient semantic similarity retrieval, enabling them to quickly find the vector most similar to the query vector.

[0115] Specifically, the extracted medical diagnostic information is input into a text embedding model, which generates a query vector from the medical diagnostic information. In a pre-built diagnostic information knowledge base, several historical diagnostic information (such as Top-5) with high semantic similarity to the query vector are retrieved as target historical diagnostic information.

[0116] In this embodiment, the cause of the target accident is determined based on several historical diagnostic information of the target. For example, if multiple historical diagnostic information of the target points to the same cause of the accident (such as 80% or 100% of the historical diagnostic information of the target being "traffic accident injury"), then the same cause of the accident is determined as the cause of the target accident.

[0117] In this embodiment, a corresponding claims processing procedure is matched to the target cause of the accident and executed. This claims processing procedure can be pre-set and may include claims amount calculation, automatic or manual review, payment, and user notification processes.

[0118] The claims processing device 400 of this application employs image recognition technology to automatically extract medical diagnostic information from claims reimbursement documents, avoiding the tediousness and errors of manual data entry and improving the efficiency and accuracy of information extraction. By utilizing a diagnostic information knowledge base constructed using a text embedding model and historical diagnostic information, semantic similarity retrieval is performed between the extracted medical diagnostic information and the data. This allows for the rapid identification of target historical diagnostic information similar to the current medical diagnostic information. Furthermore, semantic similarity retrieval enables a more accurate understanding of the semantics of the medical diagnostic information, reducing errors in determining the cause of the accident due to misunderstandings. By determining the target cause of the accident based on several retrieved target historical diagnostic information, the subjectivity and uncertainty of manual judgment are reduced, improving the consistency of claims processing. Based on this application's solution, automated technology is used to perform image recognition, information extraction, and semantic similarity retrieval on user-uploaded claims reimbursement documents, ultimately determining the target cause of the accident, significantly reducing manual intervention and improving claims processing efficiency and accuracy.

[0119] In some optional implementations of this embodiment, the claims settlement device 400 further includes an information acquisition module, a vector conversion module, and a knowledge base construction module, wherein:

[0120] The information acquisition module is used to acquire several pre-collected historical diagnostic information.

[0121] The vector transformation module is used to analyze the first semantics of each word and the first contextual information between words in the historical diagnostic information using a text embedding model, and to map the first semantics and the first contextual information to a vector space to obtain a first vector representation;

[0122] The knowledge base construction module is used to store the first vector representation into a vector database to construct a diagnostic information knowledge base.

[0123] In this embodiment, several pre-collected historical diagnostic information entries are acquired. This historical diagnostic information can be extracted from the insurance company's historical claims database, including diagnostic information from all processed claims cases over the past few years. This historical diagnostic information covers various disease types, accidental injuries, and other medical conditions, and each entry is associated with a specific cause of the claim.

[0124] Specifically, the text embedding model employs the BGE (BAAI General Embedding) model, which is based on the Transformer architecture and can map text into low-dimensional dense vectors. It supports bilingual (Chinese and English) and mixed text, performing well across more than 100 languages. This model features multi-functionality, multilingual support, and multi-granularity, enabling dense retrieval, multi-vector retrieval, and sparse retrieval, supporting inputs of varying granularities from short sentences to long documents. Through pre-training on a large amount of medical text, this BGE model can better understand the semantics of medical terminology and diagnostic descriptions.

[0125] For each historical diagnosis record, the BGE model analyzes the semantics (first semantics) of the medical terms, symptom descriptions, disease names, and other words, as well as the logical relationships and contextual information (first contextual information) between these words. For example, for the historical diagnosis record "right femoral neck fracture caused by a fall from a height," the model understands that "fall from a height" is the cause and "right femoral neck fracture" is the result, and that there is a causal relationship between the two.

[0126] The BGE model maps these initial semantic and contextual information into a multi-dimensional vector space, generating a first vector representation that comprehensively represents the semantic features of the historical diagnostic information. In the vector space, historical diagnostic information with similar semantics (such as different types of fractures) will cluster in nearby regions, while historical diagnostic information with significant semantic differences (such as fractures and colds) will be distributed in more distant regions.

[0127] Specifically, the vector database can be the Milvus vector database. The generated first vector representation is stored in the Milvus vector database to construct a diagnostic information knowledge base, providing a foundation for subsequent semantic similarity retrieval.

[0128] In some optional implementations of this embodiment, the information retrieval module includes a weight allocation submodule, a vector transformation submodule, a similarity calculation submodule, and a diagnosis and determination submodule, wherein:

[0129] The weight allocation submodule is used to allocate weights to the medical diagnostic information according to the preset document priority.

[0130] The vector transformation submodule is used to analyze the second semantics of each word and the second contextual information between words in the weighted medical diagnostic information using a text embedding model, and to map the second semantics and the second contextual information to a vector space to obtain a second vector representation.

[0131] The similarity calculation submodule is used to obtain the first vector representation in the diagnostic information knowledge base and calculate the similarity between the first vector representation and the second vector representation.

[0132] The diagnosis determination submodule is used to determine the historical diagnosis information corresponding to the first vector representation as the target historical diagnosis information if the similarity is greater than a preset threshold.

[0133] Specifically, before performing semantic similarity retrieval, the extracted medical diagnostic information is weighted according to a preset document priority. Different types of claim reimbursement documents have different levels of importance in the claim process; for example, a formal hospital diagnosis certificate may be more authoritative than a regular outpatient medical record. Therefore, corresponding weights are assigned to the medical diagnostic information extracted from different claim reimbursement documents based on the document type. For example, a diagnosis certificate might have a weight of 1.5, while a regular outpatient medical record might have a weight of 1.0.

[0134] Specifically, after weight allocation, a text embedding model is used to process the medical diagnostic information. This text embedding model is the same as the one used when constructing the diagnostic information knowledge base. Using this text embedding model, the semantics (second semantics) of each word in the weighted medical diagnostic information and the contextual relationships between words (second contextual information) are analyzed. Then, these second semantics and second contextual information are mapped into a vector space to obtain a second vector representation.

[0135] For example, for the extracted medical diagnosis information "type 2 diabetes mellitus complicated with diabetic nephropathy (early stage)," the model will analyze the second semantics of words such as "type 2," "diabetes," "complication," "diabetic nephropathy," and "early stage" and the second contextual relationship between them, and generate a second vector representation that can represent the semantic features of the medical diagnosis information.

[0136] Specifically, the first vector representation of all historical diagnostic information is obtained from the diagnostic information knowledge base, and the similarity between the first vector representation and the second vector representation of the current medical diagnostic information is calculated. Similarity calculation can employ cosine similarity, Euclidean distance, or other vector similarity measures.

[0137] Specifically, if the similarity between the first vector representation of historical diagnostic information and the second vector representation of current medical diagnostic information is greater than a preset threshold (e.g., 0.85), then the historical diagnostic information is identified as the target historical diagnostic information. In this way, the system can find multiple target historical diagnostic information that are semantically similar to the current medical diagnostic information.

[0138] In some optional implementations of this embodiment, the above-mentioned accident determination module includes a cause extraction submodule and a first cause determination submodule, wherein:

[0139] The cause extraction submodule is used to extract the historical causes of accidents for the several target historical diagnostic information respectively, and to determine whether the several historical causes of accidents are the same;

[0140] The first cause determination submodule is used to determine the historical cause of an accident as the target cause of an accident when several historical causes of an accident are the same.

[0141] Specifically, when determining the cause of a target claim, the corresponding historical claim cause is first extracted from the historical diagnostic information of each target. For example, if three historical diagnostic records are retrieved, and their corresponding historical claim causes might be "disease," "disease," and "disease," then it is checked whether these historical claim causes are the same. If all historical claim causes are the same (such as "disease" in the example above), then this consistent historical claim cause is directly determined as the target claim cause for the current case.

[0142] In some optional implementations of this embodiment, the above-mentioned accident determination module further includes a first template acquisition submodule, a first text generation submodule, and an accident cause generation submodule, wherein:

[0143] The first template acquisition submodule is used to acquire a preset first prompt word template when several of the historical causes of accidents are different. The first prompt template is used to instruct the language model to deduce the cause of the accident in the medical diagnosis information based on the correlation between the diagnostic information.

[0144] The first text generation submodule is used to combine the medical diagnosis information, the target historical diagnosis information and the first prompt word template to generate the first prompt text;

[0145] The accident cause generation submodule is used to input the first prompt text into the language model and generate the target accident cause based on the medical diagnosis information and the target historical diagnosis information.

[0146] Specifically, if the historical causes of accidents are different, for example, if the historical causes of accidents corresponding to the three target historical diagnostic information retrieved are "illness", "accident" and "illness" respectively, then further analysis is needed to determine the most appropriate cause of accident.

[0147] Specifically, a preset first prompt word template is obtained. This first prompt word template is pre-designed and is used to instruct the language model to deduce the most appropriate cause of the accident based on the correlation between diagnostic information. The first prompt word template may be in the form of: "Please analyze the most likely cause of the accident based on the following current medical diagnostic information and historical similar cases. Current diagnosis: {current medical diagnostic information}; Historical case 1: {historical diagnostic information 1}, cause of accident: {historical cause of accident 1}; Historical case 2: {historical diagnostic information 2}, cause of accident: {historical cause of accident 2}; ...".

[0148] Specifically, the current medical diagnosis information, the target historical diagnosis information and the corresponding historical cause of the accident are filled into the first prompt word template to generate a complete first prompt text, and then the first prompt text is input into the language model.

[0149] Specifically, the language model analyzes the similarities and differences between the current diagnosis and various historical cases, taking into account medical knowledge and insurance claim rules, to deduce the most appropriate cause of the claim. For example, if the current diagnosis is "fracture caused by a fall," and historical cases include "fracture caused by a traffic accident" (caused by "accident") and "fracture caused by osteoporosis" (caused by "illness"), the language model will infer, based on the "fall" factor in the current diagnosis, that the cause of the claim is more likely to be "accident" rather than "illness."

[0150] Specifically, the first prompt text is input into the language model, and the language model infers the cause of the accident based on the instructions in the first prompt text, medical diagnosis information, and the target's historical diagnosis information.

[0151] In some optional implementations of this embodiment, the image recognition module includes an image recognition submodule, a data stitching submodule, a second template acquisition submodule, and a second text generation submodule, wherein:

[0152] The image recognition submodule is used to perform image recognition on the claim reimbursement documents using optical character recognition methods to obtain text information and image layout information.

[0153] The data splicing submodule is used to splice the text information according to the image layout information to obtain spliced ​​text data;

[0154] The second template acquisition submodule is used to acquire a preset second prompt word template, and combine the concatenated text data with the second prompt word template to generate a second prompt text. The second prompt word template is used to instruct the language model to extract key information from the input data.

[0155] The second text generation submodule is used to input the second prompt text into the language model and use the language model to extract key information from the concatenated text data as medical diagnostic information.

[0156] In this embodiment, during the image recognition process, Optical Character Recognition (OCR) technology is first used to convert the text in the claim reimbursement document (image data) into editable text. OCR technology can be achieved using tools such as Tesseract or ABBYY FineReader. The OCR tool is then used to perform image recognition on the claim reimbursement document, extracting text information and recognizing the text's position, size, font, and other image layout information within the image.

[0157] For example, when given an image of a medical expense invoice, OCR can identify textual information such as the hospital name, patient information, diagnosis name, treatment items, and cost on the invoice, while also recording the positional relationship of this textual information in the image.

[0158] Specifically, based on the identified image layout information, the text information is reasonably spliced ​​together. For example, text information in the same area is merged into a paragraph, or text information in the same table is organized according to row and column relationships, thereby obtaining structured spliced ​​text data.

[0159] Specifically, a pre-designed second prompt template is obtained. This template instructs the language model to extract key medical diagnostic information from the text. The prompt template might look like this: "Please extract the patient's main diagnostic information from the following medical document content, including the disease name, symptom description, and doctor's diagnosis. Medical document content: {concatenated text data}". Then, the concatenated text data is filled into the second prompt template to generate the second prompt text.

[0160] Specifically, the language model can be GPT-4, Claude, or other large-scale language models capable of understanding medical terminology and document structure, thereby accurately extracting key medical diagnostic information. The second prompt text is input into the language model, which then extracts key information from the concatenated text data based on the instructions of the second prompt text, serving as the medical diagnostic information.

[0161] For example, from a complex medical record, a language model may extract medical diagnostic information such as "primary diagnosis: type 2 diabetes; complication: diabetic nephropathy (early stage); symptoms: polydipsia, polyuria, weight loss".

[0162] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 6 , Figure 6 This is a basic structural block diagram of the computer device in this embodiment.

[0163] The computer device 6 includes a memory 61, a processor 62, and a network interface 63 that are interconnected via a system bus. It should be noted that only the computer device 6 with memory 61, processor 62, and network interface 63 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0164] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0165] The memory 61 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 61 may be an internal storage unit of the computer device 6, such as the hard disk or memory of the computer device 6. In other embodiments, the memory 61 may also be an external storage device of the computer device 6, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 6. Of course, the memory 61 may include both the internal storage unit and its external storage device of the computer device 6. In this embodiment, the memory 61 is typically used to store the operating system and various application software installed on the computer device 6, such as computer-readable instructions for claims processing. In addition, the memory 61 can also be used to temporarily store various types of data that have been output or will be output.

[0166] In some embodiments, the processor 62 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 62 is typically used to control the overall operation of the computer device 6. In this embodiment, the processor 62 is used to execute computer-readable instructions stored in the memory 61 or to process data, for example, to execute computer-readable instructions for the claims settlement method.

[0167] The network interface 63 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 6 and other electronic devices.

[0168] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the claims settlement method described above.

[0169] This application provides a computer device and computer-readable storage medium that uses image recognition technology with a processor to automatically extract medical diagnostic information from claims and reimbursement documents. This avoids the tediousness and errors of manual data entry, improving the efficiency and accuracy of information extraction. By utilizing a diagnostic information knowledge base constructed using a text embedding model and historical diagnostic information, semantic similarity retrieval is performed with the extracted medical diagnostic information to quickly find target historical diagnostic information similar to the current medical diagnostic information. Furthermore, semantic similarity retrieval allows for a more accurate understanding of the semantics of the medical diagnostic information, reducing errors in determining the cause of the accident due to misunderstandings. By determining the target cause of the accident based on several retrieved target historical diagnostic information, the subjectivity and uncertainty of manual judgment are reduced, improving the consistency of claims processing. Based on this application's solution, automated technology is used to perform image recognition, information extraction, and semantic similarity retrieval on user-uploaded claims and reimbursement documents, ultimately determining the target cause of the accident, significantly reducing manual intervention and improving claims processing efficiency and accuracy.

[0170] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0171] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

[0172] The software tools or components not belonging to our company that appear in the embodiments of this application are merely examples and do not represent actual use.

Claims

1. A method for handling claims after an accident, characterized in that, Includes the following steps: Obtain the claim reimbursement documents uploaded by the user. The claim reimbursement documents are image data and contain medical diagnosis information. Image recognition is performed on the claim reimbursement documents to extract the medical diagnosis information; Based on a pre-built diagnostic information knowledge base, semantic similarity retrieval is performed with the medical diagnostic information to obtain several target historical diagnostic information. The diagnostic information knowledge base is constructed using a text embedding model and historical diagnostic information. The cause of the target incident is determined based on the aforementioned historical diagnostic information of several targets; The corresponding claims processing procedure will be executed based on the cause of the incident.

2. The claims settlement method according to claim 1, characterized in that, Before the step of obtaining the user-uploaded claim reimbursement documents, the method further includes: Obtain several pre-collected historical diagnostic information entries; A text embedding model is used to analyze the first semantics of each word in the historical diagnostic information and the first contextual information between words. The first semantics and the first contextual information are then mapped to a vector space to obtain a first vector representation. The first vector representation is stored in a vector database to construct a diagnostic information knowledge base.

3. The claims settlement method according to claim 2, characterized in that, The step of performing semantic similarity retrieval between the medical diagnostic information and a pre-built diagnostic information knowledge base to obtain several target historical diagnostic information includes: The medical diagnostic information is weighted according to the preset document priority. A text embedding model is used to analyze the second semantics of each word and the second contextual information between words in the weighted medical diagnostic information. The second semantics and the second contextual information are then mapped to a vector space to obtain a second vector representation. Obtain the first vector representation from the diagnostic information knowledge base, and calculate the similarity between the first vector representation and the second vector representation; If the similarity is greater than a preset threshold, then the historical diagnostic information corresponding to the first vector representation is determined as the target historical diagnostic information.

4. The claims settlement method according to claim 1, characterized in that, The step of determining the cause of the target incident based on the aforementioned historical diagnostic information includes: Extract the historical causes of accidents from the historical diagnostic information of the several targets respectively, and determine whether the several historical causes of accidents are the same; When several historical causes of accidents are the same, the historical cause of accident is determined as the target cause of accident.

5. The claims settlement method according to claim 4, characterized in that, After determining whether the causes of several historical incidents are the same, the method further includes: When several historical causes of accidents are different, a preset first prompt word template is obtained. The first prompt template is used to instruct the language model to deduce the cause of the medical diagnosis based on the correlation between the diagnostic information. By combining the medical diagnostic information, the target historical diagnostic information, and the first prompt word template, a first prompt text is generated; The first prompt text is input into the language model, and the cause of the target accident is generated based on the medical diagnosis information and the target's historical diagnosis information.

6. The claims settlement method according to any one of claims 1-5, characterized in that, The step of performing image recognition on the claim reimbursement documents and extracting the medical diagnosis information includes: The claim reimbursement document is image-recognized using optical character recognition (OCR) to obtain text information and image layout information. The text information is concatenated based on the image layout information to obtain concatenated text data; Obtain a preset second prompt word template, and combine the concatenated text data with the second prompt word template to generate a second prompt text. The second prompt word template is used to instruct the language model to extract key information from the input data. The second prompt text is input into the language model, and the language model is used to extract key information from the concatenated text data as medical diagnostic information.

7. An accident claims processing device, characterized in that, include: The document acquisition module is used to acquire the claim reimbursement documents uploaded by the user. The claim reimbursement documents are image data and contain medical diagnosis information. The image recognition module is used to perform image recognition on the claim reimbursement documents and extract the medical diagnosis information; The information retrieval module is used to perform semantic similarity retrieval with the medical diagnostic information based on a pre-built diagnostic information knowledge base, and obtain several target historical diagnostic information. The diagnostic information knowledge base is constructed using a text embedding model and historical diagnostic information. The incident determination module is used to determine the cause of the incident based on the historical diagnostic information of the aforementioned targets; The claims processing module is used to execute the corresponding claims processing procedures based on the target cause of the accident.

8. The apparatus according to claim 7, characterized in that, The device further includes: The information acquisition module is used to acquire several pre-collected historical diagnostic information. The vector transformation module is used to analyze the first semantics of each word and the first contextual information between words in the historical diagnostic information using a text embedding model, and to map the first semantics and the first contextual information to a vector space to obtain a first vector representation; The knowledge base construction module is used to store the first vector representation into a vector database to construct a diagnostic information knowledge base.

9. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the claims settlement method as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the claims settlement method as described in any one of claims 1 to 6.