Claim settlement material processing method based on artificial intelligence and computer device

CN122415237BActive Publication Date: 2026-09-29SHANGHAI ZHONGAN XINKE INFORMATION TECH SERVICES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610857872.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-09-29
Estimated Expiration
2046-06-15

AI Technical Summary

Technical Problem

理赔案件通常涉及大量理赔材料,由于理赔材料来源多、格式差异大、内容表达不统一、字段之间交叉依赖强,导致在早期的材料识别阶段就产生错误判断,进而影响后续审核准确性

Benefits of technology

[0014]区别于相关技术,本申请中通过获取并分析理赔材料,根据材料特征对理赔材料分组并为各组分配处理通道,实现材料分类处理和通道差异化,避免不同类型材料混用导致的识别偏差,提高后续处理准确性;在各处理通道中,根据各处理通道对应的关键字段类型精准提取对应字段,各通道独立执行、互不干扰,支持差异化配置和并行处理,进一步减少无关字段干扰,提高字段抽取的针对性和准确率;基于提取字段、别名字段、候选标准字段以及理赔案件的场景信息确定目标提取字段,解决别名、同义词、OCR误识别导致的字段不一致问题,将非标准表达统一映射为标准名称,为后续质检提供可靠数据基础;基于目标提取字段更新理赔材料,并在预设维度上对更新后的理赔材料进行检验,使材料矛盾在早期阶段暴露;最终,响应于检验结果为通过,理赔案件进入理赔审核流程,响应于检验结果为不通过,根据拦截规则对理赔案件执行拦截处理以阻断其进入理赔审核流程,实现风险前置拦截,使高风险案件在正式审核前被识别和分流,降低错误自动审核风险,减少无效处理链路和人工复核成本,从而提高理赔审核的整体效率和准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415237B_ABST
    Figure CN122415237B_ABST
Patent Text Reader

Abstract

The application discloses an artificial intelligence-based claim settlement material processing method and computer equipment, relates to the technical field of artificial intelligence, and comprises the following steps: grouping claim settlement materials according to the material characteristics of the claim settlement materials and allocating processing channels; in each processing channel, extracting fields corresponding to key field types from each group of claim settlement materials according to the key field types corresponding to each processing channel to obtain extracted fields; determining target extraction fields according to the extracted fields, alias fields and scene information of the claim settlement case, wherein the alias fields are semantically matched with the extracted fields and have different expression forms; updating the claim settlement materials based on the target extraction fields, verifying the updated claim settlement materials in a preset dimension; in response to a verification result of passing, the claim settlement case enters a claim settlement audit process; and in response to a verification result of not passing, performing interception processing on the claim settlement case to block the claim settlement case from entering the claim settlement audit process, so that the reliability of claim settlement case processing can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based method for processing claims materials and a computer device. Background Technology

[0002] With the development of online insurance claims processing, more and more claims are being submitted online. Claims typically involve a large amount of documentation. Due to the diverse sources, varying formats, inconsistent content, and strong interdependencies between fields in these documents, errors can occur in the early stages of document identification, thus affecting the accuracy of subsequent reviews. Summary of the Invention

[0003] Therefore, it is necessary to provide an AI-based claims material processing method and computer equipment that can improve the reliability of claims material processing, addressing the aforementioned technical problems.

[0004] To address the aforementioned technical issues, firstly, an artificial intelligence-based method for processing claims materials is provided, including: Acquire and analyze claim materials for claims cases, group the claim materials according to their characteristics, and assign processing channels to each group of claim materials; In each processing channel, based on the key field type corresponding to each processing channel, the fields corresponding to the key field type are extracted from each group of claim materials to obtain the extracted fields. Based on the extracted fields, alias fields, and scenario information of the claims case, the target extracted fields are determined. Among them, the alias fields are semantically matched with the extracted fields but have different forms of expression. Update claim materials based on target extracted fields, and verify the updated claim materials on preset dimensions; If the inspection result is satisfactory, the claim will proceed to the claims review process; if the inspection result is unsatisfactory, the claim will be blocked from proceeding to the claims review process.

[0005] In one embodiment, the claim materials for a claim case are acquired and analyzed, and the claim materials are grouped according to their material characteristics. A processing channel is then assigned to each group of claim materials, including: Extract at least one of the following characteristics from the claim materials: material content, document name, and formatting features; Match the material characteristics with preset material characteristics, and determine the material type of the claim materials based on the matching results; Group the claim materials according to their type, and group claim materials of the same type into the same group; Based on the pre-established mapping relationship between groups and processing channels, a corresponding processing channel is assigned to each group; the claim materials of each group are sent to their respective processing channels for processing, and each processing channel is independent of the others.

[0006] In one embodiment, in each processing channel, based on the key field type corresponding to each processing channel, fields corresponding to the key field type are extracted from each group of claim materials, and the extracted fields include: Obtain the key field types corresponding to each processing channel. The key field types are pre-configured according to the group corresponding to the material type. In each processing channel, based on the key field type, locate the target position of the field that matches the key field type in each group of claim materials, and extract the claim material content corresponding to the target position; Based on the predefined business field architecture, the extracted claims materials are associated with key field types and assembled into structured fields, which are then used as extraction fields.

[0007] In one embodiment, the target extraction fields are determined based on the extracted fields, alias fields, and scenario information of the claims case, including: Search for alias fields in the preset alias mapping library that match the semantics of the extracted fields but have different expressions; The extracted fields and alias fields are assembled into a first prompt word according to the preset prompt word template. The first prompt word is then input into the first model. The first model searches for the text position that matches the first prompt word in the claim materials based on the first prompt word, determines the original text value based on the text content of the text position, and outputs it. Using the original text value as the query condition, a search is performed in the standard knowledge base. The semantic similarity between the original text value and each standard name in the standard knowledge base is calculated, and at least one standard name with the highest semantic similarity is obtained as a candidate standard field. The candidate standard fields, original text values, and application scenario information of the claims materials are assembled into second prompt words according to a preset prompt word template. The second prompt words are then input into the second model. The second model calculates the semantic similarity between each candidate standard field and the original text value based on the second prompt words, and obtains a semantic similarity score. The scenario adaptation score of each candidate standard field is determined based on the application scenario information. Based on the semantic similarity score and the scenario adaptation score, the comprehensive matching score of each candidate standard field is calculated. The candidate standard field with the highest comprehensive matching score is selected as the target extraction field.

[0008] In one embodiment, updating the claim materials based on the target extracted fields, and verifying the updated claim materials on preset dimensions, includes: Replace the corresponding original extracted fields with the target extracted fields to obtain the replaced claims material data; The replaced claim materials data are formatted and standardized to obtain standardized claim materials. The format standardization process includes unifying the format of at least one of the date field, amount field, ID number field, and gender field. Determine if there are any missing fields in the standardized claims materials. If there are missing fields, correct them based on the context information of the missing field's location to obtain updated claims materials. The updated claim materials are examined according to preset dimensions, which include at least one of the following: completeness of materials, authenticity of materials, consistency of materials, matching of liability, and time sequence.

[0009] In one embodiment, the inspection of the updated claims materials in a preset dimension includes: In response to preset dimensions, including the completeness of materials, the claims materials for the claims case are compared with the preset materials list to determine if there are any missing materials, and / or the extracted fields are compared with the preset key field types to determine if there are any missing fields, and / or the page numbers of the claims materials are identified, and the identification results are compared with the expected number of pages to determine if the claims materials are missing pages; In response to preset dimensions, including the material authenticity dimension, the target type extraction fields are obtained, and it is determined whether the target type extraction fields meet the corresponding format specification requirements; the target type includes at least one of the date field, amount field, ID number field, and gender field; In response to preset dimensions, including material consistency, select any inspection field to determine whether the target extraction field corresponding to the inspection field is consistent in different groups of claim materials; In response to preset dimensions, including liability matching dimensions, determine whether the diagnosis name in the extracted fields falls within the policy liability scope, and / or determine whether the expense items in the extracted fields match the diagnosis, and / or determine whether the accident description in the extracted fields is consistent with the accident type; In response to preset dimensions, including the time sequence dimension, the system obtains and determines whether the time of the incident, the time of medical treatment, the time of policy effectiveness, and the time of document issuance meet the preset time sequence.

[0010] In one embodiment, intercepting claims includes: The interception and processing operation for a claim case is determined based on the preset dimension type, the preset dimension quantity, and the inspection results of the preset dimensions. The interception and processing operation includes at least one of the following: suspending the processing of the claim case, waiting for supplementary materials, terminating the processing of the claim case, and transferring the claim case to the review process.

[0011] In one embodiment, the target extraction fields are determined based on the extracted fields, alias fields, and scenario information of the claims case, including: Retrieve alias fields that have semantically matching and different representations from the extracted fields; Based on the extracted fields and alias fields, multiple candidate standard fields are retrieved from the standard knowledge base; each candidate standard field is encoded as a population individual, and the population is initialized; The fitness value of each individual is calculated based on the fitness function, which is constructed with multiple candidate standard fields as the search space and the extraction field, alias field and scenario information of the claims case as constraints. Parent individuals are selected based on the fitness value of each individual; Candidate individuals are generated by performing crossover and mutation operations on the parent individuals; Calculate the difference between the fitness value of each candidate individual and the fitness value of the corresponding parent individual; Candidates whose difference is greater than or equal to a preset value are treated as new individuals; In response to a difference less than a preset value, the selection probability is calculated based on the difference and the annealing temperature. Candidates with a selection probability greater than the target value are selected as new individuals. The annealing temperature is determined based on the number of candidate standard fields and the distribution of fitness values. The annealing temperature is iteratively updated according to the cooling coefficient. Population updates based on new individuals; The process iteratively executes the steps from selecting a parent individual based on the fitness value of each individual to updating the population based on the new individual, until the convergence condition is met. The candidate criterion field corresponding to the individual with the highest fitness value at the time of convergence is used as the target extraction field.

[0012] In one embodiment, the fitness function is as follows: ; Where F(s) is the fitness value of the candidate criterion field, v is the content of the claim materials corresponding to the extracted field, s is the candidate criterion field, E() is the text vector embedding function used to map the input text to a vector, and R(s) is the scenario constraint mapping matching function. M(v,c) is the standard name selected from the standard knowledge base based on the content v of the claim materials, which matches the area c where the claim occurred. P is the penalty item, and a, b, and d are weight coefficients, with a+b+d=1.

[0013] To address the aforementioned technical problems, a second aspect provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: the processor executes the computer program to perform the steps of the method described in the first aspect.

[0014] Unlike related technologies, this application acquires and analyzes claim materials, groups them according to their characteristics, and assigns processing channels to each group. This achieves material classification and channel differentiation, avoiding recognition errors caused by mixing different types of materials and improving the accuracy of subsequent processing. Within each processing channel, corresponding fields are accurately extracted based on their key field types. Each channel executes independently without interference, supporting differentiated configurations and parallel processing, further reducing interference from irrelevant fields and improving the targeting and accuracy of field extraction. The target extraction field is determined based on the extracted fields, alias fields, candidate standard fields, and scenario information of the claim case, addressing the issues caused by aliases, synonyms, and OCR misidentification. To address inconsistencies in fields, non-standard expressions are uniformly mapped to standard names, providing a reliable data foundation for subsequent quality inspection. Claim materials are updated based on target extracted fields, and the updated materials are inspected on preset dimensions to expose inconsistencies at an early stage. Finally, in response to a pass inspection result, the claim enters the claim review process; in response to a fail inspection result, the claim is blocked according to interception rules to prevent it from entering the claim review process. This achieves proactive risk interception, allowing high-risk cases to be identified and diverted before formal review, reducing the risk of erroneous automatic review, minimizing invalid processing links and manual review costs, thereby improving the overall efficiency and accuracy of claim review. Attached Figure Description

[0015] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart illustrating an artificial intelligence-based claims material processing method in one embodiment; Figure 2 This is a flowchart illustrating an artificial intelligence-based claims material processing method in another embodiment; Figure 3 This is a flowchart illustrating an artificial intelligence-based claims material processing method in yet another embodiment; Figure 4 This is a structural block diagram of an AI-based claims material processing device in one embodiment; Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0018] With the development of online insurance claims processing, an increasing number of claims are being submitted online. Claims typically involve various documents such as outpatient and emergency room medical records, inpatient medical records, laboratory reports, receipts, expense lists, accident reports, identification documents, and policy liability information. To improve processing efficiency, existing systems often incorporate OCR, field extraction, rule validation, and automated review technologies to identify and initially assess the documents.

[0019] For example, one related technology sets up an OCR-based single-rule verification scheme. Specifically, after receiving the claim materials, the materials are uniformly OCR-recognized, and the recognition results are then directly sent to the rule engine for verification and review. This scheme has a simple process, but it often assumes accurate material classification, accurate field extraction, and no obvious contradictions between fields, thus making it less adaptable to complex cases.

[0020] The second related technical solution sets up a field extraction scheme based on material type. Specifically, some systems will first distinguish between material types such as medical records, invoices, and expense lists, and then call different extraction templates to extract fields. This scheme is an improvement over the unified identification scheme, but it usually only solves the problem of "whether extraction is possible" and lacks a complete mechanism for unified correction, cross-material comparison, and correlation verification of the extracted fields.

[0021] However, processing claims materials differs from simple form entry. The difficulty lies in the diverse sources of materials, significant format differences, inconsistent content expression, and strong interdependence between fields. Many claims are not due to a lack of review rules, but rather to incorrect material classification, incorrect field extraction, inconsistent expression of key fields, and contradictions between different materials. This leads to erroneous judgments by the system in the early stages, thus affecting the accuracy of subsequent reviews.

[0022] For example, the hospital name in the same claim may be written differently in medical records, invoices, and expense lists; the same visit date, diagnosis, and insured's identity information may differ in different materials; some key materials, although uploaded, may have abnormal formatting, incorrect page order, or incomplete content, causing subsequent rules to fail to apply them correctly. If the system lacks a systematic phased material processing mechanism and a consistency quality inspection mechanism, problems often only surface during the final review stage, increasing manual review costs and affecting review timeliness.

[0023] Therefore, a processing solution more suitable for the characteristics of claims materials is needed, so that the materials undergo multi-stage classification, extraction, correction and consistency inspection before entering formal review, and high-risk abnormal situations are identified and intercepted in advance.

[0024] To address the aforementioned technical problems, in one embodiment, such as Figure 1 As shown, an artificial intelligence-based method for processing claims materials is provided, which includes the following steps: Step 101: Obtain and analyze the claim materials for the claim case, group the claim materials according to their characteristics, and assign processing channels to each group of claim materials.

[0025] Specifically, at least one of the following characteristics of the claim materials—content, document name, and format—is extracted as material features; these material features are matched with preset material features to determine the material type of the claim materials based on the matching results; the claim materials are grouped according to their material types, with claim materials of the same type grouped together; a corresponding processing channel is assigned to each group based on the pre-established mapping relationship between groups and processing channels; and each group of claim materials is sent to its respective processing channel for processing, with each processing channel operating independently.

[0026] Obtain claim materials, which include original claim documents and relevant business information, specifically medical records, invoices, expense lists, examination and testing reports, accident report or incident certificate, insured's identity information, policy and liability information, historical claim records, and external query results. After receiving the claim materials, you can assign them a basic number and archive the case, providing a unified entry point for subsequent multi-stage processing.

[0027] The claims materials are parsed. If the materials are PDF, OFD, or Word files, the text data and layout information are directly extracted. If the materials are image files, optical character recognition (OCR) is performed to extract the text data. The content is extracted from the parsed text data, the name information is extracted from the file name, and the layout features are extracted from the layout information.

[0028] Layout features include page layout, title text, and seal position; field features include specific keywords (such as diagnosis, total amount) and specific codes (such as invoice code, medical record number).

[0029] The extracted material features are compared one by one with preset material feature templates. The matching degree between the material features and each preset material feature is calculated, and the material type of each claim material is determined based on the preset material feature with the highest matching degree. Material types include different groups such as medical records, invoices, lists, certificates, and identity documents. For example, if the material contains keywords such as chief complaint, present medical history, and diagnosis and has a layout unique to medical records, it is identified as a medical record; if the material contains keywords such as invoice code, invoice number, amount, and tax, it is identified as an invoice; if the material contains table fields such as item name, unit price, quantity, and amount, it is identified as a list.

[0030] Claim materials are grouped according to their identified material types, with materials of the same type grouped together. Based on a pre-established mapping between groups and processing channels, a corresponding processing channel is assigned to each group. Each processing channel is a data processing pipeline independently configured for a specific material type. Each channel corresponds to a set of processing parameters, including field extraction rules, field mapping rules, and verification rules. The processing channels are independent of each other and can be executed in parallel.

[0031] Claim materials are grouped according to their identified material types, with materials of the same type grouped together. Based on a pre-established mapping between groups and processing channels, a corresponding processing channel is assigned to each group. Each processing channel is a data processing pipeline independently configured for a specific material type, with each channel corresponding to a set of processing parameters, including field extraction rules, field mapping rules, and verification rules. The processing channels are independent of each other and can be executed in parallel.

[0032] Step 102: In each processing channel, based on the key field type corresponding to each processing channel, extract the fields corresponding to the key field type from each group of claim materials to obtain the extracted fields.

[0033] Specifically, the key field types corresponding to each processing channel are obtained. The key field types are pre-configured according to the group corresponding to the material type. In each processing channel, the target position of the field matching the key field type is located in each group of claim materials according to the key field type, and the claim material content corresponding to the target position is extracted. According to the predefined business field architecture, the extracted claim material content is associated with the key field type and assembled into structured fields, which are used as extraction fields.

[0034] Key field types are pre-configured fields to be extracted for specific material types. For example, in medical records, this includes the diagnosis, date of visit, and hospital name; in invoices, it includes the amount, charge item, and invoice number. For instance, the key field types configured for the medical record processing channel might include the patient's name, ID number, hospital name, date of visit, and diagnosis name; for invoices, the key field types might include charge item, amount, invoice number, and issuing agency; and for bills, the key field types might include expense details, item name, and quantity information.

[0035] In each processing channel, based on the key field type corresponding to that processing channel, the target location of the field matching the key field type is located in each group of claim materials, and the claim material content corresponding to the target location is extracted.

[0036] For example, an optical character recognition (OCR) program can be invoked to perform text recognition on claim materials, and a layout analysis program can be invoked to parse the layout structure of the claim materials, identifying text areas, table areas, and seal areas in the image, and determining the location and type information of each area. The identified text areas are then processed to convert the text in the image into computer-readable text data. Next, based on the key field types corresponding to each processing channel, the field positions matching the key field types are located in the text data, and the corresponding field content is extracted.

[0037] To transform the extracted content into structured data usable for subsequent verification, the extracted claims material content can be associated with key field types according to a predefined business field architecture, assembling them into structured fields for extraction. This business field architecture consists of a predefined set of standard field names, field key names, field types, and field extraction hints for each material type.

[0038] In one specific embodiment, prompt words can be assembled according to a predefined business field architecture. A large language model is then invoked to perform field mapping and assembly, forming specified key-value pair JSON structured data. For example, for medical record materials, the business field architecture defines a standard field name (hospital name), a field key name (hospital_name), a field type of string, and a field extraction prompt word that extracts the full name of the hospital visited from the medical record. This information is assembled into prompt words, and the large language model extracts the corresponding information from the claim materials based on the prompt words, outputting structured fields in a specified format, such as {hospital_name: First People's Hospital of a Certain City, diagnosis: acute appendicitis, visit_date: 2025-03-20}.

[0039] In this way, targeted field extraction for different types of materials can be achieved. This not only extracts content from the materials but also transforms the content into structured fields that can be used for subsequent rule verification, providing a standardized data foundation for subsequent field correction, consistency quality inspection, and risk interception.

[0040] Step 103: Based on the extracted fields, alias fields, and scenario information of the claims case, determine the target extracted fields. Among them, the alias fields and the extracted fields are semantically matched but have different expressions.

[0041] Specifically, search for and extract alias fields from the preset alias mapping library that have semantic matching with the extracted fields but different expression forms; The extracted fields and alias fields are assembled into a first prompt word according to the preset prompt word template. The first prompt word is then input into the first model. The first model searches for the text position that matches the first prompt word in the claim materials based on the first prompt word, determines the original text value based on the text content of the text position, and outputs it. Using the original text value as the query condition, a search is performed in the standard knowledge base. The semantic similarity between the original text value and each standard name in the standard knowledge base is calculated, and at least one standard name with the highest semantic similarity is obtained as a candidate standard field. The candidate standard fields, original text values, and application scenario information of the claims materials are assembled into second prompt words according to a preset prompt word template. The second prompt words are then input into the second model. The second model calculates the semantic similarity between each candidate standard field and the original text value based on the second prompt words, and obtains a semantic similarity score. The scenario adaptation score of each candidate standard field is determined based on the application scenario information. Based on the semantic similarity score and the scenario adaptation score, the comprehensive matching score of each candidate standard field is calculated. The candidate standard field with the highest comprehensive matching score is selected as the target extraction field.

[0042] The alias mapping library is a pre-built database that stores the mapping relationship between each standard field name and its common aliases, abbreviations, and synonyms. For example, for the hospital name field, its alias fields include the hospital, medical institution, and hospital name.

[0043] The system searches for alias fields in a pre-defined alias mapping library that semantically match the extracted field but have a different expression. The extracted field and alias field are then assembled into a first prompt word according to a pre-defined prompt word template. This first prompt word instructs the first model to search for text content in the claim materials that matches either the extracted field or the alias field. The first prompt word is then input into the first model. The first model can be a large visual model or a multimodal model. The first model searches for the text location in the claim materials that matches the first prompt word, determines the original text value based on the text content at that location, and outputs it. For example, for the extracted field "Hospital Name" and its alias "Visiting Hospital," the first model locates the visiting hospital in the medical record materials. The location of "Visiting Hospital: Municipal First Hospital" is extracted and output as the original text value.

[0044] The system uses the original text values ​​as query criteria to perform a search within a standard knowledge base. This standard knowledge base is a pre-built repository that stores standard names and their mappings, including standard hospital name databases, standard diagnosis name databases, standard drug name databases, and standard billing item name databases.

[0045] Calculate the semantic similarity between the original text value and each standard name in the standard knowledge base, and select at least one standard name with the highest semantic similarity as a candidate standard field. For example, using the original text value "City First People's Hospital" to search the standard hospital name database, the semantic similarity is calculated as 0.95 for "City First People's Hospital", 0.85 for "City Municipal Hospital", and 0.80 for "City People's Hospital". The first three are then selected as candidate standard fields.

[0046] The candidate standard fields, original text values, and application scenario information from claim materials are assembled into a second prompt word according to a preset prompt word template. The application scenario information includes information such as the location of the accident, the location of treatment, the type of insurance, and the type of treatment.

[0047] The second prompt word is input into the second model. The second model can be a large language model. The second model performs the following operations based on the second prompt word: First, it calculates the semantic similarity between each candidate standard field and the original text value to obtain a semantic similarity score; second, it determines the scenario adaptation score of each candidate standard field based on the application scenario information; then, it calculates the comprehensive matching score of each candidate standard field based on the semantic similarity score and the scenario adaptation score; finally, it selects the candidate standard field with the highest comprehensive matching score as the target extraction field.

[0048] This application utilizes alias fields and scenario information to correct and standardize extracted fields, solving problems such as OCR misidentification, alias misuse, and inconsistent terminology. It maps non-standard expressions to standard names, providing a reliable data foundation for subsequent quality inspection.

[0049] In one embodiment, determining the target extraction field based on the extraction field, alias field, and scenario information of the claim case includes: obtaining an alias field that semantically matches the extraction field but has a different expression; retrieving at least one candidate standard field from a standard knowledge base based on the extraction field and alias field; encoding each candidate standard field into a population individual and initializing the population; calculating the fitness value of each individual based on a fitness function, wherein the fitness function is constructed with multiple candidate standard fields as the search space and the extraction field, alias field, and scenario information of the claim case as constraints; selecting a parent individual based on the fitness value of each individual; generating candidate individuals by performing crossover and mutation operations based on the parent individuals; and calculating the difference between the fitness value of each candidate individual and the fitness value of the parent individual corresponding to that candidate individual.

[0050] Candidate individuals with a difference greater than or equal to a preset value are selected as new individuals; in response to a difference less than a preset value, the selection probability is calculated based on the difference and the annealing temperature, and candidate individuals with a selection probability greater than the target value are selected as new individuals. The annealing temperature is determined based on the number of candidate standard fields and the distribution of fitness values, and is iteratively updated according to the cooling coefficient; the population is updated based on the new individuals. The process iteratively executes the steps from selecting a parent individual based on the fitness value of each individual to updating the population based on the new individual, until the convergence condition is met. The candidate criterion field corresponding to the individual with the highest fitness value at the time of convergence is used as the target extraction field.

[0051] In this implementation, an alias field that semantically matches the extracted field but has a different expression is obtained. A set of aliases corresponding to the extracted field is searched from a preset alias mapping library. For example, aliases for the extracted field "hospital name" include "hospital," "medical institution," and "hospital name." Then, based on the extracted field and the alias field, at least one candidate standard field is retrieved from a standard knowledge base. The standard knowledge base stores standard names and their semantic vectors. By calculating the semantic similarity between the extracted field and the standard names, multiple standard names with high similarity are obtained as a set of candidate standard fields.

[0052] Each candidate criterion field is encoded into an individual in the population, and the population is initialized. Each candidate criterion field corresponds to one individual, and the individual contains the name, semantic vector, and geographic attribute information of that field. The initial population contains all individuals in the set of candidate criterion fields.

[0053] The fitness value of each individual is calculated based on the fitness function. The fitness function is shown below: ; Where F(s) is the fitness value of the candidate criterion field, v is the content of the claim materials corresponding to the extracted field, s is the candidate criterion field, E() is the text vector embedding function used to map the input text to a vector, and R(s) is the scenario constraint mapping matching function. M(v,c) is the standard name selected from the standard knowledge base based on the content v of the claim materials, which matches the area c where the claim occurred. P is the penalty item, and a, b, and d are weight coefficients, with a+b+d=1.

[0054] The fitness value represents the semantic similarity between v and s. It is set to 1 when the semantic similarity is higher than the similarity threshold but the candidate standard field does not match the regional mapping result; otherwise, it is set to 0. Higher semantic similarity results in a higher fitness value. The fitness value increases when the candidate standard field equals the mapping result under regional constraints; conversely, the fitness value decreases when the semantic similarity is high but contradicts the regional mapping result. In this way, the algorithm can select a standard name that is semantically matching and conforms to regional constraints.

[0055] In this application, the weight coefficients a, b, and d in the fitness function are dynamically adjusted based on the scenario information of the claims case. The specific formula is as follows: a=a0·(1-α·region_ambiguity); Here, `region_ambiguity` represents the degree of regional ambiguity, defined as follows: when the alias of the extracted field corresponds to different standard names in different regions, `region_ambiguity` is 1; otherwise, it is 0. For example, if "People's Hospital" corresponds to "People's Hospital of City A" in City A and "People's Hospital of City B" in City B, then `region_ambiguity` = 1; while "acute appendicitis" corresponds to the same diagnostic standard in different regions, then `region_ambiguity` = 0. `a0` is the baseline weight, defaulting to 0.6, and `α` is the adjustment coefficient, defaulting to 0.2.

[0056] b=b0·(1+β·region_confidence); Here, `region_confidence` represents the reliability of regional information, defined as follows: if the claim materials explicitly include information about the region of the incident or the location of medical treatment (e.g., the materials clearly indicate the city of A or city of B), then `region_confidence` is 1; if the region is inferred from policy information or historical records (inference confidence ≥ 0.8), then `region_confidence` is 0.6; if no regional information can be obtained, then `region_confidence` is 0. `b0` is the baseline weight, defaulting to 0.2, and `β` is the adjustment coefficient, defaulting to 0.5.

[0057] d=d0·(1-γ·history_accuracy); Here, `history_accuracy` represents the historical matching accuracy, defined as: for the same extracted field (e.g., People's Hospital), in past claims cases, the ratio of the target extracted field selected by this method to the correct standard name confirmed by humans. The value of `history_accuracy` ranges from [0,1], with an initial value of 0.5, and is updated every 100 cases processed. `d0` is the baseline weight, defaulting to 0.2, and `γ` is the adjustment coefficient, defaulting to 0.3.

[0058] Parent individuals can be selected using roulette wheel selection or tournament selection. Individuals with higher fitness values ​​have a greater probability of being selected as parents, thus preserving superior genes. Based on the parent individuals, crossover and mutation operations are performed to generate candidate individuals. Crossover involves exchanging some genes between two parent individuals to generate a new individual; mutation randomly perturbs the genes of an individual, introducing new genetic diversity and preventing the algorithm from getting trapped in local optima.

[0059] For example, the crossover operation uses a uniform crossover method. For the encoding vectors of two parent individuals, each gene dimension is independently decided to be swapped with a preset crossover probability to generate two offspring individuals. The mutation operation uses a non-uniform mutation method. For the encoding vector of an individual, selected genes are mutated with a preset mutation probability. The mutation is asynchronous and decays with the number of iterations.

[0060] Calculate the difference between the fitness value of each candidate individual and the fitness value of its corresponding parent individual. If the difference is greater than or equal to a preset value (usually 0), it indicates that the candidate individual is superior to or equal to its parent individual, and the candidate individual is added to the next generation population as a new individual. If the difference is less than the preset value, it indicates that the candidate individual is inferior to its parent individual, and the selection probability is calculated based on the difference and the current annealing temperature, where the selection probability P = exp(ΔF / T), and T is the current annealing temperature. Generate a random number (target value) in the interval [0,1]. If the random number is less than the selection probability, the candidate individual is accepted as a new individual; otherwise, the candidate individual is rejected. The probability acceptance mechanism allows the algorithm to accept inferior solutions with reduced fitness values ​​with a certain probability, thereby escaping local optima and improving global search capabilities.

[0061] The initial annealing temperature is determined based on the number of candidate criteria fields and the distribution of fitness values. For example, a larger number of candidate criteria fields or a more dispersed distribution of fitness values ​​indicates greater differences in fitness among individuals, allowing for a higher initial temperature to enhance the initial exploration capability. Conversely, a smaller number of candidate criteria fields or a more concentrated distribution of fitness values ​​indicates smaller differences in fitness among individuals, allowing for a lower initial temperature. After each iteration, the annealing temperature is updated according to a cooling strategy: T{k+1}=η×Tk, where T is the current annealing temperature, k is the iteration number, and η is the cooling coefficient, ranging from 0.8 to 0.99. A cooling coefficient closer to 1 results in slower cooling and a more thorough search by the algorithm; a smaller cooling coefficient results in faster cooling and a faster convergence speed.

[0062] The population is updated based on new individuals. Accepted new individuals are added to the next generation of the population to form a new population. The steps of selecting parent individuals, crossover and mutation, calculating fitness difference, probability acceptance, annealing temperature update, and population update are iteratively executed until convergence conditions are met. Convergence conditions may include: reaching a preset maximum number of iterations, the fitness value change of the best individuals across multiple generations being less than a preset threshold, or the annealing temperature being lower than a preset minimum temperature (e.g., 0.01). When convergence conditions are met, the candidate standard field corresponding to the individual with the highest fitness value at convergence is output as the target extraction field. This achieves global optimization search in the candidate standard field space. Combined with semantic similarity and regional constraints, it effectively solves problems such as homonymous addresses and alias misuse, improving the accuracy and robustness of field correction.

[0063] Step 104: Update the claim materials based on the target extracted fields, and verify the updated claim materials on the preset dimensions.

[0064] Specifically, the target extracted fields are used to replace the corresponding original extracted fields to obtain the replaced claim material data; the replaced claim material data is then subjected to format standardization processing to obtain standardized claim materials; the format standardization processing includes unifying the format of at least one of the date field, amount field, ID number field, and gender field; it is determined whether there are any missing fields in the standardized claim materials, and in response to the existence of missing fields, the missing fields are corrected according to the context information of the missing field's location to obtain the updated claim materials; the updated claim materials are then verified on preset dimensions, which include at least one of the following dimensions: material completeness, material authenticity, material consistency, liability matching, and time sequence.

[0065] Replacing the corresponding original extracted fields with the target extracted fields to obtain the replaced claim material data. Performing format standardization processing on the replaced claim material data to obtain standardized claim materials. The format standardization processing includes unifying the format of at least one of a date field, an amount field, an ID number field and a gender field. For example, for a date field, "March 5, 2023" is uniformly converted into the standard format of "2023-03-05"; for an amount field, currency symbols are removed and Chinese uppercase amounts are converted into Arabic numerals; for an ID number field, spaces and connecting symbols are removed and the result is uniformly converted into uppercase letters; for a gender field, "M" and "Male" are uniformly converted into "Male", and "F" and "Female" are uniformly converted into "Female". For an invoice field, an invoice code and an invoice number are extracted and combined, the combined character string has a length of 10 to 17 bits, the original characters are retained when letters are contained, and an invalid format is marked as abnormal.

[0066] In response to that the preset dimension includes a complete material dimension, comparing the claim materials of a claim case with a preset material list to determine whether there is any missing material, and / or comparing the extracted fields with preset key field types to determine whether there is any missing field, and / or performing page number identification on the claim materials, and comparing an identification result with an expected number of pages to determine whether any page of the claim materials is missing.

[0067] Comparing the claim materials of a claim case with a preset material list to determine whether there is any missing material. The necessary material list is preset according to the case type (such as outpatient claim, inpatient claim, accident insurance claim) and policy clauses. For example, the necessary materials for an inpatient claim case include admission records, discharge summaries, inpatient expense lists, invoices, etc. A set comparison is performed between the actually submitted material type list and the necessary material list, and if any material type in the necessary material list does not appear in the actually submitted materials, a missing material list is generated.

[0068] Comparing the extracted fields with preset key field types to determine whether there is any missing field. The key field list template is preset according to the case type. For example, for an outpatient claim case, the key fields include the name of the patient, the visit date, the diagnosis name, the invoice amount, etc. Traverse the extracted structured fields, compare them one by one with the key field list template, record null or Null fields, and generate a missing field list.

[0069] Page number identification is performed on claim materials, and the identification results are compared with the expected number of pages to determine if any pages are missing. For multi-page materials (such as PDF files or multi-page images), page number markers are identified through layout analysis or the actual number of pages is determined through page count. This actual number is then compared with the expected number of pages (which can be preset based on the material type or extracted from the material's own information). If the actual number of pages is less than the expected number, it is determined that pages are missing. If either of the above determinations indicates that there are missing pages, the case is marked as having incomplete materials, and the specific missing item information is recorded.

[0070] In response to preset dimensions, including the material authenticity dimension, the target type extraction field is obtained, and it is determined whether the target type extraction field meets the corresponding format specification requirements; the target type includes at least one of the date field, amount field, ID number field and gender field.

[0071] For date fields, check if the date is a valid calendar date (e.g., February 30th is invalid); if the date is no later than the current date; and if the date format is parsable. If the date is March 5, 2023, parse it into the standard format of 2023-03-05; if the date is an invalid format or an invalid date, mark it as an exception.

[0072] For the amount field, the system verifies whether the amount is a valid positive number; whether the amount exceeds a preset reasonable range (e.g., outpatient expenses exceeding 1 million yuan trigger an anomaly flag); and the consistency between the total amount and the sum of the detailed item amounts. A difference exceeding a preset threshold, such as 5%, is flagged as an anomaly. For example, if the total invoice amount is 1000 yuan and the sum of the detailed expense list is 950 yuan, the difference is 5%. If the preset threshold is 5%, it is flagged as an anomaly; if the difference is 6%, it exceeds the threshold and is also flagged as an anomaly.

[0073] For the ID number field, after removing spaces and connectors, the ID number is uniformly converted to uppercase letters. If it is an ID card number, the length is checked to ensure it is 18 digits, and the check digits are correct. For example, if the ID card number is 18 digits long after removing spaces and the check digits are calculated correctly, the check passes; if the check rules are not met, it is marked as abnormal.

[0074] For the gender field, determine whether the gender field is either male or female; if there are English expressions such as M, F, Male, Female, convert them to Chinese male or female; if the value is neither Chinese gender nor a common English gender expression, mark it as an exception.

[0075] If any validation fails, the field is marked as having a format error, and the error details are recorded.

[0076] In response to preset dimensions, including material consistency, select any inspection field to determine whether the target extraction field corresponding to the inspection field is consistent in different groups of claim materials.

[0077] Please see Figure 2 Because the supporting documents for a claim can come from multiple sources, Figure 2 The document shows sources including medical records, invoices, expense lists, and insurance policies. Many fields have different semantic expressions across these sources. This application allows selecting fields with the same semantic meaning (such as hospital name, insured's name, and date of visit) as verification fields. The target field value is obtained from different groups of claim materials, and these values ​​are compared pairwise. If the same field value is completely identical in different materials, it is considered consistent. If the same field value differs in different materials but is the same after standard name mapping (e.g., the medical record says "City 1 Hospital," the invoice says "First People's Hospital," both after standard name mapping are "City 1 People's Hospital"), it is also considered consistent. If the same field value has substantial differences in different materials (e.g., the medical record says "City A People's Hospital," the invoice says "City B People's Hospital"), it is considered inconsistent, the field is marked as a cross-material inconsistency anomaly, and the specific inconsistent value is recorded.

[0078] In one embodiment, the same type of fields extracted from different materials in the same claim case and the preset consistency verification rules are assembled into quality inspection prompt words. The large language model is called to determine whether the core fields between different materials are consistent based on the quality inspection prompt words, and the execution result, judgment basis and confidence information of each rule are output.

[0079] In response to preset dimensions, including liability matching dimensions, determine whether the diagnosis name in the extracted fields falls within the policy's liability scope, and / or determine whether the expense items in the extracted fields match the diagnosis, and / or determine whether the accident description in the extracted fields is consistent with the accident type.

[0080] A pre-defined mapping database between insurance liabilities and diagnoses can be established. This database uses insurance liabilities as indexes to associate corresponding ICD-10 diagnostic codes or standard diagnostic names. For example, if the policy liability is inpatient medical care and the diagnosis is acute appendicitis (ICD-10: K35), the extracted diagnostic name is matched against the mapping database. If the diagnoses included in the liability include the extracted diagnosis, it is considered a reasonable association; otherwise, it is marked as a mismatch between diagnosis and liability coverage.

[0081] A pre-defined knowledge base can be established to link diagnostic treatments with reasonable cost items. This knowledge base uses diagnoses, surgical procedures, and treatment methods as entries, associating them with corresponding reasonable cost categories, such as examination fees, medication fees, and surgical fees. For example, reasonable cost items associated with a diagnosis of acute bronchitis include blood tests, chest X-rays, antibiotics, and cough suppressants / expectorants. The various cost items in the expense list are matched against the set of reasonable costs associated with the corresponding diagnosis in this knowledge base: if a cost item falls within this set or its subcategory, it is considered reasonable; if a cost item is found that is significantly unrelated to the diagnosis (e.g., orthopedic implant material fees appear in the expense list for an acute bronchitis case), it is marked as a mismatch between cost and diagnosis.

[0082] The accident report and the accident type declared at the time of the accident can be input into a large language model. The model then performs a step-by-step judgment according to a preset reasoning logic. First, it extracts the core accident elements (including the cause of injury, the object or environment that caused the injury, and the nature of the injury) from the accident report. Then, it performs a semantic comparison between the extracted elements and the typical features of the declared accident type. At the same time, it identifies whether there are descriptions in the accident report that negate or exclude the declared accident type (e.g., a traffic accident is declared but the accident report clearly states that the person slipped in their own bathroom). Finally, it comprehensively determines whether the accident report and the declared accident type are consistent in core facts. If the model determines that there is a contradiction, the case is marked as abnormal.

[0083] In response to preset dimensions, including the time sequence dimension, the system obtains and determines whether the time of the incident, the time of medical treatment, the time of policy effectiveness, and the time of document issuance meet the preset time sequence.

[0084] The timing of the medical visit should be greater than or equal to the policy's effective date (i.e., the visit occurs after the policy takes effect); the time of the incident should be less than or equal to the medical visit time (i.e., the incident occurs before or on the day of the medical visit); the time of the incident should be greater than or equal to the policy's effective date (i.e., the incident occurs after the policy takes effect); and the time of document issuance should be greater than or equal to the medical visit time (i.e., the document issuance occurs after or on the day of the medical visit). If all timing rules are met, the timing is deemed reasonable; if any timing rule is not met, such as the incident time being earlier than the policy's effective date, or the document issuance time being earlier than the medical visit time, the timing is deemed abnormal, marked as abnormal, and the specific non-compliance is recorded.

[0085] After completing the pre-defined dimension checks, summarize all check results, including the pass status, anomaly markers, and anomaly details for each dimension. If all dimensions pass the check, the overall check result is considered passable; if any dimension fails the check, the overall check result is considered failable, and the information of the failed dimension and anomaly details are recorded for subsequent risk interception or manual review.

[0086] By setting up multi-dimensional inspections, the completeness, authenticity, consistency, liability matching, and time sequence of the updated claims materials can be verified, exposing contradictions in the claims materials at an early stage and providing a reliable basis for subsequent risk interception or approval.

[0087] Step 105: If the inspection result is passed, the claim case enters the claim review process; if the inspection result is failed, the claim case is intercepted to prevent it from entering the claim review process.

[0088] When all preset dimensions pass the inspection, the inspection result is deemed satisfactory. The updated claim materials data, along with the inspection results for each dimension, are then output to the subsequent claim review process. This subsequent review process can be either an automated review engine or a manual review queue. Cases that pass the inspection are marked as eligible for review, triggering the subsequent review process for formal review and processing of that claim.

[0089] When any preset dimension fails the inspection, the inspection result is determined to be unsuccessful. Based on the preset dimension type, the number of preset dimensions, and the inspection result of the preset dimensions, the interception and processing operation of the claim case is determined. The interception and processing operation includes at least one of the following: suspending the processing of the claim case, waiting for supplementary materials, terminating the processing of the claim case, and transferring the claim case to the review process.

[0090] This interception operation is used when the inspection fails due to missing materials, missing fields, or missing pages. A request for supplementary documents is generated, clearly stating the types of materials that need to be supplemented or the field information that needs to be corrected. The claim is then suspended, awaiting resubmission by the user or sales personnel after supplementing the materials. For example, if the materials completeness inspection finds a missing expense list, a missing expense list request is generated, and a request for supplementary documents is sent. The case status is then set to "Materials to be Supplemented."

[0091] When the reason for failure to pass the inspection involves serious anomalies, such as inconsistencies in core identity fields (different names or ID numbers of the insured in different materials), diagnoses that are clearly not covered by the policy, or incidents occurring before the policy took effect without a reasonable explanation, the claim application will be directly returned, the processing will be terminated, and a statement of the reasons for rejection may be attached. For example, if the incident occurred before the policy took effect, it will be determined that the incident was not within the period of insurance coverage, and the case will be directly returned.

[0092] When the reason for the failure of the inspection requires manual judgment, such as doubts about the matching between cost and diagnosis, uncertainty in the semantic comparison results between the accident description and the accident type, or inconsistencies across materials but the differences may be caused by aliases, the claim case will be transferred to the manual review queue with an abnormal information report, including the dimension that failed, the specific value of the abnormal field, and the basis for the abnormal judgment, for the manual reviewers to make further judgments.

[0093] In one specific embodiment, interception operations are categorized based on the severity of the anomaly. For example, missing materials (minor anomaly) trigger a document replacement process; a difference in amount exceeding a threshold (moderate anomaly) triggers manual review; and inconsistent identity fields (severe anomaly) trigger termination of processing. Multiple claims within a short period or frequent incidents at the same hospital are directly classified as high-risk and intercepted. This tiered interception approach allows for differentiated handling strategies for cases of varying risks, ensuring effective interception of high-risk cases while avoiding excessive interception that could decrease the efficiency of processing normal cases.

[0094] By using a risk-prevention mechanism, the identification and interception of high-risk cases are moved from the post-review stage to the material processing stage, avoiding the risk of erroneous automatic review, reducing invalid processing links and manual review costs, and improving the overall efficiency and accuracy of claims review.

[0095] In related technologies, risk alerts typically occur after the review rule engine has completed execution, i.e., when the entire automated review process has run smoothly and generated a conclusion. At this point, erroneous judgments have already been made and may have been output to downstream processes. In contrast, the pre-emptive interception in this application occurs during the material processing stage, executing the interception before the case formally enters the claims review rule engine. This allows high-risk cases to be identified and isolated before entering formal review, preventing the formation and spread of erroneous judgments.

[0096] Alarms from related technologies are mostly based on OCR recognition results for a single material or simple rules, which are prone to generating a large number of false alarms due to issues such as misuse of aliases, inconsistent terminology, and abnormal formatting. The risk interception module of this application receives high-quality structured data that has undergone field correction standardization and cross-material consistency quality checks. Only after completing cross-material field alignment, alias unification, and contradiction detection can an inconsistency be accurately determined as either a genuine fraud risk or merely a difference in data representation. Therefore, the interception conditions of this application have higher accuracy and a significantly lower false alarm rate.

[0097] Alarms from related technologies typically only output a warning message or risk score, still requiring manual judgment to suspend the process. This application's pre-interception module has automatic triage and blocking capabilities, automatically triggering different actions based on the anomaly type and severity: minor anomalies automatically generate a supplementary document notification and suspend the case; moderate anomalies automatically transfer to the manual review queue with an anomaly field location report; severe anomalies directly freeze the case and block all subsequent processes. This significantly reduces the burden of manual review and improves the overall efficiency and automation level of claims processing.

[0098] Please see Figure 3 In one embodiment, the AI-based claims material processing method provided in this application can be as follows: S1: Receive claim materials.

[0099] Receive original materials and related business information uploaded for claims and complete case archiving. Original materials include, but are not limited to, medical records, invoices, expense lists, examination and test reports, accident reports or proof of incident, and the insured's identity information; business information includes policy and liability information, historical claims records, and external query results.

[0100] S2: Filing of claims materials.

[0101] Upon receipt, the materials are assigned a basic number and the case is archived, providing a unified entry point for subsequent multi-stage processing.

[0102] S3: Grouping of claim materials.

[0103] The system identifies the material type and assigns it to the corresponding processing channel. Specifically, based on the material content, file name, format characteristics, and field characteristics of each claim material, the system identifies the material type of each claim material. Claim materials of the same type are grouped together, and a corresponding processing channel is assigned to each group according to a pre-established mapping relationship between groups and processing channels. Each processing channel is independent and can be executed in parallel.

[0104] S4: Extract fields based on grouping of claim materials.

[0105] Key review fields are extracted based on material type to form a structured field set. In each processing channel, based on the key field type corresponding to that channel, the field position matching the key field type is located from each group of claim materials, the corresponding field content is extracted, and assembled into a structured field according to a predefined business field architecture.

[0106] S5: Perform field correction.

[0107] The extracted fields are standardized, aliased, and formatted. Specifically, alias fields that semantically match the extracted fields are obtained. The extracted fields and alias fields are assembled into a first prompt word. The first model is then used to extract the original text value from the claims materials. The original text value is used to retrieve candidate standard fields from the standard knowledge base. The candidate standard fields, the original text value, and the scenario information are then assembled into a second prompt word. The second model is then used to select the optimal standard name from the candidate standard fields as the target extracted field.

[0108] S6: Inspection of claim materials.

[0109] The system verifies the completeness of materials, the authenticity of fields, consistency across materials, liability relevance, and the reasonableness of the chronological order. Specifically, it checks for missing materials, fields, or pages in the completeness dimension; it checks whether fields meet format specifications in the authenticity dimension; it checks whether core fields are consistent across different materials in the consistency dimension; it checks whether the diagnosis falls within the scope of liability, whether the cost matches the diagnosis, and whether the accident report matches the accident type in the liability matching dimension; and it checks whether the time of the accident, the time of medical treatment, the policy effective date, and the time of material issuance meet the reasonable order in the chronological order dimension.

[0110] S7: Determine the handling strategy based on the test results.

[0111] If the abnormal risk conditions are met, the system will output a result of requesting supplementary documents, returning the case, or transferring it to a human review process. If the risk conditions are not met, the system will proceed to the subsequent claims review process. Specifically, the interception level and interception strategy are determined based on the inspection results of each dimension: for minor abnormalities, a supplementary document notification is automatically generated and the case is suspended; for moderate abnormalities, the case is automatically transferred to the human review queue with an abnormal field location report; for severe abnormalities, the case is directly frozen and all subsequent processes are blocked.

[0112] S8: Output the audit results and archive them.

[0113] Output the processing conclusion, and save the field extraction results, correction results, quality inspection results, and risk hit information. The processing conclusion includes, but is not limited to: complete materials allow for automatic review; materials need to be supplemented; field anomalies need to be corrected; risk hit conditions require manual review; responsibility conditions are not met; preliminary review passed, etc. The saved information can be used for manual review or subsequent rule optimization.

[0114] This application employs a phased processing approach—material classification, field extraction, and field correction—to more stably identify and organize claim materials. Completeness, authenticity, relevance, and consistency checks are performed before formal review, providing a more reliable data foundation for subsequent reviews. Risk interception is implemented during the material processing stage, allowing many high-risk cases to be identified and triaged before the final review stage. Because anomaly types, contradictory fields, and risk conditions are structured and identified, human staff do not need to sift through all materials from scratch and can directly address anomalies. Since problems are categorized and triaged during the pre-processing stage, the subsequent formal review process is shorter, more stable, and involves less rework.

[0115] It should be understood that, although Figures 1-3 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figures 1-3 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0116] In one embodiment, such as Figure 4 As shown, an artificial intelligence-based claims material processing device is provided, comprising: an acquisition module 20, a determination module 21, an inspection module 22, and a processing module 23, wherein: The acquisition module 20 is used to acquire and analyze the claim materials of the claim case, group the claim materials according to the material characteristics, and allocate processing channels to each group of claim materials; The determination module 21 is used to extract fields corresponding to key field types from each group of claims materials in each processing channel according to the key field types corresponding to each processing channel, and obtain the extracted fields; and to determine the target extracted fields according to the extracted fields, alias fields and the scenario information of the claims case, wherein the alias fields are semantically matched with the extracted fields but have different expression forms. The inspection module 22 is used to update the claim materials based on the target extracted fields and to inspect the updated claim materials on preset dimensions. Processing module 23 is used to respond to the inspection result as passing, allowing the claim case to enter the claim review process; and to respond to the inspection result as failing, to intercept the claim case to prevent it from entering the claim review process.

[0117] Specific limitations regarding the AI-based claims processing device can be found in the above description of the limitations on AI-based claims processing methods, and will not be repeated here. Each module in the aforementioned AI-based claims processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can invoke and execute the corresponding operations of each module.

[0118] In one embodiment, this application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is able to execute the artificial intelligence-based claims material processing method provided by the above methods.

[0119] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data used in the AI-based claims processing method. The network interface communicates with external terminals via a network connection. When the processor executes the computer program, it implements an AI-based claims processing method.

[0120] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0121] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0122] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0123] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for processing claims materials based on artificial intelligence, characterized in that, include: Acquire and analyze claim materials for claims cases, group the claim materials according to their characteristics, and assign processing channels to each group of claim materials; In each processing channel, based on the key field type corresponding to each processing channel, the fields corresponding to the key field type are extracted from each group of claim materials to obtain the extracted fields. Based on the extracted fields, alias fields, and scenario information of the claims case, the target extracted fields are determined, including: searching for alias fields in a preset alias mapping library that semantically match the extracted fields but have different expressions; assembling the extracted fields and alias fields into a first prompt word according to a preset prompt word template; inputting the first prompt word into a first model; the first model searches for text positions in the claims materials that match the first prompt word based on the first prompt word; determining and outputting the original text value based on the text content of the text position; using the original text value as a query condition, searching in a standard knowledge base; and calculating the difference between the original text value and each standard name in the standard knowledge base. Semantic similarity is assessed by obtaining at least one standard name with the highest semantic similarity as a candidate standard field. The candidate standard field, the original text value, and the application scenario information of the claim materials are then assembled into a second prompt word according to a preset prompt word template. This second prompt word is input into a second model, which calculates the semantic similarity between each candidate standard field and the original text value based on the second prompt word, obtaining a semantic similarity score. A scenario adaptation score for each candidate standard field is determined based on the application scenario information. A comprehensive matching score for each candidate standard field is calculated based on the semantic similarity score and the scenario adaptation score. The candidate standard field with the highest comprehensive matching score is selected as the target extraction field. The claims materials are updated based on the target extracted fields, and the updated claims materials are verified on a preset dimension. If the inspection result is satisfactory, the claim case enters the claim review process; if the inspection result is unsatisfactory, the claim case is blocked from entering the claim review process.

2. The method according to claim 1, characterized in that, The process of acquiring and analyzing claim materials for claims cases, grouping claim materials according to their characteristics, and assigning processing channels to each group of claim materials includes: Extract at least one of the following characteristics from the claim materials: material content, document name, and formatting features; Match the material characteristics with preset material characteristics, and determine the material type of the claim materials based on the matching results; Claim materials are grouped according to their material type, with claims materials of the same type grouped together. Based on the pre-established mapping relationship between groups and processing channels, a corresponding processing channel is assigned to each group; the claim materials of each group are sent to their respective processing channels for processing, and the processing channels are independent of each other.

3. The method according to claim 1, characterized in that, In each processing channel, based on the key field type corresponding to each processing channel, fields corresponding to the key field type are extracted from each group of claim materials. The extracted fields include: Obtain the key field types corresponding to each processing channel, wherein the key field types are pre-configured according to the group corresponding to the material type; In each processing channel, based on the key field type, the target location of the field matching the key field type is located in each group of claim materials, and the claim material content corresponding to the target location is extracted; According to the predefined business field architecture, the extracted claims materials are associated with key field types and assembled into structured fields, which are then used as the extracted fields.

4. The method according to claim 1, characterized in that, The step of updating the claim materials based on the target extracted fields, and verifying the updated claim materials on preset dimensions, includes: Replace the corresponding original extraction field with the target extraction field to obtain the replaced claim material data; The replaced claim materials data are standardized to obtain standardized claim materials; the standardization process includes unifying the format of at least one of the date field, amount field, ID number field and gender field. Determine whether the standardized claim materials have missing fields. If there are missing fields, correct the missing fields according to the context information of the location of the missing fields to obtain updated claim materials. The updated claim materials are examined according to preset dimensions, which include at least one of the following: material completeness, material authenticity, material consistency, liability matching, and time sequence.

5. The method according to claim 4, characterized in that, The verification of the updated claims materials in the preset dimensions includes: In response to preset dimensions, including the completeness of materials, the claims materials for the claims case are compared with the preset materials list to determine if there are any missing materials, and / or the extracted fields are compared with the preset key field types to determine if there are any missing fields, and / or the page numbers of the claims materials are identified, and the identification results are compared with the expected number of pages to determine if the claims materials are missing pages; In response to preset dimensions including material authenticity, the target type extraction field is obtained, and it is determined whether the target type extraction field meets the corresponding format specification requirements; the target type includes at least one of date field, amount field, ID number field and gender field; In response to preset dimensions, including material consistency, select any inspection field to determine whether the target extraction field corresponding to the inspection field is consistent in different groups of claim materials; In response to preset dimensions, including liability matching dimensions, determine whether the diagnosis name in the extracted fields falls within the policy liability scope, and / or determine whether the expense items in the extracted fields match the diagnosis, and / or determine whether the accident description in the extracted fields is consistent with the accident type; In response to preset dimensions, including the time sequence dimension, the system obtains and determines whether the time of the incident, the time of medical treatment, the time of policy effectiveness, and the time of document issuance meet the preset time sequence.

6. The method according to claim 1, characterized in that, The interception process for the aforementioned claims includes: The interception and processing operation for a claim case is determined based on the preset dimension type, the preset dimension quantity, and the inspection results of the preset dimension. The interception and processing operation includes at least one of the following: suspending the processing of the claim case, waiting for supplementary materials, terminating the processing of the claim case, and transferring the claim case to the review process.

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Data generation method and device

    CN117851585A

  • Intelligent financial bookkeeping method and system based on large model

    CN121685174A