Insurance policy data processing method and device, equipment and storage medium

By converting insurance policies into text data and using large language models for matching and screening, the problem of time-consuming and labor-consuming manual operations in insurance policies is solved, and an efficient and accurate automated order recording process is achieved, which improves the overall efficiency and accuracy of insurance business.

CN120494985APending Publication Date: 2025-08-15太保科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510580935.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the entry of handwritten attachment clauses of different types of insurance policies during the insurance policy entry process consumes a lot of manpower and is prone to errors, affecting the efficiency and accuracy of insurance business.

Method used

By obtaining insurance policies and converting them into text data, selecting pages with additional terms, using a large language model for matching and filtering, combining preset terms names and content search database for double matching, ensuring the accuracy and completeness of information association, and finally the large language model for final filtering.

Benefits of technology

The automated process of insurance policy entry is realized, which reduces manual operation delays and errors, improves the efficiency and accuracy of order recording, adapts to the filling habits of different salesmen, and optimizes the insurance business process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494985A_ABST
    Figure CN120494985A_ABST
Patent Text Reader

Abstract

The invention provides an insurance policy data processing method and device, equipment and a storage medium, and relates to the technical field of data processing. A to-be-processed insurance policy is firstly obtained and converted into text data, a page with an additional clause name is selected from the text data, then the additional clause name and corresponding content in the page are extracted, and a first mapping relation is constructed. Through step-by-step extraction and selection, the selection accuracy is ensured, a comprehensive basis is provided for subsequent matching from the two dimensions of name and content, and it is ensured that information association is accurate and complete. Then, the first mapping relation is matched in a preset clause name retrieval library and a preset clause name and content retrieval library, a first matching result and a second matching result are obtained, double matching makes up for errors of a single mode, and matching accuracy and comprehensiveness are improved. And finally, introducing a large language model to screen double matching results. According to the method, manual screening and inputting are converted into an automatic process, the insurance policy inputting delay and errors can be reduced, and the insurance policy inputting efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to an insurance policy data processing method, apparatus, device and storage medium. Background Art

[0002] In the insurance industry, policy entry is a critical step in the business process. After communicating with clients, insurance agents will fill out insurance applications based on their needs. The resulting applications come in a variety of formats, including Word, PDF, scanned copies of printed versions, or mobile phone photos. When entering policies, these various application forms must be accurately entered into the structured policy entry page.

[0003] However, in practice, when entering handwritten addendum clauses for different types of policies, recorders must screen the information from a database of standard additional clauses based on the application form. This process is not only labor-intensive but also prone to errors, severely impacting the efficiency and accuracy of insurance operations.

[0004] Therefore, there is an urgent need for a technical solution that can effectively solve the above problems to improve the quality and efficiency of policy entry, reduce labor costs, reduce error rates, and thus optimize the overall process of insurance business. Summary of the Invention

[0005] To address the above issues, this application provides an insurance policy data processing method, which includes the following contents:

[0006] In a first aspect, the present application provides a method for processing insurance policy data, the method comprising:

[0007] Obtaining a pending insurance policy, and converting the pending insurance policy into text data;

[0008] Selecting a page with a name of an additional clause from the text data;

[0009] Extracting the additional clause name and the content corresponding to the additional clause name from the page with the additional clause name to obtain a first mapping relationship;

[0010] Matching the additional clause name in the first mapping relationship in a preset clause name search library to obtain a first matching result, and matching the additional clause name and corresponding content in the first mapping relationship in a preset clause name and content search library to obtain a second matching result;

[0011] Based on the first matching result and the second matching result, a large language model is used for screening to obtain a final matching result.

[0012] Optionally, obtaining the pending insurance policy and converting the pending insurance policy into text data includes:

[0013] Convert the insurance policy to be processed into image format page by page;

[0014] The image is recognized using an OCR algorithm to obtain text data corresponding to each page of the insurance policy to be processed.

[0015] Optionally, extracting the additional term name and the content corresponding to the additional term name from the page with the additional term name to obtain the first mapping relationship includes:

[0016] Obtain historical insurance policy data of the insurance type corresponding to the to-be-processed insurance policy, determine the number of occurrences of the additional clauses therein, count all additional clauses whose number of occurrences exceeds a preset occurrence threshold, and generate a list of historical common additional clauses;

[0017] Utilizing a large language model in combination with the historical common additional terms list, extracting the additional terms name from the page with the additional terms name;

[0018] Using embedding technology to obtain the semantic vector of the additional clause name, and calculating the similarity, and removing duplicates of similar clause names;

[0019] The large language model is used to extract the content corresponding to the additional clause name, and a mapping relationship between the additional clause name and the content is obtained as the first mapping relationship.

[0020] Optionally, before obtaining a first matching result by matching the additional term name in the first mapping relationship in a preset term name search library and obtaining a second matching result by matching the additional term name and corresponding content in the first mapping relationship in a preset term name and content search library, the method further includes:

[0021] Generalizing and expanding the additional clause names in the additional clause standard database to obtain generalized names, and storing the generalized names as standard names in the quasi-database, wherein the additional clause names in the standard database include the original standard names and the newly stored standard names;

[0022] Using embedding technology to obtain a vector of the additional term name, and constructing a retrieval library containing only the additional term names;

[0023] After concatenating the additional clause name and the corresponding content, embedding technology is used to obtain the vector of the generalized name and the corresponding content, and a retrieval library including the additional clause name and content is constructed.

[0024] Optionally, the first matching result obtained by matching the additional term name in the first mapping relationship in a preset term name search library, and the second matching result obtained by matching the additional term name and corresponding content in the first mapping relationship in a preset term name and content search library include:

[0025] In the preset term name search library, based on the pre-recall of the additional term name in the first mapping relationship, a first matching result is obtained, wherein the matching result also includes a corresponding similarity;

[0026] In the preset clause name and content retrieval library, a pre-recall is performed based on the additional clause name and corresponding content in the first mapping relationship to obtain a second matching result, which also includes the corresponding similarity.

[0027] Optionally, based on the first matching result and the second matching result, a large language model is used for screening to obtain a final matching result including:

[0028] According to a preset first similarity threshold, retaining data in the first matching results that is greater than the first similarity threshold;

[0029] According to a preset second similarity threshold, retaining data in the second matching result that is greater than the second similarity threshold; the first similarity threshold and the second similarity threshold are the same or different;

[0030] Merging the data in the first matching result that is greater than a first similarity threshold and the data in the second matching result that is greater than a second similarity threshold, and removing duplicates to obtain a third matching result;

[0031] Based on the third matching result, the content corresponding to the additional term name in the third matching result is completed according to the additional term standard database to obtain a second mapping relationship, and the second mapping relationship is filtered using a large language model to obtain a final matching result.

[0032] Optionally, based on the third matching result, the content corresponding to the additional term name in the third matching result is supplemented according to the additional term standard database to obtain a second mapping relationship, and the second mapping relationship is filtered using a large language model to obtain a final matching result including:

[0033] Based on the preset insurance clause review regulations and the review process history records, a screening plan for the second mapping relationship is formulated using a large language model;

[0034] According to the screening plan, the second mapping relationship is screened step by step using the large language model to obtain an analysis conclusion for each step;

[0035] Based on the insurance clause review guidelines and the review process history, the large language model is used to review the screening results, determine whether the task has been completed correctly, and obtain the final matching results;

[0036] If the final matching result is no match, the standard additional terms with the highest matching degree in the screening process will be recommended to the order recorder for manual review.

[0037] In a second aspect, the present application provides an insurance policy data processing device, the device comprising:

[0038] An acquiring unit, configured to acquire an insurance policy to be processed and convert the insurance policy to be processed into text data;

[0039] A first screening unit, configured to select pages with additional clause names from the text data;

[0040] A second screening unit is configured to extract the additional term name and the content corresponding to the additional term name from the page with the additional term name, to obtain a first mapping relationship;

[0041] a matching unit configured to match the additional term name in the first mapping relationship in a preset term name search library to obtain a first matching result, and to match the additional term name and corresponding content in the first mapping relationship in a preset term name and content search library to obtain a second matching result;

[0042] The third screening unit is configured to perform screening based on the first matching result and the second matching result using a large language model to obtain a final matching result.

[0043] Optionally, the acquisition unit is specifically configured to convert the to-be-processed insurance policy page by page into a picture format;

[0044] The image is recognized using an OCR algorithm to obtain text data corresponding to each page of the insurance policy to be processed.

[0045] Optionally, the second screening unit is specifically configured to obtain historical insurance policy data of the insurance type corresponding to the to-be-processed insurance policy, determine the number of occurrences of the additional clauses therein, count all additional clauses whose number of occurrences is greater than a preset occurrence threshold, and generate a list of historical common additional clauses;

[0046] Utilizing a large language model in combination with the historical common additional terms list, extracting the additional terms name from the page with the additional terms name;

[0047] Using embedding technology to obtain the semantic vector of the additional clause name, and calculating the similarity, and removing duplicates of similar clause names;

[0048] The large language model is used to extract the content corresponding to the additional clause name, and a mapping relationship between the additional clause name and the content is obtained as the first mapping relationship.

[0049] Optionally, before obtaining a first matching result by matching the additional term name in the first mapping relationship in a preset term name search library and obtaining a second matching result by matching the additional term name and corresponding content in the first mapping relationship in a preset term name and content search library, the method further includes:

[0050] Generalizing and expanding the additional clause names in the additional clause standard database to obtain generalized names, and storing the generalized names as standard names in the quasi-database, wherein the additional clause names in the standard database include the original standard names and the newly stored standard names;

[0051] Using embedding technology to obtain a vector of the additional term name, and constructing a retrieval library containing only the additional term names;

[0052] After concatenating the additional clause name and the corresponding content, embedding technology is used to obtain the vector of the generalized name and the corresponding content, and a retrieval library including the additional clause name and content is constructed.

[0053] Optionally, the matching unit is specifically configured to perform a pre-recall of the additional term name in the first mapping relationship in the preset term name search library to obtain a first matching result, wherein the matching result also includes a corresponding similarity;

[0054] In the preset clause name and content retrieval library, a pre-recall is performed based on the additional clause name and corresponding content in the first mapping relationship to obtain a second matching result, which also includes the corresponding similarity.

[0055] Optionally, based on the first matching result and the second matching result, a large language model is used for screening to obtain a final matching result including:

[0056] According to a preset first similarity threshold, retaining data in the first matching results that is greater than the first similarity threshold;

[0057] According to a preset second similarity threshold, retaining data in the second matching result that is greater than the second similarity threshold; the first similarity threshold and the second similarity threshold are the same or different;

[0058] Merging the data in the first matching result that is greater than a first similarity threshold and the data in the second matching result that is greater than a second similarity threshold, and removing duplicates to obtain a third matching result;

[0059] Based on the third matching result, the content corresponding to the additional term name in the third matching result is completed according to the additional term standard database to obtain a second mapping relationship, and the second mapping relationship is filtered using a large language model to obtain a final matching result.

[0060] Optionally, based on the third matching result, the content corresponding to the additional term name in the third matching result is supplemented according to the additional term standard database to obtain a second mapping relationship, and the second mapping relationship is filtered using a large language model to obtain a final matching result including:

[0061] Based on the preset insurance clause review regulations and the review process history records, a screening plan for the second mapping relationship is formulated using a large language model;

[0062] According to the screening plan, the second mapping relationship is screened step by step using the large language model to obtain an analysis conclusion for each step;

[0063] Based on the insurance clause review guidelines and the review process history, the large language model is used to review the screening results, determine whether the task has been completed correctly, and obtain the final matching results;

[0064] If the final matching result is no match, the standard additional terms with the highest matching degree in the screening process will be recommended to the order recorder for manual review.

[0065] In a third aspect, the present application provides a device comprising a memory and a processor, wherein the memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes the insurance policy data processing method introduced in any implementation of the first aspect.

[0066] In a fourth aspect, the present application provides a computer-readable storage medium storing a code. When the code is executed, the device executing the code implements the insurance policy data processing method described in any implementation of the first aspect.

[0067] The present application provides a method for processing insurance policy data. When executing the method, first obtain a to-be-processed insurance policy, convert the to-be-processed insurance policy into text data, select a page with an additional clause name from the text data, extract the additional clause name and the content corresponding to the additional clause name from the page with the additional clause name, and obtain a first mapping relationship. By gradually extracting the pages with the additional clause name and then selecting the additional clause name from the page with the additional clause name, the accuracy of the selection result is ensured. By respectively extracting the additional clause name and its corresponding content, the first mapping relationship is constructed, providing a comprehensive basis for subsequent matching from the two dimensions of name and content, thereby ensuring the accuracy and completeness of information association.

[0068] Then, the name of the additional clause in the first mapping relationship is matched in the preset clause name search library to obtain a first matching result, and the name of the additional clause in the first mapping relationship and the corresponding content are matched in the preset clause name and content search library to obtain a second matching result. A dual matching is performed using the preset clause name search library and the preset clause name and content search library. The first matching result quickly locates potential standard clauses based on the clause name, and the second matching result is an in-depth comparison of the clause name and the specific content, effectively compensating for the errors that may exist in a single matching method, significantly improving the accuracy and comprehensiveness of the matching. Finally, a large language model is introduced to screen the dual matching results. With its powerful semantic understanding and logical reasoning capabilities, it can accurately judge the rationality of the matching results, effectively resolve matching ambiguities caused by differences in clause wording, non-standard names, and other issues, and obtain the final accurate and reliable matching results. In this way, when recording insurance policies, the originally tedious manual screening and entry work is transformed into an efficient and accurate automated process, reducing delays and errors caused by the policy entry link in the insurance business and improving the efficiency of insurance policy recording. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions in this embodiment or the prior art, the following briefly introduces the drawings required for use in the embodiment or the prior art description. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0070] Figure 1 A flowchart of a method for processing insurance policy data provided in an embodiment of the present application;

[0071] Figure 2 A schematic diagram of the structure of an insurance policy data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0072] In order to make the purpose, technical solutions and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0073] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0074] Figure 1 A flowchart of a method for processing insurance policy data provided in an embodiment of the present application. Figure 1 As shown, the insurance policy data processing method provided in the embodiment of the present application may include:

[0075] S101: Obtain an insurance policy to be processed, and convert the insurance policy to be processed into text data.

[0076] The pending insurance policy is an insurance policy with additional clauses, which are handwritten by the insurance agent. The pending insurance policy can be in various formats, including a Word version, a PDF version, or a scanned copy or mobile phone photo of a printed version.

[0077] After obtaining the pending insurance policy, the policy is processed using a large language model (LLM) to convert it into text data. In one implementation of the embodiment of the present application, obtaining the pending insurance policy and converting it into text data includes: converting the pending insurance policy page by page into an image format; and recognizing the image using an OCR algorithm to obtain text data corresponding to each page of the pending insurance policy.

[0078] For example, for various insurance application documents, we convert them page by page into image format, using {P1, P2, ..., Pn} (where n represents the page number) to represent the conversion result. For insurance application documents originally in image format, we directly organize them and use {P1, P2, ..., Pn} (where n represents the page number) to represent the conversion result. We then use an OCR algorithm to recognize each page of {P1, P2, ..., Pn}, obtaining the text information for each image. The resulting text information is represented by {T1, T2, ..., Tn}.

[0079] S102: Select a page with an additional clause name from the text data.

[0080] Because insurance policies may contain more than one page, and supplementary clauses may only appear on certain pages, accurate extraction of supplementary clauses requires filtering all policies to identify pages with supplementary clause names. Specifically, LLM is used to determine, page by page, whether content Ti represents a policy that may contain supplementary clauses.

[0081] Command example:

[0082]

Your role

[0083] [Task] Determine whether the page content is an insurance policy with additional terms

[0084]

Page Content

[0085] [Answer format] {"Whether additional terms are included":"Include"} or {"Whether additional terms are included":"Not included"};

[0086] Then keep the text containing the additional terms, denoted by {T1,T3,...,Tn-1}.

[0087] S103: Extract the additional clause name and the content corresponding to the additional clause name from the page with the additional clause name to obtain a first mapping relationship.

[0088] The first mapping relationship is the correspondence between the additional clause name and the corresponding content. By using LLM for the text that may contain additional clauses in the pages with additional clause names identified in the previous step, combined with the list of common additional clauses in the history of this type of insurance, the additional clause names are extracted page by page.

[0089] The process for obtaining a list of historically common additional clauses involves obtaining historical policy data for the insurance type corresponding to the policy being processed, determining the number of occurrences of the additional clauses, and then counting all additional clauses with a number of occurrences exceeding a preset threshold to generate a list of historically common additional clauses. For example, if the policy being processed is a motor vehicle insurance policy, historical policy data for all similar motor vehicle policies from the past period is collected. The additional clauses in these collected historical policies are analyzed one by one, and the number of occurrences of each additional clause is counted. For example, if "No Deductible Special Coverage" appears 80 times in 100 historical motor vehicle insurance policies, and "Glass Breakage Coverage" appears 60 times, etc. A preset threshold of occurrence, such as 50, is set. All additional clauses with a number of occurrences exceeding this threshold are counted to form a list. This list is the historically common additional clause list, which indicates which additional clauses are most common and frequently appear under this insurance type.

[0090] LLM instruction example:

[0091]

Your role

[0092] [Task] You need to extract the names of all additional terms based on [Page Content].

[0093] Output format

[0094] 1. Output a list string, where the elements are the names of additional clauses, for example: "["name1","name2",...,"namen"]"

[0095] 2. Make sure the name of the additional terms is exactly the same as in [Page Content]

[0096] 3. If there are no additional terms, output an empty list string, for example: "[]"

[0097] 4. Only output the list string, do not output anything else

[0098]

Page Content

[0099]

Historical common additional clause names

[0100] Document: List of Common Additional Clauses in Insurance History

[0101] Then, using a large language model combined with the historically common additional clause list, the additional clause names are extracted from the pages bearing the additional clause names. Using embedding technology, semantic vectors for the additional clause names are obtained, and similarity is calculated to remove duplicates from similar clause names. Specifically, embedding is used to obtain additional clause name vectors, and pairwise similarity is calculated. Combined with the results of text similarity algorithms such as Jaccard, similar clause names are removed to obtain the deduplicated additional clause name list ["name1", "name3", ..., "namen"].

[0102] Finally, the large language model is used to extract the content corresponding to the additional clause name, and the mapping relationship between the additional clause name and the content is obtained as the first mapping relationship. Specifically, for the additional clause list["name1","name3",...,"namen"] obtained in the previous step, the LLM is used to further extract the additional clause content of the page, and the mapping relationship between the additional clause name and the content is obtained as Map = {"name1":"content1","name2":"content2",...,"namen"}. n ":"content n "}. Example of LLM instruction:

[0103]

Your role

[0104] [Task] needs to extract the corresponding clause content in [Additional Clause Name] based on [Page Content].

[0105] Output format

[0106] 1. Output the dictionary consisting of the name and content of the additional clauses. The content must be the original text describing the additional clauses in [Insurance Policy Content]. For example, {"name1":"Clause Content 1","name2":"Clause Content 2",,...,"name n ":"Terms content"}

[0107] 2. Make sure that the "Terms and Conditions" are exactly the same as the original text in the "Insurance Policy"

[0108] 3. If there is no additional clause, output the corresponding "name i "The judgment result: {"name1":"no"}

[0109]

Page Content

[0110] S104. Match the additional clause name in the first mapping relationship in a preset clause name search library to obtain a first matching result, and match the additional clause name and corresponding content in the first mapping relationship in a preset clause name and content search library to obtain a second matching result.

[0111] After obtaining the additional clause name and the content corresponding to the additional clause name from the insurance policy to be processed, it is necessary to search from the preset additional clause standard database (hereinafter referred to as the "database") to obtain the corresponding additional clause. It should be noted that although the additional clause name and the content corresponding to the additional clause name have been obtained in the above steps, these contents are handwritten on the insurance policy by business personnel, and there may be non-standard expressions, diverse wording, etc. Direct matching by the original name is prone to missed matches or mismatches, so it is necessary to pre-process the database first and build different search libraries to improve matching accuracy.

[0112] Specifically, the additional clause name in the additional clause standard database is generalized and expanded using LLM to obtain the generalized name name_gene i , expand the database and change name_gene i Each name in is also a standard name, that is, name i . In this way, the additional clause names in the standard database include the original standard names and the newly stored standard names. The generalized extension is to expand some similar expressions, synonyms, common abbreviations, etc. into standard names based on the semantics of the clauses. For example, "vehicle theft and robbery insurance" is generalized and expanded to "car theft and robbery insurance", "vehicle theft and robbery insurance", etc. as standard names, and then these new standard names are stored in the database. At this time, the additional clause names in the database include the original standard names and these newly stored generalized standard names. Example of LLM instruction:

[0113]

Your role

[0114] [Task] It is necessary to generalize the [Standard Name of Additional Terms] and provide some abbreviations and common names.

[0115] Output format

[0116] 1. Output list string, where the elements are the names of additional clauses, for example: "["namei_1","namei_2",...,"namei_j"]"

[0117]

Additional clause standard name

[0118] Then two new retrieval libraries need to be built, namely the clause name retrieval library F1 and the clause name and content retrieval library F2.

[0119] To build a clause name search database, we first used embedding technology to convert all additional clause names (including original standard names and generalized standard names) into vector form. Embedding technology converts text information into mathematical vectors that can be processed by computers. The vectors carry the semantic information of the names and are used to match the additional clause names extracted from the insurance policies.

[0120] To establish a retrieval library for clause names and contents, we first concatenate the generalized and expanded additional clause names and corresponding contents in the database to form a complete text. Then, we use embedding technology to convert the concatenated text into a vector for matching the additional clause names and corresponding contents extracted from the insurance policy.

[0121] After completing the above preparations, in the preset clause name search library, based on the pre-recall of the additional clause name in the first mapping relationship, a first matching result is obtained, and the matching result also includes; in the preset clause name and content search library, based on the additional clause name in the first mapping relationship and the corresponding content, a second matching result is obtained, and the matching result also includes the corresponding similarity. Through this dual matching method, the probability of accurately finding the corresponding additional clause from the database is increased. The matching process specifically includes:

[0122] Based on the first mapping relationship, use name in turn i Perform pre-recall in the retrieval library F1 and obtain the first m recall results and their similarities R1 = {"r1_recall1":similarity1,"r1_recall2":similarity2,...,"r1_recallm":similaritym}

[0123] Based on the first mapping relationship, use name in turn i After splicing "Term Content 1" (if it does not exist, do not splice it), perform pre-recall in the retrieval library F2 to obtain the first m recall results and their similarity R2, where R2 = {"r2_recall1":similarity1,"r2_recall2":similarity2,...,"r2_recallm":similaritym}.

[0124] S105: Based on the first matching result and the second matching result, a large language model is used to perform screening to obtain a final matching result.

[0125] According to a preset first similarity threshold, data in the first matching result that is greater than the first similarity threshold is retained; according to a preset second similarity threshold, data in the second matching result that is greater than the second similarity threshold is retained; the first similarity threshold and the second similarity threshold are the same or different; the data in the first matching result that is greater than the first similarity threshold and the data in the second matching result that is greater than the second similarity threshold are merged, and a third matching result is obtained after deduplication; based on the third matching result, the content corresponding to the additional term name in the third matching result is completed according to the additional term standard database to obtain a second mapping relationship, and the second mapping relationship is filtered using a large language model to obtain a final matching result.

[0126] Specifically, for R1, set the similarity threshold Threshold1, retain the values in R1 with a similarity greater than Threshold1 (here it is assumed that there are 2 values that meet the condition), denoted as RT1 = {"r1_recall1":similarity1,"r1_recall2":similarity2}; for R2, set the similarity threshold Threshold2, retain the values in R2 with a similarity greater than Threshold2 (here it is assumed that there is 1 value that meets the condition), denoted as RT2 = {"r2_recall1":similarity1}; merge RT1 and RT2, remove duplicates, and obtain the third matching result, denoted as RT.

[0127] For the clauses in RT, the mapping relationship between the additional clause name and the clause content is obtained from the database R_Map = {"r1_recall1":"r1_recall1 corresponding clause content","r1_recall2":"r1_recall2 corresponding clause content",...,"r2_recall1":"r2_recall1 corresponding clause content"}, where R_Map is the second mapping relationship.

[0128] In one implementation of the embodiment of the present application, based on the third matching result, the content corresponding to the additional term name in the third matching result is supplemented according to the additional term standard database to obtain a second mapping relationship, and the obtained second mapping relationship is filtered using a large language model to obtain a final matching result including:

[0129] Based on the preset insurance clause review regulations and the historical records of the review process, a screening plan for the second mapping relationship is formulated using a large language model.

[0130] Specifically, the large model is used based on the relevant provisions of the insurance clause review, and the clause name is extracted i and terms and content i, formulate a screening plan for the retrieved standard terms R_Map i . LLM instruction example:

[0131]

Your role

[0132] [Task] Based on the Insurance Clauses Review Guidelines and the Review Process History, provide a subsequent screening plan to review the matching of the terms in the Underwriting Clauses Text and the Retrieved Standard Clauses.

[0133] Output format

[0134] 1. Planning steps are numbered starting from 1, with each step separated by a single line break.

[0135] 2. Steps with earlier numbers cannot depend on the results of steps with later numbers

[0136] 3. After completing the last number, the results of all steps can be combined to determine whether the match occurs.

[0137]

Insurance Terms Text

[0138] Clause name: name i

[0139] Terms and Conditions: i

[0140]

Standard Terms Retrieved

[0141] R_Map

[0142] Insurance Clauses Review Guidelines

[0143] Document Content: "Guidelines for Reviewing Insurance Terms"

[0144]

Audit process history

[0145] Record content: Audit history

[0146] According to the screening plan, the second mapping relationship is screened step by step using the large language model to obtain the analysis conclusion of each step. The result action of each step is obtained by executing the screening plan formulated by the large model in the previous step. i . LLM instruction example:

[0147]

Your role

[0148] [Task] You have already analyzed and developed a subsequent screening plan to check the matching between the "underwriting clause text" and the "retrieved standard clauses". Please follow the screening plan and give the analysis conclusions of each step one by one. In the last step, give the most suitable standard clause name.

[0149] [Output format] Only outputs the parseable json format, the key is the step name, the value is the conclusion corresponding to the step, and there is no need to output the analysis process

[0150]

Insurance Terms Text

[0151] Clause name: name i

[0152] Terms and Conditions: i

[0153]

Standard Terms Retrieved

[0154] R_Map

[0155]

Screening Plan

[0156] plan i

[0157] Based on the insurance clause review guidelines and the historical records of the review process, the screening results are reviewed using a large language model to determine whether the task has been completed correctly and obtain the final matching result. If the final matching result is no match, the standard additional clause with the highest matching degree in the screening process will be recommended to the order recorder for manual review.

[0158] LLM instruction example:

[0159]

Your role

[0160] [Task] Please review and determine whether the execution phase tasks are completed correctly according to the `Insurance Clauses Review Guidelines`, `Review Process History` (if any) and the step-by-step execution results.

[0161] Output format

[0162] 1. Only output parseable JSON format, the key contains the status field to indicate whether the execution phase task is completed, the value is 1 for completion, and 0 for incomplete;

[0163] 2. If the review is completed, the key in the result contains the result field to indicate the review conclusion. The value is 1 for a match and 0 for a non-match.

[0164] 3. If the audit is not completed, the result key contains the reason field, and the value is a brief summary of the reason for the audit failure.

[0165]

Insurance Terms Text

[0166] Clause name: name i

[0167] Terms and Conditions: i

[0168]

Standard Terms Retrieved

[0169] R_Map

[0170]

Screening Plan

[0171] plan i

[0172]

Results of executing each step

[0173] action i

[0174] Insurance Clauses Review Guidelines

[0175] Document Content: "Guidelines for Reviewing Insurance Terms"

[0176]

Audit process history

[0177] Record content: Audit history

[0178] If critic i If the result field audit conclusion is 1, then i The last step is to obtain the matching standard additional terms as the final result of the order recorder filling in the system. If there is no match, the plan i The final matching standard additional terms are recommended to the order recorder for manual review. i 、content i Execute the above method separately until the final corresponding standard clause names of all additional clauses are obtained, and the corresponding ID is obtained according to the database to complete the entry of the insurance policy.

[0179] The above embodiment of the present application introduces a method for processing insurance policy data, which significantly improves the efficiency and accuracy of insurance policy entry through automated text conversion, precise extraction of additional clauses and their contents, and matching and screening in combination with a preset search library and a large language model. It can quickly process insurance policies in various formats, reduce manual operations, and lower error rates, while adapting to the filling habits of different salesmen, thereby enhancing the flexibility and adaptability of data processing. In addition, this method provides a high-quality data foundation for the subsequent data analysis and management of insurance companies, optimizes the entire insurance business process, improves customer experience, enhances the competitiveness of insurance companies in the market, and provides strong technical support for the digital transformation and business development of the insurance industry.

[0180] The above are some specific implementations of the insurance policy data processing method provided in the embodiment of the present application. Based on this, the present application also provides a corresponding device. The device provided in the embodiment of the present application will be introduced from the perspective of functional modularization.

[0181] Figure 2 A schematic diagram of the structure of an insurance policy data processing device provided in an embodiment of the present application. Figure 2 As shown, the insurance policy data processing device 200 provided in the embodiment of the present application includes:

[0182] An acquisition unit 210 is configured to acquire an insurance policy to be processed and convert the insurance policy to be processed into text data;

[0183] A first screening unit 220 is configured to select pages with additional clause names from the text data;

[0184] The second screening unit 230 is configured to extract the additional term name and the content corresponding to the additional term name from the page with the additional term name, to obtain a first mapping relationship;

[0185] A matching unit 240 is configured to match the additional term name in the first mapping relationship with a preset term name search library to obtain a first matching result, and to match the additional term name and corresponding content in the first mapping relationship with a preset term name and content search library to obtain a second matching result;

[0186] The third screening unit 250 is configured to perform screening based on the first matching result and the second matching result using a large language model to obtain a final matching result.

[0187] In one implementation of the embodiment of the present application, the acquisition unit is specifically configured to convert the to-be-processed insurance policy page by page into an image format;

[0188] The image is recognized using an OCR algorithm to obtain text data corresponding to each page of the insurance policy to be processed.

[0189] In one implementation of the embodiment of the present application, the second screening unit is specifically configured to obtain historical insurance policy data of the insurance type corresponding to the to-be-processed insurance policy, determine the number of occurrences of the additional clauses therein, count all additional clauses whose number of occurrences exceeds a preset occurrence threshold, and generate a list of historical common additional clauses;

[0190] Utilizing a large language model in combination with the historical common additional terms list, extracting the additional terms name from the page with the additional terms name;

[0191] Using embedding technology to obtain the semantic vector of the additional clause name, and calculating the similarity, and removing duplicates of similar clause names;

[0192] The large language model is used to extract the content corresponding to the additional clause name, and a mapping relationship between the additional clause name and the content is obtained as the first mapping relationship.

[0193] In one implementation of the embodiment of the present application, before obtaining a first matching result by matching the additional term name in the first mapping relationship in a preset term name search library and obtaining a second matching result by matching the additional term name and corresponding content in the first mapping relationship in a preset term name and content search library, the method further includes:

[0194] Generalizing and expanding the additional clause names in the additional clause standard database to obtain generalized names, and storing the generalized names as standard names in the quasi-database, wherein the additional clause names in the standard database include the original standard names and the newly stored standard names;

[0195] Using embedding technology to obtain a vector of the additional term name, and constructing a retrieval library containing only the additional term names;

[0196] After concatenating the additional clause name and the corresponding content, embedding technology is used to obtain the vector of the generalized name and the corresponding content, and a retrieval library including the additional clause name and content is constructed.

[0197] In one implementation of the embodiment of the present application, the matching unit is specifically configured to obtain a first matching result based on the pre-recall of the additional term name in the first mapping relationship in the preset term name search library, wherein the matching result also includes a corresponding similarity;

[0198] In the preset clause name and content retrieval library, a pre-recall is performed based on the additional clause name and corresponding content in the first mapping relationship to obtain a second matching result, which also includes the corresponding similarity.

[0199] In one implementation of the embodiment of the present application, based on the first matching result and the second matching result, a large language model is used for screening to obtain a final matching result including:

[0200] According to a preset first similarity threshold, retaining data in the first matching results that is greater than the first similarity threshold;

[0201] According to a preset second similarity threshold, retaining data in the second matching result that is greater than the second similarity threshold; the first similarity threshold and the second similarity threshold are the same or different;

[0202] Merging the data in the first matching result that is greater than a first similarity threshold and the data in the second matching result that is greater than a second similarity threshold, and removing duplicates to obtain a third matching result;

[0203] Based on the third matching result, the content corresponding to the additional term name in the third matching result is completed according to the additional term standard database to obtain a second mapping relationship, and the second mapping relationship is filtered using a large language model to obtain a final matching result.

[0204] In one implementation of the embodiment of the present application, based on the third matching result, the content corresponding to the additional term name in the third matching result is supplemented according to the additional term standard database to obtain a second mapping relationship, and the obtained second mapping relationship is filtered using a large language model to obtain a final matching result including:

[0205] Based on the preset insurance clause review regulations and the review process history records, a screening plan for the second mapping relationship is formulated using a large language model;

[0206] According to the screening plan, the second mapping relationship is screened step by step using the large language model to obtain an analysis conclusion for each step;

[0207] Based on the insurance clause review guidelines and the review process history, the large language model is used to review the screening results, determine whether the task has been completed correctly, and obtain the final matching results;

[0208] If the final matching result is no match, the standard additional terms with the highest matching degree in the screening process will be recommended to the order recorder for manual review.

[0209] The embodiments of the present application also provide corresponding devices and computer storage media for implementing the solutions provided by the embodiments of the present application.

[0210] The device includes a memory and a processor, the memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes the method described in any embodiment of the present application.

[0211] The computer storage medium stores code, and when the code is executed, the device executing the code implements the method described in any embodiment of the present application.

[0212] Through the description of the above embodiments, it can be known that those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a general hardware platform. Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as a read-only memory (ROM) / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network communication device such as a router) to execute the methods described in each embodiment or certain parts of the embodiments of the present application.

[0213] It is understandable that in the specific implementation of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved, when the above embodiments of this application are applied to specific products or technologies, need to obtain user permission or consent, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.

[0214] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0215] It should also be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The device and apparatus embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components indicated as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0216] The above is merely one specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for processing insurance policy data, characterized in that: The method comprises: Obtaining a pending insurance policy, and converting the pending insurance policy into text data; Selecting a page with a name of an additional clause from the text data; Extracting the additional clause name and the content corresponding to the additional clause name from the page with the additional clause name to obtain a first mapping relationship; Matching the additional clause name in the first mapping relationship in a preset clause name search library to obtain a first matching result, and matching the additional clause name and corresponding content in the first mapping relationship in a preset clause name and content search library to obtain a second matching result; Based on the first matching result and the second matching result, a large language model is used for screening to obtain a final matching result.

2. The method according to claim 1, characterized in that The obtaining of the pending insurance policy and converting the pending insurance policy into text data comprises: Convert the insurance policy to be processed into image format page by page; The image is recognized using an OCR algorithm to obtain text data corresponding to each page of the insurance policy to be processed.

3. The method according to claim 1, characterized in that Extracting the additional clause name and the content corresponding to the additional clause name from the page with the additional clause name, respectively, to obtain a first mapping relationship includes: Obtain historical insurance policy data of the insurance type corresponding to the to-be-processed insurance policy, determine the number of occurrences of the additional clauses therein, count all additional clauses whose number of occurrences exceeds a preset occurrence threshold, and generate a list of historical common additional clauses; Utilizing a large language model in combination with the historical common additional terms list, extracting the additional terms name from the page with the additional terms name; Using embedding technology to obtain the semantic vector of the additional clause name, and calculating the similarity, and removing duplicates of similar clause names; The large language model is used to extract the content corresponding to the additional clause name, and a mapping relationship between the additional clause name and the content is obtained as the first mapping relationship.

4. The method according to claim 1, wherein Before obtaining a first matching result by matching the additional term name in the first mapping relationship in a preset term name search library and obtaining a second matching result by matching the additional term name and corresponding content in the first mapping relationship in a preset term name and content search library, the method further includes: Generalizing and expanding the additional clause names in the additional clause standard database to obtain generalized names, and storing the generalized names as standard names in the quasi-database, wherein the additional clause names in the standard database include the original standard names and the newly stored standard names; Using embedding technology to obtain a vector of the additional term name, and constructing a retrieval library containing only the additional term names; After concatenating the additional clause name and the corresponding content, embedding technology is used to obtain the vector of the generalized name and the corresponding content, and a retrieval library including the additional clause name and content is constructed.

5. The method according to claim 4, characterized in that The first matching result is obtained by matching the additional clause name in the first mapping relationship in the preset clause name search library, and the second matching result is obtained by matching the additional clause name and corresponding content in the first mapping relationship in the preset clause name and content search library, including: In the preset term name search library, based on the pre-recall of the additional term name in the first mapping relationship, a first matching result is obtained, wherein the matching result also includes a corresponding similarity; In the preset clause name and content retrieval library, a pre-recall is performed based on the additional clause name and corresponding content in the first mapping relationship to obtain a second matching result, which also includes the corresponding similarity.

6. The method according to claim 5, characterized in that Based on the first matching result and the second matching result, a large language model is used for screening to obtain a final matching result including: According to a preset first similarity threshold, retaining data in the first matching results that is greater than the first similarity threshold; According to a preset second similarity threshold, retaining data in the second matching result that is greater than the second similarity threshold; the first similarity threshold and the second similarity threshold are the same or different; Merging the data in the first matching result that is greater than a first similarity threshold and the data in the second matching result that is greater than a second similarity threshold, and removing duplicates to obtain a third matching result; Based on the third matching result, the content corresponding to the additional term name in the third matching result is completed according to the additional term standard database to obtain a second mapping relationship, and the second mapping relationship is filtered using a large language model to obtain a final matching result.

7. The method according to claim 6, characterized in that Based on the third matching result, the content corresponding to the additional term name in the third matching result is supplemented according to the additional term standard database to obtain a second mapping relationship, and the second mapping relationship is filtered using a large language model to obtain a final matching result including: Based on the preset insurance clause review regulations and the review process history records, a screening plan for the second mapping relationship is formulated using a large language model; According to the screening plan, the second mapping relationship is screened step by step using the large language model to obtain an analysis conclusion for each step; Based on the insurance clause review guidelines and the review process history, the large language model is used to review the screening results, determine whether the task has been completed correctly, and obtain the final matching results; If the final matching result is no match, the standard additional terms with the highest matching degree in the screening process will be recommended to the order recorder for manual review.

8. An insurance policy data processing device, characterized in that: The device comprises: An acquiring unit, configured to acquire an insurance policy to be processed and convert the insurance policy to be processed into text data; A first screening unit, configured to select pages with additional clause names from the text data; A second screening unit is configured to extract the additional term name and the content corresponding to the additional term name from the page with the additional term name, to obtain a first mapping relationship; a matching unit configured to match the additional term name in the first mapping relationship in a preset term name search library to obtain a first matching result, and to match the additional term name and corresponding content in the first mapping relationship in a preset term name and content search library to obtain a second matching result; The third screening unit is configured to perform screening based on the first matching result and the second matching result using a large language model to obtain a final matching result.

9. A computing device, characterized in that The computing device includes: a memory and a processor; The memory is used to store computer programs; The processor is configured to implement the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.