Insurance claim settlement method and device and electronic equipment
Through a multimodal large model, a fusion vector is generated using visual and text encoders, and a neural network is used for inference, solving the problems of low efficiency and low accuracy in the traditional insurance claim process, and achieving efficient and accurate claims processing.
Patent Information
- Application Number
- CN202510269229.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-08
AI Technical Summary
There are problems in the traditional insurance claims process with many links, slow processing speed, low efficiency and accumulated information errors, resulting in low accuracy.
The multimodal large model is used to process insurance claims, and the prompt text of the claims rules is obtained by obtaining pictures of the materials to be claimed and the prompt text of the claims rules. The visual encoder and text encoder are used to generate a fusion vector, and the neural network of the attention mechanism is combined for thinking chain reasoning to generate a claim conclusion that meets the requirements of the prompt text.
It improves the processing efficiency and accuracy of insurance claims, reduces error accumulation, reduces manual review costs, supports the integration of multimodal information and improves information coverage.
Smart Images

Figure CN120278823A_ABST
Abstract
Description
Technical Field
[0001] This document belongs to the technical field of data processing, and specifically relates to an insurance claim settlement method, device, and electronic device. Background Art
[0002] With the development of medical informatization and insurance claim settlement, the digital informatization processing of claim settlement processes has become increasingly important. Traditional insurance claim settlement processes are usually split into multiple links: medical image OCR (Optical Character Recognition), medical entity recognition and standardization, relationship extraction of medical information, and rule parsing of claim liability. There are many intermediate links in this process, slow processing speed, low efficiency, and it will lead to the accumulation of information errors and low accuracy. Summary of the Invention
[0003] Embodiments of this specification provide an insurance claim settlement method, device, and electronic device to improve the processing efficiency and accuracy of insurance claim settlement.
[0004] In a first aspect, embodiments of this specification provide an insurance claim settlement method, including: Obtain a picture of the material to be claimed and a prompt text of the claim settlement rule, where the prompt text requires outputting a claim settlement conclusion according to the liability clause; Input the picture and the prompt text into a multi-modal large model, and the multi-modal large model generates a fusion vector based on the picture and the prompt text, and performs chain-of-thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt text.
[0005] In a second aspect, embodiments of this specification provide an insurance claim settlement device, including: An obtaining module, which obtains a picture of the material to be claimed and a prompt text of the claim settlement rule, where the prompt text requires outputting a claim settlement conclusion according to the liability clause; A multi-modal large model, which receives the picture and the prompt text, generates a fusion vector based on the picture and the prompt text, and performs chain-of-thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt text.
[0006] In a third aspect, embodiments of this specification provide an electronic device, including: a processor, and a memory arranged to store computer-executable instructions, which when the executable instructions are executed, can enable the processor to: Obtain a picture of the material to be claimed and a prompt text of the claim settlement rule, where the prompt text requires outputting a claim settlement conclusion according to the liability clause; Input the picture and the prompt text into a multi-modal large model. The multi-modal large model generates a fusion vector based on the picture and the prompt text, and performs chain-of-thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt text.
[0007] Fourthly, an embodiment of this specification provides a storage medium for storing a computer program, and the computer program can be executed by a processor to implement the following process: Obtain a picture of the material to be claimed and the prompt text of the claim settlement rules, and the prompt text requires outputting a claim settlement conclusion according to the liability clause; Input the picture and the prompt text into a multi-modal large model. The multi-modal large model generates a fusion vector based on the picture and the prompt text, and performs chain-of-thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt text.
[0008] Fifthly, an embodiment of this specification provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the following process: Obtain a picture of the material to be claimed and the prompt text of the claim settlement rules, and the prompt text requires outputting a claim settlement conclusion according to the liability clause; Input the picture and the prompt text into a multi-modal large model. The multi-modal large model generates a fusion vector based on the picture and the prompt text, and performs chain-of-thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt text. Description of the Drawings
[0009] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in one or more embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0010] Figure 1 It is a schematic flowchart of an insurance claim settlement method provided by an embodiment of this specification; Figure 2 It is a schematic flowchart of another insurance claim settlement method provided by an embodiment of this specification; Figure 3 It is a schematic diagram of the processing process of a multi-modal large model provided by an embodiment of this specification.
[0011] Figure 4 It is a schematic structural diagram of an insurance claim settlement device provided by an embodiment of this specification; Figure 5It is a schematic structural diagram of an electronic device provided by an embodiment of this specification. Specific implementation manners
[0012] Next, the technical solutions in the embodiments of this specification will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are some but not all of the embodiments of this specification. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this document.
[0013] The terms "first", "second", etc. in the specification and claims of this document are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this specification can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.
[0014] Next, in conjunction with the accompanying drawings, a detailed description will be given to the insurance claim settlement method, device, and electronic device provided by the embodiments of this specification through specific embodiments and their application scenarios.
[0015] The insurance claim settlement method and device provided by the embodiments of this specification can be applied to any electronic device, including but not limited to: terminals or servers, etc. Among them, the terminal device can be a mobile terminal device such as a mobile phone or a tablet computer, or can also be a computer device such as a laptop computer or a desktop computer, or can also be an IoT device (specifically such as a smart watch, a vehicle-mounted device, etc.). The server can be an independent server or a server cluster composed of multiple servers, etc.
[0016] Figure 1 An insurance claim settlement method provided by an embodiment of the present invention is shown, and the method includes the following steps: Step S102: Obtain the picture of the material to be claimed and the prompt text of the claim settlement rule, and the prompt text is required to output the claim settlement conclusion according to the liability clause.
[0017] Among them, the liability clause refers to the clause in the claim settlement rule that restricts and stipulates the user and can be used to determine whether the user meets the claim settlement standard. The claim settlement rule can include one or more liability clauses, which are not specifically limited. The claim settlement involved in the embodiments of the present invention can be claim settlements in multiple fields, such as medical treatment, vehicle insurance, etc., which are not specifically limited.
[0018] The prompt text for the claim settlement rules refers to the prompt words that guide the output of the large model. It describes the claim settlement rules and can be refined from the claim settlement rules. For example, in the medical claim settlement scenario, the prompt text includes descriptions of pre-existing disease exemptions and some exempted diseases, etc., and the specific content is not limited. In addition, through the prompt text, the format, order, or other requirements of the large model output can also be limited, such as outputting the claim settlement conclusion according to the liability clause, etc.
[0019] Step S104: Input the picture and the prompt text into the multi-modal large model. The multi-modal large model generates a fusion vector based on the picture and the prompt text, and performs a chain of thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt text.
[0020] Among them, the multi-modal large model refers to an artificial intelligence model that can process and integrate multiple types of data. The multiple types of data include at least one of the following: text, images, etc.
[0021] Chain of Thought (CoT) reasoning refers to the thinking path and method of the large model during the generation process. By gradually decomposing complex problems into a series of simple sub-problems, and then solving these sub-problems one by one, the final answer is obtained. This method simulates the process of human thinking to solve problems through step-by-step reasoning, rather than directly generating the final answer, and is particularly suitable for tasks that require complex reasoning, multi-step derivation, or logical analysis, such as logical reasoning, mathematical calculations, and the solution of complex problems, etc.
[0022] In one implementation, the multi-modal large model generating a fusion vector based on the picture and the prompt text in the above step S104 may include: The multi-modal large model generates a visual vector based on the picture, the multi-modal large model generates a text vector based on the prompt text, and the multi-modal large model fuses the visual vector and the text vector to obtain a fusion vector.
[0023] In one implementation, the multi-modal large model generating a visual vector based on the picture may include: The visual encoder in the multi-modal large model extracts feature information from the picture and generates a visual vector containing the feature information.
[0024] In one implementation, the process of the visual encoder generating a visual vector may further include: The visual encoder controls the length of the visual vector according to the resolution of the picture.
[0025] In one implementation, the multi-modal large model generating a text vector based on the prompt text may include: The text encoder in the multi-modal large model looks up the corresponding index in the vocabulary according to the prompt text, and encodes the index to obtain a text vector containing text semantic information.
[0026] In one implementation, the above multi-modal large model fuses the visual vector and the text vector to obtain a fused vector, which may include: The neural network architecture based on the attention mechanism in the multi-modal large model fuses the visual vector and the text vector to obtain a fused vector.
[0027] In one implementation, the above step of performing chain-of-thought reasoning based on the fused vector in step S104 to obtain a claim settlement conclusion that meets the requirements of the prompt text may include: The decoder of the multi-modal large model decodes and performs chain-of-thought reasoning on the fused vector in the way of word prediction to generate a claim settlement conclusion that meets the requirements of the prompt text.
[0028] Among them, word prediction refers to predicting the next word of the current word according to the probability distribution of words appearing in the claim settlement scenario.
[0029] In one implementation, the decoder of the above multi-modal large model decodes and performs chain-of-thought reasoning on the fused vector in the way of word prediction to generate a claim settlement conclusion that meets the requirements of the prompt text, which may include: The decoder of the multi-modal large model decodes the fused vector in the way of word prediction and extracts the key information of the claim settlement; Based on the key information, perform chain-of-thought reasoning respectively according to each liability clause of the claim settlement rules to obtain the reasoning results of each liability clause; According to the reasoning results of each liability clause, determine the claim settlement conclusion according to the claim settlement rules and output the claim settlement conclusion in a manner that meets the requirements of the prompt text.
[0030] The above method provided by the embodiments of this specification, by obtaining the picture of the material to be claimed and the prompt text of the claim settlement rules, which requires outputting the claim settlement conclusion according to the liability clause, inputting the picture and the prompt text into the multi-modal large model, the multi-modal large model generates a fused vector according to the picture and the prompt text, and performs chain-of-thought reasoning based on the fused vector to obtain a claim settlement conclusion that meets the requirements of the prompt text, can improve the processing efficiency and accuracy of insurance claim settlement.
[0031] Figure 2 Another embodiment of the present invention provides an insurance claim settlement method, which includes the following steps: Step S202: Obtain the picture of the material to be claimed and the prompt text of the claim settlement rules, and the prompt text requires outputting the claim settlement conclusion according to the liability clause.
[0032] Among them, pictures of materials to be claimed can be uploaded by users, such as users uploading multiple pictures of medical information, including hospitalization records, medical record front pages, payment invoices and other material information. The prompt text of the claim rules refers to the prompt word prompt that guides the output of the large model. It describes the claim rules and can be extracted from the claim rules. For example, the prompt text in the medical claim scenario includes the description of the exemption of previous illnesses and some exempted diseases, etc. The specific content is not limited. In addition, the format, order or other requirements of the output content of the large model can also be limited through the prompt text, such as outputting the claim conclusion according to the liability clause.
[0033] Step S204: Input the image and the prompt text into the multimodal large model, and the visual encoder in the multimodal large model extracts feature information from the image to generate a visual vector containing the feature information.
[0034] The input of the visual encoder is a picture, and the output is a visual vector. Through deep learning technology, the visual encoder can identify and extract key visual information in the picture. The visual vector contains the feature information of the picture and can represent the key information in the claim materials.
[0035] In one implementation manner, the process of the visual encoder generating a visual vector may further include: The visual encoder controls the length of the visual vector according to the resolution of the image.
[0036] Among them, the length of the visual vector corresponding to the high-resolution image is greater than the length of the visual vector corresponding to the low-resolution image.
[0037] Specifically, for high-resolution images, the visual encoder can control the length of the visual vector to be longer, so that the visual vector can include more image information, especially text-intensive images, which can minimize the loss of text information, thereby helping large multimodal models to improve the accuracy of reasoning.
[0038] For low-resolution images, the visual encoder can control the length of the visual vector to be lower, thereby saving more resources and helping to improve the reasoning efficiency of large multimodal models.
[0039] The above method enables the visual encoder to adapt to images of different resolutions. This dynamic resolution method has flexible control and takes into account the accuracy and reasoning efficiency of large multimodal models.
[0040] Step S206: The text encoder in the multimodal large model searches for a corresponding index in the vocabulary according to the prompt text, and encodes the index to obtain a text vector containing text semantic information.
[0041] Among them, the input of the text encoder is text, and the output is a text vector. The text vector contains the semantic information of the prompt text and can represent the key information in the prompt text.
[0042] The above-mentioned vocabulary table includes the mapping relationship between text and index. The text can be text in various fields, such as medical or car insurance, etc. The index refers to the numerical value that can uniquely identify the text and is used for quantization coding. The text encoder can search the vocabulary table according to the content in the prompt text and find the index corresponding to the prompt text. For example, if the prompt text is a paragraph, the text encoder can first perform word segmentation, then use each segmented word to search the vocabulary table to obtain the index corresponding to each segmented word, and then perform quantization coding on each obtained index respectively to obtain the text vector.
[0043] Step S208: The neural network architecture based on the attention mechanism in the multi-modal large model fuses the visual vector and the text vector to obtain a fused vector.
[0044] Among them, the neural network architecture based on the attention mechanism refers to the Transformer, which allows the model to compare each element in the input sequence with other elements when processing the sequence, so as to correctly process each element in different contexts. The Transformer has a wide range of applications in natural language processing, such as machine translation, text summarization, language generation and other fields. Due to the self-attention mechanism, the Transformer can capture the global dependencies in the input sequence, rather than relying only on local information, can process all elements in the input sequence in parallel, greatly improving the computational efficiency, and can better process long sequence data, and is not prone to problems such as gradient disappearance or explosion.
[0045] Among them, the Transformer can undergo multiple layers of mutual learning and perception to fully fuse and learn the visual vector and the text vector, and finally obtain the fused vector.
[0046] Step S210: The decoder of the multi-modal large model decodes and performs chain-of-thought reasoning on the fused vector in the way of word prediction to generate a claim settlement conclusion that meets the requirements of the prompt text.
[0047] Among them, the above-mentioned word prediction refers to predicting the next word of the current word according to the probability distribution of the words appearing in the claim settlement scenario. For example, in the medical claim settlement scenario, if the current word is "A", then the probability of the next word "thyroid" of the current word predicted by word prediction is much higher than the probability of "B" appearing, because in the medical field, thyroid is a high-frequency vocabulary, while A and B are low-frequency vocabularies.
[0048] During the decoding process, word prediction is adopted to predict the subsequent text that better conforms to the claim settlement scenario, which can improve the accuracy of multi-modal large model reasoning and the accuracy of claim settlement conclusions, so that claim settlement can be carried out fully and reasonably.
[0049] In one implementation, the above step S210 may include: The decoder of the multi-modal large model decodes the fusion vector in the way of word prediction to extract the key information of claim settlement. Based on the key information, chain-of-thought reasoning is respectively carried out according to each liability clause of the claim settlement rules to obtain the reasoning results of each liability clause. According to the reasoning results of each liability clause, the claim settlement conclusion is determined according to the claim settlement rules, and the claim settlement conclusion is output in a manner that meets the requirements of the prompt text.
[0050] Among them, the above key information for claim settlement includes the necessary information in the claim settlement materials to determine whether claim settlement can be carried out. For example, in the medical claim settlement scenario, the key information may include the admission time, treatment drugs, diagnosis certificate, payment invoice, etc. Based on these key information, it can help the multi-modal large model judge whether the user meets each claim settlement liability and facilitate the model to carry out reasoning.
[0051] The above chain-of-thought reasoning according to each liability clause of the claim settlement rules can adopt parallel or serial processing methods, and there is no specific limitation. In the case of serial processing, the reasoning of each liability clause does not need to be in a specific order, and the respective reasoning results can be obtained.
[0052] For example, the claim settlement rules include 3 liability clauses: general exemption, pre-existing condition exemption, and health declaration. Chain-of-thought reasoning can be respectively carried out on these 3 liability clauses based on the key information. General exemptions such as dental and maternity are general exemption cases, pre-existing condition exemptions such as diseases occurring before insurance coverage are pre-existing condition exemptions, and health declarations such as the obligation to inform of pre-existing diseases before insurance. According to the key information, it is respectively judged whether the user meets each liability clause, and the respective reasoning results are obtained, that is, agreeing to claim settlement or rejecting claim settlement. If the information submitted by the user is dental disease information, the reasoning result of the general exemption is to reject claim settlement. If the information submitted by the user is the treatment information of thyroid tumor, but the onset date of the disease occurred before insurance coverage, the reasoning result of the pre-existing condition exemption is to reject claim settlement. If the user has informed the detailed information about whether they have been ill within two years before insurance, the reasoning result of the health declaration is to agree to claim settlement.
[0053] After obtaining the inference results of each liability clause, the final claim settlement conclusion can be determined according to the claim settlement rules. Usually, the claim settlement rules stipulate how to determine the final claim settlement conclusion based on the results of multiple liability clauses. For example, if there is one or more cases of rejecting the claim among the results of multiple liability clauses, the final claim settlement conclusion is to reject the claim. If the results of multiple liability clauses are all to approve the claim, the final claim settlement conclusion is to approve the claim.
[0054] In addition, since the prompt text requires outputting the claim settlement conclusion according to the liability clause, after obtaining the final claim settlement conclusion, output the claim settlement conclusion in the manner required by the prompt text. For example, the manner required by the prompt text is to output in three major sections: key information, inference results of liability clauses, and claim settlement conclusion. Among them, the inference results of liability clauses are output in multiple small sections for the inference results of each liability clause. The content output in the above manner includes in sequence: key information, inference results of each liability clause, and claim settlement conclusion.
[0055] Figure 3 The schematic diagram of the processing process of the multi-modal large model provided by another embodiment of the present invention is shown. Among them, the multi-modal large model includes a visual encoder, a text encoder, transformers, and a decoder. The materials uploaded by the user are pictures, which are input into the multi-modal large model and processed by the visual encoder to generate visual vectors. The claim settlement rules of the products insured by the user are input into the multi-modal large model in the form of prompt text and processed by the text encoder to generate text vectors. Transformers fuse the visual vectors and text vectors to obtain fused vectors. The decoder decodes and performs chain-of-thought reasoning on the fused vectors in the way of word prediction to generate a claim settlement conclusion that meets the requirements of the prompt text. The claim settlement conclusion may include 5 contents: key information, inference result of liability clause 1, inference result of liability clause 2, inference result of liability clause 3, and claim settlement conclusion.
[0056] The above method provided by the embodiments of this specification obtains the pictures of the materials to be claimed and the prompt text of the claim rules. The prompt text requires outputting the claim conclusion according to the liability clause. The pictures and the prompt text are input into the multi-modal large model. The visual encoder in the multi-modal large model extracts feature information from the pictures to generate visual vectors containing the feature information. The text encoder looks up the corresponding indexes in the vocabulary according to the prompt text and encodes the indexes to obtain text vectors containing the text semantic information. The neural network architecture based on the attention mechanism fuses the visual vectors and the text vectors to obtain the fused vectors. The decoder decodes and performs chain-of-thought reasoning on the fused vectors in the way of word prediction to generate the claim conclusion that meets the requirements of the prompt text. Using a multi-modal large model realizes end-to-end insurance claims. Inputting the pictures and the prompt text can directly obtain the claim conclusion without intermediate links, greatly improving the claim processing speed and response speed, avoiding error accumulation, improving the accuracy rate, and reducing the cost of manual review. Moreover, it can support the pictures of the materials to be claimed, improving the information coverage rate, and the fusion method of multi-modal information also improves the accuracy rate of claim reasoning.
[0057] Figure 4 is a schematic structural diagram of an insurance claim device according to an embodiment of the present invention. As Figure 4 shown, the device includes: An acquisition module 401 that acquires pictures of the materials to be claimed and the prompt text of the claim rules, where the prompt text requires outputting the claim conclusion according to the liability clause; A multi-modal large model 402 that receives the pictures and the prompt text, generates a fused vector according to the pictures and the prompt text, and performs chain-of-thought reasoning based on the fused vector to obtain the claim conclusion that meets the requirements of the prompt text.
[0058] In an implementation manner, the multi-modal large model 402 generates a fused vector according to the pictures and the prompt text, including: Generating visual vectors according to the pictures, generating text vectors according to the prompt text, and fusing the visual vectors and the text vectors to obtain the fused vectors.
[0059] In an implementation manner, the multi-modal large model 402 includes: A visual encoder that extracts feature information from the pictures and generates visual vectors containing the feature information.
[0060] Among them, the visual encoder can also control the length of the visual vectors according to the resolution of the pictures.
[0061] In an implementation manner, the multi-modal large model 402 includes: A text encoder that looks up the corresponding indexes in the vocabulary according to the prompt text and encodes the indexes to obtain text vectors containing the text semantic information.
[0062] In one implementation, the multimodal large model 402 includes: A neural network architecture based on the attention mechanism that fuses visual vectors and text vectors to obtain a fused vector.
[0063] In one implementation, the multimodal large model 402 includes: A decoder that decodes the fused vector and performs chain-of-thought reasoning in the form of word prediction to generate a claim settlement conclusion that meets the requirements of the prompt text; where word prediction refers to predicting the next word of the current character according to the probability distribution of words appearing in the claim settlement scenario.
[0064] Among them, the decoder can decode the fused vector in the form of word prediction to extract the key information of the claim; based on the key information, perform chain-of-thought reasoning for each liability clause according to the claim settlement rules to obtain the reasoning results of each liability clause; according to the reasoning results of each liability clause, determine the claim settlement conclusion according to the claim settlement rules and output the claim settlement conclusion in a manner that meets the requirements of the prompt text.
[0065] The above device provided by the embodiments of this specification can execute the method provided by any of the above method embodiments. For the detailed process, refer to the description in the method embodiments and will not be elaborated here.
[0066] The above device provided by the embodiments of this specification, by obtaining the picture of the materials to be claimed and the prompt text of the claim settlement rules, which requires outputting the claim settlement conclusion according to the liability clauses, inputs the picture and the prompt text into the multimodal large model. The multimodal large model generates a fused vector based on the picture and the prompt text, and performs chain-of-thought reasoning based on the fused vector to obtain a claim settlement conclusion that meets the requirements of the prompt text, which can improve the processing efficiency and accuracy of insurance claim settlement.
[0067] It should be noted that the embodiments of the insurance claim settlement device in this specification and the embodiments of the insurance claim settlement method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the corresponding insurance claim settlement method described above, and the repeated parts will not be elaborated.
[0068] Each module in the above insurance claim settlement device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the terminal device or the processor in the server in the form of hardware or be independent of them, or can be stored in the memory in the terminal device or the memory in the server in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0069] Furthermore, corresponding to the above-described insurance claim settlement method, based on the same technical concept, one or more embodiments of this specification also provide an electronic device for executing the above insurance claim settlement method.Figure 5 A schematic structural diagram of an electronic device provided for one or more embodiments of this specification.
[0070] Based on the same concept, one or more embodiments of this specification also provide an electronic device, as Figure 5 shown. The electronic device can vary significantly due to configuration or performance differences, and may include one or more processors 501 and a memory 502. One or more application programs or data can be stored in the memory 502. Among them, the memory 502 can be short-term storage or persistent storage. The application programs stored in the memory 502 can include one or more modules (not shown in the figure), and each module can include a series of computer-executable instructions for the electronic device. Further, the processor 501 can be configured to communicate with the memory 502 and execute a series of computer-executable instructions in the memory 502 on the electronic device. The electronic device can also include one or more power supplies 503, one or more wired or wireless network interfaces 504, one or more input / output interfaces 505, and one or more keyboards 506.
[0071] In a specific embodiment, the electronic device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs can include one or more modules. Each module can include a series of computer-executable instructions for the electronic device and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions for: Obtain pictures of materials to be claimed and prompt text of claim rules. The prompt text requires outputting a claim conclusion according to liability clauses; input the pictures and prompt text into a multi-modal large model. The multi-modal large model generates a fusion vector based on the pictures and prompt text, and performs chain-of-thought reasoning based on the fusion vector to obtain a claim conclusion that meets the requirements of the prompt text.
[0072] It should be noted that the embodiments of the electronic device in this specification and the embodiments of the insurance claim settlement method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the corresponding implementation of the aforementioned insurance claim settlement method, and the repeated parts will not be elaborated.
[0073] Furthermore, corresponding to the above-described insurance claim settlement method, based on the same technical concept, one or more embodiments of this specification also provide a storage medium for storing computer-executable instructions. In a specific embodiment, the storage medium can be a USB flash drive, an optical disc, a hard disk, etc. When the computer-executable instructions stored in the storage medium are executed by a processor, the following process can be implemented: Obtain the pictures of the materials to be claimed and the prompt text of the claim settlement rules. The prompt text is required to output the claim settlement conclusion according to the liability clauses; input the pictures and the prompt text into a multi-modal large model. The multi-modal large model generates a fusion vector based on the pictures and the prompt text, and performs chain-of-thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt text.
[0074] It should be noted that the embodiments of the storage medium in this specification and the insurance claim settlement method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding insurance claim settlement method described above, and the repeated parts will not be elaborated.
[0075] Furthermore, corresponding to the above-described insurance claim settlement method, based on the same technical concept, one or more embodiments of this specification also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it can implement the following processes: Obtain the pictures of the materials to be claimed and the prompt text of the claim settlement rules. The prompt text is required to output the claim settlement conclusion according to the liability clauses; input the pictures and the prompt text into a multi-modal large model. The multi-modal large model generates a fusion vector based on the pictures and the prompt text, and performs chain-of-thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt text.
[0076] It should be noted that the embodiments of the computer program product in this specification and the embodiments of the insurance claim settlement method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding insurance claim settlement method described above, and the repeated parts will not be elaborated.
[0077] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0078] In the 1990s, it was obvious to distinguish whether an improvement to a technology was an improvement in hardware (e.g., improvement to the circuit structure of diodes, transistors, switches, etc.) or an improvement in software (improvement to the method flow). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to the hardware circuit structure. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. The designer can program by himself to "integrate" a digital system on a piece of PLD without asking the chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL). And there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that as long as the method flow is slightly logically programmed with the above-mentioned several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit implementing the logical method flow.
[0079] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0080] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0081] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0082] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0083] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.
[0084] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.
[0085] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.
[0086] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0087] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0088] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0089] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0090] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0091] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0092] The above are only examples of this document and are not intended to limit this document. For those skilled in the art, this document may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this document shall be included within the scope of the claims of this document.
Claims
1. An insurance claim settlement method, comprising: Obtaining pictures of materials to be claimed and prompt texts of claim settlement rules, where the prompt texts require outputting claim settlement conclusions according to liability clauses; Inputting the pictures and prompt texts into a multi-modal large model, and the multi-modal large model generates a fusion vector based on the pictures and prompt texts, and performs chain-of-thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt texts.
2. The method according to claim 1, wherein the multi-modal large model generates a fusion vector based on the pictures and prompt texts, comprising: The multi-modal large model generates a visual vector based on the pictures; The multi-modal large model generates a text vector based on the prompt texts; The multi-modal large model fuses the visual vector and the text vector to obtain a fusion vector.
3. The method according to claim 2, wherein the multi-modal large model generates a visual vector based on the pictures, comprising: The visual encoder in the multi-modal large model extracts feature information from the pictures and generates a visual vector containing the feature information.
4. The method according to claim 3, further comprising: The visual encoder controls the length of the visual vector according to the resolution of the pictures.
5. The method according to claim 2, wherein the multi-modal large model generates a text vector based on the prompt texts, comprising: The text encoder in the multi-modal large model looks up corresponding indexes in the vocabulary according to the prompt texts, and encodes the indexes to obtain a text vector containing text semantic information.
6. The method according to claim 2, wherein the multi-modal large model fuses the visual vector and the text vector to obtain a fusion vector, comprising: The neural network architecture based on the attention mechanism in the multi-modal large model fuses the visual vector and the text vector to obtain a fusion vector.
7. The method according to any one of claims 1-6, wherein performing chain-of-thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt texts, comprising: The decoder of the multi-modal large model decodes and performs chain-of-thought reasoning on the fusion vector in the way of word prediction to generate a claim settlement conclusion that meets the requirements of the prompt texts; Wherein, the word prediction refers to predicting the next word of the current word according to the probability distribution of words appearing in the claim settlement scenario.
8. The method according to claim 7, wherein the decoder of the multi-modal large model decodes and performs chain-of-thought reasoning on the fusion vector in the way of word prediction to generate a claim settlement conclusion that meets the requirements of the prompt texts, comprising: The decoder of the multi-modal large model decodes the fusion vector in the way of word prediction and extracts the key information of the claim; Performing chain-of-thought reasoning respectively according to each liability clause of the claim settlement rules based on the key information to obtain the reasoning results of each liability clause; According to the reasoning results of each liability clause, determining a claim settlement conclusion according to the claim settlement rules and outputting the claim settlement conclusion in a manner that meets the requirements of the prompt texts.
9. An insurance claim settlement device, comprising: An acquisition module that acquires images of materials to be claimed and prompt texts of claim settlement rules, where the prompt texts require claim settlement conclusions to be output according to liability clauses; A multimodal large model that receives the images and prompt texts, generates a fusion vector based on the images and prompt texts, and performs chain-of-thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt texts.
10. An electronic device, comprising: A processor, and A memory arranged to store computer-executable instructions, which, when the executable instructions are executed, can cause the processor to: Acquire images of materials to be claimed and prompt texts of claim settlement rules, where the prompt texts require claim settlement conclusions to be output according to liability clauses; Input the images and prompt texts into a multimodal large model, and the multimodal large model generates a fusion vector based on the images and prompt texts, and performs chain-of-thought reasoning based on the fusion vector to obtain a claim settlement conclusion that meets the requirements of the prompt texts.