Insurance information extraction method and device, electronic device, and storage medium

Through artificial intelligence technology, semantic recognition and intent recognition are performed on insurance consultation data to generate personalized consultation response scripts, which solves the problem of low accuracy in insurance information extraction by intelligent robots in insurance promotion and achieves more efficient information extraction.

CN119671752BActive Publication Date: 2025-09-30CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411765344.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-09-30
Estimated Expiration
2044-12-03

AI Technical Summary

Technical Problem

In the existing technology, the accuracy of intelligent robots in extracting insurance information in insurance promotion scenarios is low, and the user description method varies greatly, which makes the semantic analysis method prone to errors and difficult to meet the complex and changeable insurance promotion scenarios.

Method used

Artificial intelligence technology is used to obtain insurance consultation data for semantic recognition, and the preset speech generation model and intention recognition sub-model are used to generate consultation response speech to improve the accuracy of insurance information extraction.

Benefits of technology

It improves the accuracy of insurance information extraction, can process insurance consultation sub-data of different formats, adapt to complex and changeable insurance promotion scenarios, and generate more accurate consultation response scripts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119671752B_ABST
    Figure CN119671752B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a method and device for extracting insurance information, an electronic device, and a storage medium, which belongs to the field of information extraction technology and is applicable to the field of financial technology. The method includes: performing semantic recognition based on the insurance consultation sub-data of the target object to obtain semantic recognition data; generating consultation prompt data based on the semantic recognition data and the target insurance data model that matches the target insurance type; performing intention recognition on the consultation prompt data and the semantic recognition data based on the intention recognition sub-model of the preset speech generation model to obtain the insurance intention label of the target object; performing speech generation on the insurance intention label and the semantic recognition data based on the speech generation sub-model of the preset speech generation model to obtain the consultation response speech; obtaining the object response data of the target object; extracting insurance information from the object response data to obtain the key information of the object's insurance. The embodiment of the present application can improve the accuracy of insurance information extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information extraction technology, and is applicable to the field of financial technology, and in particular to a method and device for extracting insurance information, an electronic device, and a storage medium. Background Art

[0002] Insurance information extraction refers to the process of identifying the type of insurance coverage based on the user's input descriptive information in insurance promotion scenarios, thereby extracting information related to that type of insurance coverage. Currently, intelligent robots can recognize user-entered descriptive information. However, since intelligent robots often use command-based settings, they need to exhaustively enumerate possible user input commands or use natural language processing (NLP) to analyze the correlation between the user's input descriptive information and the command, and then use a decision tree to determine the appropriate action to take.

[0003] However, the development complexity of intelligent robots using command-based systems is high, making it difficult to adapt to the complex and ever-changing scenarios of actual insurance promotion, which affects the accuracy of insurance information extraction. Furthermore, the semantic analysis methods used by related technologies are prone to errors, due to the wide variation in user descriptions. Therefore, improving the accuracy of insurance information extraction has become a pressing technical challenge. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a method and device for extracting insurance information, an electronic device, and a storage medium, aiming to improve the accuracy of extracting insurance information.

[0005] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes a method for extracting insurance information, the method comprising:

[0006] Obtaining insurance consultation data of a target object, wherein the insurance consultation data includes insurance consultation sub-data, and the format types of any two insurance consultation sub-data are different;

[0007] Performing semantic recognition based on the insurance consultation sub-data to obtain semantic recognition data of the target insurance type;

[0008] Acquire a target insurance data model that matches the target insurance type;

[0009] Generate prompt data based on the semantic recognition data and the target insurance data model to obtain consultation prompt data;

[0010] An intention recognition sub-model based on a preset speech generation model performs intention recognition on the consultation prompt data and the semantic recognition data to obtain the insurance intention label of the target object;

[0011] Based on the speech generation sub-model of the preset speech generation model, speech generation is performed on the insurance intention label and the semantic recognition data to obtain consultation response speech;

[0012] Acquiring object response data fed back by the target object based on the consultation response script;

[0013] The insurance information of the object response data is extracted to obtain the key insurance information of the object, and the key insurance information of the object matches the target insurance type.

[0014] In some embodiments, the intent recognition sub-model includes an encoder, a context extraction unit, a position encoder, a self-attention unit, and a decoder;

[0015] The intention recognition sub-model based on the preset speech generation model performs intention recognition on the consultation prompt data and the semantic recognition data to obtain the insurance intention label of the target object, including:

[0016] Encoding the consultation prompt data based on the encoder to obtain a consultation prompt sequence;

[0017] Performing data encoding on the semantic recognition data based on the encoder to obtain a semantic recognition sequence;

[0018] Performing sequence fusion on the consultation prompt sequence and the semantic recognition sequence to obtain an initial fused sequence;

[0019] Performing feature extraction on the initial fusion sequence based on the context extraction unit to obtain initial fusion features;

[0020] Performing position encoding on the initial fusion sequence based on the position encoder to obtain a position encoding feature;

[0021] Performing self-attention processing on the initial fusion feature and the position encoding feature based on the self-attention unit to obtain a target fusion feature;

[0022] The target fusion feature is decoded based on the decoder to obtain the insurance intention label.

[0023] In some embodiments, the speech generation sub-model based on the preset speech generation model generates speech based on the insurance intention label and the semantic recognition data to obtain the consultation response speech, including:

[0024] Selecting insurance rule data for the target insurance type from a preset insurance knowledge graph based on the insurance intention label, where the insurance rule data includes insurance rule sub-data;

[0025] Acquiring initial recognition sub-data from the semantic recognition data;

[0026] Performing data matching on the initial identification sub-data and the insurance rule sub-data to obtain a data matching result;

[0027] Determining unuploaded sub-data from the insurance rule sub-data based on the data matching result;

[0028] The initial recognition sub-data and the non-uploaded sub-data are subjected to speech generation based on the speech generation sub-model to obtain the consultation response speech.

[0029] In some embodiments, generating the speech based on the speech generation sub-model for the initial recognition sub-data and the non-uploaded sub-data to obtain the consultation response speech includes:

[0030] If the unuploaded sub-data is empty, generating a speech for the initial identification sub-data based on the speech generation sub-model to obtain the consultation response speech, which is used to guide the target object to confirm the information of the initial identification sub-data;

[0031] If the unuploaded sub-data is not empty, speech generation is performed on the unuploaded sub-data based on the speech generation sub-model to obtain the consultation response speech, which is used to guide the target object to supplement information on the unuploaded sub-data.

[0032] In some embodiments, performing semantic recognition based on the insurance consultation sub-data to obtain semantically recognized data of the target insurance type includes:

[0033] Obtain the data format type of the insurance consultation sub-data;

[0034] Determining a target semantic detection model for the insurance consultation sub-data based on the data format type;

[0035] Performing semantic detection on the insurance consultation sub-data based on the target semantic detection model to obtain a detection text;

[0036] Performing text enhancement on the detected text to obtain enhanced text;

[0037] Identify the consultation intention of the insurance consultation sub-data to obtain a consultation intention label and the target insurance type;

[0038] Generate a hypothetical document based on the consultation intention label and the enhanced text to obtain a target hypothetical document;

[0039] Data integration is performed based on the enhanced text, the consultation intention label, the target insurance type and the target hypothetical document to obtain the semantic recognition data.

[0040] In some embodiments, the method further comprises:

[0041] Obtaining an underwriting rule decision tree for the target insurance type, the underwriting rule decision tree including underwriting rule nodes and node information types of the underwriting rule nodes;

[0042] Performing node information matching on the subject's key insurance information and the underwriting rule node based on the node information type, and determining a node label of the underwriting rule node, wherein the node label includes a node matching label;

[0043] Generate an insurance application form for the target object based on the node label and the key insurance information of the object.

[0044] In some embodiments, generating the insurance application form for the target object based on the node label and the key insurance information of the object includes:

[0045] Selecting valid extracted information that matches the underwriting rule node from the subject's key information;

[0046] Classifying the effectively extracted information based on a preset information classification model to obtain an information category;

[0047] Performing information verification on the information category and the valid extracted information based on a preset information verification model to obtain a verification result;

[0048] If the verification result is that the verification is passed and the node labels of the underwriting rule nodes are all the node matching labels, the insurance application form of the target object is generated based on the key insurance information of the object.

[0049] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a device for extracting insurance information, the device comprising:

[0050] The first acquisition module is used to acquire the insurance consultation data of the target object, wherein the insurance consultation data includes insurance consultation sub-data, and the format types of any two insurance consultation sub-data are different;

[0051] A semantic recognition module, configured to perform semantic recognition based on the insurance consultation sub-data to obtain semantic recognition data of the target insurance type;

[0052] A second acquisition module is used to acquire a target insurance data model that matches the target insurance type;

[0053] A generation module, configured to generate prompt data based on the semantic recognition data and the target insurance data model to obtain consultation prompt data;

[0054] An intention recognition module is used to perform intention recognition on the consultation prompt data and the semantic recognition data based on the intention recognition sub-model of the preset speech generation model to obtain the insurance intention label of the target object;

[0055] A speech generation module is used to generate speech based on the speech generation sub-model of the preset speech generation model for the insurance intention label and the semantic recognition data to obtain the consultation response speech;

[0056] A third acquisition module is configured to acquire object response data fed back by the target object based on the consultation response script;

[0057] The extraction module is used to extract insurance information from the object response data to obtain key insurance information of the object, and the key insurance information of the object matches the target insurance type.

[0058] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.

[0059] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.

[0060] The present application proposes a method and device for extracting insurance information, an electronic device, and a storage medium. The method obtains insurance consultation data of a target object, wherein the insurance consultation data includes insurance consultation sub-data, and the format types of any two insurance consultation sub-data are different; further, semantic recognition is performed based on the insurance consultation sub-data to obtain semantic recognition data having a target insurance type; a target insurance data model matching the target insurance type is obtained; prompt data is generated based on the semantic recognition data and the target insurance data model to obtain consultation prompt data; further, an intention recognition sub-model of a preset speech generation model is used to perform intent recognition on the consultation prompt data and the semantic recognition data to obtain an insurance intention label for the target object; a speech generation sub-model of a preset speech generation model is used to generate speech for the insurance intention label and the semantic recognition data to obtain a consultation response speech; further, object response data fed back by the target object based on the consultation response speech is obtained; insurance information is extracted from the object response data to obtain key insurance information for the object, and the key insurance information for the object matches the target insurance type. Compared to related technologies that require an exhaustive list of possible user input commands and struggle to meet the complex and ever-changing scenarios encountered in actual insurance promotion, this application can perform semantic recognition on insurance consultation sub-data of varying formats and generate more accurate consultation prompt data based on the matching target insurance data model. This allows for more accurate consultation response scripts to be generated based on this consultation prompt data, facilitating information extraction in complex and ever-changing scenarios. Therefore, embodiments of this application can improve the accuracy of insurance information extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 This is the first flow chart of the insurance information extraction method provided in the embodiment of the present application;

[0062] Figure 2 yes Figure 1 A flowchart of step S120 in FIG.

[0063] Figure 3 This is a schematic diagram of a framework flow of the insurance information extraction method provided in an embodiment of the present application;

[0064] Figure 4 yes Figure 1 A flowchart of step S150 in FIG.

[0065] Figure 5 yes Figure 1 A flowchart of step S160 in FIG.

[0066] Figure 6 This is the second flow chart of the insurance information extraction method provided in the embodiment of the present application;

[0067] Figure 7 yes Figure 6A flowchart of step S630 in FIG.

[0068] Figure 8 This is a specific flowchart of generating an insurance application provided in an embodiment of the present application;

[0069] Figure 9 This is a structural diagram of the insurance information extraction device provided in an embodiment of the present application;

[0070] Figure 10 This is a hardware structure diagram of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0071] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0072] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0074] First, let’s analyze some of the terms used in this application:

[0075] Artificial Intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and create new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses theories, methods, technologies, and application systems that use digital computers or digital computer-controlled machines to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0076] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages ​​(such as Chinese and English). NLP is a branch of artificial intelligence and an interdisciplinary subject between computer science and linguistics. It is often referred to as computational linguistics. Natural language processing includes grammatical analysis, semantic analysis, and text understanding. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining related to language processing, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computing.

[0077] Intent recognition is the process of understanding user intent expressed through natural language (such as text or voice) and classifying it into predefined intent categories. In the FinTech insurance consulting scenario, intent recognition in conversations between users and intelligent robots can output responses that better match user intent, thereby enabling the user to receive the response based on the response.

[0078] A knowledge graph is an information system used to store and represent knowledge. It organizes entities and their relationships in a graph format. A knowledge graph typically consists of nodes (representing entities such as people, places, and things) and edges (representing relationships between entities). This structure enables knowledge to be represented and searched in an intuitive manner.

[0079] Insurance information extraction refers to the process of determining the type of insurance coverage based on the descriptive information entered by the user in insurance promotion scenarios, thereby extracting insurance information for that type. In traditional insurance sales processes, the collection and recording of customer information often relies on manual operations, which are inefficient and prone to errors. To improve work efficiency and service quality, a more intelligent solution is needed. Currently, intelligent robots can recognize the descriptive information entered by users. However, since intelligent robots mostly use command-based settings, they need to exhaustively enumerate possible user commands, or use natural language processing to analyze the correlation between the descriptive information entered by the customer and the command, and then use a decision tree to determine what action to perform.

[0080] However, the development complexity of intelligent robots using command-based systems is high, making it difficult to adapt to the complex and ever-changing scenarios of actual insurance promotion, which affects the accuracy of insurance information extraction. Furthermore, the semantic analysis methods used by related technologies are prone to errors, due to the wide variation in user descriptions. Therefore, improving the accuracy of insurance information extraction has become a pressing technical challenge.

[0081] Based on this, the embodiments of the present application provide a method and device for extracting insurance information, an electronic device, and a storage medium, aiming to improve the accuracy of extracting insurance information.

[0082] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0083] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0084] The insurance information extraction method provided in the embodiment of the present application relates to the field of artificial intelligence technology. The insurance information extraction method provided in the embodiment of the present application can be applied to the terminal, can also be applied to the server side, and can also be software running in the terminal or the server side. In some embodiments, the terminal can be a smart phone, tablet computer, laptop computer, desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the insurance information extraction method, etc., but is not limited to the above forms.

[0085] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0086] It should be noted that in each specific implementation of this application, when it comes to the need to perform relevant processing based on data related to the identity or characteristics of the subject, such as the subject's insurance data, the subject's voice data, and the subject's feedback data, the subject's permission or consent will be obtained first, and the collection, use and processing of these data will comply with relevant laws, regulations and standards. In addition, when the embodiment of this application needs to obtain the sensitive personal information of the subject, the subject's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the subject's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of this application will be obtained.

[0087] See also Figure 1 , Figure 1 This is an optional flow chart of the method for extracting insurance information provided in the embodiments of this application. In some embodiments of this application, Figure 1 The method may specifically include but is not limited to steps S110 to S180.

[0088] Step S110, obtaining insurance consultation data of the target object;

[0089] Step S120: performing semantic recognition based on the insurance consultation sub-data to obtain semantic recognition data of the target insurance type;

[0090] Step S130, obtaining a target insurance data model that matches the target insurance type;

[0091] Step S140, generating prompt data based on the semantic recognition data and the target insurance data model to obtain consultation prompt data;

[0092] Step S150: performing intent recognition on the consultation prompt data and semantic recognition data based on the intent recognition sub-model of the preset speech generation model to obtain the target subject's insurance intention label;

[0093] Step S160: Generate a speech based on the insurance intention label and semantic recognition data based on the speech generation sub-model of the preset speech generation model to obtain the consultation response speech;

[0094] Step S170, obtaining the target object's response data based on the consultation response script;

[0095] Step S180: extract the insurance information of the subject's response data to obtain the key insurance information of the subject.

[0096] In step S110 of some embodiments, the insurance consultation data refers to the information provided by the target subject in the insurance system regarding insurance consultation. For example, in the insurance field of FinTech, the target subject is a user seeking information about insurance products (such as social insurance and medical insurance). The target subject can then exchange information with the intelligent robot in the insurance system, and the information provided by the target subject in response is equivalent to the insurance consultation data. In this case, the insurance consultation data may include the target subject's questions, needs, preferences, etc.

[0097] It should be noted that the insurance consultation data includes multiple insurance consultation sub-data, and each insurance consultation sub-data can represent a part or an aspect of the consultation data. For example, the insurance consultation sub-data can be inquiries about insurance types, premium calculations, claims processing, insurance information, etc., without specific limitation. In addition, the insurance consultation sub-data of this application can contain different format types, that is, the format types of any two insurance consultation sub-data are different to meet the diversity of data. For example, the format type of the insurance consultation sub-data can be text format (such as questions sent by users through chat or email), voice format (such as voice records left by users through telephone consultations), image format (such as pictures of insurance documents uploaded by users), PDF or Word documents (such as insurance-related documents that users may upload), video format (such as questions that users may leave through video consultations), etc., without specific limitation.

[0098] It should be noted that due to the diversity of insurance consultation sub-data, this application can use different technologies and methods to parse and understand each format of data for subsequent semantic recognition and data processing. The specific processing process will be explained in detail later and will not be repeated here. In this way, the flexibility and adaptability of the system can be improved, enabling it to handle various forms of user input.

[0099] In step S120 of some embodiments, the present application may further analyze the collected insurance consultation sub-data to identify and understand the intentions and needs of the target object, so as to identify semantic recognition data with the target insurance type, thereby analyzing whether the target object is consulting about relevant information of the target insurance type and whether it has the intention to take out insurance. The target insurance type is used to indicate different types of pre-set insurance, such as health insurance, car insurance, pension insurance, etc. By identifying and understanding the semantic recognition data of the target object with the target insurance type, the target object's consultation can be matched with a specific insurance product or service to provide more accurate and relevant information.

[0100] It should be noted that when performing semantic recognition based on the insurance consultation sub-data, the insurance consultation sub-data can first be cleaned, formatted and other pre-processing operations to improve data quality. Furthermore, key features such as keywords, phrases, entities, etc. can be extracted from the pre-processed data, and the extracted features can be analyzed using machine learning models (such as classifiers) to identify the consulting intentions of the target object. Furthermore, the specific insurance types, terms, conditions and other entity information mentioned in the insurance consultation sub-data can be identified. By combining the results of intent recognition and entity recognition, the consulting content of the target object can be understood and matched with the target insurance type to generate semantic recognition data containing the target insurance type, providing a basis for subsequent data processing and response generation.

[0101] See also Figure 2 , Figure 2 This is a specific flow chart of step S120 provided in an embodiment of the present application. In some embodiments of the present application, step S120 may specifically include but is not limited to steps S210 to S270.

[0102] Step S210, obtaining the data format type of the insurance consultation sub-data;

[0103] Step S220, determining a target semantic detection model for the insurance consultation sub-data based on the data format type;

[0104] Step S230, performing semantic detection on the insurance consultation sub-data based on the target semantic detection model to obtain a detection text;

[0105] Step S240, performing text enhancement on the detected text to obtain enhanced text;

[0106] Step S250: Identify the consultation intention of the insurance consultation sub-data to obtain the consultation intention label and target insurance type;

[0107] Step S260 , generating a hypothetical document based on the consultation intention label and the enhanced text to obtain a target hypothetical document;

[0108] Step S270 , performing data integration based on the enhanced text, the consultation intention label, the target insurance type, and the target hypothetical document to obtain semantic recognition data.

[0109] In step S210 of some embodiments, since the insurance consultation sub-data may include multiple format types such as text, voice, and image, the data format type of the insurance consultation sub-data can be obtained first to clarify the specific type of each insurance consultation sub-data for subsequent processing.

[0110] In step S220 of some embodiments, after identifying the data format type of the insurance consultation sub-data, the present application selects an appropriate target semantic detection model based on each data format type. Each target semantic detection model is equivalent to a model trained using historical insurance data of different data format types. This means that different data formats can employ different processing methods and models to extract semantic information, thereby improving the accuracy of semantic recognition of the data and, in turn, the accuracy of extracting insurance information.

[0111] Please note that Figure 3 , Figure 3 It is a framework flow chart of the insurance information extraction method provided in the embodiment of the present application. Among them, the insurance consultation sub-data of the present application may include text data input by text, image data input by pictures and other multimodal input data, that is, a plurality of insurance consultation sub-data with different data format types are obtained. Furthermore, a target semantic detection model constructed based on the NLP model structure can be used for text data matching, a target semantic detection model constructed based on the Optical Character Recognition (OCR) model structure can be used for picture data matching, a target semantic detection model constructed based on the Computer Vision (CV) model structure can be used for multimodal input data matching, and these models can also be flexibly replaced according to actual needs without limitation. Furthermore, semantic detection of the insurance consultation sub-data can be performed based on the corresponding target semantic detection model.

[0112] It should be noted that the target semantic detection model constructed based on OCR in this application can learn open source frameworks such as PaddleOCR and EasyOCR, develop its own insurance OCR model, and use a large number of object cases for training and verification, which can accurately extract the target object’s valid information and insurance plan information. The target semantic detection model constructed based on CV in this application can be based on the SimpleCV free machine vision framework, access high-performance computer vision libraries such as openCV, and build a free insurance document and information recognition model. The target semantic detection model constructed based on NLP in this application can use open source library packages and related algorithms to build an NLP module by itself, assist in splitting the semantic components and parts of speech of the sentence input by the target object, and assist in judging the intention and emotion of the target object.

[0113] In step S230 of some embodiments, the insurance consultation sub-data may be further analyzed based on the determined target semantic detection model to extract and understand its semantic content, ultimately generating detection text. The semantic detection of the insurance consultation sub-data based on the target semantic detection model in this application may include semantic understanding, synonym rewriting, text hypothesis, etc., to obtain the detection text, without limitation.

[0114] In step S240 of some embodiments, after obtaining the detected text, the present application may perform text enhancement on the detected text to improve the quality and usability of the text, thereby obtaining enhanced text. Text enhancement may include removing noise, filling in missing information, correcting errors, etc., to obtain more accurate and complete enhanced text.

[0115] In step S250 of some embodiments, the present application can further analyze the enhanced text to identify the target user's inquiry intent and categorize these intents into inquiry intent tags. The application also identifies the specific insurance type the user is inquiring about, paving the way for providing accurate insurance information later. By performing inquiry intent identification on the insurance consultation sub-data, the specific content of the target user's inquiry can be understood, including questions, needs, concerns, etc. Inquiry intent identification results in the assignment of one or more tags to the target user's inquiry, representing their inquiry intent. For example, if the target user inquires about medical insurance coverage, the system might assign this inquiry an intent tag of "medical insurance coverage." This allows the system to more effectively respond to the target user's inquiries. By understanding the target user's intent and needs, the system can provide more targeted answers and advice. For example, if the system identifies a user's interest in health insurance, it can provide detailed information about health insurance, including the features, terms, and precautions of different products. This personalized service can improve user satisfaction and increase conversion rates.

[0116] In step S260 of some embodiments, the present application may further utilize the consultation intent tag and enhanced text to generate a hypothetical document that simulates possible answers or related information to assist in subsequent data processing and answer generation.

[0117] In step S270 of some embodiments, the present application may further integrate the enhanced text, the inquiry intent label, the target insurance type, and the hypothetical document to form a comprehensive semantic recognition data set, which will be used for subsequent decision support and response generation.

[0118] In step S130 of some embodiments, the present application can also simultaneously obtain a target insurance data model that matches the target insurance type. This target insurance data model is trained based on insurance product data related to the target insurance type and can include detailed information, terms, conditions, etc. of the insurance product for subsequent data processing and response generation.

[0119] In step S140 of some embodiments, further, as Figure 3 As shown, after determining the semantic recognition data and the target insurance data model, prompt data can be generated based on the semantic recognition data and the target insurance data model to obtain consulting prompt data, that is, personalized customized prompt data. In this way, the generated consulting prompt data can include key information about the insurance product and questions that the target person may need to know, which can be used to guide the target person to further consultation.

[0120] In step S150 of some embodiments, the present application may further use the intent recognition sub-model in the preset speech generation model to analyze the consultation prompt data and semantic recognition data to determine the target subject's insurance enrollment intention label. This insurance enrollment intention label can help the system understand the user's specific needs and goals.

[0121] It should be noted that the preset speech generation model of this application can be based on the model structure of artificial intelligence generated content (AIGC), and use the quantized low-rank adapter (QLoRA) method to fine-tune the open source general language model (such as the Qwen72B model, etc.) to form a special large model of this application in the insurance field. This model can understand insurance expertise and accurately extract information related to the insurance model. Among them, Figure 3As shown, this application can first use insurance professional knowledge training (such as insurance clauses, insurance knowledge, classic cases, etc.) to extract professional data, mark proper nouns, etc., and vectorize the above data, and optimize the AIGC model in combination with QLoRA and other large language models to obtain a preset speech generation model. At the same time, in practice, the model algorithm can also be optimized to record the process of intelligent customer service as a historical record of the target object, and use the historical record as input data for the optimization algorithm to optimize the service of intelligent customer service.

[0122] It should be noted that the intent recognition sub-model includes an encoder, a context extraction unit, a position encoder, a self-attention unit and a decoder.

[0123] See also Figure 4 , Figure 4 This is a specific flow chart of step S150 provided in an embodiment of the present application. In some embodiments of the present application, step S150 may specifically include but is not limited to steps S410 to S470.

[0124] Step S410, encoding the consultation prompt data based on the encoder to obtain a consultation prompt sequence;

[0125] Step S420, encoding the semantic recognition data based on the encoder to obtain a semantic recognition sequence;

[0126] Step S430, performing sequence fusion on the consultation prompt sequence and the semantic recognition sequence to obtain an initial fused sequence;

[0127] Step S440, performing feature extraction on the initial fusion sequence based on the context extraction unit to obtain initial fusion features;

[0128] Step S450, position encoding the initial fusion sequence based on the position encoder to obtain a position encoding feature;

[0129] Step S460, performing self-attention processing on the initial fusion feature and the position encoding feature based on the self-attention unit to obtain the target fusion feature;

[0130] Step S470: decode the target fusion features based on the decoder to obtain the insurance intention label.

[0131] In steps S410 and S420 of some embodiments, when performing intent recognition on the consultation prompt data and the semantic recognition data, the present application may first encode the consultation prompt data and the semantic recognition data based on an encoder to obtain a consultation prompt sequence and a semantic recognition sequence. The encoder may be a neural network model, such as a recurrent neural network (RNN), a long short-term memory network (LSTM), or a transformer, etc., to convert the consultation prompt data and the semantic recognition data into a series of numerical representations, i.e., a consultation prompt sequence and a semantic recognition sequence.

[0132] In step S430 of some embodiments, the consultation prompt sequence and the semantic recognition sequence may be further subjected to sequence fusion, such as sequence merging or sequence weighted merging, to create an initial fused sequence containing information of both.

[0133] In step S440 of some embodiments, further, a context extraction unit (such as a special neural network layer) can be used to extract key features from the initial fusion sequence, and these features can represent the context information in the sequence to obtain initial fusion features.

[0134] In step S450 of some embodiments, a position encoder may further be used to add position information to each element in the sequence, which helps the model understand the order of the elements in the sequence. Thus, the resulting position encoding feature is a feature with added position information.

[0135] In step S460 of some embodiments, the initial fused features and the position encoding features are further processed by the self-attention unit to obtain a target fused feature. The self-attention mechanism (such as the self-attention in the Transformer) is used to process the initial fused features and the position encoding features to enhance the model's understanding of the importance of different parts of the sequence. Furthermore, the target fused features obtained after processing contain the context and position information of the sequence.

[0136] In step S470 of some embodiments, a decoder may be used to decode the target fused features to generate a final output, namely, a label of the insurance intention. This label represents the model's recognition result of the user's consultation intention.

[0137] In the above embodiments, the present application can construct an intent recognition sub-model based on the encoding-fusion-decoding process. In this way, the model can understand and process complex language information and generate accurate intent recognition results.

[0138] In step S160 of some embodiments, further, the present application can generate speech for the insurance intention label and semantic recognition data based on the speech generation sub-model of the preset speech generation model to obtain consultation response speech. In other words, through the above continuously refined intention recognition and semantic recognition, the present application can deeply understand the conversation habits of the target object, thereby helping the system to generate business development speech that better simulates human conversation. And in the process of conversation, the needs of the target object are excavated, so that the required insurance products can be accurately recommended to the target object. For example, if the target object consults about a certain insurance product, the generated consultation response speech can also guide the target object to upload and complete the information of the relevant insurance product, thereby improving the efficiency of insurance information collection.

[0139] See also Figure 5 , Figure 5 This is a specific flow chart of step S160 provided in an embodiment of the present application. In some embodiments of the present application, step S160 may specifically include but is not limited to steps S510 to S550.

[0140] Step S510: selecting insurance rule data of the target insurance type from the preset insurance knowledge graph based on the insurance intention label;

[0141] Step S520, obtaining initial recognition sub-data from the semantic recognition data;

[0142] Step S530: performing data matching on the initial identification sub-data and the insurance rule sub-data to obtain a data matching result;

[0143] Step S540, determining the sub-data that has not been uploaded from the insurance rule sub-data based on the data matching result;

[0144] Step S550: Generate speech for the initial recognition sub-data and the unuploaded sub-data based on the speech generation sub-model to obtain consultation response speech.

[0145] In step S510 of some embodiments, when generating the consultation response, the application may first search for insurance rule data that matches the user's consultation intent in a preset insurance knowledge graph based on the previously identified insurance intention tag. This insurance rule data includes specific insurance rules and conditions, and the insurance rule data includes insurance rule sub-data, each of which can contain different rules or terms.

[0146] In step S520 of some embodiments, the present application may further obtain initial recognition sub-data from the semantic recognition data. These initial recognition sub-data are the system's preliminary understanding of the consulting content of the target object, and may include key information such as the target object's problems and needs.

[0147] In step S530 of some embodiments, the present application can further perform data matching on the initial identification sub-data and the insurance rule sub-data to determine which rules are directly related to the target object's consultation and obtain data matching results. The data matching results can help the system understand the needs of the target object more accurately and provide a basis for generating consultation response scripts.

[0148] In step S540 of some embodiments, the present application can further identify which insurance rule sub-data have not been uploaded or mentioned by the target object based on the data matching results, and these unuploaded sub-data may contain information related to the target insurance type that the target object needs to know.

[0149] In step S550 of some embodiments, further, based on a speech generation sub-model (e.g., a machine learning model specifically for generating natural language responses), speech generation can be performed on the initial recognition sub-data and the unuploaded sub-data to obtain consultation response speech. Furthermore, these consultation response speech are intended to answer the consultation questions of the target subject and provide additional information that may be required for the consultation response speech.

[0150] In the above embodiment, this application generates personalized insurance consultation responses by combining knowledge graphs and natural language processing technology. In this way, the system can provide more accurate and comprehensive answers, thereby improving user satisfaction and consultation efficiency.

[0151] It should be noted that this application can analyze the target subject's conversation context through the intent recognition sub-model, explore the target subject's insurance needs, and use the consultation response dialogue generated by the dialogue generation sub-model to guide the target subject to enter relevant insurance information. Furthermore, according to the insurance terms and conditions, the target subject's valid information and coverage information can be extracted, and relevant dialogue can be generated for the target subject to confirm or to modify or supplement the information.

[0152] It should be noted that this application generates conversational text based on the conversational text generation sub-model for the initial recognition sub-data and the unuploaded sub-data to obtain consultation response conversational text, which may specifically include the following situations:

[0153] If the uploaded sub-data is empty, the initial identification sub-data is used to generate a speech based on the speech generation sub-model to obtain the consultation response speech. The consultation response speech is used to guide the target object to confirm the information of the initial identification sub-data;

[0154] If the unuploaded sub-data is not empty, the unuploaded sub-data is subjected to speech generation based on the speech generation sub-model to obtain the consultation response speech, which is used to guide the target object to supplement the information of the unuploaded sub-data.

[0155] It should be noted that if the unuploaded sub-data is empty, this means that the target object has not provided additional information or data, and the system cannot identify the content that needs to be supplemented. At this time, the initial identification sub-data can be generated based on the speech generation sub-model, that is, based on the initial identification sub-data of the target object, a preliminary understanding is made, and consultation response speech is generated to guide the target object to confirm the information of the initial identification sub-data. If the unuploaded sub-data is not empty, this means that the system has identified that the target object has missed some information or data in the consultation. At this time, the speech generation sub-model can be used to generate speech for the unuploaded sub-data to guide the target object to supplement the information of the unuploaded sub-data, that is, to help the target object supplement the information missed in the consultation, so that the system can better understand the needs of the target object and provide accurate services.

[0156] In the above embodiment, the application can adjust the response strategy based on the status of the information provided by the user (whether there is any unuploaded sub-data) to ensure that the target party can confirm or supplement the necessary information. In other words, the application can explore the insurance needs of the target party, guide the target party to enter insurance information, and extract the target party's valid information and coverage information according to the insurance terms and conditions.

[0157] In step S170 of some embodiments, after the system provides the consultation response script, the target subject can provide feedback based on the received consultation response script, thereby obtaining the target subject's response data. The target subject's response data can include questions answered by the target subject, supplementary images, documents, etc., without limitation.

[0158] In step S180 of some embodiments, the application may further analyze all the target subject's response data to extract key insurance information that matches the target insurance type. This information may include the target subject's personal information, selected insurance product, insurance term, and required insurance product information, etc., to complete the insurance consultation and purchase process.

[0159] See also Figure 6 , Figure 6 This is another optional flow chart of the method for extracting insurance information provided by the embodiment of the present application. In some embodiments of the present application, after step S180, the method for extracting insurance information provided by the present application may further include but is not limited to steps S610 to S630.

[0160] Step S610, obtaining an underwriting rule decision tree for the target insurance type;

[0161] Step S620: Match the key information of the insurance subject with the underwriting rule node based on the node information type to determine the node label of the underwriting rule node;

[0162] Step S630: Generate an insurance application form for the target object based on the node label and the key insurance information of the object.

[0163] In step S610 of some embodiments, after the key information of all insured persons is extracted, the present application can also intelligently initiate insurance application, information review, document quality inspection, etc., and finally complete the insurance application and send the insurance policy to the target person. Specifically, the underwriting rule decision tree of the target insurance type can be obtained first. The underwriting rule decision tree is a decision support tool for automating the underwriting process, consisting of multiple nodes, each node representing a decision point, that is, the underwriting rule decision tree includes underwriting rule nodes and node information types of underwriting rule nodes, and the node information type indicates the specific information or conditions required for each underwriting rule node.

[0164] In step S620 of some embodiments, the present application may further perform node information matching between the subject's key information and the underwriting rule node based on the node information type to determine the node label of the underwriting rule node. Through this matching, the system can determine the node label of each underwriting rule node to determine whether the underwriting rule node matches the corresponding subject information, that is, whether the information acquisition is complete, and how to make the next decision based on the information provided by the target subject.

[0165] It should be noted that node labels include node matching labels and node non-matching labels. Node matching labels are used to indicate that the corresponding underwriting rule node has obtained the corresponding object information. Node non-matching labels are used to indicate that the corresponding underwriting rule node has not obtained the corresponding object information.

[0166] In some embodiments, step S630 can further generate an insurance application for the target subject based on the node label and the subject's key insurance information. In other words, in this final step, the system can utilize the determined node label and key insurance information to generate an insurance application. This application is automatically generated based on the underwriting rule decision tree and can include all necessary information to complete the user's insurance application. This automated process can reduce manual input errors and improve efficiency and accuracy.

[0167] In the above embodiment, the present application provides an automated underwriting process, which guides underwriting decisions through a decision tree to ensure the accuracy and completeness of the insurance application.

[0168] See also Figure 7 , Figure 7 This is a specific flow chart of step S630 provided in an embodiment of the present application. In some embodiments of the present application, step S630 may specifically include but is not limited to steps S710 to S740.

[0169] Step S710: Select valid extracted information that matches the underwriting rule node from the key information of the insurance subject;

[0170] Step S720, classifying the effectively extracted information based on a preset information classification model to obtain an information category;

[0171] Step S730: performing information verification on the information category and the valid extracted information based on a preset information verification model to obtain a verification result;

[0172] Step S740: If the verification result is that the verification is passed and the node labels of the underwriting rule nodes are all node matching labels, an insurance application form for the target object is generated based on the key information of the object.

[0173] In step S710 of some embodiments, the present application may first select multiple valid extracted information that matches the underwriting rule node from the key information of the insurance subject. These valid extracted information are key data required in the underwriting process, such as health status, work status, etc.

[0174] In step S720 of some embodiments, the valid extracted information may be further classified based on a preset information classification model to identify the information category of each valid extracted information. For example, health status may be classified as "basic personal information" and work status may be classified as "occupational information."

[0175] In step S730 of some embodiments, the present application may further perform information verification on the information category and valid extracted information based on a preset information verification model to verify whether the classified information is accurate, complete, and complies with the requirements of the underwriting rules, and obtain a verification result. Among them, information verification may include checking the format, logical consistency, and matching degree of the information with the underwriting rules. The verification result is used to indicate whether the valid extracted information has passed the verification. If all valid extracted information has passed the verification, the verification result is verification passed. If there is valid extracted information that has not passed the verification, the verification result is verification failed.

[0176] In some embodiments, in step S740, if the verification result is passed and the node labels of the underwriting rule nodes are all node matching labels, the application can generate an insurance application form for the target subject based on the verified key information of the target subject. This insurance application form will contain all the necessary information to complete the insurance application process for the target subject.

[0177] It should be noted that if the verification result includes a failure, the valid information extracted will be reclassified based on the preset information classification model to determine whether accurate information was not extracted. Alternatively, a corresponding consultation response script can be generated to guide the target person to supplement relevant information.

[0178] It should be noted that if the node label of the underwriting rule node contains a node unmatched label, it is necessary to re-match the node information of the object insurance key information and the underwriting rule node based on the node information type until the verification result is passed and the node labels of the underwriting rule node are all node matching labels.

[0179] In the above embodiments, the present application provides an automated information processing and verification system, which aims to ensure that the information submitted by the target object is accurate and meets the preset underwriting requirements, thereby improving underwriting efficiency and accuracy.

[0180] See also Figure 8 , Figure 8 It is a specific flow chart of generating an insurance application provided by an embodiment of the present application. Specifically, after obtaining the key information of the object's insurance participation, the present application can obtain valid extracted information (i.e., the object's insurance information) from the key information of the object's insurance participation, and match the node information of the object's insurance key information and the underwriting rule node based on the node information type of the underwriting rule node in the underwriting rule decision tree, and judge whether the node labels of the underwriting rule node are all node matching labels. If so, it means that the match is passed; if not, it means that the match is not passed. At the same time, the present application can classify the valid extracted information based on the preset information classification model to obtain the information category. Further, based on the preset information verification model, it is judged whether the verification result corresponding to the valid extracted information is verification passed (i.e., all valid extracted information has passed the verification). If the verification result is verification passed and the node labels of the underwriting rule nodes are all node matching labels, the insurance application of the target object is generated based on the key information of the object's insurance participation. Further, the target object can confirm whether to agree with the insurance application and complete the payment to receive the insurance business.

[0181] It should be noted that the non-Company's software tools or components that appear in the embodiments of this application are merely examples and do not represent actual use.

[0182] The insurance information extraction method provided in the embodiment of the present application can automatically generate sales talk based on the conversation habits of the subject, completely simulate human conversation, and explore the needs of customers during the conversation, and accurately recommend the insurance products needed by the subject. According to the conversation context, the effective information of the subject is extracted, and the insurance information extraction and entry are automatically completed, which improves work efficiency and reduces the time and errors of manual operations. The object information is recorded in real time to facilitate subsequent analysis and management, which can help companies improve service quality and increase customer satisfaction. Compared with the related technology that requires exhaustive enumeration of possible user input instructions and is difficult to meet the complex and changeable scenarios in actual insurance promotion, the present application can perform semantic recognition on insurance consultation sub-data of different format types, and generate more accurate consultation prompt data based on the matching target insurance data model, so that more accurate consultation response talk can be generated based on the consultation prompt data, which is convenient for information extraction in complex and changeable scenarios. Therefore, the embodiment of the present application can improve the accuracy of insurance information extraction.

[0183] See also Figure 9 The embodiment of the present application further provides an insurance information extraction device that can implement the above-mentioned insurance information extraction method, and the device includes:

[0184] The first acquisition module 910 is used to acquire the insurance consultation data of the target object, wherein the insurance consultation data includes insurance consultation sub-data, and the format types of any two insurance consultation sub-data are different;

[0185] Semantic recognition module 920, for performing semantic recognition based on the insurance consultation sub-data to obtain semantic recognition data of the target insurance type;

[0186] The second acquisition module 930 is used to acquire a target insurance data model that matches the target insurance type;

[0187] A generation module 940 is used to generate prompt data based on the semantic recognition data and the target insurance data model to obtain consultation prompt data;

[0188] Intent recognition module 950, used to perform intent recognition on the consultation prompt data and semantic recognition data based on the intent recognition sub-model of the preset speech generation model to obtain the insurance intention label of the target object;

[0189] A speech generation module 960 is used to generate speech based on the speech generation sub-model of the preset speech generation model based on the insurance intention label and semantic recognition data to obtain the consultation response speech;

[0190] The third acquisition module 970 is used to obtain the object response data fed back by the target object based on the consultation response script;

[0191] The extraction module 980 is used to extract insurance information from the object response data to obtain key insurance information of the object, and the key insurance information of the object matches the target insurance type.

[0192] The specific implementation of the insurance information extraction device in the embodiment of the present application is basically the same as the specific implementation of the above-mentioned insurance information extraction method, and will not be repeated here.

[0193] The present application also provides an electronic device comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described insurance information extraction method. The electronic device can be any smart terminal, including a tablet computer and an in-vehicle computer.

[0194] See also Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0195] The processor 1010 can be implemented as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0196] The memory 1020 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called by the processor 1010 to execute the insurance information extraction method of the embodiments of this application;

[0197] Input / output interface 1030, used to implement information input and output;

[0198] Communication interface 1040, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0199] bus 1050 , which transmits information between various components of the device (e.g., processor 1010 , memory 1020 , input / output interface 1030 , and communication interface 1040 );

[0200] The processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 are connected to each other in communication within the device via a bus 1050 .

[0201] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned insurance information extraction method.

[0202] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0203] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0204] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0205] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0206] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0207] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0208] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0209] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0210] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0211] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0212] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0213] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A method for extracting insurance information, characterized in that: The method comprises: Obtaining insurance consultation data of a target object, wherein the insurance consultation data includes insurance consultation sub-data, and the format types of any two insurance consultation sub-data are different; Performing semantic recognition based on the insurance consultation sub-data to obtain semantic recognition data of the target insurance type; Acquire a target insurance data model that matches the target insurance type; Generate prompt data based on the semantic recognition data and the target insurance data model to obtain consultation prompt data; The intention recognition sub-model based on the preset speech generation model performs intention recognition on the consultation prompt data and the semantic recognition data to obtain the insurance intention label of the target object; the intention recognition sub-model includes an encoder, a context extraction unit, a position encoder, a self-attention unit and a decoder. The intention recognition sub-model based on the preset speech generation model performs intention recognition on the consultation prompt data and the semantic recognition data to obtain the insurance intention label of the target object, including: data encoding the consultation prompt data based on the encoder to obtain a consultation prompt sequence; data encoding the semantic recognition data based on the encoder to obtain a semantic recognition sequence; sequence fusion of the consultation prompt sequence and the semantic recognition sequence to obtain an initial fusion sequence; feature extraction of the initial fusion sequence based on the context extraction unit to obtain an initial fusion feature; position encoding of the initial fusion sequence based on the position encoder to obtain a position encoding feature; self-attention processing of the initial fusion feature and the position encoding feature based on the self-attention unit to obtain a target fusion feature; decoding processing of the target fusion feature based on the decoder to obtain the insurance intention label; The speech generation sub-model based on the preset speech generation model generates speech for the insurance intention label and the semantic recognition data to obtain consultation response speech; the speech generation sub-model based on the preset speech generation model generates speech for the insurance intention label and the semantic recognition data to obtain consultation response speech, including: selecting the insurance rule data of the target insurance type from the preset insurance knowledge graph based on the insurance intention label, the insurance rule data including insurance rule sub-data; obtaining initial recognition sub-data from the semantic recognition data; performing data matching on the initial recognition sub-data and the insurance rule sub-data to obtain a data matching result; determining the unuploaded sub-data from the insurance rule sub-data based on the data matching result; performing speech generation on the initial recognition sub-data and the unuploaded sub-data based on the speech generation sub-model to obtain the consultation response speech; Acquiring object response data fed back by the target object based on the consultation response script; The insurance information of the object response data is extracted to obtain the key insurance information of the object, and the key insurance information of the object matches the target insurance type.

2. The method according to claim 1, characterized in that The step of generating speech for the initial recognition sub-data and the non-uploaded sub-data based on the speech generation sub-model to obtain the consultation response speech includes: If the unuploaded sub-data is empty, generating a speech for the initial identification sub-data based on the speech generation sub-model to obtain the consultation response speech, which is used to guide the target object to confirm the information of the initial identification sub-data; If the unuploaded sub-data is not empty, speech generation is performed on the unuploaded sub-data based on the speech generation sub-model to obtain the consultation response speech, which is used to guide the target object to supplement information on the unuploaded sub-data.

3. The method according to any one of claims 1 to 2, characterized in that The semantic recognition based on the insurance consultation sub-data to obtain semantic recognition data of the target insurance type includes: Obtain the data format type of the insurance consultation sub-data; Determining a target semantic detection model for the insurance consultation sub-data based on the data format type; Performing semantic detection on the insurance consultation sub-data based on the target semantic detection model to obtain a detection text; Performing text enhancement on the detected text to obtain enhanced text; Identify the consultation intention of the insurance consultation sub-data to obtain a consultation intention label and the target insurance type; Generate a hypothetical document based on the consultation intention label and the enhanced text to obtain a target hypothetical document; Data integration is performed based on the enhanced text, the consultation intention label, the target insurance type and the target hypothetical document to obtain the semantic recognition data.

4. The method according to any one of claims 1 to 2, characterized in that The method further comprises: Obtaining an underwriting rule decision tree for the target insurance type, the underwriting rule decision tree including underwriting rule nodes and node information types of the underwriting rule nodes; Performing node information matching on the subject's key insurance information and the underwriting rule node based on the node information type, and determining a node label of the underwriting rule node, wherein the node label includes a node matching label; Generate an insurance application form for the target object based on the node label and the key insurance information of the object.

5. The method according to claim 4, characterized in that Generating the insurance application form of the target object based on the node label and the key insurance information of the object includes: Selecting valid extracted information that matches the underwriting rule node from the subject's key information; Classifying the effectively extracted information based on a preset information classification model to obtain an information category; Performing information verification on the information category and the valid extracted information based on a preset information verification model to obtain a verification result; If the verification result is that the verification is passed and the node labels of the underwriting rule nodes are all the node matching labels, the insurance application form of the target object is generated based on the key insurance information of the object.

6. A device for extracting insurance information, characterized in that: The device comprises: The first acquisition module is used to acquire the insurance consultation data of the target object, wherein the insurance consultation data includes insurance consultation sub-data, and the format types of any two insurance consultation sub-data are different; A semantic recognition module, configured to perform semantic recognition based on the insurance consultation sub-data to obtain semantic recognition data of the target insurance type; A second acquisition module is used to acquire a target insurance data model that matches the target insurance type; A generation module, configured to generate prompt data based on the semantic recognition data and the target insurance data model to obtain consultation prompt data; An intention recognition module is used to perform intention recognition on the consultation prompt data and the semantic recognition data based on the intention recognition sub-model of the preset speech generation model to obtain the insurance intention label of the target object; the intention recognition sub-model includes an encoder, a context extraction unit, a position encoder, a self-attention unit and a decoder. The intention recognition sub-model based on the preset speech generation model performs intention recognition on the consultation prompt data and the semantic recognition data to obtain the insurance intention label of the target object, including: data encoding the consultation prompt data based on the encoder to obtain a consultation prompt sequence; data encoding the semantic recognition data based on the encoder to obtain a semantic recognition sequence; sequence fusion of the consultation prompt sequence and the semantic recognition sequence to obtain an initial fusion sequence; feature extraction of the initial fusion sequence based on the context extraction unit to obtain an initial fusion feature; position encoding of the initial fusion sequence based on the position encoder to obtain a position encoding feature; self-attention processing of the initial fusion feature and the position encoding feature based on the self-attention unit to obtain a target fusion feature; decoding processing of the target fusion feature based on the decoder to obtain the insurance intention label; A speech generation module is used to generate speech for the insurance intention label and the semantic recognition data based on the speech generation sub-model of the preset speech generation model to obtain consultation response speech; the speech generation sub-model based on the preset speech generation model generates speech for the insurance intention label and the semantic recognition data to obtain consultation response speech, including: selecting the insurance rule data of the target insurance type from the preset insurance knowledge graph based on the insurance intention label, the insurance rule data including insurance rule sub-data; obtaining initial recognition sub-data from the semantic recognition data; performing data matching on the initial recognition sub-data and the insurance rule sub-data to obtain a data matching result; determining the unuploaded sub-data from the insurance rule sub-data based on the data matching result; performing speech generation on the initial recognition sub-data and the unuploaded sub-data based on the speech generation sub-model to obtain the consultation response speech; A third acquisition module is configured to acquire object response data fed back by the target object based on the consultation response script; The extraction module is used to extract insurance information from the object response data to obtain key insurance information of the object, and the key insurance information of the object matches the target insurance type.

7. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.