Intelligent insurance slip generation system based on semantic recognition and extraction
Through the intelligent insurance policy generation system based on semantic recognition, the problem of multi-modal data processing difficulties in traditional insurance intermediary services is solved, and automated insurance policy generation and self-service inspection and verification are realized, improving the accuracy and efficiency of information collection.
Patent Information
- Application Number
- CN202510512033.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-29
AI Technical Summary
There are problems in traditional insurance intermediary business scenarios such as fragmentation, difficulty in multimodal data processing, high manual entry error rate, and inability to achieve end-to-end automation, which affects customer conversion efficiency.
The intelligent insurance policy generation system based on semantic recognition is adopted, and through multi-source data collection, intelligent semantic analysis, intelligent picture processing and dynamic order generation, combined with a collaborative inspection system, the automated processing of multi-modal data and self-service inspection and verification are realized.
It improves the accuracy of information collection, analyzes effective information from multiple angles, assists in manual inspection and verification, and improves the efficiency and accuracy of insurance policy generation.
Smart Images

Figure CN120388383A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent data processing, and particularly to an intelligent insurance application form generation system based on semantic recognition and extraction. Background Art
[0002] In the traditional insurance intermediary business scenario, the back-office personnel responsible for issuing policies need to manually process the car insurance application requirements transmitted by sales agents through WeChat / Enterprise WeChat, etc. This involves repetitive tasks such as manually screening chat records, extracting information, classifying documents, and entering insurance application forms. The existing technologies have the following pain points: (1) fragmented and colloquial conversation information; (2) multi-modal data sources and data integration, including cross-modal correlation processing of text, voice, and pictures, etc.; (3) error tolerance cost. When manually entering, problems such as missing or incorrect policy fields are likely to occur, and secondary checks are required; (4) traditional systems cannot achieve end-to-end automation, which affects the customer conversion efficiency. Summary of the Invention
[0003] In view of the above problems, the present invention provides an intelligent insurance application form generation system based on semantic recognition and extraction, which intelligently analyzes multi-modal data, obtains and extracts useful fields in the communication, and matches templates to generate insurance application forms.
[0004] In a first aspect, the present invention provides an intelligent insurance application form generation system based on semantic recognition and extraction, including the following: a multi-source data collection module, which obtains communication records through instant messaging software, including text, pictures, and voice, and performs preprocessing on the data to obtain Data 1;
[0005] An intelligent semantic analysis module, which uses a BERT pre-trained model to learn terms in the insurance field to obtain a dedicated word segmentation model, and performs semantic recognition on the text information in Data 1 through a multi-level semantic recognition module to obtain Feature Data 1; wherein, the multi-level semantic recognition module includes an intention recognition layer, an entity extraction layer, and a context association layer;
[0006] An intelligent image processing engine, which extracts the picture information in Data 1, classifies and adaptively recognizes it to obtain Feature Data 2;
[0007] A dynamic insurance application order generation engine, which fuses Feature Data 1 and Feature Data 2, and through a rule validator and a template adaptation system, connects to an API interface to automatically adapt the message format to obtain an insurance application order;
[0008] A collaborative inspection system, which identifies abnormal values of fields in the insurance application order based on a regular algorithm, and displays them in comparison with the original data through a visual interface for manual inspection and confirmation of information.
[0009] Furthermore, for the BERT pre-training, the related operations are as follows:
[0010] Obtain a corpus of relevant terms in the insurance field, construct a BERT model and adopt multiple transformer attention heads;
[0011] Perform pre-training using a masking strategy, randomly mask 15% of the Tokens, where 80% are replaced with the mask value, 10% are replaced with random Tokens, and 10% remain the original words;
[0012] Construct a cross-entropy loss to predict the masked words.
[0013] Furthermore, for the classification, YOLOv5 is used to classify the document types in the picture information, and the document types include identity cards, driving licenses, and vehicle operation licenses; the specific operations are as follows:
[0014] Construct a detection network model containing a backbone network, a neck connection network, and a detection head network;
[0015] Use data augmentation on the collected document photos to train the detection network model, and obtain a converged detection network model;
[0016] After pruning and compressing the detection network model, deploy it to the intelligent picture processing engine, and perform forward inference on the real-time input picture to obtain the picture detection and classification results.
[0017] Furthermore, the regular expression identifies the abnormal field values in the insurance application order, and the operations are as follows:
[0018] Convert the regular expression into a non-deterministic finite automaton for efficient string matching;
[0019] Adopt the greedy matching principle, that is, each character can correspond to multiple state transitions, and when the current path cannot complete character matching, backtrack to the nearest selection point to try other paths; among them, atomic groups and possessive quantifiers are used to control the backtracking and re-matching behavior in the search.
[0020] Furthermore, the context association layer associates the scattered insurance elements through the Attention mechanism, where the attention mechanism adopts Characterize the association relationship between the insurance elements in the feature data.
[0021] Through the above implementation solutions, the following advantages or beneficial effects are achieved:
[0022] (1) The collection and use of multi-modal data sources increase the accuracy of information collection in the intelligent system;
[0023] (2) Multi-level semantic parsing, parsing and extracting effective information in multi-modal data from multiple perspectives;
[0024] (3) Self-check and verification of the insurance application form, assisted by manual check and verification, to improve efficiency.
[0025] Other features and advantages of the present application will be described in the subsequent specification, and in part will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings
[0026] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0027] Figure 1 It is a schematic diagram of an intelligent insurance application form generation system based on semantic recognition and extraction provided by the present invention. Detailed Embodiments
[0028] To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0029] An atomic group is a syntactic structure used to control backtracking in a regular expression, where the matched content is regarded as an indivisible atom;
[0030] A possessive quantifier is a simplified control for a single quantifier in a regular expression.
[0031] Embodiment 1
[0032] As Figure 1 shown, an intelligent insurance application form generation system based on semantic recognition and extraction, a multi-source data collection module, obtains communication records through instant messaging software, including text, pictures, and voice, and performs preprocessing on the data to obtain Data 1;
[0033] The forms of instant software communication are diverse, including text, pictures, voice, and video; but the more commonly used ones are text, picture, and voice information. The multi-source data collection module can collect multi-modal data forms to form a training corpus.
[0034] Through data text cleaning, word segmentation, sentence segmentation and other processing methods, noise and some irrelevant interference symbols are removed, the text format is standardized, and regular expression tools such as regex are used during the cleaning process; through the dictionary-based word segmentation tool Jieba, the communication text is regularly segmented. Since it is daily communication, the word segmentation granularity does not need to be too fine to avoid losing semantics.
[0035] The intelligent semantic parsing module uses the BERT pre-trained model to learn the terms in the insurance field to obtain a dedicated word segmentation model, and the multi-level semantic recognition module performs semantic recognition on the Chinese text information of the data one to obtain the feature data one; among them, the multi-level semantic recognition module includes an intention recognition layer, an entity extraction layer and a context association layer.
[0036] Semantic parsing is performed on the text and voice data in the instant messaging software; first, the voice data is converted into voice text information and fused with the text information.
[0037] The BERT model uses the multi-head self-attention mechanism and bidirectional context modeling: MultiHead(Q,K,V)=Concat(head1,…,head h )W O ;
[0038] Among them, each attention head is expressed as:
[0039] The feed-forward network composed of fully connected layers enhances the non-linear expression ability:
[0040] FFN(x)=ReLU(xW1+b1)W2+b2
[0041] The task of the awareness recognition layer: Fine-tune the pre-trained BERT model to classify the types of conversations, including tasks such as inquiry, insurance application, order modification, etc.
[0042] Among them, the related operations of the pre-training of the BERT are as follows:
[0043] Obtain the corpus of relevant terms in the insurance field, build a BERT model and use multiple transformer attention heads;
[0044] Use the masking strategy for pre-training, randomly mask 15% of the Tokens, where 80% are replaced with the mask value, 10% are replaced with random Tokens, and 10% remain the original words; construct the cross-entropy loss to predict the masked words.
[0045] The entity extraction layer task extracts structured fields such as insurance types, insured amounts, license plate numbers, etc., and associates the key insurance elements in the context; the context association layer associates the scattered insurance elements through the Attention mechanism, where the attention mechanism uses To represent the association relationship between insurance elements in the feature data.
[0046] The intelligent image processing engine extracts, classifies, and adaptively identifies the image information in Data One to obtain Feature Data Two;
[0047] The classification uses YOLOv5 to classify the document types in the image information, and the document types include ID cards, driving licenses, and vehicle registration certificates.
[0048] The diversity of sample information is increased by performing operations such as rotation and flipping.
[0049] The YOLOv5 framework consists of three network components: the backbone network for feature extraction, the neck connection network for fusing features at different stages, and the detection head network for detecting target categories; a four-classification model with ID cards, driving licenses, vehicle registration certificates, and invalid documents is constructed; the specific operations are as follows:
[0050] Construct a detection network model containing a backbone network, a neck connection network, and a detection head network;
[0051] The collected document photos are used to train the detection network model by data augmentation to obtain a converged detection network model;
[0052] After pruning and compressing the detection network model, it is deployed to the intelligent image processing engine for real-time forward inference to obtain the detection result.
[0053] The dynamic insurance order generation engine fuses the Feature Data One and Feature Data Two, and then connects to the API interface through a rule validator and a template adaptation system to automatically adapt the message format to obtain an insurance order;
[0054] The dynamic insurance order generation engine adopted here mainly automatically fills in the insurance order by means of data verification and template adaptation; among them, the rule validator first sets the verification rules, and then matches and verifies the important fields in the collected Feature Data One and Feature Data Two; and adapts to the corresponding message format through the API interface, and uses the method of joint template field recognition to fill the verified important fields into the template; generate the corresponding insurance order.
[0055] The collaborative inspection system identifies field outliers based on a regular algorithm, and displays them in comparison with the original data through a visual interface for manual inspection and confirmation of information.
[0056] The regular expression recognizes the abnormal field values in the insurance application order, and the operation is as follows:
[0057] Convert the regular expression into a non-deterministic finite automaton for efficient string matching;
[0058] Adopt the greedy matching principle, that is, each character can correspond to multiple state transitions, and when the current path cannot complete character matching, backtrack to the nearest selection point to try other paths; among them, the atomic group and possessive quantifier are used to control the backtracking and re-matching behavior in the search.
[0059] Display the recognized abnormal values on the visualization interface, compare them with the original data pairs matched before, and conduct manual secondary auxiliary inspections to check whether there are abnormal values and confirm the information.
[0060] Advantages of the implementation of the present invention: Through the above solutions, the following advantages or beneficial effects are achieved:
[0061] (1) The collection and use of multi-modal data sources increase the accuracy of information collection in intelligent systems;
[0062] (2) Multi-level semantic parsing, parsing and extracting effective information in multi-modal data from multiple angles;
[0063] (3) Self-checking and verification of insurance application forms, assisting manual checking and verification to improve efficiency.
[0064] As mentioned above, the above are only the specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An intelligent generation system for insurance application forms based on semantic recognition extraction, characterized in that, These include: The multi-source data acquisition module obtains communication records, including text, pictures and voice, through instant messaging software, and pre-processes the data to obtain data 1; An intelligent semantic parsing module uses the BERT pre-trained model to learn insurance-related terminology to obtain a specialized word segmentation model. The module then performs semantic recognition on the text information in the data 1 through a multi-level semantic recognition module to obtain feature data 1. The multi-level semantic recognition module includes an intent recognition layer, an entity extraction layer, and a context association layer. The intelligent image processing engine extracts the image information from data 1 and performs classification and adaptive recognition to obtain feature data 2; The dynamic insurance order generation engine merges the feature data one and the feature data two, connects to the API interface through the rule validator and the template adaptive system to automatically adapt the message format and obtain the insurance order; the collaborative inspection system identifies abnormal values of the fields in the insurance order based on the regular algorithm, and compares them with the original data through a visual interface, and manually checks and confirms the information.
2. The intelligent generation system of insurance application forms extracted based on semantic recognition as described in claim 1, wherein The BERT pre-training related operations are as follows: Obtain a corpus of insurance-related terms and build a BERT model using multiple transformer attention heads. Use the masking strategy for pre-training, randomly masking 15% of the tokens, of which 80% are replaced with mask values, 10% are replaced with random tokens, and 10% remain the original words; Construct a cross entropy loss to predict the masked words.
3. The intelligent insurance application form generation system based on semantic recognition extraction according to claim 1, wherein The classification uses YOLOv5 to classify the document types in the image information. The document types include ID cards, vehicle registration certificates, and driver's licenses. The specific operations are as follows: Construct a detection network model containing a backbone network, a neck connection network, and a detection head network; The detection network model is trained using the collected ID photo using a data augmentation method to obtain a converged detection network model; After pruning and compressing the detection network model, it is deployed to the intelligent image processing engine, and the real-time input image is used for forward reasoning to obtain the image detection and classification results.
4. The intelligent generation system of insurance application forms extracted based on semantic recognition as described in claim 1, characterized in that, The regular expression identifies abnormal field values in the insurance order, and the operation is as follows: Convert regular expressions into non-deterministic finite automata for efficient string matching; The greedy matching principle is adopted, that is, each character can correspond to multiple state transitions, and when the current path cannot complete the character matching, it backtracks to the nearest selection point to try other paths; among them, the atomic group and possessive quantifier matching search are used to control the backtracking and rematching behavior.
5. The intelligent generation system of insurance application forms extracted based on semantic recognition as described in claim 1, wherein The context correlation layer correlates the scattered insurance elements through the Attention mechanism, where the attention mechanism uses to represent the correlation relationship between the insurance elements in the feature data.
Citation Information
Patent Citations
Script data extraction method and device, computer equipment and storage medium
CN111626585A
Waste steel classification and identification method based on YOLOv5 network
CN116206155A
Fruit identification method based on lightweight neural network
CN117218643A
Artificial intelligence domain entity recognition method and system based on multi-semantic feature fusion
CN117669574A
Personalized insurance scheme customization method and device, equipment and storage medium
CN119323483A