Data enhancement and generalization method and system for vertical large model in insurance field
By structurally refining and reorganizing multi-source heterogeneous data in the insurance field, mapping semantic space, and conducting multi-dimensional quality assessments, combined with multi-turn dialogue state transition trees, the problems of inaccurate data extraction and insufficient adaptability in existing technologies have been solved, enabling efficient training and generalization of large-scale vertical models in the insurance field.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI HENGGE INFORMATION TECH CO LTD
- Filing Date
- 2026-03-09
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies in the insurance field lack a systematic structured extraction and reorganization mechanism, making it difficult to accurately extract key business entities and attribute descriptions from multi-source heterogeneous raw data. Semantic space mapping lacks accurate relational network support, resulting in insufficient effectiveness and adaptability of training data, and failing to meet the application needs of complex business scenarios.
By structurally refining and reorganizing multi-source heterogeneous raw data in the insurance field, a synonym and hierarchical relationship network is constructed, semantic space mapping is performed, and multi-dimensional quality assessment criteria and multi-turn dialogue state transition trees are combined to achieve optimal data selection and authenticity enhancement. Finally, the model is trained through an efficient parameter fine-tuning paradigm and a three-layer guided prompting engineering framework.
It significantly improved the accuracy and standardization of training data, enhanced the model's adaptability to business scenarios in the insurance field, ensured the stability and accuracy of the model in handling single-turn and multi-turn dialogues, and met the diverse and complex application needs of the insurance field.
Smart Images

Figure CN121786191B_ABST
Abstract
Description
A Data Augmentation and Generalization Method and System for Vertical Large Models in the Insurance Field Technical Field
[0001] This invention relates to the field of machine learning technology, and in particular to a data augmentation and generalization method and system for a large vertical model in the insurance field. Background Technology
[0002] Existing technologies lack a systematic structured extraction and reorganization mechanism when processing multi-source heterogeneous raw data in the insurance field. This makes it difficult to accurately extract key business entities and attribute descriptions from texts of different sources and formats, such as insurance contracts, customer service records, and internal training materials. As a result, the generated training data has problems such as information redundancy and logical confusion, and cannot provide high-quality seed data support for large vertical models.
[0003] In the process of data augmentation and generalization of large-scale models in the insurance field, existing technologies have not fully integrated the semantic relationship features within the domain for targeted expansion. Semantic space mapping lacks accurate relational network support, and the quality assessment dimension is too singular to fully verify the semantic consistency and linguistic standardization of the augmented data. Furthermore, the construction of multi-turn dialogue data does not follow the inherent logic of business processes, resulting in insufficient effectiveness and adaptability of training data. Ultimately, this limits the generalization ability of the model and fails to meet the application needs of complex business scenarios in the insurance field. Therefore, how to improve the efficiency of data augmentation and generalization of large-scale vertical models has become an urgent problem to be solved. Summary of the Invention
[0004] This invention provides a data augmentation and generalization method and system for a large vertical model in the insurance field to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, this invention provides a data augmentation and generalization method for a vertical large-scale model in the insurance field, comprising:
[0006] S1. Structure and reorganize the multi-source heterogeneous raw data in the insurance field to obtain the seed dataset of the insurance field;
[0007] S2. Based on the synonym and hyponymy relationship network of the vertical large model in the insurance field, perform semantic space mapping on the question text of the standardized question-answer pairs in the seed dataset to obtain the basic enhanced question text set of the insurance field;
[0008] S3. Based on the multi-dimensional quality evaluation criteria of the vertical large model, the basic augmentation problem text set is selected to obtain the augmentation dataset in the insurance field.
[0009] S4. Based on the multi-turn dialogue state transition tree of the vertical large model, perform realism polishing on the question-answer pairs in the augmented dataset to obtain the multi-turn dialogue training data in the insurance field.
[0010] S5. Perform attribute adaptation fusion between the single-turn question-answering data in the augmented dataset and the multi-turn dialogue training data, and perform hierarchical labeling according to the business scenario classification system and dialogue complexity definition rules of the vertical large model to obtain the final training corpus of the insurance field.
[0011] S6. Based on the final training corpus, the preset parameter efficient fine-tuning paradigm, and the three-layer guided prompting engineering framework, the training process of the basic large language model in the insurance field is collaboratively shaped to obtain the generalized vertical large model in the insurance field.
[0012] In a preferred embodiment, the step of structurally refining and reorganizing the multi-source heterogeneous raw data in the insurance field to obtain the seed dataset for the insurance field includes:
[0013] By collecting the textual content of insurance contract documents, customer service records, and internal training materials in the insurance field, multi-source heterogeneous raw data in the insurance field is obtained.
[0014] Core information is extracted from the multi-source heterogeneous raw data to obtain purified text materials in the insurance field.
[0015] Based on the list of insurance business entities and relationships in the vertical large model of the insurance field, the purified text material is scanned in a targeted manner to identify the key business entities and attribute descriptions in the insurance field;
[0016] Based on the question-answering logic framework of the vertical large model, the key business entities and attribute descriptions are structured and arranged to obtain the seed dataset of the insurance field.
[0017] In a preferred embodiment, the semantic space mapping of the question texts of standardized question-answer pairs in the seed dataset, based on the synonym and hyponym relationship network of the vertical large model in the insurance domain, is performed to obtain the basic enhanced question text set of the insurance domain, including:
[0018] Based on the synonym glossary and concept hierarchy of the vertical large model in the insurance field, a relational knowledge base for the insurance field is constructed.
[0019] Based on the term entries and association network in the relational knowledge base, term index matching is performed on the question text of standardized question-answer pairs in the seed dataset to obtain specific business terms in the insurance field.
[0020] Semantic expansion retrieval is performed on the specific business term to obtain synonyms and subordinate and superior concept expressions of the specific business term;
[0021] Based on the grammar and context of the specific business terms, the synonymous expressions and the lower-level and higher-level concept expressions are rewritten in a contextualized manner to obtain new problem texts in the insurance field;
[0022] The new problem texts are collected and organized to obtain the basic enhanced problem text set in the insurance field.
[0023] In a preferred embodiment, the multi-dimensional quality assessment criteria based on the vertical large model are used to select the best from the basic augmentation problem text set to obtain the augmentation dataset in the insurance field, including:
[0024] Based on the semantic consistency judgment rules of the multi-dimensional quality assessment criteria in the vertical large model, the compliance judgment is performed on the basic enhanced question text set to obtain the semantic consistency verification result in the insurance field.
[0025] Based on the language standardization judgment rules in the multi-dimensional quality assessment criteria, the expression review of the basic enhancement problem text set is carried out to obtain the grammatical compliance verification results in the insurance field.
[0026] In the basic enhanced question text set, new question texts that pass both the semantic consistency check result and the syntax compliance check result are determined to be high-confidence texts in the insurance field;
[0027] The high-confidence text is paired and recombined with the original answer of the high-confidence text, and new question texts that are not judged as high-confidence texts are removed to form the augmented dataset in the insurance field.
[0028] In a preferred embodiment, the step of applying realism-enhancing effects to the question-answer pairs in the augmented dataset based on the multi-turn dialogue state transition tree of the vertical large model to obtain multi-turn dialogue training data for the insurance domain includes:
[0029] Based on the standardized process knowledge in the insurance field, the internal logical structure of multi-turn dialogue in the vertical large model is formally modeled to obtain the state transition tree of multi-turn dialogue in the insurance field.
[0030] Based on the business stage features in the multi-turn dialogue state transition tree, intent parsing is performed on the question-answer pairs in the enhanced dataset, and the dialogue stage of the question-answer pairs is determined according to the business domain of the question-answer pairs, thereby obtaining the logical location information of the insurance domain.
[0031] Based on the process order rules and state transition logic in the multi-turn dialogue state transition tree and the logical position information, the augmented dataset is logically skeletonized to obtain the multi-turn dialogue content framework in the insurance field.
[0032] Within the multi-turn dialogue content framework, the question-and-answer content of the augmented dataset is calibrated according to the business language habits in the insurance field to obtain multi-turn dialogue training data for the insurance field.
[0033] In a preferred embodiment, the formal modeling of the inherent logical structure of multi-turn dialogues in the vertical large model based on standardized process knowledge in the insurance field, to obtain the state transition tree of multi-turn dialogues in the insurance field, includes:
[0034] The standardized process knowledge in the insurance field is decomposed and sorted out to obtain the dialogue state nodes in the insurance field.
[0035] Based on the inter-step dependencies and advancement rules in the insurance business operation procedures, the dialogue state nodes are conditionally defined to obtain the transfer path in the insurance field.
[0036] Based on the weight information elements in the transfer path, information load requirements are added to the internal logical structure of the multi-turn dialogue in the vertical large model to establish the information dependency relationship in the insurance field.
[0037] A topological construction is performed on the dialogue state nodes, the transition paths, and the information dependencies to obtain a multi-turn dialogue state transition tree for the insurance domain.
[0038] In a preferred embodiment, the formula for calculating the path weight in the transfer path is as follows:
[0039] ;
[0040] In the formula, For the nodes in the dialogue state node To the node Path weights, For the nodes in the dependency relationship To the node The tightness of business logic, For the dependent node in the above dependency relationship To the node Information dependency satisfaction For the slave node in the dialogue state node To the node The number of dialogue state nodes on the path. This parameter represents the maximum allowed number of dialogue turns in the vertical large model. The preset business logic density weighting coefficient, The preset information dependency satisfaction weight coefficients, The preset path simplicity penalty strength coefficient, It is a natural constant.
[0041] In a preferred embodiment, the single-turn question-answering data and the multi-turn dialogue training data in the augmented dataset are fused using attribute adaptation, and then layered and labeled according to the business scenario classification system and dialogue complexity definition rules of the vertical large model to obtain the final training corpus for the insurance domain, including:
[0042] The format of the single-turn question-answering data and the multi-turn dialogue training data in the augmented dataset is standardized to obtain the original data collection in the insurance field.
[0043] Based on the business scenario classification system of the vertical big model, the business types of the data collection are classified and mapped to obtain the business scenario tags of the insurance field.
[0044] Based on the dialogue complexity definition rules of the vertical large model, the data collection is hierarchically divided to obtain the complexity level labels of the insurance field.
[0045] The data content of the business scenario labels and the complexity level labels are archived and organized to obtain the final training corpus for the insurance field.
[0046] In a preferred embodiment, the training process of the basic large language model in the insurance domain is collaboratively shaped based on the final training corpus, a preset parameter-efficient fine-tuning paradigm, and a three-layer guided prompting engineering framework to obtain a generalized vertical large model in the insurance domain, including:
[0047] The final training corpus will be used as the core training content for the insurance field.
[0048] Based on the efficient fine-tuning of parameters of the vertical large model and the specific requirements of model stability and adaptation efficiency in the insurance field, the parameter update range and adjustment rules in the training process of the vertical large model are explicitly designed and specified to obtain the operating procedures in the insurance field.
[0049] The functional positioning, compliance requirements, and standardized output specifications of customer service assistants in the insurance field are hierarchically coded and encapsulated to obtain a three-layer guided benchmark for the insurance field;
[0050] Based on the three-layer guided benchmark, the operating procedures, and the core training content, the vertical large model is updated in a paradigm-coordinated manner to obtain the generalized vertical large model in the insurance field.
[0051] To address the aforementioned problems, this invention also provides a data augmentation and generalization system for a large vertical model in the insurance field, the system comprising:
[0052] The data structuring module is used to extract and reorganize multi-source heterogeneous raw data in the insurance field into a structured dataset to obtain the seed dataset for the insurance field.
[0053] The semantic mapping enhancement module is used to perform semantic space mapping on the question texts of standardized question-answer pairs in the seed dataset based on the synonym and hyponymy relationship network of the vertical large model in the insurance field, so as to obtain the basic enhanced question text set in the insurance field.
[0054] The quality selection module is used to select the best from the basic augmentation problem text set based on the multi-dimensional quality evaluation criteria of the vertical large model, so as to obtain the augmentation dataset in the insurance field.
[0055] The dialogue state polishing module is used to polish the question-answer pairs in the augmented dataset to a realistic level based on the multi-turn dialogue state transition tree of the vertical large model, so as to obtain the multi-turn dialogue training data in the insurance field.
[0056] The fusion layered labeling module is used to perform attribute adaptation fusion of single-turn question-answering data and multi-turn dialogue training data in the augmented dataset, and to perform layered labeling according to the business scenario classification system and dialogue complexity definition rules of the vertical large model to obtain the final training corpus of the insurance field.
[0057] The collaborative fine-tuning and shaping module is used to collaboratively shape the training process of the basic large language model in the insurance domain based on the final training corpus, the preset parameter efficient fine-tuning paradigm, and the three-layer guided prompting engineering framework, so as to obtain the generalized vertical large model in the insurance domain.
[0058] Compared with the prior art, the present invention has the following beneficial effects:
[0059] 1. This invention systematically and structurally extracts and reorganizes multi-source heterogeneous raw data in the insurance field, accurately extracts key business entities and attribute descriptions, combines synonym and hierarchical relationship networks to achieve semantic space mapping, and then selects the best through multi-dimensional quality evaluation criteria, effectively improving the accuracy, richness and standardization of training data, ensuring that the augmented dataset has both sufficient semantic coverage and strict conformity to the business logic and language norms of the insurance field.
[0060] 2. This invention relies on a multi-turn dialogue state transition tree to refine the authenticity of question-answer pairs. It combines a business scenario classification system and dialogue complexity rules to achieve attribute adaptation fusion and hierarchical labeling of data. Furthermore, through the collaborative shaping of an efficient parameter fine-tuning paradigm and a three-layer guided prompting engineering framework, it significantly improves the adaptability of the final generalized vertical large model to business scenarios in the insurance field. It enhances the stability and accuracy of the model in handling single-turn and multi-turn dialogues, fully meeting the diverse and complex application needs of the insurance field. Attached Figure Description
[0061] Figure 1 is a flowchart illustrating a data augmentation and generalization method for a vertical large model in the insurance field according to an embodiment of the present invention;
[0062] Figure 2 is a functional block diagram of a data augmentation and generalization system for a vertical large model in the insurance field provided by an embodiment of the present invention;
[0063] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0064] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0065] This application provides a data augmentation and generalization method for a large vertical model in the insurance field. The execution entity of this data augmentation and generalization method for a large vertical model in the insurance field includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the data augmentation and generalization method for a large vertical model in the insurance field can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0066] Referring to Figure 1, this is a flowchart illustrating a data augmentation and generalization method for a vertical large-scale model in the insurance field according to an embodiment of the present invention. In this embodiment, the data augmentation and generalization method for a vertical large-scale model in the insurance field includes:
[0067] S1. Structure and reorganize the multi-source heterogeneous raw data in the insurance field to obtain the seed dataset of the insurance field;
[0068] In this embodiment of the invention, the step of structurally refining and reorganizing multi-source heterogeneous raw data in the insurance field to obtain the seed dataset for the insurance field includes:
[0069] By collecting the textual content of insurance contract documents, customer service records, and internal training materials in the insurance field, multi-source heterogeneous raw data in the insurance field is obtained.
[0070] Core information is extracted from the multi-source heterogeneous raw data to obtain purified text materials in the insurance field.
[0071] Based on the list of insurance business entities and relationships in the vertical large model of the insurance field, the purified text material is scanned in a targeted manner to identify the key business entities and attribute descriptions in the insurance field;
[0072] Based on the question-answering logic framework of the vertical large model, the key business entities and attribute descriptions are structured and arranged to obtain the seed dataset of the insurance field.
[0073] We collected relevant texts in various written and recorded forms from insurance business scenarios, including the main text and attachments of formally signed insurance contracts, written records of communication between customers and service personnel, and internal training materials and handouts for employees. We then integrated all these texts from different sources and in different formats to form multi-source heterogeneous raw data in the insurance field.
[0074] Each text in the integrated multi-source heterogeneous raw data is reviewed sentence by sentence to remove redundant information, repetitive expressions and meaningless statements that are irrelevant to the core content of insurance business. Key information such as insurance product terms, business processing procedures, customer consultation questions and training core points are accurately extracted. The extracted information is then organized into coherent text fragments according to the original business relationship logic of the text to obtain purified text materials in the insurance field.
[0075] First, clarify the various business-related entities, business projects, and their relationships included in the insurance business entity and relationship list. Based on this list, set up a keyword library and identification rules for targeted scanning. Then, compare and analyze each text fragment in the purified text material word by word to accurately locate the key business entities such as business entities and business projects in the list. At the same time, extract the attribute descriptions that describe the characteristics, attribute parameters, and applicable conditions of these entities to clearly identify the key business entities and attribute descriptions in the insurance field.
[0076] Following the question-and-answer structure, information presentation order, and expression standards defined by the question-and-answer logic framework of the vertical large model, the identified key business entities are used as core elements, paired with corresponding attribute descriptions, to construct standardized expressions that conform to the question-and-answer interaction scenario. This ensures that each piece of content contains clear business-related questions and corresponding core information answers. All standardized question-and-answer related content is systematically organized and archived to obtain a seed dataset for the insurance field.
[0077] The beneficial effects are as follows: by comprehensively collecting various relevant textual content in the insurance field, the integrity and coverage of multi-source heterogeneous original data are ensured. After extracting core information and removing irrelevant and redundant content, the purified text material is obtained, ensuring the refinement of the data. Subsequently, the purified text material is targeted by scanning the list of insurance business entities and relationships, which can accurately identify key business entities and attribute descriptions. Finally, it is structured and arranged according to the question-and-answer logic framework, so that the generated seed dataset not only conforms to the business logic of the insurance field, but also has a standardized question-and-answer structure. This provides high-quality and highly adaptable basic data support for the subsequent training of large-scale vertical models in the insurance field, effectively improving data utilization efficiency and the quality of data preparation in the early stage of model training.
[0078] S2. Based on the synonym and hyponymy relationship network of the vertical large model in the insurance field, perform semantic space mapping on the question text of the standardized question-answer pairs in the seed dataset to obtain the basic enhanced question text set of the insurance field;
[0079] In this embodiment of the invention, the semantic space mapping of the question texts of standardized question-answer pairs in the seed dataset based on the synonym and hyponym relationship network of the vertical large model in the insurance domain, to obtain the basic enhanced question text set of the insurance domain, includes:
[0080] Based on the synonym glossary and concept hierarchy of the vertical large model in the insurance field, a relational knowledge base for the insurance field is constructed.
[0081] Based on the term entries and association network in the relational knowledge base, term index matching is performed on the question text of standardized question-answer pairs in the seed dataset to obtain specific business terms in the insurance field.
[0082] Semantic expansion retrieval is performed on the specific business term to obtain synonyms and subordinate and superior concept expressions of the specific business term;
[0083] Based on the grammar and context of the specific business terms, the synonymous expressions and the lower-level and higher-level concept expressions are rewritten in a contextualized manner to obtain new problem texts in the insurance field;
[0084] The new problem texts are collected and organized to obtain the basic enhanced problem text set in the insurance field.
[0085] Based on the established synonym glossary and concept hierarchy in the insurance field, we first systematically analyze each term in the synonym glossary, clarify all synonymous expressions corresponding to each term and establish their associations. Then, we analyze the hierarchical relationships of each concept in the concept hierarchy, determine the superordinate and subordinate concepts of each concept, and finally integrate the synonym association information with the concept hierarchy to construct a relational knowledge base for the insurance field that includes synonym associations of terms and hierarchical associations of concepts.
[0086] Extract all terminology entries and established association networks from the relational knowledge base. Segment the question text of standardized question-answer pairs in the seed dataset sentence by sentence to extract business-related terms. Compare these extracted terms with terminology entries in the relational knowledge base one by one to find terms that are completely matched or semantically consistent. Through this precise indexing and matching method, specific business terms in the insurance field are obtained.
[0087] Using a specific business term as the core of the search, all synonyms corresponding to the term are searched in the relational knowledge base. At the same time, according to the concept hierarchy, the superordinate concept and subordinate concept are traced to its parent and child concepts. The retrieved expressions are comprehensively collected to ensure that no key related expressions are missed, and finally, the synonyms and subordinate superordinate concepts of the specific business term are obtained.
[0088] By deeply analyzing the grammatical functions of specific business terms in the original question text, clarifying their components in sentences and their collocation relationships with other words, and combining them with the insurance business scenario context involved in the original question, the specific business terms in the original question are replaced with synonymous expressions and subordinate-superior concept expressions, and the word order and wording of the sentences are adjusted to make the replaced sentences grammatically correct, semantically coherent and in line with the expression habits of the insurance field, thus obtaining a new question text in the insurance field.
[0089] All generated new question texts are collected, and the text content is repeatedly screened to remove identical new question texts. The remaining new question texts are then classified and organized according to the core categories of insurance business or the degree of semantic relevance to ensure that the text set has a clear structure and complete content, ultimately resulting in a basic enhanced question text set for the insurance field.
[0090] The beneficial effects are as follows: By constructing a relational knowledge base based on a synonym glossary and concept hierarchy structure in the insurance field, a solid domain knowledge support is provided for semantic space mapping. Relying on this knowledge base for term indexing and matching, specific business terms in standardized question-answer pairs in seed data can be accurately located. Then, through semantic expansion retrieval, synonymous expressions and hypernym / hyper ...
[0091] S3. Based on the multi-dimensional quality evaluation criteria of the vertical large model, the basic augmentation problem text set is selected to obtain the augmentation dataset in the insurance field.
[0092] In this embodiment of the invention, the multi-dimensional quality assessment criteria based on the vertical large model are used to select the best from the basic augmentation problem text set to obtain the augmentation dataset in the insurance field, including:
[0093] Based on the semantic consistency judgment rules of the multi-dimensional quality assessment criteria in the vertical large model, the compliance judgment is performed on the basic enhanced question text set to obtain the semantic consistency verification result in the insurance field.
[0094] Based on the language standardization judgment rules in the multi-dimensional quality assessment criteria, the expression review of the basic enhancement problem text set is carried out to obtain the grammatical compliance verification results in the insurance field.
[0095] In the basic enhanced question text set, new question texts that pass both the semantic consistency check result and the syntax compliance check result are determined to be high-confidence texts in the insurance field;
[0096] The high-confidence text is paired and recombined with the original answer of the high-confidence text, and new question texts that are not judged as high-confidence texts are removed to form the augmented dataset in the insurance field.
[0097] The semantic consistency judgment rules included in the multi-dimensional quality assessment criteria in the vertical large model are clearly defined. These rules are formulated around the core business semantics of the insurance field and specify the specific requirement that the new question text must be consistent with the core business meaning of the original standardized question and answer pair. Semantic analysis is performed on each new question text in the basic enhanced question text set, and the core business semantics expressed are compared with the core semantics of the corresponding original standardized question and answer pair. If the semantics are completely consistent, the verification is passed. If there are semantic deviations or deviations, the verification is failed. Finally, the semantic consistency verification result of the insurance field is formed.
[0098] Based on the language standardization judgment rules set in the multi-dimensional quality assessment criteria, which cover specific requirements for grammatical correctness, standardized word choice, fluent expression, and conformity with common expression habits in the insurance field, each new question text in the basic enhancement question text set is reviewed sentence by sentence to check for grammatical errors, whether the word choice conforms to the professional standards of the insurance field, and whether the expression of the sentences is logically fluent. If all the above requirements are fully met, the verification is deemed to have passed; if any one of them is not met, the verification is deemed to have failed, thus obtaining the grammatical compliance verification result in the insurance field.
[0099] In the basic enhanced question text set, the semantic consistency check result and the grammatical compliance check result corresponding to each new question text are extracted one by one. The two check results are checked simultaneously. Only when a new question text passes both the semantic consistency check and the grammatical compliance check is it clearly determined as a high-confidence text in the insurance field, ensuring that the judgment process strictly follows the sole standard of passing both checks.
[0100] The original answers in the original standardized question-answer pairs corresponding to each high-confidence text are traced back. According to the correspondence between "question and answer", the high-confidence texts and the matching original answers are accurately paired and recombined to form a standardized question-answer pair structure. At the same time, new question texts that are not judged as high-confidence texts in the basic enhanced question text set are comprehensively screened and completely removed. All the paired and recombined compliant question-answer pairs are systematically organized and collected to form an enhanced dataset in the insurance field.
[0101] The beneficial effects are as follows: By conducting compliance judgments based on the semantic consistency judgment rules in the multi-dimensional quality assessment criteria, it is ensured that the new question texts in the basic augmented question text set are consistent with the core business semantics of the original standardized question-answer pairs. Then, by combining the expression review with the language standardization judgment rules, it is ensured that the text expression conforms to the grammatical norms and professional expression habits in the insurance field. Through double verification, high-confidence texts are screened out, and unqualified texts with semantic deviations and grammatical errors are effectively removed. Subsequently, the high-confidence texts are accurately paired and recombined with the original answers, and unqualified texts are removed. The resulting augmented dataset has high reliability, strong standardization, and good business adaptability, providing high-quality and compliant core data support for the subsequent training of large-scale vertical models in the insurance field, and improving the utilization value of training data and the effectiveness of model training.
[0102] S4. Based on the multi-turn dialogue state transition tree of the vertical large model, perform realism polishing on the question-answer pairs in the augmented dataset to obtain the multi-turn dialogue training data in the insurance field.
[0103] In this embodiment of the invention, the step of applying realism-enhancing effects to the question-answer pairs in the augmented dataset based on the multi-turn dialogue state transition tree of the vertical large model to obtain multi-turn dialogue training data in the insurance domain includes:
[0104] Based on the standardized process knowledge in the insurance field, the internal logical structure of multi-turn dialogue in the vertical large model is formally modeled to obtain the state transition tree of multi-turn dialogue in the insurance field.
[0105] Based on the business stage features in the multi-turn dialogue state transition tree, intent parsing is performed on the question-answer pairs in the enhanced dataset, and the dialogue stage of the question-answer pairs is determined according to the business domain of the question-answer pairs, thereby obtaining the logical location information of the insurance domain.
[0106] Based on the process order rules and state transition logic in the multi-turn dialogue state transition tree and the logical position information, the augmented dataset is logically skeletonized to obtain the multi-turn dialogue content framework in the insurance field.
[0107] Within the multi-turn dialogue content framework, the question-and-answer content of the augmented dataset is calibrated according to the business language habits in the insurance field to obtain multi-turn dialogue training data for the insurance field.
[0108] Based on standardized process knowledge in the insurance field, the inherent logical structure of multi-turn dialogues in the vertical large model is formally modeled to obtain the state transition tree of multi-turn dialogues in the insurance field, including:
[0109] The standardized process knowledge in the insurance field is decomposed and sorted out to obtain the dialogue state nodes in the insurance field.
[0110] Based on the inter-step dependencies and advancement rules in the insurance business operation procedures, the dialogue state nodes are conditionally defined to obtain the transfer path in the insurance field.
[0111] Based on the weight information elements in the transfer path, information load requirements are added to the internal logical structure of the multi-turn dialogue in the vertical large model to establish the information dependency relationship in the insurance field.
[0112] A topological construction is performed on the dialogue state nodes, the transition paths, and the information dependencies to obtain a multi-turn dialogue state transition tree for the insurance domain.
[0113] The formula for calculating the path weight in the transfer path is as follows:
[0114] ;
[0115] In the formula, For the nodes in the dialogue state node To the node Path weights, For the nodes in the dependency relationship To the node The tightness of business logic, For the dependent node in the above dependency relationship To the node Information dependency satisfaction For the slave node in the dialogue state node To the node The number of dialogue state nodes on the path. This parameter represents the maximum allowed number of dialogue turns in the vertical large model. The preset business logic density weighting coefficient, The preset information dependency satisfaction weight coefficients, The preset path simplicity penalty strength coefficient, It is a natural constant.
[0116] This study delves into the standardized processes in the insurance industry, covering the standardized operating procedures and process connections for various businesses such as insurance application, claims processing, policy inquiry, and terms consultation. Each business link is broken down into independent dialogue state nodes. Based on the dependencies and advancement rules between steps in the business operation procedures, the triggering conditions and constraints for each node to transfer to other nodes are clarified. Information dependencies are then established in conjunction with the information transmission needs between nodes. Finally, through a topology construction method, the dialogue state nodes, transfer paths, and information dependencies are integrated into a multi-turn dialogue state transition tree for the insurance industry with clear logical hierarchy and transfer rules.
[0117] The core features of each business stage in the multi-turn dialogue state transition tree are extracted, including business objectives, key operations, and information needs. The core intent of each question-answer pair in the augmented dataset is analyzed one by one to clarify the business demands and core content carried by the question-answer pair. Then, the insurance business domain category to which the question-answer pair belongs is accurately matched with the business stage features in the transition tree to determine the specific business stage corresponding to the question-answer pair in the entire multi-turn dialogue process, thereby obtaining the logical position information of the insurance domain.
[0118] The pre-defined process sequence rules and state transition logic in the multi-turn dialogue state transition tree are retrieved. The process sequence rules clarify the order of advancement of insurance business dialogues, and the state transition logic defines the rational basis for the jump between nodes. Combined with the obtained logical position information, the question-answer pairs in the enhanced dataset are arranged according to the corresponding business stage order. Based on the state transition logic, the connection relationship between the question-answer pairs is sorted out, and a multi-turn dialogue skeleton that conforms to the internal logic of the insurance business process is constructed, resulting in a multi-turn dialogue content framework in the insurance field.
[0119] We collect standard scripts, professional expression habits, and compliant expression norms formed in daily business communication in the insurance field. Based on the multi-round dialogue content framework, we adjust and optimize the wording and expression of the questions and answers within the framework one by one to ensure that the questions and answers not only conform to the professional terminology usage norms in the insurance field, but also fit the language habits in actual business communication, so that the dialogue content is natural, fluent, accurate and compliant, and finally obtain multi-round dialogue training data in the insurance field.
[0120] We comprehensively collect standardized process knowledge in the insurance field, covering the standardized processes of various core businesses such as insurance application, underwriting, claims, policy changes, and consulting services. We break down each business process into its steps, clarifying each independent and key business operation or communication link as an independent unit, defining the core business objectives, the information to be processed, and the corresponding dialogue functions of each unit, and finally obtaining the dialogue status nodes in the insurance field.
[0121] By deeply analyzing the inherent dependencies between various business steps in the insurance business operation procedures, clarifying the support requirements of the preceding steps for the subsequent steps, and sorting out the legal or industry-agreed rules for the advancement of each step, specific conditions for triggering the transfer are set for each dialogue state node based on these dependencies and advancement rules, including information completeness conditions, operation completion conditions, etc. When a node meets the corresponding conditions, it can transfer to the designated next node. Through this condition definition, the transfer path in the insurance field is formed.
[0122] Identify the weighted information elements in the transfer path. These elements include the business importance of the path, the critical level of information transmission, and the urgency of user needs. Based on the specific circumstances of these elements, clarify the information load requirements to be carried by each transfer path, define the core business information, user interaction data, and process status information that must be transmitted in the path, and clarify the dependency standards of subsequent nodes on the information transmitted by preceding nodes, thereby establishing information dependency relationships in the insurance field.
[0123] All dialogue state nodes are arranged according to the logical order of the transition paths. The direction and content of information transmission between nodes are clarified based on information dependencies. A topological structure is constructed, with nodes as the core, transition paths as the connections, and information dependencies as constraints, to build a hierarchical and logically coherent tree structure. This ensures that the transition path of each node is unique and conforms to the business process logic, and that each information dependency is reflected in the structure, ultimately resulting in a multi-turn dialogue state transition tree for the insurance field.
[0124] The tightness of business logic stems from an in-depth analysis of the dependencies between nodes in the operational procedures of the insurance industry. Focusing on core business processes such as insurance application, claims processing, and policy changes, each business step is broken down into its operational purpose and execution conditions. This clarifies whether a preceding node is a necessary prerequisite for subsequent nodes or merely provides auxiliary support. By determining whether the completion of a preceding node directly determines the initiation of subsequent nodes and whether business operations revolve around the same core business objective, the depth of business connections between nodes is analyzed. For example, in the claims process, the "claims application submission" node is the sole prerequisite for initiating the "claims information review" node; the two have a very high degree of business connection, thus clarifying the tightness of the business logic between nodes.
[0125] Information dependency satisfaction is determined based on the information load requirements of the transfer paths between nodes. First, according to insurance business rules, the core business information that must be transmitted in each transfer path is identified, including policy number, insured information, accident details, etc. User interaction data covers user-inputted inquiries and confirmation instructions, while process status information includes node completion status and review results. Simultaneously, dependency standards for information completeness and standardization are established. Each type of information transmitted from preceding nodes to subsequent nodes is checked item by item to confirm whether the information completely covers the required items and conforms to the data format and accuracy requirements of the insurance field. No missing or biased information is considered fully satisfied; partial missing or slight bias is considered partially satisfied; and severely missing or significantly biased information is considered unsatisfied. Based on this, the corresponding degree of satisfaction is determined.
[0126] The number of dialogue state nodes on the path is directly counted as the actual number of independent dialogue state nodes included in the transfer path from one dialogue state node to another. In the counting process, the starting node of the transfer path is taken as the starting point and the target node is taken as the ending point. Each node corresponding to each independent business link passed through in the path is enumerated one by one. The same node is not counted repeatedly, ensuring that the statistical results accurately reflect the total number of nodes covered by the path.
[0127] The maximum allowed number of dialogue rounds is a pre-set upper limit for dialogue rounds in the vertical large model, which is designed based on common scenarios of multi-turn dialogues in the insurance field, including complex claims consultations and interpretation of policy terms. By analyzing a large amount of historical dialogue data, the range of dialogue rounds required to complete user needs in different business scenarios is determined. At the same time, the user's patience threshold in business consultation and the efficiency requirements of business processing are taken into account. This ensures that the dialogue can fully cover business needs without causing a decline in user experience due to too many rounds.
[0128] The business logic tightness weight coefficient and information dependency satisfaction weight coefficient are ranked according to the importance of insurance business, distinguishing the priority of core businesses such as claims and underwriting from auxiliary businesses such as consultation and inquiry. Simultaneously, considering the data adaptation requirements for training large-scale vertical models, the impact of business logic tightness and information completeness on the accuracy of model dialogue is clarified. For core business processes, a higher weight coefficient is assigned to business logic tightness; for scenarios with extremely high information transmission requirements, the weight coefficient for information dependency satisfaction is increased. Values are pre-set to adjust the influence of corresponding factors in the weight calculation.
[0129] The path simplicity penalty intensity coefficient is designed to meet the efficiency requirements of dialogue processes in the insurance industry. It references the optimal number of nodes in multi-turn dialogues within the industry, considers user expectations for business processing efficiency, and avoids excessively long dialogue processes and user frustration due to too many path nodes. For transfer paths with more nodes than a reasonable range, corresponding penalty standards are set; the more nodes and the longer the path, the greater the penalty. Pre-set values for penalizing such transfer paths ensure that the final selected paths are both logically sound and efficient.
[0130] Natural constants are fixed constants in the field of mathematics. Their values are fixed and do not change with the business scenarios or calculation conditions in the insurance field. In the process of weight calculation, they are used to adjust the numerical range of various calculation results, so that the calculation results of different dimensions can be effectively integrated and compared, adapt to the needs of logical operations, and ensure the rationality and comparability of the final weight values.
[0131] The core significance of this calculation is to quantify the rationality and priority of the transfer paths between dialogue state nodes. By comprehensively considering the tightness of business logic, it ensures that the path conforms to the inherent order of insurance business operations, considers the satisfaction of information dependence, ensures the effectiveness of information transmission in the path, considers the simplicity of the path, and takes into account the efficiency of the dialogue. The evaluation results of these three dimensions are integrated through quantitative calculation to obtain the weight value of each transfer path.
[0132] The weight values directly reflect the adaptability and priority of transition paths in multi-turn dialogue processes within the insurance field. Paths with higher weights better align with business operation logic and dialogue progression needs. In the topology construction of the multi-turn dialogue state transition tree, high-weight paths are prioritized as core backbone paths, while low-weight paths serve as alternative paths or optimization targets, providing a crucial basis for the topology construction of the transition tree. This process ensures that the constructed transition tree accurately matches the inherent logic of the insurance business process, thereby guaranteeing that the multi-turn dialogue training data, after being realistically refined based on this transition tree, possesses rigorous logical coherence in dialogue order, information transmission, and business connections. Simultaneously, it is highly adaptable to insurance business scenarios, improving the accuracy and rationality of multi-turn dialogues in vertical large-scale models.
[0133] The beneficial effect is that a multi-turn dialogue state transition tree is constructed based on standardized process knowledge in the insurance field. By decomposing and sorting out the process to obtain dialogue state nodes, defining conditions to form transition paths, adding information loads to establish information dependencies and constructing topology, the transition tree accurately matches the internal logic of insurance business, providing solid logical support for subsequent question-and-answer pair processing. Then, based on the business stage characteristics in the transition tree, the intent of the question-and-answer pair is analyzed and the logical position is determined. The logical skeleton is arranged by combining process sequence rules and state transition logic, and the language habits of insurance business are used for word calibration. The final multi-turn dialogue training data has both rigorous logical coherence and business scenario adaptability, and conforms to the language expression norms in the field. It can provide high-quality multi-turn dialogue training materials for large-scale vertical models in the insurance field, effectively improving the accuracy and rationality of the model in handling complex multi-turn dialogues.
[0134] S5. Perform attribute adaptation fusion between the single-turn question-answering data in the augmented dataset and the multi-turn dialogue training data, and perform hierarchical labeling according to the business scenario classification system and dialogue complexity definition rules of the vertical large model to obtain the final training corpus of the insurance field.
[0135] In this embodiment of the invention, the step of performing attribute-adaptive fusion of the single-turn question-answering data and the multi-turn dialogue training data in the augmented dataset, and performing hierarchical labeling according to the business scenario classification system and dialogue complexity definition rules of the vertical large model, to obtain the final training corpus for the insurance domain, includes:
[0136] The format of the single-turn question-answering data and the multi-turn dialogue training data in the augmented dataset is standardized to obtain the original data collection in the insurance field.
[0137] Based on the business scenario classification system of the vertical big model, the business types of the data collection are classified and mapped to obtain the business scenario tags of the insurance field.
[0138] Based on the dialogue complexity definition rules of the vertical large model, the data collection is hierarchically divided to obtain the complexity level labels of the insurance field.
[0139] The data content of the business scenario labels and the complexity level labels are archived and organized to obtain the final training corpus for the insurance field.
[0140] The core components of single-turn question-and-answer data and multi-turn dialogue training data in the augmented dataset were extracted. The binary structure of "question-answer" in single-turn question-and-answer data and the multivariate structure of "turn-question-answer-context association" in multi-turn dialogue training data were clarified. A unified format specification was formulated, requiring all data to include key fields such as core business information, interaction content, and corresponding results. Context association identifiers were added to each single-turn question-and-answer data entry. The multi-turn dialogue training data was split and the independent field boundaries of each turn were clarified. It was ensured that the field names, expression formats, and information arrangement order of the two types of data were completely consistent. All data after the format was unified were integrated and collected to obtain the original data collection in the insurance field.
[0141] The core categories covered by the vertical large-scale model's business scenario classification system are clearly defined, including key business scenarios in the insurance field such as insurance consultation, claims processing, policy changes, terms interpretation, and premium inquiries. Each category has a clearly defined business scope and core characteristics. Each piece of data in the original data collection is analyzed to extract the core business actions, service requirements, and business objects involved. The analysis results are then precisely compared with the category characteristics in the business scenario classification system to find perfectly matching business scenario categories. Each piece of data is then assigned a corresponding scenario identifier, resulting in business scenario tags for the insurance field.
[0142] Based on a vertical large-scale model, the dialogue complexity definition rule explicitly uses the business logic hierarchy, information correlation strength, and user need depth involved in the dialogue as the core judgment criteria, dividing it into three distinct levels: simple, medium, and complex. Each piece of data in the original dataset undergoes complexity analysis. Single-round dialogues involving only a single business point and with no information correlation are classified as simple; multi-round dialogues involving a small amount of business correlation and clearly defined needs are classified as medium; and multi-round dialogues involving multiple layers of business logic, cross-correlation of information, and needs requiring in-depth decomposition are classified as complex. Based on the analysis results, each piece of data is assigned a corresponding level label, resulting in complexity level tags for the insurance field.
[0143] The business scenario tags and complexity level tags corresponding to each data point are bound to the data ontology to establish a correspondence of "data content - business scenario tag - complexity level tag". All bound data are first-level classified and collected according to business scenario categories. Under each business scenario category, the data is then organized into second-level hierarchical structures according to complexity level tags to ensure that data of the same category and level are stored together. At the same time, a classification index is established to clarify the storage location and relationship of data of each category and level, forming a final training corpus for the insurance field that is well-structured, clearly classified, and easy to retrieve.
[0144] The beneficial effects include unifying the format representation of single-turn question-and-answer data and multi-turn dialogue training data in the augmented dataset, effectively eliminating format differences between the two types of data, ensuring the consistency and compatibility of the original data collection, and classifying and mapping data based on the business scenario classification system of the vertical large model, so that each data can accurately match the corresponding business scenario. The resulting business scenario labels make the data classification clear and explicit. Combined with the dialogue complexity definition rules, hierarchical division is carried out, and the data is graded by complexity level labels. Then, the two types of labels and data content are archived and organized. The final training corpus in the insurance field is well-structured, accurately classified, and clearly hierarchical. It covers various business scenarios in the insurance field and includes question-and-answer data of different complexities. It can provide comprehensive, adaptable, and easily accessible high-quality corpus support for the subsequent training of basic large language models, helping the model to learn business content of different scenarios and complexities in a targeted manner, and improving the accuracy and efficiency of training.
[0145] S6. Based on the final training corpus, the preset parameter efficient fine-tuning paradigm, and the three-layer guided prompting engineering framework, the training process of the basic large language model in the insurance field is collaboratively shaped to obtain the generalized vertical large model in the insurance field.
[0146] In this embodiment of the invention, the training process of the basic large language model in the insurance field is collaboratively shaped based on the final training corpus, a preset parameter efficient fine-tuning paradigm, and a three-layer guided prompting engineering framework to obtain a generalized vertical large model in the insurance field, including:
[0147] The final training corpus will be used as the core training content for the insurance field.
[0148] Based on the efficient fine-tuning of parameters of the vertical large model and the specific requirements of model stability and adaptation efficiency in the insurance field, the parameter update range and adjustment rules in the training process of the vertical large model are explicitly designed and specified to obtain the operating procedures in the insurance field.
[0149] The functional positioning, compliance requirements, and standardized output specifications of customer service assistants in the insurance field are hierarchically coded and encapsulated to obtain a three-layer guided benchmark for the insurance field;
[0150] Based on the three-layer guided benchmark, the operating procedures, and the core training content, the vertical large model is updated in a paradigm-coordinated manner to obtain the generalized vertical large model in the insurance field.
[0151] The final training corpus in the insurance field was comprehensively reviewed, clarifying the various business scenario tags, complexity level tags, and corresponding data content. The training priority of different categories of data in the corpus was determined to ensure that the training process can cover core business and auxiliary business, simple scenarios and complex scenarios in sequence. The reviewed complete corpus was directly used as the core training content for training the basic large language model in the insurance field.
[0152] The core idea behind the efficient fine-tuning of vertical large model parameters is to focus on core model components that are relevant to insurance business adaptation. Combining the requirements of the insurance field for model stability and adaptation efficiency, the range of updatable parameters in the model is defined one by one. The upper limit of the adjustment range, the adjustment frequency, and the business data matching conditions that trigger the adjustment for each updatable parameter are clarified. These definitions are systematically organized in plain text to form operating procedures in the insurance field.
[0153] This study delves into the core functional positioning of customer service assistants in the insurance field, clarifies the specific responsibilities the model must undertake, such as business consultation, process guidance, and information verification, outlines the compliance expression boundaries required by insurance industry regulations and company business systems, formulates compliance principles such as truthfulness, accuracy, and non-misleading content in responses, and clarifies standardized output norms such as tone, sentence structure, and information presentation order. Based on three levels—functional positioning, compliance requirements, and standardized output norms—these contents are transformed into structured rules that the model can follow, coded, integrated, and encapsulated to obtain a three-layered guiding benchmark in the insurance field.
[0154] Based on the core training content in the final training corpus, the training process of a basic large language model for the insurance field is initiated. During training, the established operating procedures are strictly followed, and the range and magnitude of parameter updates are precisely controlled to ensure that parameter adjustments meet the requirements of model stability and adaptation efficiency. At the same time, a three-layer guided benchmark is embedded in the training process to constrain the output direction of the model in real time, ensuring that the model's functional performance, compliance performance, and output format all meet the preset standards. Through the synergistic effect of the core training content, operating procedures, and three-layer guided benchmark, the model is continuously iterated and optimized, ultimately resulting in a generalized vertical large model for the insurance field that is adapted to various business scenarios and has stable output capabilities.
[0155] The beneficial effects include using a well-structured and precisely categorized final training corpus as the core training content, providing the model with high-quality learning materials that comprehensively cover business scenarios and varying degrees of complexity in the insurance field. Operational procedures based on efficient parameter fine-tuning and domain-specific requirements enable precise control over the parameter update range and adjustment rules during training, ensuring the stability and adaptability of model training. A three-layered guided benchmark, formed by hierarchically encoding and encapsulating the customer service assistant's functional positioning, compliance requirements, and standardized output specifications, defines clear compliance boundaries and standards for model output. Through the synergistic effect of these three elements, the basic large-scale language model is trained and shaped, resulting in a generalized vertical large-scale model for the insurance field that deeply aligns with the business logic and language habits of the insurance industry, possesses stable output performance and strong generalization capabilities, and can accurately respond to the interaction needs of various business scenarios, ensuring the compliance, accuracy, and standardization of the output content.
[0156] Figure 2 shows a functional block diagram of a data augmentation and generalization system for a vertical large model in the insurance field provided by an embodiment of the present invention.
[0157] The data augmentation and generalization system 100 for a vertical large-scale model in the insurance field described in this invention can be installed in an electronic device. Depending on the functions implemented, the data augmentation and generalization system 100 for a vertical large-scale model in the insurance field may include a data structuring module 101, a semantic mapping enhancement module 102, a quality selection module 103, a dialogue state polishing module 104, a fusion layered tagging module 105, and a collaborative fine-tuning shaping module 106. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0158] In this embodiment, the functions of each module / unit are as follows:
[0159] The data structuring module 101 is used to perform structured extraction and reorganization of multi-source heterogeneous raw data in the insurance field to obtain the seed dataset of the insurance field.
[0160] The semantic mapping enhancement module 102 is used to perform semantic space mapping on the question texts of standardized question-answer pairs in the seed dataset based on the synonym and hyponymy relationship network of the vertical large model in the insurance field, so as to obtain the basic enhanced question text set in the insurance field.
[0161] The quality selection module 103 is used to select the best from the basic enhancement problem text set based on the multi-dimensional quality evaluation criteria of the vertical large model to obtain the enhanced dataset in the insurance field.
[0162] The dialogue state polishing module 104 is used to polish the question-answer pairs in the augmented dataset according to the multi-turn dialogue state transition tree of the vertical large model, so as to obtain the multi-turn dialogue training data in the insurance field.
[0163] The fusion layer labeling module 105 is used to perform attribute adaptation fusion of single-turn question-answering data and multi-turn dialogue training data in the augmented dataset, and to perform layer labeling according to the business scenario classification system and dialogue complexity definition rules of the vertical large model to obtain the final training corpus of the insurance field.
[0164] The collaborative fine-tuning and shaping module 106 is used to collaboratively shape the training process of the basic large language model in the insurance field based on the final training corpus, the preset parameter efficient fine-tuning paradigm and the three-layer guided prompting engineering framework, so as to obtain the generalized vertical large model in the insurance field.
[0165] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0166] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0167] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0168] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0169] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A data augmentation and generalization method for a large vertical model in the insurance field, characterized in that, The method includes: S1, performing structured extraction and reorganization of multi-source heterogeneous raw data in the insurance field to obtain a seed dataset for the insurance field; S2, based on the synonym and hierarchical relationship network of the vertical large model in the insurance field, performing semantic space mapping on the question texts of standardized question-answer pairs in the seed dataset to obtain a basic enhanced question text set for the insurance field; S3, based on the multi-dimensional quality evaluation criteria of the vertical large model, selecting the best from the basic enhanced question text set to obtain the enhanced dataset for the insurance field; S4, according to the multi-turn dialogue state transition tree of the vertical large model, processing the enhanced... S5. The question-answer pairs in the enhanced dataset are augmented for realism to obtain multi-turn dialogue training data for the insurance domain; S6. The single-turn question-answer data in the enhanced dataset are fused with the multi-turn dialogue training data for attribute adaptation, and layered labeling is performed according to the business scenario classification system and dialogue complexity definition rules of the vertical big model to obtain the final training corpus for the insurance domain; S7. Based on the final training corpus, the preset parameter efficient fine-tuning paradigm and the three-layer guided prompting engineering framework, the training process of the basic big language model in the insurance domain is collaboratively shaped to obtain the generalized vertical big model for the insurance domain.
2. The data augmentation and generalization method for a vertical large-scale model in the insurance field as described in claim 1, characterized in that, The process of structuring and reorganizing multi-source heterogeneous raw data in the insurance field to obtain a seed dataset for the insurance field includes: collecting text content from insurance contract documents, customer service records, and internal training materials in the insurance field to obtain multi-source heterogeneous raw data; extracting core information from the multi-source heterogeneous raw data to obtain purified text materials for the insurance field; performing targeted scanning on the purified text materials based on the list of insurance business entities and relationships in the vertical large model of the insurance field to identify key business entities and attribute descriptions in the insurance field; and structuring and arranging the key business entities and attribute descriptions based on the question-and-answer logic framework of the vertical large model to obtain the seed dataset for the insurance field.
3. The data augmentation and generalization method for a vertical large-scale model in the insurance field as described in claim 1, characterized in that, The process involves semantic space mapping of the question texts of standardized question-answer pairs in the seed dataset, based on the synonym and hierarchical relationship network of the vertical large model in the insurance domain, to obtain a basic enhanced question text set for the insurance domain. This includes: constructing a relational knowledge base for the insurance domain based on the synonym terminology list and concept hierarchy structure of the vertical large model in the insurance domain; performing terminology index matching on the question texts of standardized question-answer pairs in the seed dataset based on the terminology entries and association network in the relational knowledge base to obtain specific business terms for the insurance domain; performing semantic expansion retrieval on the specific business terms to obtain synonym expressions and hierarchical / superior concept expressions for the specific business terms; rewriting the synonym expressions and hierarchical / superior concept expressions in a contextualized manner based on the syntax and context of the specific business terms to obtain new question texts for the insurance domain; and collecting and organizing the new question texts to obtain the basic enhanced question text set for the insurance domain.
4. The data augmentation and generalization method for a vertical large-scale model in the insurance field as described in claim 1, characterized in that, The method of selecting the best from the basic augmented question text set based on the multi-dimensional quality assessment criteria of the vertical large model to obtain the augmented dataset for the insurance domain includes: performing compliance judgment on the basic augmented question text set based on the semantic consistency judgment rules of the multi-dimensional quality assessment criteria in the vertical large model to obtain the semantic consistency verification result for the insurance domain; performing expression review on the basic augmented question text set according to the language standardization judgment rules in the multi-dimensional quality assessment criteria to obtain the grammatical compliance verification result for the insurance domain; identifying new question texts that pass both the semantic consistency verification result and the grammatical compliance verification result as high-confidence texts for the insurance domain in the basic augmented question text set; pairing and recombining the high-confidence texts with their original answers, and removing new question texts that were not identified as high-confidence texts to form the augmented dataset for the insurance domain.
5. The data augmentation and generalization method for a vertical large-scale model in the insurance field as described in claim 1, characterized in that, The process of refining the question-answer pairs in the augmented dataset to obtain multi-turn dialogue training data in the insurance domain, based on the multi-turn dialogue state transition tree of the vertical large model, includes: formally modeling the inherent logical structure of multi-turn dialogues in the vertical large model based on standardized process knowledge in the insurance domain, to obtain the multi-turn dialogue state transition tree in the insurance domain; performing intent parsing on the question-answer pairs in the augmented dataset based on the business stage features in the multi-turn dialogue state transition tree, and determining the dialogue stage of the question-answer pairs according to the business domain of the question-answer pairs, to obtain the logical position information in the insurance domain; arranging the logical skeleton of the augmented dataset according to the process sequence rules and state transition logic in the multi-turn dialogue state transition tree and the logical position information, to obtain the multi-turn dialogue content framework in the insurance domain; and performing domain-specific language calibration on the question-answer content of the augmented dataset within the multi-turn dialogue content framework according to the business language habits in the insurance domain, to obtain the multi-turn dialogue training data in the insurance domain.
6. The data augmentation and generalization method for a vertical large-scale model in the insurance field as described in claim 5, characterized in that, The method involves formally modeling the internal logical structure of multi-turn dialogues in the vertical large model based on standardized process knowledge in the insurance field, resulting in a multi-turn dialogue state transition tree for the insurance field. This includes: decomposing and organizing the standardized process knowledge in the insurance field to obtain dialogue state nodes; defining conditions for the dialogue state nodes based on the inter-step dependencies and progression rules in the insurance business operation procedures to obtain transition paths; adding information load requirements to the internal logical structure of multi-turn dialogues in the vertical large model according to the weight information elements in the transition paths to establish information dependencies in the insurance field; and constructing a topology for the dialogue state nodes, the transition paths, and the information dependencies to obtain the multi-turn dialogue state transition tree for the insurance field.
7. The data augmentation and generalization method for a vertical large-scale model in the insurance field as described in claim 6, characterized in that, The formula for calculating the path weight in the transfer path is as follows: In the formula, For the nodes in the dialogue state node To the node Path weights, For the nodes in the dependency relationship To the node The tightness of business logic, For the dependent node in the above dependency relationship To the node Information dependency satisfaction For the slave node in the dialogue state node To the node The number of dialogue state nodes on the path. This parameter represents the maximum allowed number of dialogue turns in the vertical large model. The preset business logic density weighting coefficient, The preset information dependency satisfaction weight coefficients, The preset path simplicity penalty strength coefficient, It is a natural constant.
8. The data augmentation and generalization method for a vertical large-scale model in the insurance field as described in claim 1, characterized in that, The process of fusing single-turn question-answering data and multi-turn dialogue training data from the augmented dataset with attribute adaptation, and performing hierarchical labeling based on the business scenario classification system and dialogue complexity definition rules of the vertical big model to obtain the final training corpus for the insurance domain, includes: unifying the format representation of the single-turn question-answering data and multi-turn dialogue training data from the augmented dataset to obtain the original data collection for the insurance domain; classifying and mapping the business types of the data collection based on the business scenario classification system of the vertical big model to obtain business scenario labels for the insurance domain; hierarchically dividing the data collection based on the dialogue complexity definition rules of the vertical big model to obtain complexity level labels for the insurance domain; and archiving and organizing the data content of the business scenario labels and complexity level labels to obtain the final training corpus for the insurance domain.
9. A data augmentation and generalization method for a vertical large-scale model in the insurance field as described in claim 1, characterized in that, The method, based on the final training corpus, a pre-defined parameter fine-tuning paradigm, and a three-layer guided prompting engineering framework, collaboratively shapes the training process of the basic large language model in the insurance domain to obtain a generalized vertical large model in the insurance domain. This includes: using the final training corpus as the core training content for the insurance domain; based on the efficient parameter fine-tuning of the vertical large model and the specific requirements for model stability and adaptation efficiency in the insurance domain, explicitly designing and specifying the parameter update range and adjustment rules during the training process of the vertical large model to obtain the operating procedures for the insurance domain; layering and encapsulating the functional positioning, compliance requirements, and standardized output specifications of customer service assistants in the insurance domain to obtain a three-layer guided benchmark for the insurance domain; and performing paradigm-based collaborative updates on the vertical large model based on the three-layer guided benchmark, the operating procedures, and the core training content to obtain the generalized vertical large model for the insurance domain.
10. A data augmentation and generalization system for a large vertical model in the insurance field, characterized in that, The system, used to implement the data augmentation and generalization method for a vertical large-scale model in the insurance field as described in claim 1, comprises: a data structuring module for refining and reorganizing multi-source heterogeneous raw data in the insurance field to obtain a seed dataset for the insurance field; a semantic mapping enhancement module for performing semantic space mapping on the question texts of standardized question-answer pairs in the seed dataset based on the synonym and hyponymy relationship network of the vertical large-scale model in the insurance field to obtain a basic enhanced question text set for the insurance field; a quality selection module for performing quality selection on the basic enhanced question text set based on the multi-dimensional quality evaluation criteria of the vertical large-scale model to obtain an enhanced dataset for the insurance field; and a dialogue state polishing module for... Based on the multi-turn dialogue state transition tree of the vertical large model, the question-answer pairs in the augmented dataset are realistically polished to obtain multi-turn dialogue training data for the insurance domain. A fusion hierarchical labeling module is used to perform attribute-adaptive fusion of the single-turn question-answer data in the augmented dataset with the multi-turn dialogue training data, and to perform hierarchical labeling according to the business scenario classification system and dialogue complexity definition rules of the vertical large model to obtain the final training corpus for the insurance domain. A collaborative fine-tuning and shaping module is used to collaboratively shape the training process of the basic large language model in the insurance domain based on the final training corpus, a preset parameter efficient fine-tuning paradigm, and a three-layer guided prompting engineering framework to obtain a generalized vertical large model for the insurance domain.
Citation Information
Patent Citations
Method and device for determining training data of large model in insurance field, equipment and medium
CN119226519A
Large model customized training method and system for industry application
CN120492599A