Insurance product data processing method and device, equipment and storage medium

By using multi-agent collaborative processing of multi-source data to generate optimal insurance product data packages, the problems of single creative ideas, long cycles, and insufficient accuracy in traditional insurance product development are solved, thus achieving efficient and accurate product development.

CN121883176APending Publication Date: 2026-04-17PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-06
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In traditional insurance product development, creative ideas rely on human experience, lack data analysis, have low verification efficiency, and insufficient analytical accuracy, resulting in slow product innovation speed and low precision.

Method used

By accessing multi-source heterogeneous data, cleaning and standardizing it, calling the creative generation agent for semantic analysis, and combining it with the competitor comparison agent, actuarial evaluation agent, compliance review agent and market insight agent for data processing, the optimal insurance product data package is generated using a multi-objective optimization algorithm.

Benefits of technology

It has achieved full automation from data collection to product output, significantly improving product innovation efficiency, shortening the R&D cycle, and enhancing the market accuracy and compliance reliability of products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883176A_ABST
    Figure CN121883176A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing of financial and medical scenes, and discloses an insurance product data processing method and device, computer equipment and a storage medium, and the method comprises the steps: accessing multi-source heterogeneous data, and processing the multi-source heterogeneous data into standardized structured data resources and unstructured data resources; calling a plurality of agents to process the structured data resources and the unstructured data resources to generate a plurality of items of intermediate data related to the insurance product; and receiving the multiple items of intermediate data, performing comprehensive optimization on the multiple items of intermediate data by using a multi-objective optimization algorithm, and generating and outputting an optimal insurance product data packet. According to the invention, a complete intelligent product research and development closed loop is constructed, and full-process coverage from data acquisition to product output is realized through multi-agent division cooperation; the product innovation efficiency is remarkably improved, the research and development period is greatly shortened, and meanwhile the market precision and compliance reliability of products are effectively improved through data-driven intelligent analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology in the medical and health field, and in particular to a data processing method, apparatus, equipment and storage medium for insurance products. Background Technology

[0002] In the financial insurance and healthcare insurance sectors, the speed and accuracy of product innovation determine a company's market competitiveness. Traditional insurance product development follows a linear process that relies heavily on manual labor, a process that has revealed many inherent flaws in today's fast-paced, data-driven market environment.

[0003] First, the current creative generation mechanism is relatively rigid; mainly because product ideas rely heavily on human experience and lack systematic analysis of unstructured information such as medical data and regulatory rules, resulting in insufficient effective creative output and difficulty in meeting the market's demand for innovative financial and health insurance products.

[0004] Secondly, the current verification process is inefficient; mainly because actuarial calculations, compliance reviews and other steps are carried out in a sequential manual operation, which makes product verification time-consuming. Actuaries need to manually process structured data such as life tables, and compliance personnel need to compare massive amounts of regulatory provisions one by one, which seriously restricts the speed of product launch.

[0005] Finally, current analytical methods lack precision; primarily, competitor analysis relies on manual comparison using Excel, making it difficult to achieve semantic understanding at the clause level, and market analysis lacks the ability to mine sentiment from user comments, leading to a disconnect between products and real needs.

[0006] Therefore, there is an urgent need for a technical solution that can achieve full-process automated collaboration to solve the problems of single source of ideas, long R&D cycle and insufficient analysis accuracy. Summary of the Invention

[0007] The purpose of this invention is to provide a method, apparatus, device, and storage medium for processing insurance product data, aiming to solve the problems of limited sources of product ideas, lengthy R&D cycles, and insufficient analytical accuracy.

[0008] In a first aspect, embodiments of the present invention provide an insurance product data processing method, comprising: Access multi-source heterogeneous data, and clean and standardize the multi-source heterogeneous data to generate standardized structured data resources and unstructured data resources; The creative generation agent is invoked to perform semantic analysis and creative reasoning on the unstructured data resources to obtain preliminary product solution data; The competitor comparison agent is invoked, and semantic retrieval and comparison processing of the clause text in the unstructured data resources is performed in conjunction with the preliminary product plan data to obtain differentiated analysis data; The actuarial assessment agent is invoked to perform actuarial simulation and risk assessment on the structured data resources, generating actuarial assessment data. The compliance review intelligent agent is invoked to perform rule matching and risk identification processing on the clause text in the unstructured data resources to obtain compliance review data; The market insight agent is invoked to perform sentiment analysis and topic mining on the market commentary text in the unstructured data resources to obtain market trend profile data; By optimizing and integrating the data received by the intelligent agent, including preliminary product plan data, differentiation analysis data, actuarial evaluation data, compliance review data, and market trend profile data, and using a multi-objective optimization algorithm to perform comprehensive optimization, the optimal insurance product data package is generated and output.

[0009] Secondly, embodiments of the present invention provide an insurance product data processing apparatus, comprising: The data processing unit is used to access multi-source heterogeneous data, and to clean and standardize the multi-source heterogeneous data to generate standardized structured data resources and unstructured data resources. The first intelligent agent unit is used to call the creative generation intelligent agent to perform semantic analysis and creative reasoning processing on the unstructured data resources to obtain preliminary product solution data; The second intelligent agent unit is used to call the competitor comparison intelligent agent and combine the preliminary product plan data to perform semantic retrieval and comparison processing on the clause text in the unstructured data resources to obtain differentiated analysis data. The third intelligent agent unit is used to call the actuarial evaluation intelligent agent to perform actuarial simulation and risk assessment processing on the structured data resources and generate actuarial evaluation data; The fourth intelligent agent unit is used to call the compliance review intelligent agent to perform rule matching and risk identification processing on the clause text in the unstructured data resources to obtain compliance review data; The fifth intelligent agent unit is used to call the market insight intelligent agent to perform sentiment analysis and topic mining on the market comment text in the unstructured data resources to obtain market trend profile data; The optimization unit is used to receive the preliminary product plan data, differentiation analysis data, actuarial evaluation data, compliance review data, and market trend profile data from the intelligent agent through optimization and integration, and to perform comprehensive optimization using a multi-objective optimization algorithm to generate and output the optimal insurance product data package.

[0010] Thirdly, embodiments of the present invention provide a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the insurance product data processing method described in the first aspect above.

[0011] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the insurance product data processing method described in the first aspect.

[0012] In the aforementioned insurance product data processing method, apparatus, equipment, and storage medium, the solution involves accessing multi-source heterogeneous data, cleaning and standardizing the multi-source heterogeneous data to generate standardized structured and unstructured data resources; invoking a creative generation intelligence agent to perform semantic analysis and creative reasoning processing on the unstructured data resources to obtain preliminary product plan data; invoking a competitor comparison intelligence agent, and combining the preliminary product plan data to perform semantic retrieval and comparison processing on the clause text in the unstructured data resources to obtain differential analysis data; and invoking an actuarial evaluation intelligence agent to process the structured data resources... The system performs actuarial simulation and risk assessment on the source data to generate actuarial assessment data. A compliance review agent is invoked to perform rule matching and risk identification on the clause text in the unstructured data resources, resulting in compliance review data. A market insight agent is invoked to perform sentiment analysis and topic mining on the market commentary text in the unstructured data resources, resulting in market trend profile data. An optimization and integration agent receives the preliminary product plan data, differentiation analysis data, actuarial assessment data, compliance review data, and market trend profile data, and uses a multi-objective optimization algorithm to perform comprehensive optimization, generating and outputting the optimal insurance product data package. In this invention, a complete intelligent product development closed loop is constructed for insurance product data processing applications in financial and healthcare scenarios. Through multi-agent collaboration, it achieves full-process coverage from data collection to product output; significantly improving product innovation efficiency and greatly shortening the development cycle; simultaneously, data-driven intelligent analysis effectively improves the market accuracy and compliance reliability of the product. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1This is a schematic diagram of an application environment for the insurance product data processing method provided in an embodiment of the present invention; Figure 2 This is a flowchart illustrating an insurance product data processing method provided in an embodiment of the present invention. Figure 3 A schematic block diagram of an insurance product data processing device provided in an embodiment of the present invention; Figure 4 A schematic diagram of a computer device is provided for an embodiment of the present invention; Figure 5 Another structural schematic diagram of a computer device is provided for an embodiment of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] The insurance product data processing method provided in this embodiment of the invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The server can access multi-source heterogeneous data through the client, clean and standardize the data to generate standardized structured and unstructured data resources. It then invokes a creative generation agent to perform semantic analysis and creative reasoning on the unstructured data resources to obtain preliminary product plan data. Next, it invokes a competitor comparison agent to perform semantic retrieval and comparison of the clause text in the unstructured data resources, based on the preliminary product plan data, to obtain differential analysis data. Finally, it invokes an actuarial assessment agent to perform actuarial simulation and risk assessment on the structured data resources, generating actuarial assessment data. A compliance review agent performs rule matching and risk identification on the clause text in the unstructured data resources, obtaining compliance review data. A market insight agent performs sentiment analysis and topic mining on the market commentary text in the unstructured data resources, obtaining market trend profile data. Finally, by optimizing and integrating the preliminary product plan data, differential analysis data, actuarial assessment data, compliance review data, and market trend profile data, the server uses a multi-objective optimization algorithm to perform comprehensive optimization, generating and outputting the optimal insurance product data package. This invention constructs a complete intelligent product development closed loop for insurance product data processing applications in financial and healthcare scenarios. Through multi-agent collaboration, it achieves full-process coverage from data collection to product output, significantly improving product innovation efficiency, greatly shortening the development cycle, and effectively enhancing the market accuracy and compliance reliability of products through data-driven intelligent analysis. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a dedicated server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0017] Please see Figure 2 As shown, Figure 2 This is a flowchart illustrating the insurance product data processing method provided in an embodiment of the present invention.

[0018] like Figure 2 As shown, the method includes steps S201-S207.

[0019] S201. Access multi-source heterogeneous data, and clean and standardize the multi-source heterogeneous data to generate standardized structured data resources and unstructured data resources. The multi-source heterogeneous data in step S201 can include data from the financial and healthcare sectors, such as regulatory documents issued by financial regulatory agencies, publicly available product terms documents from peer insurance companies, and user reviews and survey reports obtained from social media and health community platforms. This data is cleaned and standardized using natural language processing tools. Structured data resources mainly refer to actuarial parameter tables (such as life tables and disease incidence tables in healthcare insurance) and market statistical indicators that can be directly stored and calculated in a database; unstructured data resources include original regulatory rules, insurance terms texts, user review texts, and other content requiring further semantic analysis.

[0020] S202. Call the creative generation intelligent agent to perform semantic analysis and creative reasoning processing on unstructured data resources to obtain preliminary product solution data; Step S202 uses a creative agent to perform deep semantic understanding of unstructured data, aiming to transform scattered market demands and rule information into concrete product ideas.

[0021] Step S202 specifically includes: Extract market demand signals and regulatory trend information from unstructured data resources; Based on the large language model, a pre-set Prompt template is used to perform chain reasoning on market demand signals and regulatory trend information to generate multiple structured product plans containing information on coverage, rate structure and target customer groups, which serve as preliminary product plan data.

[0022] In a specific embodiment of step S202, the creative generating agent can be a natural language processing module based on deep learning algorithms. The core components of this module include a semantic understander and a text generator.

[0023] Specifically, creative generative agents can extract key information from unstructured data using named entity recognition technology combined with the TF-IDF algorithm. Specifically, the entity recognition module in the spaCy toolkit can be used to automatically identify medical project requirements such as "early cancer screening" and "chronic disease management" from regulatory documents in unstructured data (such as unstructured data in the medical field). Simultaneously, it can capture trending keywords such as "telemedicine" and "home care" from market dynamics through word frequency statistics, forming a structured demand feature vector.

[0024] Specifically, the ChatGPT series of large language models can be used as the core generation engine. The specific working process is as follows: First, the extracted market demand signals and regulatory trend information are converted into specific Prompt instruction sequences, such as "Based on the current requirements of early cancer screening rules, design a health insurance product for the middle-aged and elderly population, which needs to include the following elements...". The model first generates a draft of the coverage scope through multi-round chain reasoning, including core coverage items such as inpatient medical care, outpatient surgery, and specific drugs; then, based on the principles of risk actuarial science, it generates the corresponding rate structure, distinguishing between basic premiums and additional service fees; finally, it clarifies the characteristics of the target customer group, including dimensions such as age distribution, occupation type, and health status, forming a complete product solution matrix.

[0025] Based on this, this embodiment utilizes the generation capabilities of a large language model and a chain-reasoning mechanism to systematically transform fragmented market demands and rule requirements into executable product solutions. It can generate multiple differentiated solutions within a single cycle, effectively solving the problems of reliance on personal experience and low efficiency in traditional creative generation. Simultaneously, the structured output format ensures the integrity and comparability of product solutions in key elements such as coverage scope and fee structure.

[0026] For example, in specific applications in the health insurance field, when the system identifies the regulatory requirements for coverage of "diabetic complications" and the emerging trend of "continuous blood glucose monitoring" technology, it can automatically generate a comprehensive insurance plan through the above workflow. This plan includes elements such as coverage for new monitoring equipment, medical expenses for complications, and personalized health management. The coverage explicitly includes the cost of the continuous blood glucose meter, the premium structure is set in tiers according to the monitoring frequency, and the target customer group is accurately positioned as type 2 diabetes patients, significantly improving the accuracy and efficiency of product innovation.

[0027] S203. Call the competitor comparison agent and combine it with the preliminary product solution data to perform semantic retrieval and comparison processing on the clause text in the unstructured data resources to obtain differentiated analysis data. Step S203 involves using a competitor comparison agent to perform semantic-level clause analysis, with the aim of establishing a differentiated product positioning from the early stages of the creative process.

[0028] Step S203 specifically includes: Key elements are extracted from product terms text in unstructured data resources to generate term vectors; Perform semantic similarity matching between the clause vectors and the product clause vectors pre-stored in the vector database; Based on the matching results, a competitive analysis report containing differential indicators is generated and used as differential analysis data.

[0029] In a specific embodiment of step S203, the competitor comparison agent is an analysis engine specifically designed for comparing insurance terms, and its core functions include text vectorization and similarity calculation.

[0030] First, a BERT-based pre-trained language model can be used as a tool for extracting key elements. During the extraction process, the product terms text from unstructured data resources is first input into the model. A 12-layer Transformer encoder extracts textual features, and then a mean pooling layer is used to convert the variable-length sequence into a fixed-length 768-dimensional vector representation. This process effectively captures the deep semantic features of key elements such as "coverage scope," "disclaimer clauses," and "compensation conditions" to generate term vectors.

[0031] Subsequently, one or more similarity matching algorithms can be used for analysis, including cosine similarity, Euclidean distance, and Manhattan distance. In practice, the similarity calculation module in the competitor comparison agent performs a multi-dimensional comparison between the clause vector to be compared and the pre-stored product clause vectors conforming to industry standards in the vector database. The matching results are mainly divided into three categories: complete match (similarity > 0.9), partial match (similarity 0.5-0.9), and significant difference (similarity < 0.5).

[0032] Finally, based on the matching results, a competitive analysis report is generated using a pre-defined quantitative analysis model. Specifically, the Pandas data analysis library in Python can be used to standardize the similarity scores and, combined with the importance weighting coefficients of the clauses, calculate the difference index for each clause. For example, for medical insurance clauses, the system will automatically generate specific quantitative indicators such as "This product's cancer hospitalization reimbursement ratio is 12% lower than the industry average" or "It has broader coverage in terms of special drug reimbursement," and integrate these indicators into a structured competitive analysis report.

[0033] Therefore, this embodiment upgrades the traditional coarse-grained analysis based on keyword matching to a semantic-level precise comparison. Through multi-dimensional similarity algorithms, it can discover subtle but crucial differences between clauses, providing a reliable basis for product differentiation positioning. Simultaneously, the quantitative report generation mechanism based on statistical analysis ensures the objectivity and operability of the analysis results.

[0034] For example, in practical applications within the medical insurance field, this method can accurately compare the special drug insurance terms of different companies. The system can identify subtle differences in certain terms regarding "targeted drug coverage" or "immunotherapy reimbursement ratio." Through semantic similarity calculations, it can discover that a certain product has a coverage rate difference of X% compared to the industry standard in the PD-1 inhibitor reimbursement terms, and automatically generate a quantitative analysis report containing specific difference data and improvement suggestions, providing clear direction for product optimization.

[0035] S204. Call the actuarial evaluation intelligent agent to perform actuarial simulation and risk assessment on the structured data resources, and generate actuarial evaluation data; Step S204 uses an actuarial evaluation agent to quickly simulate structured risk data, aiming to complete risk pricing and profitability verification during the product design phase.

[0036] Step S204 specifically includes: Extract actuarial parameter tables from structured data resources; Extract the rate assumptions and warranty information from the preliminary product plan data; Based on the actuarial parameter table, rate assumptions and coverage information, the premium and reserve are calculated using a pre-set actuarial model to obtain the calculation results; An assessment report containing profitability indicators and risk exposure forecasts is generated based on the calculation results and used as actuarial assessment data.

[0037] In a specific embodiment of step S204, the actuarial evaluation agent is a professional analysis system that integrates actuarial principles and computational technology.

[0038] Firstly, in the data extraction phase, a multi-tool collaborative approach can be adopted. For extracting actuarial parameter tables, Python's Pandas library can be used to directly read core parameters such as life tables and disease incidence tables from structured data (e.g., structured data in the medical field) via an SQL query interface. During this process, a connection to the actuarial database is first established, and then parameterized query statements are executed to convert the CSV-formatted actuarial basic data into a DataFrame structure for storage. For extracting rate assumptions and coverage information, a JSON parser can be used to structure the preliminary product plan data, automatically identifying and extracting key fields related to "rate assumptions and coverage information," such as "base rate," "age coefficient," and "list of coverage items," through key-value matching.

[0039] Secondly, in the actuarial calculation phase, an actuarial model based on a compound Poisson process can be used as the core calculation engine. The specific calculation process of this actuarial model (taking the healthcare field as an example) includes the following steps: First, a risk exposure matrix is ​​established based on the coverage information, and the expected loss frequency and magnitude for each coverage item are calculated; second, based on mortality and morbidity data in the actuarial parameter table, a preset number of random scenario tests are generated using Monte Carlo simulation; then, the aggregated loss distribution is calculated through convolution operations, and combined with the expense ratio and profit margin parameters in the rate assumptions, the final calculation results for net premium, supplementary premium, and total premium are obtained.

[0040] Finally, during the report generation phase, a standardized actuarial assessment report is automatically generated based on the calculation results. The calculation results are divided into multiple loss rate ranges, each corresponding to an assessment conclusion. For example, when the loss rate is in the 0%-10% range, the assessment conclusion is "good profitability and controllable risk exposure"; when the loss rate is in the 10%-20% range, the assessment conclusion is "basically stable profitability, but potential risks need to be monitored." The report specifically includes the following core components: expected loss rate analysis, profit sensitivity test results, reserve adequacy assessment, and risk capital requirement calculation modules, each accompanied by corresponding data visualizations and textual explanations.

[0041] Therefore, this embodiment fully automates the complex actuarial calculation process, providing accurate risk assessment and profit forecasts through rigorous random simulation and statistical calculations. Compared to traditional manual calculation methods, this embodiment shortens the actuarial evaluation time while avoiding biases caused by subjective judgment, ensuring the scientific nature and accuracy of product pricing.

[0042] S205. Call the compliance review intelligent agent to perform rule matching and risk identification processing on the clause text in the unstructured data resources to obtain compliance review data; Step S205 uses a compliance review agent to automate rule matching, aiming to identify and correct compliance risks in the terms in real time.

[0043] Step S205 specifically includes: Construct a compliance knowledge graph based on product regulatory rules; Extract product terms text from unstructured data resources; By matching product terms and conditions with nodes and relationships in a compliance knowledge graph, clause fragments or risky terms that do not comply with product regulatory rules can be identified. Generate compliance reports containing specific risk locations and proposed modifications based on clause snippets or risk terminology, and use these reports as compliance review data.

[0044] In a specific embodiment of step S205, the compliance review agent is an intelligent review system based on knowledge graphs and rule reasoning. Its core components include three main modules: a knowledge graph construction module, a text parsing module, and a rule matching engine. The knowledge graph construction module is responsible for transforming regulatory provisions into structured knowledge; the text parsing module is used to process insurance clause texts; and the rule matching engine executes specific compliance check logic.

[0045] First, in the knowledge graph construction phase, Neo4j graph database can be used as the storage tool, in conjunction with the Stanford CoreNLP toolkit for semantic parsing of regulatory texts. The construction process specifically includes: first, identifying rule documents related to insurance products, such as the "Insurance Law" and "Health Insurance Management Measures," from unstructured data resources (e.g., unstructured data resources in the medical field) using named entity recognition, and extracting key concepts such as "insurance period," "exclusions," and "waiting period" as graph nodes; then, identifying relationships between nodes through dependency parsing, such as "inclusion," "prohibition," and "restriction"; finally, importing the nodes and relationships into the Neo4j database to form a compliance knowledge graph containing multiple nodes and relationships.

[0046] Secondly, the specific process of node and relationship matching is as follows: First, the extracted product terms text is converted into a vector representation using the BERT model. Then, cosine similarity is calculated between this vector and the node vectors in the compliance knowledge graph to identify the most relevant regulatory rule nodes. Next, the Cypher query language is used to perform path queries in the compliance knowledge graph to check whether the product terms text violates restrictive relationships in regulatory rules. For example, when the expression "exemption of insurer's liability" is detected in the product terms text, the system will query the graph to see if this expression is associated with regulatory-prohibited exemption clauses, thereby identifying clause fragments or risky terms that do not comply with product regulatory rules.

[0047] Finally, during compliance report generation, based on the detection results of the rule engine, a detailed report can be automatically generated using templated natural language generation technology. The specific process includes: first, locating the specific position (including chapter, paragraph, and line number) of the content that does not comply with product regulatory rules in the original terms; then, extracting the corresponding regulatory basis from the compliance knowledge graph; and finally, generating specific modification suggestions based on a predefined modification rule base. For example, when it detects that the expression "XX disease" in health insurance terms violates regulatory requirements, the system will automatically suggest modifying it to the standardized expression "XX deficiency syndrome" and specify the relevant regulatory document clauses.

[0048] Therefore, this embodiment transforms traditional compliance reviews that rely on human experience into a systematic and standardized automated detection process. Through the semantic understanding capabilities of the compliance knowledge graph, the system can not only identify literal violations but also uncover potential compliance risks, significantly improving the accuracy and efficiency of the review. Simultaneously, the automated report generation mechanism ensures accurate identification and corrective action guidance for each issue.

[0049] For example, in medical insurance applications, this embodiment can automatically detect the compliance of health insurance terms. For instance, when reviewing a critical illness insurance policy, the system can use knowledge graph matching to discover discrepancies between the definition of coverage for "XX deficiency syndrome" in the policy and the latest regulatory requirements. It then automatically generates a detailed report containing specific location identifiers, regulatory references, and standardized modification suggestions, effectively avoiding compliance risks caused by inappropriate wording in the policy terms.

[0050] S206. Call the market insight intelligent agent to perform sentiment analysis and topic mining on the market commentary text in the unstructured data resources to obtain market trend profile data; Step S206 uses market insight agents to uncover user sentiment and demand hotspots, aiming to ensure that product design keeps pace with market dynamics.

[0051] Step S206 specifically includes: Extract market commentary text from unstructured data resources; Data cleaning and preprocessing of market commentary texts; The preprocessed market comment text is input into a sentiment analysis model to perform sentiment analysis and obtain the sentiment data of user comments in the market comment text; at the same time, the preprocessed market comment text is input into a topic model to extract the core discussion topics. Based on sentiment data and core discussion topics, user profile reports describing market preferences and demand hotspots are generated and used as market trend profile data.

[0052] In a specific embodiment of step S206, the market insight agent is an analysis system specifically designed to process user feedback data. Its core architecture includes four key components: a data acquisition module, a sentiment analysis engine, a topic modeling engine, and a report generation module. The data acquisition module is responsible for acquiring market commentary texts from multiple channels; the sentiment analysis engine is responsible for quantifying user sentiment tendencies; the topic modeling engine is responsible for identifying core discussion topics; and the report generation module is responsible for integrating the analysis results to form user profiles.

[0053] First, in the data extraction stage, the Scrapy framework can be used as a web crawler tool, in conjunction with the BeautifulSoup library for webpage parsing. The specific implementation process includes: first, configuring crawler rules to selectively extract insurance product-related content from unstructured data resources; then, removing noisy data through HTML tag parsing and text cleaning; finally, using the same spaCy library as in the aforementioned examples for text preprocessing, including word segmentation, stop word removal, and lemmatization, to generate standardized market commentary text.

[0054] Secondly, in the sentiment analysis stage, a BERT-based deep learning model is employed, specifically the RoBERTa-base version. The sentiment analysis process of this model is as follows: First, the preprocessed market commentary text is input into a 12-layer Transformer encoder, which captures the sentiment semantic features in the text through a self-attention mechanism. Then, the classification layer at the top of the model outputs the probability distributions of positive, negative, and neutral sentiments. Finally, the final sentiment score is calculated using the Softmax function. For example, for a comment like "This insurance claim processing is too slow," the model can accurately identify its negative sentiment and provide a corresponding negative probability score. Simultaneously, the topic model employs an LDA (Latent Dirichlet Distribution) model. The specific topic extraction process of this model includes: first, constructing a document-term matrix and iteratively calculating the distribution of latent topics using the Gibbs sampling method; then, extracting the keywords with the highest weights under each topic based on the topic-term distribution; and finally, calculating the weight distribution of each topic in the overall corpus. For example, from bank wealth management and insurance reviews, core topics such as "yield," "risk level," "purchase threshold," and "service attitude" may be extracted.

[0055] Finally, in the user profile report generation stage, a structured report can be automatically generated using multi-source data fusion technology. The specific process includes: first, correlation analysis of sentiment analysis results and topic modeling results to calculate the sentiment distribution for each topic; then, using clustering algorithms to identify user groups with similar preferences (such as investor groups); and finally, generating a comprehensive report based on the analysis results, including a demand hotspot map, user group satisfaction analysis, and product improvement suggestions. The report will automatically highlight key findings such as "Investors have a strong demand for highly liquid products" and "Satisfaction with the returns of medium- and long-term wealth management products is low."

[0056] Based on this, this embodiment can systematically mine investment preferences from market commentary texts (i.e., user feedback), accurately capture investor sentiment through the semantic understanding capabilities of deep learning, and achieve multi-level market demand insights by combining probabilistic topic model clustering analysis. Compared with traditional research methods, this implementation method can track market dynamics in real time, discover potential investor needs, and provide data-driven decision support for financial product innovation.

[0057] For example, in practical applications in the financial investment field, by analyzing tens of thousands of user comments in investment communities, the system can accurately identify investors' demand trends for product features such as "flexible redemption mechanisms" and "principal protection clauses," and generate detailed profile reports that include the specific intensity of demand, investor risk preference characteristics, and product optimization suggestions, providing clear directional guidance for the research and development and innovation of wealth management products.

[0058] S207. By optimizing and integrating the intelligent agent's reception of preliminary product plan data, differential analysis data, actuarial evaluation data, compliance review data, and market trend profile data, a multi-objective optimization algorithm is used to perform comprehensive optimization, generating and outputting the optimal insurance product data package. Step S207 aims to optimize and integrate the intelligent agent by using a multi-objective optimization algorithm to comprehensively balance various types of output data, and generate an insurance product data package that can be directly applied.

[0059] Step S207 specifically includes: The optimization objectives include market demand matching, profitability, and compliance. Using genetic algorithms or reinforcement learning algorithms, iterative calculations are performed on preliminary product plan data, differentiation analysis data, actuarial evaluation data, compliance review data, and market trend profile data to find the optimal combination of product parameters that balances multiple objectives. Based on the optimal combination of product parameters, a solution is generated that includes the final terms and conditions, actuarial calculation results, and product differentiation positioning, serving as an insurance product data package.

[0060] In a specific embodiment of step S207, the optimized and integrated intelligent agent is a core system that integrates multi-source data fusion and intelligent decision-making. Its architecture includes three key components: a data fusion engine, a multi-objective optimization algorithm library, and a decision support module. The data fusion engine is responsible for standardizing and extracting features from the output data of five intelligent agents: idea generation, competitor comparison, actuarial evaluation, compliance review, and market insight. The multi-objective optimization algorithm library provides implementations of various optimization algorithms. The decision support module is responsible for generating the final product solution and decision recommendations.

[0061] First, a multi-objective optimization method based on genetic algorithms can be employed. The specific optimization process of this algorithm is as follows: Initial product plan data is encoded as chromosomes, with each gene representing a product parameter (such as coverage scope, premium level, etc.). Then, differentiation analysis data is transformed into competitiveness indicators in the fitness function, actuarial evaluation data into profitability indicators, compliance review data into compliance indicators, and market trend profile data into market demand matching indicators. The algorithm iteratively optimizes product parameters through operations such as selection, crossover, and mutation: selection operations preserve non-dominated solutions based on Pareto dominance relationships; crossover operations exchange superior genes among different product plans; and mutation operations randomly adjust certain parameters with a certain probability to ensure population diversity. After iterative evolution, the algorithm converges to a set of Pareto optimal solutions, which achieve an optimal balance among multiple objectives.

[0062] Then, in the comprehensive balancing phase, a multi-criteria decision-making method based on entropy weighting can be adopted. Specific balancing criteria include: market demand matching degree weight (0.4), profitability weight (0.35), and compliance weight (0.25). The balancing process first normalizes the data for each indicator, and then calculates the weighted score for each option. For example, if an option's market demand matching degree score is 0.9 (out of 1.0), its profitability score is 0.8, and its compliance score is 1.0, then its comprehensive score is 0.9 × 0.4 + 0.8 × 0.35 + 1.0 × 0.25 = 0.89. The system automatically selects the option with the highest comprehensive score as the optimal output, and also provides sensitivity analysis for each indicator, showing the changes in the ranking of options under different weight configurations. Based on this, this embodiment can find a reasonable balance point among multiple mutually constraining objectives. Through the global search capability of intelligent algorithms, it avoids the local optima and subjective biases that may occur in traditional manual decision-making. At the same time, the weight-based quantitative balancing method ensures the transparency and interpretability of the decision-making process.

[0063] For example, in practical applications in the financial and insurance field, taking the development of investment-linked insurance products as an example, the system can simultaneously optimize the product's investment return characteristics (profitability), risk control requirements (compliance), and investor preference matching (market demand), automatically generating the optimal product solution that meets investors' return expectations and is competitive in the market while ensuring compliance, significantly improving the scientific nature and efficiency of product design.

[0064] As can be seen, the above solution constructs a complete intelligent product development closed loop for insurance product data processing applications in financial and healthcare scenarios. Through the division of labor and collaboration among multiple intelligent agents, it achieves full-process coverage from data collection to product output. It significantly improves product innovation efficiency, greatly shortens the R&D cycle, and effectively enhances the market accuracy and compliance reliability of products through data-driven intelligent analysis.

[0065] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0066] In one embodiment, an insurance product data processing apparatus is provided, which corresponds one-to-one with the insurance product data processing method described in the above embodiments. For example... Figure 3 As shown, the insurance product data processing device includes a data processing unit 301, a first intelligent agent unit 302, a second intelligent agent unit 303, a third intelligent agent unit 304, a fourth intelligent agent unit 305, a fifth intelligent agent unit 306, and an optimization unit 307. Detailed descriptions of each functional module are as follows: The data processing unit 301 is used to access multi-source heterogeneous data, and to clean and standardize the multi-source heterogeneous data to generate standardized structured data resources and unstructured data resources. The first intelligent agent unit 302 is used to call the creative generation intelligent agent to perform semantic analysis and creative reasoning processing on unstructured data resources to obtain preliminary product solution data; The second intelligent agent unit 303 is used to call the competitor comparison intelligent agent and combine the preliminary product plan data to perform semantic retrieval and comparison processing on the clause text in the unstructured data resources to obtain differentiated analysis data. The third intelligent agent unit 304 is used to call the actuarial evaluation intelligent agent to perform actuarial simulation and risk assessment processing on structured data resources and generate actuarial evaluation data; The fourth intelligent agent unit 305 is used to call the compliance review intelligent agent to perform rule matching and risk identification processing on the clause text in the unstructured data resources to obtain compliance review data; The fifth intelligent agent unit 306 is used to call the market insight intelligent agent to perform sentiment analysis and topic mining on market commentary texts in unstructured data resources to obtain market trend profile data; The optimization unit 307 is used to optimize and integrate the data received by the intelligent agent, including preliminary product plan data, differentiation analysis data, actuarial evaluation data, compliance review data, and market trend profile data, and to use a multi-objective optimization algorithm to perform comprehensive optimization, thereby generating and outputting the optimal insurance product data package.

[0067] In one embodiment, the first intelligent agent unit 302 is specifically used for: Extract market demand signals and regulatory trend information from unstructured data resources; Based on the large language model, a pre-set Prompt template is used to perform chain reasoning on market demand signals and regulatory trend information to generate multiple structured product plans containing information on coverage, rate structure and target customer groups, which serve as preliminary product plan data.

[0068] In one embodiment, the second intelligent agent unit 303 is specifically used for: Key elements are extracted from product terms text in unstructured data resources to generate term vectors; Perform semantic similarity matching and comparison between the clause vectors and the product clause vectors pre-stored in the vector database; Based on the matching results, a competitive analysis report containing differential indicators is generated and used as differential analysis data.

[0069] In one embodiment, the third intelligent agent unit 304 is specifically used for: Extract actuarial parameter tables from structured data resources; Extract the rate assumptions and warranty information from the preliminary product plan data; Based on the actuarial parameter table, rate assumptions and coverage information, the premium and reserve are calculated using a pre-set actuarial model to obtain the calculation results; An assessment report containing profitability indicators and risk exposure forecasts is generated based on the calculation results and used as actuarial assessment data.

[0070] In one embodiment, the fourth intelligent agent unit 305 is specifically used for: Construct a compliance knowledge graph based on product regulatory rules; Extract product terms text from unstructured data resources; By matching product terms and conditions with nodes and relationships in a compliance knowledge graph, clause fragments or risky terms that do not comply with product regulatory rules can be identified. Generate compliance reports containing specific risk locations and proposed modifications based on clause snippets or risk terminology, and use these reports as compliance review data.

[0071] In one embodiment, the fifth intelligent agent unit 306 is specifically used for: Extract market commentary text from unstructured data resources; Data cleaning and preprocessing of market commentary texts; The preprocessed market comment text is input into a sentiment analysis model to perform sentiment analysis and obtain the sentiment data of user comments in the market comment text; at the same time, the preprocessed market comment text is input into a topic model to extract the core discussion topics and their weight distribution. Based on sentiment data and core discussion topics and their weight distribution, user profile reports describing market preferences and demand hotspots are generated and used as market trend profile data.

[0072] In one embodiment, the optimization unit 307 is specifically used for: The optimization objectives include market demand matching, profitability, and compliance. Using genetic algorithms or reinforcement learning algorithms, iterative calculations are performed on preliminary product plan data, differentiation analysis data, actuarial evaluation data, compliance review data, and market trend profile data to find the optimal combination of product parameters that balances multiple objectives. Based on the optimal combination of product parameters, a solution is generated that includes the final terms and conditions, actuarial calculation results, and product differentiation positioning, serving as an insurance product data package.

[0073] Specific limitations regarding the insurance product data processing device can be found in the limitations regarding the insurance product data processing method described above, and will not be repeated here. Each module in the aforementioned insurance product data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0074] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a server-side insurance product data processing method.

[0075] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of an insurance product data processing method.

[0076] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Access multi-source heterogeneous data, and clean and standardize the multi-source heterogeneous data to generate standardized structured data resources and unstructured data resources; The creative generation agent is invoked to perform semantic analysis and creative reasoning on unstructured data resources to obtain preliminary product solution data; The competitor comparison agent is invoked, and semantic retrieval and comparison processing of the clause text in the unstructured data resources are performed in combination with the preliminary product plan data to obtain differentiated analysis data. The actuarial assessment agent is invoked to perform actuarial simulation and risk assessment on structured data resources, generating actuarial assessment data. The compliance review intelligent agent is invoked to perform rule matching and risk identification processing on the clause text in unstructured data resources to obtain compliance review data; The market insight agent is invoked to perform sentiment analysis and topic mining on market commentary texts in unstructured data resources to obtain market trend profile data; By optimizing and integrating the data received by the intelligent agent, including preliminary product plan data, differentiation analysis data, actuarial evaluation data, compliance review data, and market trend profile data, and using a multi-objective optimization algorithm to perform comprehensive optimization, the optimal insurance product data package is generated and output.

[0077] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Access multi-source heterogeneous data, and clean and standardize the multi-source heterogeneous data to generate standardized structured data resources and unstructured data resources; The creative generation agent is invoked to perform semantic analysis and creative reasoning on unstructured data resources to obtain preliminary product solution data; The competitor comparison agent is invoked, and semantic retrieval and comparison processing of the clause text in the unstructured data resources are performed in combination with the preliminary product plan data to obtain differentiated analysis data. The actuarial assessment agent is invoked to perform actuarial simulation and risk assessment on structured data resources, generating actuarial assessment data. The compliance review intelligent agent is invoked to perform rule matching and risk identification processing on the clause text in unstructured data resources to obtain compliance review data; The market insight agent is invoked to perform sentiment analysis and topic mining on market commentary texts in unstructured data resources to obtain market trend profile data; By optimizing and integrating the data received by the intelligent agent, including preliminary product plan data, differentiation analysis data, actuarial evaluation data, compliance review data, and market trend profile data, and using a multi-objective optimization algorithm to perform comprehensive optimization, the optimal insurance product data package is generated and output.

[0078] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0079] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0080] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0081] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. An insurance product data processing method, characterized by, include: Access multi-source heterogeneous data, and clean and standardize the multi-source heterogeneous data to generate standardized structured data resources and unstructured data resources; The creative generation agent is invoked to perform semantic analysis and creative reasoning on the unstructured data resources to obtain preliminary product solution data; The competitor comparison agent is invoked, and semantic retrieval and comparison processing of the clause text in the unstructured data resources is performed in conjunction with the preliminary product plan data to obtain differentiated analysis data; The actuarial assessment agent is invoked to perform actuarial simulation and risk assessment on the structured data resources, generating actuarial assessment data. The compliance review intelligent agent is invoked to perform rule matching and risk identification processing on the clause text in the unstructured data resources to obtain compliance review data; The market insight agent is invoked to perform sentiment analysis and topic mining on the market commentary text in the unstructured data resources to obtain market trend profile data; By optimizing and integrating the data received by the intelligent agent, including preliminary product plan data, differentiation analysis data, actuarial evaluation data, compliance review data, and market trend profile data, and using a multi-objective optimization algorithm to perform comprehensive optimization, the optimal insurance product data package is generated and output.

2. The insurance product data processing method of claim 1, wherein: The process of invoking the creative generation agent to perform semantic analysis and creative reasoning on the unstructured data resources to obtain preliminary product solution data includes: Extract market demand signals and regulatory trend information from the unstructured data resources; Based on the large language model, a pre-set Prompt template is used to perform chain reasoning on the market demand signals and regulatory trend information to generate multiple structured product plans containing information on coverage, rate structure and target customer groups, which serve as the preliminary product plan data.

3. The insurance product data processing method of claim 1, wherein, The process involves invoking a competitor comparison agent to perform semantic retrieval and comparison of the clause text in the unstructured data resources, resulting in differential analysis data, including: Key elements are extracted from the product terms text in the unstructured data resources to generate a terms vector; Perform semantic similarity matching between the clause vector and the product clause vectors pre-stored in the vector database; Based on the matching results, a competitive analysis report containing differential indicators is generated and used as the differential analysis data.

4. The insurance product data processing method according to claim 1, characterized in that, The process of invoking the actuarial evaluation intelligent agent to perform actuarial simulation and risk assessment on the structured data resources, generating actuarial evaluation data, includes: Extract the actuarial parameter table from the structured data resources; Extract the rate assumptions and warranty information from the preliminary product plan data; Based on the actuarial parameter table, rate assumptions, and coverage information, the premium and reserve are calculated using a preset actuarial model to obtain the calculation results. An assessment report containing profitability indicators and risk exposure forecasts is generated based on the calculation results and used as the actuarial assessment data.

5. The insurance product data processing method according to claim 1, characterized in that, The compliance review agent is invoked to perform rule matching and risk identification processing on the clause text in the unstructured data resources to obtain compliance review data, including: Construct a compliance knowledge graph based on product regulatory rules; Extract the product terms text from the unstructured data resources; The product terms and conditions text is matched with the compliance knowledge graph by nodes and relationships to identify clause fragments or risky terms that do not comply with product regulatory rules; A compliance report containing specific risk locations and proposed modifications is generated based on the aforementioned clause fragments or risk terms, and this report serves as the compliance review data.

6. The insurance product data processing method according to claim 1, characterized in that, The market insight agent is invoked to perform sentiment analysis and topic mining on the market commentary text in the unstructured data resources to obtain market trend profile data, including: Extract market commentary text from the unstructured data resources; The market commentary text was cleaned and preprocessed. The preprocessed market comment text is input into a sentiment analysis model to perform sentiment analysis and obtain the sentiment data of user comments in the market comment text; at the same time, the preprocessed market comment text is input into a topic model to extract the core discussion topics. Based on the sentiment data and the core discussion topics, a user profile report describing market preferences and demand hotspots is generated and used as the market trend profile data.

7. The insurance product data processing method according to claim 1, characterized in that, The process involves optimizing and integrating the preliminary product plan data, differentiation analysis data, actuarial evaluation data, compliance review data, and market trend profile data received by the intelligent agent. A multi-objective optimization algorithm is then used to perform comprehensive optimization, generating and outputting the optimal insurance product data package, including: The optimization objectives include market demand matching, profitability, and compliance. Using genetic algorithms or reinforcement learning algorithms, iterative calculations are performed on the preliminary product plan data, differentiation analysis data, actuarial evaluation data, compliance review data, and market trend profile data to find the optimal combination of product parameters that balances multiple objectives. Based on the optimal combination of product parameters, a solution is generated that includes the final terms text, actuarial calculation results, and product differentiation positioning, serving as the insurance product data package.

8. An insurance product data processing device, characterized in that, include: The data processing unit is used to access multi-source heterogeneous data, and to clean and standardize the multi-source heterogeneous data to generate standardized structured data resources and unstructured data resources. The first intelligent agent unit is used to call the creative generation intelligent agent to perform semantic analysis and creative reasoning processing on the unstructured data resources to obtain preliminary product solution data; The second intelligent agent unit is used to call the competitor comparison intelligent agent and combine the preliminary product plan data to perform semantic retrieval and comparison processing on the clause text in the unstructured data resources to obtain differentiated analysis data. The third intelligent agent unit is used to call the actuarial evaluation intelligent agent to perform actuarial simulation and risk assessment processing on the structured data resources and generate actuarial evaluation data; The fourth intelligent agent unit is used to call the compliance review intelligent agent to perform rule matching and risk identification processing on the clause text in the unstructured data resources to obtain compliance review data; The fifth intelligent agent unit is used to call the market insight intelligent agent to perform sentiment analysis and topic mining on the market comment text in the unstructured data resources to obtain market trend profile data; The optimization unit is used to receive the preliminary product plan data, differentiation analysis data, actuarial evaluation data, compliance review data, and market trend profile data from the intelligent agent through optimization and integration, and to perform comprehensive optimization using a multi-objective optimization algorithm to generate and output the optimal insurance product data package.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the insurance product data processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the insurance product data processing method as described in any one of claims 1 to 7.