Insurance scheme recommendation method and device based on multi-source data fusion, equipment and medium

CN122656776APending Publication Date: 2026-08-28CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610801790.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0004]本发明的主要目的在于提供多源数据融合的保险方案推荐方法、装置、设备与介质,旨在解决现有技术中保险方案推荐同质化明显,难以形成精准个性化的保险方案推荐的技术问题,提高保险方案推荐的准确性

Benefits of technology

[0009]Beneficial Effects: This invention discloses a method, apparatus, device, and medium for recommending insurance plans based on multi-source data fusion. Compared to existing technologies, this invention, in response to an insurance plan recommendation request, collects multi-source heterogeneous data from target customers. This multi-source heterogeneous data includes structured data, unstructured text data, time-series data, and graph data. Corresponding feature extraction processing is performed on the multi-source heterogeneous data to obtain structured feature vectors, text feature vectors, time-series feature vectors, and graph feature vectors. After multi-dimensional fusion processing of these feature vectors, a unified customer risk profile vector is output. This customer risk profile vector is then compared with a preset product knowledge base for product matching evaluation, and a corresponding insurance recommendation plan is generated based on the matching evaluation results. This invention can be applied to business scenarios such as fintech and healthcare. By collecting multi-source heterogeneous data and performing multi-source fusion processing, it accurately and comprehensively captures customer risk profiles, thereby achieving product matching and recommendation based on accurate customer risk profiles, improving the personalization and accuracy of the recommendation plan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122656776A_ABST
    Figure CN122656776A_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence and discloses an insurance scheme recommendation method, device, equipment and medium based on multi-source data fusion, which comprises the following steps: collecting multi-source heterogeneous data of a target customer, including structured data, unstructured text data, time series data and graph data; performing corresponding feature extraction processing on the multi-source heterogeneous data to obtain a structured feature vector, a text feature vector, a time series feature vector and a graph feature vector; performing multi-dimensional fusion processing on the structured feature vector, the text feature vector, the time series feature vector and the graph feature vector to output a unified customer risk portrait vector; and performing product matching evaluation on the customer risk portrait vector and a preset product knowledge base to generate a corresponding insurance recommendation scheme. The application can be applied to business scenes such as financial technology and medical health, can accurately and comprehensively capture a customer risk portrait through multi-source heterogeneous data, can perform product matching and recommendation, and can improve the personalization and accuracy of the recommendation scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology and can be applied to business areas such as fintech and healthcare, particularly to a method, apparatus, equipment, and medium for recommending insurance solutions based on multi-source data fusion. Background Technology

[0002] The insurance industry is a crucial component of the social risk management and health protection system, and personalized insurance plan recommendations are a core element in insurance business development, customer service improvement, and risk compliance management. Especially with the rapid development of the healthcare sector, insurance products deeply integrated with medical scenarios, such as health insurance, medical insurance, and critical illness insurance, are becoming increasingly popular. Information such as medical examination reports, medical records, chronic disease management, follow-up data, medical public opinion, industry treatment standards, and medical insurance policies have become key bases for insurance recommendations and underwriting risk control. At the same time, traditional life insurance, property insurance, and corporate insurance businesses are also experiencing continuous growth, leading to increasingly diverse data sources and more complex data types, including customer intentions, industry reports, market dynamics, and underwriting policies.

[0003] Currently, insurance plan recommendations still largely rely on traditional big data cleaning and simple rule engines. While some institutions have introduced recommendation systems, they mostly employ traditional algorithms such as collaborative filtering and content-based recommendation, which can only perform shallow matching on structured insurance information. The recommended results are highly homogenized and cannot form accurate and personalized underwriting plans. Therefore, how to achieve multi-source data fusion analysis and personalized recommendations has become an urgent problem to be solved in the process of digital and intelligent upgrading of health insurance and all types of insurance businesses. Summary of the Invention

[0004] The main objective of this invention is to provide a method, apparatus, device, and medium for recommending insurance plans based on multi-source data fusion, aiming to solve the technical problem that existing insurance plan recommendations are highly homogeneous and difficult to form accurate and personalized insurance plan recommendations, thereby improving the accuracy of insurance plan recommendations.

[0005] The technical solution of the present invention is as follows: The first aspect of this invention provides a method for recommending insurance schemes based on multi-source data fusion, comprising: In response to insurance plan recommendation requests, multi-source heterogeneous data of target customers are collected, including structured data, unstructured text data, time-series data, and graph data. The multi-source heterogeneous data is subjected to corresponding feature extraction processing to obtain structured feature vectors, text feature vectors, time-series feature vectors, and graph feature vectors, respectively. After performing multi-dimensional fusion processing on the structured feature vector, text feature vector, time-series feature vector, and graph feature vector, a unified customer risk profile vector is output. The customer risk profile vector is matched with a preset product knowledge base for product matching evaluation, and a corresponding insurance recommendation plan is generated based on the matching evaluation results.

[0006] A second aspect of the present invention provides an insurance scheme recommendation device based on multi-source data fusion, comprising: The data acquisition module is used to collect multi-source heterogeneous data of target customers in response to insurance plan recommendation requests. The multi-source heterogeneous data includes structured data, unstructured text data, time series data, and graph data. The data processing module is used to perform corresponding feature extraction processing on the multi-source heterogeneous data to obtain structured feature vectors, text feature vectors, time-series feature vectors, and graph feature vectors, respectively. The multi-dimensional fusion module is used to perform multi-dimensional fusion processing on the structured feature vector, text feature vector, time-series feature vector and graph feature vector, and output a unified customer risk profile vector. The solution recommendation module is used to perform product matching evaluation between the customer risk profile vector and the preset product knowledge base, and generate corresponding insurance recommendation solutions based on the matching evaluation results.

[0007] A third aspect of the present invention provides a computer device including at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform the above-described method for recommending an insurance scheme based on multi-source data fusion.

[0008] A fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the above-described method for recommending a multi-source data fusion insurance scheme.

[0009] Beneficial Effects: This invention discloses a method, apparatus, device, and medium for recommending insurance plans based on multi-source data fusion. Compared to existing technologies, this invention, in response to an insurance plan recommendation request, collects multi-source heterogeneous data from target customers. This multi-source heterogeneous data includes structured data, unstructured text data, time-series data, and graph data. Corresponding feature extraction processing is performed on the multi-source heterogeneous data to obtain structured feature vectors, text feature vectors, time-series feature vectors, and graph feature vectors. After multi-dimensional fusion processing of these feature vectors, a unified customer risk profile vector is output. This customer risk profile vector is then compared with a preset product knowledge base for product matching evaluation, and a corresponding insurance recommendation plan is generated based on the matching evaluation results. This invention can be applied to business scenarios such as fintech and healthcare. By collecting multi-source heterogeneous data and performing multi-source fusion processing, it accurately and comprehensively captures customer risk profiles, thereby achieving product matching and recommendation based on accurate customer risk profiles, improving the personalization and accuracy of the recommendation plan. Attached Figure Description

[0010] To more clearly illustrate the solutions in this invention, the accompanying drawings used in the description of the embodiments of this invention will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0011] Figure 1 A schematic diagram of an application environment for the insurance scheme recommendation method based on multi-source data fusion provided in this embodiment of the invention; Figure 2 A flowchart illustrating a method for recommending insurance schemes based on multi-source data fusion provided in an embodiment of the present invention; Figure 3 A flowchart of step S202 in the insurance scheme recommendation method based on multi-source data fusion provided in the embodiments of the present invention; Figure 4 This is another flowchart of step S202 in the insurance scheme recommendation method for multi-source data fusion provided in the embodiments of the present invention; Figure 5 This is another flowchart of step S202 in the insurance scheme recommendation method for multi-source data fusion provided in the embodiments of the present invention; Figure 6 This is another flowchart of step S202 in the insurance scheme recommendation method for multi-source data fusion provided in the embodiments of the present invention; Figure 7 This is a flowchart of step S203 in the insurance scheme recommendation method based on multi-source data fusion provided in an embodiment of the present invention; Figure 8 A flowchart of step S204 in the insurance scheme recommendation method for multi-source data fusion provided in the embodiments of the present invention; Figure 9 A schematic diagram of the functional modules of the insurance scheme recommendation device for multi-source data fusion provided in an embodiment of the present invention; Figure 10 A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0012] To make the objectives, technical solutions, and effects of this invention clearer and more explicit, the invention is further described in detail below. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. The embodiments of the invention are described below in conjunction with the accompanying drawings.

[0013] The multi-source data fusion insurance scheme recommendation method provided in this invention can be applied to, for example... Figure 1 In the application environment, it includes a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0014] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as financial clients, healthcare clients, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software, etc. (for example only).

[0015] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0016] Server 105 can be a server providing various services, such as a backend server supporting the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend server can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices. Server 105 can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the shortcomings of traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"), such as high management difficulty and weak business scalability. Server 105 can also be a server for a distributed system or a server combined with blockchain.

[0017] It should be noted that the multi-source data fusion insurance scheme recommendation method provided in this application embodiment can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the multi-source data fusion insurance scheme recommendation device provided in this embodiment can also be located in the first terminal device 101, the second terminal device 102, or the third terminal device 103. Alternatively, the multi-source data fusion insurance scheme recommendation method provided in this embodiment can generally be executed by the server 105. Correspondingly, the multi-source data fusion insurance scheme recommendation device provided in this embodiment can generally be located in the server 105.

[0018] It should be understood that the number of terminal devices, networks, and servers listed above is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be used.

[0019] like Figure 2 As shown, the insurance scheme recommendation method based on multi-source data fusion provided in this embodiment of the invention specifically includes the following steps: S201. In response to the insurance plan recommendation request, collect multi-source heterogeneous data of the target customer, including structured data, unstructured text data, time series data, and graph data.

[0020] In this embodiment, the insurance plan recommendation request can be initiated by the sales terminal or by the customer through the customer terminal. The request includes the target customer's basic identification information (such as the ID number of an individual customer or the unified social credit code of a corporate customer) to accurately locate the target customer and collect corresponding data. Based on the basic identification information carried in the request, corresponding multi-source heterogeneous data is collected from different data sources as the core foundation for personalized insurance recommendations. This includes structured data, unstructured text data, time-series data, and graph data. These four types of data represent the target customer's risk characteristics, needs and preferences, regulatory policies, and market environment from different dimensions, ensuring the comprehensiveness and accuracy of the recommended plan.

[0021] Specifically, structured data refers to tabular data with a fixed format that can be directly quantified. It mainly includes basic information of target customers (age, occupation, income, etc. of individual customers; business registration information, registered capital, industry classification, number of employees, etc. of corporate customers), past insurance / claims records, insurance product parameter data (insured amount, premium, deductible, scope of coverage), etc. After collection, it is uniformly stored according to preset field formats for easy subsequent processing.

[0022] Unstructured text data refers to text data that has no fixed format and cannot be directly quantified, including policy documents issued by regulatory authorities (such as environmental protection and work safety policies, and the China Banking and Insurance Regulatory Commission document

[2024] X), industry research reports (such as risk analysis reports for various industries), and news and public opinion texts (such as negative news and market dynamics of the target customer's industry). It can be collected through methods such as web search, database retrieval, and manual input.

[0023] Time series data refers to sequence data that changes over time, including monthly / quarterly sales data of various insurance products, time series of claims amount and frequency, time series data of risk index of the target customer's industry, and time series data of market premium fluctuations. The collection period can be set to monthly or quarterly according to actual needs, and the data is sorted by timestamp after collection to ensure the integrity of time series information.

[0024] Graph data refers to network-based data that represents various entities and their relationships, including target customer relationship network graphs (such as supplier-customer dependencies for corporate customers, and insurance relationships for relatives of individual customers), supplier-customer dependency graphs (upstream and downstream relationships in the supply chain), and related enterprise graphs.

[0025] For example, when a manufacturing company initiates an insurance plan recommendation request, the multi-source heterogeneous data collected includes: structured data (company registration information, insurance claims records for the past 3 years, product parameter tables); unstructured text data (new safety production policy documents, manufacturing industry risk reports, and related news and public opinion about the company); time-series data (time-series of manufacturing claims rates for the past 24 months, and time-series data of the company's premium payments); graph data (dependency graphs between the company and its upstream and downstream suppliers, and customer relationship network diagrams), etc. Through multi-dimensional data collection, the risk characteristics, regulatory policies, and market environment of the corporate client are comprehensively covered.

[0026] This embodiment collects heterogeneous data from multiple sources to cover structured and unstructured text, time series, and other data. Figure 4 This data comprehensively captures customer risk characteristics, demand preferences, regulatory policies, and market environment, solving the problem of one-sided information from a single data source and laying a data foundation for subsequent accurate recommendations.

[0027] S202. Perform corresponding feature extraction processing on the multi-source heterogeneous data to obtain structured feature vectors, text feature vectors, time-series feature vectors, and graph feature vectors, respectively.

[0028] In this embodiment, since the modal and format differences of multi-source heterogeneous data are large, they cannot be directly fused. Therefore, it is necessary to adopt the corresponding feature extraction method for the characteristics of each type of data to transform the original data into computable and fused feature vectors. Each type of feature vector corresponds to the core features of a type of data, laying the foundation for subsequent multi-dimensional fusion.

[0029] The feature extraction process for structured data focuses on weighted processing, extracting core features related to customer risk, enterprise profile, and product matching through feature engineering, and transforming them into standardized structured feature vectors. Feature extraction for unstructured text data emphasizes semantic extraction and quantification, transforming textual information into textual feature vectors containing sentiment and policy impact information through key information extraction, sentiment analysis, and policy impact analysis. Feature extraction for time-series data focuses on trend and fluctuation analysis, extracting time-series feature vectors containing risk fluctuations and market sentiment information through time-series feature extraction and predictive analysis. Feature extraction for graph data focuses on structure and correlation analysis, extracting corresponding graph feature vectors by constructing corresponding graph structures and performing risk propagation analysis and high-risk subgraph identification.

[0030] For example, for the multi-source heterogeneous data of the aforementioned manufacturing enterprise customers, after feature extraction processing, we obtain structured feature vectors (including features such as enterprise registered capital, claims ratio in the past 3 years, and unit premium payout ratio), text feature vectors (including features such as policy sensitivity, public opinion sentiment score, and policy impact index), time series feature vectors (including features such as risk volatility index and market heat index), and graph feature vectors (including features such as node centrality, risk transmission index, and high-risk subgraph labels). These four types of feature vectors are all high-dimensional numerical vectors with fixed dimensions, which can be directly used for subsequent fusion processing.

[0031] This embodiment performs targeted data processing on raw data of different modalities and formats, transforming various multi-source heterogeneous data into computable and fusionable feature vectors, eliminating differences in data modality and format, extracting core information of various data types, and providing a reliable data foundation for multi-dimensional feature fusion.

[0032] S203. After performing multi-dimensional fusion processing on the structured feature vector, text feature vector, time-series feature vector, and graph feature vector, a unified customer risk profile vector is output.

[0033] In this embodiment, after obtaining the feature vectors corresponding to various multi-source heterogeneous data, the feature vectors of different modalities and dimensions are further subjected to multi-dimensional fusion processing, mapping them to the same semantic space to obtain a unified customer risk profile vector. Specifically, during the multi-dimensional fusion processing, the structured feature vector, text feature vector, temporal feature vector, and graph feature vector are first input into dedicated domain encoders for deep encoding and dimensional unification to solve the problems of inconsistent feature dimensions and semantic misalignment among different modalities. Then, based on the outputs of each domain encoder, dynamic weighted fusion using an attention mechanism is performed to integrate the core information of various data, thereby generating a unified customer risk profile vector that can comprehensively and accurately represent the risk status, demand preferences, and policy impact of target customers. This fusion method can adaptively adjust the fusion weights of various data. For example, when policy documents are updated, the weight of text data will automatically increase to ensure the real-time performance and accuracy of the customer risk profile.

[0034] For example, the structured feature vectors, text feature vectors, time-series feature vectors, and graph feature vectors of the aforementioned manufacturing enterprise customers are input into the MLP encoder, BERT+Pooling encoder, Temporal Transformer encoder, and GAT encoder, respectively, to obtain four types of encoded vectors. Then, the dynamic scores of each type of data are calculated through an attention scoring network to obtain attention weights. Among them, text data (policy-related) has the highest weight because the policy documents have recently been updated. Finally, a unified customer risk profile vector is obtained by weighted summation through attention. This vector integrates multiple dimensions of information such as the enterprise's risk level, policy sensitivity, and market environment.

[0035] This embodiment maps features from different dimensions to the same semantic space through multi-dimensional fusion, integrates core information from multiple sources, generates a comprehensive and accurate customer risk profile, avoids the limitations of a single feature dimension, and improves the accuracy of risk assessment.

[0036] S204. Perform product matching evaluation between the customer risk profile vector and the preset product knowledge base, and generate corresponding insurance recommendation schemes based on the matching evaluation results.

[0037] In this embodiment, the preset product knowledge base contains detailed information on various insurance products, including product attributes (coverage amount, premium, scope of coverage, underwriting conditions), applicable population / industry, and product risk suitability standards, which can be updated in real time according to market changes and policy adjustments. When generating an insurance recommendation plan, the customer's risk profile vector is compared with the preset product knowledge base for product matching evaluation. The compatibility between the customer and various insurance products is quantified from multiple evaluation dimensions. Finally, the product or combination of products with the highest compatibility is selected as the recommended product. The recommendation reasons are generated in natural language form based on the customer's risk profile, and a complete insurance recommendation plan is generated and pushed to the sales terminal or customer terminal.

[0038] For example, for the aforementioned manufacturing enterprise clients, a pre-set product knowledge base includes various products such as property insurance, liability insurance, and work safety insurance. Based on the client's risk profile vector (high risk transmission, high policy sensitivity), the matching degree of each product is calculated. Among them, work safety insurance scores the highest in the dimensions of coverage matching degree and policy compliance, ranking first in matching score. The generated recommendation reason is "We recommend XX work safety insurance because work safety policies in your manufacturing industry have recently tightened, and the supply chain dependency diagram shows high risk transmission potential. The coverage of this product is close to the industry's risk points and complies with current regulatory requirements." Finally, the recommended products and recommendation reasons are combined into an insurance recommendation plan and pushed to the sales terminal for sales personnel to introduce to customers.

[0039] This embodiment uses a precise customer risk profile and product knowledge base to conduct multi-dimensional product matching assessment to achieve accurate product matching, select the best recommended products or combinations and generate targeted recommendation reasons, which can better adapt to the actual needs of customers, changes in regulatory policies and market fluctuations, and improve the personalization and accuracy of the recommendation solution.

[0040] In the above embodiments, this invention discloses an insurance scheme recommendation method based on multi-source data fusion. Responding to an insurance scheme recommendation request, it collects multi-source heterogeneous data of the target customer, including structured data, unstructured text data, time-series data, and graph data. The method performs corresponding feature extraction processing on the multi-source heterogeneous data to obtain structured feature vectors, text feature vectors, time-series feature vectors, and graph feature vectors. After multi-dimensional fusion processing of the structured feature vectors, text feature vectors, time-series feature vectors, and graph feature vectors, a unified customer risk profile vector is output. The customer risk profile vector is then compared with a preset product knowledge base for product matching evaluation, and a corresponding insurance recommendation scheme is generated based on the matching evaluation results. This invention can be applied to business scenarios such as fintech and healthcare. By collecting multi-source heterogeneous data and performing multi-source fusion processing, it accurately and comprehensively captures customer risk profiles, thereby achieving product matching and recommendation based on accurate customer risk profiles, improving the personalization and accuracy of the recommendation schemes.

[0041] In one embodiment, such as Figure 3 As shown, when the multi-source heterogeneous data is structured data, step S202 includes: S301. Perform data cleaning and outlier removal on the structured data to obtain compliant and valid data; S302. Perform feature engineering on the compliant and valid data to extract risk factors, enterprise profile features, and product matching features from the compliant and valid data to obtain a set of business features. S303. The structured feature vector is generated by standardizing and mapping the set of business features.

[0042] In this embodiment, when extracting features from structured data, data cleaning and outlier removal are performed first to ensure data compliance and validity. Specifically, data cleaning can include field standardization and missing value imputation. Field standardization unifies the naming and data types of structured data from different sources according to preset fields. For example, "Company Registration Date" is standardized to "YYYY-MM-DD", and "Premium" is standardized to "Yuan", etc. Missing value imputation can adopt different strategies according to different field types. For example, numerical fields (such as income and premium) can be filled with the mean or median, categorical fields (such as occupation and industry) can be filled with the mode, and business-sensitive fields (such as past claims records) can be filled with model prediction, such as based on claims data of customers in the same industry or age group, etc., thereby obtaining cleaned structured data.

[0043] Outlier detection and removal can employ the Z-score method to detect anomalies in continuous numerical fields (such as claim amounts and premiums) in the cleaned structured data. When a value deviates from the mean by more than three times the standard deviation, it is identified as an outlier. Extreme values ​​(such as premiums significantly higher than the average level of similar customers) can be removed using the IQR quartile method. Furthermore, business rule verification can be combined with this method. For example, logically incorrect data, such as insurance application dates earlier than the customer's birth date, can be directly marked as outliers and removed, thereby obtaining compliant and valid data.

[0044] Subsequently, feature engineering was performed on the compliant and valid data to extract core features relevant to insurance recommendations, including risk factors, corporate profile features, and product matching features. Risk factors address the customer's risk profile, such as the number of claims in the past 1 / 3 / 5 years, claim rate, payout ratio, and interval between claims. Corporate profile features include industry classification, registered capital, years of establishment, employee size, registered capital growth rate, and industry concentration. Product matching features include coverage rate and the ratio of deductible to corporate assets. Based on risk factors, corporate profile features, and product matching features, a set of business features is constructed that characterizes customer risk, corporate status, and product suitability, thereby discarding irrelevant features and improving the relevance and effectiveness of the features.

[0045] Because the various features in the business feature set have inconsistent dimensions and large differences in numerical range, they are standardized and mapped to the same numerical range. Specifically, the Min-Max normalization method can be used to map all features in the business feature set to the [0,1] interval, eliminating the impact of differences in dimension and numerical range. All standardized features are then concatenated in a preset order to generate a fixed-dimensional structured feature vector, which accurately represents the customer's risk status, corporate profile, and product suitability.

[0046] In this embodiment, data cleaning, feature engineering and standardized mapping processing are performed to ensure the compliance and effectiveness of structured data. Core features related to customer risk, enterprise portrait and product matching are extracted, and then structured feature vectors are generated, which improves the accuracy and standardization of structured data feature extraction and provides high-quality data support for subsequent fusion.

[0047] In one embodiment, as Figure 4 shows, when the multi-source heterogeneous data is unstructured text data, step S202 comprises: S401, performing text preprocessing on the unstructured text data, and extracting key information from the preprocessed text data to obtain policy keywords and risk events; S402, performing sentiment analysis according to the policy keywords and risk events to obtain corresponding sentiment polarity; S403, establishing a mapping relationship between the policy keywords and the risk events, performing weighted quantization, and generating a corresponding policy impact index; S404, fusing the sentiment polarity and the policy impact index into a unified semantic feature, and then generating the text feature vector.

[0048] In this embodiment, since unstructured text data has disordered formats and contains a large amount of redundant information such as irrelevant statements and special symbols, text preprocessing needs to be performed first to purify the text data. The specific text preprocessing comprises word segmentation, stop word removal and lemmatization. That is, the pure text is first segmented into independent words through the jieba word segmentation tool, for example, the sentence "strengthen work safety supervision and reduce claim settlement risks of manufacturing enterprises" is segmented into "strengthen, work safety, supervision, reduce, manufacturing industry, enterprise, claim settlement, risk"; then, words without practical significance such as "of,地,得,ah" and words irrelevant to the insurance field are removed through stop word removal, so as to retain core words. Finally, words after word segmentation and stop word removal are unified to their original forms through lemmatization, for example, "claim settlements" is unified into "claim settlement", so as to ensure the consistency of vocabulary.

[0049] Key information is extracted from preprocessed text data to obtain policy keywords and risk events. Specifically, the key information extraction uses a BERT+CRF model, which includes a BERT pre-training layer and a CRF decoding layer. The BERT pre-training layer uses a finely tuned model from the insurance field to accurately identify professional terms and entities in the insurance industry. The CRF decoding layer optimizes entity recognition results and improves the accuracy of key information extraction. The extracted policy keywords mainly include core terms from regulatory policies (such as "safety production supervision," "environmental compliance," and "tightening underwriting"), while risk events mainly include risk-related events in the customer's industry (such as "frequent photovoltaic module fires" and "increased claims for equipment failures in manufacturing"). This process extracts core policy keywords and risk events related to insurance recommendations, laying the foundation for subsequent sentiment analysis and policy impact modeling.

[0050] After obtaining policy keywords and risk events, sentiment analysis is further performed to determine the sentiment tendency (positive, neutral, negative) of unstructured text related to policy keywords and risk events. This sentiment polarity can influence the customer's risk level assessment; for example, risk events with negative sentiment will increase the customer's risk level, while policies with positive sentiment will reduce the customer's compliance risk. Specifically, sentiment analysis can use a domain-adaptive BERT model. This model is fine-tuned on an insurance domain text dataset (including policy documents, industry reports, and public opinion texts) to accurately identify the sentiment tendency of insurance domain texts. The model's input consists of preprocessed plain text and extracted policy keywords and risk events, and the output is the sentiment polarity (positive, neutral, negative) and a sentiment intensity score (0-10 points, with higher scores indicating a stronger sentiment tendency).

[0051] Furthermore, when extracting features from unstructured text data, policy impact modeling is performed to reflect the correlation between policy keywords and risk events. Specifically, a "policy-risk event" mapping table is first established, associating extracted policy keywords with risk events and clarifying the risk event type and correlation impact degree corresponding to each policy keyword (e.g., "tightening safety production supervision" corresponds to "equipment failure and safety hazard related risk events," with a high correlation impact). Then, the mapping relationship is weighted and quantified to obtain the corresponding policy impact index. The specific policy impact index = Σ(policy weight × risk event correlation degree), where the correlation degree indicates the closeness of the association between policy keywords and risk events (between 0 and 1), which can be calculated through semantic similarity. The policy weight can be dynamically determined based on the timeliness, influence, and severity of the risk event. The newer the policy and the greater its influence, the higher the weight; the more severe the risk event, the higher the weight. Finally, a higher policy impact index indicates a greater impact of the policy on customer risk and a higher policy compliance risk faced by the customer, and vice versa.

[0052] After obtaining the sentiment polarity and policy impact index, they are merged into a unified semantic feature to generate a text feature vector. Specifically, the sentiment polarity is converted into a quantitative value, for example, positive = 1, neutral = 0.5, negative = 0, and multiplied by the normalized sentiment intensity score to obtain the sentiment quantification value. Then, the sentiment quantification value and the policy impact index are normalized respectively, converted into a unified scale quantification feature, and then the features are concatenated to obtain a unified semantic feature. Finally, the unified semantic feature is converted into a fixed-dimensional dense vector representation through linear mapping, thus obtaining a text feature vector. This text feature vector contains sentiment information and policy impact information, and can accurately represent the core connotation of unstructured text data.

[0053] This embodiment extracts key information from unstructured text data and further performs sentiment analysis and policy impact quantification to obtain text feature vectors. This transforms messy unstructured text data into text feature vectors containing sentiment information and policy impact indices, thereby accurately capturing the impact of policies and public opinion on customer risks.

[0054] In one embodiment, such as Figure 5 As shown, when the multi-source heterogeneous data is time-series data, step S202 includes: S501. Perform time-series decomposition processing on the time-series data to obtain corresponding decomposed components; S502. Extract the corresponding time-series features based on the decomposed components, and predict the claims rate based on the time-series features to obtain claims trend data; S503. Based on the claims trend data, perform risk fluctuation analysis and market popularity analysis to obtain the corresponding risk fluctuation index and market popularity index; S504. The risk volatility index and the market popularity index are integrated through time-series correlation to generate the time-series feature vector.

[0055] In this embodiment, when extracting features from time-series data, a time-series decomposition process is first performed. Specifically, the time-series data can be decomposed into three independent components using the STL decomposition method: trend classification, seasonal classification, and residual classification. The trend component reflects the long-term trend of the time-series data, such as the year-on-year increase or decrease in claims ratios; the seasonal component reflects the periodic changes in the time-series data, such as the peak of claims ratios in a certain quarter of each year; and the residual component reflects random fluctuations in the time-series data that cannot be explained by the trend and seasonal components, such as fluctuations in claims ratios caused by sudden events. Through time-series decomposition, the different patterns of change in the time-series data can be clearly separated, facilitating subsequent targeted feature extraction.

[0056] Based on the obtained decomposed components, features that can characterize the time series change patterns and customer risk status are extracted. For example, features such as trend slope and trend change rate are extracted for the trend component to reflect the long-term change trend of the time series data; features such as seasonal cycle, seasonal peak and seasonal fluctuation amplitude are extracted for the seasonal component to reflect the periodic changes of the time series data; and features such as residual variance, residual maximum value and residual fluctuation frequency are extracted for the residual component to reflect the random fluctuation of the time series data. The features of the three components are integrated to obtain complete time series features.

[0057] Based on the extracted time-series features, claims ratio prediction is further performed. Specifically, claims ratio prediction can use a pre-built and trained Prophet model or Transformer-Encoder model. During the prediction process, the extracted time-series features are used as model input. The model captures the trend and seasonal features and long-term time-series dependencies in the time-series features and outputs the monthly claims ratio prediction value for a specified future time, such as 6 months, thereby obtaining claims trend data to reflect changes in customers' future claims risks.

[0058] Furthermore, risk volatility analysis and market heat analysis are conducted on the claims trend data to transform the predicted claims trends into quantifiable risk and market indicators. Risk volatility analysis calculates a risk volatility index based on the predicted claims rate for the next 6 months: Risk Volatility Index = (Maximum Claim Rate for the Next 6 Months - Minimum Claim Rate for the Next 6 Months) / Average Claim Rate for the Next 6 Months. A higher index indicates greater volatility in the customer's future claims risk and stronger uncertainty. The index is also adjusted based on the fluctuation of residual components to ensure accuracy. Market heat analysis calculates a market heat index based on claims trend data combined with concurrent market insurance product sales data: Market Heat Index = (Average Claim Rate for the Next 6 Months / Average Industry Claim Rate for the Same Period) × (Average Industry Sales of Related Products in the Past 3 Months / Average Industry Sales of All Insurance Types in the Past 3 Months). The market heat index ranges from 0 to 1; a higher index indicates higher demand for insurance in the customer's market.

[0059] After obtaining the risk volatility index and market popularity index based on time series analysis, the two are then integrated through time series correlation. This involves binding the risk volatility index and market popularity index with their corresponding time information to form a time series feature combination with timestamps and arranged in chronological order. This ensures that each index has a clear time series position and change relationship. Finally, normalization processing is performed to generate a fixed-dimensional time series feature vector, which can simultaneously represent the risk volatility state and market popularity state that change over time.

[0060] This embodiment captures the trend, cycle, and fluctuation characteristics in time series data by performing time series decomposition and feature extraction. It also generates a time series feature vector containing risk fluctuation and market popularity by combining risk fluctuation analysis and market popularity analysis. This enables accurate prediction of customers' future claims risks and improves the foresight of the recommended plan.

[0061] In one embodiment, such as Figure 6 As shown, when the multi-source heterogeneous data is graph data, step S202 includes: S601. Construct a heterogeneous graph from the graph data, and embed the initial graph structure obtained by the construction to obtain the node embedding of each node in the initial graph structure. S602. Perform centrality and risk propagation analysis based on the node embedding to obtain node centrality and risk transmission data for each node; S603. Perform subgraph mining on the initial graph structure to identify risky subgraphs and add risk labels to the subgraphs. S604. The node centrality, risk transmission data, and subgraph risk labels are concatenated and normalized to obtain the graph feature vector.

[0062] In this embodiment, after collecting the graph data, it only contains the raw information of entities and relationships, which cannot be directly extracted as features. Therefore, heterogeneous graph construction is first performed on the graph data to transform the raw relationship information into a structured graph structure. Specifically, when constructing the heterogeneous graph, the node types include customers (individuals / enterprises), insurance products, risk events, and suppliers. Each node contains corresponding attribute information. For example, customer nodes include risk level and insurance records, while product nodes include coverage scope and premiums. The edge relationship types include cooperation relationships (customers and suppliers), insurance relationships (customers and products), and claims transmission relationships (customers and customers, customers and risk events). Based on the above node types and edge relationship types, nodes are set up and connecting edges are established between nodes to construct an initial graph structure, which is stored in the form of node-edge-attribute to ensure the integrity and accuracy of the graph structure.

[0063] The initial graph structure is then embedded using a Graph Attention Network (GAT) model to obtain the node embeddings for each node. The GAT model includes a graph attention layer, an activation layer, and a normalization layer. The graph attention layer calculates attention weights between nodes based on their neighbor information, highlighting the influence of important neighbors. The activation layer uses the ReLU activation function to enhance the model's non-linear expressive power. The normalization layer normalizes the node embedding vectors to ensure their stability. After graph embedding, the low-dimensional embedding vectors of each node are obtained, which preserve the structural and attribute features of the nodes, facilitating subsequent graph analysis.

[0064] Specifically, based on node embedding, centrality and risk propagation analysis are performed on the initial graph structure. Node centrality can be calculated using the PageRank algorithm. By calculating the in-degree and out-degree of nodes and combining the edge association strength, the importance of a node in the graph structure is quantified. Node centrality reflects the importance of a node in the graph structure; higher centrality indicates a greater influence of the node (e.g., core supplier nodes have higher centrality). Risk propagation analysis can be performed using the PageRank algorithm combined with Monte Carlo simulation. First, the PageRank algorithm calculates the risk transmission capability of each node. Then, Monte Carlo simulation simulates the risk transmission process in the graph structure, calculating the probability and intensity of risk transmission to customer nodes, i.e., risk transmission data. This risk transmission data includes the probability of risk transmission (the probability that a customer node will be subject to risk transmission from other nodes) and the intensity of risk transmission (the severity of risk transmission). Higher values ​​indicate greater risk transmission pressure faced by the customer.

[0065] In addition, subgraph mining is performed on the initial graph structure to identify high-risk subgraphs related to customer nodes (such as multiple customers concentrated in the same supplier, and that supplier has negative public opinion). Specifically, the nodes in the initial graph structure can be clustered into multiple subgraphs using graph clustering algorithms. Each subgraph contains nodes with strong correlations (such as customer and supplier nodes in the same supply chain). Then, each subgraph is screened for risk. The screening criteria include the number of risk event nodes in the subgraph, the average risk level of the nodes, and the risk transmission intensity. If a subgraph meets the preset risk conditions, such as the number of risk event nodes ≥ 1, the average risk level ≥ 0.6, and the risk transmission intensity ≥ 0.5, it is determined to be a high-risk subgraph. Corresponding subgraph risk labels are added to the high-risk subgraphs obtained by subgraph mining to reflect the local risks faced by each node.

[0066] The node centrality, risk transmission data, and subgraph risk labels obtained from the initial graph structure are all core quantitative features of graph data. Therefore, these three are spliced ​​and normalized to integrate the structural features, risk transmission features, and local risk features of the graph structure, generating a fixed-dimensional graph feature vector that comprehensively represents the customer's risk status in the graph structure.

[0067] This embodiment analyzes the node centrality and risk propagation of the constructed graph structure and mines high-risk subgraphs to extract the structural features and risk transmission features of customers in the associated network, generating a comprehensive graph feature vector, thereby accurately capturing the associated risks faced by customers.

[0068] In one embodiment, such as Figure 7 As shown, step S203 includes: S701. The structured feature vector, text feature vector, temporal feature vector and graph feature vector are respectively input into an independent domain encoder for encoding processing to obtain structured encoding vector, text encoding vector, temporal encoding vector and graph encoding vector; S702. Dynamic scoring is performed on structured data, unstructured text data, time-series data and graph data according to preset indicators through a preset attention scoring network to obtain dynamic scores for each data item. S703. Calculate the attention score for each data item based on the dynamic score, and perform attention-weighted fusion processing on the structured coding vector, text coding vector, temporal coding vector and graph coding vector based on the attention score to output a unified customer risk profile vector.

[0069] In this embodiment, when performing multi-source fusion processing on feature vectors extracted from multi-source heterogeneous data, they are first input into independent domain encoders for deep encoding and dimensional unification, resulting in corresponding structured encoding vectors, text encoding vectors, temporal encoding vectors, and graph encoding vectors, enabling these four types of encoding vectors to be mapped to the same semantic space. The selection and structure of the four types of domain encoders must be adapted to the characteristics of the corresponding feature vectors. For example, structured feature vectors are input into an MLP encoder, which includes three fully connected layers, a ReLU activation layer, and a normalization layer. This encoder extracts the deep semantics of the structured features through nonlinear transformations and outputs structured encoding vectors. Text feature vectors are input into a BERT+Pooling encoder, which is based on a BERT model fine-tuned for the insurance domain. The encoder inputs the text feature vectors into the encoding layer of the BERT model to extract deep semantic features, and then uses a Pooling layer (global average pooling) to compress the variable-length semantic features into a fixed dimension before outputting the text encoding vector. Temporal feature vectors are input into a Temporal encoder. The Transformer encoder comprises a temporal position encoding layer, a multi-head self-attention layer, and a fully connected layer. The temporal position encoding layer injects temporal information into the temporal feature vector, the multi-head self-attention layer captures the long temporal dependencies of the temporal features, and the fully connected layer maps the features to a fixed dimension before outputting the temporal encoded vector. The graph feature vector is then input into the GAT encoder, which comprises a graph attention layer, an activation layer, and a fully connected layer. The graph attention layer captures the structural relationships of the graph features, the activation layer enhances the nonlinear expression, and the fully connected layer maps the graph features to a fixed dimension before outputting the graph encoded vector.

[0070] At nodes that fuse various encoded vectors, a learnable attention fusion mechanism is employed:

[0071] in It is a dynamic score calculated by a lightweight attention scoring network based on the timeliness, completeness, and confidence of the data source. For the attention score of the i-th data class, This is the encoding vector for the i-th type of data.

[0072] The system first dynamically scores data based on pre-defined indicators for structured data, unstructured text data, time-series data, and graph data. These pre-defined indicators include data timeliness, data completeness, and data confidence. Timeliness is scored based on the data's update time; the more recent the update, the higher the score. For example, data updated within the last month receives 10 points, data updated within the last three months receives 8 points, and so on. Completeness is scored based on the missing data rate; the lower the missing rate, the higher the score. For example, a missing rate of 0% receives 10 points, a missing rate of ≤5% receives 8 points, and so on. Confidence is scored based on the reliability of the data source; the more reliable the source, the higher the score. For example, official data sources receive 10 points, third-party data sources receive 8 points, and manually entered data receives 6 points, and so on. Attention scoring networks can employ lightweight fully connected neural networks. By using fully connected projection layers to weight and fuse the scores of three preset indicators for various data types, an initial feature representation for each data source is obtained. The learnable weights within the network automatically strengthen the indicators that are more critical to the recommendation results (such as strengthening timeliness when policies are updated and strengthening confidence when claims are predicted), thereby focusing attention on the quality dimension of the data source. Then, the projection results are non-linearly transformed through activation layers to output a single scalar score corresponding to each data source type, thus outputting the dynamic scores for each data item.

[0073] Then, attention scores are calculated for each data point based on the dynamic scores. The dynamic scores are converted into attention scores to represent the weights of various coding vectors. Finally, attention-weighted fusion processing is performed on the four types of coding vectors based on the attention scores to generate a unified customer risk profile vector. This vector can integrate the core information of various data to accurately represent the customer's risk profile, so as to obtain the best product recommendation solution.

[0074] This embodiment achieves dimensional unification and semantic alignment of four types of feature vectors through a domain encoder, and combines an attention scoring network to dynamically allocate weights, realizing adaptive weighted fusion of multi-source features to generate a more accurate and comprehensive customer risk profile vector, improving the accuracy of risk profiles, and thus enabling more accurate personalized insurance plan recommendations.

[0075] In one embodiment, such as Figure 8 As shown, step S204 includes: S801. Calculate the matching degree between the customer risk profile vector and each insurance product in the preset product knowledge base from several evaluation dimensions to obtain the matching degree of the corresponding evaluation dimensions. S802. The matching degree of each evaluation dimension is weighted and summed according to the preset weight strategy to obtain the matching score between each insurance product and the customer. S803. Select the top specified number of insurance products with the highest matching scores as recommended products, and generate the recommendation reasons for each recommended product; S804. The insurance recommendation scheme is generated by combining the recommended products with the corresponding recommendation reasons.

[0076] In this embodiment, the matching degree between the customer risk profile vector and each insurance product in the preset product knowledge base is calculated from different evaluation dimensions. The specific evaluation dimensions include coverage matching degree, risk and underwriting strategy fit degree, cost-effectiveness, and regulatory risk. Among them, coverage matching degree (CoverageFit) can be calculated based on the semantic similarity between the customer risk profile and the coverage of each insurance product to obtain the similarity between the customer's core risk points and the product's coverage, which is the coverage matching degree. Risk and underwriting strategy fit degree (RiskMatch) can be calculated based on the customer risk profile vector and the underwriting strategy of each product (such as the underwriting risk level range, deductible conditions) to calculate the degree of matching between the customer's risk characteristics and the product's underwriting strategy, which is the risk and underwriting strategy fit degree. Cost-effectiveness (CostEfficiency) can be obtained by the formula cost-effectiveness = unit sum assured / unit risk cost. Regulatory risk (RegulatoryRisk) assesses whether the insurance product violates the current regulatory policy. Based on the extracted policy keywords, policy impact index, and other information, the core attributes of each insurance product are extracted from the product knowledge base and confirmed through semantic matching and rule verification to confirm whether the product violates the current policy.

[0077] The matching degree of each evaluation dimension is weighted and summed according to a preset weighting strategy, that is, the matching score between each insurance product and the customer is: S(p,c)=ω1 CoverageFit+ω2 RiskMatch+ω3 CostEfficiency ω4 Regulatory Risk ω1, ω2, ω3, and ω4 correspond to the weights of coverage matching, risk and underwriting strategy fit, cost-effectiveness, and regulatory risk. The specific weighting strategy can be a pre-allocated fixed weight or a dynamic weighting method, that is, combining the core characteristics of the customer risk profile vector to adaptively adjust the weights of each assessment dimension, avoiding matching bias caused by fixed weights and ensuring the accuracy of matching scores. For example, the higher the customer's risk level, the higher the weight of risk and underwriting strategy fit; the higher the policy impact index, the higher the weight of policy compliance; the higher the customer's sensitivity to premiums, the higher the weight of cost-effectiveness, and so on.

[0078] After calculating the matching score between each insurance product and the customer, a Top-K recommended product is output by sorting. This means selecting the top specified number of insurance products with the highest matching scores as recommended products, and generating a recommendation reason for each product. The specific number of recommended products can be adjusted according to actual needs; this embodiment does not limit this. When generating the recommendation reason, it is output in natural language. For example, it can retrieve specific data about the recommended product from a preset product knowledge base, including core data such as the product's coverage, underwriting conditions, premium and sum assured parameters, and policy compliance attributes. Secondly, it simultaneously retrieves core feature data from the customer's risk profile vector (such as risk transmission probability, policy impact index, risk level, and claims trends). Finally, it correlates these two types of data, expressing the fit between the core data of the recommended product and the customer's risk characteristics in easily understandable natural language, such as "XX insurance is recommended because of rising negative public opinion in your industry and a high transmission risk shown in the supply chain dependency graph."

[0079] Finally, the selected recommended products are combined with their corresponding reasons to generate an insurance recommendation plan. For example, a personalized list of recommended insurance plans can be output, with a reason for recommendation in natural language attached to each recommended product. Furthermore, a complete insurance recommendation plan can also include detailed product information (coverage amount, premium, coverage period, underwriting conditions), premium calculation details, and claims process guidance. The content is adaptively adjusted based on the terminal initiating the recommendation request. For example, when pushed to a sales terminal, sales staff prompts can be generated simultaneously to assist sales staff in accurately recommending products to customers; when pushed to a customer terminal, technical terms are simplified and visual charts (such as product matching comparison charts) are added to improve the customer's reading experience.

[0080] It should be noted that there is no necessary order between the above steps. Those skilled in the art will understand from the description of the embodiments of the present invention that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.

[0081] Further reference Figure 9 As a response to the above Figure 2 The present invention provides an embodiment of a multi-source data fusion insurance scheme recommendation device, which, in accordance with the implementation of the method shown, provides an embodiment of such a device. Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0082] like Figure 9 As shown, the insurance scheme recommendation device 90 for multi-source data fusion described in this embodiment includes: The data acquisition module 901 is used to collect multi-source heterogeneous data of the target customer in response to the insurance plan recommendation request. The multi-source heterogeneous data includes structured data, unstructured text data, time series data and graph data. Data processing module 902 is used to perform corresponding feature extraction processing on the multi-source heterogeneous data to obtain structured feature vectors, text feature vectors, time-series feature vectors and graph feature vectors respectively; The multi-dimensional fusion module 903 is used to perform multi-dimensional fusion processing on the structured feature vector, text feature vector, time-series feature vector and graph feature vector, and output a unified customer risk profile vector. The solution recommendation module 904 is used to perform product matching evaluation between the customer risk profile vector and the preset product knowledge base, and generate corresponding insurance recommendation solutions based on the matching evaluation results.

[0083] The module referred to in this invention is a series of computer program instruction segments that can perform specific functions. It is more suitable than a program for describing the recommended execution process of insurance schemes for multi-source data fusion. For the specific implementation of each module, please refer to the corresponding method embodiments above, which will not be repeated here.

[0084] In one embodiment, the data processing module 902 includes: The data compliance processing unit is used to perform data cleaning and outlier removal on the structured data to obtain compliant and valid data. The feature engineering unit is used to perform feature engineering processing on the compliant and valid data, extracting risk factors, enterprise profile features, and product matching features from the compliant and valid data to obtain a set of business features. The standardization mapping unit is used to generate the structured feature vector by standardizing and mapping the set of business features.

[0085] In one embodiment, the data processing module 902 includes: The preprocessing and key information processing unit is used to preprocess the unstructured text data and extract key information from the preprocessed data to obtain policy keywords and risk events. The sentiment analysis unit is used to perform sentiment analysis based on the policy keywords and risk events to obtain the corresponding sentiment polarity. The policy impact unit is used to establish a mapping relationship between the policy keywords and risk events and to perform weighted quantification to generate a corresponding policy impact index. The text feature unit is used to generate the text feature vector by fusing the emotion polarity and the policy impact index into a unified semantic feature.

[0086] In one embodiment, the data processing module 902 includes: The time-series decomposition unit is used to perform time-series decomposition processing on the time-series data to obtain corresponding decomposed components; The trend prediction unit is used to extract corresponding time-series features based on the decomposed components, and to predict the claims rate based on the time-series features to obtain claims trend data. The trend analysis unit is used to perform risk fluctuation analysis and market popularity analysis based on the claims trend data, and obtain the corresponding risk fluctuation index and market popularity index. The time-series feature unit is used to integrate the risk volatility index and the market popularity index in a time-series manner to generate the time-series feature vector.

[0087] In one embodiment, the data processing module 902 includes: The graph construction unit is used to construct a heterogeneous graph from the graph data and embed the constructed initial graph structure to obtain the node embedding of each node in the initial graph structure. The risk propagation unit is used to perform centrality and risk propagation analysis based on the node embedding to obtain the node centrality and risk transmission data of each node. The subgraph mining unit is used to perform subgraph mining on the initial graph structure, identify risky subgraphs, and add risk labels to the subgraphs. The graph feature unit is used to concatenate and normalize the node centrality, risk transmission data, and subgraph risk labels to obtain the graph feature vector.

[0088] In one embodiment, the multidimensional fusion module 903 includes: The domain coding unit is used to input the structured feature vector, text feature vector, temporal feature vector and graph feature vector into an independent domain encoder for encoding processing, so as to obtain structured coding vector, text coding vector, temporal coding vector and graph coding vector respectively; The dynamic scoring unit is used to dynamically score the data based on preset indicators of structured data, unstructured text data, time-series data, and graph data through a preset attention scoring network, and obtain dynamic scores for each data item. The multi-source fusion unit is used to calculate the attention score of each data item according to the dynamic score, and to perform attention-weighted fusion processing on the structured coding vector, text coding vector, temporal coding vector and graph coding vector according to the attention score, and output a unified customer risk profile vector.

[0089] In one embodiment, the scheme recommendation module 904 includes: The matching calculation unit is used to calculate the matching degree between the customer risk profile vector and each insurance product in the preset product knowledge base from several evaluation dimensions, and obtain the matching degree of the corresponding evaluation dimensions. The weighted summation unit is used to perform weighted summation on the matching degree of each evaluation dimension according to a preset weighting strategy to obtain the matching score between each insurance product and the customer. The recommendation selection unit is used to select the top specified number of insurance products with the highest matching scores as recommended products, and generate the recommendation reasons for each recommended product; The scheme generation unit is used to combine the recommended products with the corresponding recommendation reasons to generate the insurance recommendation scheme.

[0090] In the above embodiments, the present invention discloses an insurance scheme recommendation device based on multi-source data fusion. By collecting heterogeneous data from multiple sources and performing multi-source fusion processing, it accurately and comprehensively captures customer risk profiles, and then realizes product matching and recommendation based on accurate customer risk profiles, thereby improving the personalization and accuracy of the recommendation schemes.

[0091] Specific limitations regarding the insurance scheme recommendation device for multi-source data fusion can be found in the limitations of the insurance scheme recommendation method for multi-source data fusion described above, and will not be repeated here. Each module in the aforementioned insurance scheme recommendation device for multi-source data fusion can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0092] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0093] Another embodiment of the present invention provides a computer device, such as... Figure 10 As shown, the computer device 100 includes: One or more processors 101 and memory 102, Figure 10 The following description uses a processor 101 as an example. The processor 101 and the memory 102 can be connected via a bus or other means. Figure 10 Taking the example of a connection between China and Israel via a bus.

[0094] The processor 101 is used to perform various control logics of the computer device 100. It can be any conventional processor, microprocessor, state machine, general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), microcontroller, ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components.

[0095] The memory 102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the multi-source data fusion insurance scheme recommendation method in the embodiments of the present invention. The processor 101 executes various functional applications and data processing of the computer device 100 by running the non-volatile software programs, instructions, and units stored in the memory 102, thereby implementing the multi-source data fusion insurance scheme recommendation method in the above method embodiments.

[0096] Another embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by one or more processors, perform the steps of the insurance scheme recommendation method for multi-source data fusion in any of the above method embodiments.

[0097] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0098] Based on the above description of the embodiments, those skilled in the art will understand that the methods described in the embodiments can be implemented using software plus necessary general-purpose hardware platforms. Of course, they can also be implemented using hardware, but in many cases, the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0099] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The computer program can be stored in a non-volatile, computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The storage medium can be a memory, magnetic disk, floppy disk, flash memory, optical storage, etc.

[0100] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

[0101] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for recommending insurance schemes based on multi-source data fusion, characterized in that, include: In response to insurance plan recommendation requests, multi-source heterogeneous data of target customers are collected, including structured data, unstructured text data, time-series data, and graph data. The multi-source heterogeneous data is subjected to corresponding feature extraction processing to obtain structured feature vectors, text feature vectors, time-series feature vectors, and graph feature vectors, respectively. After performing multi-dimensional fusion processing on the structured feature vector, text feature vector, time-series feature vector, and graph feature vector, a unified customer risk profile vector is output. The customer risk profile vector is matched with a preset product knowledge base for product matching evaluation, and a corresponding insurance recommendation plan is generated based on the matching evaluation results.

2. The insurance scheme recommendation method based on multi-source data fusion according to claim 1, characterized in that, When the multi-source heterogeneous data is structured data, the corresponding feature extraction processing of the multi-source heterogeneous data includes: The structured data is cleaned and outlier removed to obtain compliant and valid data; The compliant and valid data is subjected to feature engineering to extract risk factors, enterprise profile features, and product matching features from the compliant and valid data to obtain a set of business features. The structured feature vector is generated by standardizing and mapping the set of business features.

3. The insurance scheme recommendation method based on multi-source data fusion according to claim 1, characterized in that, When the multi-source heterogeneous data is unstructured text data, the corresponding feature extraction processing of the multi-source heterogeneous data includes: The unstructured text data is preprocessed, and key information is extracted from the preprocessed data to obtain policy keywords and risk events. Sentiment analysis is performed based on the policy keywords and risk events to obtain the corresponding sentiment polarity; Establish a mapping relationship between the policy keywords and risk events and perform weighted quantification to generate a corresponding policy impact index; The text feature vector is generated by integrating the emotional polarity and the policy impact index into a unified semantic feature.

4. The insurance scheme recommendation method based on multi-source data fusion according to claim 1, characterized in that, When the multi-source heterogeneous data is time-series data, the corresponding feature extraction processing of the multi-source heterogeneous data includes: The time-series data is subjected to time-series decomposition processing to obtain the corresponding decomposed components; Based on the decomposed components, corresponding time-series features are extracted, and claims rate prediction is performed based on the time-series features to obtain claims trend data; Based on the claims trend data, risk volatility analysis and market popularity analysis are performed to obtain the corresponding risk volatility index and market popularity index. The risk volatility index and the market popularity index are correlated and integrated over time to generate the time-series feature vector.

5. The insurance scheme recommendation method based on multi-source data fusion according to claim 1, characterized in that, When the multi-source heterogeneous data is graph data, the step of performing corresponding feature extraction processing on the multi-source heterogeneous data includes: Heterogeneous graph construction is performed on the graph data, and graph embedding is performed on the initial graph structure obtained to obtain the node embedding of each node in the initial graph structure. Based on the node embedding, a centrality and risk propagation analysis is performed to obtain the node centrality and risk transmission data of each node; Subgraph mining is performed on the initial graph structure to identify risky subgraphs and add risk labels to them; The node centrality, risk transmission data, and subgraph risk labels are concatenated and normalized to obtain the graph feature vector.

6. The insurance scheme recommendation method based on multi-source data fusion according to claim 1, characterized in that, After performing multi-dimensional fusion processing on the structured feature vector, text feature vector, time-series feature vector, and graph feature vector, a unified customer risk profile vector is output, including: The structured feature vector, text feature vector, temporal feature vector, and graph feature vector are respectively input into an independent domain encoder for encoding processing to obtain structured encoded vector, text encoded vector, temporal encoded vector, and graph encoded vector; The preset attention scoring network dynamically scores structured data, unstructured text data, time-series data, and graph data based on preset indicators, and obtains dynamic scores for each data item. Based on the dynamic scores, attention scores are calculated for each data item. Then, attention-weighted fusion processing is performed on the structured coding vector, text coding vector, temporal coding vector, and graph coding vector based on the attention scores to output a unified customer risk profile vector.

7. The insurance scheme recommendation method based on multi-source data fusion according to claim 1, characterized in that, The step of performing product matching assessments by comparing the customer risk profile vector with a preset product knowledge base, and generating corresponding insurance recommendation schemes based on the matching assessment results, includes: The matching degree between the customer risk profile vector and each insurance product in the preset product knowledge base is calculated from several evaluation dimensions to obtain the matching degree of the corresponding evaluation dimensions. The matching degree of each evaluation dimension is weighted and summed according to the preset weighting strategy to obtain the matching score between each insurance product and the customer. Select the top specified number of insurance products with the highest matching scores as recommended products, and generate a recommendation reason for each recommended product; The insurance recommendation scheme is generated by combining the recommended products with the corresponding reasons for recommendation.

8. An insurance scheme recommendation device based on multi-source data fusion, characterized in that, include: The data acquisition module is used to collect multi-source heterogeneous data of target customers in response to insurance plan recommendation requests. The multi-source heterogeneous data includes structured data, unstructured text data, time series data, and graph data. The data processing module is used to perform corresponding feature extraction processing on the multi-source heterogeneous data to obtain structured feature vectors, text feature vectors, time-series feature vectors, and graph feature vectors, respectively. The multi-dimensional fusion module is used to perform multi-dimensional fusion processing on the structured feature vector, text feature vector, time-series feature vector and graph feature vector, and output a unified customer risk profile vector. The solution recommendation module is used to perform product matching evaluation between the customer risk profile vector and the preset product knowledge base, and generate corresponding insurance recommendation solutions based on the matching evaluation results.

9. A computer device, characterized in that, Includes at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the insurance scheme recommendation method for multi-source data fusion as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform the insurance scheme recommendation method for multi-source data fusion as described in any one of claims 1-7.