Insurance recommendation method, system and terminal based on large language model and multi-source data
By employing a large language model and multi-source data-driven insurance recommendation method, combined with trending news and user behavior data, personalized insurance recommendation texts are generated. This addresses the problem of the limited range of existing insurance recommendation methods and achieves highly accurate and real-time personalized recommendations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AIA LIFE INSURANCE CO LTD
- Filing Date
- 2026-04-28
- Publication Date
- 2026-07-24
AI Technical Summary
Existing insurance recommendation methods are simplistic, the recommended content is out of touch with users' actual needs, the recommendation accuracy is low, the personalization is insufficient, and there is a lack of real-time scenario awareness, resulting in low user click-through rates and high customer churn rates.
By using large language models and multi-source data, insurance risk events in trending news texts are identified. In-depth customer profiles are drawn by combining user behavior data, and multi-dimensional matching is performed in the insurance knowledge graph to generate personalized insurance recommendation texts. After compliance review, the texts are sent to the user's terminal.
It enables personalized insurance recommendations that accurately reach users, improves recommendation accuracy and user interest, reduces customer churn and information overload, and enhances the real-time nature and professionalism of recommendations.
Smart Images

Figure CN122453531A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology, and in particular to an insurance recommendation method, system and terminal based on large language models and multi-source data. Background Technology
[0002] With the improvement of people's living standards and the enhancement of risk awareness, insurance planning has become an indispensable part of family financial planning. However, the existing online insurance information is abundant but of varying quality, making it difficult for users to select insurance types and products that truly meet their needs. Furthermore, it is even more difficult to obtain timely insurance information that is most crucial for their risk protection. Therefore, in the insurance field, it is possible to recommend insurance information to users based on their basic data and insurance needs. However, current technology can only recommend a certain type of insurance or a specific insurance product based on certain characteristics such as risk preference. The recommendation method is relatively simple, and the recommended content is disconnected from the user's actual needs, making it difficult to attract user interest and resulting in low recommendation accuracy. Specifically, the existing insurance recommendation text content and recommendation methods have the following main shortcomings.
[0003] ① Insurance companies' customer management systems, insurance product databases, and insurance market intelligence systems have long operated independently, creating a "data silo" effect. Furthermore, these systems primarily use traditional relational databases, supporting only linear queries and unable to establish unstructured relationship networks linking "users, scenarios, and products." This results in user behavior patterns, policy holding status, and real-time social trends being scattered across different platforms, lacking cross-dimensional correlation mechanisms. Consequently, insurance recommendation texts rely on fragmented data, typically generating recommendation text based only on historical user behavior data, lacking real-time scenario awareness and impacting recommendation accuracy. For example, user browsing behavior data on children's insurance pages cannot be linked in real-time to education policy reform trends, leading to a disconnect between text content and users' actual needs, resulting in a continuous decline in click-through rates.
[0004] ② The recommendation response to social emergencies (such as influenza outbreaks or policy adjustments) requires a four-stage manual process: "event identification, demand analysis, content creation, and compliance review." This high reliance on manual labor and long average response times in the industry result in poor timeliness of recommendations, making it impossible to promptly deliver relevant insurance information to users in need. It also leads to current insurance recommendations relying solely on scheduled recommendations, resulting in a limited and simplistic recommendation method.
[0005] ③ Existing recommendation engines cannot distinguish user interest categories, and each user receives the same recommendation text, resulting in low personalization and accuracy of insurance recommendations. This, in turn, affects user click-through rates, leading to high customer churn rates and low policy renewal rates. Furthermore, it can also cause information intrusion on non-target customers.
[0006] ④ The existing recommendation engine does not optimize for insurance terminology when generating recommendation text, resulting in a high error rate in interpreting terms and conditions, which poses a compliance risk. Summary of the Invention
[0007] In view of the shortcomings of the prior art described above, the purpose of this application is to provide an insurance recommendation method, system and terminal based on a large language model and multi-source data, to solve the technical problems of existing insurance recommendation methods being singular, the recommended content being out of touch with the user's real needs and the recommendation accuracy being low.
[0008] To achieve the above and other related objectives, the first aspect of this application provides an insurance recommendation method based on a large language model and multi-source data. The method includes: periodically acquiring multiple trending news texts and, based on a pre-trained insurance risk event identification model, identifying insurance risks in each trending news text to obtain one or more news events with insurance risks; acquiring user behavior data of a target user and, based on the user behavior data, identifying the target user's lifecycle events and / or behavioral events to map the lifecycle events, behavioral events, and each news event to the same risk event space to obtain a target risk event; based on a pre-built insurance recommendation text generation model, drawing a deep customer profile of the target user based on the user behavior data, and, combined with the target risk event, performing multi-dimensional matching and retrieval in a pre-built insurance knowledge graph to filter suitable recommended insurance types and recommendation text templates for the target user, thereby generating an initial insurance recommendation text; based on a pre-built insurance clause rule base, conducting compliance review on the initial insurance recommendation text, correcting any violations, and sending the obtained target insurance recommendation text to the target user's terminal via a communication interface.
[0009] In some embodiments of the first aspect of this application, the method of identifying insurance risks in various hot news texts based on a pre-trained insurance risk event identification model to obtain one or more news events with insurance risks includes: performing natural language processing on each hot news text to obtain one or more news keywords and news description data for each hot news text; identifying one or more candidate news texts related to insurance risks based on each news keyword, and marking the associated insurance types of each candidate news text; evaluating the relevance of each candidate news text to its associated insurance type, eliminating candidate news texts with low relevance, and obtaining one or more news texts; obtaining news description data for each news text, and extracting news event elements from each news text based on each news description data to obtain multiple corresponding news events.
[0010] In some embodiments of the first aspect of this application, the lifecycle events, behavioral events, and news events are mapped to the same risk event space to obtain target risk events. This includes: treating the lifecycle events, behavioral events, and news events as risk events, and performing event element analysis on each risk event to extract the theme features, object features, time features, regional features, and risk features of each risk event to generate risk event feature vectors for each risk event; performing dimensional alignment and semantic normalization on each risk event feature vector to generate multiple standardized risk event representations within the same risk event space; and then, based on the target user's... User behavior data and the thematic, object, temporal, geographical, and risk characteristics of each standardized risk event are used to assess the object proximity, temporal proximity, geographical proximity, and risk proximity of each risk event to the target user, and to calculate the scenario relevance of each risk event to the target user. Based on the target user's user behavior data, the policy status, coverage, exclusions, coverage range, and coverage period of the target user's current policies are obtained, and the target user's protection gap is calculated to screen one or more candidate risk events that match the protection gap. The candidate risk event with the highest scenario relevance to the target user is identified as the target risk event.
[0011] In some embodiments of the first aspect of this application, the method for generating initial insurance recommendation text based on the insurance recommendation text generation model includes: drawing a deep customer profile of the target user based on the user behavior data; performing multi-dimensional matching retrieval and cross-entity association operations based on a pre-built insurance knowledge graph, according to the deep customer profile of the target user and the target risk event, to screen one or more candidate recommendable insurance products suitable for the target user; evaluating the relevance of each candidate recommendable insurance product to the needs of the target user, and determining the candidate recommendable insurance product with the highest relevance to the needs as the recommendable insurance product; obtaining a recommendation text template for the recommendable insurance product based on the insurance knowledge graph, and filling the recommendation text template to generate basic insurance recommendation text for the recommendable insurance product; selecting an appropriate insurance recommendation text generation strategy based on the deep customer profile of the target user, and constructing a contextualized narrative structure based on the insurance recommendation text generation strategy, and embedding the target risk event into the storyline and injecting emotional expression to optimize the basic insurance recommendation text and generate initial insurance recommendation text; wherein, the insurance recommendation text generation strategy includes at least: narrative template, story map, emotional temperature, language style, and rhetoric library.
[0012] In some embodiments of the first aspect of this application, the in-depth customer profile includes: attribute tags, including but not limited to: age tags, city tags, occupation tags, marital and childbearing tags, asset and liability tags, health tags, and protection gap tags, or one or more combinations thereof; event tags, including but not limited to: behavioral event tags and / or life cycle event tags; psychological and need tags, including but not limited to: emotion tags, interest and preference tags, consumption and risk preference tags, and value tags, or one or more combinations thereof.
[0013] In some embodiments of the first aspect of this application, the method of constructing the insurance knowledge graph includes: acquiring user behavior data of multiple historical insured users, and drawing in-depth customer profiles for each historical insured user based on the user behavior data to construct a customer profile library; defining multiple risk events, and defining one or more associated insurance types for each risk event to construct a risk scenario library; acquiring insurance product data, and obtaining multiple insurance products for multiple insurance types, as well as the insured objects, insured responsibilities, exclusion clauses, and recommended text templates for each insurance product based on the insurance product data, and establishing an insurance type classification tree; defining the entity types of the insurance knowledge graph as insurance products, in-depth customer profiles, risk events, and recommended text templates, and instantiating each insurance product, each in-depth customer profile, each risk event, and each recommended text template as multiple entity nodes, reasoning the relationships between each entity node, and constructing the insurance knowledge graph.
[0014] In some embodiments of the first aspect of this application, the construction method of the insurance recommendation text generation model includes: acquiring user behavior data of multiple historical insured users and multiple historical news events to construct a training sample dataset; inputting the training sample dataset into a large language model and introducing professional corpus in the insurance field, training the large language model using a self-supervised learning method to obtain an insurance recommendation text generation model; labeling and sorting the insurance recommendation texts of some historical insured users output by the insurance recommendation text generation model according to their matching degree to generate insurance recommendation text preference data, so as to train a text encoder to obtain a reward model for evaluating each insurance recommendation text output by the insurance recommendation text generation model; iteratively optimizing the insurance recommendation text generation model through a reinforcement learning algorithm so that the insurance recommendation text generation model generates insurance recommendation texts with high rewards.
[0015] In some embodiments of the first aspect of this application, the insurance recommendation method based on a large language model and multi-source data further includes: monitoring the click-through rate and conversion rate of the target insurance recommendation text, and optimizing the insurance recommendation text generation model accordingly.
[0016] To achieve the above and other related objectives, a second aspect of this application provides an insurance recommendation system based on a large language model and multi-source data. The system includes: a news event identification module, used to periodically acquire multiple trending news texts and, based on a pre-trained insurance risk event identification model, identify the insurance risks of each trending news text to obtain one or more news events with insurance risks; and a risk event determination module, connected to the news event identification module, used to acquire user behavior data of a target user and, based on the user behavior data, identify the target user's lifecycle events and / or behavioral events, so as to map the lifecycle events, the behavioral events, and each news event to the same risk event space to obtain the target risk. The system comprises the following modules: an event and a text generation module, connected to the risk event determination module. The event generation module generates an in-depth customer profile of the target user based on a pre-built insurance recommendation text generation model and user behavior data. It then performs multi-dimensional matching and retrieval within a pre-built insurance knowledge graph, selecting suitable insurance types and text templates to generate initial insurance recommendation text. A text compliance review module, connected to the text generation module, performs compliance review on the initial insurance recommendation text based on a pre-built insurance clause rule base, correcting any violations to obtain the target insurance recommendation text. Finally, a text push module, connected to the text compliance review module, sends the target insurance recommendation text to the target user's terminal via a communication interface.
[0017] To achieve the above and other related objectives, a third aspect of this application provides an insurance recommendation terminal based on a large language model and multi-source data. The insurance recommendation terminal includes a memory and a processor; the memory stores a computer program; and the processor executes the computer program stored in the memory to enable the terminal to perform the insurance recommendation method based on a large language model and multi-source data as described in any of the above embodiments.
[0018] As described above, this application has the following beneficial effects: This application provides an insurance recommendation method, system, and terminal based on a large language model and multi-source data. By identifying insurance risks from multiple periodically acquired hot news texts, one or more news events with insurance risks are obtained; and by combining the target user's lifecycle events and / or behavioral events identified based on user behavior data, target risk events are selected to trigger a scenario-based insurance recommendation task; simultaneously, based on a pre-built insurance recommendation text generation model and an insurance knowledge graph, personalized initial insurance recommendation text that meets the target user's real needs is generated according to the target user's in-depth customer profile and target risk events. After the initial insurance recommendation text passes compliance review, the obtained more professional target insurance recommendation text is sent to the target user's terminal to achieve precise reach. Attached Figure Description
[0019] Figure 1 The diagram shown is a flowchart illustrating an insurance recommendation process based on a large language model and multi-source data in one embodiment of this application.
[0020] Figure 2 The diagram shown is a flowchart of a news event identification method in one embodiment of this application.
[0021] Figure 3 The diagram shown is a flowchart illustrating a method for determining target risk events in one embodiment of this application.
[0022] Figure 4 The diagram shown is a flowchart illustrating an insurance recommendation text generation method in one embodiment of this application.
[0023] Figure 5 The diagram shown is a structural schematic of an insurance recommendation system based on a large language model and multi-source data in one embodiment of this application.
[0024] Figure 6 The diagram shown is a structural schematic of an insurance recommendation terminal based on a large language model and multi-source data in one embodiment of this application. Detailed Implementation
[0025] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.
[0026] In the embodiments of this application, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect, without limiting their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" do not necessarily imply that they are different.
[0027] To address the problems mentioned above, this application provides an insurance recommendation method, system, and terminal based on a large language model and multi-source data, aiming to solve the technical problems of existing insurance recommendation methods being simplistic, having content that is out of touch with users' actual needs, and having low recommendation accuracy.
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application are further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining this application and are not intended to limit this application.
[0029] like Figure 1 The diagram illustrates a flowchart of an insurance recommendation method based on a large language model and multi-source data, as described in this embodiment. The insurance recommendation method based on a large language model and multi-source data in this embodiment mainly includes the following steps.
[0030] Step S1: Periodically acquire multiple trending news texts, and based on a pre-trained insurance risk event identification model, identify insurance risks in each trending news text to obtain one or more news events with insurance risks.
[0031] Among them, trending news mainly refers to recent domestic news reports that have garnered significant attention. In one embodiment, multiple news data can be filtered based on their discussion volume, interaction volume, rate of change in dissemination popularity, affected area, and impact effectiveness. One or more news data that meet preset conditions are designated as trending news, and trending news text is generated. In this embodiment, the acquired news data includes one of the following modalities: text, audio, and video. When the determined trending news is in audio or video modal, automatic speech recognition technology can be used to convert the audio or video trending news into text modal to generate trending news text.
[0032] The insurance risk event identification model can be trained based on a large language model. It should be understood that a large language model (LLM) is a deep learning-based natural language processing technique designed to train large-scale models capable of understanding, processing, and generating natural language text. The core idea of LLM is to use deep neural networks to pre-train models on large-scale text data, and then use these pre-trained models for fine-tuning or direct application to downstream tasks. In this embodiment, the insurance risk event identification model is trained based on a large language model, and then the model performs a news text classification task, thereby identifying whether various trending news texts contain insurance risks and outputting the corresponding news events.
[0033] In one embodiment, the method of training a large language model to obtain the insurance risk event recognition model includes: acquiring multiple historical hot news texts, and associating some of the historical hot news texts with insurance types to construct a training sample dataset; inputting the training sample dataset into the large language model, training the large language model using a semi-supervised learning method to obtain the insurance risk event recognition model; defining a loss function, and iteratively optimizing the insurance risk event recognition model using an optimization algorithm until the model converges.
[0034] Preferably, when marking certain historical hot news texts with associated insurance types, the associated insurance types marked include, but are not limited to: one type of personal insurance, such as critical illness insurance, medical insurance, accident insurance, life insurance, annuity insurance (including: retirement annuity insurance, education annuity insurance, and term annuity insurance), and one type of increasing whole life insurance, so that the insurance risk event identification model can identify insurance risk types including: health risk, accident risk, education risk, and retirement risk, etc.
[0035] It should be noted that this application does not specifically limit the types of insurance risks that the model can identify; users can set them according to their needs. In other preferred embodiments, users can also add a type of property insurance as an associated insurance category, specifically including one of auto insurance, household property insurance, liability insurance, and corporate property insurance, thereby enabling the insurance risk event identification model to identify a wide range of property risks.
[0036] In one embodiment, such as Figure 2 As shown, based on the insurance risk event identification model, the method of identifying insurance risks in various hot news texts and obtaining one or more news events with insurance risks includes the following steps.
[0037] Step S11: Perform natural language processing on each hot news text to obtain one or more news keywords and news description data for each hot news text.
[0038] Each news keyword is used to reflect the news details and background of each trending news text. Specifically, each news keyword can be an entity keyword, including the people involved in the news (such as the subject of the loss), the location of the news, and the news-specific matters (such as the laws and policies involved); each news keyword can also be an action keyword used to describe the core event of the news, and / or a theme keyword used to summarize the topic to which the news belongs, etc.
[0039] The news description data shall include at least the following detailed information: news source, time of occurrence, location of occurrence, background of occurrence, content of occurrence, volume of discussion, volume of interaction, rate of change in the popularity of news dissemination, area of influence, and impact of news.
[0040] Step S12: Based on each news keyword, identify one or more candidate news texts related to insurance risks, and mark the associated insurance types of each candidate news text.
[0041] In this embodiment, based on the insurance risk event identification model, semantic understanding, analysis, and reasoning are performed on each news keyword to determine whether each trending news text carries an insurance risk. One or more trending news texts with insurance risks are selected as candidate news texts, and the associated insurance types and corresponding confidence levels are labeled for each candidate news text. It should be noted that each candidate news text may be labeled with multiple associated insurance types; this application does not specifically limit this.
[0042] Step S13: Evaluate the relevance of each candidate news text to its associated insurance product, eliminate candidate news texts with low relevance, and obtain one or more news texts.
[0043] Specifically, the correlation between each candidate news text and its associated insurance products can be assessed based on the confidence level of the associated insurance products tagged with each candidate news text. When the confidence level is lower than a preset confidence threshold, the corresponding candidate news text is considered to have a low correlation with its associated insurance products and is therefore removed. The remaining candidate news texts are then considered to meet the correlation requirements. When a candidate news text has multiple associated insurance products, those with a confidence level lower than the preset confidence threshold are first removed. Furthermore, if the confidence levels of the candidate news text with all associated insurance products are lower than the preset confidence threshold, the candidate news text is then removed. In this case, the resulting news text may also be tagged with multiple associated insurance products.
[0044] Step S14: Obtain the news description data of each news text, and extract the news event elements from each news text according to the news description data to obtain the corresponding news event.
[0045] In one embodiment, the method of extracting news event elements from the news text includes: performing entity recognition, event extraction, regional recognition, timeliness recognition, and risk classification on the news text to obtain the object, occurrence time, duration of impact, affected area, occurrence process, subject of loss, risk type, and one or more related insurance types of the corresponding news event, thereby obtaining the news event.
[0046] Once the subject of the loss is identified as a person, relevant information about the subject (such as age, gender, occupation, marital status, etc.) can be extracted to create a profile of the subject. This profile helps in subsequent screening of risk events suitable for target users and generating target insurance recommendation text. It should be noted that in other embodiments, the subject of the loss may also be property such as houses, vehicles, or pets; this application is not specifically limited to these categories.
[0047] The risk types of the news events mentioned include, but are not limited to, one of the following: epidemic diseases, policies and regulations, natural disasters, and accidents.
[0048] Step S2: Obtain user behavior data of the target user, and based on the user behavior data, identify the target user's lifecycle events and / or behavioral events, so as to map the lifecycle events, the behavioral events and each news event to the same risk event space to obtain the target risk event.
[0049] The user behavior data includes, but is not limited to, one or more combinations of the target user's basic data, physical examination data, family data, insurance policy data, and behavioral data.
[0050] Specifically, the family data includes, but is not limited to, one or more combinations of the following: basic data, medical examination data, policy data, and behavioral data of one or more immediate family members of the target user; the basic data includes, but is not limited to, one or more combinations of the following: the target user's age, city, occupation, marital and childbearing status, and asset and liability status; the policy data includes, but is not limited to, one or more combinations of the following: the policy status, coverage, exclusions, coverage range, coverage period, and historical claims records of one or more policies currently held by the target user; the behavioral data includes, but is not limited to, one or more combinations of the following: the target user's browsing history, search history, consultation history, insurance application history, consumption history, medical examination history, travel history, and other behavioral records.
[0051] In this embodiment, the user behavior data can be obtained from different external systems, such as insurance company customer management systems, cooperative medical institution systems, internet platforms and / or social platforms, thereby integrating structured and unstructured data from multiple data sources, enriching the data content, solving the data silo problem, and helping to form a comprehensive and three-dimensional in-depth customer profile, which is the foundation for realizing personalized and accurate insurance text recommendations.
[0052] The types of lifecycle events identified based on the multi-source heterogeneous user behavior data include: birth of a new family member, family member starting school, adulthood, independence, first employment, marriage, home purchase, significant debt, birth of a child, career advancement, income increase, middle age, family maturity, retirement, illness, accident, death, and total disability. The identified behavioral events are used to describe specific behaviors of the target user, and their types include, but are not limited to, purchasing airline tickets, undergoing medical examinations, seeking medical treatment, traveling, paying attention to insurance, and policy expiration.
[0053] In one embodiment, the user behavior data can be semantically understood and analyzed based on a large language model to identify the target user's lifecycle events and / or behavioral events. Specifically, a pre-built insurance recommendation text generation model can be used to extract the target user's behavioral features and analyze the target user's behavioral patterns by performing semantic understanding and analysis on the user behavior data, thereby identifying the target user's lifecycle events and / or behavioral events.
[0054] In this embodiment, after identifying the lifecycle events and / or behavioral events of the target user, they are mapped to the same risk event space along with various news events, and target risk events are selected and used as the timing to trigger the insurance recommendation task, thereby realizing scenario-based insurance recommendations that are as appealing to the target user as possible and as close as possible to the target user's real needs.
[0055] In one embodiment, such as Figure 3As shown, the specific methods for screening target risk events include the following steps.
[0056] Step S21: The life cycle event, the behavioral event, and the news event are respectively regarded as risk events, and the event element analysis is performed on each risk event to extract the theme feature, object feature, time feature, regional feature and risk feature of each risk event to generate the risk event feature vector of each risk event.
[0057] The subject feature is used to describe the subject or field to which the risk event belongs; the object feature is used to describe the subject of the risk event, including the subject's age, gender, occupation, marital status, etc.; the time feature is used to describe the time factor such as the occurrence time and duration of the impact of the risk event; the geographical feature is used to describe the geographical factor such as the location of the risk event and the area affected; and the risk feature is used to describe the risk type of the risk event.
[0058] Step S22: Perform dimensional alignment and semantic normalization on the feature vectors of each risk event to generate multiple standardized risk event representations in the same risk event space.
[0059] Step S23: Based on the user behavior data of the target user and the theme characteristics, object characteristics, time characteristics, geographical characteristics and risk characteristics represented by each standardized risk event, evaluate the object proximity, time proximity, geographical proximity and risk proximity between each risk event and the target user, and calculate the scenario correlation between each risk event and the target user.
[0060] Specifically, the method for calculating the scenario relevance between the risk event and the target user includes: defining weights for object proximity, time proximity, geographical proximity, and risk proximity, and using a weighted average method to calculate the scenario relevance. The weight coefficients for object proximity, time proximity, geographical proximity, and risk proximity can be pre-set based on statistical analysis results of historical high-conversion recommendation cases and calibrated and updated based on online A / B testing results.
[0061] Step S24: Based on the target user's user behavior data, obtain the policy status, coverage, exclusions, coverage range, and coverage period of the target user's current policy, and calculate the target user's protection gap to screen one or more candidate risk events that match the protection gap.
[0062] The coverage gap is mainly used to describe characteristics such as insurance types that the target user has not yet deployed, missing coverage responsibilities, and insurance products that are about to expire.
[0063] Step S25: Identify the candidate risk events with the highest relevance to the target user's scenario as the target risk events.
[0064] In this embodiment, this application integrates the lifecycle events and behavioral events of the target user identified by user behavior data, as well as one or more news events determined by various trending news texts. It selects the target risk events most likely to resonate with the target user as the triggering conditions for scenario-based insurance recommendations, thereby triggering insurance push tasks and responding to key business signals. This enables personalized insurance recommendations to be made when each user needs them most, such as triggering an insurance recommendation task for the user 30 days before the policy needs to be renewed, or triggering an insurance recommendation task for users in the affected areas during an outbreak, or triggering a scenario-based insurance recommendation task for the user when the user searches for "children's education insurance" multiple times. This achieves precise, layered outreach and breaks through the homogeneity bottleneck of traditional insurance text recommendations.
[0065] Step S3: Based on the pre-built insurance recommendation text generation model, draw a deep customer profile for the target user according to the user behavior data, and combine the target risk event to perform multi-dimensional matching and retrieval in the pre-built insurance knowledge graph, filter the recommended insurance types and recommendation text templates that are suitable for the target user, and then generate the initial insurance recommendation text.
[0066] In one embodiment, the insurance recommendation text generation model can be trained based on a large language model, and uses natural language processing technology to understand and generate natural language text. The specific training method of the insurance recommendation text generation model mainly includes the following steps.
[0067] ① Obtain user behavior data from multiple historical insured users and multiple historical news events to construct a training sample dataset.
[0068] The user behavior data of each historical insured user can come from the insurance company's customer management system, the cooperative medical institution system, the Internet platform and / or social platform; the historical news events can come from the insurance market intelligence system or various news public platforms, etc., thereby integrating data from multiple data sources, enriching the data content, enabling the large language model to learn more subtle semantic features and discover deep semantic relationships, improving the accuracy and reliability of the model in text generation tasks, enhancing the model's generalization ability, reducing bias and unfairness, and improving robustness.
[0069] ② Input the training sample dataset into the large language model, and introduce professional corpus in the insurance field. Train the large language model using a self-supervised learning method to obtain an insurance recommendation text generation model.
[0070] The insurance-related corpus includes insurance-related legal documents, insurance industry normative documents, insurance product brochures, insurance contracts, protection plans, insurance product promotional materials, and claims rules. In this embodiment, by introducing the insurance-related corpus, the large language model is fine-tuned for the specific domain, enabling the trained insurance recommendation text generation model to accurately understand insurance terminology, improving the model's semantic understanding accuracy, enhancing the professionalism of the initial insurance recommendation text generated by the model, and reducing the error rate in interpreting insurance clauses.
[0071] ③ The insurance recommendation texts of some historical insured users output by the insurance recommendation text generation model are labeled and sorted according to their matching degree to generate insurance recommendation text preference data, so as to train the text encoder to obtain a reward model for evaluating each insurance recommendation text output by the insurance recommendation text generation model.
[0072] In this embodiment, by designing a reward signal in the reward model, the optimization direction of the insurance recommendation text generation model can be accurately guided, enabling the model to generate answers that are better than the training data, thereby improving the creativity, accuracy, and reliability of the generated insurance recommendation text.
[0073] ④ The insurance recommendation text generation model is iteratively optimized through reinforcement learning algorithms, so that the insurance recommendation text generation model generates insurance recommendation text with high rewards.
[0074] In one embodiment, such as Figure 4 As shown, the method for generating initial insurance recommendation text based on the insurance recommendation text generation model mainly includes the following steps.
[0075] Step S31: Based on the user behavior data, create an in-depth customer profile for the target user.
[0076] In one embodiment, the method of creating a deep customer profile for a target user based on the user behavior data includes: extracting attribute features of the target user based on the user behavior data, including the target user's age, city, occupation, marital and childbearing status, asset and liability status, and health status; obtaining the policy status, coverage, exclusions, coverage range, and coverage period of the target user's current insurance policies, and calculating the target user's coverage gap to obtain the target user's coverage gap features; identifying one or more behavioral events and lifecycle events of the target user based on the user behavior data; deeply mining the target user's behavioral patterns based on the user behavior data, and reasoning and analyzing the target user's needs and motivations, thereby extracting features such as the target user's emotions, interests, consumption and risk preferences, and values; and labeling each extracted feature to generate a deep customer profile of the target user.
[0077] Therefore, the in-depth customer profile includes: attribute tags, event tags, and psychological and need tags. Specifically, the attribute tags include, but are not limited to, one or more combinations of: age tags, city tags, occupation tags, marital and childbearing tags, asset and liability tags, health tags, and protection gap tags; wherein, the health tags are used to describe the target user's health status, current illnesses, and potential disease risks; the protection gap tags are used to describe the target user's unused insurance products, missing coverage, and expiring policies, etc. The event tags include, but are not limited to: behavioral event tags (such as recently browsing "newborn care," recently searching / consulting about a certain insurance product, recently undergoing a physical examination, etc.) and / or life cycle event tags. The psychological and needs tags include, but are not limited to: emotional tags (such as anxiety about one's own health or parents' retirement security, or distrust of insurance products and aversion to complex insurance terms), interest and preference tags (such as liking travel or extreme sports), consumption and risk preference tags (such as impulsive, easily influenced by current emotions or situations in making decisions; conservative, preferring savings and certainty; or pragmatic, focusing on cost-effectiveness), and value tags (such as family responsibility, concerned about parents' retirement and children's education; or freedom, valuing life experiences). For example, the in-depth customer profile might be: male, 35 years old, married with children, has a mortgage, and has repeatedly searched for "children's education."
[0078] Step S32: Based on the pre-built insurance knowledge graph, perform multi-dimensional matching retrieval and cross-entity association operations according to the target user's in-depth customer profile and the target risk event, and filter one or more candidate recommendable insurance products that are suitable for the target user.
[0079] The insurance knowledge graph includes: multiple insurance product entity nodes, multiple deep customer profile entity nodes, multiple risk event entity nodes, multiple recommended text template entity nodes, and relational edges connecting the entity nodes. The attributes of the insurance product entity nodes include: the type of insurance to which the insurance product belongs, the insured object, coverage, and exclusions of the insurance product; the attributes of the deep customer profile entity nodes include: multiple feature tags describing the deep customer profile; the attributes of the risk event entity nodes include: the theme, object, time, location, and risk characteristics of the risk event, as well as one or more related insurance types.
[0080] In one embodiment, the insurance knowledge graph is constructed by the following steps.
[0081] ① Obtain user behavior data from multiple historical policyholders, and based on each user behavior data, create in-depth customer profiles for each historical policyholder to build a customer profile database.
[0082] It should be noted that the method for creating in-depth customer profiles for each historical insured user is the same as the method for creating in-depth customer profiles for target users, and will not be repeated here for the sake of brevity.
[0083] ② Define multiple risk events and one or more associated insurance types for each risk event to build a risk scenario library.
[0084] Specifically, the types of definable risk events include: customer lifecycle events, behavioral events, and news events. The defined types of lifecycle events include: birth of a new family member, family member starting school, reaching adulthood, becoming independent, first employment, marriage, home purchase, significant debt, birth of a child, career advancement, income increase, middle age, family maturity, retirement, illness, accident, death, and total disability. The defined types of behavioral events include, but are not limited to: purchasing airline tickets, medical checkups, seeking medical treatment, travel, checking insurance, and policy expiration. The defined types of news events include: epidemics, policies and regulations, natural disasters, and accidents.
[0085] It should be noted that the risk events and their associated insurance types defined in the insurance knowledge graph are consistent with the associated insurance types of hot news texts of each risk type marked and identified by the insurance risk event identification model in step S1, as well as the life cycle events and behavioral events that can be identified based on user behavior data in step S2.
[0086] ③ Obtain insurance product data, and based on the insurance product data, obtain multiple insurance products for multiple insurance types, as well as the insured objects, insured responsibilities, exclusion clauses and recommended text templates for each insurance product, and establish an insurance type classification tree.
[0087] In one embodiment, insurance product brochures, insurance contracts, insurance terms and actuarial data can be obtained from an external insurance product database to obtain the insurance product data, thereby obtaining multiple insurance products for multiple insurance types, as well as the insured objects, coverage responsibilities, exclusion clauses and recommended text templates for each insurance product.
[0088] In this embodiment, the established insurance product classification tree clearly defines each type of insurance, as well as multiple insurance products to be recommended under each type of insurance. It also clearly defines the insured objects, coverage responsibilities, exclusion clauses, and recommendation text templates for each insurance product, as well as the relationships between each insurance product.
[0089] The types of insurance include, but are not limited to, one type of personal insurance, such as critical illness insurance, medical insurance, accident insurance, life insurance, annuity insurance (including retirement annuity insurance, education annuity insurance, and term annuity insurance), and one type of increasing whole life insurance, thereby enabling the insurance recommendation text generation model to recommend multiple insurance products belonging to the corresponding insurance type. In other embodiments, the types of insurance may also include one type of property insurance, such as one type of auto insurance, household property insurance, liability insurance, and corporate property insurance, in which case the insurance recommendation text generation model can recommend multiple insurance products under the corresponding insurance type. The types of insurance are not specifically limited in this application.
[0090] ④ Define the entity types of the insurance knowledge graph as insurance products, in-depth customer profiles, risk events, and recommended text templates, and instantiate each insurance product, each in-depth customer profile, each risk event, and each recommended text template as multiple entity nodes, infer the relationships between each entity node, and construct the insurance knowledge graph.
[0091] Specifically, based on the insured objects, insured liabilities, and exclusion clauses of each insurance product, the relationship between each insurance product and each in-depth customer profile can be inferred and established; based on the insurance types to which each insurance product belongs and the related insurance types of each risk event, the relationship between each insurance product and each risk event can be inferred and established; and based on the established relationships, in-depth reasoning can be continued to connect the entities and construct the insurance knowledge graph.
[0092] In this embodiment, a multi-dimensional insurance knowledge system of "customer-product-scenario" is constructed through the entity nodes in the insurance knowledge graph and the relational edges connecting these nodes. The customer dimension integrates structured data (age / occupation / policy details) and unstructured data (consultation records / risk preferences) from multiple historical insured users. The product dimension defines an insurance product classification tree, linking the insured objects, coverage responsibilities, exclusions, and recommendation text templates of multiple insurance products. The scenario dimension covers standardized lifecycle events, behavioral events, and different types of news events, allowing real-time hot news or user behavior to be linked to insurance products to support scenario-based insurance recommendation triggering logic. Furthermore, a real-time hot news analysis pipeline is deployed, giving the constructed multi-dimensional insurance knowledge system a hot topic dimension.
[0093] Therefore, based on the insurance knowledge graph, multi-dimensional matching and retrieval operations are performed according to the target user's deep customer profile and the target risk event to filter out recommendable insurance types suitable for the current scenario-based insurance recommendation task and meeting the target user's real needs, as well as one or more candidate recommendable insurance products belonging to the recommendable insurance types. Simultaneously, cross-entity association operations can be performed to deeply infer and mine more candidate recommendable insurance products belonging to the recommendable insurance types that meet the target user's real needs. It should be noted that the recommendable insurance types are consistent with the associated insurance types of the target risk event determined in step S2, to achieve deep integration of scenario-based insurance recommendations, such as incorporating descriptive information of news events, lifecycle events, or behavioral events that serve as the target risk event into the insurance recommendation text.
[0094] For example, based on a deep customer profile of the target user, if the target user's occupation is a programmer, then a cross-entity association operation is performed based on the insurance knowledge graph, associating it with the fact that programmers usually sit for long periods of time and have high work intensity, and may have the risk of cervical spondylosis and critical illness, thus associating it with medical insurance and critical illness insurance; if the target user is pregnant, then it is associated with the fact that the target user may need prenatal check-ups and medical treatment and education after the child is born, thus associating it with medical insurance and education annuity insurance.
[0095] In one embodiment, after obtaining each candidate recommended insurance product, it is necessary to perform duplicate protection liability screening based on the target user's protection gap label, filter out candidate recommended insurance products that overlap with the target user's already deployed protection liability, and obtain one or more candidate recommended insurance products that truly meet the target user's needs. This avoids the technical problem of poor recommendation accuracy and impact on user experience caused by repeatedly recommending existing insurance products to the target user, and further ensures accurate insurance recommendations.
[0096] In one embodiment, the insurance knowledge graph can be updated periodically to improve the customer profile database, risk scenario database, and insurance product classification tree, so that the insurance knowledge graph contains the latest multi-dimensional insurance knowledge system.
[0097] Step S33: Evaluate the relevance of each candidate recommended insurance product to the needs of the target user, and determine the candidate recommended insurance product with the highest relevance to the needs as the recommended insurance product.
[0098] In one embodiment, when selecting candidate recommended insurance products based on the insurance knowledge graph, the insurance recommendation text generation model can perform semantic understanding and analysis on the target user's in-depth customer profile and each candidate recommended insurance product, and simultaneously output the confidence score of each candidate recommended insurance product as a score of the relevance between each candidate recommended insurance product and the target user's needs, thereby evaluating the relevance between each candidate recommended insurance product and the target user's needs based on each confidence score.
[0099] In another embodiment, when constructing the insurance knowledge graph, the relationship edges connecting each entity node can be set with a correlation score based on the degree of association between the two entity nodes. Based on the correlation scores on the retrieval path of each candidate recommendable insurance product, the demand correlation of each candidate recommendable insurance product can be calculated. For example, "Education annuity insurance demand correlation 92% > Critical illness insurance demand correlation 85% > Retirement annuity insurance demand correlation 68%", thereby obtaining the recommendable insurance products with the highest demand correlation.
[0100] Step S34: Based on the insurance knowledge graph, obtain the recommendation text template of the recommendable insurance type, and fill in the recommendation text template to generate the basic insurance recommendation text of the recommendable insurance type.
[0101] In one embodiment, the recommended text template includes at least: the target user's name, the insurance product name, and one or more scenario keywords. The scenario keywords are obtained through the target risk event.
[0102] Step S35: Based on the in-depth customer profile of the target user, select an appropriate insurance recommendation text generation strategy, and based on the insurance recommendation text generation strategy, construct a contextualized narrative structure, embed the target risk event into the storyline, and inject emotional expression to optimize the basic insurance recommendation text and generate the initial insurance recommendation text.
[0103] The insurance recommendation text generation strategy includes at least: narrative templates, story maps, emotional temperature, language style, and rhetoric library.
[0104] In one embodiment, the insurance recommendation text generation strategy can construct a contextualized narrative structure based on corresponding narrative templates and story maps, embedding the target risk event and improving the storyline. Simultaneously, the temperature parameter of the insurance recommendation text generation model can be adjusted to amplify or reduce the weight of preset high-probability words, controlling the emotional temperature and language style of the generated initial insurance recommendation text. This allows the initial insurance recommendation text to be diverse, enabling the model to generate different versions of the initial insurance recommendation text for different target users. For example, for impulsive and responsible target users, a warm version of the initial insurance recommendation text is generated, emphasizing emotional resonance and family care; for conservative target users who distrust insurance products, a data-driven version of the initial insurance recommendation text is generated, focusing on data analysis and clause interpretation, balancing professionalism and readability; for anxious target users, the generated initial insurance recommendation text incorporates emotionally reassuring language.
[0105] In this embodiment, based on the insurance recommendation text generation model, the customer characteristics of the target user can be dynamically analyzed according to user behavior data, including the target user's psychological and demand characteristics, thereby drawing a comprehensive and three-dimensional in-depth customer profile. Combined with a multi-dimensional insurance knowledge system (i.e., insurance knowledge graph), it can accurately output recommended insurance products that meet the target user's real needs. It can also integrate social hot news and user behavior to generate personalized initial insurance recommendation text that matches the target user's emotional stage, consumption and risk preferences, and values, and matches the current recommendation scenario, thereby improving the recommendation accuracy of the initial insurance recommendation text.
[0106] Step S4: Based on the pre-built insurance terms and rules base, conduct a compliance review on the initial insurance recommendation text, correct any violations, and send the obtained target insurance recommendation text to the target user terminal through the communication interface.
[0107] The insurance clause rule base contains a variety of insurance clause rules. In one specific embodiment, the insurance clause rule base contains 823 insurance clause rules, used for compliance review of the generated initial insurance recommendation text. Specifically, the compliance review of the initial insurance recommendation text includes: real-time detection of illegal content, including real-time detection of exaggerated returns (such as the statement "100% return guarantee"), concealed disclaimers (such as the statement "zero risk"), and misinterpretation of clauses, and automatic correction of illegal content, such as replacing "absolutely safe" with "risk level rating AA", thereby avoiding misinterpretation of clauses and obtaining a more professional target insurance recommendation text, ensuring the target user's trust in the target insurance recommendation text.
[0108] In one embodiment, the insurance recommendation method based on a large language model and multi-source data further includes: monitoring the click-through rate and conversion rate of the target insurance recommendation text and sending them to the insurance recommendation text generation model; optimizing the insurance recommendation text generation model to form a closed-loop optimization mechanism of "data-driven decision-decision verification effect-effect feedback data", which effectively improves the recommendation accuracy of the final target insurance recommendation text and improves the user experience.
[0109] In one specific embodiment, optimizing the insurance recommendation text generation model includes optimizing the insurance recommendation text generation strategy to improve the suitability of the initial insurance recommendation text with the target user, thereby increasing the click-through rate and conversion rate of the target insurance recommendation text.
[0110] In another specific embodiment, the correlation score of each relation edge in the insurance knowledge graph can be periodically optimized based on the click-through rate and conversion rate of the target insurance recommendation text, so as to improve the accuracy of selecting recommended insurance products based on the insurance knowledge graph, thereby improving the recommendation accuracy of the final target insurance recommendation text, that is, improving the click-through rate and conversion rate of the target insurance recommendation text.
[0111] like Figure 5 The diagram illustrates the structure of an insurance recommendation system 500 based on a large language model and multi-source data, as described in this embodiment. The insurance recommendation system 500 based on a large language model and multi-source data in this embodiment mainly includes: a news event identification module 501, a risk event determination module 502, a text generation module 503, a text compliance review module 504, and a text push module 505, connected sequentially.
[0112] Specifically, the news event identification module 501 is used to periodically acquire multiple hot news texts and, based on a pre-trained insurance risk event identification model, identify insurance risks in each hot news text to obtain one or more news events with insurance risks.
[0113] The risk event determination module 502 is used to acquire user behavior data of the target user, and identify lifecycle events and / or behavioral events of the target user based on the user behavior data, so as to map the lifecycle events, the behavioral events and various news events to the same risk event space to obtain the target risk event.
[0114] The text generation module 503 is used to draw a deep customer profile of the target user based on the user behavior data, and combine the target risk event to perform multi-dimensional matching and retrieval in the pre-built insurance knowledge graph, filter the recommended insurance types and recommended text templates that are suitable for the target user, and then generate the initial insurance recommendation text.
[0115] The text compliance review module 504 is used to conduct compliance review on the initial insurance recommendation text based on a pre-built insurance clause rule base, and to correct any violations to obtain the target insurance recommendation text.
[0116] The text push module 505 is used to send the target insurance recommendation text to the target user terminal through a communication interface.
[0117] In one embodiment, the insurance recommendation system 500 based on a large language model and multi-source data further includes a text click monitoring module. The text click monitoring module is connected to the text generation module 503 and is used to monitor the click-through rate and conversion rate of the target insurance recommendation text, and optimize the insurance recommendation text generation model accordingly.
[0118] It should be understood that the insurance recommendation system based on large language models and multi-source data and the insurance recommendation method based on large language models and multi-source data belong to the same inventive concept. The implementation methods of each module have been described in detail in the above method embodiments, and will not be repeated here for the sake of brevity.
[0119] It should also be understood that the module division in the embodiments of this application is illustrative and only represents a logical functional division; in actual implementation, there may be other division methods. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor or functional module, or they can exist as separate physical entities, or they can be divided into more functional modules. The integrated modules or units described above can be implemented in hardware or as software functional modules.
[0120] The insurance recommendation method based on large language models and multi-source data provided in this invention can be implemented on the terminal side or the server side. For the hardware structure of the insurance recommendation terminal based on large language models and multi-source data, please refer to [link to relevant documentation]. Figure 6 This is a schematic diagram of an optional hardware structure of an insurance recommendation terminal 600 based on a large language model and multi-source data provided in an embodiment of the present invention. The insurance recommendation terminal 600 based on a large language model and multi-source data can be a mobile phone, computer device, tablet device, personal digital processing device, factory back-end processing device, etc. The insurance recommendation terminal 600 based on a large language model and multi-source data includes: at least one processor 601, a memory 602, at least one network interface 604, and a user interface 606. The various components in the device are coupled together through a bus system 605. It is understood that the bus system 605 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 605 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 6 The general will label all buses as bus systems.
[0121] The user interface 606 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.
[0122] It is understood that memory 602 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.
[0123] In this embodiment of the invention, the memory 602 is used to store various categories of data to support the operation of the insurance recommendation terminal 600 based on a large language model and multi-source data. Examples of this data include: any executable program for operating on the insurance recommendation terminal 600 based on a large language model and multi-source data, such as operating system 6021 and application program 6022; operating system 6021 includes various system programs, such as framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. Application program 6022 may include various applications, such as media player, browser, etc., for implementing various application services. The insurance recommendation method based on a large language model and multi-source data provided in this embodiment of the invention can be included in application program 6022.
[0124] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 601. Processor 601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 601 or by instructions in software form. The processor 601 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 601 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. General-purpose processor 601 may be a microprocessor or any conventional processor, etc. The steps of the insurance recommendation method based on large language models and multi-source data provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.
[0125] In an exemplary embodiment, the insurance recommendation terminal 600 based on a large language model and multi-source data can be used by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs) to execute the aforementioned insurance recommendation method based on a large language model and multi-source data.
[0126] In summary, this application provides an insurance recommendation method, system, and terminal based on a large language model and multi-source data. It identifies insurance risks from multiple periodically acquired trending news texts to obtain one or more news events with insurance risks. Combined with target user lifecycle events and / or behavioral events identified based on user behavior data, it filters to obtain target risk events to trigger scenario-based insurance recommendation tasks. Simultaneously, based on a pre-built insurance recommendation text generation model and an insurance knowledge graph, and according to the target user's in-depth customer profile and target risk events, it generates personalized initial insurance recommendation text that meets the target user's actual needs. After the initial insurance recommendation text passes compliance review, the obtained more professional target insurance recommendation text is sent to the target user's terminal, achieving precise reach.
[0127] Therefore, this application effectively overcomes the various shortcomings of the prior art and has high industrial application value.
[0128] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. An insurance recommendation method based on a large language model and multi-source data, characterized in that, include: Regularly acquire multiple trending news texts, and based on a pre-trained insurance risk event identification model, identify insurance risks in each trending news text to obtain one or more news events with insurance risks. Acquire user behavior data of the target user, and based on the user behavior data, identify the target user's lifecycle events and / or behavioral events, so as to map the lifecycle events, the behavioral events and various news events to the same risk event space to obtain the target risk event; Based on the pre-built insurance recommendation text generation model, a deep customer profile is drawn for the target user according to the user behavior data, and combined with the target risk event, multi-dimensional matching and retrieval are performed in the pre-built insurance knowledge graph to filter the insurance types and recommendation text templates that are suitable for the target user, and then generate the initial insurance recommendation text. Based on a pre-built insurance terms and rules library, the initial insurance recommendation text undergoes compliance review, and any violations are corrected. The resulting target insurance recommendation text is then sent to the target user terminal via a communication interface.
2. The insurance recommendation method based on a large language model and multi-source data according to claim 1, characterized in that, Based on a pre-trained insurance risk event identification model, insurance risk identification can be performed on various trending news texts to obtain one or more news events with insurance risks. The methods include: Perform natural language processing on each trending news text to obtain one or more news keywords and news description data for each trending news text; Based on each news keyword, identify one or more candidate news texts related to insurance risks, and mark the associated insurance types for each candidate news text; Evaluate the relevance of each candidate news text to its associated insurance product, eliminate candidate news texts with low relevance, and obtain one or more news texts. Obtain news description data from each news text, and extract news event elements from each news text based on the news description data to obtain multiple corresponding news events.
3. The insurance recommendation method based on a large language model and multi-source data according to claim 1, characterized in that, The methods for mapping the lifecycle events, behavioral events, and various news events to the same risk event space to obtain the target risk event include: The life cycle events, behavioral events, and news events are respectively regarded as risk events, and the event elements of each risk event are analyzed to extract the theme features, object features, time features, regional features, and risk features of each risk event to generate a risk event feature vector for each risk event. The feature vectors of each risk event are dimensionally aligned and semantically normalized to generate multiple standardized risk event representations in the same risk event space. Based on the user behavior data of the target users and the thematic characteristics, object characteristics, time characteristics, geographical characteristics and risk characteristics represented by each standardized risk event, the object proximity, time proximity, geographical proximity and risk proximity of each risk event to the target users are evaluated respectively, and the scenario correlation between each risk event and the target users is calculated. Based on the target user's user behavior data, obtain the policy status, coverage, exclusions, coverage range, and coverage period of the target user's current policies, and calculate the target user's protection gap in order to screen one or more candidate risk events that match the protection gap; The candidate risk event with the highest relevance to the target user's scenario is identified as the target risk event.
4. The insurance recommendation method based on a large language model and multi-source data according to claim 1, characterized in that, Based on the aforementioned insurance recommendation text generation model, the methods for generating initial insurance recommendation text include: Based on the user behavior data, create in-depth customer profiles for the target users; Based on a pre-built insurance knowledge graph, multi-dimensional matching retrieval and cross-entity association operations are performed according to the deep customer profile of the target user and the target risk event, respectively, to filter one or more candidate recommendable insurance products that are suitable for the target user. Each candidate recommended insurance product is evaluated for its relevance to the needs of the target users, and the candidate recommended insurance product with the highest relevance to needs is identified as the recommended insurance product. Based on the insurance knowledge graph, a recommendation text template for the recommendable insurance products is obtained, and the recommendation text template is filled in to generate the basic insurance recommendation text for the recommendable insurance products. Based on the in-depth customer profile of the target user, an appropriate insurance recommendation text generation strategy is selected. Based on the insurance recommendation text generation strategy, a contextualized narrative structure is constructed, and the target risk event is embedded into the storyline and injected with emotional expression to optimize the basic insurance recommendation text and generate the initial insurance recommendation text. The insurance recommendation text generation strategy includes at least: narrative templates, story maps, emotional temperature, language style, and rhetoric library.
5. The insurance recommendation method based on a large language model and multi-source data according to claim 4, characterized in that, The in-depth customer profile includes: Attribute tags, including but not limited to: age tags, city tags, occupation tags, marital and childbearing tags, asset and liability tags, health tags, and one or more combinations of insurance gap tags; Event tags, including but not limited to: behavioral event tags and / or lifecycle event tags; Psychological and need labels, including but not limited to: one or more combinations of emotion labels, interest and preference labels, consumption and risk preference labels, and value labels.
6. The insurance recommendation method based on a large language model and multi-source data according to claim 4, characterized in that, The methods for constructing the insurance knowledge graph include: Acquire user behavior data from multiple historical policyholders, and based on each user behavior data, create in-depth customer profiles for each historical policyholder to build a customer profile database; Define multiple risk events and one or more associated insurance products for each risk event to build a risk scenario library; Acquire insurance product data, and based on the insurance product data, obtain multiple insurance products for multiple insurance types, as well as the insured objects, coverage responsibilities, exclusion clauses and recommended text templates for each insurance product, and establish an insurance type classification tree; The entity types of the insurance knowledge graph are defined as insurance products, in-depth customer profiles, risk events, and recommended text templates. Each insurance product, in-depth customer profile, risk event, and recommended text template is instantiated as multiple entity nodes. The relationships between the entity nodes are inferred to construct the insurance knowledge graph.
7. The insurance recommendation method based on a large language model and multi-source data according to claim 1, characterized in that, The construction method of the insurance recommendation text generation model includes: Collect user behavior data from multiple historical insured users and multiple historical news events to construct a training sample dataset; The training sample dataset is input into the large language model, and professional corpus in the insurance field is introduced. The large language model is trained using a self-supervised learning method to obtain an insurance recommendation text generation model. The matching degree of insurance recommendation texts of some historical insured users output by the insurance recommendation text generation model is marked and sorted to generate insurance recommendation text preference data, so as to train the text encoder to obtain a reward model for evaluating each insurance recommendation text output by the insurance recommendation text generation model; The insurance recommendation text generation model is iteratively optimized using a reinforcement learning algorithm, enabling it to generate insurance recommendation texts with high rewards.
8. The insurance recommendation method based on a large language model and multi-source data according to claim 7, characterized in that, Also includes: Monitor the click-through rate and conversion rate of the target insurance recommendation text, and optimize the insurance recommendation text generation model accordingly.
9. An insurance recommendation system based on a large language model and multi-source data, characterized in that, include: The news event identification module is used to periodically acquire multiple trending news texts and, based on a pre-trained insurance risk event identification model, identify insurance risks in each trending news text to obtain one or more news events with insurance risks. The risk event determination module is connected to the news event identification module. It is used to acquire user behavior data of the target user and identify lifecycle events and / or behavioral events of the target user based on the user behavior data, so as to map the lifecycle events, the behavioral events and each news event to the same risk event space to obtain the target risk event. The text generation module, connected to the risk event determination module, is used to draw a deep customer profile of the target user based on the pre-built insurance recommendation text generation model and the user behavior data, and to perform multi-dimensional matching and retrieval in the pre-built insurance knowledge graph in combination with the target risk event, to filter the recommended insurance types and recommendation text templates that are suitable for the target user, and then generate the initial insurance recommendation text. The text compliance review module, connected to the text generation module, is used to conduct compliance review on the initial insurance recommendation text based on a pre-built insurance clause rule base, and correct any violations to obtain the target insurance recommendation text; The text push module, connected to the text compliance review module, is used to send the target insurance recommendation text to the target user terminal through a communication interface.
10. An insurance recommendation terminal based on a large language model and multi-source data, characterized in that, Memory and processor; The memory is used to store computer programs; The processor is configured to execute the computer program stored in the memory, so that the terminal performs the insurance recommendation method based on a large language model and multi-source data as described in any one of claims 1 to 8.