An insurance short video content generation method, device, medium and program product

CN121711545BActive Publication Date: 2026-09-15BEIJING ZHIBAO HUIZHONG DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511900430.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-09-15
Estimated Expiration
2045-12-16

AI Technical Summary

Technical Problem

相关技术中,模板的通用性往往难以满足专业性和时效性的要求,在合规标准和专业术语方面可能生成语义模糊、信息不全以及存在合规风险的内容

Benefits of technology

[0015]通过采用上述技术方案,从目标保险产品对应的承保责任子图中,通过语义匹配算法计算各保障条款与目标核心风险要素的语义相似度,同时通过重要性排序算法基于缺口责任项目清单确定各条款的重要性权重,双重筛选机制确保了核心保障亮点既与热点事件高度相关又能有效填补用户保障缺口。根据用户子群体的推荐流程阶段标签从预设策略库中确定策略模板,并基于目标核心风险要素中公众关注度指数的数值区间动态调整策略模板中的情感倾向参数和紧急程度参数,使提示词策略在内容侧重和表达风格上与热点关注度相适应。将核心保障亮点、调整后的策略模板以及目标保险产品的产品特征进行整合,提升了生成内容的关联度和准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121711545B_ABST
    Figure CN121711545B_ABST
Patent Text Reader

Abstract

An insurance short video content generation method, device, medium and program product, relating to the technical field of artificial intelligence. The core risk elements are extracted by semantic analysis of hot event data, so that the system can capture the risk characteristics of the current social concern. The core risk elements are matched with the protection needs of potential users and the protection gap is determined, establishing the correlation mapping between the hot event risk and the protection needs, so that the determination of the target insurance product has a risk-oriented. The prompt word generation strategy is determined from the preset strategy library, and the content draft is generated by inputting the insurance field large language model, realizing the automatic generation of content. Combined with the pre-established compliance detection mechanism, the compliance of the generated content is ensured. The rapid response and generation of insurance short video content are realized, and the efficiency of content generation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, device, medium, and program product for generating short insurance video content. Background Technology

[0002] With the popularization of digital media, short videos have become an important medium for the insurance industry to conduct marketing and reach users.

[0003] In related technologies, a short video generation system based on a material library and templates is typically employed. This system pre-builds a library containing various high-quality video clips, images, text descriptions, and other materials, and designs multiple preset templates suitable for different marketing scenarios. When a new short video needs to be created, the system retrieves relevant materials from the material library based on current social trends or marketing themes, automatically populates them into the selected template, and quickly generates short video content through editing and combination.

[0004] However, in highly specialized fields like insurance, when creating content targeting specific trending events, niche user groups, or complex insurance products, the content must not only accurately convey product information but also strictly adhere to all relevant laws, regulations, and compliance requirements. In related technologies, the generality of templates often fails to meet the demands of professionalism and timeliness, potentially generating content that is semantically ambiguous, incomplete, or carries compliance risks regarding compliance standards and terminology. This can lead to significant manual review and revisions, thereby reducing the efficiency of generating insurance-related short videos. Summary of the Invention

[0005] This application provides a method, device, medium, and program product for generating insurance short video content, which can improve the timeliness and relevance of short video content to trending topics.

[0006] Firstly, this application provides a method for generating short insurance video content. The method includes: semantically parsing hot topic event data within a preset time period to obtain the core risk elements of the corresponding hot topic events. The hot topic event data includes multiple hot topic events, and the core risk elements include risk type identifiers, impact range parameters, and public attention indexes. If it is determined that the public attention index of a detected target hot topic event exceeds a preset heat value threshold, and the core risk elements of the target hot topic event match the preset protection needs of a potential user, then the potential user's policy data is obtained and a protection gap is determined. The protection gap includes the potential user's existing insurance coverage under the risk type corresponding to the risk type identifier. The system generates a list of liability items that fill the gap between the single coverage limit and the recommended coverage limit calculated based on user profile data, along with a coverage gap score for the potential user. Based on this coverage gap, it identifies target insurance products that match the core risk elements of the target hot topic. According to the product characteristics of the target insurance product and the core risk elements, it determines a corresponding prompt word generation strategy for the potential user from a pre-set prompt word strategy library. The prompt words generated based on this strategy are input into a pre-trained insurance domain large language model to obtain a draft of the creative content. If the draft of the creative content passes compliance checks, a target recommendation video is generated based on the draft.

[0007] By employing the aforementioned technical solutions, semantic analysis is performed on hot topic event data to extract core risk elements, enabling the system to capture risk characteristics currently of public concern. These core risk elements are matched with the protection needs of potential users to identify protection gaps, establishing a correlation mapping between hot topic event risks and protection needs, thus making the selection of target insurance products risk-oriented. A prompt word generation strategy is determined from a pre-set strategy library and input into a large-scale language model in the insurance field to generate an initial content draft, achieving automated content generation. Combined with a pre-implemented compliance detection mechanism, the compliance of the generated content is ensured. This enables rapid response and generation of insurance short video content, improving the efficiency of content creation.

[0008] In conjunction with some embodiments of the first aspect, in some embodiments, the step of determining a target insurance product that matches the target core risk element of the target hot topic event based on the potential user's protection gap specifically includes: using the target core risk element as the starting node set in a preset insurance product knowledge graph, and through a graph traversal algorithm, determining nodes whose matching degree with the protection gap exceeds a preset threshold as candidate insurance product nodes; performing language complexity analysis on the product terms text corresponding to the candidate insurance product nodes through natural language processing to calculate the terms' comprehensibility score; determining the list of underwriting liability items and coverage parameters from the structured product data of the candidate insurance products, and comparing the list of underwriting liability items with the list of liability items for the gap to calculate the underwriting liability items. The coverage ratio of the list of items covering the gap in liability is used as the underwriting matching score; a first set of keywords is determined from the descriptive text corresponding to the core risk element using a keyword extraction algorithm, and a second set of keywords is determined from the product introduction text and terms text of the candidate insurance product using the same keyword extraction algorithm; the ratio of the number of intersection elements to the number of union elements of the first and second keyword sets is calculated to obtain the event-product relevance score; the terms comprehensibility score, the underwriting matching score, and the event-product relevance score are weighted according to preset weights to obtain the comprehensive matching score of the candidate insurance product; based on the comprehensive matching score, the insurance product with the highest comprehensive matching score is selected from the candidate insurance product nodes as the target insurance product.

[0009] By employing the aforementioned technical solution, a graph traversal is performed within a pre-defined insurance product knowledge graph, starting with the core risk elements of the target product. This enables structured retrieval from the risk dimension to the product dimension, initially identifying candidate product nodes related to coverage gaps. The linguistic complexity of the product terms is quantitatively analyzed through a terms-understanding score; the coverage matching score measures the degree to which product liability meets the user's coverage gap; and the event-product relevance score assesses the semantic connection between the product and trending events. These three dimensions characterize the product's suitability from different perspectives. A weighted average of these three scores yields a comprehensive matching score, ensuring that product selection simultaneously considers terms-understanding, coverage, and event relevance. This resolves the product matching bias caused by single-dimensional evaluation, improving the accuracy and comprehensiveness of identifying target insurance products.

[0010] In conjunction with some embodiments of the first aspect, in some embodiments, before the step of determining, through graph traversal algorithm, nodes with a matching degree exceeding a preset threshold for the protection gap as candidate insurance product nodes in a preset insurance product knowledge graph, starting with the target core risk element, the method further includes: performing natural language processing on the descriptive text of the target core risk element to determine key risk entities and risk attributes, wherein the key risk entities include event type, affected region, and affected population, and the risk attributes include risk level and duration; based on the user's policy protection gap score and the list of liability items for the gap, combined with preset risk protection demand mapping rules, converting the policy protection gap score into query conditions for insurance product underwriting liability nodes and protection amount nodes; and generating a structured query instruction for the preset insurance product knowledge graph using the key risk entities, the risk attributes, and the query conditions through a preset knowledge graph query language template.

[0011] By employing the aforementioned technical solutions, natural language processing is used to extract key risk entities and risk attributes from the descriptive text of the target core risk elements, achieving the transformation from unstructured text to structured risk features. Based on the user's policy coverage gap score and gap liability item list, combined with preset risk protection demand mapping rules, the user's coverage gap is converted into specific query conditions targeting the underwriting liability nodes and coverage amount nodes of insurance products, establishing a mapping relationship between user needs and graph queries. The above information is integrated into structured query instructions through preset knowledge graph query language templates, improving the accuracy and execution efficiency of knowledge graph retrieval.

[0012] In conjunction with some embodiments of the first aspect, in some embodiments, after determining the target insurance product matching the target core risk element of the target hot event based on the protection gap of the potential user, the method further includes: obtaining the access trajectory sequence of each user among the potential users for the target insurance product and similar products within a preset time window, wherein the similar products are determined based on a preset insurance product knowledge graph, and the access trajectory sequence includes a time sequence record of the user's access time, browsing behavior, and operation behavior on the page corresponding to the target insurance product; calculating the interaction depth feature value of the corresponding user among the potential users by weighting and summing the page dwell time, page scrolling depth, and page element click frequency in the access trajectory sequence according to a preset weight; dividing the potential user into multiple user subgroups according to the interaction depth feature value, and assigning corresponding recommendation process stage labels to the multiple user subgroups, wherein the recommendation process stage labels include the recommendation content generation focus features and form templates for the risk characteristics of the target hot event, which are used to guide the determination of subsequent prompt word generation strategies.

[0013] By employing the aforementioned technical solution, the system acquires the access trajectory sequence of potential users to the target insurance product and similar products within a preset time window, capturing complete user interaction data on the product page. By weighting and summing three indicators—page dwell time, page scrolling depth, and page element click frequency—according to preset weights, the multi-dimensional behavioral data is converted into a single interaction depth feature value, achieving a quantitative representation of user interaction levels. Based on this feature value, potential users are divided into multiple user subgroups and assigned corresponding recommendation process stage labels, enabling the system to identify differences in the decision-making stages of different users. These recommendation process stage labels include recommendation content generation focus features and format templates tailored to the risk characteristics of target hot events, improving the targeting and adaptability of content generation.

[0014] In conjunction with some embodiments of the first aspect, in some embodiments, the step of determining the prompt word generation strategy corresponding to the potential user from a preset prompt word strategy library based on the product characteristics of the target insurance product and the target core risk element specifically includes: calculating the semantic similarity between each protection clause and the target core risk element from the underwriting liability sub-graph corresponding to the target insurance product, based on the target core risk element and the gap liability item list, using a semantic matching algorithm, and determining the importance weight of the protection clause based on the gap liability item list using an importance ranking algorithm; taking the protection clause with a semantic similarity exceeding a preset semantic threshold and the importance weight exceeding a preset weight threshold as the core protection highlight, the core protection highlight including the protection description of specific risks in the insurance product terms; determining a strategy template from the preset prompt word strategy library based on the recommendation process stage tags of the user subgroup, the strategy template defining the structure, length, and information focus of the prompt words; determining the sentiment tendency parameter and urgency parameter in the strategy template based on the numerical range of the public attention index in the target core risk element; and integrating the core protection highlight, the adjusted strategy template, and the product characteristics of the target insurance product to determine the prompt word generation strategy corresponding to the user subgroup.

[0015] By employing the aforementioned technical solution, a semantic matching algorithm is used to calculate the semantic similarity between each coverage clause and the target core risk element from the underwriting liability sub-graph corresponding to the target insurance product. Simultaneously, an importance ranking algorithm is used to determine the importance weight of each clause based on the gap liability item list. This dual screening mechanism ensures that the core coverage highlights are both highly relevant to trending events and effectively fill user protection gaps. Based on the recommendation process stage tags of user subgroups, strategy templates are determined from a pre-set strategy library. The sentiment and urgency parameters in the strategy templates are dynamically adjusted based on the numerical range of the public attention index in the target core risk element, ensuring that the prompt word strategy is adapted to the trending topic in terms of content emphasis and expression style. Integrating the core coverage highlights, the adjusted strategy templates, and the product characteristics of the target insurance product improves the relevance and accuracy of the generated content.

[0016] In conjunction with some embodiments of the first aspect, in some embodiments, the step of acquiring potential users' policy data and determining protection gaps specifically includes: based on the risk type identifier of the target hot event, extracting the coverage scope and corresponding existing policy coverage amount from the potential user's policy data; based on the potential user's user profile data, calculating the recommended coverage amount for the risk type identifier using a preset risk assessment model; comparing the preset risk type standard liability list corresponding to the risk type identifier with the coverage scope of the existing policy to generate a list of gap liability items; calculating the difference between the recommended coverage amount and the existing policy coverage amount, and combining the number and importance level of the gap liability item list, obtaining the protection gap score through a preset scoring function.

[0017] By employing the aforementioned technical solution, the coverage scope and corresponding coverage amount of existing policies are extracted from the policy data of potential users based on risk type identifiers of target hot events, thus clarifying the user's current protection baseline. Based on user profile data, a recommended coverage amount for this risk type identifier is calculated using a pre-set risk assessment model, establishing a quantitative relationship between individual user characteristics and protection needs. A pre-set standard liability list for this risk type identifier is compared with the coverage scope of existing policies to generate a gap liability item list, identifying protection deficiencies at the liability level. The difference between the recommended coverage amount and the existing policy coverage amount is calculated, and a protection gap score is obtained through a pre-set scoring function, combining the number and importance level of the gap liability item list. This achieves a comprehensive quantitative assessment of protection gaps from two dimensions: liability deficiencies and insufficient coverage, improving the accuracy and comprehensiveness of protection needs analysis.

[0018] In conjunction with some embodiments of the first aspect, in some embodiments, the step of generating a target recommended video based on the initial draft of the creative content, after determining that the initial draft of the creative content has passed the compliance test, specifically includes: after determining that the initial draft of the creative content has passed the compliance test, converting the text script in the initial draft of the creative content into audio material with a preset emotional tone; using virtual digital human technology to generate a lip-sync sequence, facial expression sequence, and body movement sequence of a virtual human from the phoneme time sequence and speech frequency features in the audio material, and synchronizing the sequence with the timeline of the audio material to output a virtual human audio-visual synchronized video clip; retrieving and matching corresponding multimedia materials from a preset material library according to the keywords in the initial draft of the creative content; and combining the audio material, the virtual human audio-visual synchronized video clip, and the multimedia material according to a preset video synthesis template to generate the target recommended video.

[0019] By adopting the above technical solution, after confirming that the initial draft of the creative content passes the compliance test, the text script is converted into audio material with a preset emotional tone, realizing the automated conversion from text to speech and endowing the content with emotional expression capabilities. The phoneme time series and speech frequency features in the audio material are used to generate lip-sync sequences, facial expression sequences, and body movement sequences for a virtual human using virtual digital human technology. These sequences are then synchronized with the timeline of the audio material to output a synchronized audio-visual video clip of the virtual human, achieving automated generation of visual presentation. Based on the keywords in the initial draft of the creative content, corresponding multimedia materials are retrieved and matched from a preset material library, enriching the visual elements of the video. Following a preset video synthesis template, the audio material, the synchronized audio-visual video clip of the virtual human, and the multimedia materials are combined to generate a target recommended video, solving the problems of manual reliance and low efficiency in existing video production technologies, and realizing a fully automated generation process from compliant text to complete video.

[0020] In a second aspect, this application provides a content generation device comprising: one or more processors and a memory; the memory is coupled to the one or more processors and is used to store computer program code including computer instructions, wherein the one or more processors invoke the computer instructions to cause the content generation device to perform the method described in the first aspect and any possible implementation thereof.

[0021] Thirdly, this application provides a computer program product containing instructions that, when run on a content generation device, cause the content generation device to perform the method described in the first aspect and any possible implementation thereof.

[0022] Fourthly, this application provides a computer-readable storage medium including instructions that, when executed on a content generating device, cause the content generating device to perform the method described in the first aspect and any possible implementation thereof. One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. By adopting the above technical solutions, and by using dynamic matching of core risk elements of hot events with potential user protection gaps, as well as technical means of generating content based on a large language model in the insurance field and conducting compliance testing, rapid response and compliance control of insurance short video content have been achieved, improving the overall efficiency and quality of content generation.

[0023] 2. By adopting the above technical solution, the accuracy and execution efficiency of knowledge graph retrieval are improved by using natural language processing to extract key risk entities and attributes from the core risk elements of the target, and by converting the user's protection gap into query conditions for the knowledge graph's underwriting responsibility nodes and coverage amount nodes, thereby generating structured query instructions. This provides a query basis for subsequent product matching.

[0024] 3. By adopting the above technical solution, the core protection highlights are selected from the target insurance product's underwriting liability sub-graph based on semantic matching and importance ranking. The sentiment tendency and urgency parameters in the strategy template are dynamically adjusted according to the recommendation process stage tags of user subgroups and the public attention index. Finally, the prompt word generation strategy is integrated to form a prompt word generation strategy. Therefore, the prompt word generation strategy is made intelligent, refined and personalized, and the relevance and accuracy of the generated content are improved. Attached Figure Description

[0025] Figure 1 This is a flowchart illustrating a method for generating insurance short video content in an embodiment of this application; Figure 2 This is another flowchart illustrating a method for generating insurance short video content in an embodiment of this application; Figure 3 This is an exemplary hardware structure diagram of a content generation device in the embodiments of this application. Detailed Implementation

[0026] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.

[0027] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0028] Please see Figure 1 This is a flowchart illustrating a method for generating insurance short video content in an embodiment of this application.

[0029] It should be noted that the collection and processing of user data involved in this application are all carried out with the user's authorization and in compliance with relevant data privacy protection laws and regulations.

[0030] S101. Obtain the core risk elements of the corresponding hot events by semantic parsing the hot event data within the preset time period.

[0031] Among them, hot topic event data refers to event information with high social attention collected from multiple sources such as news media, social platforms, and government announcements within this time period, including structured and unstructured data such as event title, description, time of occurrence, geographical location, and scope of impact.

[0032] This step is executed when the content generation device periodically scans various information sources and detects hot events with potential insurance relevance. During execution, the content generation device activates its built-in event acquisition interface to synchronously acquire raw data of hot events from multiple data sources (such as data crawlers or third-party data APIs), extracting hot information within a specified time period from multiple preset data sources. After acquiring the raw data, the device activates its semantic parsing engine, using named entity recognition technology to identify key entities in the event, such as location, time, people, and loss type. Then, it uses dependency parsing and semantic role labeling technology to analyze the relationships between entities in the event description text, thereby extracting the core risk-related elements. This includes text preprocessing, such as word segmentation and stop word removal. Named entity recognition technology identifies key entities in the text, such as event name (e.g., "Super Typhoon XX"), location (e.g., "Southeast Coastal Provinces"), organization (e.g., "An Airlines Company"), and people (e.g., "Outdoor Workers"). Through event extraction and relationship extraction models, the relationships between these entities are analyzed to construct a complete picture of the event. For example, the model might identify that a typhoon "makes landfall" in a "southern coastal province" and "affects" "outdoor workers" and "air transport." Based on this information, the device will categorize the event and generate standardized "risk type identifiers" using a pre-set risk classification system (which could be a tree structure or ontology). Simultaneously, the content generation device will classify and standardize the extracted elements based on a pre-trained insurance domain knowledge base, converting natural language descriptions into structured risk element data. It will also calculate a public attention index (such as a weighted scoring model) by combining multiple dimensions of indicators, including the event's popularity, media coverage frequency, and social media discussion volume. This forms a complete set of core risk elements.

[0033] In some embodiments, semantic parsing and extraction of core risk elements from hot topic event data can be achieved in various ways. Optionally, the device maintains a dynamically updated risk keyword dictionary, where each keyword (such as "earthquake" or "hacker attack") is pre-associated with a risk type identifier. Secondly, the device uses regular expressions and a geographic information dictionary to match and extract information such as place names and city clusters from the text as the scope of influence. The public attention index is quantified by statistically analyzing the frequency and growth rate of keywords on major social media platforms and combining this with a preset calculation formula. Optionally, the content generation device can also employ a knowledge graph-based semantic reasoning method. It pre-constructs an insurance domain knowledge graph containing various risk events, insurance products, geographical regions, and their relationships. When a new hot topic event occurs, graph embedding technology maps the event text to the vector space of the knowledge graph. Then, a graph neural network is used for multi-hop reasoning, gradually deriving core elements such as the relevant risk type, scope of influence, and potential losses from the event entity. For example, starting from the event entity "an earthquake occurred in a certain area," related property insurance risks, personal insurance risks, and loss estimates based on historical data can be inferred. It is understandable that other rule-based or hybrid intelligence-based methods can also be used to extract risk factors from hot events, and no specific method is specified here.

[0034] S102. If it is determined that the public attention index of the target hot event exceeds the preset heat value threshold, and the core risk elements of the target hot event match the preset protection needs of potential users, obtain the policy data of potential users and determine the protection gap.

[0035] Potential users refer to target customer groups with active behavior records on the insurance platform, whose user profile data shows a potential need for protection against specific risk types. Preset protection needs represent the user's predicted insurance protection inclination and demand intensity under various risk scenarios, based on dimensions such as historical behavior, geographical location, occupational characteristics, and family structure. Policy data refers to all insurance contract information currently held by the user, including complete policy lifecycle data such as insurance product type, coverage scope, coverage amount, effective date, and claims records.

[0036] Specifically, the content generation device first checks whether the public attention index of each trending event reaches a preset heat value threshold. This threshold is usually dynamically adjusted based on the platform's user base, industry characteristics, and marketing goals. For example, for an insurance platform with millions of users, the heat value threshold might be set to an attention index greater than 10,000. After passing the threshold screening, the content generation device intelligently matches the core risk elements of the event with the protection needs of potential users in the user database. It identifies user groups highly relevant to the event through semantic similarity calculations and risk type correspondence retrieval. After identifying target users, the content generation device conducts in-depth analysis of these users' existing insurance policy configurations, including information on various insurance products they hold, such as life insurance, property insurance, and liability insurance. It then calculates the user's ideal protection needs under the current event's risk type using a risk assessment model, compares this with the coverage and coverage amount of their existing policies, and identifies specific gaps in the user's risk protection, including missing coverage items and insufficient coverage amounts.

[0037] In one embodiment, based on the risk type identifier of the target hotspot event (such as "public health - acute infectious disease"), the device retrieves and matches information from the potential user's policy data to accurately extract the coverage scope related to this risk from all existing policies (e.g., whether it includes "infectious disease hospitalization allowance" or "infectious disease death benefit") and the corresponding existing policy coverage amount (e.g., allowance of 200 yuan / day, death benefit of 100,000 yuan). Next, the device retrieves the user's profile data (including age, occupation, income, family structure, and location) and inputs it into a preset risk assessment model. Based on the user's specific circumstances, the model calculates a reasonable recommended coverage amount suggested by an expert under the current risk type identifier (e.g., based on income and family responsibilities, the recommended death benefit for infectious disease should be 500,000 yuan). Subsequently, the device compares a pre-defined standard liability list (an ideal list of coverage items, such as "pay upon diagnosis," "hospitalization allowance," "ICU allowance," "death / total disability," etc.) corresponding to the risk type identifier ("Public Health - Acute Infectious Disease") with the coverage scope of the user's existing policy item by item, generating a gap liability item list that clearly lists the user's missing coverage items. The device calculates the difference between the calculated recommended coverage amount and the user's existing policy coverage amount to obtain the coverage gap, and combines the number of gap liability items on the list with the pre-defined importance level of each item (e.g., "pay upon diagnosis" is "high"), using a pre-defined scoring function to comprehensively calculate the coverage gap score.

[0038] In other embodiments, a protection demand prediction method based on user behavior analysis can be adopted. By analyzing users' browsing history, search keywords, and dwell time on insurance-related websites, a collaborative filtering algorithm is used to identify user groups with similar behavioral patterns. Then, based on the policy configuration of similar users, the potential protection needs of target users are inferred. Next, a risk propagation model is used to analyze the potential impact of hot events on different user groups. Combined with static attribute information such as users' geographical location, occupation characteristics, and lifestyle, a personalized protection gap score is calculated. This is not limited here.

[0039] S103. Based on the protection gaps of potential users, identify target insurance products that match the core risk elements of the target hot events.

[0040] Based on the risk characteristics of target trending events and users' protection gap information, the device locates relevant product nodes in a pre-defined insurance product knowledge graph. Using a graph traversal algorithm, it explores all possible paths from risk element nodes to insurance product nodes, identifying candidate products that highly match user needs in terms of coverage, coverage amount, and applicable scenarios. Subsequently, the content generation device conducts a multi-dimensional suitability assessment of each candidate product, including analyzing the linguistic complexity of the product terms to assess user comprehension difficulty, calculating the coverage ratio between the product's liability and the user's protection gap, and measuring the correlation between product features and trending events. By weighted and integrated these assessment results, the content generation device generates a comprehensive matching score for each candidate product and selects the product with the highest score as the final recommendation target.

[0041] In one embodiment, unstructured risk information and protection needs need to be transformed into query instructions that can be executed within a knowledge graph. It performs natural language processing on the descriptive text of the target core risk elements to identify key risk entities (e.g., event type: flight delay, affected area: nationwide, affected population: business travelers) and risk attributes (e.g., risk level: medium, duration: short-term). Simultaneously, based on the user's policy protection gap score and gap liability item list, combined with preset risk protection need mapping rules, this gap information is converted into query conditions for insurance product underwriting liability nodes (e.g., needing to include "flight delay allowance" and "lost baggage compensation") and coverage amount nodes (e.g., coverage amount must be greater than 500 yuan). The device then uses these key risk entities, risk attributes, and query conditions through a preset knowledge graph query language template (e.g., Cypher or SPARQL template) to generate a structured query instruction. Next, the content generation device uses the target core risk element (such as "flight delay") as the starting node set. Within a pre-defined insurance product knowledge graph, it executes the generated structured query instructions and uses graph traversal algorithms (such as breadth-first or depth-first search) to find all insurance product nodes that meet the query conditions. These initially qualified nodes are identified as candidate insurance product nodes. Subsequently, natural language processing is used to perform language complexity analysis on the product terms text corresponding to the candidate insurance product nodes (such as calculating average sentence length and frequency of professional terms) to obtain a terms comprehensibility score. A list of insured liability items is determined from the structured product data of the candidate insurance products and compared with the user's list of missing liability items. The coverage ratio of the former to the latter is calculated as the underwriting matching score. Furthermore, it extracts keywords from the description text of the core risk element and the introduction and terms text of the candidate products, calculates the Jaccard similarity coefficient between the two keyword sets, and obtains an event-product relevance score. The device calculates a weighted sum of the three scores (understanding of terms, underwriting matching, and relevance of the event to the product) based on preset weights to obtain a comprehensive matching score for each candidate insurance product, and selects the insurance product with the highest comprehensive matching score as the final target insurance product.

[0042] In other embodiments, the content generation device first performs feature engineering on all products in the insurance product knowledge base from multiple dimensions (such as coverage, product terms, target customer group, and price range), and uses a deep learning model (such as the dual-tower model) to encode each product into a high-dimensional vector, which is then stored in a vector database. Secondly, when matching is required, the device encodes the user's coverage gap (gap liability list, coverage gap) and the target core risk elements (risk type, scope of impact, etc.) into a query vector. An efficient near-nearest neighbor search is performed in the vector database to recall a batch of candidate products that are semantically most similar to the user's needs and risk scenarios. A re-ranking model (such as LambdaMART) then intervenes, using more multi-dimensional cross-features (such as the product's historical conversion rate, the matching degree between the user's profile features and the product's target customer group, etc.) to re-score and rank the recalled candidate products, ultimately outputting the top-ranked product as the target insurance product. It is understandable that other methods can be used to determine the target product, such as collaborative filtering to find which products other user groups with similar protection gaps to the current user ultimately purchased, and making recommendations based on trending events. This is not limited here.

[0043] S104. Based on the product characteristics and core risk elements of the target insurance product, determine the corresponding prompt word generation strategy for potential users from the preset prompt word strategy library.

[0044] Among them, product features represent the key attribute information of the target insurance product, including multi-dimensional product description elements such as coverage scope, coverage amount, payment method, claims process, and product advantages. The prompt word strategy library represents a pre-built knowledge base containing prompt word templates and generation rules corresponding to various marketing scenarios, user groups, and product types. Each strategy template defines the prompt word's structural framework, content focus, language style, and other elements.

[0045] By conducting in-depth semantic analysis of the underwriting terms of target insurance products, the semantic similarity between each term and the risk elements of the target trending event is calculated. Combined with the importance weight of the gap coverage items in the user's protection gap, core protection highlights that are highly relevant to the current trending event and can effectively fill the user's protection gap are selected. Simultaneously, the content generation device segments potential users into different decision-making stages based on their historical behavioral data and interaction patterns. For example, users in the information gathering stage are more concerned with a comprehensive product introduction, while users in the decision-making stage require a prominent display of key advantages. Based on these analysis results, the content generation device selects the most suitable strategy template from a prompt word strategy library and adjusts the sentiment and urgency parameters in the strategy template according to the public attention index of the trending event, ultimately integrating them to form a personalized prompt word generation strategy for specific user groups.

[0046] In one embodiment, the device acquires the access trajectory sequence of each potential user within a preset time window for the target insurance product and similar products. The similar products are determined based on a preset insurance product knowledge graph. The access trajectory sequence includes a time-series record of the user's access time, browsing behavior, and operational behavior on the corresponding page of the target insurance product. This data is collected through website analytics tools and a user behavior tracking system. Next, the page dwell time, page scroll depth, and page element click frequency in the access trajectory sequence are weighted and summed according to preset weights to calculate the interaction depth feature value for the corresponding user among the potential users. Based on this feature value, the device divides the potential users into multiple user subgroups, such as the "Initial Awareness Group" (shallow interaction), the "Interest Exploration Group" (intermediate interaction), and the "Deep Comparison Group" (deep interaction), and assigns them corresponding recommendation process stage tags. These recommendation process stage tags include the recommended content generation focus features and format templates based on the risk characteristics of the target hot topic event, used to guide the determination of subsequent prompt word generation strategies. The device uses a semantic matching algorithm to calculate the semantic similarity between each coverage clause and current hot risks from the underwriting liability subgraph (part of the knowledge graph) corresponding to the target insurance product, based on the target's core risk elements and the user's list of gap liability items. It then uses an importance ranking algorithm (such as a variant of PageRank) to determine the importance weight of each clause based on the gap liability item list. Clauses with both high semantic similarity and high importance weight are selected as "core coverage highlights." The device selects the most suitable strategy template from a pre-set prompt word strategy library based on the recommendation process stage tags of the user subgroup. For example, it selects the "risk education + product introduction" template for the "initial awareness group" and the "advantage comparison + Q&A" template for the "in-depth comparison group." Subsequently, the device dynamically adjusts the sentiment parameters (e.g., using a "concerned" tone if the attention is high) and urgency parameters (e.g., increasing "urgency" if the attention is extremely high) in the selected strategy template based on the public attention index value range of the target's core risk elements. The device integrates the extracted "core protection highlights", the "strategy template" with adjusted parameters, and other "product features" of the target insurance product (such as brand and price) to form a complete prompt word generation strategy.

[0047] S105. Input the prompt words generated based on the prompt word generation strategy into the pre-trained insurance domain big language model to obtain the initial draft of the creative content.

[0048] The device constructs structured prompts based on a pre-defined generation strategy. These prompts detail various aspects of the target content, including the overall framework, key benefits to be highlighted, appropriate language style, target user characteristics, and relevance to current trending events. The prompt design adheres to the input specifications of a large language model, ensuring accurate model understanding of the generation task through clear instruction hierarchies and output requirements. The content generation device then inputs the constructed prompts into a pre-trained insurance-related large language model. This model, trained on extensive data from insurance literature, product brochures, and marketing cases, possesses a deep understanding of insurance industry knowledge and the ability to utilize professional terminology. Based on the input prompts, the model generates marketing content that meets the requirements.

[0049] In some embodiments, the strategy-based content generation process can be implemented in several ways: Optionally, a hierarchical prompt word construction method can be adopted. First, a first-level prompt word containing basic background information is constructed, including descriptions of hot events, user profile summaries, and basic product information. Then, specific generation requirements are added to form second-level prompt words, including content structure requirements, language style settings, and key content to be highlighted. Next, personalized elements are added according to the specific needs of the user group to form third-level prompt words. Finally, the content interacts with the large language model through multi-turn dialogues to gradually improve and optimize the quality and relevance of the generated content. Optionally, a method combining template filling and intelligent generation can be adopted. A content template containing placeholders is pre-designed, and the content generation device automatically fills in structured content such as product data, user information, and event elements. Then, the large language model is used to generate unstructured content such as creative descriptive text, vivid scenario examples, and personalized expressions. Finally, the structured and unstructured content are organically integrated to generate complete marketing content. No limitations are imposed here.

[0050] S106. If the initial draft of the creative content passes the compliance test, generate the target recommended video based on the initial draft of the creative content.

[0051] The content generation device conducts a comprehensive compliance check on the initial draft of the generated creative content. This process includes multiple levels of review: checking for prohibited words or exaggerated statements at the vocabulary level; analyzing for misleading descriptions or inappropriate promises at the semantic level; and verifying the clarity and rationality of the content's logic at the structural level. Only content that passes the compliance check will proceed to the subsequent video production process. Next, the content generation device initiates the multimedia content generation process, converting the text script into emotionally expressive audio material and generating natural and fluent narration audio through speech synthesis technology. Simultaneously, the content generation device utilizes virtual digital human technology to create a visual avatar of a virtual anchor, automatically generating corresponding lip movements, facial expressions, and body movements by analyzing the audio's speech characteristics to achieve a synchronized audio-visual visual effect. Furthermore, the content generation device retrieves matching background images, animation effects, charts, and other auxiliary materials from a pre-set multimedia material library based on the content's keywords and themes. Finally, it organically combines all elements according to a pre-set video production template to generate a complete recommended video.

[0052] In one embodiment, after confirming that the initial draft of the creative content passes the compliance check, the text script from the compliance-tested initial draft is input into an advanced text-to-speech (TTS) engine. This engine not only generates fluent and natural speech but also automatically assigns a preset emotional tone to the generated audio material based on the emotional tendency parameters (such as "concern" and "professionalism") set in strategy S104. Next, to generate the "human" element in the video, the device utilizes virtual digital human technology. It takes the phoneme time sequence and speech frequency features from the audio material generated in the previous step as input and, through a deep learning model, generates in real time a virtual human's lip-sync sequence, facial expression sequence (such as smiling and frowning), and body movement sequence (such as gestures and nodding) that highly match the speech content. The content generation device precisely synchronizes these change sequences with the timeline of the audio material, outputting a synchronized audio-visual video clip of the virtual human anchor reading the script. The device may also retrieve and match corresponding multimedia materials from a pre-set, authorized library of video, image, and animation materials based on keywords in the initial draft of the creative content (such as "typhoon" or "hospitalization"). Following a pre-set video compositing template (which defines the layout, transitions, background music, etc.), the device intelligently combines, edits, and renders the generated audio materials, synchronized virtual human audio-visual video clips, and the retrieved multimedia materials to ultimately generate a complete target recommendation video.

[0053] In other embodiments, the generation of targeted recommendation videos can also be a lightweight method based on dynamic templates and intelligent material filling. The device maintains a dynamic video template library containing various visual styles (such as animated explanation style, graphic flash style). After confirming that the initial draft of the creative content passes compliance testing, the device first selects the most suitable template based on the user profile and content theme. Secondly, the device segments the text script and extracts core keywords for each segment. The content generation device does not generate virtual humans, but directly renders the text into artistic text or subtitles with dynamic effects, serving as the main information carrier of the image. Simultaneously, based on the keywords of each segment, it retrieves matching background videos or images from the material library and fills them into the corresponding positions in the template. The dynamic text layer, background material layer, and a background music (BGM) matched to the content's emotional content are synthesized to generate a targeted recommendation video with high information density and a strong sense of rhythm. This method can achieve automated video content production without using virtual humans. It is understood that other methods can also be used to generate targeted recommendation videos, such as sending the script to a third-party video production platform integrated with the system and calling its services via API to complete the video synthesis; this is not limited here.

[0054] In this embodiment, by employing the technical feature of dynamically matching the core risk elements of hot events with potential user protection gaps, and by combining real-time social hotspots with users' personalized risk protection needs, and by adopting the technical feature of generating initial drafts based on a large language model in the insurance field and conducting compliance checks in advance, the existing technologies effectively solve the problems of semantic ambiguity and compliance risks that are easily generated in template-based content and require a large amount of manual review, thereby improving the generation efficiency and professional compliance level of insurance short video content.

[0055] However, in practical applications, when implementing the above methods, the efficiency of user interaction may be reduced due to the simplistic presentation strategy for potential users at different decision-making stages or with varying depths of interest. This insurance short video content generation method can achieve intelligent content presentation strategies by integrating user interaction depth analysis, user subgroup segmentation, and dynamic prompt word generation strategies based on subgroups, thereby improving the system's efficiency in content distribution and user interaction. This, in turn, enhances the accuracy of content association with trending topics.

[0056] Please see Figure 2 This is another flowchart illustrating a method for generating insurance short video content in an embodiment of this application.

[0057] S201. Semantic analysis is used to obtain the core risk elements of the corresponding hot events from the hot event data within the preset time period.

[0058] S202. If it is determined that the public attention index of the target hot event exceeds the preset heat value threshold, and the core risk elements of the target hot event match the preset protection needs of potential users, obtain the policy data of potential users and determine the protection gap.

[0059] S203. Based on the protection gaps of potential users, identify target insurance products that match the core risk elements of the target hot events.

[0060] Steps S201~S203 and Figure 1 Steps S101 to S103 in the illustrated embodiment are similar and will not be repeated here.

[0061] S204. Obtain the access trajectory sequence of each potential user to the target insurance product and similar products within a preset time window.

[0062] The content generation device first determines the start and end times of a preset time window. The start time is typically set to the moment the target hot event is detected by the system or slightly earlier, and the end time can be set to the moment the current step is executed. It then extracts raw access log data for each user in the potential user list from the user behavior database. This log data records all user interactions on the product display platform. The raw access log data is then filtered and cleaned, retaining access records related to the target insurance product and similar products, while filtering out access records for other pages unrelated to the insurance product, such as product details pages, product comparison pages, premium calculation pages, and terms and conditions viewing pages. During the filtering process, the content generation device uses technologies such as product identifier matching, page URL parsing, and page content tag recognition to identify which access records belong to the target insurance product or similar products. The filtered access records are then sorted by timestamps to form a chronological access trajectory sequence, and each access node in the sequence is labeled with detailed attribute information, including the accessed product identifier, page type, source path to the page, and redirection path to the page. For each user in the potential user list, a complete access trajectory sequence data structure is generated. This data structure fully records all the user's interaction behavior trajectories related to the target insurance product and similar products within a preset time window.

[0063] In some embodiments, the acquisition of access trajectory sequences can be achieved in multiple ways: Optionally, the content generation device adopts a session-level trajectory tracking method. First, when a user accesses the product display platform, a unique session identifier is assigned to each user. This session identifier remains unchanged throughout the user's access process. By embedding a behavior tracking script in the front-end page, every user's page access, click, scroll, input, and other interactive behaviors are captured in real time. These behavioral data, along with the session identifier, user identifier, timestamp, and other information, are sent to the behavior data collection server in real time. The behavior data collection server merges and sorts the received behavioral data according to the user identifier and timestamp to form the original behavior stream for each user. Then, a session segmentation operation is performed on the original behavior stream, based on two consecutive interactions between the user and the session. The time interval between interactions determines whether they belong to the same access session. When the time interval exceeds a preset session timeout threshold (e.g., 30 minutes), they are divided into different sessions. Then, within each session, access behaviors related to the target insurance product and similar products are identified. By maintaining a product identifier mapping table, which records information such as the product number, product name, and associated page URL pattern of the target insurance product and all similar products, the access records of each page in the session are matched and judged using this mapping table. Successfully matched access records are retained and other irrelevant records are filtered out. Finally, the retained access records are organized into an access trajectory sequence according to time order, and metadata labels are added to the sequence, including statistical information such as the start time, end time, total number of access nodes, and number of products involved.Optionally, the content generation device employs a trajectory reconstruction method based on event tracking. First, event tracking points are pre-set on key pages and key operation locations of the product display platform. These tracking points include various types such as page view tracking, button click tracking, form submission tracking, and content expansion tracking. Each tracking point carries specific event parameters; for example, page view tracking points carry parameters such as page type, product identifier, and source channel, while button click tracking points carry parameters such as button function, click location, and product identifier. When a user triggers a tracking point, the front end reports the tracking event data to the data analysis platform. Then, the content generation device extracts all tracking event data for users within a specified preset time window and a specified list of potential users from the data analysis platform. Event correlation analysis is then performed on the tracking event data for each user. The system identifies which events belong to the same product access behavior. For example, "product details page browsing event" + "coverage expansion event" + "premium calculation button click event" can be associated as a complete product deep access behavior. Then, an event sequence pattern matching algorithm is used to identify event sequence fragments related to the target insurance product and similar products. These fragments are extracted and sorted according to their occurrence time to reconstruct the user's access trajectory sequence. During the reconstruction process, missing contextual information is supplemented. For example, if it detects that a user visited a product comparison page but the event tracking data lacks information about the specific products visited before entering the comparison page, the product list information carried in the comparison page event parameters can be analyzed to infer the product pages the user may have visited before, thus generating a more complete access trajectory sequence. It is understood that other methods can also be used to obtain the access trajectory sequence, such as directly analyzing the web server's access log files using server-side log parsing methods, or using API interfaces provided by third-party user behavior analysis tools to obtain user behavior data; this is not limited here.

[0064] S205. The page dwell time, page scrolling depth and page element click frequency in the access trajectory sequence are weighted and summed according to preset weights to calculate the interaction depth feature value of the corresponding user among the potential users.

[0065] The content generation device iteratively processes the access trajectory sequence of each potential user. When processing a single user's sequence, the device first extracts three key interaction metrics from the sequence: First, it calculates the dwell time on each page and sums or averages the dwell times of all related pages to obtain the total dwell time or average dwell time; second, it extracts the maximum scroll depth for each page visit and calculates the average scroll depth for all page visits; third, it counts the total click frequency of all preset key page elements in the sequence. After extracting these three values ​​(or normalized values), the device reads the preset weight values ​​for these three metrics from the configuration library. For example, the page dwell time weight is w1, the page scroll depth weight is w2, and the page element click frequency weight is w3. Finally, the device performs a weighted summation calculation using the formula: Interaction Depth Feature Value = (Page Dwell Time × w1) + (Page Scroll Depth × w2) + (Page Element Click Frequency × w3). Through this calculation, each potential user will obtain a corresponding interaction depth feature value.

[0066] In some embodiments, the calculation of interaction depth feature values ​​can be implemented in several ways: Optionally, a dynamic weight calculation method based on machine learning can be used. First, the content generation device collects a batch of labeled user behavior data, where the labels indicate whether the user will subsequently convert (e.g., make a purchase, schedule a consultation, etc.). Then, the device uses this data to train a logistic regression or gradient boosting tree model. The input features of the model are raw interaction indicators such as page dwell time, page scroll depth, and page element click frequency, and the output of the model is the conversion probability. Finally, when calculating the interaction depth feature values, the model is directly used to predict the interaction indicators of new users, and the resulting conversion probability value is used as its interaction depth feature value. Optionally, a more granular event weighting method can also be used. First, the device predefines a detailed event value scoring table, assigning a specific score to each specific event that may occur in the access trajectory (e.g., "clicking 'health declaration'", "viewing 'claims cases'", "using the premium calculator for more than 30 seconds"). Then, when processing the user's access trajectory sequence, the device no longer extracts macro indicators, but instead traverses each event in the sequence and finds its corresponding score in the scoring table. Finally, the scores of all events in the user trajectory sequence are summed to obtain the total score, which is the interaction depth feature value. This method can more accurately capture high-intent behaviors. It is understood that other methods can also be used to calculate the interaction depth feature value; this is not limited here.

[0067] S206. Based on the interaction depth feature value, potential users are divided into multiple user subgroups, and corresponding recommendation process stage labels are assigned to multiple user subgroups.

[0068] The content generation device first collects interaction depth feature values of all potential users to form a one-dimensional numerical data set. Then, the device performs a user division operation based on the data set. The division can be performed based on preset thresholds. For example, if the device predefines two thresholds T1 and T2 (T1<T2), users with interaction depth feature values less than T1 can be divided into a first sub-group, users with feature values between T1 and T2 can be divided into a second sub-group, and users with feature values greater than T2 can be divided into a third sub-group. After the division is completed, the device assigns a corresponding recommendation process stage label to each formed user sub-group. This assignment process is based on a preset mapping rule. For example, the sub-group with the lowest feature values is marked as "interest perception stage", the middle one is marked as "information exploration stage", and the highest one is marked as "comparison decision stage". Finally, the original heterogeneous potential user group is converted into several relatively homogeneous user sub-groups with clear stage identifiers.

[0069] In some embodiments, user division and label assignment can be implemented in a variety of ways: alternatively, an unsupervised clustering algorithm can be used to discover user groups. First, the content generation device takes the interaction depth feature values of all potential users (and possible other user portrait features, such as age, region, etc., which form a multi-dimensional feature vector) as input. Then, the device applies a clustering algorithm, such as the K-Means algorithm. The algorithm aggregates users with close distances in the feature space to form K clusters, and each cluster is a user sub-group. Finally, the device assigns a recommendation process stage label to the cluster by analyzing the centroid of each cluster (that is, the average value of all user feature vectors in the cluster), especially the average value of the interaction depth feature values. For example, the cluster with the lowest centroid interaction depth value is marked as "interest perception stage". Alternatively, a dynamic division method based on a rule engine can also be adopted. First, a rule engine is embedded in the device, and the rule base of the engine can be dynamically configured by operation personnel. The form of the rule can be "IF [interaction depth feature value > X AND there is a price comparison behavior within the last 7 days] THEN [stage label = 'comparison decision stage']". Then, the device inputs the interaction depth feature value of each user and other relevant behavior data into the rule engine. Finally, the engine directly assigns a corresponding stage label to the user according to the first matched rule, and all users who obtain the same label naturally form a user sub-group. This method has higher flexibility and can integrate more complex business logic. It can be understood that other methods can also be used to implement user division, for example, a density-based clustering algorithm such as DBSCAN can be used to cope with unevenly distributed data, which is not limited herein.

[0070] S207: Determine a strategy template from a preset prompt word strategy base according to the recommendation process stage label of the user sub-group.

[0071] For each subgroup, the device first obtains its assigned recommendation process stage label. Then, using this label as a query key, the device searches for a match in a pre-defined prompt word strategy library. After finding a matching strategy template, the device extracts it, using it as the basic framework for generating specific prompt words for that particular user subgroup in subsequent steps.

[0072] In some embodiments, strategy templates can be determined in several ways: Optionally, a dynamic template selection mechanism based on A / B test feedback can be used. First, for the same recommendation process stage label, multiple candidate strategy templates (A, B, C, etc.) can be preset in the strategy library. When a template needs to be determined for the user subgroup corresponding to the label, the device will select the template that has shown the highest conversion rate or click-through rate in the past based on historical A / B test data. The device will periodically run A / B tests, push the content generated by different templates to different user samples under the same stage label, collect effect data, and then update the "winning" status of the template. Optionally, a combined template generation method based on multi-tag matching can also be used. The user's tag system can be more complex, and in addition to recommendation process stage labels, it may also include user profile tags (such as "young parents" or "high-net-worth individuals"). The templates in the strategy library are also associated with multiple tag combinations. When determining a strategy template, the device will obtain all the tags common to the user subgroup (such as "information exploration stage" + "young parents"), and then search for a template in the strategy library that completely matches or is most similar to this tag combination. If no perfect match can be found, the device can dynamically generate a more targeted combined strategy template by splicing or merging the basic template matching the "information exploration stage" and the supplementary template matching "young parents" according to preset rules. No limitation is made here.

[0073] S208. Based on the numerical range of the public attention index in the core risk elements of the target, determine the sentiment tendency parameter and urgency parameter in the strategy template.

[0074] The content generation device first acquires the core risk elements related to the current hot topic and reads their corresponding public attention index values. Next, the device compares this index value with several preset numerical ranges to determine its level. For example, the preset ranges might be: 0-30 for "low attention," 31-70 for "medium attention," and 71-100 for "high attention." Then, the device maps this attention level to specific sentiment and urgency parameters according to a preset mapping rule. For example, the rule could be defined as: low attention corresponds to "neutral" sentiment and "normal" urgency; medium attention corresponds to "concern" sentiment and "relatively urgent" urgency; and high attention corresponds to "warning" sentiment and "extremely urgent" urgency. Finally, the device associates these two parameters (such as "concern" and "relatively urgent") with a strategy template or directly replaces the corresponding placeholders in the template, forming a dynamically adjusted strategy template with clear sentiment and urgency guidance.

[0075] S209. From the underwriting liability sub-graph corresponding to the target insurance product, based on the target core risk elements and the gap liability item list, the semantic similarity between each protection clause and the target core risk elements is calculated by semantic matching algorithm, and the importance weight of the protection clause is determined based on the gap liability item list by importance ranking algorithm.

[0076] This step involves two parallel computational processes. The first process calculates semantic similarity: the content generation device iterates through the text of each coverage clause in the target insurance product's coverage subgraph. For each clause, the device uses a semantic matching algorithm to calculate a semantic similarity score between its text content and the target core risk element (e.g., the text description "acute myocardial infarction"). The second process calculates importance weights: the device similarly iterates through each coverage clause. For each clause, the device checks whether its coverage matches any item in the potential user's list of gap coverage items. The device quantifies this match using an importance ranking algorithm; for example, the more gap coverage items a clause covers, or the higher the priority of the covered gap coverage items in the list, the higher its importance weight. Ultimately, each coverage clause receives two independent scores: a semantic similarity score and an importance weight score.

[0077] In some embodiments, this step can be implemented in several ways: Optionally, when calculating semantic similarity, a knowledge graph-based path-based correlation calculation can be used. First, the content generation device links the target core risk element (such as "lung cancer") and the coverage terms in a large-scale medical or risk knowledge graph, locating the corresponding nodes. Then, the device calculates the shortest path length between the two nodes in the graph or calculates their correlation score using an algorithm (such as PersonalizedPageRank). The shorter the graph path or the higher the correlation score, the higher the semantic similarity is considered. This approach can utilize external knowledge to discover deep connections not obvious on the surface of the text. Optionally, when calculating importance weights, a user profile-based weighted ranking algorithm can be used. First, the device considers not only the list of gap liability items but also the profile tags of user subgroups (such as "family breadwinner" or "mortgage holder"). Then, the importance ranking algorithm assigns a base weight to different gap liability items and dynamically adjusts this weight according to the user profile tags. For example, for "family breadwinner" users, the importance weight of "death / total disability liability" will be significantly increased. Finally, the importance weight of a coverage clause is the sum of the weights of all gap liability items it covers, weighted by the profiling. Understandably, this step can also be achieved in other ways, such as by integrating semantic similarity and importance weights into a unified model for end-to-end joint learning and computation; this is not limited here.

[0078] S210. The core protection feature is the protection clause where the semantic similarity exceeds the preset semantic threshold and the importance weight exceeds the preset weight threshold.

[0079] The content generation device iterates through all the guarantee clauses and their corresponding two scores processed in S209. For each guarantee clause, the device performs a dual-condition judgment: first, it checks whether its semantic similarity score is greater than or equal to a preset semantic threshold; second, it checks whether its importance weight score is greater than or equal to a preset weight threshold. Only when a guarantee clause meets both conditions is it deemed qualified by the device. The device collects all guarantee clauses that pass this dual screening, forming a set, which is defined as the core guarantee highlights. These highlights will serve as key information that needs to be emphasized and described in detail when generating subsequent content.

[0080] S211. Integrate the core protection highlights, the adjusted strategy template, and the product characteristics of the target insurance product to determine the prompt word generation strategy corresponding to the user subgroup.

[0081] The content generation device performs an integration operation for each user subgroup. First, the device obtains the strategy template corresponding to that subgroup, whose parameters have been adjusted in S208. Then, the device fills in the core coverage highlights (which may be one or more paragraphs of text) identified in S210 and other product features of the target insurance product extracted from the product library according to the placeholders preset in the strategy template. For example, the [Product Name] placeholder in the template will be replaced with the actual product name, and the [Core Coverage Highlights] placeholder will be replaced with the selected specific clause descriptions. In addition to simple text replacement, the device also attaches the adjusted sentiment parameters and urgency parameters as meta-instructions to the strategy. Finally, the device generates a complete prompt generation strategy for that user subgroup. This strategy may be a complex prompt that includes background, role, task, format requirements, sentiment orientation, and specific content.

[0082] In some embodiments, the cue word generation strategy can be determined in several ways: Optionally, a structured data input method can be used. Instead of integrating all information into a single long text cue, the device constructs a structured data object (such as JSON) and uses it as input. This JSON object contains multiple fields, such as "role": "insurance expert", "emotion": "concern", "urgency": "more urgent", "target_audience_profile": "young parents in the information exploration stage", "key_highlights": ["description of clause A", "description of clause B"], "product_info": {"name": "XX Fortune", "company": "YY Life Insurance"}. Optionally, a role-based multi-turn dialogue strategy can also be used. The device does not generate the final content all at once, but generates the cue word strategy through a preset dialogue flow. For example, the device first sends an initial instruction to the large language model: "You are now an insurance consultant and need to introduce a product to users in the 'information exploration stage'." Then, the model may return a confirmation or a question. Next, the device inputs a second round of instructions: "This user's focus is on 'cardiovascular and cerebrovascular diseases,' and the product's core highlight is 'protection against mild stroke sequelae.' Please organize the key points of your introduction." Through this simulated dialogue, the model is gradually guided to focus on the core task, and the resulting dialogue history itself constitutes a complex prompt word generation strategy. It is understandable that other methods can be used to determine the prompt word generation strategy; this is not limited here.

[0083] S212. Input the prompt words generated based on the prompt word generation strategy into the pre-trained insurance domain large language model to obtain the initial draft of the creative content.

[0084] The content generation device iterates through all user subgroups identified in S206. For each subgroup, the device uses the corresponding prompt word generation strategy generated in S211. Then, the device takes this specific strategy as input and calls the generation interface of a pre-trained insurance domain big language model. Since the prompt word generation strategies for different user subgroups differ in templates, sentiment parameters, urgency parameters, and even the emphasis on core coverage highlights, the big language model generates a draft of creative content for each subgroup. For example, for the subgroup in the "interest perception stage," the prompt words received by the model emphasize popular science and guidance, and the generated draft might be a short, easy-to-understand article that sparks thought; while for the subgroup in the "comparison and decision-making stage," the prompt words received by the model emphasize urgency and product advantages, and the generated draft might be a piece of copy containing clear data comparisons and calls to action. The device stores these different drafts of creative content generated for each user subgroup and associates them with the corresponding user subgroup tags.

[0085] S213. If the initial draft of the creative content passes the compliance test, generate the target recommended video based on the initial draft of the creative content.

[0086] The content generation device first acquires a user subgroup and its corresponding initial draft of creative content. Next, the device sends this draft to a built-in compliance detection module. This module uses keyword matching, regular expressions, and semantic analysis to check the text for prohibited words or exaggerated promises such as "guaranteed returns," "maximum payout," or "only choice." Only when the initial draft of creative content completely passes all detection rules and is deemed compliant will the device proceed to the next step. After compliance is confirmed, the device invokes a text-to-video synthesis engine. It uses the compliant initial draft of creative content as the video's narration or subtitle script. Understandably, the video's visual style, background music, and pacing can also be differentiated based on the user subgroup's recommendation process stage tag. For example, a video generated for users in the "interest perception stage" might use animation and infographics with a relaxed pace; a video generated for users in the "comparison and decision-making stage" might use live expert commentary with a fast pace and highlight key data. Finally, the device generates a corresponding target recommendation video for each user subgroup.

[0087] In this embodiment, by employing technical features that divide potential users into subgroups at different stages of the recommendation process based on the user's interaction depth features with insurance products, and by integrating user subgroup tags, event popularity, and the importance of protection gaps, highly customized prompts are generated through multi-dimensional information fusion. This effectively solves the problems of lack of targeting and insufficient content relevance in existing technologies, thereby improving the accuracy of the association between marketing content and hot topics while maintaining timeliness, and enhancing the accuracy and matching degree of the generated content. The exemplary content generation device 300 provided in the embodiments of this application is described below. Figure 3 This is an exemplary hardware structure diagram of the content generation device 300 provided in the embodiments of this application.

[0088] In some embodiments, the content generation device 300 is a computer device or includes a computer device. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores data. The network interface of the computer device is used to communicate with other external terminals or servers via a network connection. In some embodiments, the network interface can be a wired network interface; in some embodiments, the network interface can also be a wireless network interface. When the computer program is executed by the processor, it implements the methods in the embodiments of this application.

[0089] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0090] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".

[0091] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0092] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for generating short insurance video content, characterized in that, The method includes: The core risk elements of the corresponding hot events are obtained by semantic parsing the hot event data within a preset time period. The hot event data includes multiple hot events, and the core risk elements include risk type identifier, impact scope parameter and public attention index. If it is determined that the public attention index of the target hot event exceeds the preset heat value threshold, and the core risk elements of the target hot event match the preset protection needs of potential users, the policy data of the potential users is obtained and the protection gap is determined. The protection gap includes a list of liability items that are the gap between the existing policy protection amount and the recommended protection amount calculated based on user profile data under the risk type corresponding to the risk type identifier, and the protection gap score of the potential user. Based on the protection gaps of the potential users, target insurance products matching the core risk elements of the target hot-spot events are identified. Specifically, this includes: using the core risk elements as the starting node set in a pre-defined insurance product knowledge graph, a graph traversal algorithm is used to identify nodes whose matching degree with the protection gap exceeds a preset threshold as candidate insurance product nodes; natural language processing is used to perform language complexity analysis on the product terms text corresponding to the candidate insurance product nodes to calculate the terms' comprehensibility score; a list of insured liability items and coverage parameters are determined from the structured product data of the candidate insurance products, and the list of insured liability items is compared with the list of insured liability items to calculate whether the list of insured liability items covers the list of insured liability items. The coverage ratio is used as the underwriting matching score; a first set of keywords is determined from the descriptive text corresponding to the core risk elements using a keyword extraction algorithm, and a second set of keywords is determined from the product introduction text and terms text of the candidate insurance products using the same keyword extraction algorithm; the ratio of the number of intersection elements to the number of union elements of the first and second keyword sets is calculated to obtain the event-product relevance score; the terms comprehensibility score, the underwriting matching score, and the event-product relevance score are weighted according to preset weights to obtain the comprehensive matching score of the candidate insurance products; based on the comprehensive matching score, the insurance product with the highest comprehensive matching score is selected from the candidate insurance product nodes as the target insurance product; Obtain the access trajectory sequence of each potential user for the target insurance product and similar products within a preset time window. The similar products are determined based on a preset insurance product knowledge graph. The access trajectory sequence includes a time sequence record of the user's access time, browsing behavior, and operation behavior on the corresponding page of the target insurance product. The page dwell time, page scrolling depth and page element click frequency in the access trajectory sequence are weighted and summed according to preset weights to calculate the interaction depth feature value of the corresponding user among the potential users. Based on the interaction depth feature value, the potential users are divided into multiple user subgroups, and corresponding recommendation process stage labels are assigned to the multiple user subgroups. The recommendation process stage labels include the recommendation content generation focus features and form templates for the risk characteristics of the target hot event. From the underwriting liability sub-graph corresponding to the target insurance product, based on the target core risk element and the gap liability item list, the semantic similarity between each protection clause and the target core risk element is calculated by a semantic matching algorithm, and the importance weight of the protection clause is determined based on the gap liability item list by an importance ranking algorithm. The insurance clauses whose semantic similarity exceeds a preset semantic threshold and whose importance weight exceeds a preset weight threshold are designated as core insurance highlights. The core insurance highlights include the insurance product terms' description of coverage for specific risks. Based on the recommendation process stage tags of user subgroups, a strategy template is determined from a preset prompt word strategy library. The strategy template defines the structure, length, and information focus of the prompt words. Based on the numerical range of the public attention index among the core risk elements of the target, the sentiment tendency parameter and urgency parameter in the strategy template are determined; By integrating the core protection highlights, the adjusted strategy template, and the product features of the target insurance product, a prompt word generation strategy corresponding to the user subgroup is determined. The prompts generated based on the aforementioned prompt generation strategy are input into a pre-trained insurance domain large language model to obtain an initial draft of the creative content; If the initial draft of the creative content passes the compliance check, a target recommended video is generated based on the initial draft of the creative content.

2. The method according to claim 1, characterized in that, Before the step of determining, using a graph traversal algorithm in a preset insurance product knowledge graph with the target core risk element as the starting node set, nodes whose matching degree with the protection gap exceeds a preset threshold as candidate insurance product nodes, the method further includes: Natural language processing is performed on the descriptive text of the target core risk elements to determine key risk entities and risk attributes. The key risk entities include event type, affected area and affected population, and the risk attributes include risk level and duration. Based on the user's policy coverage gap score and the list of gap liability items, and combined with the preset risk protection demand mapping rules, the policy coverage gap score is converted into query conditions for insurance product underwriting liability nodes and protection amount nodes; The key risk entities, risk attributes, and query conditions are used to generate structured query instructions for the preset insurance product knowledge graph using a preset knowledge graph query language template.

3. The method according to claim 1, characterized in that, The step of obtaining the policy data of the potential users and determining the coverage gap specifically includes: Based on the risk type identifier of the target hot event, the coverage scope of the existing policy and the corresponding existing policy coverage amount are extracted from the policy data of the potential users; Based on the user profile data of the potential users, the recommended coverage amount for the risk type is calculated using a preset risk assessment model; The list of standard liabilities for the preset risk type corresponding to the risk type identifier is compared with the coverage of the existing policy to generate the list of gap liability items. The difference between the recommended coverage amount and the existing policy coverage amount is calculated, and the coverage gap score is obtained by combining the number and importance level of the gap liability item list through a preset scoring function.

4. The method according to claim 1, characterized in that, The step of generating a target recommended video based on the initial draft of the creative content, after determining that the initial draft of the creative content has passed the compliance test, specifically includes: If the initial draft of the creative content passes the compliance test, the text script in the initial draft of the creative content will be converted into audio material with a preset emotional tone. The phoneme time sequence and speech frequency features in the audio material are used to generate a virtual human's lip shape change sequence, facial expression change sequence and body movement sequence through virtual digital human technology. The change sequence is then synchronized with the timeline of the audio material to output a virtual human audio-visual synchronized video clip. Based on the keywords in the initial draft of the creative content, retrieve and match corresponding multimedia materials from the preset material library; The audio material, the virtual human audio-visual synchronized video clip, and the multimedia material are combined according to the preset video synthesis template to generate the target recommended video.

5. A content generation device, characterized in that, The content generation device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors invoke the computer instructions to cause the content generation device to perform the method as described in any one of claims 1-4.

6. A computer program product containing instructions, characterized in that, When the computer program product is run on a content generation device, the content generation device performs the method as described in any one of claims 1-4.

7. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on the content generation device, the content generation device performs the method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Information processing method and device

    CN120512587A

  • Content generation method and device, electronic equipment and storage medium

    CN120579523A