An insurance user portrait generation method and device based on PageRank and mutual information

By constructing a multi-layered hypergraph and performing gated diffusion and bounce injection, the problem that traditional sequence modeling cannot take into account both long and short cycle behaviors is solved, improving the accuracy and reliability of insurance user profiles and enabling timely reflection of users' potential risk status.

CN121190112BActive Publication Date: 2026-04-21CHINA LIFE INSURANCE CO LTD HUBEI BRANCH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA LIFE INSURANCE CO LTD HUBEI BRANCH
Filing Date
2025-09-08
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Traditional sequence modeling methods struggle to simultaneously capture both the long-term steady-state and short-term explosive characteristics of insurance user behavior, resulting in low accuracy in user profile generation and an inability to promptly reflect users' potential risk profiles.

Method used

By constructing a multi-layer hypergraph based on PageRank and mutual information and performing gated diffusion, steady-state and burst diffusion vectors are generated. The mutual information tensor is combined with back-injection to generate user profile vectors. The feature dependencies are verified by causal auditing, and spurious features are eliminated.

Benefits of technology

It enables dynamic modeling across time scales, improves the accuracy and credibility of user profiles, ensures consistency between profiles and user behavior, and enhances the timely response to potential risk situations and practical value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190112B_ABST
    Figure CN121190112B_ABST
Patent Text Reader

Abstract

This application provides a method and apparatus for generating insurance user profiles based on PageRank and mutual information, relating to the field of data processing. The method first segments the insurance event stream to form time-phase slice packets containing long-term and short-term phase windows. A multi-layer hypergraph is constructed based on these slice packets, and gated diffusion is performed at the long-phase and short-phase layers respectively to obtain steady-state diffusion vectors and burst diffusion vectors. The multi-layer hypergraph is updated by calculating the mutual information tensor and introducing a transport graph combined with bounce injection. A second gated diffusion is then performed, transforming the hypergraph into feature-level response vectors, which are fused within the feature neighborhood of user nodes to generate a user profile vector. Finally, feature dependencies are verified through causal auditing and counterfactual replay, and spurious correlations are eliminated to obtain a target user profile vector containing steady-state contributions, burst contributions, and explanatory chains. Implementing the technical solution provided in this application facilitates improvements in the accuracy of insurance user profile generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, specifically to a method and apparatus for generating insurance user profiles based on PageRank and mutual information. Background Technology

[0002] In the insurance business scenario, user interaction data often exhibits dynamic characteristics across scales. That is, over a long period, it presents sparsely distributed low-frequency behavioral trajectories, such as renewal inquiries, policy payments, or long-term claims records, while over a short period, there may be concentrated bursts of high-frequency behavioral peaks, such as batch clicks during marketing campaigns, multiple insurance intentions within a short period, or high-frequency claims applications triggered by sudden events.

[0003] Traditional sequence modeling methods typically rely on fixed windows or single time scales, making it difficult to simultaneously account for both long-term steady-state behavior and short-term explosive behavior. This can easily lead to the neglect of important cross-scale dependencies during the modeling process. Due to this deficiency, the generated user profiles cannot reflect users' potential risk profiles in a timely manner, resulting in low accuracy in insurance user profile generation.

[0004] Therefore, there is an urgent need for a method and apparatus for generating insurance user profiles based on PageRank and mutual information. Summary of the Invention

[0005] This application provides a method and apparatus for generating insurance user profiles based on PageRank and mutual information, which can help improve the accuracy of insurance user profile generation.

[0006] The first aspect of this application provides a method for generating insurance user profiles based on PageRank and mutual information. The method includes: segmenting an acquired insurance event stream according to time sequence to generate a time phase slice packet containing a long-term phase window and a short-term phase window; constructing a multi-layer hypergraph containing user nodes, feature nodes, and phase nodes based on the time phase slice packet; performing gated diffusion at the long-term phase layer of the multi-layer hypergraph to output a steady-state diffusion vector, and performing gated diffusion at the short-term phase layer of the multi-layer hypergraph to output a burst diffusion vector; calculating a mutual information tensor based on the time phase slice packet, and then... The mutual information tensor is introduced into the transport graph, and a bounce injection is performed in combination with the steady-state diffusion vector and the burst diffusion vector to output a multi-layer hypergraph after bounce injection. Secondary gated diffusion is then performed using the multi-layer hypergraph after bounce injection to transform the steady-state diffusion vector and the burst diffusion vector into feature-level response vectors. These feature-level response vectors are then fused within the feature neighborhoods already held by user nodes to generate user profile vectors. Causal auditing and counterfactual replay are used to verify the feature dependencies in the user profile vectors, eliminating spurious correlation features to obtain a target user profile vector containing steady-state contributions, burst contributions, and explanatory chains.

[0007] A second aspect of this application provides an insurance user profile generation apparatus based on PageRank and mutual information. The apparatus includes an acquisition module and a processing module. The acquisition module is used to segment the acquired insurance event stream according to time sequence to generate a time phase slice package containing long-term phase windows and short-term phase windows. The processing module is used to construct a multi-layer hypergraph containing user nodes, feature nodes, and phase nodes based on the time phase slice package, and to perform gated diffusion in the long-term phase layer of the multi-layer hypergraph to output a steady-state diffusion vector, and to perform gated diffusion in the short-term phase layer of the multi-layer hypergraph to output a burst diffusion vector. The processing module is further used to calculate the mutual information tensor based on the time phase slice package. The processing module further performs secondary gated diffusion using the multi-layered hypergraph after back-injection, transforming the steady-state diffusion vector and the burst diffusion vector into feature-level response vectors. It also fuses the feature-level response vectors within the feature neighborhoods already held by the user nodes to generate user profile vectors. Finally, it verifies the feature dependencies in the user profile vectors through causal auditing and counterfactual replay, eliminating spurious correlation features to obtain a target user profile vector containing steady-state contributions, burst contributions, and explanatory chains.

[0008] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, and both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method described above.

[0009] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described above.

[0010] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages:

[0011] 1. By sequentially segmenting the insurance event stream and generating time-phase slice packages containing long-term and short-term phase windows, it is possible to simultaneously capture both long-term stable behavioral patterns and short-term explosive behavioral characteristics of users. This avoids the loss of important information caused by traditional single-time-scale modeling, thus achieving dynamic modeling across time scales. By constructing a multi-layered hypergraph containing user nodes, feature nodes, and phase nodes, and performing gated diffusion at the long-phase and short-phase layers respectively, a hierarchical representation of steady-state and explosive behaviors can be formed within the graph structure. This ensures that the propagation paths and contributions of user behavior at different time scales can be independently modeled and quantified.

[0012] 2. By calculating the mutual information tensor based on time-phase slice packets and introducing it into the transport graph, combined with the steady-state diffusion vector and the burst diffusion vector for bounce injection, statistical correlation and structural relationships can be applied to the propagation process. This ensures that feature coupling relationships are strengthened during cross-scale propagation, thereby enhancing the accuracy and robustness of feature dependency modeling. By performing secondary gated diffusion on the multi-layered hypergraph after bounce injection, the steady-state diffusion vector and the burst diffusion vector are transformed into feature-level response vectors. This enables the formation of high-resolution response representations at the feature level, resulting in fine-grained feature contribution structures in the generation of user profile vectors, thus improving the accuracy of the profile.

[0013] 3. By fusing feature-level response vectors within the feature neighborhood already held by user nodes to generate user profile vectors, the consistency between the generated user profile and the user's historical behavior and current features is ensured. This avoids interference from isolated or irrelevant features, making the profile results more targeted and credible. Causal auditing and counterfactual replay are used to verify the feature dependencies in the user profile vectors, eliminating spurious features and ensuring that the final target user profile vector possesses causal consistency and interpretability. Simultaneously, steady-state contributions, burst contributions, and explanatory chains are preserved, thereby enabling timely reflection of potential user risk profiles and enhancing the practical value and traceability of the profile results in insurance business scenarios. Attached Figure Description

[0014] Figure 1 A flowchart illustrating an insurance user profile generation method based on PageRank and mutual information, provided for an embodiment of this application;

[0015] Figure 2 A schematic diagram of a module for generating insurance user profiles based on PageRank and mutual information, provided in an embodiment of this application;

[0016] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0017] Explanation of reference numerals in the attached figures: 21. Acquisition module; 22. Processing module; 31. Processor; 32. Communication bus; 33. User interface; 34. Network interface; 35. Memory. Detailed Implementation

[0018] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0019] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0020] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0021] To address the aforementioned technical issues, this application provides a method for generating insurance user profiles based on PageRank and mutual information, referring to... Figure 1 , Figure 1 This is a flowchart illustrating a method for generating insurance user profiles based on PageRank and mutual information, provided in an embodiment of this application. The method is applied to a server and includes steps S110 to S160, as follows:

[0022] S110. The obtained insurance event stream is segmented according to time sequence to generate a time phase slice package containing long time phase windows and short time phase windows.

[0023] Specifically, a server refers to a computing entity that handles event access, time base alignment, data cleaning, and window segmentation, possessing the ability to continuously consume message queues, perform sequential writing, and generate batch slices. It can be deployed on cloud host clusters or container platforms, responsible for pulling raw events from message middleware, standardizing and writing them to disk, and then triggering the segmentation task. An example is an event access node in the Tokyo area receiving raw records from mobile applications, call centers, and third-party payment gateways, and transforming them into data rows with a unified structure. An insurance event flow refers to a time-based sequence of discrete events generated throughout the entire insurance business process. The smallest unit of an event is a record with a unified timestamp, event type label, data source label, and payload summary. Typical events include insurance application calculation requests, successful policy payment, policy amendment taking effect, claims acceptance, claims settlement, initiation of a human customer service session, in-app marketing clicks, and SMS receipt arrival. An example is the same user triggering an insurance application calculation request within a week, successfully paying the policy premium two hours later, and submitting a claims acceptance three days later.

[0024] Time order refers to the global sequential relationship based on a unified timestamp as the unique sorting key. It requires time zone standardization, out-of-order correction, and deduplication to ensure that events from the same user and across different sources are strictly incremental on the same timeline. An example is that local time from a mobile application and Coordinated Universal Time (UTC) from a payment gateway are unified to the Tokyo time zone and sorted by second, then payment confirmations arriving five seconds late are inserted into their correct positions. Segmentation refers to dividing an insurance event stream into adjacent, non-overlapping sets within predefined boundaries without disrupting the event content and order, and generating traceable slice identifiers and metadata for each set. An example is splitting a long-term event stream into August and September windows using the calendar month as the boundary, retaining event counts, source distribution, and missing percentages within each window.

[0025] Long-term phase windows refer to time intervals covering billing cycles, policy cycles, or seasonal cycles, used to characterize low-frequency and stable behavioral patterns and risk accumulation. Boundaries typically coincide with the start and end of the billing period or the policy's expiration date, allowing for boundary buffers to absorb cross-boundary events. An example is the entire month-long interval from 00:00 on August 1, 2025 to 24:00 on August 31, 2025, used to statistically analyze renewal inquiry rates and claims settlement rates within the billing period. Short-term phase windows refer to time intervals covering marketing campaign periods, intraday rhythms, or emergency response periods, used to capture high-frequency and concentrated behavioral pulses. Boundaries align with the start and end of the campaign or fixed hourly granularity, emphasizing rapid changes and peak positioning. An example is the two-hour campaign window from 18:00 to 20:00 on August 26, 2025, used to identify multiple short-term insurance purchase intentions and intensive customer service conversations.

[0026] A temporal phase slice package refers to a structured packaged product generated by the server after segmentation based on time sequence. It consists of long-term phase window slice sub-packages and short-term phase window slice sub-packages, accompanied by a slice index table, bridging records, and statistical summaries. Each slice contains metadata such as window identifier, phase label, coverage start and end, event count, source distribution, missing percentage, and quality annotation. An example is generating a slice package for all events of a single user in August. The long-term phase window slice summarizes behavior during the billing period, while the short-term phase window slice covers a two-hour marketing window. The two types of slices are mapped through an index table, which is used for subsequent gated diffusion and generation of steady-state and burst diffusion vectors in the multi-layered hypergraph.

[0027] Furthermore, the insurance event flow is first subjected to out-of-order correction and duplicate event cleaning, with unified timestamps and a clean event sequence generated. A unified timestamp is calculated for each original event, and deduplication is performed within a sliding window to form a clean event sequence with unified timestamps and event tags. Specifically, the following two types of calculations are used:

[0028]

[0029] in, Indicates the first The source timestamp of the original event. This represents the time difference correction value between the source time zone and the target time zone. This indicates the clock offset correction value for the message access link. Indicates the first A unified timestamp for each event.

[0030]

[0031] in, Indicates user identifier, Indicates the event type label, Indicates the event payload fingerprint, Indicates the duration of the deduplication sliding window. For indicator functions, Indicates the first The event is a repeating event. This represents a non-repeating event. For events that satisfy... The events are arranged according to Output in ascending order to form a cleaning event sequence, and use this cleaning event sequence as input for the next step of time phase labeling.

[0032] Subsequently, a time phase database was constructed, and phase labels were assigned to cleaning event sequences to form a time phase mapping table. The time phase database consists of four types of business timescales: circadian rhythm, billing cycle, policy milestones, and marketing campaign periods. For each event, its belonging phase and phase position confidence were estimated, specifically using phase posterior calculations.

[0033]

[0034] in, Indicates an event Belongs to the phase category Phase score, Indicates the relationship between the billing cycle and the category The corresponding set of time intervals Represents the set of policy node times. Indicates the set of times for marketing campaigns. Represents the set of time periods of the diurnal rhythm. These represent the weight coefficients for the four types of business time scales. The scores for each category are normalized to obtain the phase label and phase position confidence, and the results are summarized into a time phase mapping table. The time phase mapping table and the cleaning event sequence are used together as inputs for window determination and aggregation.

[0035] Next, based on the time phase mapping table, the sets of long-term phase windows and short-term phase windows are determined, and window aggregation is performed within these two types of windows to generate a set of structured slices. For each window, metadata such as event counts, source distribution, and missing percentage are statistically analyzed using the following aggregation calculations:

[0036]

[0037] in, Display window The number of events within.

[0038]

[0039] in, Display window Internal source The proportion of sources, Indicates an event Data source tags.

[0040]

[0041] in, Display window The percentage of missing values, Represents a collection of fields. Indicates an event In the field Is it missing? Check the window identifier, phase label, and start / end of coverage. , , Write the data to form a structured slice set, and use this structured slice set as input for cross-scale alignment and consistency verification.

[0042] Next, cross-scale alignment and consistency checks are performed on the structured slice set, outputting the slice set that passes the checks and the bridging record. Cross-scale alignment is based on maximizing temporal overlap, establishing a mapping between short-time phase windows and the long-time phase windows they cover, and generating bridging records for anomaly boundaries. Specifically, the following two types of calculations are used:

[0043]

[0044] in, Indicates short-time phase window With long phase window The normalized overlap rate.

[0045]

[0046] in, Indicates an event Whether to trigger cross-border bridging This represents the minimum time step. A mapping is established for window pairs that meet the overlap rate and consistency threshold. Bridging records are generated for events that trigger bridging. The set of verified slices and the bridging records are used as input for encapsulation and packaging.

[0047] Finally, using the validated slice set and bridging records as input, the slices are grouped and encapsulated according to user identifiers to obtain time-phase slice packages. To facilitate subsequent retrieval and propagation calculations, a slice index is created and written into a statistical summary, specifically using the following index and packaging definitions:

[0048]

[0049] in, Indicates user The set of slices Denotes the set of all slices. Indicates user slice index, Indicates slice identifier, Indicates whether it is a long-term phase or a short-term phase. and Indicates the start and end times of the coverage. , The corresponding bridging records and statistical summaries are encapsulated together into a temporal phase slice package, which is then used as the direct input for subsequent multi-layer hypergraph construction and gated diffusion, ensuring that the output of each step corresponds one-to-one with the input of the next step and is traceable.

[0050] S120. Construct a multi-layered hypergraph containing user nodes, feature nodes, and phase nodes based on the time phase slice package. Perform gated diffusion in the long phase layer of the multi-layered hypergraph to output the steady-state diffusion vector, and perform gated diffusion in the short phase layer of the multi-layered hypergraph to output the burst diffusion vector.

[0051] Specifically, a multi-layered hypergraph refers to a graph structure that connects multiple nodes with hyperedges and is layered according to temporal phase semantics. It consists of three types of nodes: user nodes, feature nodes, and phase nodes. It can simultaneously represent cross-user relationships, cross-feature relationships, and cross-temporal phase relationships, and allows independent propagation and joint constraints within different layers. An example is a hyperedge that connects a user node, two feature nodes, and a phase node, indicating that the user jointly triggered two behavioral features under that phase.

[0052] A user node is an abstract node in a multi-layered hypergraph that uniquely corresponds to a natural person. It carries attributes such as user identification, channel distribution, and profile fragment indexes, and is used for aggregation, distribution, recording, interpretation, feedback, and auditing. An example is user node U123, which corresponds to a unified identifier for a mobile application account and a call center number. A feature node is an abstract node in a multi-layered hypergraph that uniquely corresponds to a behavioral or attribute feature. It carries attributes such as feature definition, source statistics, and quality labeling, and is used for cross-temporal phase propagation calculations. Examples are feature node F (insurance trial calculation), feature node F (successful policy payment), and feature node F (claims acceptance). A phase node is an abstract node in a multi-layered hypergraph that uniquely corresponds to a time phase window. It carries attributes such as phase label, coverage start and end, and phase position confidence, and is used to bind the semantics of long-term or short-term phase windows. Examples are phase node P (billing period August 2025) or phase node P (activity from 6:00 PM to 8:00 PM on August 26, 2025).

[0053] A long-phase layer refers to a hierarchical subgraph consisting of phase nodes corresponding to all long-term phase windows and their associated user nodes and feature nodes. It is used to represent the cumulative effect of low-frequency stable behavior related to billing terms or seasons. An example is a subgraph containing phase node P with the billing term in August 2025 and the feature nodes and involved user nodes related to renewal consultations within that month. A short-phase layer refers to a hierarchical subgraph consisting of phase nodes corresponding to all short-term phase windows and their associated user nodes and feature nodes. It is used to represent the sudden effect of high-frequency burst behavior related to activity responses or intraday rhythms. An example is a subgraph containing phase node P with the activity from 6 PM to 8 PM on August 26, 2025, and the feature nodes and involved user nodes related to marketing clicks within that period. Gated diffusion refers to the information propagation process within a multi-layered hypergraph constrained by a gating strategy. The gating strategy is jointly determined by structural constraints, channel constraints, and phase constraints, controlling which hyperedges are passable, the pass rate and propagation priority, and the inactivation conditions, thereby avoiding noise path amplification and maintaining a clear interpretation chain. The example allows only user feature superedges with a phase position confidence higher than the threshold and a high source quality to participate in the first propagation, and reduces the weight of superedges across product lines.

[0054] The steady-state diffusion vector, after gated diffusion at the long-phase layer, is a vector representation describing the degree of long-term pattern influence for each feature node or user node. It emphasizes stable contribution and slow cumulative effect, while retaining the corresponding upstream path index and phase label trajectory. For example, in the billing term dimension, the steady-state diffusion vector of user node U123 for renewal consultation with feature node F is significant, reflecting its stable long-term renewal tendency. The burst diffusion vector, after gated diffusion at the short-phase layer, is a vector representation describing the degree of short-term burst influence for each feature node or user node. It emphasizes rapid propagation and peak response effect, while retaining the corresponding upstream path index and phase label trajectory. For example, within a two-hour activity window, the burst diffusion vector of user node U123 for marketing clicks with feature node F is significant, reflecting its concentrated interaction behavior within a short period.

[0055] Furthermore, the user identifier, event features, and phase tags in the time phase slice package are first mapped to user nodes, feature nodes, and phase nodes, respectively. Unique identifiers and attribute descriptions are then established for these three types of nodes, enabling them to be searchable and traceable. To support subsequent edge weight calculation and diffusion initialization, the occurrence frequency and phase coverage of nodes are calculated, and a node attribute table is generated. Specifically, the following calculations are used:

[0056]

[0057] in, Represents a node The frequency of occurrence, Indicates the first This event, Indicates an event Belonging to nodes The determination, Indicates the event timestamp. This represents the time phase window corresponding to the phase node. Represents a node In phase The coverage is calculated based on the frequency obtained from the event-to-node attribution count, and then normalized within a subset of the phase window to obtain the coverage. The aforementioned node attribute table serves as input for edge construction and edge weight allocation.

[0058] Subsequently, user feature edges are constructed based on the event attribution relationship between user nodes and feature nodes, user phase edges are constructed based on the temporal coupling relationship between user nodes and phase nodes, and feature phase edges are constructed based on the occurrence dependency relationship between feature nodes and phase nodes. Edge weights are then assigned to these three types of edges, forming a weighted structure of a multi-layered hypergraph. The edge weight assignment employs a combination of frequency normalization and quality weighting, specifically:

[0059]

[0060]

[0061]

[0062] in, Represents user node With feature nodes The right to the border, Indicates user Triggering features Event count, This indicates a score for the quality of the data source for this relationship. These are weighting coefficients; Represents user node With phase node The right to the border, Indicates user events in phase Time coverage ratio This represents the average confidence level of the phase label on the user side. These are weighting coefficients; Represents feature nodes With phase node The right to the border, Representation of features In phase The occurrence count within, This represents the temporal freshness score of a feature within a phase. These are weighting coefficients. The calculation principle is to determine the connection strength using both frequency ratio and quality factor, and to normalize the results during source comparisons to suppress scale differences. The aforementioned weighted edge set serves as the input for extracting long-phase and short-phase subgraphs and generating the transition matrix.

[0063] Phase nodes with long-term phase labels and their associated user nodes and feature nodes are extracted from the multi-layer hypergraph to form a long-phase layer subgraph. A column-normalized transition matrix and a gating matrix are constructed based on weighted edges, and gated diffusion is performed to output the steady-state diffusion vector. The diffusion adopts a gated random walk with restart, specifically:

[0064]

[0065] in, Indicates the first The steady-state diffusion vector of the next iteration. Represents the column-normalized transition matrix of the long-phase subgraph. This represents the gating matrix generated jointly by the edge weight threshold and the phase position confidence. This indicates element-wise multiplication. Indicates the diffusion retention coefficient. This represents the initialization vector anchored to the target user node. The computation principle involves probability propagation along traversable edges, gating to suppress low-confidence paths, and using a restart term to ensure convergence and anchor consistency. Iteration until convergence yields the steady-state diffusion vector. This steady-state diffusion vector serves as a baseline for comparison between parallel computation of the short-phase layer and subsequent multi-channel fusion.

[0066] Phase nodes with short-term phase labels and their associated user nodes and feature nodes are extracted from the multi-layer hypergraph to form a short-phase layer subgraph. Similarly, a column-normalized transition matrix and a gating matrix are constructed, and gated diffusion is performed to output the burst diffusion vector. The diffusion employs a fast response and sparsity preservation strategy, specifically:

[0067]

[0068] in, Indicates the first The burst diffusion vector of the next iteration. Represents the column-normalized transition matrix of the short-phase subgraph. This represents the gating matrix generated by the edge weight threshold, phase position confidence, and event density peak index. Indicates the diffusion retention coefficient. This represents the initialization vector anchored to the short-term activity characteristics of the target user node. The calculation principle involves prioritizing the propagation of high-intensity paths within a short timescale and suppressing backflow noise through gating. A restart term is used to lock in the short-term anchored features, and iteration until convergence yields the burst diffusion vector. The burst diffusion vector, together with the steady-state diffusion vector, serves as the input for subsequent mutual information tensor introduction, bounce injection, and secondary gated diffusion.

[0069] S130. Calculate the mutual information tensor based on the time phase slice packet, introduce the mutual information tensor into the transport graph, combine the steady-state diffusion vector and the burst diffusion vector to perform a bounce injection, and output the multi-layer hypergraph after the bounce injection.

[0070] Specifically, the mutual information tensor refers to a multidimensional array formed by stacking feature pairs along the time phase dimension after measuring information sharing under different time phase windows. It is used to quantify the strength and differences in dependency between feature nodes in long-term and short-term phases. For example, in the billing period window, "Policy Payment Successful" and "Renewal Consultation" have high mutual information, while in the activity window, "Marketing Click" and "Insurance Trial Calculation Request" have high mutual information; both are recorded in the same mutual information tensor for later use. The transport graph refers to a propagation graph with random walk and bounce capabilities constructed on a multi-layered hypergraph. It is used to describe the transition probabilities of information between user nodes, feature nodes, and phase nodes, and supports the injection of external priors into the walk process. For example, transition probabilities are generated by normalizing edge weights between user nodes and feature nodes, while retaining a bounce path to allow for a return to key nodes at any time.

[0071] Jump-back injection refers to injecting mutual information tensors, steady-state diffusion vectors, and burst diffusion vectors as prior signals into the jump mechanism of the propagation graph. This allows random walks to return to high-value nodes based on statistical correlation and cross-scale weights during propagation, thereby improving the throughput of effective paths and suppressing noisy paths. Examples include the walk jumping back to nodes with high mutual information and high steady-state weights, such as "successful policy payment," with a set probability when encountering low-confidence edges, or jumping back to nodes with high burst weights, such as "marketing clicks," within the activity window to accelerate focusing. The multi-layered hypergraph after jump-back injection refers to the graph structure after parameterized updates. This is reflected in the dynamic reweighting of transition probabilities, gating thresholds, or edge weights by the mutual information tensor and the two types of diffusion vectors, thus changing the manifold morphology of subsequent propagation and aggregation. Examples include the amplification of paths strongly correlated with "renewal consultation" at the billing layer and paths strongly correlated with "insurance trial calculation request" at the activity layer after injection, making the effective paths of the graph more closely aligned with cross-scale dependencies.

[0072] Furthermore, using the time-phase slice package as input, a joint occurrence frequency table of feature pairs is constructed based on the joint distribution of user nodes, feature nodes, and phase nodes in each time-phase window, and the corresponding time-phase labels are marked in the joint occurrence frequency table of feature pairs. For each time-phase window, the joint occurrence count, marginal occurrence count, and total number of events within the window are calculated as the basis for subsequent probability estimation, specifically as follows:

[0073]

[0074] in, Indicates the time phase window, Indicates the first This event, This represents a node identifier, which can be a user node or a feature node. Indicates in window internal nodes With nodes The joint occurrence count, Indicates in window internal nodes The marginal occurrence count, Display window The total number of events, This indicates the indicator function, and the event labeling rules are determined by the mapping relationship provided by the time phase slice package.

[0075] Within each time phase window, the degree of information sharing is calculated and a mutual information matrix is ​​formed. These matrices are then concatenated along the time dimension to form a mutual information tensor. To obtain robust probabilities, additive smoothing is applied to the joint and marginal distributions, and the mutual information is calculated accordingly. The specific calculation is as follows:

[0076]

[0077]

[0078] in, Display window internal nodes With nodes The joint probability, Display window internal nodes The marginal probability, Represents the additive smoothing coefficient. Display window The equivalent count of node pairs participating in the statistics. Display window mutual information value, The mutual information tensor represents the relationship between nodes. With time phase window The value of the position. The above calculation principle is to estimate the mutual information using a smoothed empirical distribution, and stack the mutual information matrix into a three-dimensional tensor using windows as slices.

[0079] The mutual information tensor is introduced as a correction factor for edges into the transmission graph of a multi-layered hypergraph, so that the transmission probability is simultaneously determined by the graph structure, edge weights, and statistical correlation. First, the mutual information tensor is normalized within each window, then the edge weights are reweighted and column normalization is performed to obtain the transmission probability. The specific calculation is as follows:

[0080]

[0081]

[0082] in, Display window Normalized values ​​of internal mutual information This indicates the avoidance of stable terms with a denominator of zero. Indicates in window Next to the edge Corrected edge weights, This represents the original weighted edge weights. Indicates the mutual information correction strength coefficient. Indicates in window Next node Transfer to node via random walk The transmission probability is calculated by linearly enhancing the original edge weights with normalized mutual information and normalizing them in the column directions, so that the transmission graph takes into account both structural strength and statistical dependence.

[0083] On the transport graph after introducing mutual information tensors, a bounce injection is performed by combining steady-state diffusion vectors and burst diffusion vectors to achieve dynamic equilibrium diffusion across time scales, and outputs a multi-layered hypergraph after bounce injection. The bounce injection is achieved by constructing a bounce vector within each time phase window and updating the node influence distribution with a restarted random walk, as specifically calculated below:

[0084]

[0085]

[0086] in, Display window The bounce vector, Represents the steady-state diffusion vector. Represents the burst diffusion vector. Describes the norm, Indicates in window The dynamic balance coefficient between the homeostatic channel and the burst channel. Indicates the first The distribution of node influence in the next iteration. Display window The retention coefficient, This represents the aforementioned transmission probability matrix after mutual information correction and normalization. The calculation principle involves diffusion based on transmission probabilities and restarting using a bounce vector, balancing steady-state and burst contributions, iterating until convergence, and then updating the matrix. With convergent distribution We jointly define the time phase layer parameters of the multi-layer hypergraph after bounce injection, which are used as direct inputs for subsequent secondary gated diffusion and feature-level response vector generation.

[0087] S140. A second-gated diffusion is performed using a multi-layered hypergraph after bounce-back injection to transform the steady-state diffusion vector and the burst diffusion vector into a characteristic-level response vector.

[0088] Specifically, secondary gated diffusion refers to a constrained propagation process executed again on a multi-layered hypergraph after bounce injection. Gating strategies control the traversable edges, their proportion, and propagation priority. Channels distinguish between steady-state and burst channels, ultimately converging the influence of the node layer to the usable representation of the feature layer. For example, in the long-phase layer, user feature edges with reliable sources and high phase position confidence are prioritized; in the short-phase layer, feature phase edges with dense events and high burstiness are prioritized. Simultaneously, cross-layer backflow is suppressed to ensure stable and interpretable propagation. The feature-level response vector refers to the response representation generated for each feature node after secondary gated diffusion. It integrates steady-state and burst contributions and retains the path interpretation index for subsequent fusion within the feature neighborhood already held by the user node. For example, the feature-level response vector for the feature "renewal consultation" simultaneously includes the steady-state contribution component from the billing period layer and the burst contribution component from the activity layer, pointing to the upstream key phase nodes and user node paths, facilitating weighted aggregation and traceable display when generating user profile vectors.

[0089] Furthermore, using the multi-layered hypergraph after bounce injection, the steady-state diffusion vector, and the burst diffusion vector as input, the user node set is determined as the diffusion anchor set, and the diffusion initiation state and anchor edge set are generated. First, the user node anchor score is calculated, and anchor users are selected based on this. Then, edges connected to anchor users and meeting the credibility criteria are extracted from the weighted edge set to form the anchor edge set. Simultaneously, the anchor score is injected as the diffusion initiation state. The calculation formula is as follows:

[0090]

[0091] in, This represents the value of the steady-state diffusion vector at the user node. This represents the value of the burst diffusion vector on the user node. and This represents the anchoring coefficient between the two channels. Indicates the anchoring threshold. Represents the diffusion anchor set, This represents the initialization vector with user nodes as unit vectors. Indicates the diffusion initiation state. Represents the anchored edge set, Indicates the edge weight, This represents the lower bound of the edge weights. The calculation principle is to obtain the anchoring strength by linearly superimposing the influence of the two channels, and use the normalized anchoring strength as the initial probability distribution. The high-confidence path is truncated by the edge weight threshold, and the output diffusion initiation state and anchoring edge set are used as the input of the gating strategy.

[0092] A gating strategy comprising structure gates, channel gates, and phase gates is constructed based on the diffusion initiation state and anchored edge set, and a gating weight configuration is generated. Three types of gating factors are calculated for each edge, and a comprehensive gating weight is obtained through a convex combination. The original edge weights are then reweighted to form the gating weight configuration. The calculation formula is as follows:

[0093]

[0094]

[0095] in, Represents the gate factor. Indicates the structure gate threshold; Indicates the channel gate factor. Indicates the index of the target node of the edge. and This represents the value of the influence degree of the two channels after max-min normalization. Indicates the channel ratio coefficient; Represents the phase gate factor. This represents the phase position confidence of the phase node to which the edge belongs. This indicates the freshness score of the edge within the phase. This represents the phase-to-freshness tradeoff coefficient; Indicates the overall gating weight. This indicates the combined weight of the three doors. This represents the edge weights after gating. The calculation principle is to use three dimensions—structural strength, channel influence, and phase reliability—to screen edges, suppressing noisy paths and highlighting critical paths. The output gating weight configuration serves as the input for hierarchical diffusion.

[0096] In the long-phase subgraph of the multi-layer hypergraph, diffusion under gated weight configuration is performed to form a steady-state contribution pool. First, a column-normalized transition matrix is ​​generated in the long-phase layer using the gated edge weights. Then, iterative propagation with restart is performed until convergence. The components distributed on the feature nodes at convergence are written into the pool as steady-state contributions. The calculation formula is as follows:

[0097]

[0098] in, The gated edge weight matrix represents the long-phase subgraph. Represents the long phase layer transfer matrix. Indicates the long phase layer preservation coefficient. Indicates the first The propagation state of the next iteration. Indicates the diffusion initiation state. Indicates steady-state distribution, Represents feature nodes The steady-state contribution is calculated by performing Markov propagation on the gating path and periodically restarting to the anchored distribution to ensure stable convergence of long-term patterns, and outputting a steady-state contribution pool for use by cross-gating.

[0099] In the short-phase subgraph of the multi-layer hypergraph, diffusion under gated weight configuration is performed to form an burst contribution pool. The method is consistent with that of the long-phase layer, but stronger response and sparsity constraints are adopted to ensure that the peak is rapidly preserved in short-time diffusion. The calculation formula is as follows:

[0100]

[0101] in, This represents the gated edge weight matrix of the short-phase layer subgraph. Represents the short phase layer transfer matrix. Indicates the short phase layer preservation coefficient. Indicates a convergent distribution. Represents feature nodes The burst contribution is calculated by enhancing the fidelity of short-term peaks with a lower preservation coefficient and sparser pathways, and outputting a burst contribution pool for use by cross-gating.

[0102] Cross-gating and conflict resolution are performed on the steady-state contribution pool and the burst contribution pool to generate a channel-level feature contribution table and path interpretation records. First, phase consistency and path independence are calculated. Then, weights are assigned to the contributions of the two channels using a soft allocation method. This automatically reduces the weights and retains the most interpretable path when phase conflicts and path re-entry exist. The calculation formula is as follows:

[0103]

[0104]

[0105] in, Representation of features Phase consistency score, Representation of features Path independence score, Represents the weight mapping coefficients. and This indicates the weight allocation between the two channels. and This represents the contributions of the two channels after cross-gating. The calculation principle is to obtain higher weights by using a log-linear mapping to the more consistent and independent channel contributions, while recording the corresponding upstream paths to form path interpretation records. The output channel-level feature contribution table serves as the input for residual synthesis and scale alignment.

[0106] The initial response vector is obtained by combining the channel-level feature contribution table, path interpretation record, and bounce injection residual with residual synthesis and scale alignment. First, global scale normalization is performed on the contributions of the two channels. Then, fusion coefficients are automatically assigned according to the overall channel energy and synthesized together with the residuals. The calculation formula is as follows:

[0107]

[0108]

[0109]

[0110] in, This represents the global mean and standard deviation of the steady-state channel. This represents the global mean and standard deviation of the burst channel. and This represents the normalized contribution of the two channels. This indicates that the bounce-injection residual is in the feature The amount on, Indicates the three-way fusion coefficient. Indicates the initial response vector in the feature The value is taken from the above. The calculation principle is to eliminate the scale difference by standardizing within the channel, then adaptively allocate the fusion ratio with the total channel energy, and inject residuals to retain the high-confidence local influence of the bounce stage. The output initial response vector is used as the input for robustness and sparsity.

[0111] The initial response vector is subjected to robustness and sparsity processing, outputting a feature-level response vector. Robustness uses quantile pruning to suppress extreme values, while sparsity uses soft thresholding to compress long-tail contributions, resulting in a stable and interpretable final representation under noise. The calculation formula is as follows:

[0112]

[0113]

[0114] in, and These represent the lower and upper quantiles of the initial response vector. and This represents the lower and upper thresholds for robustness. This indicates the response after clipping. Indicates the soft threshold strength. Indicates the feature-level response vector in the feature The final value is determined by quantile pruning to weaken abnormal spikes and background noise, followed by sparsification using a soft threshold. This ensures that the output is statistically robust and focuses on high-value features, forming a feature-level response vector that serves as the direct input for subsequent fusion within the feature neighborhood already held by the user node.

[0115] S150. Fuse feature-level response vectors within the feature neighborhood already held by the user node to generate a user profile vector.

[0116] Specifically, the already held feature neighborhood refers to a local subset consisting of feature nodes that have established valid connections between user nodes in the historical and current time phases, emphasizing relationships that have actually occurred and passed data quality verification. For example, user node A's already held feature neighborhood includes features such as policy renewal inquiries, successful policy payments, and marketing clicks, but does not include features such as policy cancellation applications that have not yet occurred or have only been reported once.

[0117] Furthermore, using the feature-level response vector as input, a set of feature nodes connected to the user node is retrieved to form a user node feature neighborhood. The corresponding steady-state contribution component, burst contribution component, and residual component are then extracted from this neighborhood. To ensure comparability of components from different features, intra-channel standardization is performed on the three types of components to obtain normalized components for subsequent fusion. The calculation formula is as follows:

[0118]

[0119] in, Representation of features The steady-state contribution component, Representation of features The contribution of the outbreak, Representation of features The residual components, These represent the mean values ​​of the three components in the feature neighborhood of the user node, respectively. These represent the standard deviations of the three components in the user node feature neighborhood. The above calculation principle is based on the zero mean and unit scale transformation within the channel to eliminate dimensional differences and improve the stability of subsequent aggregation.

[0120] The steady-state contribution component, burst contribution component, and residual component of each feature node in the feature neighborhood of the user node are concatenated to form a multi-dimensional feature contribution vector. This vector is then combined with time phase labels and path interpretation records to generate a weight allocation table. To ensure that the weights are constrained by phase consistency and path reliability, while also reflecting source quality, a fusion weight is calculated for each feature. The calculation formula is as follows:

[0121]

[0122] in, Representation of features Phase consistency score, Represents the set of time phases. Representation of features In phase Phase label confidence, Indicates the user in phase Activity weight; Representation of features Path reliability score, Representation of features Explanation of the chain path set, Representing a path The intensity of the impact, Representing a path Independence score; Representation of features Source quality score; This represents the weighting coefficients of the three rating items; Represents the feature neighborhood of a user node; This represents the normalized fusion weights. The calculation principle is based on phase consistency and path reliability as the primary signals, supplemented by source quality, and a fully neighborhood-comparable weight allocation table is obtained through soft maximum normalization.

[0123] Aggregation operations are performed within the feature neighborhood of user nodes based on a weighted allocation table to generate a user feature aggregation table. First, channel energy is adaptively allocated, then weighted aggregation is performed on the feature dimensions to obtain the channel-level user representation and the feature-level fusion value. The calculation formula is as follows:

[0124]

[0125]

[0126]

[0127] in, These represent the weighted total contributions of the three channels, These represent the adaptive matching coefficients of the three channels, This indicates that the stable term avoids a denominator of zero. Representation of features The fused contribution value is calculated by first summarizing the energy within each channel according to weights, then determining the channel allocation based on the energy percentage, and finally performing linear fusion along the feature dimension according to a uniform allocation to form the fused contribution of each feature in the user feature aggregation table.

[0128] Consistency correction and interpretability annotation are performed on the user feature aggregation table to remove cross-phase conflicting feature contributions and embed interpretation chains, outputting user profile vectors. To resolve phase conflicts and path reentry, each feature is corrected and traceable annotations are generated, calculated using the following formula:

[0129]

[0130]

[0131]

[0132] in, Representation of features Phase contradiction degree, Representation of features The degree of path contradiction, The threshold representing the discrepancy between phase and path. For indicator functions, This represents the characteristic contribution value after consistency correction. Represents a user profile vector. This represents a vector concatenation operation. The calculation principle is to eliminate conflicting contributions using a consistency threshold while retaining compliant paths, concatenating the channel allocation and corrected feature contributions together to form the final output, and simultaneously recording each feature in the user feature aggregation table. The upstream path identifiers are used to form an explanatory chain, thereby obtaining a user profile vector with steady-state contribution, burst contribution and explanatory chain.

[0133] S160. Verify the feature dependencies in the user profile vector through causal auditing and counterfactual replay, remove spurious correlation features, and obtain the target user profile vector containing steady-state contribution, burst contribution and explanatory chain.

[0134] Specifically, causal auditing refers to examining the dependency edges in a user profile vector based on conditional independence, temporal sequence, and business causal direction, retaining dependency edges that meet causal constraints and marking or removing those that do not. For example, in an event sequence where "successful policy payment" precedes "renewal consultation," it is confirmed that "payment changes" have a positive impact on "renewal consultation," while "frequency of late-night visits" is no longer significant for "claims processing" after controlling for "incident occurrence time," and is therefore determined to be non-causal. Counterfactual replay refers to intervening in the contribution of a certain feature while keeping other conditions unchanged, replacing it with a control state, and recalculating the changes in the user profile vector to examine the actual strength of the feature dependency path. For example, "high-frequency marketing clicks" is set to zero in the activity phase and replayed; if the contribution of "insurance trial calculation request" decreases significantly, the path is confirmed as a genuine contribution path; otherwise, it is marked as a pseudo-related path.

[0135] Steady-state contribution refers to the long-term pattern contribution component at feature nodes after gated diffusion from the long-phase layer, followed by secondary gated diffusion and fusion, emphasizing slow accumulation and stable preferences. An example is the stable weight of "policy maintenance" and "renewal consultation" on the whole-month billing cycle dimension. Burst contribution refers to the short-term pulse contribution component at feature nodes after gated diffusion from the short-phase layer, followed by secondary gated diffusion and fusion, emphasizing rapid peaks and concentrated responses. An example is the peak weight of "marketing clicks" and "insurance calculation requests" within a two-hour activity window. The explanatory chain refers to the minimum traceable path set composed of user nodes, feature nodes, and phase nodes, used to explain the source and path of each feature contribution in the user profile vector. An example is the steady-state path of "billing cycle phase → user node A → feature renewal consultation" and the burst path of "activity phase → user node A → feature marketing click" together forming the explanatory chain. The target user profile vector refers to the final profile representation output after causal auditing and counterfactual replay, and the removal of spurious features. It explicitly includes steady-state contribution, burst contribution, and explanatory chains, and is directly used for risk assessment, targeted outreach, and compliance auditing. For example, the final vector for a user retains a high steady-state contribution from "renewal consultation" and a medium burst contribution from "marketing clicks," along with two corresponding explanatory chains for traceability.

[0136] Furthermore, using user profile vectors as input, a feature dependency graph is constructed by combining steady-state contribution channels, burst contribution channels, residual compensation channels, and path interpretation records. First, the set of features already held by user nodes is used as the node set, and the co-occurrence relationship of the same path and the intra-channel collaboration strength in the path interpretation records are used as the source of edge weights to obtain the adjacency matrix of the weighted undirected graph. This matrix is ​​used for statistical testing and directed processing of subsequent causal audits. The specific calculation is as follows:

[0137]

[0138]

[0139] in, Represents feature nodes, Represents the edge weights of feature pairs. This represents the common path weighting coefficient of the three channels. Indicates the channel type, with values ​​of long-phase channel, short-phase channel, or residual channel. and Indicates features in the channel The following path interpretation set, Representing a path In features The path on the surface affects the intensity. This represents a stable term. The calculation principle is to characterize the cooperative tightness of feature pairs in the same explanatory chain using weighted common path overlap, and to obtain the initial edge weights of the feature dependency graph by convex combination of different channels.

[0140] In the causal audit process, conditional independence detection and temporal consistency verification are performed on the feature dependency graph to obtain a causal verification graph. First, the temporal phase slice is projected onto the feature layer to construct binary active variables and conditional sets. Conditional mutual information is used as the test statistic, and spurious dependencies are screened out using threshold rules. Simultaneously, the direction is determined by the temporal partial order relationship. The specific calculations are as follows:

[0141]

[0142]

[0143]

[0144] in, Indicates in the condition set Lower features With features Conditional mutual information estimation, This represents the empirical probability after additive smoothing. The chi-square statistic represents the approximation of the likelihood ratio. Indicates the sample count. Indicates the significance threshold. This represents the median time difference under the constraints of path interpretation and time phase. Indicates the timestamp of the corresponding event. This represents the time phase window. The calculation principle is to test the significance of residual dependencies using conditional mutual information under the condition of controlling for common antecedent factors, and to assign directions to edges using the sign of the time difference, thereby obtaining a causal verification graph that passes both statistical and temporal verification.

[0145] During the counterfactual replay process, an intervention experiment is performed on the feature dependency paths in the causal verification graph to detect whether there are significant changes in the user profile vector output. Based on this, the true contributing paths and pseudo-correlated paths are labeled, and a causal reinforcement graph is generated. Minimal intervention is applied to each candidate path, while keeping the other paths and weights unchanged. The changes in the user profile vector before and after the intervention are compared, and labels are assigned based on a threshold. The specific calculation is as follows:

[0146]

[0147] in, Represents the original user profile vector. This represents the image recalculation function that maintains the aggregation and proportioning mechanisms unchanged. Indicates path Counterfactual intervention operations, Representing a path Triggering characteristics, Indicates the characteristic contribution before correction. Indicates the contribution compared to the baseline. This represents the user profile vector after intervention. Indicates the strength of the counterfactual effect. This represents the significance threshold. The calculation principle is to zero out or set the baseline for a single path and recalculate the output. The absolute magnitude of the output change approximates the causal effect of the path, and then the path is marked as a true contribution or a spurious correlation based on the threshold.

[0148] Based on the causal reinforcement graph, spurious features in the user profile vector are deleted or downweighted, and the steady-state contribution channels, burst contribution channels, and explanatory chains are reorganized. The resulting target user profile vector is corrected by causal auditing and counterfactual playback. Each feature is causally weighted according to the effect strength of its true path set, while contributions supported only by spurious paths are removed. Finally, the channel allocation is updated and the explanatory index is reconstructed. The specific calculations are as follows:

[0149]

[0150]

[0151]

[0152]

[0153] in, Representation of features The set of real contribution paths Representation of features The set of pseudo-related paths, Representing a path The strength of the counterfactual effect The causal reinforcement coefficient representing the feature. Indicates the characteristic contribution after causal correction. This represents the component of the characteristic contribution in the steady-state channel, the burst channel, and the residual channel. This indicates the corrected channel ratio. Represents the target user profile vector. This represents vector concatenation. This represents the stable term. The calculation principle is to proportionally amplify or zero out the feature contribution by using the cumulative effect of the real path, and re-estimate the global proportion using the total energy at the channel level. At the same time, the set of real paths is retained as an index for the explanatory chain, forming a final traceable representation that includes steady-state contribution, burst contribution and explanatory chain.

[0154] This application also provides an insurance user profile generation device based on PageRank and mutual information, referring to... Figure 2 , Figure 2 This is a schematic diagram of a module for generating insurance user profiles based on PageRank and mutual information, provided in an embodiment of this application. The device is a server, comprising an acquisition module 21 and a processing module 22. The acquisition module 21 is used to segment the acquired insurance event stream according to time sequence, generating time phase slice packets containing long-term and short-term phase windows. The processing module 22 is used to construct a multi-layer hypergraph containing user nodes, feature nodes, and phase nodes based on the time phase slice packets, and to perform gated diffusion in the long-term phase layer of the multi-layer hypergraph to output a steady-state diffusion vector, and to perform gated diffusion in the short-term phase layer of the multi-layer hypergraph to output a burst diffusion vector. The processing module 22 is also used to calculate a mutual information tensor based on the time phase slice packets and to introduce the mutual information tensor into... The processing module 22 transmits the hypergraph, combines the steady-state diffusion vector and the burst diffusion vector to perform a bounce-back injection, and outputs a multi-layer hypergraph after the bounce-back injection. The processing module 22 is also used to perform secondary gated diffusion using the multi-layer hypergraph after the bounce-back injection, and transform the steady-state diffusion vector and the burst diffusion vector into feature-level response vectors. The processing module 22 is also used to fuse the feature-level response vectors in the feature neighborhood already held by the user node to generate a user profile vector. The processing module 22 is also used to verify the feature dependencies in the user profile vector through causal auditing and counterfactual replay, remove pseudo-correlated features, and obtain a target user profile vector containing steady-state contribution, burst contribution and explanatory chain.

[0155] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0156] This application also provides an electronic device, with reference to... Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: at least one processor 31, at least one network interface 34, a user interface 33, a memory 35, and at least one communication bus 32.

[0157] The communication bus 32 is used to enable communication between these components.

[0158] The user interface 33 may include a display screen and a camera. Optionally, the user interface 33 may also include a standard wired interface and a wireless interface.

[0159] The network interface 34 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0160] The processor 31 may include one or more processing cores. The processor 31 connects to various parts of the server via various interfaces and lines, executing instructions, programs, code sets, or instruction sets stored in the memory 35, and calling data stored in the memory 35 to perform various server functions and process data. Optionally, the processor 31 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 31 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed on the screen; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 31 and may be implemented as a separate chip.

[0161] The memory 35 may include random access memory (RAM) or read-only memory. Optionally, the memory 35 may include a non-transitory computer-readable storage medium. The memory 35 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 35 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 35 may also be at least one storage device located remotely from the aforementioned processor 31. Figure 3As shown, the memory 35, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for generating insurance user profiles based on PageRank and mutual information.

[0162] exist Figure 3 In the electronic device shown, the user interface 33 is mainly used to provide an input interface for the user and obtain the user input data; while the processor 31 can be used to call an application stored in the memory 35 that is an insurance user profile generation method based on PageRank and mutual information. When executed by one or more processors, the electronic device executes one or more methods as described in the above embodiments.

[0163] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0164] This application also provides a computer-readable storage medium storing instructions. When executed by one or more processors, these instructions cause an electronic device to perform one or more of the methods described in the above embodiments.

[0165] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A method for generating insurance user profiles based on PageRank and mutual information, characterized in that, The method includes: The acquired insurance event stream is segmented in chronological order to generate a time phase slice package containing long-term phase windows and short-term phase windows; A multi-layer hypergraph containing user nodes, feature nodes, and phase nodes is constructed based on the time phase slice package. Gated diffusion is performed in the long phase layer of the multi-layer hypergraph to output a steady-state diffusion vector, and gated diffusion is performed in the short phase layer of the multi-layer hypergraph to output an burst diffusion vector. The mutual information tensor is calculated based on the time phase slice packet, and the mutual information tensor is introduced into the transport graph. The steady-state diffusion vector and the burst diffusion vector are combined to perform a bounce injection, and the multi-layer hypergraph after the bounce injection is output. The multi-layer hypergraph after the bounce-injection is used to perform secondary gated diffusion, transforming the steady-state diffusion vector and the burst diffusion vector into a feature-level response vector. The feature-level response vector is fused within the feature neighborhood already held by the user node to generate a user profile vector; The feature dependencies in the user profile vector are verified by causal auditing and counterfactual replay, and spurious correlation features are eliminated to obtain the target user profile vector containing steady-state contribution, burst contribution and explanatory chain. The process of calculating the mutual information tensor based on the time-phase slice packet, incorporating the mutual information tensor into the transport graph, performing a bounce-back injection in conjunction with the steady-state diffusion vector and the burst diffusion vector, and outputting a multi-layered hypergraph after the bounce-back injection specifically includes: Using the time phase slice package as input, a joint occurrence frequency table of feature pairs is constructed based on the joint distribution of user nodes, feature nodes, and phase nodes in each time phase window, and the corresponding time phase labels are marked in the joint occurrence frequency table of feature pairs. Within the time phase window, the degree of information sharing between the user node and the feature node, as well as between the feature nodes, is calculated to generate a mutual information matrix. The mutual information matrices under different time phase windows are then concatenated along the time dimension to form a mutual information tensor. The mutual information tensor is introduced as a correction factor for the edges into the transmission graph of the multilayer hypergraph, so that the feature transmission probability in the transmission graph is jointly determined by the edge weights of the graph structure and the statistical correlation. In the transport graph after introducing mutual information tensor, the steady-state diffusion vector and the burst diffusion vector are combined to perform a bounce injection, so as to achieve dynamic equilibrium diffusion across time scales in the multi-layer hypergraph and output the multi-layer hypergraph after bounce injection.

2. The insurance user profile generation method based on PageRank and mutual information according to claim 1, characterized in that, The process of segmenting the acquired insurance event stream according to time sequence to generate a time phase slice package containing long-term phase windows and short-term phase windows specifically includes: The insurance event stream is subjected to out-of-order correction and duplicate event cleaning to generate a clean event sequence with a unified timestamp and event tag; A time phase library is constructed based on circadian rhythms, billing cycles, policy milestones, and marketing activity periods. A time phase mapping table is formed by labeling events in the cleaning event sequence with phase tags according to the time phase library. The set of long-term phase windows and the set of short-term phase windows are determined according to the time phase mapping table. The cleaning event sequence is then aggregated within the set of long-term phase windows and the set of short-term phase windows to generate a set of structured slices containing window identifiers, phase labels, coverage start and end, event counts, source distribution and missing percentages. Based on the structured slice set, cross-scale alignment and consistency verification are performed, and the slice set that passes the verification and the bridging record are output. The slice set that passes the verification and the bridging record are then grouped and encapsulated to obtain the time phase slice package.

3. The insurance user profile generation method based on PageRank and mutual information according to claim 1, characterized in that, The process of constructing a multi-layered hypergraph containing user nodes, feature nodes, and phase nodes based on the time-phase slice package, and performing gated diffusion at the long-phase layer of the multi-layered hypergraph to output a steady-state diffusion vector, and performing gated diffusion at the short-phase layer of the multi-layered hypergraph to output an burst diffusion vector, specifically includes: Map the user identifier in the time phase slice package to the user node, map the event feature in the time phase slice package to the feature node, map the phase label in the time phase slice package to the phase node, and establish a unique identifier and attribute description for the user node, the feature node and the phase node respectively; User feature edges are established based on the event attribution relationship between the user node and the feature node, user phase edges are established based on the time coupling relationship between the user node and the phase node, and feature phase edges are established based on the occurrence dependency relationship between the feature node and the phase node. Edge weights are assigned to the user feature edges, the user phase edges, and the feature phase edges respectively to form the multi-layer hypergraph. The phase nodes and associated nodes with long-term phase labels are extracted from the multi-layer hypergraph to form a long-phase subgraph, and gated diffusion is performed in the long-phase subgraph to output the steady-state diffusion vector; The phase nodes and associated nodes with short-time phase labels are extracted from the multi-layer hypergraph to form a short-phase layer subgraph, and gated diffusion is performed in the short-phase layer subgraph to output the burst diffusion vector.

4. The insurance user profile generation method based on PageRank and mutual information according to claim 1, characterized in that, The step of performing secondary gated diffusion using the multi-layer hypergraph after bounce-back injection to transform the steady-state diffusion vector and the burst diffusion vector into a feature-level response vector specifically includes: Using the multi-layered hypergraph after the bounce-injection, the steady-state diffusion vector, and the burst diffusion vector as input, the user node set is determined as the diffusion anchor set, and the diffusion initiation state and anchor edge set are generated. Based on the diffusion initiation state and the anchored edge set, a gating strategy including structure gates, channel gates and phase gates is constructed, and a gating weight configuration is generated; Diffusion under the control of the gated weight configuration is performed in the long phase subgraph of the multi-layer hypergraph to form a steady-state contribution pool, and diffusion under the control of the gated weight configuration is performed in the short phase subgraph of the multi-layer hypergraph to form an burst contribution pool. Cross-gating and conflict resolution are performed on the steady-state contribution pool and the burst contribution pool to generate a channel-level feature contribution table and path interpretation record; The channel-level feature contribution table, the path interpretation record, and the bounce injection residual are combined and scaled to obtain the initial response vector; The initial response vector is subjected to robustness and sparsity processing to output the feature-level response vector.

5. The insurance user profile generation method based on PageRank and mutual information according to claim 4, characterized in that, The step of fusing the feature-level response vector within the feature neighborhood already held by the user node to generate a user profile vector specifically includes: Using the feature-level response vector as input, a set of feature nodes connected to the user node is retrieved to form a user node feature neighborhood, and the corresponding steady-state contribution component, burst contribution component and residual component are extracted from the user node feature neighborhood. The steady-state contribution component, burst contribution component and residual component of each feature node in the feature neighborhood of the user node are concatenated to form a feature multidimensional contribution vector, and a weight allocation table is generated by combining the time phase label and the path interpretation record. An aggregation operation is performed within the user node feature neighborhood based on the weight allocation table to generate a user feature aggregation table. The user feature aggregation table is subjected to consistency correction and interpretability annotation, cross-phase contradictory feature contributions are removed, and an interpretation chain is embedded to output the user profile vector.

6. The insurance user profile generation method based on PageRank and mutual information according to claim 4, characterized in that, The process of verifying the feature dependencies in the user profile vector through causal auditing and counterfactual replay, eliminating spurious correlation features, and obtaining a target user profile vector containing steady-state contributions, burst contributions, and explanatory chains specifically includes: Using the user profile vector as input, a feature dependency graph is constructed by combining the steady-state contribution channel, the burst contribution channel, the residual compensation channel, and the path interpretation record; During the causal audit process, conditional independence detection and temporal consistency verification are performed on the feature dependency graph to obtain a causal verification graph; During the counterfactual replay process, an intervention experiment is performed on the feature dependency path in the causal verification graph to detect whether the user profile vector output changes significantly, and the true contribution path and the pseudo-correlation path are marked accordingly to generate a causal reinforcement graph. Based on the causal reinforcement graph, the pseudo-correlation features in the user profile vector are deleted or downweighted, and the steady-state contribution channel, burst contribution channel and explanation chain are reorganized to output the target user profile vector after causal audit and counterfactual playback correction.

7. An insurance user profile generation device based on PageRank and mutual information, characterized in that, The apparatus is used to execute the insurance user profile generation method based on PageRank and mutual information as described in any one of claims 1 to 6, the apparatus comprising an acquisition module (21) and a processing module (22), wherein, The acquisition module (21) is used to segment the acquired insurance event stream according to the time sequence to generate a time phase slice package containing a long time phase window and a short time phase window; The processing module (22) is used to construct a multi-layer hypergraph containing user nodes, feature nodes and phase nodes according to the time phase slice package, and to perform gated diffusion in the long phase layer of the multi-layer hypergraph to output a steady-state diffusion vector, and to perform gated diffusion in the short phase layer of the multi-layer hypergraph to output an burst diffusion vector. The processing module (22) is also used to calculate the mutual information tensor based on the time phase slice packet, introduce the mutual information tensor into the transmission graph, combine the steady-state diffusion vector and the burst diffusion vector to perform a back-injection, and output the multi-layer hypergraph after the back-injection. The processing module (22) is also used to perform secondary gated diffusion using the multi-layer hypergraph after the bounce injection, and to convert the steady-state diffusion vector and the burst diffusion vector into a feature-level response vector; The processing module (22) is also used to fuse the feature-level response vector within the feature neighborhood already held by the user node to generate a user profile vector; The processing module (22) is also used to verify the feature dependencies in the user profile vector through causal auditing and counterfactual replay, remove pseudo-correlation features, and obtain a target user profile vector containing steady-state contribution, burst contribution and explanatory chain.

8. An electronic device, characterized in that, The electronic device includes a processor (31), a memory (35), a user interface (33), and a network interface (34). The memory (35) is used to store instructions. The user interface (33) and the network interface (34) are both used to communicate with other devices. The processor (31) is used to execute the instructions stored in the memory (35) to cause the electronic device to perform the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Track portrait generation method and system based on artificial intelligence

    CN120448743A

  • Insurance intelligent decision-making engine system based on multi-modal user portraits

    CN120525615A