Operator user portrait construction method and system based on rules and vector features

By using a rule-based and vector feature-based approach, and leveraging graph neural networks and gated fusion networks to process operator user data, the accuracy and generalization ability of user profiling are improved. This solves the problem of low accuracy in existing user profiling technologies and achieves better user behavior adaptability and service support.

CN122022862APending Publication Date: 2026-05-12CHINA UNITED NETWORK COMM CO LTD SOFTWARE RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNITED NETWORK COMM CO LTD SOFTWARE RES INST
Filing Date
2025-12-24
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in user profiling for operators, and the generalization ability of the construction process is low, making it difficult to adapt to changes in user behavior under new models.

Method used

By using rule-based and vector feature-based methods, de-identified user data from operator databases is obtained, structured feature vectors are constructed, and graph neural networks and gated fusion networks are used for feature matching and fusion to determine rule hit information and feature contribution, thereby improving the accuracy and generalization ability of user profiles.

Benefits of technology

It improves the accuracy of operator user profiles and the generalization ability of the construction process, enabling it to better adapt to changes in user behavior and new patterns, and provide precise marketing and service support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122022862A_ABST
    Figure CN122022862A_ABST
Patent Text Reader

Abstract

The invention discloses a rule and vector feature-based operator user portrait construction method and system. The method comprises the following steps of: obtaining desensitized user data based on an operator database, and processing the desensitized user data to obtain a structured feature vector; matching the structured feature vector based on a preset rule base to obtain a rule hit score, constructing a graph neural network based on the structured feature vector, and mapping the vector feature into a vector feature score after determining the vector feature of the user node in the graph neural network; inputting the rule hit score and the vector feature score into a gating fusion network to obtain a fusion score; performing feature attribution analysis and rule path tracking based on the fusion score to determine rule hit information and a feature contribution degree; and the fusion score, the rule hit information, the abstract of the vector feature and the feature contribution degree are written into the user portrait, so that the accuracy of the operator user portrait can be improved, and the generalization ability of the user portrait construction process is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication technology, specifically relating to a method and system for constructing operator user profiles based on rules and vector features. Background Technology

[0002] User profiling, or user information tagging, is crucial for telecom operators. Establishing accurate user profiles is essential for managing existing users, maintaining customer loyalty, and enhancing their value. It can effectively improve the efficiency and quality of product marketing and services.

[0003] Existing solutions for building user profiles for telecom operators largely rely on rule-based tagging systems. While the rules are logically clear, their generalization capabilities are insufficient. When new patterns emerge beyond the predetermined rules, these new patterns can lead to changes in user behavior patterns caused by adjustments in business strategies, changes in product forms, changes in application entry points, or changes in user behavior migration, as well as changes in cross-entity relationship structures or relationship strength distributions. This can result in existing rule bases being unable to cover these patterns in a timely manner, leading to a decrease in the accuracy of user profiles. Existing technical solutions are ill-suited to these new patterns.

[0004] Therefore, how to improve the accuracy of operator user profiles and enhance the generalization ability of the user profile construction process are technical problems that need to be solved by those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to solve the technical problems of low accuracy of operator user profiles and low generalization ability in the existing technology.

[0006] To achieve the above technical objectives, on the one hand, the present invention provides a method for constructing operator user profiles based on rules and vector features, the method comprising: The user data is obtained after being de-identified from the operator's database, and the user data is processed to obtain a structured feature vector. The structured feature vectors are matched based on a preset rule base to obtain a rule hit score. At the same time, a graph neural network is constructed based on the structured feature vectors, and after determining the vector features of user nodes in the graph neural network, the vector features are mapped to vector feature scores. The rule hit score and vector feature score are input into a gated fusion network to obtain a fusion score; Based on the fusion score, feature attribution analysis and rule path tracking are performed to determine rule hit information and feature contribution. The fusion score, rule hit information, vector feature summary, and feature contribution are written into the user profile.

[0007] Furthermore, the rule hit score is obtained by matching the structured feature vector based on a preset rule base, specifically through the following formula: ;

[0008] In the formula, To score points for hitting the rules, Let k be the weight of the k-th rule. For the number of rules, For the hit result of the k-th rule, This is a bias term.

[0009] Furthermore, the construction of the graph neural network based on the structured features specifically involves: using user-related entities as nodes and user-related behaviors as edges to construct a graph structure, wherein the edges record timestamps.

[0010] Furthermore, the vector features are mapped to vector feature scores using the following formula: ;

[0011] In the formula, For vector feature scores, For the weight vector, For transpose, For vector features, This is a bias term.

[0012] Furthermore, the fusion score is specifically determined using the following formula: ; ;

[0013] In the formula, To integrate scores, To integrate weights, To score points for hitting the rules, For vector feature scores, For the gated weight matrix, This is the concatenated vector obtained by concatenating the rule vector and the vector features. For bias terms, This is the rule vector obtained by vectorizing the rules. These are vector features.

[0014] Furthermore, the feature attribution analysis is performed on the fusion score, whereby the features include rule-based, vector features, and key behavioral features. The feature contribution is determined using the following formula: ; ;

[0015] In the formula, The feature contribution of the i-th feature. For a subset that does not contain the i-th feature, For a set containing all features, The function that outputs the fusion score, This is the baseline term.

[0016] Furthermore, the rule-based path tracking of the fusion score specifically includes: Rules based on the fusion score query hit; The feature contribution of the hit rule is determined, as well as the original feature, rule condition, rule label and fusion score corresponding to the hit rule. The feature contribution, original feature, rule condition, rule label and fusion score corresponding to the hit rule are combined into a tracking path. The tracking path is saved to the rule hit information corresponding to the hit rule.

[0017] On the other hand, the present invention also provides a user profile construction system for operators based on rules and vector features, the system comprising: The acquisition module is used to acquire de-identified user data based on the operator's database and process the user data to obtain a structured feature vector. The scoring module is used to match the structured feature vectors based on a preset rule base to obtain a rule hit score. At the same time, it constructs a graph neural network based on the structured feature vectors and maps the vector features to vector feature scores after determining the vector features of user nodes in the graph neural network. The fusion module is used to input the rule hit score and vector feature score into the gated fusion network to obtain the fusion score; The determination module is used to determine the rule hit information and feature contribution based on the fusion score by performing feature attribution analysis and rule path tracking. The profile module is used to write the fusion score, rule hit information, vector feature summary, and feature contribution into the user profile.

[0018] The present invention provides a method and system for constructing operator user profiles based on rules and vector features. Compared with existing technologies, this method includes: obtaining anonymized user data from an operator database and processing the user data to obtain structured feature vectors; matching the structured feature vectors based on a preset rule base to obtain rule hit scores, simultaneously constructing a graph neural network based on the structured features, and mapping the vector features to vector feature scores after determining the vector features of user nodes in the graph neural network; inputting the rule hit scores and vector feature scores into a gated fusion network to obtain a fusion score; performing feature attribution analysis and rule path tracking based on the fusion score to determine rule hit information and feature contribution; and writing the fusion score, rule hit information, vector feature summary, and feature contribution into the user profile. This method can improve the accuracy of operator user profiles and enhance the generalization ability of the user profile construction process. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 The diagram shown is a flowchart illustrating the operator user profile construction method based on rules and vector features provided in the embodiments of this specification. Figure 2 The diagram shown is a structural schematic of the operator user profile construction system based on rules and vector features provided in the embodiments of this specification. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] like Figure 1The diagram illustrates a flowchart of a method for constructing operator user profiles based on rules and vector features, as provided in an embodiment of this specification. While this specification provides the method operation steps or system structure shown in the following embodiments or figures, based on conventional methods or without creative effort, the method or system may include more or fewer operation steps or module units after partial merging. In steps or structures where there is no logically necessary causal relationship, the execution order of these steps or the module structure of the system is not limited to the execution order or module structure shown in the embodiments or figures of this specification. When the method or module structure is applied in actual systems, servers, or terminal products, it can be executed sequentially or in parallel according to the method or module structure shown in the embodiments or figures (e.g., in a parallel processor or multi-threaded processing environment, or even in a distributed processing or server cluster implementation environment).

[0023] The operator user profile construction method based on rules and vector features provided in the embodiments of this specification can be applied to terminal devices such as clients and servers, for example... Figure 1 As shown, the method specifically includes the following steps: Step S101: Obtain de-identified user data based on the operator's database, and process the user data to obtain a structured feature vector.

[0024] Specifically, multimodal user data is obtained from the operator's database, including but not limited to: C1. Call and SMS data: voice call duration, number of calls, caller-to-call ratio, number of SMS messages, etc. C2. Application Usage Logs: Number of app launches, number of page views, search activity, dwell time, button click events, etc. C3. Product ordering data: current package, historical package change records, value-added service ordering status, participation in promotional activities, etc. C4. Basic User Attributes: City / Region of Origin, Term of Service, Terminal Brand and Model, Billing Type, etc.

[0025] After anonymizing the obtained multimodal user data, preprocessing operations such as field alignment, time alignment, anomaly detection, and missing value imputation are performed. Numerical features are scaled using normalization, standardization, and bucketing, while categorical features are transformed using one-hot encoding, target encoding, or embedding encoding to obtain a unified structured feature vector, as shown in the following formula: ;

[0026] In the formula, d is the feature dimension.

[0027] Step S102: Match the structured feature vectors based on the preset rule base to obtain the rule hit score. At the same time, construct a graph neural network based on the structured feature vectors, and after determining the vector features of user nodes in the graph neural network, map the vector features to vector feature scores.

[0028] Specifically, a rule engine can be used to match structured feature vectors. The feature vector x is input into the rule engine. Multiple business rules are pre-configured in the rule base, for example: D1. If "the average monthly consumption in the past three months is greater than threshold 1 and the total voice call duration is greater than threshold 2", then the user is labeled as a "high-value user". D2. If “the amount of recharged has been declining continuously in recent months and the number of complaints has increased”, then the user will be labeled as a “potential churned user”.

[0029] The rule engine performs rule matching and condition evaluation for each user, and outputs a rule feature vector: , in For the number of rules, Indicates the first The hit result or hit degree of a rule can take the value 0 or 1, or the value in the range [0,1].

[0030] In this embodiment of the application, the rule hit score is obtained by matching the structured feature vector based on a preset rule base, specifically through the following formula: ;

[0031] In the formula, To score points for hitting the rules, Let k be the weight of the k-th rule. For the number of rules, For the hit result of the k-th rule, For bias terms, weight parameters It can be set based on business experience, or it can be trained in a data-driven manner when there are labeled samples, so as to reflect the relative importance of different rules to the target task.

[0032] The construction of a graph neural network based on the structured feature vectors specifically involves: using user-related entities as nodes and user-related behaviors as edges to construct a graph structure. Nodes include user nodes (carrying features such as city, network duration, current plan, average monthly call cost, and terminal brand); product nodes (carrying features such as product category, price tier, whether on promotion, and whether a bundled plan is available); and application nodes (carrying features such as APP category, function tags, and penetration rate level). Edges include: user-user call edges (with attributes such as call duration, number of calls, and cost); user-product order edges (with attributes such as order time, order status, and usage duration); and user-application interaction edges (with attributes such as usage duration, number of launches, and last usage time).

[0033] The edges contain timestamps. Used for subsequent time encoding, the node type set is denoted as The set of edge types is Each directed edge The meta-relation is represented as a triple. ,in Mapping for node types, Mapping for edge types.

[0034] The graph neural network encoding module adopts an L-layer heterogeneous graph Transformer (HGT) structure, with each layer including: A21. Linear projection related to node type (Query / Key / Message); A22. Multi-head attention computation based on meta-relations; A23. Message transformation based on edge type; A24. Target-node-oriented aggregation and type-specific mapping; A25, Residual Connections and Nonlinear Activation.

[0035] Let the first Layer user node The representation of is . For example, in the l-th layer, for each attention head... , Represented as nodes within the layer. Here, 'a' represents the final user vector feature, 'h' represents the attention head number, and 'h' represents the number of attention heads.

[0036] A21, Type-Specific Projection (Node Level)

[0037] For the source node (The commas in the subscripts represent index tuples, indicating that parameters are selected by node type, layer number, or header number.) ;

[0038] For the target node : ;

[0039] in These are the learnable matrices related to node type, layer number, and header number, respectively. for, For the first Layer user node The expression, , The first A vector of key, value, and query elements.

[0040] A22. Multi-head attention computation based on meta-relations

[0041] For the edge and its edge types Define the score for the a-th head: ;

[0042] in This is the edge type correlation matrix. This is the importance coefficient tensor of the meta-relation, used to distinguish the prior importance of different types of relations. For opposite sides and its edge types , No. Layer Score based on height.

[0043] The attention weights are obtained by performing softmax normalization on all neighbors u of the same target node v. ;

[0044] A23. Message Transformation Based on Edge Type

[0045] The message vector of the a-th head corresponding to edge type r: ;

[0046] in This is a message transformation matrix bound to the edge type, used to perform differentiated linear transformations on different business relationships (calls, orders, APP usage).

[0047] A24. Target Node-Oriented Aggregation and Type-Specific Mapping

[0048] Aggregate messages from all relations r and all neighbors u according to their attention weights: ;

[0049] By piecing together all the heads, we get: ;

[0050] A25, Residual Connectivity and Nonlinear Activation

[0051] Restore to a uniform dimension using a linear transformation related to the node type, and add residuals: ;

[0052] in For type correlation matrix, Use ReLU or GELU activation functions. For nodes In the The output of the layer represents a vector.

[0053] By stacking L layers of HGT units, the user's L-jump high-order neighborhood information can be encoded into the final representation. middle, That is to say h in the text.

[0054] To characterize the temporal evolution of user behavior, a relative time encoding vector is introduced. ,in, This represents the time difference between the target node v and the source node u.

[0055] Constructing the basic time code: ;

[0056] In the formula, Base(ΔT) represents the difference between the relative time differences. The resulting basic time-coded vector (sine and cosine positional coding form) is constructed.

[0057] After a linear transformation, we obtain: ;

[0058] Finally, it is added to the source node representation: ;

[0059] Use in attention and message computation Alternative This allows for the modeling of behavioral temporal sequences, where... To represent the learnable parameter matrix (temporal projection matrix) that linearly maps the underlying temporal code, used to... Mapped to the temporal embedding of the same dimension as the node representation.

[0060] The HGT layers are concatenated in series according to the L layers, and the layers are connected through the aforementioned residual connections. This represents the final representation of all user nodes. It is provided to the fusion module through an interface.

[0061] Specifically, vector features are mapped to vector feature scores using the following formula: ;

[0062] In the formula, For vector feature scores, For the weight vector, For transpose, For vector features, This is a bias term.

[0063] Step S103: Input the rule hit score and vector feature score into the gated fusion network to obtain the fusion score.

[0064] In this embodiment of the application, the fusion score is specifically determined by the following formula: ; ;

[0065] In the formula, To integrate scores, To integrate weights, To score points for hitting the rules, For vector feature scores, For the gated weight matrix, This is the concatenated vector obtained by concatenating the rule vector and the vector features. For bias terms, This is the rule vector obtained by vectorizing the rules. These are vector features.

[0066] Specifically, when near When the overall score relies more on rule scores, the system output is closer to traditional rule-based decisions; when near In this case, the overall score relies more heavily on the vector score, and the system output makes fuller use of the expressive power of the deep model. By jointly training the overall model (including the representation learning layer and the gating fusion layer), the gating network can automatically learn appropriate fusion strategies for different samples and scenarios.

[0067] In risk control and compliance scenarios, regularization terms or prior constraints can be added to make g more biased towards larger values ​​during training, allowing the rule part to dominate. In recommendation and marketing scenarios, the constraints on g can be relaxed, allowing the model to rely more on the latent pattern mining capabilities of the vector part.

[0068] Specifically, the gated converged network adopts a structure consistent with existing technologies, including: a rules engine, a converged core (gated network), and hybrid services.

[0069] B1 Rule Engine and Rule Vectors: B11. The rule engine maintains a set of configurable rules. Each rule includes a name, weight, and condition expression; B12. Apply all rules sequentially to the raw feature dictionary (raw_features) of a single user to obtain... Then use The rule hit score is obtained, and the rule details (rule_details) record the contribution of each rule.

[0070] B2 gated converged network: The structure of the FusionCore core is as follows: B21. Embedding and concatenating rule vectors and vector features: ;

[0071] in, These are vector features.

[0072] B22. Calculate the gating value, i.e., the fusion weight, through the linear layer: ;

[0073] in, , which represents "gated fusion weights / rule-side weights", used to balance rule scores and vector scores in the final score.

[0074] B23. Calculate vector scores using a two-layer fully connected network. That is, vector feature scores : ;

[0075] B24. Final Portrait Score: ;

[0076] In its implementation, FusionCore.forward(h_rule,h_nn,f_rule) returns (score,gate,f_nn), which correspond to the three quantities mentioned above. This is also known as rule hit score, or rule score.

[0077] The training process for graph neural networks and gated fusion networks includes: C21, Subgraph Sampling.

[0078] On the full heterogeneous graph, heterogeneous subgraph sampling is performed on each batch of user sets to make the number of each node type in the subgraph as balanced as possible, thus ensuring training stability. The HGSampling approach can be referenced.

[0079] C22, Forward computation.

[0080] For each batch: The HGT module is used to perform L-level message passing on the sampled subgraph to obtain the batch user embeddings. ; Calculate the raw_features of users in the same batch using the rule engine. and ; The above results are fed into FusionCore to obtain... Gating value Sum of vector scores .

[0081] C23. Loss function design.

[0082] In binary classification tasks (such as churn prediction), cross-entropy loss can be used: ;

[0083] In the formula, Here, the task loss function is cross-entropy loss, used in binary classification tasks, to measure the difference between the model's predictions and the true labels. This represents the number of samples (or user nodes) participating in training within a batch. These are the actual labels for the samples / users.

[0084] To control gating behavior and model stability, regularization terms can be introduced: C231. Gating constraint: Prevents the gating value from being excessively biased to one side.

[0085] ;

[0086] In the formula, This is a gating constraint regularization term used to constrain the gating coefficient g(v) from being excessively biased towards 0 or 1 over a long period, thereby preventing the fusion result from being completely dominated by a single path (rule or vector). The desired gating average value, for example, 0.5; C232, Embedding Smoothing Term: Constrains user embeddings to not change too much between adjacent training rounds.

[0087] Total loss : ;

[0088] In the formula, For gating regularization terms The weighting coefficients are used to adjust the influence of gating constraints on the total loss. For smoothing regularization terms The weighting coefficients are used to adjust the influence of the embedding smoothing constraint on the total loss. The embedding smoothing regularization term is used to constrain the variation of the same user's embedding vector in adjacent training rounds (or adjacent time windows) to prevent it from changing too much, thereby improving training stability and suppressing representation drift.

[0089] C24, Backpropagation and Parameter Update.

[0090] C241. Use AdamW or the Adam optimizer to jointly update parameters such as HGT layer, FusionCore, and time coding.

[0091] C242, learning rate, for example The batch size is, for example, 512 to 1024, the number of HGT layers is L=2 to 3, the number of attention heads is h=4 to 8, and the hidden dimension is d=128 or 256.

[0092] C25. Verification and Early Termination.

[0093] C251. Use more recent data as the validation set to monitor metrics such as AUC and F1.

[0094] C252. If the validation set metrics no longer improve within several rounds, stop training and roll back to the best model.

[0095] In specific application scenarios, the settings for the execution process of graph neural networks and gated fusion networks also include the following: A1 default hyperparameters and data shape.

[0096] Hidden Dimensions: Number of heads to focus on: Single-head dimension: Number of floors: ; Node and edge input tensor: for any node Node input features: , No. Layer node representation: ,initialization: For any directed edge Edge type: (Discrete ID), edge attribute vector: Time information: or (Numerical value, subsequently encoded as a vector).

[0097] A2 Parameter List and Shapes (grouped by "Layer l, Head a, Type / Relationship").

[0098] For clarity, the following parameters are indexed by "node type τ", "edge type r", "layer number l", and "head number a". A dedicated Q / K / V projection matrix is ​​used for each node type: for each layer... Each Each node type : , , ,in, , and Node types In the Layer, First The Query / Key / Value projection matrix of each attention head is used to linearly map the node's previous layer representation (128 dimensions) to the head's subspace (32 dimensions), i.e. Mapped to Relationship Attention Matrix (Edge Type Enters Attention Relationship Layer): For each layer Each Each edge type , , For the first Layer, First Size, edge type (relationship) Relational attention transformation matrix; relational message transformation matrix (edge-type message transformation): for each layer Each Each edge type , , For the first Layer, First Size, side type Relationship message transformation matrix; edge attribute mapping matrix (mapping 128-dimensional edge attributes to a single-head 32-dimensional matrix): for each layer Each Each edge type (Optional: Distinguish by relationship) , For the first Layer, First Size, side type Edge attribute mapping matrix; Gated network parameters (gated fusion: message + edge attribute): for each layer Each Input concatenation dimension: , , Here, "dimensional gating" is used (outputting a 32-dimensional gated vector). For further simplification, the gating can be replaced with a scalar: , To hide the dimension for a single head, and The first Layer, First Weight matrix and bias vector of each size The weight matrix is ​​for scalar gating; the output mapping matrix (maps the 128-dimensional aggregated vector after multi-head concatenation back to 128 dimensions): for each layer l, each node type τ: , , and Node types In the The layer's output mapping matrix and bias.

[0099] "Connection order and tensor shape" of A3 single-layer computation (layer l).

[0100] This section is from arrive The complete link.

[0101] For each edge and each head Execute A3.1 through A3.5 in parallel; then process each node... Execute A3.6 to A3.8.

[0102] A3.1 Node Projection: Calculate Query / Key / Value (project the start / end point separately); For each head : Target node Query: enter: Output: ;but: ;

[0103] Source node Key / Value: enter: Output: , ,but ; ;

[0104] Connection relationships: .

[0105] A3.2 Relational Attention Scoring: Edge type r participates (resulting in scalar e); opposite side ,head : enter: , Edge type Output: ,but: ;

[0106] Connection relationships: .

[0107] A3.3 Softmax Normalization: Same target node Normalization on the same neighbor set (to obtain) ); For fixed target nodes Fixed relationship , fixed head : enter: Output: and ,but: ;

[0108] Connection relationships: (Normalization range: same) ,same Same head ).

[0109] A3.4 Relational Message Transformation: Obtain the attention message vector m (32-dimensional); opposite side ,head : enter: , , Output: ,but: ;

[0110] Connection relationships: .

[0111] A3.5 Edge Attribute Gating Fusion: Injecting edge attribute vectors into messages (output) ); opposite side ,head : (1) Edge attribute mapping (128→32): enter: Output: ,but ;

[0112] In the formula, For the edge The original edge attribute vector (edge ​​features). For the first Layer, First Under each attention head, the edge attribute mapping matrix For the original edge attributes The edge attribute embedding after linear mapping (32-dimensional per head) is used to participate in subsequent gating fusion and aggregation calculations together with the message vector.

[0113] (2) Gated vector calculation (output 32 dimensions): Concatenated input: Output: ,but: ;

[0114] In the formula, For the first Layer, First A focus of attention on the side The gating vector (dimensional gating coefficient). For the first Layer, First The weight matrix of the gating network for each attention head. This is the bias vector (32-dimensional) corresponding to the above gated linear transformation.

[0115] (3) Fusion output (final edge message, 32-dimensional): ;

[0116] Connection relationships: and Output after "mapping → gating → fusion" And serve as input to the aggregation module.

[0117] A3.6 Head-in-Head Aggregation: Summing all incoming edges of the target node v (resulting in...) ); For fixed nodes ,head : enter: Output: ,but: ;

[0118] In the formula, For the first Layer, First Under one's attention, on the side The final fused message vector (32-dimensional). For the first Layer, First Target node under attention The aggregate representation (32-dimensional). For nodes The set of incoming edges ending at (i.e., all edges satisfying) (the set of edges).

[0119] Connection relationship: Same node All →(Σ)→ .

[0120] A3.7 Multi-head splicing: Obtaining a 128-dimensional aggregated vector ; enter: Output: ,but: ;

[0121] A3.8 Output Mapping + Residual + Activation: Obtain the output of this layer. ; For nodes : enter: The next layer Output: ,but: ; ;

[0122] Connection relationships: , and then with Residual summation output .

[0123] The specific meanings of A4 cross-layer connections (vertical stacking) and neighbor connections (horizontal parallelism).

[0124] A4.1 Horizontal parallelism (within the same layer); For the same node its many neighbors Side message In A3.6, a unified summation is performed.

[0125] A3.2 to A3.5 of each edge are calculated independently (in parallel); A3.6 Perform an aggregation (summary).

[0126] A4.2 Vertical stacking (across layers); Layer Layer 0 output As the input of layer 1, the output of layer 1 As the final source of user vectors, ultimately: .

[0127] A5 time information enters the link.

[0128] Specifically, this involves incorporating time encoding into the edge attribute vector: A5.1 time-coded vector; enter: (Scalar or Discrete bucket), Output: .

[0129] A5.2 merged into edge attributes; ;

[0130] In the formula, For the edge The time-encoded vector is used to represent the source node. With the target node The time interval information corresponding to the behavior / interaction, merged Directly enters the edge attribute gating fusion branch of A3.5, affecting... , and .

[0131] A6 From End-User Vectors to Fusion Scores: Parameter Shape and Connection Order.

[0132] A6.1 VectorScoreHead. enter: ,parameter: , Output: ,but: .

[0133] A6.2 RuleEngine; Input: Rule hit vector (or ),parameter: , Output: .but: ;

[0134] A6.3 FusionGate; enter: ,parameter: , Output: ,but: ;

[0135] A6.4 fused output; Output: ; ;

[0136] In summary, the entire link sequence is as follows: initialization: (All nodes); Layer 0: Calculated according to A3.1 to A3.8 ; Layer 1: Calculated according to A3.1 to A3.8 ; Get the target user vector: ; calculate : VectorScoreHead; calculate Original structural features Rule engine; Computational gating And merge: obtain .

[0137] Step S104: Based on the fusion score, perform feature attribution analysis and rule path tracking to determine rule hit information and feature contribution.

[0138] In this embodiment of the application, the feature attribution analysis is performed on the fusion score. The features include rule-based features, vector features, and key behavioral features. The feature contribution is determined by the following formula: ; ;

[0139] In the formula, The feature contribution of the i-th feature. For a subset that does not contain the i-th feature, For a set containing all features, The function that outputs the fusion score, This is the baseline term.

[0140] Specifically, This indicates the output fusion score under fixed parameter conditions. The function takes a feature set of user samples as input (including rule-based features and vector-based features) and outputs a fusion score. .in Indicates using only a subset of features The expected value of the fusion score output when (other features are masked or replaced with baseline values), Indicates that in the existing feature set Introducing features based on This is the marginal contribution to the model output. The aforementioned coefficients are used to weight all features according to the order in which they are added, ensuring fairness in contribution distribution.

[0141] Rules based on the fusion score query hit; The feature contribution of the hit rule is determined, as well as the original feature, rule condition, rule label and fusion score corresponding to the hit rule. The feature contribution, original feature, rule condition, rule label and fusion score corresponding to the hit rule are combined into a tracking path. The tracking path is saved to the rule hit information corresponding to the hit rule.

[0142] Specifically, during the execution of the rules engine, the system records: E21. Whether each rule is met; E22. The original features and threshold conditions upon which each rule depends; E23. Logical combination relationships (AND, OR, NOT, etc.) between multiple hit rules.

[0143] Based on the above information, a decision path can be constructed from "original features → rule conditions → rule labels → comprehensive score". Combined with the Shapley value attribution results, business personnel can clearly understand: E24. What rules primarily drive a given profile result? E25, how much do the rule-based part and the vector part contribute to the overall score respectively; E26. Which behavioral characteristics played a key role in this decision-making process?

[0144] This invention organizes feature contribution, rule hit information, fusion weight and other content into a structured interpretation result and visualizes it for business personnel for auditing and decision support.

[0145] Step S105: Write the fusion score, rule hit information, vector feature summary, and feature contribution into the user profile.

[0146] Specifically, the system integrates scores, rule hit information, vector features, and feature contribution into a unified user profile database. This database is provided as a service through a unified interface layer, offering real-time or near real-time online interfaces to provide user profile query capabilities for business systems such as precision marketing, personalized recommendations, customer retention, and high-risk churn warnings. Furthermore, the profile results can be associated with the model and rule versions used for storage, supporting subsequent result traceability and compliance auditing.

[0147] Based on the above-described method for constructing operator user profiles based on rules and vector features, one or more embodiments of this specification also provide a platform or terminal for constructing operator user profiles based on rules and vector features. This platform or terminal may include a system, software, modules, plug-ins, servers, clients, etc., using the methods described in the embodiments of this specification, combined with necessary implementation hardware. Based on the same innovative concept, the systems in one or more embodiments provided in this specification are as described in the following embodiments. Since the implementation schemes and methods for solving the system problem are similar, the specific system implementations in the embodiments of this specification can refer to the implementation of the aforementioned methods. Repeated descriptions will not be repeated. The terms "unit" or "module" used below can refer to a combination of software and / or hardware that achieves a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, hardware implementations, and a combination of software and hardware, are also possible and contemplated.

[0148] Specifically, Figure 2 This is a schematic diagram of the module structure of an embodiment of the operator user profile construction system based on rules and vector features provided in this specification, as shown below. Figure 2 As shown, the operator user profile construction system based on rules and vector features provided in this specification includes: The acquisition module 201 is used to acquire de-identified user data based on the operator's database, and process the user data to obtain a structured feature vector; The score module 202 is used to match the structured feature vector based on a preset rule base to obtain a rule hit score, and at the same time construct a graph neural network based on the structured feature vector, and map the vector features to vector feature scores after determining the vector features of user nodes in the graph neural network. The fusion module 203 is used to input the rule hit score and vector feature score into the gated fusion network to obtain a fusion score; The determination module 204 is used to determine the rule hit information and feature contribution based on the fusion score by performing feature attribution analysis and rule path tracking. The profile module 205 is used to write the fusion score, rule hit information, vector feature summary and feature contribution into the user profile.

[0149] It should be noted that the system described above may include other implementation methods based on the description of the corresponding method embodiments. The specific implementation methods can be referred to the description of the corresponding method embodiments above, and will not be elaborated here.

[0150] This application also provides an electronic device, including: processor; Memory used to store the processor's executable instructions; The processor is configured to perform the methods provided in the embodiments described above.

[0151] The electronic device provided in this application embodiment stores executable instructions of the processor in a memory. When the processor executes the executable instructions, it can obtain de-identified user data based on the operator's database and process it to obtain a structured feature vector. It then matches the structured feature vector based on a preset rule base to obtain a rule hit score. Simultaneously, it constructs a graph neural network based on the structured feature vector and maps the vector features to vector feature scores after determining the vector features of user nodes in the graph neural network. The rule hit score and vector feature score are input into a gated fusion network to obtain a fusion score. Based on the fusion score, feature attribution analysis and rule path tracking are performed to determine rule hit information and feature contribution. Finally, the fusion score, rule hit information, vector feature summary, and feature contribution are written into a user profile, which can improve the accuracy of the operator's user profile and enhance the generalization ability of the user profile construction process.

[0152] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0153] The methods or systems described in the embodiments provided in this specification can implement business logic through computer programs and record it on a storage medium. The storage medium can be read and executed by a computer to achieve the effects of the solutions described in the embodiments of this specification, such as: The user data is obtained after being de-identified from the operator's database, and the user data is processed to obtain a structured feature vector. The structured feature vectors are matched based on a preset rule base to obtain a rule hit score. At the same time, a graph neural network is constructed based on the structured feature vectors, and after determining the vector features of user nodes in the graph neural network, the vector features are mapped to vector feature scores. The rule hit score and vector feature score are input into a gated fusion network to obtain a fusion score; Based on the fusion score, feature attribution analysis and rule path tracking are performed to determine rule hit information and feature contribution. The fusion score, rule hit information, vector feature summary, and feature contribution are written into the user profile.

[0154] The storage medium can include a physical system for storing information, typically digitizing the information and then storing it using media employing electrical, magnetic, or optical methods. The storage medium can include: systems that store information using electrical energy, such as various types of memory, like RAM and ROM; systems that store information using magnetic energy, such as hard disks, floppy disks, magnetic tapes, magnetic core memory, bubble memory, and USB flash drives; and systems that store information using optical methods, such as CDs or DVDs. Of course, there are other readable storage media, such as quantum memories and graphene memories.

[0155] The embodiments in this specification are not limited to conforming to industry communication standards, standard computer resource data update and data storage rules, or the situations described in one or more embodiments of this specification. Slightly modified implementations based on certain industry standards or custom methods or embodiments can also achieve the same, equivalent, or similar, or predictable, implementation effects as described above. Embodiments that utilize these modified or modified methods for data acquisition, storage, judgment, and processing still fall within the scope of optional implementations of the embodiments in this specification.

[0156] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs SC8051F320. Memory controllers can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the system included therein for implementing various functions can also be considered a structure within the hardware component. Alternatively, the system for implementing various functions can be considered as both a software module implementing the method and a structure within the hardware component.

[0157] The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or plug-ins may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0158] These computer program instructions can also be loaded onto a computer or other programmable resource data updating device, causing a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0159] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0160] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A method for constructing operator user profiles based on rules and vector features, characterized in that, The method includes: The user data is obtained after being de-identified from the operator's database, and the user data is processed to obtain a structured feature vector. The structured feature vectors are matched based on a preset rule base to obtain a rule hit score. At the same time, a graph neural network is constructed based on the structured feature vectors, and after determining the vector features of user nodes in the graph neural network, the vector features are mapped to vector feature scores. The rule hit score and vector feature score are input into a gated fusion network to obtain a fusion score; Based on the fusion score, feature attribution analysis and rule path tracking are performed to determine rule hit information and feature contribution. The fusion score, rule hit information, vector feature summary, and feature contribution are written into the user profile.

2. The method for constructing operator user profiles based on rules and vector features as described in claim 1, characterized in that, The rule hit score is obtained by matching the structured feature vector based on a preset rule base, specifically using the following formula: ; In the formula, To score points for hitting the rules, Let k be the weight of the k-th rule. For the number of rules, For the hit result of the k-th rule, This is a bias term.

3. The method for constructing operator user profiles based on rules and vector features as described in claim 1, characterized in that, The construction of the graph neural network based on the structured feature vector specifically involves: using user-related entities as nodes and user-related behaviors as edges to construct a graph structure, wherein the edges record timestamps.

4. The method for constructing operator user profiles based on rules and vector features as described in claim 3, characterized in that, Specifically, vector features are mapped to vector feature scores using the following formula: ; In the formula, For vector feature scores, For the weight vector, For transpose, For vector features, This is a bias term.

5. The method for constructing operator user profiles based on rules and vector features as described in claim 1, characterized in that, The fusion score is determined using the following formula: ; ; In the formula, To integrate scores, To integrate weights, To score points for hitting the rules, For vector feature scores, For the gated weight matrix, This is the concatenated vector obtained by concatenating the rule vector and the vector features. For bias terms, This is the rule vector obtained by vectorizing the rules. These are vector features.

6. The method for constructing operator user profiles based on rules and vector features as described in claim 5, characterized in that, The feature attribution analysis performed on the fusion score includes rule-based, vector-based, and key behavioral features. The feature contribution is determined using the following formula: ; ; In the formula, The feature contribution of the i-th feature. For a subset that does not contain the i-th feature, For a set containing all features, The function that outputs the fusion score, This is the baseline term.

7. The method for constructing operator user profiles based on rules and vector features as described in claim 6, characterized in that, The rule-based path tracking of the fusion score specifically includes: Rules based on the fusion score query hit; The feature contribution of the hit rule is determined, as well as the original feature, rule condition, rule label and fusion score corresponding to the hit rule. The feature contribution, original feature, rule condition, rule label and fusion score corresponding to the hit rule are combined into a tracking path. The tracking path is saved to the rule hit information corresponding to the hit rule.

8. A user profiling system for telecom operators based on rules and vector features, characterized in that, The system includes: The acquisition module is used to acquire de-identified user data based on the operator's database and process the user data to obtain a structured feature vector. The scoring module is used to match the structured feature vectors based on a preset rule base to obtain a rule hit score. At the same time, it constructs a graph neural network based on the structured feature vectors and maps the vector features to vector feature scores after determining the vector features of user nodes in the graph neural network. The fusion module is used to input the rule hit score and vector feature score into the gated fusion network to obtain the fusion score; The determination module is used to determine the rule hit information and feature contribution based on the fusion score by performing feature attribution analysis and rule path tracking. The profile module is used to write the fusion score, rule hit information, vector feature summary, and feature contribution into the user profile.