A dual verification customer group diffusion method fusing graph structure and semantic concept

By constructing a heterogeneous information graph and performing path coupling analysis, the problem of insufficient differentiation of the importance of user association paths in existing technologies is solved, achieving high-precision customer diffusion and providing interpretable evidence and business credibility.

CN121542478BActive Publication Date: 2026-04-17ZHEJIANG FULIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG FULIN TECH CO LTD
Filing Date
2026-01-21
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing methods for calculating customer diffusion cannot effectively integrate users' long-term semantic features, real-time dynamic intent data, and geospatial information, resulting in limited accuracy and depth of data analysis. Furthermore, they cannot distinguish the importance of different types of related paths, and the output results lack interpretability and accuracy.

Method used

By constructing a heterogeneous information graph, combining user nodes and intermediate nodes, constructing high-order relationship edges, using the graph data processing module to generate high-order embedding vectors, calculating the similarity score between candidate users and seed customer groups, and performing double verification through the path coupling analysis module to determine the target diffusion customer group.

Benefits of technology

It significantly improves the signal-to-noise ratio of interest similarity measurement, enhances graph connectivity and data processing accuracy, provides interpretable evidence and commercial credibility, and ensures the high purity and accuracy of the diffused customer base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542478B_ABST
    Figure CN121542478B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of information retrieval, in particular to a dual-verification guest group diffusion method fusing a graph structure and semantic concepts, which comprises the following steps: collecting multi-source heterogeneous data of candidate users, and constructing a heterogeneous information graph; the heterogeneous information graph comprises user nodes and intermediate nodes for establishing association, and initial feature vectors are configured for the user nodes; the heterogeneous information graph is processed, the initial feature vectors of the user nodes are aggregated, and high-order relationship information and structural positions in the heterogeneous information graph are generated to generate high-order embedding vectors; a center vector is calculated based on the high-order embedding vectors of various sub-users in a seed guest group; a similarity score between the high-order embedding vectors of the candidate users and the center vector is calculated; a path coupling analysis module is constructed, the heterogeneous information graph is searched according to the similarity score, and an interpretable high-order path between the candidate users and the seed guest group is quantified; and the application is rechecked based on the interpretable high-order path to determine a target diffusion guest group.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information retrieval technology, specifically to a dual-verification customer diffusion method that integrates graph structure and semantic concepts. Background Technology

[0002] Against the backdrop of rapid development in digital information processing and network technology, analyzing massive amounts of user data to identify potential target groups—a process known as customer diffusion computation—has become an important application in the field of data processing. Existing methods for customer diffusion computation mainly face several technical challenges. On the one hand, they focus on using user profile data or semantic tags for vector similarity calculation. For example, they retrieve similar data by analyzing users' static attributes, interest tags, or query logs. On the other hand, they attempt to utilize graph data structures, such as graph models constructed from social network topologies or user behavior logs. For example, they perform graph traversal or recommendation based on collaborative filtering algorithms, or perform multi-hop path retrieval within the graph.

[0003] However, such methods often overlook the complex relationships between data nodes, resulting in limited accuracy and depth of data analysis. It is difficult to design effective fusion mechanisms to handle multi-dimensional data inputs such as users' long-term semantic features, real-time dynamic intent data, and geospatial information. At the same time, they cannot distinguish the importance of different types of association paths in the graph structure. For example, they lack adaptive weighting mechanisms for nodes or edges with different attributes.

[0004] In summary, the output of existing computing systems often lacks interpretability and fails to provide clear data traceability paths and correlation logic. This can lead to the final data processing results containing a large number of pseudo-similar data points, affecting the accuracy and reliability of the entire data processing flow.

[0005] To address this, a dual-verification customer diffusion method integrating graph structure and semantic concepts is proposed. Summary of the Invention

[0006] The purpose of this invention is to provide a dual-verification customer group diffusion method that integrates graph structure and semantic concepts. By using the similarity score between the higher-order embedding vector and the center vector of candidate users, the heterogeneous information graph is retrieved, and the interpretable higher-order path between candidate users and seed customer groups is quantified. Based on the interpretable higher-order path, the target diffusion customer group is determined through verification.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A dual-verification customer diffusion method integrating graph structure and semantic concepts includes:

[0009] Collect multi-source heterogeneous data of candidate users and construct a heterogeneous information graph; the heterogeneous information graph includes user nodes and intermediate nodes for establishing associations, the intermediate nodes include opinion leader nodes, product nodes, content tag nodes and geographic information nodes; based on user nodes and intermediate nodes, construct high-order relationship edges representing the structural similarity between users; configure initial feature vectors for user nodes, the initial feature vectors include semantic information representing user profiles;

[0010] The graph data processing module processes the heterogeneous information graph, aggregates the initial feature vectors of user nodes, as well as their structural positions and higher-order relationship information in the heterogeneous information graph, and generates a higher-order embedding vector for each user node.

[0011] Receive a seed customer group, calculate a center vector based on the high-order embedding vectors of various sub-users in the seed customer group, and calculate the similarity score between the high-order embedding vector of the candidate user and the center vector.

[0012] A path coupling analysis module is constructed to retrieve heterogeneous information graphs based on similarity scores and quantify the interpretable higher-order paths between candidate users and seed customer groups; the target diffusion customer groups are determined based on the interpretable higher-order paths.

[0013] The steps for configuring initial feature vectors for user nodes include:

[0014] Semantic concept recognition: The natural language processing module is used to perform multi-level semantic analysis on user-generated content, search terms, and comment data in the multi-source heterogeneous data, and extract and encode semantic concept vectors that represent long-term interest preferences;

[0015] Dynamic intent capture: The application time-series data processing module encodes the user's recent browsing and click behavior sequences in the multi-source heterogeneous data to generate dynamic intent vectors representing immediate needs;

[0016] Feature fusion: The semantic concept vector and the dynamic intent vector are concatenated to form the initial feature vector.

[0017] The construction of the higher-order relationship edges includes mutually concerned relationship edges;

[0018] The construction of the shared attention relationship edge is specifically as follows: it is calculated based on the meta-path of "user-follower-opinion leader node-followed-user"; its weight calculation not only considers the number of opinion leader nodes that users jointly follow, but also weights the popularity and scarcity of the opinion leader nodes by inverse user frequency, and assigns high weight to the shared attention of opinion leader nodes in niche vertical fields.

[0019] The construction of the higher-order relationship edges includes co-purchase relationship edges; the construction of the co-purchase relationship edges is based on the meta-path of "user-purchase-product node-purchased-user", and its calculation steps specifically include:

[0020] The category profile similarity score is calculated by constructing purchase vectors of candidate users in terms of product categories and brands, and calculating the cosine similarity between the purchase vectors.

[0021] Calculate the scarcity item similarity score, which is obtained by quantifying the scarcity of each item by calculating the inverse product frequency and summing the scarcity scores of all items jointly purchased by two users.

[0022] The weight of the co-purchase relationship edge is determined by weighted fusion of the category profile similarity score and the rare single product similarity score.

[0023] The construction of the higher-order relation edges includes content similarity relation edges;

[0024] The construction of the content similarity relationship edge is specifically as follows: extract the semantic information part, i.e., the semantic concept vector, from the initial feature vector configured for each user node; calculate the cosine similarity of the semantic concept vectors between any two user nodes; and establish a content similarity relationship edge for user pairs with a similarity higher than a preset content threshold.

[0025] The construction of the higher-order relation edges includes geographical proximity relation edges;

[0026] The construction of the geographic proximity relationship edge is specifically as follows: the geographic information nodes, including the user's check-in points based on location services, are converted into standardized geohash codes; by calculating the overlap of the geohash trajectories of two users within a preset time window and the frequency of shared geographic location nodes, the similarity between offline activity spaces and life scenarios is quantified, and geographic proximity relationship edges are established.

[0027] The graph data processing module includes a heterogeneous graph attention network;

[0028] The graph data processing module is equipped with an edge type-aware attention mechanism, which can automatically learn and distinguish different types of nodes and the different contributions of different types of high-order relation edges to the generation of the final high-order embedding vector, thereby realizing the intelligent fusion of heterogeneous information.

[0029] The specific methods for calculating the seed customer group center vector include:

[0030] Obtain the high-order embedding vectors corresponding to all seed users in the seed customer group, and perform mean pooling operation to calculate the average value of all vectors in the high-order embedding vector set of the seed customer group in each dimension. The resulting average vector is the center vector of the seed customer group.

[0031] The path coupling analysis module determines the target customer group by including the following steps:

[0032] Vector initial screening: Based on similarity scores, a preliminary candidate customer group with scores higher than a preset threshold is selected from the candidate users;

[0033] Path verification: For candidate users in the preliminary candidate customer group, interpretable high-order paths are retrieved and quantified through heterogeneous information graphs, and a path reachability score is calculated for each preliminary candidate user;

[0034] The interpretable higher-order path is further decomposed into multiple independent dimension-specific path scores, specifically including interest path scores, consumption path scores, and geographical path scores. The path accessibility score is obtained by quantifying the multi-dimensional coupling relationship between candidate users and seed customer groups through the path coupling analysis module. The interest path score is determined by the common attention relationship edge and the content similarity relationship edge, the consumption path score is obtained by the common purchase relationship edge, and the geographical path score is determined by the geographical proximity relationship edge.

[0035] Fusion determination: By using a fusion function, the similarity score and the path reachability score are combined to calculate the final confirmation score; and the target diffusion customer group is determined by sorting according to the final confirmation score.

[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0037] 1. This invention significantly improves the signal-to-noise ratio of interest similarity measurement by introducing common interest relationship edges based on inverse user frequency weighting. First, it effectively suppresses noise interference from super nodes, greatly reducing their influence in the graph structure and thus filtering out a large number of false weak connections. Second, it significantly amplifies the influence of niche vertical fields by assigning high weights to scarce opinion leaders, enabling the keen capture of deep commonalities among users in specific vertical fields. This provides higher-quality topological information for graph neural networks, greatly improving the accuracy of subsequent customer base diffusion.

[0038] 2. The content similarity relationship edges constructed in this invention solve the problems of data sparsity and domain isolation in graph structures. First, it realizes cross-domain connections and cold-start associations. In real-world scenarios, users may lack common opinion leaders or products due to differences in platforms or behaviors. This solution starts directly from the user's most essential semantic profile, enabling two users with no overlap in behavior but highly consistent interests to be connected, greatly enriching the connectivity of the graph. Second, it provides pure interest similarity. Compared to joint purchases or shared interests, similarity calculated based on semantic concept vectors better reflects the user's true and stable intrinsic preferences, thus generating a deeper and more essential embedded representation of the user.

[0039] 3. The path coupling analysis module of this invention realizes the core concept of dual verification, providing extremely high business credibility; it provides interpretable evidence for each diffused user, greatly enhancing the business's trust in and adoption of the model; it significantly improves the accuracy of diffusion, effectively filtering out pseudo-similar users and ensuring the high purity of the final customer group; at the same time, it provides rich business insights, and by analyzing the path score composition of the final customer group, it can accurately understand the core commonalities of users. Attached Figure Description

[0040] Figure 1 This is a flowchart illustrating a dual-verification customer diffusion method that integrates graph structure and semantic concepts according to the present invention.

[0041] Figure 2 This is a schematic diagram illustrating the construction process of the heterogeneous information graph of the present invention;

[0042] Figure 3 This is a schematic diagram of the target customer group identification process of the present invention. Detailed Implementation

[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0044] Example 1:

[0045] This invention proposes a dual-verification customer diffusion method that integrates graph structure and semantic concepts. The process of the method is as follows: Figure 1 As shown, it includes:

[0046] Collect multi-source heterogeneous data of candidate users and construct a heterogeneous information graph; the heterogeneous information graph includes user nodes and intermediate nodes for establishing associations, the intermediate nodes include opinion leader nodes, product nodes, content tag nodes and geographic information nodes; based on user nodes and intermediate nodes, construct high-order relationship edges representing the structural similarity between users; configure initial feature vectors for user nodes, the initial feature vectors include semantic information representing user profiles;

[0047] The graph data processing module processes the heterogeneous information graph, aggregates the initial feature vectors of user nodes, as well as their structural positions and higher-order relationship information in the heterogeneous information graph, and generates a higher-order embedding vector for each user node.

[0048] Receive a seed customer group, calculate a center vector based on the high-order embedding vectors of various sub-users in the seed customer group, and calculate the similarity score between the high-order embedding vector of the candidate user and the center vector.

[0049] A path coupling analysis module is constructed to retrieve heterogeneous information graphs based on similarity scores and quantify the interpretable higher-order paths between candidate users and seed customer groups; the target diffusion customer groups are determined based on the interpretable higher-order paths.

[0050] The construction process of the heterogeneous information graph is as follows: Figure 2 As shown.

[0051] Collect multi-source heterogeneous data of candidate users. The sources of the multi-source heterogeneous data include, but are not limited to: obtaining basic user profiles through CRM systems, obtaining user-generated content, following lists and comment data through social media platforms, obtaining purchase records, browsing sequences, search terms, click behaviors and product comment data through e-commerce platforms, and obtaining geographical location check-in points and timestamps through location services (LBS).

[0052] The steps for configuring initial feature vectors for user nodes include:

[0053] Semantic concept recognition: The natural language processing module is used to perform multi-level semantic analysis on user-generated content, search terms, and comment data in the multi-source heterogeneous data, and extract and encode semantic concept vectors that represent long-term interest preferences;

[0054] The natural language processing module specifically includes employing a Transformer-based pre-trained language model and constructing a domain knowledge graph to define multi-level semantic concepts; inputting user-related raw text, such as user-generated content, search terms, and comment data, into the pre-trained language model, and performing word segmentation and entity recognition; mapping the identified entities / phrases to predefined semantic concepts; and summarizing the semantic concepts mapped to all user texts to output a semantic concept vector.

[0055] Specifically, the natural language processing module adopts BERT-base, a 12-layer transformer structure with 768 hidden units per layer, and uses the cross-entropy loss function for intent classification training.

[0056] Dynamic intent capture: The application time-series data processing module encodes the user's recent browsing and click behavior sequences in the multi-source heterogeneous data to generate dynamic intent vectors representing immediate needs;

[0057] The time-series data processing module employs recurrent neural networks such as GRU and LSTM, taking a time-ordered sequence of user behaviors as input; each behavior can be encoded as its ID embedding. The model processes the sequence sequentially, learning the temporal dependencies between behaviors and outputting a dynamic intent vector. This vector is typically taken from the hidden state of the last time step of the sequence model and is a compressed representation of the user's immediate needs.

[0058] Feature fusion: The semantic concept vector and the dynamic intent vector are concatenated to form the initial feature vector.

[0059] This invention greatly enriches the semantic information of user nodes by constructing an initial feature vector that integrates long-term and short-term interests, making the extracted semantic concepts more in-depth and hierarchical. At the same time, the dynamic intent captured by the temporal model ensures the timeliness of the profile and provides high-quality input for subsequent graph neural networks. When aggregating neighbor information, it can more accurately determine the similarity between nodes, thereby generating higher-order embedding vectors with greater representational power.

[0060] The construction of the higher-order relationship edges includes mutually concerned relationship edges;

[0061] The construction of the shared attention relationship edge is specifically as follows: it is calculated based on the meta-path of "user-follower-opinion leader node-followed-user"; its weight calculation not only considers the number of opinion leader nodes that users jointly follow, but also weights the popularity and scarcity of the opinion leader nodes by inverse user frequency, and assigns high weight to the shared attention of opinion leader nodes in niche vertical fields.

[0062] The first step is to identify common interests. For any two users, find the set of opinion leaders they both follow. If the set is empty, the weight is 0. If the set of opinion leaders is not empty, the weight of the common interest relationship edge is weighted by scarcity based on the number and scarcity of opinion leader nodes.

[0063] Inverse user frequency coefficient of each opinion leader The calculation formula is:

[0064] ;

[0065] in, The inverse user frequency coefficient of opinion leader k; This refers to the total number of users in the system. For example, on an online shopping platform, it refers to the total number of people registered on that platform. This indicates the number of users who follow opinion leader k.

[0066] Then, normalization is performed based on the inverse user frequency coefficient to obtain the inverse user frequency weight; the number of opinion leader nodes is then weighted based on the inverse user frequency weight.

[0067] In social networks, the degree distribution of nodes follows a power-law distribution, with a few super nodes possessing a massive number of followers. If we simply count the number of opinion leader nodes they both follow, then two users might be considered highly similar merely because they both follow a few top celebrities, but this is clearly a weak correlation or even noise.

[0068] This invention significantly improves the signal-to-noise ratio of interest similarity measurement by introducing shared interest relationship edges based on inverse user frequency weighting. First, it effectively suppresses noise interference from supernodes, greatly reducing their influence in the graph structure and thus filtering out a large number of false weak connections. Second, it significantly amplifies the influence of niche vertical fields by assigning high weights to scarce opinion leaders, enabling the keen capture of deep commonalities among users in specific vertical fields. This provides higher-quality topological information for graph neural networks, greatly improving the accuracy of subsequent customer base diffusion.

[0069] Furthermore, the process for identifying opinion leaders is as follows:

[0070] Node Influence Calculation: In the user-attention subgraph of the heterogeneous information graph, the initial influence score of each user node is calculated based on the graph centrality algorithm;

[0071] Content verticality calculation: The natural language processing module is used to perform semantic analysis on the content generated by user nodes, extract its semantic concept vector, and calculate the topic concentration of the vector in a specific vertical domain to obtain the content verticality score.

[0072] Node confirmation: The initial influence score and the content verticality score are weighted and fused to generate the final leader score; user nodes whose final leader scores are higher than the preset leader threshold are identified as opinion leader nodes.

[0073] Specifically, after constructing the heterogeneous information graph, for the "user-follower-user" topology, graph centrality algorithms such as PageRank or LeaderRank are used to calculate the initial influence score of all nodes, quantifying their structural importance in the network. Next, the natural language processing module is invoked to process the user-generated content of each candidate node, obtaining its semantic concept vector representing long-term interests. The content verticality of the node is quantified by calculating the cosine similarity between this vector and the concepts in the predefined vertical domain knowledge graph, or by using entropy to measure the distribution concentration of its semantic concept vector. Finally, a fusion function, such as linear weighting, is set to combine the initial influence score with the content verticality score.

[0074] A node must possess both high influence and high content verticality for its final leader score to exceed the threshold and be confirmed as an opinion leader node. This scheme ensures that the identified opinion leader nodes are not only highly popular but also of high quality and high relevance, avoiding the misclassification of super nodes that are enthusiastic about all fields but not proficient in any as opinion leaders in a vertical field, and significantly improving the signal-to-noise ratio of shared attention relationships.

[0075] The construction of the higher-order relationship edges includes co-purchase relationship edges; the construction of the co-purchase relationship edges is based on the meta-path of "user-purchase-product node-purchased-user", and its calculation steps specifically include:

[0076] The category profile similarity score is calculated by constructing purchase vectors of candidate users in terms of product categories and brands, and calculating the cosine similarity between the purchase vectors.

[0077] Calculate the scarcity item similarity score, which is obtained by quantifying the scarcity of each item by calculating the inverse product frequency and summing the scarcity scores of all items jointly purchased by two users.

[0078] The calculation and use of the inverse product frequency is similar to that of the inverse user frequency, and is used to quantify the importance of scarce niche products.

[0079] The weight of the co-purchase relationship edge is determined by weighted fusion of the category profile similarity score and the rare single product similarity score.

[0080] This invention constructs a highly robust consumer relationship edge by integrating category profiles and scarce individual products. Firstly, it balances the breadth and depth of consumption. Category profile similarity, at a macro level, identifies user groups with similar lifestyles and consumption levels; while scarce individual product similarity, through weighting, accurately identifies users with shared unique preferences at a micro level. Secondly, it effectively filters out consumer noise. By suppressing the weight of high-frequency, low-information-content goods such as mass-market fast-moving consumer goods, the construction of the relationship edge focuses on scarce and niche consumption behaviors that truly reflect user differences, providing stronger commercial interpretability.

[0081] The construction of the higher-order relation edges includes content similarity relation edges;

[0082] The construction of the content similarity relationship edge is specifically as follows: extract the semantic information part, i.e., the semantic concept vector, from the initial feature vector configured for each user node; calculate the cosine similarity of the semantic concept vectors between any two user nodes; and establish a content similarity relationship edge for user pairs with a similarity higher than a preset content threshold.

[0083] In heterogeneous information graphs, relationships between users, such as shared interests and joint purchases, typically rely on common intermediary nodes. If two users have the same deep interests but follow different opinion leaders and purchase similar but different types of goods, they may be isolated and unable to establish connections in a graph structure based on shared behavior.

[0084] To address the aforementioned shortcomings, this solution constructs content similarity edges that bypass intermediate nodes and directly establish connections based on the user's intrinsic "semantic concepts" expressed through language. As long as the user's semantic concept vectors are identified and encoded as similar, a content similarity edge can be established between them.

[0085] Furthermore, in order to reduce the computational cost of cosine similarity of semantic concept vectors between user nodes, locality-sensitive hashing or inverted indexing techniques are applied to perform coarse-grained bucketing or initial screening of the semantic concept vectors of all user nodes to obtain a candidate neighbor set for each user node; only the cosine similarity between the user node and the nodes in its candidate neighbor set is calculated.

[0086] This invention constructs content-similar relationship edges, solving the problems of data sparsity and domain isolation in graph structures. First, it achieves cross-domain connections and cold-start associations. In real-world scenarios, users may lack common opinion leaders or products due to platform differences or behavioral variations. This solution directly starts from the user's most fundamental semantic profile, enabling connections between two users whose behaviors have no overlap but whose interests are highly consistent, greatly enriching the graph's connectivity. Second, it provides pure interest similarity. Compared to shared purchases or shared interests, similarity calculated based on semantic concept vectors better reflects users' true and stable intrinsic preferences, thus generating a deeper and more fundamental embedded representation of the user.

[0087] The construction of the higher-order relation edges includes geographical proximity relation edges;

[0088] The construction of the geographic proximity relationship edge is specifically as follows: the geographic information nodes, including the user's check-in points based on location services, are converted into standardized geohash codes; by calculating the overlap of the geohash trajectories of two users within a preset time window and the frequency of shared geographic location nodes, the similarity between offline activity spaces and life scenarios is quantified, and geographic proximity relationship edges are established.

[0089] The geographic information nodes include, but are not limited to, user location-based check-in points, food delivery addresses, and sign-in locations, all accompanied by latitude, longitude, and timestamps. A geohash algorithm is applied to convert each (longitude, latitude) coordinate into a standardized Geohash string; the precision of the Geohash (string length) is pre-set; ultimately, the user's activity trajectory within a preset time window is converted into a trajectory sequence.

[0090] The user's trajectory sequence is converted into a geohash set. This process removes duplicate geohash locations and retains only the list of locations the user has visited.

[0091] The Jaccard similarity coefficient formula is used to calculate the geohash sets of two users.

[0092] The first step is to find the intersection: find the intersection of two geohash sets, which represents the number of geohash locations that the two users have visited together.

[0093] The second step is to find the union: find the union of two geohash sets, which represents the total number of all geohash locations visited by at least one of the two users.

[0094] The third step is to calculate the score: divide the number of intersections by the number of unions to get a ratio between 0 and 1, which is used to measure the degree of overlap between the activity ranges of the two users.

[0095] The quantification process of the frequency of shared geographic location nodes includes: acquiring the user's original trajectory sequence within a time window and retaining repeated access records; identifying the set of locations jointly visited by the two users; and measuring the activity level of the two users within their shared activity range by the ratio of the total number of times the two users visited common locations to the sum of the number of times locations were visited in the original trajectory sequence.

[0096] The weights are calculated based on the degree of overlap in users' activity ranges and the ratio of total activity within the shared activity range.

[0097] This invention incorporates user offline profiles into a graph model by introducing Geohash-based geographic proximity edges, providing a novel and high-value dimension for customer base expansion. Secondly, the application of Geohash encoding balances efficiency and privacy, simplifying complex geospatial calculations into efficient string matching. Simultaneously, its controllable accuracy level provides technical assurance for data compliance, greatly enhancing the model's ability to expand its reach across various scenarios.

[0098] Furthermore, the construction of the higher-order relation edges also includes geographic behavior pattern relation edges; the construction of the geographic behavior pattern relation edges specifically involves:

[0099] Location semantic mapping: Mapping the geographic information nodes to predefined location semantic categories through a point of interest database;

[0100] Trajectory sequence generation: Converts a user's geohash trajectory into a time-ordered sequence of location semantic categories;

[0101] Pattern similarity calculation: Apply sequence similarity algorithm to calculate the similarity of location semantic category sequences between two user nodes, quantify the similarity of their offline life behavior patterns, and establish the geographical behavior pattern relationship edge accordingly.

[0102] Specifically, the geographic behavior pattern relationship edge is a supplement or replacement to the original geographic proximity relationship edge, aiming to solve the problem that the original solution relies on the overlap of geographic hash trajectories and cannot compare users in different cities.

[0103] First, define a set of location semantic categories, such as: {residential area, office area, shopping mall, restaurant, park / green space, educational institution}; second, for each geographic information node, query the commercial POI database and map its (latitude / longitude / Geohash) to one of the above semantic categories. For example, it might be mapped to "shopping mall".

[0104] Next, the user's original trajectory within the preset time window is converted into a semantic trajectory sequence, such as "residential area - office area - shopping mall"; finally, the similarity between the semantic trajectory sequences of two users is calculated.

[0105] For example, user a is in city A, and their trajectory is "residential area - office area - restaurant"; user b is in city B, and their trajectory is also "residential area - office area - restaurant". Although their Geohash does not overlap at all, this scheme determines that they have extremely high similarity in geographic behavior patterns by calculating sequence similarity, such as edit distance or dynamic time warping (DTW), thereby establishing high-weight geographic behavior pattern relationship edges.

[0106] Since the distribution of user data is difficult to predict, there may be situations where the overlap of geohash trajectories and the frequency of shared geographical location nodes among users are low, which makes it difficult for geographical proximity relationship edges to play a significant role. Therefore, the similarity of geographical behavior patterns can be used as a supplement to the overlap of geohash trajectories and the frequency of shared geographical location nodes to jointly calculate geographical proximity relationship edges; alternatively, geographical behavior patterns can be used to replace geographical proximity relationship edges to form geographical behavior pattern relationship edges.

[0107] This invention elevates geographical similarity from physical spatial overlap to the dimension of behavioral pattern convergence, enabling the model to uncover potential customer groups across different cities with highly consistent lifestyles and work patterns, significantly expanding the applicability of the original solution. It complements and replaces the construction of geographical proximity relationships, achieving more accurate geographical information identification.

[0108] The graph data processing module includes a heterogeneous graph attention network;

[0109] The graph data processing module is equipped with an edge type-aware attention mechanism, which can automatically learn and distinguish different types of nodes and the different contributions of different types of high-order relation edges to the generation of the final high-order embedding vector, thereby realizing the intelligent fusion of heterogeneous information.

[0110] The core of heterogeneous graph attention network is the hierarchical attention mechanism, which includes node-level attention and edge-type attention.

[0111] Node-level attention learns the importance of neighboring nodes within the same type of meta-path. For a central user node, it is connected to neighbors through "co-purchase" relationships. The model calculates an attention coefficient representing the importance of neighboring users' information to the central user in the "purchase" dimension. The embedding value of the central user in the "purchase" dimension is a weighted sum of all its "purchase" neighbor features.

[0112] The edge type attention mechanism needs to learn the contribution of different edge (semantic) types to the generation of the final embedding. An attention weight is calculated for each edge type; the final higher-order embedding vector is a weighted fusion of all semantically specific embeddings.

[0113] The training process of the heterogeneous graph attention network includes: pre-training the heterogeneous graph attention network using self-supervised learning; using a link prediction task, randomly masking some higher-order relation edges and having the model predict whether these edges exist; and using contrastive learning to bring closer nodes with strong correlations on the graph (positive samples) and push away unrelated nodes (negative samples). The loss function used is cross-entropy loss.

[0114] User profiles are multi-dimensional, and the importance of different dimensions changes dynamically across different diffusion tasks. Simply piecing together or averaging all information will result in significant information loss and an inability to adapt to task variations. Employing an edge-type-aware attention mechanism, which allows the model to automatically learn how to assign weights to different data sources (edge ​​types), is a learnable and dynamic fusion strategy.

[0115] The heterogeneous graph attention network module employed in this invention achieves efficient and intelligent fusion of multi-source heterogeneous information. First, it solves the heterogeneity problem; through node-level attention, the model can handle different types of neighbors; through semantic-level attention, the model can distinguish different types of relationships. Second, this module has high adaptability; the edge-type-aware attention mechanism allows the model to automatically learn the relative importance of each relationship based on the characteristics of the data itself and the training task; it generates highly condensed high-order embedding vectors; providing a solid foundation for subsequent vector similarity calculation and path analysis, and is a core prerequisite for achieving high-precision dual-verification customer diffusion.

[0116] The specific methods for calculating the seed customer group center vector include:

[0117] Obtain the high-order embedding vectors corresponding to all seed users in the seed customer group, and perform mean pooling operation to calculate the average value of all vectors in the high-order embedding vector set of the seed customer group in each dimension. The resulting average vector is the center vector of the seed customer group.

[0118] This invention employs mean pooling to calculate the centroid vector, significantly improving the efficiency of diffusion calculation. During similarity calculation, mean pooling compresses the seed customer group into a centroid vector, reducing the complexity of similarity calculation and greatly enhancing the scalability and real-time performance of customer diffusion. Furthermore, this centroid vector exhibits strong representativeness and robustness. In a high-quality embedding space, the centroid obtained through mean pooling is the optimal estimate of the average prototype of the group, integrating the characteristics of all seed users and automatically smoothing out the individual characteristics and noise of individual seed users.

[0119] The path coupling analysis module determines the target customer group by including the following steps:

[0120] Vector initial screening: Based on similarity scores, a preliminary candidate customer group with scores higher than a preset threshold is selected from the candidate users;

[0121] Path verification: For candidate users in the preliminary candidate customer group, interpretable high-order paths are retrieved and quantified through heterogeneous information graphs, and a path reachability score is calculated for each preliminary candidate user;

[0122] The interpretable higher-order path is further decomposed into multiple independent dimension-specific path scores, specifically including interest path scores, consumption path scores, and geographical path scores. A path coupling analysis module quantifies the multi-dimensional coupling relationships between candidate users and seed customer groups to obtain path reachability scores. The interest path score is determined by shared attention relationships and content similarity relationships, such as through weighted fusion, with a default initial weight of 0.5 for each. The consumption path score is obtained through shared purchase relationships, and the geographical path score is determined through geographical proximity relationships.

[0123] Fusion determination: By using a fusion function, the similarity score and the path reachability score are combined to calculate the final confirmation score; and the target diffusion customer group is determined by sorting according to the final confirmation score.

[0124] The process for identifying the target customer group is as follows: Figure 3 As shown.

[0125] The specific steps include: selecting a preliminary candidate user group based on the similarity between the candidate user and the center vector; selecting several representative seed user nodes that are closest to the center vector from the seed user group, or constructing virtual nodes based on the center vector; and in the heterogeneous information graph, retrieving the connection paths from the candidate user nodes to the representative seed user nodes or virtual nodes, and calculating the path reachability score.

[0126] To achieve intelligent path coupling analysis, this invention constructs a neural network model, which is essentially a multilayer perceptron. Its input layer receives four key features: vector similarity score, interest path score, consumption path score, and geographical path score. These features are then passed through multiple hidden layers, undergoing complex nonlinear transformations and deep coupling using activation functions, enabling the model to automatically learn the complex AND or OR relationships between features. The output layer is a single node that produces a probability value between 0 and 1, which is the final confirmation score for the candidate user.

[0127] Training the aforementioned neural network requires a labeled dataset. Positive samples come from the seed customer group itself, or from customers who have already been converted in past marketing campaigns. Negative samples include random users, as well as more crucial "hard negative samples," such as "pseudo-similar" users with high vector similarity but lacking interpretable paths. Using these positive and negative sample data, the difference between the model's predictions and the true labels is measured using a binary cross-entropy loss function. Then, backpropagation and optimization algorithms (such as Adam) are used to iteratively update the network's weight parameters until the model can accurately distinguish between target and non-target customers.

[0128] Once the model is trained, its recognition logic will far exceed simple manual weighting. During the inference phase, the model receives the four-dimensional scores of candidate users and performs forward propagation through the complex weights it has learned and trained. It can automatically identify various coupling patterns. For example, it will assign extremely high weights to users with both high scores in "vector similarity" and "consumption path", while imposing significant "rejection" penalties on users with high "vector similarity" but zero scores in "all paths".

[0129] Ultimately, the model outputs a highly reliable confirmation probability, which is used to rank the candidate customer groups, thereby achieving automated and high-precision dual verification.

[0130] The path coupling analysis module of this invention implements the core concept of dual verification, providing extremely high business credibility; it provides interpretable evidence for each diffused user, greatly enhancing the business's trust in and adoption of the model; it significantly improves the accuracy of diffusion, effectively filtering out pseudo-similar users and ensuring the high purity of the final customer group; at the same time, it provides rich business insights, and by analyzing the path score composition of the final customer group, it can accurately understand the core commonalities of users.

[0131] Furthermore, in the step of determining the target diffusion customer group by the path coupling analysis module, between the path verification and the fusion determination, the following steps are also included:

[0132] Seed customer profile analysis: Based on the central vector of the seed customer group, analyze the feature distribution of the seed customer group in different dimensions such as interests, consumption and geography, and determine its core profile type;

[0133] Dynamic weighting of paths: Based on the core profile type, dynamic weights are assigned to the interest path score, consumption path score and geographical path score to highlight the associated paths that match the core profile of the seed customer group;

[0134] Dynamic score fusion: The weighted path score and the similarity score are input into the fusion function to calculate the final confirmation score.

[0135] Specifically, in addition to using neural networks to fuse similarity scores and three types of path scores, a dynamic weighting layer can be added to make the fusion more adaptive.

[0136] Upon receiving a seed customer group, the first step is to analyze the composition of its central vector. For example, if the vector has extremely high weights on the dimension corresponding to the "co-purchase relationship edge," the customer group is determined to be "consumption-driven." If it is highly concentrated on the dimension of the "geographical proximity relationship edge," it is classified as "geographically clustered." Based on this, the path coupling analysis module dynamically adjusts the characteristics of the input fusion function. For "consumption-driven" seed customer groups, the weight of the "consumption path score" will be automatically increased during diffusion; for "geographically clustered" customer groups, the weight of the "geographical path score" will be increased.

[0137] This invention makes the dual verification process more intelligent and adaptive, so that the model no longer treats all paths equally, but determines which interpretable higher-order path is more critical verification evidence based on the preferences of the seed customer group, thereby greatly improving the accuracy of diffusion tasks for different types of customer groups.

[0138] Example 2:

[0139] This invention proposes a dual verification method for customer diffusion that integrates graph structure and semantic concepts. The following section describes the specific implementation process of the invention based on customer diffusion in the outdoor sports category.

[0140] The method includes: collecting multi-source heterogeneous data of candidate users and constructing a heterogeneous information graph; the heterogeneous information graph includes user nodes and intermediate nodes for establishing associations, the intermediate nodes including opinion leader nodes, product nodes, content tag nodes and geographic information nodes; based on user nodes and intermediate nodes, constructing high-order relationship edges representing the structural similarity between users; configuring initial feature vectors for user nodes, the initial feature vectors including semantic information representing user profiles;

[0141] The graph data processing module processes the heterogeneous information graph, aggregates the initial feature vectors of user nodes, as well as their structural positions and higher-order relationship information in the heterogeneous information graph, and generates a higher-order embedding vector for each user node.

[0142] Receive a seed customer group, calculate a center vector based on the high-order embedding vectors of various sub-users in the seed customer group, and calculate the similarity score between the high-order embedding vector of the candidate user and the center vector.

[0143] A path coupling analysis module is constructed to retrieve heterogeneous information graphs based on similarity scores and quantify the interpretable higher-order paths between candidate users and seed customer groups; the target diffusion customer groups are determined based on the interpretable higher-order paths.

[0144] Preferably, the data preparation process includes:

[0145] Target customer group: Find potential users with high spending power who are similar to the existing "seed customer group" and are interested in niche vertical outdoor sports, such as skiing.

[0146] Candidate user set: 100,000 active users on the e-commerce platform.

[0147] Seed customer base: 1,000 users on the platform who have purchased high-end ski equipment in the past year and follow opinion leaders in specific skiing fields on social media.

[0148] Multi-source heterogeneous data collection includes e-commerce platform data: purchase records (category, brand, individual product), browsing / click sequences. Social media data: user-generated content, search terms, comments, following lists. LBS / geographic data: recent check-in points, with latitude, longitude, and timestamps.

[0149] Preferably, the construction of heterogeneous information graphs includes:

[0150] Based on user nodes and intermediate nodes, a high-order relation edge representing the structural similarity between users is constructed; the node types include: user nodes, opinion leader nodes, product nodes, content tag nodes, and geographic information nodes.

[0151] The construction of the higher-order relationship edges includes mutually concerned relationship edges;

[0152] The construction of the shared attention relationship edge is specifically as follows: it is calculated based on the meta-path of "user-follower-opinion leader node-followed-user"; its weight calculation not only considers the number of opinion leader nodes that users jointly follow, but also weights the popularity and scarcity of the opinion leader nodes by inverse user frequency, and assigns high weight to the shared attention of opinion leader nodes in niche vertical fields.

[0153] The construction of the higher-order relationship edges includes joint purchase relationship edges;

[0154] The construction of the co-purchase relationship edge is based on the meta-path of "user-purchase-product node-purchased-user". The calculation steps include: calculating the category profile similarity score, which is obtained by constructing purchase vectors of candidate users in product categories and brands, and calculating the cosine similarity between purchase vectors; calculating the scarcity single-item similarity score, which is obtained by calculating the inverse product frequency of each product to quantify its scarcity, and accumulating the scarcity scores of all products jointly purchased by two users; the weight of the co-purchase relationship edge is determined by weighted fusion of the category profile similarity score and the scarcity single-item similarity score. The default weights of both the category profile similarity score and the scarcity single-item similarity score are 0.5, which can be dynamically adjusted based on data learning and training.

[0155] The construction of the higher-order relation edges includes content similarity relation edges;

[0156] The construction of the content similarity relationship edge is specifically as follows: extract the semantic information part, i.e., the semantic concept vector, from the initial feature vector configured for each user node; calculate the cosine similarity of the semantic concept vectors between any two user nodes; and establish a content similarity relationship edge for user pairs with a similarity higher than a preset content threshold. The initial default value of the preset content threshold is set to 0.8.

[0157] The construction of the higher-order relation edges includes geographical proximity relation edges;

[0158] The construction of the geographic proximity relationship edge involves: converting the geographic information nodes, including user check-in points based on location services, into standardized geohash codes; quantifying the similarity between offline activity spaces and life scenarios by calculating the overlap of geohash trajectories of two users within a preset time window and the frequency of shared geographic location nodes, thus establishing a geographic proximity relationship edge. The quantification process of the frequency of shared geographic location nodes includes: obtaining the user's original trajectory sequence within the time window, retaining repeated access records; identifying the set of locations jointly visited by the two users; and measuring the activity level of the two users within their shared activity range by the ratio of the total number of visits to common locations to the sum of the number of visits to locations in the original trajectory sequence. Weights are calculated based on the degree of overlap in user activity ranges and the ratio of total activity within the shared activity range. During the calculation process, the default weights for both the degree of overlap in user activity ranges and the ratio of total activity within the shared activity range are 0.5.

[0159] The process of configuring the initial feature vector includes:

[0160] The semantic concept vector model applies a BERT-based pre-trained language model and incorporates a knowledge graph from the outdoor sports domain. It performs multi-level semantic analysis on user-generated content, search terms, and comment data from the past year, extracting and encoding a 128-dimensional vector to represent long-term interests and preferences, such as freestyle skiing. The model's input is user-related raw text, including user-generated content, search terms, and comment data. Through word segmentation and entity recognition, the identified entities or phrases are mapped to predefined multi-level semantic concepts. Finally, the model summarizes the semantic concepts mapped to all user text and outputs a 128-dimensional semantic concept vector.

[0161] The Dynamic Intent Vector uses a 3-layer GRU model to encode a user's recent browsing and click behavior sequences over the past 7 days, generating a 128-dimensional vector representing immediate needs. The model's input is a time-sorted sequence of recent browsing and click behavior, with each behavior encoded as its ID embedding. The model processes this sequence sequentially, learning the temporal dependencies between behaviors and outputting a 128-dimensional Dynamic Intent Vector. This vector is typically taken from the hidden state of the last time step of the sequence model and represents a compressed representation of the user's immediate needs.

[0162] The semantic concept vector and the dynamic intent vector are concatenated to form an initial feature vector with a total dimension of 256.

[0163] The higher-order embedding vector generation process includes:

[0164] A heterogeneous graph attention network is used to construct the graph data processing module. An edge-type-aware attention mechanism is configured, including node-level attention (learning the importance of neighboring nodes) and edge-type attention (learning the contribution of different relation edges to the final embedding). This enables automatic learning and differentiation of the contributions of different types of nodes and different types of higher-order relation edges, thereby intelligently fusing heterogeneous information. The pre-training process of the heterogeneous graph attention network includes: using contrastive learning based on the InfoNCE loss function to perform a link prediction task, training for 200 epochs; and generating a 128-dimensional higher-order embedding vector for each user node.

[0165] The initial screening process based on vector similarity includes: receiving a seed customer group; calculating a center vector based on the high-order embedding vectors of various sub-users in the seed customer group; calculating the similarity score between the high-order embedding vector of the candidate user and the center vector; and selecting preliminary candidate customer groups from the candidate users whose scores are higher than a preset threshold based on the similarity score. The default value of the preset threshold is set to 0.8.

[0166] The process of path verification and quantification includes:

[0167] For the initial candidate customer group, explainable higher-order paths are retrieved and quantified using heterogeneous infographics; these explainable higher-order paths are further decomposed into multiple independent dimension-specific path scores:

[0168] Interest path score: determined by both the edges representing shared interests and the edges representing content similarity;

[0169] Consumption path score: obtained through joint purchase relationship edges;

[0170] Geographic path score: determined by geographic proximity edges.

[0171] The path coupling analysis module quantifies the multi-dimensional coupling relationships between candidate users and seed customer groups, obtaining path accessibility scores. Further, it analyzes the core profile types of the seed customer groups, such as consumption-driven, interest-driven, and geographically clustered. Based on the core profile type, dynamic weights are assigned to interest, consumption, and geographic path scores to highlight associated paths matching the core profile. For example, for consumption-driven users, the weight of the consumption path score is increased to 0.6, while the interest and geographic path scores are adjusted to 0.2. These weight changes can be evaluated by an expert panel or trained on historical data to achieve non-linear adjustments.

[0172] Specifically, the present invention preferably selects a dynamic weight adjustment method through cross-validation, or combines expert historical strategies with neural network algorithms trained on sample data to achieve constrained adaptive changes in weights, thereby avoiding the risk of inconsistency caused by hard-coding.

[0173] A final confirmation score is calculated by combining the similarity score and the path reachability score using a fusion function, such as a multilayer perceptron. The final confirmation scores are then used to rank the target customer groups and determine their reachability. This dual verification process effectively filters out "pseudo-similar" users with high vector similarity but lacking interpretable paths.

[0174] Specifically, the fusion function employs a three-layer multilayer perceptron. The input layer receives four types of features: similarity score, interest path score, consumption path score, and geographical path score. It also includes two hidden layers with 64 and 32 nodes respectively, both using the ReLU activation function. The output layer uses the Sigmoid activation function to output the final confirmation score. The model uses the Adam optimizer with an initial learning rate of 0.001 and is trained using the binary cross-entropy loss function.

[0175] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A dual-verification method for customer diffusion that integrates graph structure and semantic concepts, characterized in that, include: Collect multi-source heterogeneous data of candidate users and construct a heterogeneous information graph. The heterogeneous information graph includes user nodes and intermediate nodes for establishing connections. The intermediate nodes include opinion leader nodes, product nodes, content tag nodes and geographic information nodes. Based on user nodes and intermediate nodes, construct high-order relationship edges that represent the structural similarity between users. Configure initial feature vectors for user nodes. The initial feature vectors include semantic information that represents the user profile. Higher-order relationship edges include mutual interest relationship edges, joint purchase relationship edges, content similarity relationship edges, and geographical proximity relationship edges; In the user-attention subgraph of the heterogeneous information graph, the initial influence score of each user node is calculated based on the graph centrality algorithm; the natural language processing module performs semantic analysis on the content generated by the user nodes, extracts semantic concept vectors, calculates the topic concentration of the vectors in the vertical domain, and obtains the content verticality score. The initial influence score and content verticality score are weighted and merged to generate the final leader score; User nodes whose final leader score is higher than the preset leader threshold are identified as opinion leader nodes; The graph data processing module processes the heterogeneous information graph, aggregates the initial feature vectors of user nodes, and the structural position and higher-order relationship information of user nodes in the heterogeneous information graph, and generates a higher-order embedding vector for each user node. Higher-order relationship information includes mutual interest relationships, co-purchase relationships, content similarity relationships, and geographical proximity relationships; Receive the seed customer group, calculate the center vector based on the high-order embedding vectors of various sub-users in the seed customer group; calculate the similarity score between the high-order embedding vector and the center vector of the candidate user; A path coupling analysis module is constructed to retrieve heterogeneous information graphs based on similarity scores and quantify the explainable higher-order paths between candidate users and seed customer groups. The target customer base for expansion is determined by reviewing the interpretable high-order path. The construction of an explainable high-order path includes: based on interest path score, consumption path score and geographic path score, the path coupling analysis module quantifies the multi-dimensional coupling relationship between candidate users and seed customer groups to obtain path accessibility score; interest path score is determined by common attention relationship edge and content similarity relationship edge, consumption path score is obtained by common purchase relationship edge, and geographic path score is determined by geographic proximity relationship edge. Common interest relationship edges are obtained based on the meta-path of "user - following - opinion leader node - followed - user"; content similarity relationship edges are obtained based on the cosine similarity of the semantic information part in the initial feature vector configured for user nodes.

2. The dual-verification customer diffusion method integrating graph structure and semantic concept as described in claim 1, characterized in that, The steps for configuring initial feature vectors for user nodes include: Semantic concept recognition: The natural language processing module is used to perform multi-level semantic analysis on user-generated content, search terms, and comment data in the multi-source heterogeneous data, and extract and encode semantic concept vectors that represent long-term interest preferences; Dynamic intent capture: The application time-series data processing module encodes the user's recent browsing and click behavior sequences in the multi-source heterogeneous data to generate dynamic intent vectors representing immediate needs; Feature fusion: The semantic concept vector and the dynamic intent vector are concatenated to form the initial feature vector.

3. The dual-verification customer diffusion method integrating graph structure and semantic concept as described in claim 1, characterized in that, The construction of the higher-order relationship edges includes mutually concerned relationship edges; The construction of the shared attention relationship edge is specifically as follows: it is calculated based on the meta-path of "user - following - opinion leader node - followed - user"; its weight calculation not only considers the number of opinion leader nodes that users jointly follow, but also weights the popularity and scarcity of the opinion leader nodes by inverse user frequency, and assigns high weight to the shared attention of opinion leader nodes in niche vertical fields.

4. The dual-verification customer diffusion method integrating graph structure and semantic concept as described in claim 1, characterized in that, The construction of the higher-order relationship edges includes joint purchase relationship edges; The construction of the co-purchase relationship edge is based on the meta-path of "user - purchase - product node - purchased - user", and its calculation steps specifically include: The category profile similarity score is calculated by constructing purchase vectors of candidate users in terms of product categories and brands, and calculating the cosine similarity between the purchase vectors. Calculate the scarcity item similarity score. The scarcity item similarity score is obtained by calculating the inverse product frequency of each item to quantify its scarcity and summing the scarcity scores of all items jointly purchased by two users. The weight of the co-purchase relationship edge is determined by weighted fusion of the category profile similarity score and the rare single product similarity score.

5. The dual-verification customer diffusion method integrating graph structure and semantic concept as described in claim 1, characterized in that, The construction of the higher-order relation edges includes content similarity relation edges; The construction of the content similarity relationship edge is specifically as follows: extract the semantic information part, i.e., the semantic concept vector, from the initial feature vector configured for each user node; calculate the cosine similarity of the semantic concept vectors between any two user nodes; and establish a content similarity relationship edge for user pairs with a similarity higher than a preset content threshold.

6. The dual-verification customer diffusion method integrating graph structure and semantic concept as described in claim 1, characterized in that, The construction of the higher-order relation edges includes geographical proximity relation edges; The construction of the geographic proximity relationship edge specifically involves converting the geographic information nodes, including user check-in points based on location services, into standardized geographic hash codes. By calculating the overlap of geohash trajectories of two users within a preset time window and the frequency of shared geographic location nodes, the similarity between offline activity spaces and life scenarios is quantified, and geographic proximity relationships are established.

7. The dual-verification customer diffusion method integrating graph structure and semantic concept as described in claim 1, characterized in that, The graph data processing module includes a heterogeneous graph attention network; The graph data processing module is equipped with an edge type-aware attention mechanism, which can automatically learn and distinguish different types of nodes and the different contributions of different types of higher-order relation edges to the generation of the final higher-order embedding vector, thereby achieving the fusion of heterogeneous information.

8. The dual-verification customer diffusion method integrating graph structure and semantic concept as described in claim 1, characterized in that, The specific methods for calculating the seed customer group center vector include: Obtain the high-order embedding vectors corresponding to all seed users in the seed customer group, and perform mean pooling operation to calculate the average value of all vectors in the high-order embedding vector set of the seed customer group in each dimension. The resulting average vector is the center vector of the seed customer group.

9. The dual-verification customer diffusion method integrating graph structure and semantic concept as described in claim 1, characterized in that, The path coupling analysis module determines the target customer group by including the following steps: Vector initial screening: Based on similarity scores, a preliminary candidate customer group with scores higher than a preset threshold is selected from the candidate users; Path verification: For candidate users in the preliminary candidate customer group, interpretable high-order paths are retrieved and quantified through heterogeneous information graphs, and a path reachability score is calculated for each preliminary candidate user; Fusion determination: By using a fusion function, the similarity score and the path reachability score are combined to calculate the final confirmation score; and the target diffusion customer group is determined by sorting according to the final confirmation score.

Citation Information

Patent Citations

  • Fair perception social contact recommendation method and related device

    CN118747247A

  • Method, system and software for tracing path in graph

    JP2024116079A