Lottery drawing user filtering method based on knowledge graph

By building a knowledge graph and correlation graph of user multi-source data and combining machine learning models to identify the risk of lottery draws, the problem of difficulty in effectively filtering false accounts in the existing technology is solved, and the fairness and efficiency of user lottery activities are improved.

CN120256529APending Publication Date: 2025-07-04WUHAN YIBAOTONG NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332635.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively identify and filter false accounts in user lottery activities, which affects the fairness and actual effects of the activities. The existing methods are inefficient and costly, making it difficult to deal with complex cheating methods.

Method used

The knowledge graph is constructed based on the user's multi-source data, and the correlation relationship between users is analyzed through the user association graph, and the lottery risk is identified by combining the radial basis function network and the feedforward sequence model to eliminate high-risk users.

Benefits of technology

It realizes accurate identification and filtering of user lottery activities, improves activity fairness, reduces the possibility of participation of false accounts, and improves the purity of user pools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256529A_ABST
    Figure CN120256529A_ABST
Patent Text Reader

Abstract

The invention provides a lottery drawing user filtering method based on a knowledge graph, and the method comprises the steps: monitoring whether a new user triggers a lottery drawing activity participation condition, adding the new user into a user pool when the new user is monitored to trigger the lottery drawing activity participation condition, and obtaining the multi-source data of the new user, the user pool is used for storing information of users participating in lottery drawing; constructing a user knowledge graph based on the multi-source data of the user; constructing a user association graph based on the plurality of user knowledge graphs; inputting the user knowledge graph and the user association graph into a risk analysis model for processing, and identifying the lottery drawing risk of the user; and marking the users with the lottery drawing risk, and removing the marked users from the user pool during lottery drawing. According to the method, the knowledge graph based on the user multi-source data is constructed, and the graph calculation and machine learning technologies are combined, so that the high-risk users can be comprehensively and accurately identified, and the user filtering effect is improved to guarantee the activity fairness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of user screening, and particularly to a method for filtering lottery users based on a knowledge graph. Background Art

[0002] User lottery participation, as a widely used marketing method, is applied in many fields. For example, e-commerce platforms attract users to participate in lottery activities to increase user activity and purchase conversion rate; social platforms increase user stickiness and promote content dissemination through lottery activities. Although lottery activities have many advantages, there are also some potential risks: some users may participate in lotteries through methods such as brushing orders and fake accounts, resulting in prizes not being distributed to real target users, affecting the fairness and actual effect of the activity. Therefore, it is necessary to filter and screen lottery users. The common current methods for filtering and screening lottery users mainly include the following:

[0003] Rule filtering: Screening users based on preset rules (such as account registration time, participation frequency, device information, etc.);

[0004] Behavior analysis: Judging whether a user is a real user by analyzing the user's historical behavior (such as purchase records, browsing records);

[0005] Manual review: Conducting manual review on winning users to ensure that they comply with the activity rules.

[0006] Although the above methods can filter some false and invalid users to a certain extent, there are still certain defects: for example, rule filtering relies on static rules, is difficult to cope with complex cheating methods, and is prone to misjudging real users; behavior analysis requires a large amount of historical data support and has poor judgment effects on new users or low-frequency users; manual review has low efficiency, high costs, and is difficult to handle large-scale lottery activities. It can be seen that existing methods usually only focus on information in a single dimension (such as devices, behaviors), lacking comprehensive analysis of the multi-dimensional association relationships of users. Summary of the Invention

[0007] The purpose of the present invention is to provide a method for filtering lottery users based on a knowledge graph, generating a knowledge graph based on multi-source data of users, and filtering lottery users on this basis to improve the filtering effect and ensure the fairness of lottery activities.

[0008] To achieve the above object of the invention, the present invention provides a method for filtering lottery users based on a knowledge graph, and the method includes:

[0009] Monitoring whether a new user triggers the lottery activity participation condition, and when it is monitored that a new user triggers the lottery activity participation condition, adding the new user to the user pool and obtaining the multi-source data of the new user, where the user pool is used to store user information participating in the lottery;

[0010] Construct a user knowledge graph based on the user's multi-source data;

[0011] Construct a user association graph based on multiple user knowledge graphs;

[0012] Input the user knowledge graph and the user association graph into a risk analysis model for processing to identify the user's lottery risk;

[0013] Mark the users with lottery risks. When conducting a lottery, the marked users will be excluded from the user pool.

[0014] Furthermore, the multi-source data of the user includes but is not limited to: user basic information, user behavior data, social network data, device data, and third-party data.

[0015] Furthermore, constructing a user knowledge graph based on the user's multi-source data specifically includes:

[0016] Perform knowledge extraction based on multi-source data. The knowledge extraction includes entity extraction, relationship extraction, and attribute extraction;

[0017] Based on the preset user knowledge graph schema layer structure, format the knowledge extraction results into a quadruple form, and the quadruple form is expressed as <entity, relationship, entity, time>;

[0018] Parse the entities, relationships, and attributes in the quadruple, generate nodes according to the entities, generate directed edges according to the relationships, create a user knowledge graph based on the nodes and edges, and store the user knowledge graph in a graph database.

[0019] Furthermore, constructing a user association graph based on multiple user knowledge graphs specifically includes:

[0020] Calculate the similarity of the user knowledge graphs corresponding to different users in the user pool, and determine the association degree between different user knowledge graphs according to the similarity;

[0021] Identify whether the association degree between a user knowledge graph and other user knowledge graphs is greater than a preset association threshold. When the association degree between a user knowledge graph and another user knowledge graph is greater than the preset association threshold, it is considered that there is an association between the two user knowledge graphs;

[0022] Generate bidirectional edges between the associated user knowledge graphs. The bidirectional edges are used to connect the nodes of the corresponding user entities in the two user knowledge graphs.

[0023] Furthermore, calculating the similarity of the user knowledge graphs corresponding to different users in the user pool and determining the association degree between different user knowledge graphs according to the similarity specifically includes:

[0024] Calculate the structural similarity of two user knowledge graphs, and record the calculation result as R1;

[0025] Calculate the comprehensive semantic similarity of two user knowledge graphs, and record the calculation result as R2;

[0026] Calculate the correlation degree of two user knowledge graphs according to R1 and R2. The calculation formula is:

[0027] R = W1 × R1 + W2 × R2

[0028] In the above formula, W1 is the weight of the structural similarity, W2 is the weight of the comprehensive semantic similarity, and R represents the correlation degree of two user knowledge graphs.

[0029] Furthermore, calculating the structural similarity of two user knowledge graphs specifically includes:

[0030] Calculate the adjacency matrices M1 and M2 of two user knowledge graphs respectively;

[0031] Calculate the adjacency matrix of the direct product graph of two user knowledge graphs according to the adjacency matrices M1 and M2. The calculation formula is:

[0032]

[0033] Calculate the structural similarity of two user knowledge graphs based on the adjacency matrix of the direct product graph. The calculation formula is:

[0034]

[0035] In the above formula, S represents the structural similarity of user knowledge graphs m1 and m2, τ represents the attenuation factor, and tr() represents the trace of the matrix.

[0036] Furthermore, input the user knowledge graph and the user association graph into the risk analysis model for processing to identify the user's lottery risk, specifically including:

[0037] According to the time information in the quadruple of the user knowledge graph, divide the user knowledge graph into a real-time graph and a historical graph;

[0038] Analyze the historical graph through a radial basis function network to obtain the first risk prediction result;

[0039] Analyze the real-time graph through a feedforward sequence model and a gated recurrent unit to obtain the second risk prediction result;

[0040] Perform a weighted sum of the first risk prediction result and the second risk prediction result to obtain a comprehensive prediction result, and judge the user's lottery risk level according to the comprehensive prediction result.

[0041] Furthermore, the historical graph is analyzed through a radial basis function network to obtain a first risk prediction result, specifically: the historical graph information is learned and updated through a radial basis function network, the features of the target entity at each moment are aggregated, the repeated lottery risk behaviors recorded in the historical graph are identified, and the lottery risk level at the next moment is predicted, including generating a first index vector according to the target entity, the time step, and the preset learning parameters, inputting the first index vector and the historical vocabulary into the softmax function for processing, and outputting the lottery risk prediction probability of the target entity in the historical time table as the first risk prediction result, wherein the historical vocabulary is composed of the target entity and the relationship at different time steps.

[0042] Furthermore, the real-time graph is analyzed by a feedforward sequence model and a gated recurrent unit to obtain a second risk prediction result, which specifically includes:

[0043] Input the target entity, relation, time step and preset learning parameters into the feed-forward sequence model and calculate the second index vector;

[0044] Input the target entity, relation and preset learning parameters into the feedforward sequence model and calculate the third index vector;

[0045] Input the target entity, relation, and time step into the gated recurrent unit and calculate the fourth index vector;

[0046] The second index vector, the third index vector, and the fourth index vector are input into the softmax function for processing, and the lottery risk prediction probability of the target entity at the current moment is output as the second risk prediction result.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] The present invention provides a method for filtering lottery users based on a knowledge graph. When a new user is monitored to participate in a lottery, multi-source data of the corresponding user is obtained to construct a knowledge graph, and a user-related knowledge graph is further constructed on the basis of the user knowledge graph to mine possible fake lottery accounts. The user knowledge graph and the user-related graph are input into a risk analysis model for processing to identify user lottery risks and remove risky users from the user pool. The present invention constructs a knowledge graph based on user multi-source data, combines graph computing and machine learning technology, and can comprehensively and accurately identify high-risk users, improve user filtering effects and ensure activity fairness. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only the preferred embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0050] Figure 1 It is a schematic diagram of the overall process of a lottery user filtering method based on a knowledge graph provided by an embodiment of the present invention. Specific embodiments

[0051] The following describes the principles and features of the present invention in conjunction with the drawings. The listed embodiments are only used to explain the present invention and are not used to limit the scope of the present invention.

[0052] Referring to Figure 1 , this embodiment provides a lottery user filtering method based on a knowledge graph. The method includes:

[0053] S101. Monitor whether a new user triggers the participation condition of the lottery activity. When it is monitored that a new user triggers the participation condition of the lottery activity, add the new user to the user pool and obtain the multi-source data of the new user. The user pool is used to store the user information participating in the lottery.

[0054] Exemplarily, the participation condition of the lottery activity can be set according to the actual situation. For example, monitor whether the user has forwarded a specific tweet, whether the user has clicked the activity participation option, etc. The new user refers to a user who has not participated in the lottery activity before, not specifically a newly registered user.

[0055] S102. Construct a user knowledge graph based on the multi-source data of the user.

[0056] In this embodiment, the multi-source data includes but is not limited to: user basic information, user behavior data, social network data, device data, and third-party data.

[0057] Exemplarily, user basic information includes name, age, gender, etc. User behavior data includes browsing records, purchase records, lottery participation records, etc. Social network data includes friend relationships, interaction records, etc. Device data includes device ID, IP address, geographical location, etc. Third-party data includes credit scores, blacklists, etc. obtained from third-party platforms.

[0058] S103. Construct a user association graph based on multiple user knowledge graphs.

[0059] S104. Input the user knowledge graph and the user association graph into a risk analysis model for processing to identify the lottery risks of users.

[0060] S105. Mark users at risk of lottery participation, and exclude the marked users from the user pool when conducting the lottery.

[0061] In this embodiment, a user knowledge graph is constructed based on multi-source data of users, and a user association knowledge graph is further constructed based on the knowledge graphs of multiple users. The user knowledge graph and the user association knowledge graph are processed by a risk analysis model to identify the lottery risk of users, and the users at risk of lottery participation are excluded from the user pool, thereby improving the purity of the user pool and reducing the possibility of false accounts participating in the lottery activity.

[0062] As a possible implementation, constructing a user knowledge graph based on multi-source data of users specifically includes:

[0063] S201. Perform knowledge extraction based on multi-source data. The knowledge extraction includes entity extraction, relationship extraction, and attribute extraction.

[0064] Exemplarily, through natural language processing (NLP) technology, relevant entities and relationships can be automatically extracted according to the vocabulary and context information in the text, and entities, relationships, and time are associated in this process to obtain corresponding standardized time information.

[0065] S202. Based on the preset user knowledge graph schema layer structure, format the knowledge extraction result into a quadruple form, and the quadruple form is expressed as <entity, relationship, entity, time>.

[0066] Exemplarily, the user knowledge graph can be divided into a data layer and a schema layer. The data layer consists of a series of knowledge with facts as basic units, and the schema layer is based on the data layer. By deeply analyzing the fact units, these knowledge are further generalized and abstracted to construct a more systematic and standardized knowledge structure.

[0067] S203. Analyze the entities, relationships, and attributes in the quadruple, generate nodes according to the entities, generate directed edges according to the relationships, create a user knowledge graph based on the nodes and edges, and store the user knowledge graph in a graph database.

[0068] The user knowledge graph constructed by the above implementation can reflect the basic situation and association relationships of the corresponding users. In the lottery activity, in order to increase the winning probability, some users may create a large number of small accounts to participate in the lottery activity, affecting the fairness of the activity. In order to discover possible false accounts in the user pool, the present invention further analyzes the association relationships between different users by constructing a user association knowledge graph.

[0069] As a further possible implementation, constructing a user association graph based on multiple user knowledge graphs specifically includes:

[0070] S301. Calculate the similarity of the user knowledge graphs corresponding to different users in the user pool, and determine the association degree between different user knowledge graphs according to the similarity.

[0071] S302. Identify whether the association degree between the user knowledge graph and other user knowledge graphs is greater than a preset association threshold. When the association degree between the user knowledge graph and another user knowledge graph is greater than the preset association threshold, it is considered that there is an association between the two user knowledge graphs.

[0072] S303. Generate a two-way edge between the associated user knowledge graphs, and the two-way edge is used to connect the nodes of the corresponding user entities in the two user knowledge graphs.

[0073] In this embodiment, the association degree between different user knowledge graphs is determined by calculating the similarity between any two user knowledge graphs in the user pool, and when the preset conditions are met, it is determined that there is an association between the two user knowledge graphs. A new edge is generated between the associated user knowledge graphs to reflect their relationship, so as to facilitate the subsequent analysis of potential lottery risks.

[0074] As a further possible embodiment, calculating the similarity of the user knowledge graphs corresponding to different users in the user pool and determining the association degree between different user knowledge graphs specifically include:

[0075] S401. Calculate the structural similarity of the two user knowledge graphs, and record the calculation result as R1.

[0076] S402. Calculate the comprehensive semantic similarity of the two user knowledge graphs, and record the calculation result as R2.

[0077] S403. Calculate the association degree of the two user knowledge graphs according to R1 and R2, and the calculation formula is:

[0078] R = W1×R1 + W2×R2

[0079] In the above formula, W1 is the weight of the structural similarity, W2 is the weight of the comprehensive semantic similarity, and R represents the association degree of the two user knowledge graphs.

[0080] In this embodiment, by calculating the structural similarity and comprehensive semantic similarity of any two user knowledge graphs, the similarity of different user knowledge graphs is analyzed from multiple levels, which can largely reflect the association relationship between different user accounts and is beneficial to analyzing and discovering the behavior of batch-registering accounts to participate in the lottery.

[0081] Exemplarily, calculating the structural similarity of the two user knowledge graphs can be achieved through the following operations:

[0082] S501. Calculate the adjacency matrices M1 and M2 of the knowledge graphs of two users respectively.

[0083] S502. Calculate the adjacency matrix of the direct product graph of the knowledge graphs of two users according to the adjacency matrices M1 and M2. The calculation formula is as follows:

[0084]

[0085] S503. Calculate the structural similarity of the knowledge graphs of two users based on the adjacency matrix of the direct product graph. The calculation formula is as follows:

[0086]

[0087] In the above formula, S represents the structural similarity of the knowledge graphs m1 and m2 of users, τ represents the attenuation factor, and tr() represents the trace of the matrix.

[0088] The comprehensive semantic similarity of the knowledge graphs of two users can be calculated through the following method:

[0089] First, use word vectors (such as Word2Vec, BERT) or graph embeddings (such as TransE, Node2Vec) to represent entities, relationships, etc. as vectors, calculate the cosine similarity between the vectors of the two knowledge graphs, and then combine the weights of different dimensions such as entities and relationships, and sum to obtain the comprehensive semantic similarity of the knowledge graphs of two users.

[0090] As another possible implementation, input the user knowledge graph and the user association graph into a risk analysis model for processing to identify the user's lottery risk, specifically including:

[0091] S601. Divide the user knowledge graph into a real-time graph and a historical graph according to the time information in the quadruple of the user knowledge graph.

[0092] S602. Analyze the historical graph through a radial basis function network to obtain the first risk prediction result.

[0093] In this step, learn and update the information of the historical graph through a radial basis function network, aggregate the features of the target entity at each moment, identify the repeated lottery risk behaviors recorded in the historical graph, and predict the lottery risk level at the next moment, including generating a first index vector according to the target entity, time step, and preset learning parameters, inputting the first index vector and the historical vocabulary into the softmax function for processing, and outputting the lottery risk prediction probability of the target entity in the historical time table as the first risk prediction result. The historical vocabulary consists of the target entity and the relationships at different time steps.

[0094] S603. Analyze the real-time atlas through the feed-forward sequence model and the gated recurrent unit to obtain the second risk prediction result.

[0095] This step specifically includes the following operations:

[0096] S701. Input the target entity, relationship, time step, and preset learning parameters into the feed-forward sequence model to calculate the second index vector.

[0097] S702. Input the target entity, relationship, and preset learning parameters into the feed-forward sequence model to calculate the third index vector.

[0098] S703. Input the target entity, relationship, and time step into the gated recurrent unit to calculate the fourth index vector.

[0099] S704. Input the second index vector, third index vector, and fourth index vector into the softmax function for processing, and output the lottery risk prediction probability of the target entity at the current moment as the second risk prediction result.

[0100] S604. Perform weighted summation on the first risk prediction result and the second risk prediction result to obtain the comprehensive prediction result, and judge the user's lottery risk level according to the comprehensive prediction result.

[0101] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A lottery user filtering method based on a knowledge graph, characterized in that The method includes: Monitoring whether a new user triggers the participation condition of the lottery activity. When it is monitored that a new user triggers the participation condition of the lottery activity, adding the new user to the user pool and obtaining the multi-source data of the new user, where the user pool is used to store the user information participating in the lottery; Constructing a user knowledge graph based on the multi-source data of the user; Constructing a user association graph based on multiple user knowledge graphs; Inputting the user knowledge graph and the user association graph into a risk analysis model for processing to identify the lottery risks of users; Marking the users with lottery risks, and removing the marked users from the user pool when conducting the lottery.

2. The lottery user filtering method based on a knowledge graph according to claim 1, wherein The multi-source data of the user includes but is not limited to: user basic information, user behavior data, social network data, device data, and third-party data.

3. A lottery user filtering method based on a knowledge graph according to claim 1 or 2, characterized in that, Constructing a user knowledge graph based on the multi-source data of the user specifically includes: Performing knowledge extraction based on the multi-source data, where the knowledge extraction includes entity extraction, relationship extraction, and attribute extraction; Formatting the knowledge extraction result into a quadruple form based on the preset user knowledge graph schema layer structure, and the quadruple form is expressed as <entity, relationship, entity, time>; Parsing the entities, relationships, and attributes in the quadruple, generating nodes according to the entities, generating directed edges according to the relationships, creating a user knowledge graph based on the nodes and edges, and storing the user knowledge graph in a graph database.

4. A lottery user filtering method based on a knowledge graph according to claim 1, characterized in that, Constructing a user association graph based on multiple user knowledge graphs specifically includes: Calculating the similarity of the user knowledge graphs corresponding to different users in the user pool, and determining the association degree between different user knowledge graphs according to the similarity; Identifying whether the association degree between a user knowledge graph and other user knowledge graphs is greater than a preset association threshold. When the association degree between a user knowledge graph and another user knowledge graph is greater than the preset association threshold, it is considered that there is an association between the two user knowledge graphs; Generating a bidirectional edge between the associated user knowledge graphs, where the bidirectional edge is used to connect the nodes of the corresponding user entities in the two user knowledge graphs.

5. The lottery user filtering method based on a knowledge graph according to claim 4, characterized in that Calculating the similarity of the user knowledge graphs corresponding to different users in the user pool, and determining the association degree between different user knowledge graphs according to the similarity specifically includes: Calculating the structural similarity of two user knowledge graphs, and recording the calculation result as R1; Calculating the comprehensive semantic similarity of two user knowledge graphs, and recording the calculation result as R2; Calculating the association degree of two user knowledge graphs according to R1 and R2, and the calculation formula is: R = W1×R1 + W2×R2 In the above formula, W1 is the weight of the structural similarity, W2 is the weight of the comprehensive semantic similarity, and R represents the association degree of the two user knowledge graphs.

6. The lottery user filtering method based on a knowledge graph according to claim 5, characterized in that Calculating the structural similarity of two user knowledge graphs specifically includes: Respectively calculating the adjacency matrices M1 and M2 of the two user knowledge graphs; Calculating the adjacency matrix of the direct product graph of the two user knowledge graphs according to the adjacency matrices M1 and M2, and the calculation formula is: Calculating the structural similarity of the two user knowledge graphs based on the adjacency matrix of the direct product graph, and the calculation formula is: In the above formula, S represents the structural similarity of the user knowledge graphs m1 and m2, τ represents the attenuation factor, and tr() represents the trace of the matrix.

7. A lottery user filtering method based on a knowledge graph according to claim 1, characterized in that, The user knowledge graph and the user association graph are input into the risk analysis model for processing to identify the user's lottery-drawing risks, specifically including: According to the time information in the quadruple of the user knowledge graph, the user knowledge graph is divided into a real-time graph and a historical graph; The historical graph is analyzed through a radial basis function network to obtain the first risk prediction result; The real-time graph is analyzed through a feed-forward sequence model and a gated recurrent unit to obtain the second risk prediction result; The first risk prediction result and the second risk prediction result are weighted and summed to obtain a comprehensive prediction result, and the user's lottery-drawing risk level is judged according to the comprehensive prediction result.

8. The lottery user filtering method based on a knowledge graph according to claim 7, wherein The historical graph is analyzed through a radial basis function network to obtain the first risk prediction result, specifically: the historical graph information is learned and updated through a radial basis function network, the features of the target entity at each moment are aggregated, the repeated lottery-drawing risk behaviors recorded in the historical graph are identified, and the lottery-drawing risk level at the next moment is predicted, including generating a first index vector according to the target entity, the time step, and the preset learning parameters, inputting the first index vector and the historical vocabulary into the softmax function for processing, and outputting the lottery-drawing risk prediction probability of the target entity in the historical time table as the first risk prediction result, where the historical vocabulary consists of the target entity and the relationships at different time steps.

9. The lottery user filtering method based on a knowledge graph according to claim 8, characterized in that The real-time graph is analyzed through a feed-forward sequence model and a gated recurrent unit to obtain the second risk prediction result, specifically including: Inputting the target entity, the relationship, the time step, and the preset learning parameters into the feed-forward sequence model to calculate the second index vector; Inputting the target entity, the relationship, and the preset learning parameters into the feed-forward sequence model to calculate the third index vector; Inputting the target entity, the relationship, and the time step into the gated recurrent unit to calculate the fourth index vector; Inputting the second index vector, the third index vector, and the fourth index vector into the softmax function for processing, and outputting the lottery-drawing risk prediction probability of the target entity at the current moment as the second risk prediction result.