Method, device, computer equipment and storage medium for analyzing mail collection and delivery data
By constructing undirected graphs and low-dimensional feature spaces, combined with the prediction of the hidden Markov model, the problem of low accuracy of traditional fraud analysis is solved, and more accurate identification and prevention of fraud behavior is achieved.
Patent Information
- Application Number
- CN202011251262.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-11-11
AI Technical Summary
Traditional fraud possibility analysis is not very accurate in the logistics field, mainly due to the simple data source and analysis methods, making it difficult to effectively identify and prevent fraud.
By obtaining the receiving and sending data of the target user, an undirected graph is constructed and feature dimensionality reduction is performed, nodes are embedded in the low-dimensional feature space, and the feature representation sequence is predicted using the Hidden Markov Model (HMM), and the difference in the prediction results of the two times is calculated to determine the probability of fraud.
The accuracy of fraud possibility analysis is improved, and through multi-dimensional data fusion and differential comparison, the error caused by single-dimensional information is reduced, and fraud is more accurately identified and prevented.
Smart Images

Figure CN114492549B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data technology, and in particular to a method, device, computer equipment and storage medium for analyzing mail collection and delivery data. Background Art
[0002] In the field of logistics, there is a risk of fraud in every link of the core enterprise's supply chain. With the development of big data technology, the application of big data plays an important role in logistics risk control.
[0003] The information technology industry is often faced with the challenge of understanding and mining various complex networks, such as mining and utilizing logistics data information and analyzing the possibility of fraud in users' mailing and receiving behaviors.
[0004] Traditional fraud possibility analysis generally determines the fraud possibility of a user's mailing and receiving behavior by exploring whether there is fraud in the specific user's historical behavior. The data source and analysis methods are relatively simple, resulting in inaccurate analysis results. Summary of the invention
[0005] Based on this, it is necessary to provide a method, device, computer equipment and storage medium for analyzing mail receipt and delivery data that can improve the accuracy of fraud possibility analysis in response to the above technical problems.
[0006] A method for analyzing mail collection and delivery data, the method comprising:
[0007] Obtain data analysis tasks, determine the target users and target time to be analyzed, and obtain the mail collection and delivery data corresponding to the target time period according to the target time period to which the target time belongs;
[0008] Based on the mail collection and delivery data, an undirected graph is constructed, and the nodes on the undirected graph are embedded into a low-dimensional feature space through feature dimensionality reduction.
[0009] According to the target node corresponding to the target user in the low-dimensional feature space, a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at the reference time are obtained, and the reference time is associated with the target time;
[0010] Inputting the first feature representation sequence and the second feature representation sequence into an HMM (Hidden Markov Model) model respectively to obtain a first prediction result and a second prediction result;
[0011] The fraud probability of the target user at the target time is obtained based on the difference data between the first prediction result and the second prediction result.
[0012] In one embodiment, constructing an undirected graph based on the mail collection and delivery data, and embedding the nodes on the undirected graph into a low-dimensional feature space through feature dimensionality reduction includes:
[0013] Determine the user associated with the collection and delivery data based on the collection and delivery behavior corresponding to the collection and delivery data;
[0014] An undirected graph is constructed with users as nodes and collection and delivery behaviors as node associations;
[0015] The nodes in the undirected graph are subjected to feature dimensionality reduction, and the nodes are embedded into the low-dimensional feature space based on the mapping relationship between the undirected graph and the low-dimensional feature space.
[0016] In one embodiment, the feature dimension reduction is performed on the nodes in the undirected graph, and before the nodes are embedded in the low-dimensional feature space based on the mapping relationship between the undirected graph and the low-dimensional feature space, the method further includes:
[0017] According to the node relationship in the low-dimensional feature space, determine the objective function of the low-dimensional feature space;
[0018] According to the stochastic gradient descent optimization algorithm, when the objective function takes the optimal value, the mapping relationship between the undirected graph and the low-dimensional feature space is determined.
[0019] In one embodiment, according to the target node corresponding to the target user in the low-dimensional feature space, obtaining a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at the reference time includes:
[0020] According to the mapping relationship between the low-dimensional feature space and the undirected graph, determine the target node corresponding to the target user in the low-dimensional feature space;
[0021] According to the random walk algorithm, the neighboring nodes of the target node in the low-dimensional feature space are obtained;
[0022] According to any moment in the target time period, obtain the low-dimensional feature data of the subgraph composed of the target node and the neighboring nodes at that moment;
[0023] Filter the low-dimensional feature data corresponding to each time before the target time to obtain the first feature representation sequence corresponding to the target node at the target time;
[0024] A reference time associated with the target time is determined, and low-dimensional feature data corresponding to each time before the reference time of the target node is screened to obtain a second feature representation sequence corresponding to the reference time.
[0025] In one embodiment, filtering low-dimensional feature data corresponding to each time before the target time to obtain a first feature representation sequence corresponding to the target node at the target time includes:
[0026] Get the target number of features of the feature representation sequence;
[0027] According to the number of target features, recursively take the target moment as the last moment, and sequentially obtain the low-dimensional feature data corresponding to each moment before the target moment;
[0028] When the number of acquired low-dimensional feature data reaches the target number of features, a first feature representation sequence is constructed according to the acquired low-dimensional feature data.
[0029] In one embodiment, before inputting the first feature representation sequence and the second feature representation sequence into the HMM model respectively to obtain the first prediction result and the second prediction result, the method further includes:
[0030] The feature representation sequences corresponding to multiple nodes are used as the observation sequence of the initial HMM model;
[0031] The observed sequence is used as training data, and the parameters of the HMM model are trained according to the Baum-Welch algorithm to obtain the HMM model.
[0032] In one embodiment, obtaining the fraud probability of the target user at the target time according to the difference data between the first prediction result and the second prediction result includes:
[0033] Determine difference data between the first prediction result and the second prediction result;
[0034] Calculate the ratio of the difference data to the first prediction result;
[0035] When the ratio is not less than the preset threshold, an analysis result indicating the existence of fraudulent behavior is obtained;
[0036] When the ratio is less than a preset threshold, an analysis result indicating that there is no fraud is obtained.
[0037] A device for analyzing mail collection and delivery data, the device comprising:
[0038] The task acquisition module is used to acquire data analysis tasks, determine the target user and target time to be analyzed, and obtain the mail collection and delivery data corresponding to the target time period according to the target time period to which the target time belongs;
[0039] The node embedding module is used to construct an undirected graph based on the mail collection and delivery data, and embed the nodes on the undirected graph into a low-dimensional feature space through feature dimensionality reduction;
[0040] A sequence obtaining module, used to obtain, according to the target node corresponding to the target user in the low-dimensional feature space, a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at the reference time, wherein the reference time is associated with the target time;
[0041] A result prediction module, used to input the first feature representation sequence and the second feature representation sequence into the HMM model respectively to obtain a first prediction result and a second prediction result;
[0042] The fraud analysis module is used to obtain the fraud probability of the target user at the target time according to the difference data between the first prediction result and the second prediction result.
[0043] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0044] Obtain data analysis tasks, determine the target users and target time to be analyzed, and obtain the mail collection and delivery data corresponding to the target time period according to the target time period to which the target time belongs;
[0045] Based on the mail collection and delivery data, an undirected graph is constructed, and the nodes on the undirected graph are embedded into a low-dimensional feature space through feature dimensionality reduction.
[0046] According to the target node corresponding to the target user in the low-dimensional feature space, a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at the reference time are obtained, and the reference time is associated with the target time;
[0047] Inputting the first feature representation sequence and the second feature representation sequence into the HMM model respectively to obtain a first prediction result and a second prediction result;
[0048] The fraud probability of the target user at the target time is obtained based on the difference data between the first prediction result and the second prediction result.
[0049] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0050] Obtain data analysis tasks, determine the target users and target time to be analyzed, and obtain the mail collection and delivery data corresponding to the target time period according to the target time period to which the target time belongs;
[0051] Based on the mail collection and delivery data, an undirected graph is constructed, and the nodes on the undirected graph are embedded into a low-dimensional feature space through feature dimensionality reduction.
[0052] According to the target node corresponding to the target user in the low-dimensional feature space, a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at the reference time are obtained, and the reference time is associated with the target time;
[0053] Inputting the first feature representation sequence and the second feature representation sequence into the HMM model respectively to obtain a first prediction result and a second prediction result;
[0054] The fraud probability of the target user at the target time is obtained based on the difference data between the first prediction result and the second prediction result.
[0055] The above-mentioned data analysis method, device, computer equipment and storage medium for receiving and sending mails determine the analysis object by obtaining the data analysis task, and use the receiving and sending mails data of the target time period belonging to the target time as the analysis data, and realize the association between users in the supply chain corresponding to the receiving and sending mails data by constructing an undirected graph, embed the nodes on the undirected graph into a low-dimensional feature space, and use the low-dimensional feature space to obtain the feature representation sequence of the target user at the target time and the feature representation sequence at the reference time, and use the HMM model to make predictions, and use the difference data of the prediction results corresponding to the two feature representation sequences to determine the fraud probability of the target user at the target time. In the whole scheme, the feature representation sequence is obtained by using the undirected graph and the low-dimensional feature space, which not only takes into account the information of the data dimension of the target user itself, but also takes into account the topological structure information of the entire supply chain, realizes the fusion of multi-dimensional data, and avoids the errors caused by single-dimensional information by comparing the differences of the prediction results at different times, so as to obtain accurate fraud analysis results. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 A diagram of an application environment of a method for analyzing mail collection and delivery data in one embodiment;
[0057] Figure 2 It is a flowchart of a method for analyzing mail receiving and mailing data in one embodiment;
[0058] Figure 3 It is a flowchart of a method for analyzing mail collection and delivery data in another embodiment;
[0059] Figure 4 A schematic diagram of a flow chart of a method for analyzing mail delivery and receipt data in yet another embodiment;
[0060] Figure 5 It is a flowchart of a method for analyzing mail collection and delivery data in another embodiment;
[0061] Figure 6 A flowchart of a method for analyzing mail delivery data in another embodiment;
[0062] Figure 7 A data flow chart of a method for analyzing mail collection and delivery data in one embodiment;
[0063] Figure 8 is a structural block diagram of a device for analyzing mail receiving and mailing data in one embodiment;
[0064] Fig. 9 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0066] The data analysis method for receiving and sending mails provided in this application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The server 104 obtains the data analysis task configured by the terminal, determines the target user and target time to be analyzed, obtains the collection and delivery data corresponding to the target time period according to the target time period to which the target time belongs, and constructs an undirected graph according to the collection and delivery data. Through feature dimension reduction, the nodes on the undirected graph are embedded in the low-dimensional feature space. According to the target node corresponding to the target user in the low-dimensional feature space, the first feature representation sequence corresponding to the target node at the target time and the second feature representation sequence corresponding to the target node at the reference time are obtained. The reference time is associated with the target time, and the first feature representation sequence and the second feature representation sequence are respectively input into the HMM model to obtain the first prediction result and the second prediction result. According to the difference data between the first prediction result and the second prediction result, the fraud probability of the target user at the target time is obtained. And the obtained fraud probability is pushed to the terminal 102. Among them, the terminal 102 can be, but not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0067] In one embodiment, Figure 2 As shown, a method for analyzing the data of mail collection and delivery is provided, and the method is applied to Figure 1 Taking the server in as an example, the method includes the following steps 202 to 210.
[0068] Step 202, obtain a data analysis task, determine the target user and target time to be analyzed, and obtain the mail collection and delivery data corresponding to the target time period according to the target time period to which the target time belongs.
[0069] The data analysis task is to analyze the fraud possibility of the collection and delivery behavior of a specified user at a specified time. In the data analysis task, the target user to be analyzed and the information corresponding to the target time are configured. The target user is the specified user in the supply chain who needs to be analyzed for fraud risk analysis, and the target time is the time point when a specified collection and delivery behavior occurs.
[0070] The target time period to which the target moment belongs may be a time period of a preset time length with the target moment as the last time, and the time period includes multiple moments with the same time interval. For example, the time length of the target time period is one week, and the time interval is one day. When the target moment is January 7, the target time period may be a time period from January 1 to January 7. The target time period to which the target moment belongs may be a time period including the target moment among multiple time periods that have been pre-divided. For example, the time length of the target time period is one week, and the time interval is one day. Moreover, the server has pre-divided the time periods, such as January 1 to January 7 is one time period, and January 8 to January 14 is one time period. When the target moment corresponds to January 7, the target time period to which the target moment belongs is January 1 to January 7.
[0071] The collection and delivery data refers to the data information generated when users receive or send items in the logistics field. The collection and delivery data includes the recipient information, the sender information, the payment status of the user who receives and sends the items, and the collection and delivery time.
[0072] The mail collection and delivery data corresponding to the target time period refers to the mail collection and delivery data whose mail collection and delivery time is within the target time period. It should be noted that the mail collection and delivery data obtained based on the target time period is only related to the mail collection and delivery time, and has nothing to do with the specific mail collection and delivery user. In other words, as long as the mail collection and delivery time is within the target time period, the mail collection and delivery data corresponding to each user can be obtained.
[0073] Based on the recipients and senders corresponding to the collection and delivery data, a supply chain consisting of recipients and senders can be constructed. The supply chain refers to the network chain structure formed by upstream and downstream enterprises involved in the production and distribution process that provide products to end users. It can be understood that the same user can be both the recipient and the sender in the collection and delivery activities corresponding to different collection and delivery data.
[0074] Step 204, construct an undirected graph based on the mail collection and delivery data, and embed the nodes on the undirected graph into a low-dimensional feature space through feature dimensionality reduction.
[0075] An undirected graph is a graph with no direction on the edges. An undirected graph constructed based on the mail collection and delivery data can represent the connection between two users who have mail collection and delivery behaviors. An undirected graph can be used to display the relationship between multiple users in the same graph. Specifically, based on the mail collection and delivery data within the target time period, the server can store the users involved in the mail collection and delivery data and the relationship between them into an undirected graph.
[0076] Dimensionality reduction is a very important concept in machine learning. In machine learning, we will encounter some high-dimensional data sets. In the case of high-dimensional data, data samples will be sparse, and distance calculation will be difficult. In high-dimensional features, linear correlation between features is easy to appear, resulting in feature redundancy. Feature dimensionality reduction can be achieved through linear dimensionality reduction. Linear dimensionality reduction refers to the process of projecting data into a low-dimensional linear subspace through a linear combination of features. By performing dimensionality reduction on the features of each node in the undirected graph, the nodes on the undirected graph are embedded in the low-dimensional feature space.
[0077] Graph Embedding is the process of embedding nodes in an undirected graph into a low-dimensional vector space to obtain a low-dimensional vector expression. Let V be the set of all nodes in the undirected graph, E be the set of all the relationships (edges) between nodes, and G = (V, E) be the undirected graph consisting of nodes and edges. Learn the feature expression f of graph embedding, embed each node on the undirected graph into a low-dimensional feature space R d , that is, f:V→R d .
[0078] Step 206, based on the target node corresponding to the target user in the low-dimensional feature space, obtain a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at the reference time, and the reference time is associated with the target time.
[0079] By embedding the nodes on the undirected graph into the low-dimensional feature space, each node in the undirected graph has a corresponding node in the low-dimensional feature space. According to the node corresponding to the target user in the undirected graph, based on the mapping relationship between the wireless graph and the low-dimensional feature space, the target node corresponding to the node in the low-dimensional feature space can be determined, that is, the target node corresponding to the target user in the low-dimensional feature space.
[0080] The feature representation sequence is a sequence result obtained by combining multiple feature representations according to certain rules, where the combination rule can be a combination in chronological order. The feature representation refers to the representation result of the feature vector corresponding to each node at each moment. By combining the feature vectors corresponding to multiple moments of each node in chronological order, the feature representation sequence corresponding to the node can be obtained.
[0081] Among them, the sequence length of the feature sequence can be pre-configured, and the first feature representation sequence corresponding to the target moment refers to a feature representation sequence with a fixed length and the feature representation of the target moment as the sequence starting point or the sequence end point. Similarly, the second feature representation sequence corresponding to the reference moment refers to a feature representation sequence with a fixed length and the feature representation of the reference moment as the sequence starting point or the sequence end point. The sequence length of the first feature representation sequence may be the same as the sequence length of the second feature representation sequence. The reference moment refers to a specified moment used to refer to the target moment, which may specifically be the moment before the target moment, or other moments with a fixed time interval length from the target moment. In an embodiment, the association relationship between the reference moment and the target moment may be pre-configured, and when the target moment is determined, the corresponding reference moment is also determined.
[0082] Step 208: Input the first feature representation sequence and the second feature representation sequence into the HMM model respectively to obtain a first prediction result and a second prediction result.
[0083] The HMM model (Hidden Markov Model) is a probability model about time series, which describes the process of randomly generating a random sequence of unobservable states by a hidden Markov chain, and then generating an observation from each state to generate a random sequence of observations. Its state cannot be observed directly, but can be observed through a sequence of observation vectors. Each observation vector is represented by various states through certain probability density distributions, and each observation vector is generated by a state sequence with a corresponding probability density distribution. The HMM model is a dual random process, with a hidden Markov chain with a certain number of states and a set of explicit random functions. The sequence of states randomly generated by the hidden Markov chain is called a state sequence, each state generates an observation, and the random sequence of observations generated by this is called an observation sequence, and each position in the sequence can be regarded as a moment.
[0084] In one embodiment, the first feature representation sequence and the second feature representation sequence are respectively input into the HMM model, and before the first prediction result and the second prediction result are obtained, a training process of the HMM model is also included. The training process of the HMM model is as follows: the feature representation sequences corresponding to the multiple nodes are used as the observation sequence of the initial HMM model. The observation sequence is used as training data, and the parameters of the HMM model are trained according to the Baum-Welch algorithm to obtain the HMM model.
[0085] The multiple nodes can be nodes in an undirected graph or in a low-dimensional feature space. The feature representation sequence corresponding to the node can be a feature representation sequence of a fixed sequence length or a feature representation sequence of a random length. The Baum-Welch algorithm can train and fit the parameters of the model without knowing the state sequence, and obtain an HMM model with optimal parameters.
[0086] By using the feature representation sequence of each node in an undirected graph or in a low-dimensional feature space as the observation sequence of the HMM model for model training, the trained HMM model can be more suitable for actual application scenarios and obtain more accurate prediction results.
[0087] In an embodiment, the first feature representation sequence and the second feature representation sequence are respectively input into the HMM model to obtain a first prediction result and a second prediction result. The first prediction result may be a prediction result corresponding to the first feature representation sequence, and correspondingly, the second prediction result is a prediction result corresponding to the second feature representation sequence. It can be understood that in other embodiments, the first prediction result may also be a prediction result corresponding to the second feature representation sequence, and correspondingly, the second prediction result is a prediction result corresponding to the first feature representation sequence.
[0088] Step 210, obtaining the fraud probability of the target user at the target time according to the difference data between the first prediction result and the second prediction result.
[0089] The first prediction result and the second prediction result are probability data predicted by the HMM model, and the difference data between the first prediction result and the second prediction result is the numerical difference between the two probability data. Taking the prediction result of the first feature representation sequence as the first prediction result as an example, the fraud probability of the target user at the target time is obtained according to the ratio of the data difference to the first prediction result.
[0090] The above-mentioned method for analyzing the data of receiving and sending mails determines the analysis object by obtaining the data analysis task, and takes the data of receiving and sending mails in the target time period belonging to the target time as the analysis data, realizes the association between users in the supply chain corresponding to the data of receiving and sending mails by constructing an undirected graph, embeds the nodes on the undirected graph into a low-dimensional feature space, uses the low-dimensional feature space to obtain the feature representation sequence of the target user at the target time and the feature representation sequence at the reference time, predicts through the HMM model, and uses the difference data of the prediction results corresponding to the two feature representation sequences to determine the fraud probability of the target user at the target time. In the whole scheme, the feature representation sequence is obtained by using the undirected graph and the low-dimensional feature space, which not only considers the information of the data dimension of the target user itself, but also considers the topological structure information of the entire supply chain, realizes the fusion of multi-dimensional data, and avoids the error caused by single-dimensional information by comparing the difference of the prediction results at different times, so as to obtain accurate fraud analysis results.
[0091] In one embodiment, Figure 3 As shown, an undirected graph is constructed based on the mail collection and delivery data, and the nodes on the undirected graph are embedded into a low-dimensional feature space through feature dimensionality reduction, including steps 302 to 306.
[0092] Step 302: Determine the user associated with the mail receiving and sending data according to the mail receiving and sending behavior corresponding to the mail receiving and sending data.
[0093] Step 304: construct an undirected graph with users as nodes and collection and delivery behaviors as node associations.
[0094] Step 306 , performing feature dimensionality reduction on the nodes in the undirected graph, and embedding the nodes into the low-dimensional feature space based on the mapping relationship between the undirected graph and the low-dimensional feature space.
[0095] The collection and delivery behavior is used to characterize the flow process of items from user A to user B or from user B to user A. According to the collection and delivery users corresponding to the collection and delivery behaviors, the users associated with the collection and delivery data can be determined. With users as nodes and collection and delivery behaviors as the association between nodes, an undirected graph between multiple users is constructed. By constructing an undirected graph, the feature association of each user in the supply chain based on the logistics relationship is realized, which paves the way for multi-angle feature fusion.
[0096] Since the dimensions of the data corresponding to each node are numerous, based on the mapping relationship between the undirected graph and the low-dimensional feature space, the feature dimension reduction of the nodes in the undirected graph is performed to realize the mapping of the nodes in the undirected graph in the low-dimensional feature space, and the nodes are embedded in the low-dimensional feature space to realize multi-dimensional feature fusion.
[0097] In one embodiment, Figure 4 As shown, feature dimensionality reduction is performed on nodes in an undirected graph, and based on the mapping relationship between the undirected graph and the low-dimensional feature space, before the nodes are embedded in the low-dimensional feature space, steps 402 to 404 are also included.
[0098] Step 402: Determine the objective function of the low-dimensional feature space according to the node relationship in the low-dimensional feature space.
[0099] Step 404, according to the stochastic gradient descent optimization algorithm, when the objective function takes the optimal value, determine the mapping relationship between the undirected graph and the low-dimensional feature space.
[0100] Specifically, let f:V→R d is the mapping from node to d-dimensional feature space, let is the set of neighboring nodes of node v under the random walk sampling method S. Under the condition that the feature representation of v is f(v), the log-likelihood estimate is:
[0101]
[0102] Assume that given the feature representation f(v) of v, v’s neighboring nodes N S () are conditionally independent, that is:
[0103]
[0104] Further assume that in the d-dimensional feature space, any two nodes influence each other, so we have the following formula:
[0105]
[0106] Based on the above assumptions, the objective function of the low-dimensional feature space can be rewritten as follows:
[0107]
[0108] where ∑ u∈V exp(f(u)·f(v)) is estimated using negative sampling.
[0109] According to the stochastic gradient descent optimization algorithm, optimize the objective function and find the optimal value of the objective function f:V→R d , and obtain the mapping relationship between the undirected graph and the low-dimensional feature space.
[0110] In one embodiment, Figure 5 As shown, according to the target node corresponding to the target user in the low-dimensional feature space, a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at the reference time are obtained, that is, step 206, including steps 502 to 510.
[0111] Step 502: Determine a target node corresponding to the target user in the low-dimensional feature space according to a mapping relationship between the low-dimensional feature space and the undirected graph.
[0112] Step 504: Obtain neighboring nodes of the target node in the low-dimensional feature space according to the random walk algorithm.
[0113] Step 506: According to any moment in the target time period, obtain low-dimensional feature data of a subgraph formed by the target node and neighboring nodes at that moment.
[0114] Step 508: Filter the low-dimensional feature data corresponding to each time before the target time to obtain a first feature representation sequence corresponding to the target node at the target time.
[0115] Step 510, determining a reference time associated with the target time, screening low-dimensional feature data corresponding to each time before the reference time of the target node, and obtaining a second feature representation sequence corresponding to the reference time.
[0116] The basic idea of the random walk algorithm is to randomly select a node and jump to another node. Whenever you reach a node, you have two choices: stay at the current node or continue to jump to other nodes. If the probability that the user continues to visit the node is d, then the probability that the user stays at the current node is 1-d. If you continue to visit other nodes, you will randomly visit another node pointed to by the current node in a uniformly distributed manner. This is a random walk process. According to the random walk algorithm, you can get the neighboring nodes of the target node in the low-dimensional feature space.
[0117] The subgraph composed of the target node and its neighboring nodes refers to a node connection graph that only includes the target node and its neighboring nodes. The node connection graph is part of an undirected graph, so it is called a subgraph. The low-dimensional feature data of the subgraph at time is used to represent the result obtained by fusion of the features of the target node and its neighboring nodes at a specified time.
[0118] For example, for any node v, let the subgraph formed by it and all its neighboring nodes at time t0 be Then the low-dimensional feature representation of the subgraph is in, is the normalization factor, and d(v) represents the degree of node v.
[0119] In an embodiment, a node v is at t0, t1, ... t n The feature representation sequence of the subgraph sequence composed of moments As the observation sequence of HMM, the parameters of the HMM model are trained according to the Baum-Welch algorithm.
[0120] In this embodiment, we only need to pay attention to the mail collection and delivery behaviors in the most recent m moments, and set the target moment t n The feature representation sequence Input the HMM model and get the first prediction result according to the forward-backward algorithm Then the reference time (such as the current time t n The moment before t n-1 ) features represent the sequence Input the HMM model and get the second prediction result according to the forward-backward algorithm
[0121] Finally, calculate if This means that the current collection and delivery behavior is quite different from the previous one, and it can be suspected that there is fraud. Otherwise, it is considered to be normal behavior, where θ is the manually set threshold.
[0122] In one embodiment, low-dimensional feature data corresponding to each time before the target time is filtered to obtain a first feature representation sequence corresponding to the target node at the target time.
[0123] Get the target number of features for the feature representation sequence.
[0124] According to the number of target features, recursion is performed with the target moment as the last moment, and the low-dimensional feature data corresponding to each moment before the target moment is obtained in sequence.
[0125] When the number of acquired low-dimensional feature data reaches the target number of features, a first feature representation sequence is constructed according to the acquired low-dimensional feature data.
[0126] The target feature number refers to the number of feature representations used to form a feature representation sequence. The target feature number can be configured in advance. For example, if it is set to m, it means that only the collection and delivery behaviors of the last m moments need to be paid attention to. The target moment is used as the last moment for recursion, and the low-dimensional feature data corresponding to each moment before the target moment is obtained in sequence. Correspondingly, the target time t is constructed n The feature representation sequence is
[0127] In one embodiment, Figure 6 As shown, obtaining the fraud probability of the target user at the target time according to the difference data between the first prediction result and the second prediction result includes steps 602 to 608.
[0128] Step 602: Determine the difference data between the first prediction result and the second prediction result.
[0129] Step 604, calculating the ratio of the difference data to the first prediction result.
[0130] Step 606: When the ratio is not less than the preset threshold, an analysis result indicating the existence of fraudulent behavior is obtained.
[0131] Step 608: When the ratio is less than a preset threshold, an analysis result indicating that no fraudulent behavior exists is obtained.
[0132] In the embodiment, the first prediction result is obtained through the prediction of the HMM model: The first prediction result is The difference between the first prediction result and the second prediction result is By calculation Get the difference data and the first prediction result if This means that the current collection and delivery behavior is quite different from the previous one, and it can be suspected that there is fraud. Otherwise, it is considered to be normal behavior, where θ is the manually set threshold.
[0133] In one embodiment, Figure 7As shown, a data processing flow chart of a method for analyzing mail collection and delivery data is provided. First, the time period T = (t0, t1, ... t n ) and their relationships are stored as a graph G = (V, E), and then the feature expression f of the graph embedding is learned to embed each node on the graph into a low-dimensional feature space R d , that is, f:V→R d ; For node v, let the subgraph formed by it and all its neighboring nodes at time t0 be Then the low-dimensional feature representation of the subgraph is in is the normalization factor, and d(v) represents the degree of node v. 11 ,v 12 ……v 1d ),……v n =(v n1 ,v n2 ……v nd ). Find the node v at t0, t1, …t n The feature representation sequence of the subgraph sequence composed of moments As the observation sequence of HMM, the parameters of the HMM model are trained according to the Baum-Welch algorithm. In the prediction process, only the collection and delivery behaviors of the most recent m moments are considered. Input HMM and get the prediction result according to the forward-backward algorithm Then the next moment sequence Input HMM and obtain the prediction result according to the forward-backward algorithm if This means that the current collection and delivery behavior is quite different from the previous one, and it can be suspected that there is fraud. Otherwise, it is considered to be normal behavior, where θ is the manually set threshold.
[0134] It should be understood that, although the steps in each flow chart involved in the above-described embodiment are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a portion of the steps in each flow chart involved in the above-described embodiment may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of the steps or stages in other steps.
[0135] In one embodiment, Figure 8As shown, a device for analyzing mail collection and delivery data is provided, including: a task acquisition module 802, a node embedding module 804, a sequence acquisition module 806, a result prediction module 808 and a fraud analysis module 810, wherein:
[0136] Task acquisition module 802, used to acquire data analysis tasks, determine the target user and target time to be analyzed, and obtain the mail collection and delivery data corresponding to the target time period according to the target time period;
[0137] The node embedding module 804 is used to construct an undirected graph based on the mail collection and delivery data, and embed the nodes on the undirected graph into a low-dimensional feature space through feature dimensionality reduction;
[0138] A sequence obtaining module 806 is used to obtain, according to the target node corresponding to the target user in the low-dimensional feature space, a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at the reference time, wherein the reference time is associated with the target time;
[0139] A result prediction module 808, used to input the first feature representation sequence and the second feature representation sequence into the HMM model respectively to obtain a first prediction result and a second prediction result;
[0140] The fraud analysis module 810 is used to obtain the fraud probability of the target user at the target time according to the difference data between the first prediction result and the second prediction result.
[0141] In one of the embodiments, the node embedding module is also used to determine the users associated with the mail collection and delivery data based on the mail collection and delivery behaviors corresponding to the mail collection and delivery data; construct an undirected graph with users as nodes and mail collection and delivery behaviors as node association relationships; perform feature dimensionality reduction on the nodes in the undirected graph, and embed the nodes into the low-dimensional feature space based on the mapping relationship between the undirected graph and the low-dimensional feature space.
[0142] In one of the embodiments, the node embedding module is also used to determine the objective function of the low-dimensional feature space based on the node relationship in the low-dimensional feature space; and according to the stochastic gradient descent optimization algorithm, when the objective function takes the optimal value, determine the mapping relationship between the undirected graph and the low-dimensional feature space.
[0143] In one of the embodiments, the sequence acquisition module is also used to determine the target node corresponding to the target user in the low-dimensional feature space according to the mapping relationship between the low-dimensional feature space and the undirected graph; obtain the neighboring nodes of the target node in the low-dimensional feature space according to the random walk algorithm; obtain the low-dimensional feature data of the subgraph composed of the target node and the neighboring nodes at any moment in the target time period; filter the low-dimensional feature data corresponding to each moment before the target moment to obtain a first feature representation sequence corresponding to the target node at the target moment; determine a reference moment associated with the target moment, filter the low-dimensional feature data corresponding to the target node at each moment before the reference moment, and obtain a second feature representation sequence corresponding to the reference moment.
[0144] In one embodiment, the sequence acquisition module is also used to obtain a target number of features of a feature representation sequence; based on the target number of features, recursively taking the target moment as the last moment, and sequentially obtain low-dimensional feature data corresponding to each moment before the target moment; when the number of low-dimensional feature data obtained reaches the target number of features, construct a first feature representation sequence based on the obtained low-dimensional feature data.
[0145] In one of the embodiments, the parcel receiving and sending data analysis device also includes a model training module, which is used to use the feature representation sequence corresponding to multiple nodes as the observation sequence of the initial HMM model; using the observation sequence as training data, the parameters of the HMM model are trained according to the Baum-Welch algorithm to obtain the HMM model.
[0146] In one of the embodiments, the fraud analysis module is also used to determine the difference data between the first prediction result and the second prediction result; calculate the ratio of the difference data to the first prediction result; when the ratio is not less than a preset threshold, obtain an analysis result indicating the presence of fraud; when the ratio is less than the preset threshold, obtain an analysis result indicating the absence of fraud.
[0147] The above-mentioned receiving and sending data analysis device obtains data analysis tasks, determines the analysis object, and uses the receiving and sending data of the target time period belonging to the target time as the analysis data. By constructing an undirected graph, the association between users on the supply chain corresponding to the receiving and sending data is realized, and the nodes on the undirected graph are embedded in a low-dimensional feature space. The low-dimensional feature space is used to obtain the feature representation sequence of the target user at the target time and the feature representation sequence at the reference time, and the prediction is performed through the HMM model. The difference data of the prediction results corresponding to the two feature representation sequences is used to determine the fraud probability of the target user at the target time. In the whole scheme, the feature representation sequence is obtained through the undirected graph and the low-dimensional feature space, which not only takes into account the information of the data dimension of the target user itself, but also takes into account the topological structure information of the entire supply chain, realizes the fusion of multi-dimensional data, and avoids the error caused by single-dimensional information by comparing the difference of the prediction results at different times, and can obtain accurate fraud analysis results.
[0148] The specific definition of the receiving and sending data analysis device can be found in the definition of the receiving and sending data analysis method above, and will not be repeated here. Each module in the above-mentioned receiving and sending data analysis device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0149] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig. 9 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store parcel collection and delivery data analysis data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for analyzing parcel collection and delivery data is implemented.
[0150] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0151] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0152] Acquire a data analysis task, determine the target user and target time to be analyzed, and obtain the mail collection and delivery data corresponding to the target time period according to the target time period to which the target time belongs; construct an undirected graph based on the mail collection and delivery data, and embed the nodes on the undirected graph into a low-dimensional feature space through feature dimensionality reduction; obtain a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at a reference time according to the target node corresponding to the target user in the low-dimensional feature space, and associate the reference time with the target time; input the first feature representation sequence and the second feature representation sequence into an HMM (Hidden Markov Model) model respectively to obtain a first prediction result and a second prediction result; obtain the fraud probability of the target user at the target time according to the difference data between the first prediction result and the second prediction result.
[0153] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0154] According to the collection and delivery behaviors corresponding to the collection and delivery data, the users associated with the collection and delivery data are determined; an undirected graph is constructed with users as nodes and collection and delivery behaviors as node associations; feature dimensionality reduction is performed on the nodes in the undirected graph, and based on the mapping relationship between the undirected graph and the low-dimensional feature space, the nodes are embedded in the low-dimensional feature space.
[0155] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0156] According to the node relationship in the low-dimensional feature space, the objective function of the low-dimensional feature space is determined; according to the stochastic gradient descent optimization algorithm, when the objective function takes the optimal value, the mapping relationship between the undirected graph and the low-dimensional feature space is determined.
[0157] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0158] According to the mapping relationship between the low-dimensional feature space and the undirected graph, the target node corresponding to the target user in the low-dimensional feature space is determined; according to the random walk algorithm, the neighboring nodes of the target node in the low-dimensional feature space are obtained; according to any moment in the target time period, the low-dimensional feature data of the subgraph composed of the target node and the neighboring nodes at the moment is obtained; the low-dimensional feature data corresponding to each moment before the target moment is filtered to obtain the first feature representation sequence corresponding to the target node at the target moment; the reference moment associated with the target moment is determined, and the low-dimensional feature data corresponding to the target node at each moment before the reference moment is filtered to obtain the second feature representation sequence corresponding to the reference moment.
[0159] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0160] Obtain a target number of features of a feature representation sequence; based on the target number of features, recursively take the target moment as the last moment, and sequentially obtain low-dimensional feature data corresponding to each moment before the target moment; when the number of low-dimensional feature data obtained reaches the target number of features, construct a first feature representation sequence based on the obtained low-dimensional feature data.
[0161] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0162] The feature representation sequences corresponding to the multiple nodes are used as the observation sequence of the initial HMM model; the observation sequence is used as training data, and the parameters of the HMM model are trained according to the Baum-Welch algorithm to obtain the HMM model.
[0163] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0164] Determine the difference data between the first prediction result and the second prediction result; calculate the ratio of the difference data to the first prediction result; when the ratio is not less than a preset threshold, obtain an analysis result indicating the presence of fraudulent behavior; when the ratio is less than the preset threshold, obtain an analysis result indicating the absence of fraudulent behavior.
[0165] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0166] Acquire a data analysis task, determine the target user and target time to be analyzed, and obtain the mail collection and delivery data corresponding to the target time period according to the target time period to which the target time belongs; construct an undirected graph based on the mail collection and delivery data, and embed the nodes on the undirected graph into a low-dimensional feature space through feature dimensionality reduction; obtain a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at a reference time according to the target node corresponding to the target user in the low-dimensional feature space, and associate the reference time with the target time; input the first feature representation sequence and the second feature representation sequence into an HMM (Hidden Markov Model) model respectively to obtain a first prediction result and a second prediction result; obtain the fraud probability of the target user at the target time according to the difference data between the first prediction result and the second prediction result.
[0167] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0168] According to the collection and delivery behaviors corresponding to the collection and delivery data, the users associated with the collection and delivery data are determined; an undirected graph is constructed with users as nodes and collection and delivery behaviors as node associations; feature dimensionality reduction is performed on the nodes in the undirected graph, and based on the mapping relationship between the undirected graph and the low-dimensional feature space, the nodes are embedded in the low-dimensional feature space.
[0169] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0170] According to the node relationship in the low-dimensional feature space, the objective function of the low-dimensional feature space is determined; according to the stochastic gradient descent optimization algorithm, when the objective function takes the optimal value, the mapping relationship between the undirected graph and the low-dimensional feature space is determined.
[0171] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0172] According to the mapping relationship between the low-dimensional feature space and the undirected graph, the target node corresponding to the target user in the low-dimensional feature space is determined; according to the random walk algorithm, the neighboring nodes of the target node in the low-dimensional feature space are obtained; according to any moment in the target time period, the low-dimensional feature data of the subgraph composed of the target node and the neighboring nodes at the moment is obtained; the low-dimensional feature data corresponding to each moment before the target moment is filtered to obtain the first feature representation sequence corresponding to the target node at the target moment; the reference moment associated with the target moment is determined, and the low-dimensional feature data corresponding to the target node at each moment before the reference moment is filtered to obtain the second feature representation sequence corresponding to the reference moment.
[0173] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0174] Obtain a target number of features of a feature representation sequence; based on the target number of features, recursively take the target moment as the last moment, and sequentially obtain low-dimensional feature data corresponding to each moment before the target moment; when the number of low-dimensional feature data obtained reaches the target number of features, construct a first feature representation sequence based on the obtained low-dimensional feature data.
[0175] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0176] The feature representation sequences corresponding to the multiple nodes are used as the observation sequence of the initial HMM model; the observation sequence is used as training data, and the parameters of the HMM model are trained according to the Baum-Welch algorithm to obtain the HMM model.
[0177] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0178] Determine the difference data between the first prediction result and the second prediction result; calculate the ratio of the difference data to the first prediction result; when the ratio is not less than a preset threshold, obtain an analysis result indicating the presence of fraudulent behavior; when the ratio is less than the preset threshold, obtain an analysis result indicating the absence of fraudulent behavior.
[0179] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0180] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0181] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for analyzing mail collection and delivery data, characterized in that: The method comprises: Obtain a data analysis task, determine the target user and target time to be analyzed, and obtain the mail collection and delivery data corresponding to the target time period according to the target time period to which the target time belongs; Determine the user associated with the mail receiving and sending data according to the mail receiving and sending behavior corresponding to the mail receiving and sending data; An undirected graph is constructed with the user as a node and the collection and delivery behavior as a node association relationship; By reducing the dimension of features, embedding the nodes on the undirected graph into a low-dimensional feature space; According to the target node corresponding to the target user in the low-dimensional feature space, a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at a reference time are obtained, wherein the reference time is associated with the target time; Inputting the first feature representation sequence and the second feature representation sequence into an HMM model respectively to obtain a first prediction result and a second prediction result; The fraud probability of the target user at the target time is obtained according to the difference data between the first prediction result and the second prediction result.
2. The method according to claim 1, characterized in that The embedding of the nodes on the undirected graph into a low-dimensional feature space by reducing the feature dimension comprises: Feature dimensionality reduction is performed on the nodes in the undirected graph, and based on the mapping relationship between the undirected graph and the low-dimensional feature space, the nodes are embedded in the low-dimensional feature space.
3. The method according to claim 2, characterized in that The step of performing feature dimensionality reduction on the nodes in the undirected graph, based on the mapping relationship between the undirected graph and the low-dimensional feature space, before embedding the nodes into the low-dimensional feature space, further includes: Determining an objective function of the low-dimensional feature space according to a node relationship in the low-dimensional feature space; According to the stochastic gradient descent optimization algorithm, when the objective function takes an optimal value, a mapping relationship between the undirected graph and the low-dimensional feature space is determined.
4. The method according to claim 1, characterized in that The obtaining, according to the target node corresponding to the target user in the low-dimensional feature space, a first feature representation sequence corresponding to the target node at the target time and a second feature representation sequence corresponding to the target node at the reference time comprises: Determining a target node corresponding to the target user in the low-dimensional feature space according to a mapping relationship between the low-dimensional feature space and the undirected graph; Obtaining neighboring nodes of the target node in the low-dimensional feature space according to a random walk algorithm; According to any moment in the target time period, obtaining low-dimensional feature data of a subgraph formed by the target node and the neighboring nodes at the moment; Filter low-dimensional feature data corresponding to each time before the target time to obtain a first feature representation sequence corresponding to the target node at the target time; A reference time associated with the target time is determined, and low-dimensional feature data corresponding to each time before the reference time of the target node is screened to obtain a second feature representation sequence corresponding to the reference time.
5. The method according to claim 4, characterized in that The filtering of the low-dimensional feature data corresponding to each time before the target time to obtain the first feature representation sequence corresponding to the target node at the target time includes: Get the target number of features of the feature representation sequence; According to the number of target features, recursively taking the target moment as the last moment, and sequentially obtaining low-dimensional feature data corresponding to each moment before the target moment; When the number of the acquired low-dimensional feature data reaches the target number of features, the first feature representation sequence is constructed according to the acquired low-dimensional feature data.
6. The method according to claim 1, characterized in that Before inputting the first feature representation sequence and the second feature representation sequence into the HMM model respectively to obtain the first prediction result and the second prediction result, the method further includes: The feature representation sequences corresponding to multiple nodes are used as the observation sequence of the initial HMM model; The HMM model is obtained by taking the observation sequence as training data and training the parameters of the HMM model according to the Baum-Welch algorithm.
7. The method according to claim 1, characterized in that The obtaining of the fraud probability of the target user at the target time according to the difference data between the first prediction result and the second prediction result includes: Determining difference data between the first prediction result and the second prediction result; Calculating a ratio between the difference data and the first prediction result; When the ratio is not less than a preset threshold, an analysis result indicating the existence of fraudulent behavior is obtained; When the ratio is less than a preset threshold, an analysis result indicating that no fraudulent behavior exists is obtained.
8. A device for analyzing mail collection and delivery data, characterized in that: The device comprises: A task acquisition module is used to acquire data analysis tasks, determine the target user and target time to be analyzed, and obtain the mail collection and delivery data corresponding to the target time period according to the target time period to which the target time belongs; A node embedding module is used to determine the user associated with the collection and delivery data according to the collection and delivery behavior corresponding to the collection and delivery data; construct an undirected graph with the user as the node and the collection and delivery behavior as the node association relationship; and embed the nodes on the undirected graph into a low-dimensional feature space through feature dimensionality reduction; A sequence obtaining module, configured to obtain, according to a target node corresponding to the target user in the low-dimensional feature space, a first feature representation sequence corresponding to the target node at the target time, and a second feature representation sequence corresponding to the target node at a reference time, wherein the reference time is associated with the target time; A result prediction module, used to input the first feature representation sequence and the second feature representation sequence into an HMM model respectively to obtain a first prediction result and a second prediction result; The fraud analysis module is used to obtain the fraud probability of the target user at a target time according to the difference data between the first prediction result and the second prediction result.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Transaction recognition method and device
CN108830603A
User assistance coordination in anomaly detection
US20170279834A1