An E-commerce Network Abnormality Product Detection Method and System

By integrating domain knowledge into the e-commerce network, building a twin network and using graph autoencoder and graph attention network, the problem of difficult to identify abnormal products in the existing technology is solved, and more efficient and reliable abnormal product detection is achieved.

CN115564456BActive Publication Date: 2025-07-01SHANXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211204359.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-07-01
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and reliably identify abnormal products through the topology and product attributes of e-commerce networks, resulting in limited detection capabilities.

Method used

Using a method of integrating domain knowledge, we build a domain knowledge graph of e-commerce network, apply a TransR model, construct a twin network, and use graph autoencoder and graph attention network to build an abnormal product detection model.

Benefits of technology

It improves the reliability and efficiency of abnormal product detection, enhances the robustness and detection performance of the model, and can more accurately identify abnormal products in the e-commerce network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115564456B_ABST
    Figure CN115564456B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of Internet e-commerce. In order to improve the effectiveness and reliability of abnormal product detection, a method and system for detecting abnormal products in an e-commerce network are disclosed. The method includes a preprocessing step for domain knowledge graph data, a construction and optimization step for an abnormal product detection model in the e-commerce network, and an output and processing step for the detection results of abnormal products in the e-commerce network. The present invention combines the knowledge of products in the domain knowledge graph obtained by the TransR model and the e-commerce network topology to construct its twin network. After the multi-view encoder encodes the e-commerce network and the twin network respectively, the aggregator realizes the effective integration of product domain knowledge. Finally, an anomaly score is calculated for each product based on the decoded error to evaluate its anomaly degree, thereby completing the detection and identification of abnormal products. By integrating domain knowledge, the present invention can establish an accurate and reliable description of abnormal products in the network, effectively improving the detection performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet e-commerce, and particularly to a method and system for detecting abnormal products in an e-commerce network. Background Art

[0002] In recent years, due to its characteristics such as being unrestricted by time and space, having rich and complete commodity resources, and having an efficient and reliable transaction method, the e-commerce network has gradually become an important part of people's lives. While bringing convenience and speed to people's lives, the e-commerce network is also full of various fraud and abnormal risks. Many unethical merchants collude with a large number of fraudulent users to disguise inferior products as high-quality products and mix them into the co-purchase network of qualified products through malicious brushing of orders and other means, in order to deceive consumers' trust and purchases and obtain improper benefits. The existence of these products seriously threatens the safety and health of the e-commerce network and also seriously damages the interests of consumers.

[0003] In order to identify abnormal products in the network, a large number of anomaly detection methods have been proposed in the academic and industrial fields, such as methods based on community analysis or measuring ego networks, methods based on feature subspace selection, methods based on residual analysis, and methods based on deep learning. In terms of manifestation, abnormal products have product attributes similar to those of normal products. Therefore, relying solely on the topological structure and product attributes of the e-commerce network often makes it difficult to accurately and reliably describe abnormal products in the network, resulting in insufficient guidance and support for decision-making generation, and to a certain extent limiting the anomaly detection ability of the model. In fact, in addition to having attribute information, products in the e-commerce network also have domain knowledge descriptions about the products in other forms. Among them, the knowledge graph, as an important and widely available knowledge source, describes and stores various types of knowledge in the real world in the form of a graph, and can provide rich background domain knowledge for entities in the real world. The knowledge and attributes of products are descriptions and characterizations of the same object from different perspectives, and they jointly provide a strong basis for decision-making generation. Therefore, how to incorporate domain knowledge into the detection model and then improve the reliability of anomaly detection is of great significance for the research and application of e-commerce network anomaly detection technology. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: to design a method and system for detecting abnormal products in an e-commerce network, and effectively and reliably detect abnormal products in the e-commerce platform network by incorporating the domain knowledge of products.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] The present invention provides a method for detecting abnormal products in an e-commerce network. The method comprises three main steps: step S10 is a step for preprocessing the knowledge graph data in the e-commerce network field; steps S20 to S40 are a step for constructing and optimizing an abnormal product detection model in an e-commerce network; and step S50 is a step for outputting and processing the abnormal product detection results in an e-commerce network. The specific steps are as follows:

[0007] S10, collect all relevant triples for products in the e-commerce network, and perform data cleaning to form a domain knowledge graph to reduce the impact of noise on the detection results;

[0008] S20, applying the TransR model to the knowledge graph obtained in step S10 to obtain the embedding vector of the product domain knowledge, and constructing its twin network based on the topological structure of the e-commerce network;

[0009] S30, based on the twin network obtained in step S20, using the graph autoencoder and the graph attention network to build an abnormal product detection model for the e-commerce network;

[0010] S40, for the e-commerce network abnormal product detection model constructed in step S30, the model is trained by iterative calculation to determine the undetermined parameters in the model;

[0011] S50, using the e-commerce network abnormal product detection model constructed in step S30 and the model parameters determined in step S40, output the abnormal product detection results and process the abnormal products.

[0012] Further, step S10 of the method comprises the following specific steps:

[0013] S11. For each product in the e-commerce network, use SPARQL statements to match the head entity or tail entity from the DBpedia knowledge graph to extract all relevant triples for the product;

[0014] S12, cleaning the triples obtained in step S11, removing triples whose head entities or tail entities are picture links, synonymous entities expressed in other languages, or text information, to reduce the impact of redundancy and noise of triples on the detection results;

[0015] S13, using the triples preprocessed in step S12 to form a domain knowledge graph, denoted as Among them, N is the set of entities, R is the set of relations, and the relation in R connects two entities to form a triple (h, r, t)∈T, h∈N is the head entity, t∈N is the tail entity, r is the relationship between entities, and T represents the set of triples, each of which describes a piece of knowledge about a product in the real world.

[0016] Further, the specific steps included in step S20 of the method are as follows:

[0017] S21. Represent the e-commerce network data as an attribute network The products in the e-commerce network are represented as product nodes in the attribute network, and the co-purchase relationships between products are represented as edges between product nodes; among them, V = {v i |i = 1, 2,..., n} represents the set composed of n product nodes in the network, v i represents the node corresponding to the i-th product, and each product node has d attributes, x i ∈X (i = 1, 2,..., n) represents the attribute vector of the product node v i ; E = {e(i, j)|v i ∈V, v j ∈V} is the set composed of all edges in the network, and e(i, j) is the edge between the product nodes v i and v j indicating that there is a user who has purchased the products represented by these two nodes at the same time, and the attribute network contains a total of |E| = m edges; is the matrix composed of the attribute vectors of all product nodes; in addition, the adjacency matrix represents the topological structure of the attribute network. If there is an edge between the product nodes v i and v j , then A ij = 1, otherwise A ij = 0;

[0018] S22. For each triple (h, r, t) in the domain knowledge graph obtained in S13, denote the embeddings of its head entity and tail entity as h and t respectively, let r represent the embedding of the relationship between entities, and set a mapping matrix M r Project the entities from the entity space to the corresponding relationship space, and obtain the head entity projection vector h r and the tail entity projection vector t r as shown in formulas (1) and (2):

[0019] h r = hM r (1)

[0020] t r = tM r (2)

[0021] S23. Further, establish a transformation between the projected entities, calculate the distance between the head and tail entity projection vectors through the scoring function of formula (3), and use this to measure the possibility of the fact being established:

[0022]

[0023] Among them, f r (h, t) represents the distance score between the head entity h and the tail entity t;

[0024] S24. Continuously update the embeddings of triples through iterative steps S22 and S23, and extract the knowledge embeddings of product corresponding entities in the e-commerce network from the embedding results to construct the domain knowledge matrix of the attribute network, denoted as where k i ∈K represents the knowledge feature of the product node v i ;

[0025] S25. Construct a twin network of the attribute network based on the topological structure A and the domain knowledge matrix K of the attribute network, denoted as

[0026] Furthermore, the e-commerce network abnormal product detection model in step S30 includes four parts: a multi-view encoder composed of an attribute encoder and a knowledge encoder, an aggregator, a decoder, and a detector. Its overall structure is as shown in the appendix Figure 1

[0027] 1) The attribute encoder is stacked by two attention layers, and encodes the attribute network

[0028] into the low-dimensional embedding of product nodes in the latent space under the attribute view; among them, the first layer uses the multi-head attention mechanism, and the second layer uses the single-head attention mechanism. The formal representation of the encoding process is shown in Equation (4):

[0029] In the formula,

[0030] is the low-dimensional embedding of product nodes in the latent space output by the attribute encoder, h′ is the dimension of the product node embedding, and the functions f1(·, ·) and f2(·, ·) represent the first and second graph attention layers respectively; for each graph attention layer, in the process of aggregating the neighborhood information of product nodes in the attribute network, the graph attention mechanism assigns different attention coefficients to each product node in the neighborhood; the attention coefficient between the product node v i and its neighbor is calculated by Equation (5):

[0031] In the formula, e ij is the attention coefficient between the product node v i and v j , representing the importance of v j to v i ; For product node v i is the set of first-order neighbor nodes; is a learnable weight matrix applied to each product node, and || represents the concatenation operation, is the weight vector; The LeakyReLU(·) function is defined as in Equation (6):

[0032]

[0033] where the α value controls the gradient magnitude of the negative part of the linear function;

[0034] To make the attention coefficients between product nodes easy to compare, the attention coefficients are normalized by Equation (7) next:

[0035]

[0036] where α ij is the result after normalization of the attention coefficient between product nodes v i and v j ; The output feature of product node v i in the single-head attention graph attention layer is obtained by the linear combination of the normalized attention coefficient and the input features of neighbor product nodes:

[0037]

[0038] where x′ i represents the embedding of product node v i output by the graph attention layer under single-head attention, and σ(·) is the activation function to achieve the non-linear transformation of data;

[0039] The above steps update the product node embedding by aggregating neighborhood information through the single-head attention mechanism;

[0040] Furthermore, to enhance the model generalization ability, the multi-head attention mechanism is introduced in the graph attention layer. A set of attention coefficients are calculated separately by multiple independent single-head attention networks, and then the multiple embedding results are concatenated by Equation (9) to obtain the output feature of product node v i in the multi-head attention graph attention layer:

[0041]

[0042] where x i ″ is the embedding of product node v i output by the graph attention layer with multi-head attention; P represents the number of single-head attention networks, represents the attention coefficient calculated by the p-th attention network, Denote the weight matrix of the p-th attention network;

[0043] 2) The knowledge encoder has exactly the same structure as the attribute encoder, and encodes the siamese network into the low-dimensional embedding of the product node in the latent space under the knowledge view, and its composition is shown in Equation (10):

[0044]

[0045] where, is the low-dimensional embedding of the product node output by the knowledge encoder in the latent space, and the functions g1(·,·) and g2(·,·) represent the first and second graph attention layers in the knowledge encoder respectively;

[0046] 3) The aggregator fuses the network embeddings under different views to obtain the multi-view unified embedding of the product nodes in the network; either of the two view aggregation strategies can be adopted, specifically as follows:

[0047] (1) Concatenation aggregator: Vertically concatenate the embeddings of the product nodes in the two views to form a vector with a higher dimension, as shown in Equation (11):

[0048]

[0049] where, concat(·) represents the vertical concatenation operation, represents the multi-view unified embedding formed after concatenation aggregation;

[0050] (2) Summation aggregator: Add the embeddings of the two views in the corresponding positions to obtain a new vector with the same dimension, as shown in Equation (12):

[0051]

[0052] In the formula, is the multi-view unified embedding obtained by using the summation aggregator;

[0053] 4) The decoder includes a structure decoder and a feature decoder, where:

[0054] The structure decoder reconstructs the network topology by calculating the inner product between the unified embeddings of the product nodes, as shown in Equation (13):

[0055]

[0056] where, is the reconstructed network adjacency matrix output by the structure decoder, U is the multi-view unified embedding of the product nodes obtained by using any aggregator, and the sigmoid(·) function is defined as in Equation (14):

[0057]

[0058] The feature decoder uses a two - layer fully - connected network to reconstruct the features of product nodes, as shown in Equation (15):

[0059]

[0060] where, is the feature information of the reconstructed product node, and b (l) respectively represent the parameter matrix and bias vector of the l - th fully - connected layer, l ∈ {1, 2}; the product - node features output by this decoder are the reconstruction of both attribute information and domain knowledge in the same feature space;

[0061] 5) The detector uses the reconstruction errors of product nodes in the two dimensions of structure and features as an important basis for evaluating the anomaly score of products. The specific anomaly - scoring function is shown in Equation (16):

[0062]

[0063] In the formula, score(v i ) is the anomaly score of product node v i , which is used to reflect the anomaly degree of the i - th product. The larger its value, the higher the anomaly degree of the product; a i represents the topological structure of product node v i in the attribute network, and are respectively the feature and topological structure of product node v i after being reconstructed by the decoder; λ ∈ [0, 1] is a balance coefficient, which is used to balance the proportion of the reconstruction errors of the feature and topological structure of product node v i in the anomaly score.

[0064] Furthermore, step S40 of this method includes the following specific steps:

[0065] S41. Establish the joint objective function shown in Equation (17) for the e - commerce network anomaly - product detection model established in S30:

[0066]

[0067] where, is the model feature reconstruction error, which is defined by the F - norm in Equation (18):

[0068]

[0069] Since the input of the decoder is the unified embedding fused by the aggregator, the attributes in the attribute network and the twin network are used as the common reference when calculating the feature reconstruction error; is the structural reconstruction error, defined by Equation (19):

[0070]

[0071] In the formula, the balance coefficient λ is the same as λ in Equation (16), and is used to adjust the weight between the two reconstruction errors;

[0072] S42. Initialize the e-commerce network abnormal product detection model in step S30, and initialize the encoder parameters W enc and c, the decoder parameters W dec and b, the hidden layer embedding dimension dim, the number of attention heads nheads, the balance coefficient λ, the number of iterations epoch, and the learning rate learning rate;

[0073] Iteratively execute steps S43 to S48 until the number of iterations is reached, complete the training of the e-commerce network abnormal product detection model, and obtain the optimal parameters of the model:

[0074] S43. Use the domain knowledge graph obtained in S13 as the input, learn the knowledge embedding of the entity according to Equations (1), (2), and (3), extract the entity knowledge vectors corresponding to the products in the e-commerce network to form the domain knowledge matrix K, and combine the topological structure of the attribute network to construct a twin network

[0075] S44. Use Equation (4) for the attribute encoder and Equation (10) for the knowledge encoder to obtain the low-dimensional embeddings of the product nodes in the attribute network and the low-dimensional embeddings of the product nodes in the twin network

[0076] S45. Aggregate the embeddings of the two networks using Equation (11) or (12) to obtain the unified embedding U of the product nodes;

[0077] S46. Decode and reconstruct the product node features and structures using Equations (13) and (15) to obtain and

[0078] S47. Adopt the stochastic gradient descent method to complete the backpropagation by optimizing the joint objective function in Equation (17) to realize the update of the encoder weight matrix W enc and the weight vector c, the decoder weight matrix W dec and the bias vector b;

[0079] S48. Calculate the anomaly score score(v i ) for each product in the e-commerce network using the detector of formula (16).

[0080] Furthermore, step S50 of this method includes the following specific steps:

[0081] S51. After obtaining the optimal parameters of the e-commerce network anomaly product detection model by iteratively executing the training process of steps S43 - S48, use the anomaly detection result obtained in the last training as the final detection result;

[0082] S52. Output the e-commerce network anomaly product detection result to the operation and supervision personnel of the e-commerce platform to improve the efficiency and reliability of their e-commerce network anomaly product detection, and perform further targeted processing on the harm degree and risk impact of the anomaly products.

[0083] The present invention also provides an e-commerce network anomaly product detection system for implementing the above-mentioned e-commerce network anomaly product detection method, including a computer processor and memory, an e-commerce network domain knowledge graph data preprocessing unit, an e-commerce network anomaly product detection model training unit, and an e-commerce network anomaly product detection result output unit; the e-commerce network domain knowledge graph data preprocessing unit executes step S10 to preprocess the collected e-commerce network domain knowledge graph data and load it into the computer memory; the e-commerce network anomaly product detection model training unit constructs an e-commerce network anomaly product detection model according to the e-commerce network domain knowledge graph generated by the e-commerce network domain knowledge graph data preprocessing unit and executes steps S20 - S40, and determines the optimal value of the parameters in the model through iterative calculation; the e-commerce network anomaly product detection result output unit executes step S50 to output the e-commerce network anomaly product detection result to relevant staff or researchers for e-commerce network anomaly product detection and related tasks of network security detection for each e-commerce platform; the specific data processing and calculation work in all units is completed by the computer processor, and all units interact with the computer memory data.

[0084] Compared with the prior art, the present invention has the following advantages:

[0085] 1. The e-commerce network anomaly product detection method of the present invention constructs a twin network by combining domain knowledge, describes and depicts products from another perspective, which helps to enhance the robustness and effectiveness of the model and helps to obtain more reliable e-commerce network anomaly product detection results.

[0086] 2. The e-commerce network abnormal product detection method of the present invention uses a graph autoencoder, a graph attention network and domain knowledge description to establish an e-commerce network abnormal product detection model, so that the model has efficient unsupervised learning capabilities, can establish a reliable description of abnormal products in the network, and effectively improve the detection performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 This is a structural diagram of the abnormal product detection model of the e-commerce network in step S30 of the present invention;

[0088] Figure 2 This is a system structure diagram of an abnormal product detection system for an e-commerce network according to the present invention;

[0089] Figure 3 This is a flow chart of a method for detecting abnormal products on an e-commerce network according to the present invention. DETAILED DESCRIPTION

[0090] In order to further illustrate the technical solution of the present invention, the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0091] The method for detecting abnormal products on an e-commerce network described in the present invention is implemented by a computer program. Figure 3 The process shown in the figure details the specific implementation of the technical solution proposed by the present invention. Through the technical solution of the present invention, abnormal product detection is performed on the Disney dataset of the Amazon e-commerce platform co-purchase network. The dataset consists of 124 movie products, each product has 28 attributes, and these products are connected by 335 co-purchase relationships. The network contains several abnormal movie products.

[0092] The implementation method mainly includes the following key contents:

[0093] S10, collect all relevant triples for products in the e-commerce network, and perform data cleaning to form a domain knowledge graph to reduce the impact of noise on the detection results, including the following specific steps:

[0094] S11. For each product in the e-commerce network, use SPARQL statements to match the head entity or tail entity from the DBpedia knowledge graph to extract all relevant triples for the product;

[0095] S12, cleaning the triples obtained in step S11, removing triples whose head entities or tail entities are picture links, synonymous entities expressed in other languages, or text information, to reduce the impact of redundancy and noise of triples on the detection results;

[0096] S13, using the triples preprocessed in step S12 to form a domain knowledge graph, denoted as Among them, \(N\) is a set of entities, containing 29,517 entities, \(R\) is a set of relationships, containing 331 relationships. The relationships in \(R\) connect two entities to form a triple \((h, r, t)\in T\), where \(h\in N\) is the head entity, \(t\in N\) is the tail entity, and \(r\) is the relationship between entities. \(T\) represents the set of triples, with a total of 53,429 triples. Each triple describes a piece of knowledge about products in the real world.

[0097] S20. Apply the TransR model to the knowledge graph obtained in step S10 to obtain the embedding vectors of product domain knowledge, and construct its twin network based on the topological structure of the e-commerce network, including the following specific steps:

[0098] S21. Represent the Disney e-commerce network dataset as an attribute network Products in the e-commerce network are represented as product nodes in the attribute network, and the co-purchase relationship between products is represented as an edge between product nodes; among them, \(V = \{v i |i = 1, 2,..., n\}\) represents the set composed of \(n = 124\) product nodes in the network, \(v i represents the node corresponding to the \(i\)-th product, and each product node has \(d = 28\) attributes, \(x i \in X(i = 1, 2,..., n)\) represents the attribute vector of the product node \(v i ; \(E=\{e(i, j)|v i \in V, v j \in V\}\) is the set composed of all edges in the network, and \(e(i, j)\) is the edge between the product nodes \(v i and \(v j \), indicating that a user has purchased the products represented by these two nodes at the same time. The attribute network contains a total of \(|E| = m = 335\) edges; is the matrix composed of the attribute vectors of all product nodes; in addition, the adjacency matrix represents the topological structure of the attribute network. If there is an edge between the product nodes \(v i and \(v j \), then \(A ij = 1\), otherwise \(A ij = 0\);

[0099] S22. For each triple \((h, r, t)\) in the domain knowledge graph obtained in S13, denote the embeddings of its head entity and tail entity as \(h\) and \(t\) respectively, let \(r\) denote the embedding of the relationship between entities, and set a mapping matrix \(M r Project the entity from the entity space to the corresponding relationship space to obtain the head entity projection vector \(h r and the tail entity projection vector \(t r As shown in equations (1) and (2):

[0100] h r = hM r (1)

[0101] t r = tM r (2)

[0102] S23. Further, establish a transformation between projection entities, and calculate the distance between the projection vectors of the head and tail entity projections through the scoring function of Equation (3) to measure the possibility of the fact holding:

[0103]

[0104] where f r (h, t) represents the distance score between the head entity h and the tail entity t;

[0105] S24. Continuously update the embedding of the triple through iterative steps S22 and S23, and extract the knowledge embedding of the product corresponding entity in the e-commerce network from the embedding result to construct the domain knowledge matrix of the attribute network, denoted as where k i ∈ K represents the knowledge feature of the product node v i ;

[0106] S25. Construct the twin network of the attribute network based on the topological structure A and the domain knowledge matrix K of the attribute network, denoted as

[0107] S30. Based on the twin network obtained in step S20, use the graph autoencoder and the graph attention network to construct an e-commerce network abnormal product detection model. This model consists of four parts: a multi-view encoder, an aggregator, a decoder, and a detector, which are composed of an attribute encoder and a knowledge encoder. Its overall structure is as shown in the appendix Figure 1 Figure 1 as follows, where:

[0108] 1) The attribute encoder is stacked by two attention layers, and encodes the attribute network into the low-dimensional embedding of the product node in the latent space under the attribute view; among them, the first layer uses the multi-head attention mechanism, and the second layer uses the single-head attention mechanism. The formal representation of the encoding process is shown in Equation (4):

[0109]

[0110] In the formula, is the low-dimensional embedding of the product node in the latent space output by the attribute encoder, h′ = 128 is the dimension of the product node embedding, and the functions f1(·,·) and f2(·,·) represent the first and second graph attention layers respectively; for each graph attention layer, in the process of aggregating the neighborhood information of the product nodes in the attribute network, the graph attention mechanism assigns different attention coefficients to each product node in the neighborhood; the product node v i The attention coefficient between it and its neighbors is calculated by Equation (5):

[0111]

[0112] In the formula, e ij is the attention coefficient between the product node v i and v j , indicating the importance of v j to v i . is the set of first-order neighbor nodes of the product node v i ; is the learnable weight matrix applied to each product node, || represents the concatenation operation, is the weight vector; the LeakyReLU(·) function is defined as Equation (6):

[0113]

[0114] Among them, the value of α = 0.2 controls the gradient magnitude of the negative part of the linear function.

[0115] In order to make the attention coefficients between product nodes easy to compare, next, the attention coefficients are normalized by Equation (7):

[0116]

[0117] Among them, α ij is the result after normalizing the attention coefficient between the product nodes v i and v j ; the output feature of the product node v i in the single-head attention graph attention layer is obtained by the linear combination of the normalized attention coefficient and the input features of the neighbor product nodes:

[0118]

[0119] Among them, x i ′ represents the embedding of the product node v i output by the graph attention layer under single-head attention; σ(·) is the activation function, which realizes the non-linear transformation of the data;

[0120] The above steps update the product node embedding by aggregating neighborhood information through a single-head attention mechanism;

[0121] Furthermore, to enhance the generalization ability of the model, a multi-head attention mechanism is introduced in the graph attention layer. A set of attention coefficients are calculated respectively through multiple independent single-head attention networks, and then the multiple embedding results are concatenated through Equation (9) to obtain the product node v i The output features of the multi-head attention graph attention layer:

[0122]

[0123] where, x i ″ is the embedding of the product node v output by the graph attention layer with multi-head attention i ; P = 3 represents the number of single-head attention networks, represents the attention coefficient calculated by the p-th attention network, represents the weight matrix of the p-th attention network;

[0124] 2) The knowledge encoder has exactly the same structure as the attribute encoder, and encodes the siamese network into the low-dimensional embedding of the product node in the latent space, and its composition is shown in Equation (10):

[0125]

[0126] where, is the low-dimensional embedding of the product node in the latent space output by the knowledge encoder, and the functions g1(·,·) and g2(·,·) represent the first and second graph attention layers in the knowledge encoder respectively;

[0127] 3) The aggregator fuses the network embeddings under different views, and then obtains the multi-view unified embedding of the product node in the network; either of the two view aggregation strategies can be adopted, specifically as follows:

[0128] (1) Concatenation aggregator: Vertically concatenate the product node embeddings of the two views to form a vector with a higher dimension, as shown in Equation (11):

[0129]

[0130] where, concat(·) represents the vertical concatenation operation, represents the multi-view unified embedding formed after concatenation aggregation;

[0131] (2) Summation aggregator: Add the embeddings of the two views in the corresponding positions to obtain a new vector with the same dimension, as shown in Equation (12):

[0132]

[0133] wherein, is the multi-view unified embedding obtained by using the summation aggregator;

[0134] 4) The decoder includes a structure decoder and a feature decoder, wherein:

[0135] The structure decoder realizes the reconstruction of the network topology by calculating the inner product between the unified embeddings of product nodes, as shown in Equation (13):

[0136]

[0137] wherein, is the reconstructed network adjacency matrix output by the structure decoder, U is the multi-view unified embedding of product nodes obtained by using any aggregator, and the sigmoid(·) function is defined as in Equation (14):

[0138]

[0139] The feature decoder uses a two-layer fully connected network to realize the reconstruction of product node features, as shown in Equation (15):

[0140]

[0141] wherein, is the reconstructed feature information of product nodes, and b (l) respectively represent the parameter matrix and bias vector of the l-th fully connected layer, l ∈ {1, 2}; the product node features output by this decoder are the reconstruction of the two parts of information, namely attribute information and domain knowledge, in the same feature space;

[0142] 5) The detector uses the reconstruction errors of product nodes in the two dimensions of structure and feature as an important basis for evaluating the product anomaly score. The specific anomaly scoring function is as shown in Equation (16):

[0143]

[0144] wherein, score(v i ) is the anomaly score of product node v i , which is used to reflect the anomaly degree of the i-th product. The larger its value, the higher the anomaly degree of the product; a i represents the topological structure of product node v i in the attribute network, and are respectively the reconstructed product node v iThe characteristics and topological structure; λ ∈ [0, 1] is the balance coefficient, which is used to balance the product node v i The proportion of the reconstruction errors of the characteristics and topological structure of v in the anomaly score.

[0145] S40. For the e-commerce network anomaly product detection model constructed in step S30, the model is trained by an iterative calculation method to determine the undetermined parameters in the model. S40 includes the following specific steps:

[0146] S41. Establish the joint objective function shown in formula (17) for the e-commerce network anomaly product detection model established in S30:

[0147]

[0148] where is the model feature reconstruction error, which is defined by formula (18) using the F-norm:

[0149]

[0150] Since the input of the decoder is the unified embedding fused by the aggregator, the attributes in the attribute network and the twin network are used as the common reference when calculating the feature reconstruction error; is the structure reconstruction error, which is defined by formula (19):

[0151]

[0152] In the formula, the balance coefficient λ is the same as the λ in formula (16), which is used to adjust the weight between the two reconstruction errors;

[0153] S42. Initialize the e-commerce network anomaly product detection model in step S30, and initialize the encoder parameters W enc and c, the decoder parameters W dec and b, the hidden layer embedding dimension dim, the number of attention heads nheads, the balance coefficient λ, the number of iterations epoch, and the learning rate learning rate;

[0154] Iteratively execute steps S43 to S48 until the number of iterations is reached, complete the training of the e-commerce network anomaly product detection model, and obtain the optimal parameters of the model:

[0155] S43. Use the domain knowledge graph obtained in S13 as the input, learn the knowledge embedding of entities according to formulas (1), (2), and (3), extract the entity knowledge vectors corresponding to products in the e-commerce network to form the domain knowledge matrix K, and combine the topological structure of the attribute network to construct a twin network

[0156] S44, using the attribute encoder of formula (4) and the knowledge encoder of formula (10) to obtain the low-dimensional embedding of the product nodes of the attribute network Low-dimensional embedding of product nodes in twin networks

[0157] S45, using equation (11) or (12) to aggregate the embeddings of the two networks and obtain a unified embedding U of the product node;

[0158] S46. Use equations (13) and (15) to decode and reconstruct the product node features and structures to obtain and

[0159] S47, using stochastic gradient descent method, by optimizing the joint objective function in formula (17) Complete back propagation to realize the encoder weight matrix W enc and weight vector c, decoder weight matrix W dec and the update of the bias vector b;

[0160] S48, using the detector of formula (16) to calculate the abnormal score score (v) of each product in the e-commerce network i ).

[0161] S50, using the e-commerce network abnormal product detection model constructed in step S30 and the model parameters determined in step S40, outputting the abnormal product detection results and processing the abnormal products, including the following specific steps:

[0162] S51, after obtaining the optimal parameters of the abnormal product detection model of the e-commerce network by iteratively executing the training process of steps S43 to S48, the abnormal detection result obtained in the last training is used as the final detection result;

[0163] S52. Output the abnormal product detection results to the operators and supervisors of the e-commerce platform to improve the efficiency and reliability of abnormal product detection, and conduct further targeted processing based on the degree of harm and risk impact of abnormal products.

[0164] Technical effect evaluation:

[0165] In order to verify the effectiveness and advancement of the technical solution proposed in the present invention, the present invention (concatenated aggregator or summed aggregator) is compared with several classic anomaly detection methods, including AMEN, Radar, LOF, SCAN and Dominant methods. The average AUC value, Precision@K and Recall@K values ​​of 20 experiments are used as evaluation indicators, and the above matching results are compared and analyzed. The comparison results are shown in Table 1:

[0166]

[0167] As can be seen from the results in the table, when the technical solution of the present invention detects abnormal products in the e-commerce network, in most cases, it can obtain better precision and recall rates, as well as higher AUC values, in the top K detection results than the existing methods.

[0168] Such as Figure 2 shown, an e-commerce network abnormal product detection system includes a computer processor and memory, an e-commerce network domain knowledge graph data preprocessing unit, an e-commerce network abnormal product detection model training unit, and an e-commerce network abnormal product detection result output unit; the e-commerce network domain knowledge graph data preprocessing unit executes step S10 to preprocess the collected e-commerce network domain knowledge graph data and load it into the computer memory; the e-commerce network abnormal product detection model training unit constructs an e-commerce network abnormal product detection model by executing steps S20-S40 according to the e-commerce network domain knowledge graph generated by the e-commerce network domain knowledge graph data preprocessing unit, and determines the optimal values of the parameters in the model through iterative calculation; the e-commerce network abnormal product detection result output unit executes step S50 to output the e-commerce network abnormal product detection results to relevant staff or researchers for abnormal product detection of each e-commerce platform and related tasks such as network security detection; specific data processing and calculation work in all units is completed by the computer processor, and all units interact with the computer memory data.

Claims

1. An e-commerce network abnormal product detection method, characterized in that, The following steps are involved: S10, collect all relevant triples for products in the e-commerce network, and perform data cleaning to form a domain knowledge graph to reduce the impact of noise on the detection results; S20, applying the TransR model to the knowledge graph obtained in step S10 to obtain the embedding vector of the product domain knowledge, and constructing its twin network based on the topological structure of the e-commerce network; S30, based on the twin network obtained in step S20, using the graph autoencoder and the graph attention network to build an abnormal product detection model for the e-commerce network; S40, for the e-commerce network abnormal product detection model constructed in step S30, the model is trained by iterative calculation to determine the undetermined parameters in the model; S50, using the e-commerce network abnormal product detection model constructed in step S30 and the model parameters determined in step S40, outputting the abnormal product detection results and processing the abnormal products; The step S20 includes the following specific steps: S21. Represent the e-commerce network data as an attribute network The products in the e-commerce network are represented as product nodes in the attribute network, and the co-purchase relationships between products are represented as edges between product nodes; Among them, V = {v i | i = 1, 2,..., n} represents the set composed of n product nodes in the network, v i represents the node corresponding to the i-th product. Each product node has d attributes, and x i ∈ X (i = 1, 2,..., n) represents the attribute vector of the product node v i ; E = {e(i, j)| v i ∈ V, v j ∈ V} is the set composed of all edges in the network. e(i, j) is the edge between the product nodes v i and v j , indicating that a user has purchased the products represented by these two nodes simultaneously. The attribute network contains a total of |E| = m edges; is the matrix composed of the attribute vectors of all product nodes; in addition, the adjacency matrix represents the topological structure of the attribute network. If there is an edge between the product nodes v i and v j , then A ij = 1, otherwise A ij = 0; S22. For each triple (h, r, t) in the domain knowledge graph obtained in S13, denote the embeddings of its head entity and tail entity as h and t respectively, let r denote the embedding of the relationship between entities, and set a mapping matrix M for the relationship r r Project the entity from the entity space to the corresponding relationship space, and obtain the head entity projection vector h respectively r and the tail entity projection vector t r As shown in equations (1)(1) and (2)(2): h r = hM r (1) t r = tM r (2) S23. Further, a transformation is established between the projected entities, and the distance between the projection vectors of the head and tail entities is calculated by the score function of formula (3) to measure the possibility of the fact being true: Among them, f r (h, t) represents the distance score between the head entity h and the tail entity t; S24. Continuously update the embeddings of the triples through iterative steps S22 and S23, extract the knowledge embeddings of the entities corresponding to the products in the e-commerce network from the embedding results, and construct the domain knowledge matrix of the attribute network, denoted as where k i ∈K represents the knowledge features of the product node v i ; S25. Based on the topological structure A and the domain knowledge matrix K of the attribute network, construct a twin network of the attribute network, denoted as and 2. The method for detecting abnormal products in an e-commerce network according to claim 1, wherein, The step S10 of the method comprises the following specific steps: S11. For each product in the e-commerce network, use SPARQL statements to match the head entity or tail entity from the DBpedia knowledge graph and extract all relevant triples for the product; S12, cleaning the triples obtained in step S11, removing triples whose head entities or tail entities are picture links, synonymous entities expressed in other languages, or text information, to reduce the impact of redundancy and noise of triples on the detection results; S13. Use the triples preprocessed in step S12 to form a domain knowledge graph, denoted as where N is a set of entities, R is a set of relationships, and the relationships in R connect two entities to form a triple (h, r, t) ∈ T, h ∈ N is the head entity, t ∈ N is the tail entity, r is the relationship between entities, T represents the set of triples, and each triple describes a piece of knowledge about products in the real world.

3. The method for detecting abnormal products in an e-commerce network according to claim 2, wherein The abnormal product detection model of the e-commerce network in step S30 of the method comprises four parts: a multi-view encoder consisting of an attribute encoder and a knowledge encoder, an aggregator, a decoder and a detector, wherein: 1) The attribute encoder is composed of two stacked attention layers, which encode the attribute network into a low-dimensional embedding of product nodes in the latent space under the attribute view; among them, the first layer uses the multi-head attention mechanism, and the second layer uses the single-head attention mechanism. The formal representation of the encoding process is shown in Equation (4): In the formula, is the low-dimensional embedding of the product node in the hidden space output by the attribute encoder, h' is the dimension of the product node embedding, and the functions f1(·,·) and f2(·,·) represent the first and second graph attention layers respectively; for each graph attention layer, in the process of aggregating the neighborhood information of the product nodes in the attribute network, the graph attention mechanism assigns different attention coefficients to each product node in the neighborhood; the product node v i The attention coefficient between it and its neighbors is calculated by Equation (5): where, e ij is the attention coefficient between product nodes v i and v j , representing the importance of v j to v i ; is the set of first-order neighbor nodes of product node v i ; is a learnable weight matrix applied to each product node, || represents the concatenation operation, is the weight vector; The LeakyReLU(·) function is defined as in Equation (6): Among them, the α value controls the gradient size of the linear function in the negative part; In order to make the attention coefficients between product nodes easier to compare, the attention coefficients are normalized using formula (7): where α ij is the result after normalizing the attention coefficient between product nodes v i and v j ; the output feature of product node v i in the single-head attention graph attention layer is obtained by the linear combination of the normalized attention coefficient and the input features of neighbor product nodes: Among them, x i ' represents the embedding of the product node v output by the graph attention layer under single-head attention, and σ(·) is the activation function, which realizes the non-linear transformation of the data; i ​ The above steps update the product node embedding by aggregating neighborhood information through a single-head attention mechanism; Furthermore, to enhance the generalization ability of the model, a multi-head attention mechanism is introduced in the graph attention layer. A set of attention coefficients are calculated respectively by multiple independent single-head attention networks, and then the multiple embedding results are concatenated through Equation (9) to obtain the product node v i Output features of the multi-head attention graph attention layer: where x i ″ is the embedding of the product node v i output by the graph attention layer with multi-head attention; P represents the number of single-head attention networks, represents the attention coefficient calculated by the p-th attention network, represents the weight matrix of the p-th attention network; 2) The knowledge encoder has exactly the same structure as the attribute encoder and encodes the twin network into the low-dimensional embedding of the product node in the latent space, and its composition is shown in Equation (10): wherein, is the low-dimensional embedding of the product node output by the knowledge encoder in the latent space, and the functions g1(·, ·) and g2(·, ·) respectively represent the first and second graph attention layers in the knowledge encoder; 3) The aggregator fuses the network embeddings under different views to obtain a unified multi-view embedding of product nodes in the network. Either of the two view aggregation strategies can be used, as follows: (1) Concatenation Aggregator: The product node embeddings of the two views are concatenated vertically to form a vector with a higher dimension, as shown in Equation (11): where concat(·) represents a vertical concatenation operation, represents the multi-view unified embedding formed after concatenation aggregation; (2) Addition aggregator: The embeddings of the two views are added together at corresponding positions to obtain a new vector of the same dimension, as shown in Equation (12): wherein, is the multi-view unified embedding obtained using the sum aggregator; 4) The decoder includes a structure decoder and a feature decoder, wherein: The structural decoder reconstructs the network topology by calculating the inner product between the unified embeddings of the product nodes, as shown in formula (13): wherein, is the reconstructed network adjacency matrix output by the structure decoder, U is the multi-view unified embedding of product nodes obtained by using any aggregator, and the sigmoid(·) function is defined as in Equation (14): The feature decoder uses a two-layer fully connected network to reconstruct the product node features, as shown in formula (15): Among them, is the feature information of the reconstructed product node, and b (l) respectively represent the parameter matrix and the bias vector of the l-th fully connected layer, where l ∈ {1, 2}; the product node features output by this decoder are the reconstruction of the two parts of information, namely attribute information and domain knowledge, in the same feature space. 5) The detector uses the reconstruction errors of product nodes in two dimensions, namely structure and feature, as an important basis for evaluating the anomaly scores of products. The specific anomaly scoring function is shown in Equation (16): where score(v i ) is the anomaly score of product node v i , which is used to reflect the anomaly degree of the i-th product. The larger its value, the higher the anomaly degree of the product; a i represents the topological structure of product node v i in the attribute network, and are the features and topological structure of product node v i after decoder reconstruction respectively; λ ∈ [0, 1] is the balance coefficient, which is used to balance the proportion of the reconstruction errors of the features and topological structure of product node v i in the anomaly score.

4. The method for detecting abnormal products in an e-commerce network according to claim 3, characterized in that The steps S40 of this method include the following specific steps: S41. Establish the combined objective function shown in Equation (17) for the e-commerce network anomaly product detection model established in S30: where, is the model feature reconstruction error, defined by the F-norm in Equation (18): Since the input of the decoder is the unified embedding fused by the aggregator, the attributes in the attribute network and the twin network are used as the common reference when calculating the feature reconstruction error; is the structural reconstruction error, defined by Equation (19): In the formula, the balance coefficient λ is the same as λ in Equation (16) and is used to adjust the weights between the two reconstruction errors; S42. Initialize the e-commerce network anomaly product detection model in step S30, and initialize the encoder parameters W enc and c, the decoder parameters W dec and b, the hidden layer embedding dimension dim, the number of attention heads nheads, the balance coefficient λ, the number of iterations epoch, and the learning rate learning rate; Iteratively execute steps S43 to S48 until the number of iterations is reached, complete the training of the e-commerce network anomaly product detection model, and obtain the optimal parameters of the model: S43. Use the domain knowledge graph obtained in S13 as input, learn the knowledge embeddings of entities according to formulas (1), (2) and (3), extract the entity knowledge vectors corresponding to products in the e-commerce network to form a domain knowledge matrix K, and construct a siamese network in combination with the topological structure of the attribute network ​ S44. Obtain the low-dimensional embeddings of the product nodes in the attribute network using the attribute encoder of Equation (4) and the knowledge encoder of Equation (10) respectively and the low-dimensional embeddings of the product nodes in the siamese network S45. Aggregate the embeddings of the two networks using Equation (11) or (12) to obtain the unified embedding U of the product node; S46. Decode and reconstruct the product node features and structures using (13) and (15) to obtain and S47. Adopt the stochastic gradient descent method to complete the backpropagation by optimizing the joint objective function in Equation (17), and realize the update of the encoder weight matrix W and the weight vector c, and the decoder weight matrix W enc and the bias vector b; dec ​ S48. Calculate the anomaly score score(v i ) for each product in the e-commerce network using the detector of formula (16).

5. The method for detecting abnormal products in an e-commerce network according to claim 4, wherein The steps S50 of this method include the following specific steps: S51. After obtaining the optimal parameters of the e-commerce network anomaly product detection model through the training process of iteratively executing steps S43 to S48, use the anomaly detection results obtained in the last training as the final detection results; S52. Output the anomaly product detection results to the operation and supervision personnel of the e-commerce platform to improve the efficiency and reliability of their anomaly product detection, and conduct further targeted processing for the harm degree and risk impact of the anomaly products.

6. An e-commerce network abnormal product detection system, characterized in that: A device for implementing the e-commerce network anomaly product detection method according to any one of claims 1-5, comprising a computer processor and memory, a preprocessing unit for e-commerce network domain knowledge graph data, a training unit for the e-commerce network anomaly product detection model, and an output unit for the e-commerce network anomaly product detection results; the preprocessing unit for e-commerce network domain knowledge graph data executes step S10, preprocesses the collected e-commerce network domain knowledge graph data, and loads it into the computer memory; The training unit for the e-commerce network anomaly product detection model executes steps S20 to S40 according to the e-commerce network domain knowledge graph generated by the preprocessing unit for e-commerce network domain knowledge graph data, constructs the e-commerce network anomaly product detection model, and determines the optimal values of the parameters in the model through iterative calculations; the output unit for the e-commerce network anomaly product detection results executes step S50, outputs the e-commerce network anomaly product detection results to relevant staff or researchers for anomaly product detection of each e-commerce platform and related tasks of network security detection; the specific data processing and calculation work in all units is completed by the computer processor, and all units interact with the computer memory data.