A commodity classification method based on distance geometry

By optimizing the product representation vector based on distance geometry, the problem of poor performance of existing graph neural networks in different graphs is solved, high-precision product classification is achieved, and the model interpretability and computing efficiency are maintained in the airspace.

CN116975730BActive Publication Date: 2025-07-11RENMIN UNIVERSITY OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310975113.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-03
Publication Date
2025-07-11
Estimated Expiration
2043-08-03

AI Technical Summary

Technical Problem

The existing graph neural network performs poorly when processing heterogeneous graphs, and the existing general graph neural network fails to fully consider geometric angles when designing airspace, resulting in low classification accuracy when classified products and limited generalization capabilities.

Method used

The distance geometry-based method is used to calculate the similarity and distance between products through a multi-layer perceptron and hyperbolic tangent function. Combined with the Adam optimizer and L2 regular terms, the product representation vector is optimized, which is suitable for homogeneous and different graphics.

Benefits of technology

The accuracy of product classification is improved, suitable for a variety of graph structures, the model is interpretable in the airspace, and the calculation complexity does not increase.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116975730B_ABST
    Figure CN116975730B_ABST
Patent Text Reader

Abstract

The present invention realizes a commodity classification method based on distance geometry. By S1, the commodities and the relationships between the commodities are modeled into a graph, S2 obtains an initial embedding matrix, S3 learns the distance information between the commodities, S4 sets hyperparameters and the number of propagation times, S5 reduces the dimension of the matrix obtained after propagation to obtain a prediction matrix, S6 uses a cross-entropy loss function to calculate the loss and optimizes it with a backpropagation algorithm, S7 repeats until the classification accuracy of the algorithm on the validation set no longer improves within a certain number of iteration steps, then stops updating the model parameters; inputs the features of the test set commodities into the model, and finally obtains the output prediction. Thus, the technical effects of improving the classification accuracy, coping with the scenarios of homogeneous and heterogeneous matching diagrams, having interpretability in the airspace, and having advantages in terms of computational complexity are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a commodity classification method based on distance geometry. Background Art

[0002] Graph data: In the era of big data, the growth rate of data volume in social life is astonishing. The organization form of data has gradually changed from mainly structured data such as tables to mainly unstructured data such as images, texts, graphs, etc. Among them, graph data is a data structure abstracted from entities and the relationships between entities. Common graph data in real life includes social networks, molecular graphs, guarantee relationship networks, etc.

[0003] Graph neural network: A graph neural network is a special neural network, which is widely used in tasks such as traffic prediction, recommendation systems, and molecular property prediction in the industrial field. Taking the recommendation system based on graph neural network as an example, first, users and commodities are abstracted into nodes, and the attention relationships between users, the interaction relationships between users and commodities, the similarity relationships between commodities, etc. are modeled as edges of the graph. Then, the structure of the graph together with the features of users and commodities is input into the graph neural network to learn the representation vectors of users and commodities, and then ranking recommendations are made based on the similarity between the representation vectors of users and commodities.

[0004] Geometric graph neural network: A geometric graph neural network is a special type of graph neural network. The characteristic that distinguishes this type of graph neural network from other graph neural networks is that they consider geometric symmetries such as equivariance and invariance when designing the network structure. Taking the molecular classification based on graph neural network as an example, first, the atoms in the molecule are abstracted into nodes, and the chemical bonds between molecules are modeled as edges. Then, the structure of the graph together with the type information and relative positions of the molecules are used as features and input into the geometric graph neural network together to learn the representation vector of the entire molecular graph, and then it is used for the classification prediction of molecules.

[0005] Homophily and heterophily of graphs: If the nodes connected by edges in a graph tend to belong to the same category or have some same features, then this graph is said to have homophily. Conversely, a graph that does not have this property is called heterophilic. Most existing graph neural networks make the assumption that the graph has homophily unconsciously when aggregating the features of neighbor nodes. This assumption results in the representation vectors of each node learned by the graph neural network being very close and difficult to distinguish. At the same time, there are also many graphs in reality that do not satisfy homophily. For example, when there are not only positive feedback relationships such as "attention" and "like" in the user-commodity interaction graph, but also negative feedback relationships such as "do not want to see" and "dislike", the graph neural network based on the homophily assumption will ignore these negative relationships, resulting in poor performance of the graph neural network when facing these graphs.

[0006] The universality of graph neural networks has recently been widely studied. Designing a general graph neural network means that this graph neural network can handle more types of graphs, such as paper citation networks with assortativity, social media attention networks, and molecular graphs with disassortativity, etc., thus having a wider range of application scenarios. Most of the existing general graph neural networks start from the spectral domain perspective and try to design graph neural networks that can fit arbitrary filters. However, there is still no clear theoretical basis and practical method for designing a general graph neural network for commodity classification from the spatial domain perspective, especially the geometric perspective.

[0007] The first prior art is the Graph Isomorphism Network (GIN). This technology is a spatial domain graph neural network technology. A spatial domain graph neural network refers to designing the message passing process on a graph to learn the abstract representation of nodes by exchanging information between nodes. The Graph Isomorphism Network is the first model to relate the Weisfeiler-Lehman (WL) graph isomorphism test algorithm to the expressive power of graph neural networks. When this model passes messages, it aggregates the features of surrounding nodes and the features of this node, and then passes them through a multi-layer perceptron (MLP) to obtain the updated representation of this node. The authors of this model claim that the expressive power of all graph neural networks based on the message passing mechanism is bounded by the WL graph isomorphism test under certain premises, that is, the expressive power of any such model will not be stronger than the WL graph isomorphism test. The authors of this model also claim that the Graph Isomorphism Network they designed has reached this upper bound.

[0008] (1) The expressive power is actually relatively limited. Using "whether it can simulate the WL graph isomorphism test algorithm" as the criterion for measuring the expressive power of graph neural networks may not be reasonable. In fact, the graph isomorphism test only tries to determine whether the topological structures of two graphs are the same and does not care about the specific meanings of the nodes in the graph. In addition, this technology only considers the topological structure of the graph and ignores the distances between the representation vectors of each node in space, which may also limit the expressive power of this graph neural network.

[0009] (2) The universality is limited and it cannot be generalized to heterogeneous graphs. The Graph Isomorphism Network still uses the assortativity assumption when aggregating neighbor nodes. Therefore, the Graph Isomorphism Network performs well on assortative graphs, but performs poorly on disassortative graphs and has poor generalization.

[0010] The second prior art is the Equivariant Graph Neural Network (EGNN). This technology is also a spatial domain graph neural network. The Equivariant Graph Neural Network makes relevant modifications to the general spatial domain graph neural network and takes geometric symmetry into account when designing the message passing formula. More specifically, when this model aggregates the messages of surrounding neighbor nodes, it not only considers the representation vectors of neighbor nodes themselves, but also considers geometric features such as the distances between the representation vectors of the central node and neighbor nodes, thus introducing geometric symmetry to the model.

[0011] (1) In the design, only geometric symmetry is considered, without considering geometric generality. Specifically, the model considers geometric symmetry, that is, the representation vector learned after rotating the eigenvector of each node on the graph is equal to the vector obtained by directly inputting the original representation vector into the geometric graph neural network and then rotating it. However, having such symmetry does not necessarily mean that the model has generality. For example, it cannot be guaranteed that the representation vectors learned by dissimilar nodes are also dissimilar. Summary of the Invention

[0012] To this end, the present invention first proposes a commodity classification method based on distance geometry, including steps S1 - S7:

[0013] Step S1: Model the commodities and the relationships between them into a graph G, where the commodities are nodes, the relationships between the commodities are edges, and are represented by an adjacency matrix A. The features attached to all nodes form a matrix X, and the nodes in the training set are attached with labels y;

[0014] Step S2: First, use a multi - layer perceptron to perform dimensionality transformation on the original features X of the commodities, reducing from the original F dimensions to d dimensions to obtain an initial embedding matrix Z0, where Z0 = ReLU(ReLU(XW a +b a )W b +b b );

[0015] Step S3: Learn the distance information between commodities from Z0;

[0016] Step S4: Set hyperparameters α, β, γ, the number of propagations is K, and execute the following process K times. Denote the commodity embedding matrix obtained after k propagations as Z k ;

[0017] Step S5: Reduce the dimension of the matrix Z k obtained after K propagations through a linear transformation, reducing from the original d dimensions to the number of categories c dimensions to obtain a prediction matrix The mathematical form is

[0018] Step S6: Use the cross - entropy loss function to calculate the loss of the prediction matrix and the true label Y on the training set, and then use the backpropagation algorithm to optimize all learnable parameters;

[0019] Step S7: Repeat steps S2 to S7 until the classification accuracy of the algorithm no longer improves on the validation set within a certain number of iteration steps, then stop updating the model parameters; input the features of the test - set commodities into the model, and after steps S2 to S6, finally obtain the output prediction.

[0020] The specific implementation of step S3 is divided into the following sub-steps:

[0021] Step S3.1: Use another multi-layer perceptron to perform a transformation with invariant dimensions on the initial embedding matrix Z0 to obtain matrix H, where H = ReLU(ReLU(Z0W c +b c )W d +b d );

[0022] Step S3.2: For each pair of commodities (i, j) connected by an edge in the graph, concatenate the representation vectors corresponding to i and j in H into a vector [H i ||H j , then take the dot product with a learnable vector a, and finally obtain the similarity α between commodities i and j through the hyperbolic tangent function ij where α ij = tanh(a T [H i ||H j );

[0023] Step S3.3: For each pair of commodities (i, j) connected by an edge in the graph, calculate the distance M between commodities according to the following formula ij : ij : where ε is a small positive number, such as 10 -5 ; Z 0,i and Z 0,j are the representation vectors corresponding to commodities i and j in the initial embedding matrix Z0; ||x|| represents the norm of the vector.

[0024] Step S4 also includes, in the k-th propagation process, for each commodity node i, execute the following steps:

[0025] Step S4.1: For each neighbor commodity node j of commodity i, calculate and then sum them up, denoted as Z k,i and Z k,j are the representation vectors corresponding to commodities i and j in matrix Z k ;

[0026] Step S4.2: For each neighbor commodity node j of commodity i, calculate and then sum them up, denoted as

[0027] Step S4.3: Update the representation vector of commodity i to T3 = (1 - α)T1 + βT2 + αZ 0,i .

[0028] Step S4.4: Pass the representation vector T3 of product i through a linear transformation W k-1 to obtain T4, i.e., T4 = T3W k-1 ;

[0029] Step S4.5: Update the representation vector of product i to Z k,i = ReLU(γT4+(1 - γ)T3).

[0030] In step S6, the Adam optimizer is used. Define the learning rate of the optimizer and the L2 regularization term as hyperparameters, and select the hyperparameters on the validation set.

[0031] The technical effects to be achieved by the present invention are as follows:

[0032] (1) When classifying products, the present invention can improve the classification accuracy compared with the existing technologies described above.

[0033] (2) The present invention is applicable not only to homogeneous graphs but also to heterogeneous graphs. As long as the distance M_ij between product nodes of the same category is set to be as small as possible, and the distance M_ij between product nodes of different categories is set to be as large as possible, the present invention can make the learned representation vectors of products satisfy the distance requirements as much as possible, so as to handle homogeneous graph and heterogeneous graph scenarios.

[0034] (3) The present invention has interpretability in the spatial domain. When the model classifies products, the larger the learned distance M_ij, the less similar the two products are, and vice versa.

[0035] (4) Compared with other existing technologies, the present invention only adds a module for predicting the similarity and distance between products, and does not increase the complexity of the model in terms of computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Product classification method architecture based on distance geometry; DETAILED DESCRIPTION OF THE INVENTION

[0037] The following are the preferred embodiments of the present invention in combination with the drawings. The technical solutions of the present invention are further described, but the present invention is not limited to this embodiment.

[0038] The present invention proposes a product classification method based on distance geometry. It includes steps S1 - S7:

[0039] Step S1: Model the products and the relationships between products into a graph G, where the products are nodes, the relationships between products are edges, and the adjacency matrix A is used to represent them. The features attached to all nodes form a matrix X, and the training set nodes are attached with labels y;

[0040] Step S2: First, use a multi-layer perceptron (MLP) to perform a dimensionality transformation on the original features X of the commodity, reducing the dimension from the original F dimensions to d dimensions, obtaining the initial embedding matrix Z0. The mathematical form is Z0 = ReLU(ReLU(XW a +b a )W b +b b );

[0041] Step S3: Learn the distance information between commodities from Z0. Specifically, it is divided into the following sub-steps:

[0042] Step S3.1: Use another multi-layer perceptron (MLP) to perform a transformation with unchanged dimensions on the initial embedding matrix Z0 to obtain the matrix H. The mathematical form is H = ReLU(ReLU(Z0W c +b c )W d +b d ).

[0043] Step S3.2: For each pair of commodities (i, j) connected by an edge in the graph, concatenate the representation vectors corresponding to i and j in H into a vector [H i ||H j , then take the dot product with a learnable vector a, and finally obtain the similarity α ij between commodities i and j through the hyperbolic tangent function (tanh x). ij The mathematical form is α T = tanh(a i [H j ).

[0044] Step S3.3: For each pair of commodities (i, j) connected by an edge in the graph, calculate the distance M ij between commodities according to the following formula ij : where ε is a small positive number, such as 10 -5 ; Z 0,i and Z 0,j are the representation vectors corresponding to commodities i and j in the initial embedding matrix Z0; ||x|| represents the norm of the vector.

[0045] Step S4: Set the hyperparameters α, β, γ, the number of propagation times as K, and execute the following process K times. Denote the commodity embedding matrix obtained after k times of propagation as Z k .

[0046] Step S4.1: In the k-th propagation process, for each commodity node i, execute the following steps:

[0047] Step S4.1.1: For each neighbor product node j of product i, calculate Then sum them up, denoted as Z k,i And Z k,j are the representation vectors corresponding to products i and j in matrix Z k .

[0048] Step S4.1.2: For each neighbor product node j of product i, calculate Then sum them up, denoted as

[0049] Step S4.1.3: Update the representation vector of product i to T3 = (1 - α)T1 + βT2 + αZ 0,i .

[0050] Step S4.1.4: Pass the representation vector T3 of product i through a linear transformation W k-1 to get T4, that is, T4 = T3W k-1 .

[0051] Step S4.1.5: Update the representation vector of product i to Z k,i = ReLU(γT4 + (1 - γ)T3).

[0052] Step S5: Reduce the dimension of the matrix Z k obtained after K - times of propagation through a linear transformation, from the original d - dimension to the number of categories c - dimension, to get the prediction matrix The mathematical form is

[0053] Step S6: Use the cross - entropy loss function to calculate the loss of the prediction matrix and the true label Y on the training set, and then use the backpropagation algorithm to optimize all learnable parameters. The present invention uses the Adam optimizer, defines the learning rate of the optimizer and the L2 regularization term as hyperparameters, and selects hyperparameters on the validation set.

[0054] Step S7: Repeat steps S2 to S7 until the classification accuracy of the algorithm on the validation set no longer improves within a certain number of iteration steps, then stop updating the model parameters. Input the features of the test - set products into the model, and after steps S2 to S6, finally obtain the output prediction.

Claims

1. A commodity classification method based on distance geometry, characterized in that: Including steps S1 - S7: Step S1: Model the products and the relationships between products into a graph G, where the products are nodes, the relationships between products are edges, represented by an adjacency matrix A, the features attached to all nodes form a matrix X, and the nodes in the training set are attached with labels Y; Step S2: First, use a multi-layer perceptron to perform dimensionality transformation on the original features X of the commodity, reducing them from the original F dimensions to d dimensions to obtain the initial embedding matrix Z0, where Z0 = ReLU(ReLU(XW a +b a )W b +b b ); in the above formula, W a is a parameter matrix of F×d dimensions, b a is a parameter vector of d dimensions, W b is a parameter matrix of d×d dimensions, b b is a parameter vector of d dimensions; Step S3: Learn the distance information between products from Z0; Step S4: Set the hyperparameters α, β, γ as real numbers in [0, 1], set the number of propagations as K, and execute the following process K times. Denote the commodity embedding matrix obtained after k propagations as Z k ; Step S5: The matrix Z obtained after K - time propagation k is dimension - reduced through a linear transformation, from the original d dimensions to c dimensions (the number of categories), to obtain the prediction matrix The mathematical form is Step S6: Calculate the loss of the prediction matrix and the true label Y on the training set using the cross-entropy loss function, and then use the backpropagation algorithm to optimize all learnable parameters; where the cross-entropy loss function is defined as N is the number of products, M is the number of product categories, and y ic is the product category label. When product i belongs to category c, y ic = 1, otherwise 0, and p ic is the probability that the model predicts product i belongs to category c; Step S7: Repeat steps S2 to S7 until the classification accuracy on the validation set no longer improves after no more than 1500 iterations of the algorithm, then stop updating the model parameters; Input the features of the test set products into the model, and after steps S2 to S6, finally obtain the output prediction result, that is, the category to which the product belongs; Step S4 further includes that in the k-th propagation process, for each product node i, perform the following steps: Step S4.1: For each neighbor product node j of product i, calculate Then sum them up and denote it as where d i and d j are the degrees of product i and j respectively, that is, the number of their respective neighbor products, Z k,i and Z k,j are the representation vectors corresponding to products i and j in the matrix Z k ; Step S4.2: For each neighbor product node j of product i, calculate Then sum them up, denoted as where M ij is the distance between products i and j calculated in step S3.3; Step S4.3: Update the representation vector of product i to T3 = (1-α)T1 + βT2 + αZ 0,i ; where α and β are hyperparameters set in step S4; Step S4.4: Pass the representation vector T3 of product i through a linear transformation W k-1 to obtain T4, i.e., T4 = T3W k-1 ; Step S4.5: Update the representation vector of product i to Z k,i = ReLU(γT4+(1 - γ)T3); where γ is the hyperparameter set in step S4; In step S6, the Adam optimizer is used, and the learning rate of the optimizer and the L2 regularization term are defined as hyperparameters taking values in the range of {0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1}, and all hyperparameters are selected on the validation set.

2. The method for classifying commodities based on distance geometry according to claim 1, wherein: The specific implementation method of step S3 is divided into the following sub-steps: Step S3.1: Use another multi-layer perceptron to perform a dimension-invariant transformation on the initial embedding matrix Z0 to obtain matrix H, where H = ReLU(ReLU(Z0W c +b c )W d +b d ); In the above formula, W c , W d are both d×d-dimensional parameter matrices, and b c , b d are both d-dimensional parameter vectors; Step S3.2: For each pair of products (i, j) connected by an edge in the graph, concatenate the representation vectors corresponding to i and j in H into a vector [H i ||H j , then perform a dot product with a parameter vector a of dimension 2d, and finally obtain the similarity s between products i and j through the hyperbolic tangent function ij ; where s ij =tanh(a T [H i ||H j ); Step S3.3: For each pair of products (i, j) connected by an edge in the graph, calculate the distance M between the products according to the similarity s obtained in the previous step ij using the following formula ij : where ε is a small positive number, such as 10 -5 ; Z 0,i and Z 0,j are the representation vectors corresponding to products i and h in the initial embedding matrix Z0; ||x|| represents the norm of the vector.