Customs Data Classification Method, Device and Storage Medium Based on Knowledge Graph Representation
Through the customs data classification method based on knowledge graph representation, the problem that the digital expression of noun entities in customs data is difficult to retain overall information, and efficient customs data prediction is achieved, which improves processing efficiency and accuracy.
Patent Information
- Application Number
- CN202210662243.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-13
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-06-13
AI Technical Summary
The digital expression of noun entities in customs data is difficult to retain overall information. Traditional language models require complex and complex machine learning models when extracting features, resulting in low operational efficiency and unable to meet the customs' needs for improving prediction speed and accuracy.
Using a customs data classification method based on knowledge graph representation, a triple is constructed by extracting the attributes of noun entities, and an embedded representation is carried out. Combining BiLSTM and translation modules, an efficient pre-trained representation is generated, and a lightweight classification module is constructed for high-precision declaration factor prediction.
Effectively distinguishing most entities in customs data, improving the accuracy and speed of declaration factor prediction, reducing the complexity and calculation cost of the model, and improving the processing efficiency of customs data.
Smart Images

Figure CN115098694B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of customs data, and particularly relates to a knowledge graph representation model based on customs data. Background Art
[0002] With the rapid development of current international trade and cross-border e-commerce, the tax supervision business of customs faces multiple difficulties such as an unprecedented variety of commodity types and more complex and diverse trade forms. The requirements for the speed, accuracy, and risk control of tax collection are also constantly increasing. The customs' import and export commodity specification declaration catalog stipulates the declaration elements for different types of commodities. However, in actual declarations, due to factors such as some elements being related to tariffs, there are problems of incorrect filling and concealment of these elements. Therefore, based on the basic declaration elements of commodities, analyze and identify the characteristics of commodities, and use the identified characteristics to detect the basic declaration elements to trace incorrect filling and concealment. Control the risks of customs import and export.
[0003] Customs basic information includes a large amount of text information such as commodity names, store names, and commodity descriptions. However, different from the text in news or communication, the core information in customs data is contained in a large number of noun entities, rather than the logic of the context. In addition, each commodity also has a proprietary entity description in its professional field. Therefore, commodity feature recognition requires not only a comprehensive understanding of the general attributes of commodities but also professional field knowledge for different commodities. It is difficult to widely extract effective entity information through some text processing tools or models.
[0004] Currently, the following problems exist in customs multi-attribute (declaration element) data: Noun entities are easy to find, but the digital expression for noun entities can only use traditional language models. When segmenting noun entities and constructing word vector expressions word by word, the overall information of the entity is lost. In addition, in the case of using traditional language models as the expression of customs data, a redundant and complex machine learning model is required to extract features to ensure a high classification accuracy. RNN-based models cannot be applied to parallel operations, and the operation efficiency is slow. In applications, the feedback speed is also an important factor in improving office efficiency. Prediction of the country of origin, port prediction, and tariff number prediction are all important concerns in customs declaration element prediction. To improve the accuracy and speed of prediction, it is necessary to explore a representation model applicable to customs data that can distinguish most entities in customs data as much as possible to solve the current bottleneck. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a customs data classification method, device and storage medium based on knowledge graph representation. The core of the invention is an embedding representation module based on the knowledge graph. Based on the pre-trained embedding representation module, the noun entities in the customs data are emphasized and effectively distinguished. On this basis, a lightweight classification module can be constructed to predict declaration factors with high accuracy.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is:
[0007] The present invention first provides a customs data classification method based on knowledge graph representation, which is characterized by including the following steps:
[0008] Step 1: Extract the attributes containing noun entities in the customs data;
[0009] Step 2: Use the entities extracted in Step 1 to construct triples, and split each customs data into multiple triples: the head entity, relation, and tail entity in the triple respectively correspond to the data serial number, attribute name, and the value of the corresponding attribute in this data in the customs data;
[0010] Step 3: Perform word segmentation and stop word removal on the tail entity in Step 2 to obtain the serialized representation of the text;
[0011] Step 4: Pass the head entity, relation, and tail entity obtained in Step 2 through three different embedding layers to embed the head entity, relation, and tail entity into low-dimensional vectors. At the same time, perform word embedding on the serialized representation of the tail entity in Step 3 and process it through a layer of BiLSTM into equal-length features. The combination of the embedding vectors of the head entity and relation, and the embedding vector of the tail entity and the BiLSTM output features are respectively recorded as h, l, t;
[0012] The combination of the embedding vectors of the head entity and relation, and the embedding vector of the tail entity and the BiLSTM output features are respectively recorded as h, l, t;
[0013] Step 5: Construct a translation module, use h and l as inputs, and output a translation matrix;
[0014] Step 6: Multiply the translation matrix obtained in Step 5 by h to obtain the translation result t* of h, and calculate the distance between t and t*;
[0015] Step 7: Use the calculation result in Step 6 to calculate the loss function, and train the embedding layer module and the translation module under the condition of no supervision information;
[0016] Step 8: Use the embedding layer and BiLSTM layer of t as the pre-trained language representation method for each attribute of the customs data. For a single piece of data, splice the representations obtained from multiple attributes as the input data for classification;
[0017] Step 9: Input the data from Step 6 into the classification module to extract data features;
[0018] Step 10: Expand the data features obtained in Step 9 into one-dimensional vectors and splice them, perform classification through two fully connected layers, and use cross-entropy loss to train the classification module and the fully connected layers;
[0019] Step 11: Classify the customs data using the classification module and the fully connected layers trained in Step 10.
[0020] The present invention also provides a customs data classification device based on knowledge graph representation, including a processor and a memory; a program or instruction is stored in the memory, and the program or instruction is loaded and executed by the processor to implement the customs data classification method based on knowledge graph representation.
[0021] The present invention also provides a computer-readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the customs data classification method based on knowledge graph representation is implemented.
[0022] In Step 1, each piece of customs data contains multiple attributes, and the attributes may include information such as commodity name, production enterprise name, commodity description, production time, etc. In order to construct a representation model conducive to classification, descriptive attributes of customs commodities are extracted. For example, commodity name, production enterprise, commodity description, etc., and attributes that cannot describe commodity characteristics such as production time, commodity serial number, etc. are discarded.
[0023] In Step 2, the customs data consists of multiple multi-attribute data. Since the relationships between different attributes are unknown, and the final classification task does not focus on the connections between multiple attributes, but generates an input representation conducive to classification through a knowledge graph, when constructing the knowledge graph triples, the form of (data serial number, attribute name, value of the corresponding attribute in this piece of data) is used to represent the triples, and the number of triples generated for each piece of data is determined according to the number of attributes selected in Step 1.
[0024] In Step 3, the long Chinese text in the tail entity is word-segmented character by character, encoded according to the dictionary in jieba word segmentation. At the same time, for some short texts, such as attributes like commodity name and e-commerce name, all the commodity names and e-commerce names that appear in the training set are used as the dictionary, and they are encoded without word-segmentation in subsequent steps. The relationships in the triples are encoded using the method of directly encoding without word-segmentation.
[0025] In step 4, three embedding layers are used to process the direct encodings of the head entity, relation, and tail entity. Each piece of data is converted into three vectors, h, l, and t, with lengths of 40, 50, and 50 respectively. A new embedding layer processes the tokenized encoding features of the tail entity into a size of (T, 64), where T is the length of the tokenized sequence. It is processed by a BiLSTM into a feature of length 50, and this feature, together with the embedding layers of the head entity and relation, forms a new set of h, l, and t. Through the process of step 4, each triple with a tail entity in the text is split into two triples, representing the overall language representation and the temporal language representation of the tail entity respectively.
[0026] In step 5, a translation module is constructed, whose purpose is to translate h into t through l. In the constructed knowledge graph, the head entity and the tail entity are closely connected by the relation, but this relation is not simply additive. The translation module is obtained from the following model: The inputs are h and l. h is input into a layer of mlp network, and a vector with the same length as h is output, denoted as f(h); l is input into a layer of mlp network, and a vector with the same length as h is output, denoted as g(l); f(h) and g(l) are concatenated to obtain a feature of (40, 2); finally, it passes through three one-dimensional convolutional layers (2, 16), (16, 32), (32, 50) in sequence, and a matrix of size (40, 50) is output as the translation matrix. The advantage of this approach is that through training, this translation module can learn the deep relationships among the head entity, relation, and tail entity; using the method of multiplying a matrix and a vector can more accurately find the connections in different dimensions after the head entity and tail entity are embedded. For example, compared with some knowledge graph methods that use the addition of h and l as the translation process, this translation module does not limit the dimensions after the head entity, relation, and tail entity are embedded. Assuming there is a close connection between the 3rd dimension of the head entity and the 20th dimension of the tail entity, this translation module has the ability to find this relationship, but the addition method cannot find this relationship.
[0027] In step 6, the vector obtained from the translation module is multiplied by h to get a predicted entity, and this process is denoted as t* = F(h, l). The purpose of translation in the knowledge graph is to find the tail entity given the head entity and the relation. For the predicted tail entity t*, the gap between it and the real tail entity t needs to be found to train an accurate translation module. Step 6 uses the Euclidean distance to represent the gap:
[0028] d(h, l, t) = ||F(h, l) - t|| 2
[0029] In step 7, the form of the loss function is:
[0030]
[0031] Among them, N represents the number of triplets, M represents the sampling of negative example relations from M other triplets, α1 and α2 are empirical constants, and the first term of loss makes clear requirements for h, l, and t from the same triplet, that is, d(h,l,t)=0. In the same triplet, h must be translated from l to t. When d(h,l,t) is large, the gradient generated by loss is large, and the model will converge to the direction of d(h,l,t)=0 as soon as possible. The second term is to use the relations in other triplets to construct negative examples. For a triplet, the sampling of negative example relations includes: 1. Using several other attribute names of the same data as negative example relations and calculating d(h i ,l j ,t i ), the purpose of negative examples is to reduce translation errors and avoid getting many wrong entities after one entity is translated through a relation. When the relation is wrong, the head entity and the tail entity in each triple do not correspond. In the second item, when d(h i ,l j ,t i ) approaches 0, the gradient is large. i ,l j ,t i ) is relatively large, the gradient tends to 0. The purpose of this is to avoid incorrect entity relationship translation as much as possible. Compared with the first item, the second item has a smaller gradient value and a more moderate adjustment of the parameters. The purpose of this is that there may be relatively similar triplets in the triplets constructed from customs data, and the goods from the same store may be recorded repeatedly. Based on this feature, this patent is more inclined to train based on positive samples, while being somewhat inclusive of negative samples. The training process uses the Adam optimizer to train the translation module, embedding layer and BiLSTM simultaneously. In order to prevent overfitting, each layer of the neural network in each module is followed by a Dropout operation.
[0032] In step 8, the embedding layer and BiLSTM trained in step 7 are used to generate the embedding features of all attributes of each data. For attributes containing text, the above model provides the overall embedding and overall language representation and temporal language representation (provided by BiLSTM) of the attribute, all of which are 50 dimensions. This project uses a total of 3 attributes in customs data, 2 of which are text data. The embeddings of the data in these attributes are concatenated to obtain the input features of (50,5).
[0033] In step 9, a 5-layer one-dimensional convolutional network is constructed as the classification model. Each layer contains a 3-convolution with a one-dimensional convolution kernel, a BN layer, and a ReLU activation layer. The channel sizes of the convolutional layers are (5, 32), (32, 64), (64, 64), (64, 128), and (128, 128) respectively. The finally output features of (50, 128) are used for classification. Compared with many deep models or sequential models that cannot be parallelized, the volume of this classification module is relatively small and the computing speed is relatively fast, but it can still achieve relatively high accuracy. The main reason for this phenomenon is still that the previous knowledge graph representation provides an efficient pre-training representation for the classification model.
[0034] In step 10, the features of (50, 128) are expanded into a vector of 6400, and passed through fully connected layers of (6400, 1024) and (1024, the number of categories) in sequence. The first fully connected layer contains a ReLU activation function, and the result of the second fully connected layer is used to predict the category through softmax. The cross-entropy is used to train the classification module and the fully connected layers. The present invention can predict various key attributes in customs data, such as "port", "country of origin", etc., and these tasks correspond to different "numbers of categories". When switching tasks, it is necessary to retrain the classification module and the fully connected layers.
[0035] The purpose of the knowledge graph technology used in the present invention is similar to that of TransE and its derivative methods, but the method of constructing triples and calculating entity relationships in the present invention is completely different from that of TransE. There are also significant differences in the measurement method and the loss function. TransE is mainly applied to model the relationship between head and tail entities belonging to the same type of logical things, and its method of using addition as translation is obviously inconsistent with the relationships between different entities in customs data. In terms of the construction of the translation model and the design of the loss, the present invention pays more attention to the form of entities in customs data, and for the final embedded representation, the present invention comprehensively considers the overall information and the sequential information, which is a more comprehensive embedded expression. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The following further specific description of the present invention is made in conjunction with the drawings and specific embodiments, and the above or other advantages of the present invention will become clearer.
[0037] Figure 1 is the model framework diagram of the present invention;
[0038] Table 1 shows the comparison of the classification metrics of the port and country of origin attributes of customs data on the validation set with the results of other methods. The comparison is made through the average accuracy rate. By comparing with the BiLSTM, TextCNN, and Bert methods, the average accuracy rate of the present method on customs data is currently the highest. DETAILED DESCRIPTION OF THE INVENTION
[0039] The present invention will be further described below in conjunction with embodiments.
[0040] Embodiment 1
[0041] Referring to the method flow of the present invention, a customs data classification method based on knowledge graph representation of the present invention includes the following steps:
[0042] Step 1: Taking a piece of source data as an example, "Zhejiang Tmall Technology Co., Ltd.|Hangzhou Fudie E-commerce Co., Ltd.|3301968FU0|I|2019 / 3 / 5 5:54|4612|4612|2019 / 3 / 5 0:00|8f55f5f0-142b-472e-b989-6804409e00f0|5.66041E+11|revlon Revlon lipstick for female students, genuine product, long-lasting moisturizing, non-smudging, bargain-hunting dog, bean paste color 225|3304100091|Revlon Colorstay lipstick 445#|4.2g / each; zinc citrate, ozokerite, ceresin; color enhancement and brightening; Revlon; other overseas". According to the customs data, three attributes meaningful for classification are extracted: "e-commerce enterprise name", "product name", and "product description". The data is reorganized as: "Zhejiang Tmall Technology Co., Ltd.|revlon Revlon lipstick for female students, genuine product, long-lasting moisturizing, non-smudging, bargain-hunting dog, bean paste color 225|4.2g / each; zinc citrate, ozokerite, ceresin; color enhancement and brightening; Revlon; other overseas"
[0043] Step 2: Use the entities extracted in Step 1 to construct triples, and split each piece of customs data into multiple triples: (data serial number, attribute name, value of the corresponding attribute in this piece of data), corresponding to the head entity, relationship, and tail entity respectively. Assuming the serial number of the above case is 5, the data is constructed as (5, e-commerce enterprise name, Zhejiang Tmall Technology Co., Ltd.), (5, product name, revlon Revlon lipstick for female students, genuine product, long-lasting moisturizing, non-smudging, bargain-hunting dog, bean paste color 225), (5, product description, 4.2g / each; zinc citrate, ozokerite, ceresin; color enhancement and brightening; Revlon; other overseas);
[0044] Step 3: Segment the tail entity representing the product description in Step 2, remove stop words, and obtain the serialized representation of the text;
[0045] Step 4: Pass the head entity, relationship, and tail entity of the three triples obtained in Step 2 through three different embedding layers to embed the head entity, relationship, and tail entity into low-dimensional vectors. At the same time, perform word embedding on the serialized representation of the product description tail entity in Step 3 and process it through a layer of BiLSTM into features of equal length. The combination of the embedding vectors of the head entity and relationship, and the embedding vector of the tail entity and the BiLSTM output features are respectively recorded as h, l, t;
[0046] Step 5: Construct a translation module that takes h and l as inputs and outputs a translation matrix;
[0047] Step 6: Multiply the translation matrix obtained in Step 5 by h to get the translated result t* of h, and calculate the distance between t and t* using the designed metric;
[0048] Step 7: Use the result of Step 6 to calculate the loss function, and train the embedding layer module and the translation module without supervision information;
[0049] Step 8: Use the embedding layer and the BiLSTM layer of t as the pre-trained language representation method for each attribute of the customs data. For a single piece of data, concatenate the representations obtained from multiple attributes as the input data for classification;
[0050] Step 9: Input the data from Step 6 into the classification module to extract data features, where the classification module consists of 5 convolutional layers;
[0051] Step 10: Expand the features obtained in Step 9 into one-dimensional vectors and concatenate them, perform classification through two fully connected layers, and use cross-entropy loss to train the classification module and the fully connected layers.
[0052] In Step 2, the customs data consists of multiple multi-attribute data. Since the relationships between different attributes are unknown, and the final classification task does not focus on the connections between multiple attributes, but generates input representations conducive to classification through a knowledge graph, when constructing the knowledge graph triples, the form (data serial number, attribute name, value of the corresponding attribute in this piece of data) is used to represent the triples, and the number of triples generated for each piece of data is determined according to the number of attributes selected in Step 1.
[0053] In Step 3, the Chinese text in the tail entity is tokenized character by character and encoded according to the dictionary in jieba tokenization, and the above operations are performed on the "commodity description" attribute. For short text attributes such as "e-commerce enterprise name" and "commodity name", the dictionary is all the commodity names and e-commerce names that appear in the training set, and they are encoded in an untokenized form in the subsequent process.
[0054] In step 4, three embedding layers are used to process the direct encodings of the head entity, relation, and tail entity. Each piece of data is converted into three vectors, h, l, and t, with lengths of 40, 50, and 50 respectively. A new embedding layer processes the tokenized encoding features of the tail entity into a size of (T, 64), where T is the length of the tokenized sequence. It is processed by a BiLSTM into a feature of length 50, and this feature, together with the embedding layers of the head entity and relation, forms a new set of h, l, and t. Through the process of step 4, each triple with a tail entity in the text is split into two triples, representing the overall language representation and the temporal language representation of the tail entity respectively.
[0055] In step 5, a translation module is constructed, whose purpose is to translate h into t through l. In the constructed knowledge graph, the head entity and the tail entity are closely connected by the relation, but this relation is not simply additive. The translation module is obtained from the following model: The inputs are h and l. h is input into a layer of mlp network, and a vector with the same length as h is output, denoted as f(h); l is input into a layer of mlp network, and a vector with the same length as h is output, denoted as g(l); f(h) and g(l) are concatenated to obtain a feature of (40, 2); finally, it passes through three one-dimensional convolutional layers (2, 16), (16, 32), (32, 50) in sequence, and a matrix of size (40, 50) is output as the translation matrix. The advantage of this approach is that through training, this translation module can learn the deep relationships among the head entity, relation, and tail entity; using the method of multiplying the matrix and the vector can more accurately find the connections in different dimensions after the head entity and the tail entity are embedded.
[0056] In step 6, the vector obtained from the translation module is multiplied by h to get a predicted entity, and this process is denoted as t* = F(h, l). The purpose of translation in the knowledge graph is to find the tail entity given the head entity and the relation. For the predicted tail entity t*, the gap between it and the real tail entity t needs to be found to train an accurate translation module. Step 6 uses the Euclidean distance to represent the gap:
[0057] d(h, l, t) = ||F(h, l) - t|| 2
[0058] In step 7, the form of the loss function is:
[0059]
[0060] Among them, N represents the number of triplets, M represents the sampling of negative example relations from M other triplets, α1 and α2 are empirical constants, and the first term of loss makes clear requirements for h, l, and t from the same triplet, that is, d(h,l,t)=0. In the same triplet, h must be translated from l to t. When d(h,l,t) is large, the gradient generated by loss is large, and the model will converge to the direction of d(h,l,t)=0 as soon as possible. The second term is to use the relations in other triplets to construct negative examples. For a triplet, the sampling of negative example relations includes: 1. Using several other attribute names of the same data as negative example relations and calculating d(h i ,l j ,t i ), the purpose of negative examples is to reduce translation errors and avoid getting many wrong entities after one entity is translated through a relation. When the relation is wrong, the head entity and the tail entity in each triple do not correspond. In the second item, when d(h i ,l j ,t i ) approaches 0, the gradient is large. i ,l j ,t i ) is relatively large, the gradient tends to 0. The purpose of this is to avoid incorrect entity relationship translation as much as possible. Compared with the first item, the second item has a smaller gradient value and the adjustment of parameters is more moderate. The purpose of this is that there may be relatively similar triplets in the triplets constructed from customs data, and the products under the same store may be recorded repeatedly. Based on this feature, this patent is more inclined to train based on positive samples, while having a certain tolerance for negative samples. The training process uses the Adam optimizer to train the translation module, embedding layer and BiLSTM at the same time
[0061] In step 8, the embedding layer and BiLSTM trained in step 7 are used to generate the embedding features of all attributes of each data. For attributes containing text, the above model provides the overall embedding and overall language representation and temporal language representation (provided by BiLSTM) of the attribute, all of which are 50 dimensions. This project uses a total of 3 attributes in customs data, 2 of which are text data. The embeddings of the data in these attributes are concatenated to obtain the input features of (50,5).
[0062] In step 9, a 5-layer one-dimensional convolutional network is constructed as the classification model. Each layer contains a 3-convolution with a one-dimensional convolution kernel, a BN layer, and a ReLU activation layer. The channel sizes of the convolutional layers are (5,32), (32,64), (64,64), (64,128), and (128,128). The final output feature (50,128) is used for classification.
[0063] In step 10, the feature of (50, 128) is expanded into a vector of 6400, and successively passes through fully connected layers of (6400, 1024) and (1024, the number of categories). The first fully connected layer contains a ReLU activation function, and the result of the second fully connected layer is used to predict the category through softmax. The classification module and the fully connected layers are trained using cross-entropy. The present invention can predict multiple key attributes in customs data. In this example, the declaration elements of "port" and "country of origin" are used. The number of categories for the port is 96, and the number of categories for the country of origin is 93.
[0064] Training hyperparameter settings: All methods use a set of training parameters. The data is the import and export data of the customs e-commerce platform, and the training set and test set are divided according to a ratio of 7:3. During training, the optimizer is Adam, where the weight_decay parameter is set to 0.001, the batch size is set to 16, and the number of training epochs is 10. In the comparative method of the embodiment, the bidirectional long short-term memory model uses 256 hidden neurons. In addition to the convolutional layer (the first layer) with multiple convolutional kernels in the text convolutional neural network, 5 subsequent convolutional layers are connected to extract features.
[0065] Table 1 shows the test results in this embodiment. Predictions are made on two attributes, the port and the country of origin, and three current mainstream text neural networks are compared: 1. Bidirectional long short-term memory model (BiLSTM), 2. Text convolutional neural network (TextCNN), 3. Bidirectional Encoder Representations from Transformers (BERT):
[0066] Table 1
[0067] Method Accuracy of checkpoint prediction Country of origin prediction The method of the present invention 0.83 0.59 BiLSTM 0.77 0.46 TextCNN 0.78 0.48 BERT 0.80 0.58
[0068] The experimental results are evaluated using accuracy. The three comparative methods are all the classification effects under traditional word embedding methods. The dimension of word embedding is 300, and the input sequence size is (300, T). In the comparative experiment, all methods are run on the same device and framework, and the word segmentation method is also the same. As can be seen from Table 1, compared with many deep models or sequential models that cannot be parallelized, the volume of this classification module is relatively small, the calculation speed is relatively fast, but it can still achieve a relatively high accuracy. The main reason for this phenomenon is still that the previous knowledge graph representation provides an efficient pre-training representation for the classification model. Among the floating-point operation times of each model, the ranking is "the present invention" < "TextCNN" < "BiLSTM" < "BERT".
[0069] This embodiment provides a customs data classification device based on knowledge graph representation, which is characterized by including a processor and a memory; programs or instructions are stored in the memory, and the programs or instructions are loaded and executed by the processor to implement the customs data classification method based on knowledge graph representation in the embodiment.
[0070] This embodiment provides a computer-readable storage medium, on which programs or instructions are stored, and when the programs or instructions are executed by a processor, the customs data classification method based on knowledge graph representation in the embodiment is implemented.
Claims
1. A customs data classification method based on knowledge graph representation, characterized in that, It includes the following steps: Step 1: Extract the attributes containing noun entities from the customs data; Step 2: Use the entities extracted in Step 1 to construct triples, and split each piece of customs data into multiple triples: the head entity, relation, and tail entity in the triple respectively correspond to the data serial number, attribute name, and the value of the corresponding attribute in this piece of data in the customs data; Step 3: Segment the tail entity in Step 2, remove stop words, and obtain the serialized representation of the text; Step 4: Pass the head entity, relation, and tail entity obtained in Step 2 through three different embedding layers to embed the head entity, relation, and tail entity into low-dimensional vectors. At the same time, perform word embedding on the serialized representation of the tail entity in Step 3 and process it through a BiLSTM layer to obtain features of equal length. The combined embedding vectors of the head entity and relation, as well as the embedding vector of the tail entity and the output features of the BiLSTM are respectively recorded as h, l, t; Step 5: Construct a translation module, and use h and l as inputs to output a translation matrix; Step 6: Multiply the translation matrix obtained in Step 5 by h to obtain the translated result t* of h, and calculate the distance between t and t*; Step 7: Use the calculation result in Step 6 to calculate the loss function, and train the embedding layer module and the translation module under the condition of no supervision information; Step 8: Use the embedding layer and BiLSTM layer of t as the pre-trained language representation method for each attribute of the customs data. For a single piece of data, splice the representations obtained from multiple attributes as the input data for classification; Step 9: Input the data in Step 6 into the classification module to extract data features; Step 10: Expand the data features obtained in Step 9 into one-dimensional vectors and splice them, perform classification through a fully connected layer, and use cross-entropy loss to train the classification module and the fully connected layer; Step 11: Use the classification module and the fully connected layer trained in Step 10 to classify the customs data.
2. The customs data classification method based on knowledge graph representation according to claim 1, wherein In Step 3, the text is encoded according to types. Among them, for the Chinese text of "commodity description" in the tail entity, word segmentation is performed character by character and encoded according to the dictionary in jieba word segmentation; for short texts of attributes such as commodity names and e-commerce names, all the commodity names and e-commerce names that appear in the training set are used as the dictionary and encoded without word segmentation; for the relations in the triples, a method of direct encoding without word segmentation is adopted.
3. The customs data classification method based on knowledge graph representation according to claim 1, wherein In Step 4, use three embedding layers to process the direct encodings of the head entity, relation, and tail entity. Each piece of data is converted into three vectors h, l, t, with lengths of 40, 50, and 50 respectively; use a new embedding layer to process the segmented encoding features of the tail entity into a size of (T, 64), where T is the sequence length after word segmentation; process it through a BiLSTM layer to obtain features of length 50, and this feature and the embedding layers of the head entity and relation form a new set of h, l, t.
4. The customs data classification method based on knowledge graph representation according to claim 3, wherein, The translation module constructed in step 5 is obtained from the following model: The inputs are h and l. h is input into a one-layer MLP network, and a vector with the same length as h is output, denoted as f(h); l is input into a one-layer MLP network, and a vector with the same length as h is output, denoted as g(l); f(h) and g(l) are concatenated to obtain a feature of (40, 2); finally, it passes through three one-dimensional convolutional layers (2, 16), (16, 32), (32, 50) in sequence, and a matrix of size (40, 50) is output as the translation matrix.
5. The customs data classification method based on knowledge graph representation according to claim 1, wherein In step 6, the distance between t and t* is calculated as: d(h, l, t) = ||F(h, l) - t|| 2 In the formula, F(h, l) = t*, which is a predicted entity obtained by multiplying the vector obtained by the translation module with h.
6. The customs data classification method based on knowledge graph representation according to claim 1, characterized in that, In step 7, the form of the loss function is: Where N represents the number of triples, M represents sampling negative examples from M other triples, and α1 and α2 are empirical constants.
7. The customs data classification method based on knowledge graph representation according to claim 1, characterized in that In step 9, a 5-layer one-dimensional convolutional network is constructed as the classification model. Each layer contains a convolution with a one-dimensional convolution kernel of 3, a BN layer, and a ReLU activation layer; the channel sizes of the convolutional layers are (5, 32), (32, 64), (64, 64), (64, 128), (128, 128) respectively; finally, the feature of (50, 128) is output for classification.
8. The customs data classification method based on knowledge graph representation according to claim 7, characterized in that In step 10, the feature of (50, 128) is expanded into a vector of 6400, and it passes through fully connected layers of (6400, 1024) and (1024, the number of categories) in sequence. The first fully connected layer contains a ReLU activation function, and the result of the second fully connected layer is used to predict the category through softmax.
9. A customs data classification device based on knowledge graph representation, characterized in that, It includes a processor and a memory; programs or instructions are stored in the memory, and the programs or instructions are loaded and executed by the processor to implement the customs data classification method based on knowledge graph representation as described in any one of claims 1 to 8.
10. A computer-readable storage medium, on which programs or instructions are stored, and when the programs or instructions are executed by a processor, the customs data classification method based on knowledge graph representation as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Knowledge graph representation learning method for integrating text semantic features based on attention mechanism
CN110334219A
Feature tensor-based Chinese knowledge graph representation learning method
CN111160564A