A customs data checkpoint prediction method, device and storage medium
By converting customs multi-attribute data into sequence text, and combining constant convolution and attention mechanisms to extract keyword characteristics and relationship characteristics, the problem of difficulty in extracting keyword characteristics in customs data is solved, and the accuracy and efficiency of checkpoint prediction are improved.
Patent Information
- Application Number
- CN202210546521.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-05-18
AI Technical Summary
The prior art is difficult to effectively extract keyword characteristics in customs multi-attribute data and consider the relationship between keywords, resulting in low accuracy in customs prediction of customs data.
A customs data check prediction method is adopted. By combining customs multi-attribute data into sequence text with length T, encoding it into word vectors using the language presentation layer, and combining constant one-dimensional convolution and attention mechanism, the basic information of keywords and the relationship characteristics between keywords are extracted, and finally inputting the fully connected softmax classification layer for check prediction.
The long-range features and keyword attributes in customs data are effectively extracted, which improves the accuracy and efficiency of checkpoint prediction, and solves the problems of poor generalization and slow training speed of traditional models in customs data processing.
Smart Images

Figure CN114971001B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of customs data, and in particular relates to a keyword feature fusion model based on customs data. Background Art
[0002] Customs is an important unit that supervises and manages imported and exported goods, passenger luggage and postal items, inbound and outbound means of transport, collects tariffs and other taxes, and investigates smuggling. Customs tax revenue is an important source of national fiscal revenue and an important tool for the country to implement macroeconomic regulation. Therefore, the supervision of customs tax revenue is also very important.
[0003] With the rapid development of international trade and cross-border e-commerce, the tax supervision business of customs is facing multiple difficulties such as the unprecedented variety of commodities and more complex and diverse trade forms. The requirements for the speed, accuracy and risk control of taxation are also constantly increasing. At present, the customs import and export commodity standard declaration catalog stipulates the declaration elements of different types of commodities, and each element may cover one or more commodity attributes. Therefore, it is necessary to analyze and identify the commodity characteristics according to the declaration elements of the commodity, extract the key information of different commodities through commodity feature recognition, and conduct further risk analysis, so as to effectively regulate different categories of commodities. Different commodities have different characteristics. In addition to the common commodity attributes (price, brand, model, etc.), each commodity also has key attributes in its professional field. Therefore, commodity feature recognition requires not only a comprehensive understanding of the common attributes of commodities, but also professional field knowledge for different commodities. Although the description fields such as commodity attributes contain a lot of important information, they are also very irregular, and it is difficult for machines to widely extract effective keyword information.
[0004] At present, there are the following problems in customs multi-attribute data: some fields are missing or incomplete, the length difference of the same attribute sequence of different data is too large, and the time series information extracted by the same model often has poor generalization. In addition to the time series information of the context, the overall information of the keyword has a great influence on the classification results, among which the keyword features are particularly important, and the relationship between multiple keyword attributes is also very important. However, the current technical model cannot cope with customs multi-attribute data to specifically solve the problem of keyword feature extraction, and the keyword relationship may appear in multiple different attributes. It is difficult to consider the relationship between keywords using traditional classification methods. At the same time, the amount of customs data is huge, and redundant models will increase training time and reduce training speed. Country of origin prediction, customs port prediction, and tax number prediction are all important concerns for customs anomaly detection. In order to improve the accuracy of prediction, it is necessary to explore a fast and efficient model applicable to customs data to solve the current bottlenecks. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides a customs data prediction method, device and storage medium for accurate prediction of customs data.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0007] A customs data customs port prediction method, comprising:
[0008] Step 1: Merge each multi-attribute data of the customs into a sequence text of length T;
[0009] Step 2: Input the sequence text of length T obtained in step 1 into the language representation layer, encode it into a word vector, and obtain a two-dimensional matrix of size (T, k);
[0010] Step 3: Convolve the two-dimensional matrix obtained in step 2 with the constant one-dimensional matrix to obtain the combined basic keyword information;
[0011] Step 4: Input the combined basic keyword features obtained in step 3 into an individual attention module and a relationship attention module respectively to extract the features of important keywords and the relationship features between discontinuous keywords;
[0012] Step 5: Combine the two features obtained in step 4, using the idea of leapfrog connection. Leapfrog connection means that the result of individual attention in the model diagram skips one layer and is directly connected to the next layer. The connection method is to directly add the output of this layer to the output of the next layer.
[0013] Step 6: Input the two-dimensional matrix encoded in step 2 into the bidirectional long short-term neural network to extract the feature vector;
[0014] Step 7: Input the feature vector obtained in step 6 into the attention mechanism to extract the feature vector;
[0015] Step 8: Input the feature vector obtained in step 5 into the fully connected layer, and then go through a max pooling step;
[0016] Step 9: Input the feature vector obtained in step 7 into the fully connected layer;
[0017] Step 10: The feature vectors obtained in step 8 and step 9 are merged and then input into a fully connected softmax classification layer to obtain the conditional probability of the gate prediction.
[0018] The commodity description of the customs contains word features that are useful for classification tasks. Such word features are often irrelevant to the context and are sparsely present in the entire commodity description. Due to its sparsity and the disorder of effective words in the commodity description, traditional text classification models cannot efficiently extract keyword information and long-range relationships between keywords (and keywords in non-same attributes). Based on the characteristics of Chinese text, the prediction method proposed in this invention can extract useful Chinese vocabulary information and the correlation between multiple words from long sequence inputs in the customs attribute recognition task. Using these key features, our prediction method can accurately classify the prediction tasks corresponding to customs data.
[0019] In step 1, each customs e-commerce data contains multiple attributes, and all attributes are merged into a sequence text of length T. After processing, the data set is divided into a training set and a validation set.
[0020] In step 2, word2vec is used as the language representation model, and the sequence text obtained in step 1 is input into word2vec, so that each word of the text is converted into a word vector of length k, and finally the sequence text of length T is converted into a two-dimensional matrix of size (T, k).
[0021] In step 3, three constant convolution kernels are used to process the two-dimensional word vector. Considering that the length of a Chinese phrase generally does not exceed 4 Chinese characters, the convolution kernel sizes are set to (1,2), (1,3), and (1,4) respectively. Through convolution kernel processing, the input information is transformed into feature A of (3T, 32).
[0022] Assuming that the information carried by different Chinese characters in the vocabulary is of equal weight, the constant values of three convolution kernels are given in imitation of the Box blur kernel. Take the (1,3) convolution kernel as an example:
[0023]
[0024] Through this process, the input information is transformed into (T, 32, 3). This process completely saves the input information and evenly extracts all possible keyword vectors contained in the text. The features of step 3 are different from those of traditional TextCNN. The main differences include: (1). TextCNN uses (k, 2), (k, 3), (k, 4) convolution kernel sizes to mix the k-dimensional features of each character in the input information. (2). TextCNN uses learnable convolution to process the input text, destroying the original information of the input text. (3). The method of the present invention does not include a normalization layer and an activation layer. In summary, the method of the present invention integrates keyword attributes while retaining the original text information as much as possible. The feature processing is handled by the subsequent model, and its purpose and construction method are essentially different from TextCNN.
[0025] In step 4, the combined (3T, 32) keyword feature A obtained in step 3 is input into the created feature attention layer. The purpose of this layer is to find and emphasize keyword features or keyword relationship features that are useful for a given task. The present invention is divided into two modules to process keyword feature A:
[0026] (1). The first module is the individual attention module. We use a convolution layer with an output channel of 64 (the convolution kernel size is 1 times 32) and a BN layer and an activation layer to process the input feature A into feature B of (3T, 64). After that, feature B is passed through 4 similar "convolution-BN layer-activation layer" structures in succession. The final output is still (3T, 64). The signal is max pooled in 64 dimensions to become a 3T-dimensional vector. After that, the processed vector is passed through two fully connected layers (including BN layer and activation layer), and the output is still a T-dimensional vector. Max pooling is used to obtain the maximum signal value of the corresponding feature of each word, and the fully connected layer is used to determine the importance of each data in the sequence to a given task. In order to normalize the importance parameter, the activation function of the last layer is represented by the exp function, and the function form is expressed as:
[0027] y=βe -α||x||2
[0028] Among them, α and β are empirical constants. Compared with the Sigmoid function, the above function provides a smoother normalized interval, and the importance parameter will not simply fall to 0 or 1. Finally, the importance parameter (3T dimension) is multiplied by the 3T dimension of the (3T, 64)-dimensional feature before maxpooling to obtain the (3T, 64) output feature C:
[0029] C i =y i *B i
[0030] In the above formula, C i represents the i-th 64-dimensional feature, and C is the feature with local important information determined by the individual attention mechanism.
[0031] (2) The second module is a relational attention module, which is used to extract the features of the relations between words. The (3T, 64)-dimensional feature C contains all the features of 2-4 words, but the final classification target is not necessarily related to the features of a single word, but is likely to be related to the features of multiple words that are far apart. Therefore, the present invention designs a long-range relational attention mechanism to further extract features.
[0032] First, the i-th 64-dimensional vector and the j-th 64-dimensional vector in the feature C with dimension (3T, 64) are concatenated to obtain a 128-dimensional vector. Then, i and j are traversed in the 3T dimension to construct a feature D with dimension (3T, 3T, 128):
[0033] D ij =[C i ,C j ]
[0034] After that, let this feature D pass through four "1x1 convolution, BN, ReLU activation" layers, and the activation function of the last layer is the tanh function, which becomes a (3T, 3T, 1)-dimensional matrix D'. This matrix is used to represent the relationship features between words. For example, the relationship between the i-th and j-th vectors in the 3T dimension is reflected in the (3T, 64) i-th row and j-th column of the matrix. After obtaining the relationship matrix, perform matrix multiplication on the feature C of dimension (3T, 64) to obtain a feature E of size (3T, 64) containing the relationship between words:
[0035] E=D′C
[0036] The purpose of the present invention is similar to that of self-attention, but the process of the present invention is completely different from that of self-attention. In self-attention, the relationship matrix is obtained by multiplying the result of (3T, 64) after two fully connected layers. In this process, the original (3T, 64) undergoes multiple linear and nonlinear transformations, and the correctness of the training is ensured only by end-to-end training. The method of the present invention directly retains the original 64-dimensional features of each dimension, and directly predicts the relationship based on the splicing of two features that may be related. Compared with self-attention, the attention mechanism of the present invention has stronger interpretability and more direct relationship parameters. The present invention uses tanh to indicate whether the relationship between two word variables is positively correlated or negatively correlated to the final classification. Compared with self-attention, the attention mechanism of the present invention is more suitable for the relationship extraction of text vocabulary in this task, and has also achieved better results in experiments. Finally, the present invention connects the output feature C of the individual attention mechanism to the output of the relationship attention layer to emphasize the result of individual attention. The final output result is feature C+E.
[0037] In this process, the present invention uses constant convolution to extract the basic information of keywords, and uses the attention mechanism to learn useful keyword features or relationship features between keywords. Different from TextCNN, the method of the present invention efficiently finds useful keyword features in the original sequence that has not been abstracted. At the same time, the method of the present invention takes into account the long-range keyword relationship, while TextCNN can only connect long-range relationships by continuously deepening the depth. When the sequence is long or the sequence length is not fixed, TextCNN is not easy to build and has low efficiency.
[0038] In step 6, the two-dimensional matrix encoded in step 2 is input into a bidirectional long-term short-term neural network to extract the temporal feature vector. In addition to the keyword feature, the present invention extracts the temporal feature of the original text, which contains different information from the feature obtained in step 5 and is used together for the final classification.
[0039] In step 7, the feature vector obtained in step 6 is input into the traditional attention mechanism (self-attention) module, the weight is initialized, and then three sets of weight matrices K, Q, and V are derived to calculate the attention score of the input object, and then the softmax is calculated, the score is multiplied by V, and the total weighted value is obtained to obtain the output. The specific formula is expressed as:
[0040]
[0041] where d k is an empirical constant. After the output two-dimensional matrix is input into the fully connected layer, it is processed by BatchNormalization and then dropout. The output result is then input into dropout, fully connected layer, BatchNormalization, and ReLU in sequence to obtain the extracted feature vector. The features of this branch contain the temporal information of the original text. The branch represented by CNN plus the new attention mechanism module is responsible for extracting keywords and keyword relationship features that do not have temporal information. The two parts of the features are fused to obtain the final classification features.
[0042] The present invention also provides a device, including a processor and a memory; the memory stores a program or instruction, and the program or instruction is loaded and executed by the processor to implement the above-mentioned customs data checkpoint prediction method.
[0043] The present invention also provides a computer-readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the above-mentioned customs data checkpoint prediction method is implemented.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] 1. The present invention proposes a customs data checkpoint prediction method, which combines a bidirectional long-short term neural network and an attention mechanism to effectively extract long-range features. It combines individual attention and global attention mechanisms to learn the combined features of different keywords respectively, effectively extract keyword attribute features, and provides an effective method for solving the checkpoint prediction problem.
[0046] 2. The present invention uses a method for processing two-dimensional word vectors using three constant convolution kernels, which does not include a normalization layer and an activation layer. This process completely saves the input information and evenly extracts all possible keyword vectors contained in the text. The method of the present invention integrates keyword attributes while retaining the original text information as much as possible. The feature processing is handled by the subsequent model, and its purpose and construction method are essentially different from TextCNN.
[0047] 3. The present invention uses constant convolution to combine Chinese short words. TextCNN also has the ability to combine short words, but TextCNN uses convolution kernels of different sizes to directly extract features from original features, and then performs max pooling on the results of each convolution kernel to remove sequence information, and then integrates them together for feature extraction. The bottleneck of this approach is that the maxpooling part loses a lot of text information, and when extracting text features, the performance of the convolution network is often inferior to that of the recurrent network and transformer. After max pooling, the continued features have lost a lot of information that the original features have. In the present invention, constant convolution is used to retain the original features as much as possible, and no linear or nonlinear transformation is performed on them. More importantly, after constant convolution, all short word information is retained by splicing in the sequence dimension, which provides good original information for the subsequent attention module to extract features. Sequence dimension splicing also enables the attention mechanism module to extract the relationship between short words of different lengths. For example, in a sentence, the three-character word at the beginning and the two-character word at the end jointly determine the classification result. The constant convolution module in the present invention retains all such possible patterns and inputs them into the feature extraction module, solving the problem of large-scale information loss in TextCNN.
[0048] 4. The present invention provides an attention mechanism based on text sequences. The attention mechanism is divided into two parts. The first part is to directly search for sequence segments that are more important to a given task in the sequence, and the second part is to find the relationship between the more important sequence segments, that is, whether the combination of certain sequence segments is more conducive to the completion of the given task. Traditional self attention can also achieve this goal, but self attention sends the input features into three fully connected layers to obtain three intermediate features Q, K, and V, and then generates an attention matrix from Q and K. In this process, the representation of Q and K depends entirely on training by setting a reasonable loss, which also leads to slow convergence speed. It completely relies on training to regulate Q and K, which also makes it not very interpretable. The attention matrix method in the present invention is completely different from self attention. It directly uses the values of each two dimensions of the input features to directly form the initial attention tensor. This tensor does not need to be trained to directly retain the information of the original features in pairs, and then processes it into an attention matrix through neural network processing. Its initial features have good interpretability. Since the initial attention tensor to the final attention matrix only passes through several layers of convolution layers in sequence, there are no additional branches, and the convergence of parameters during training is also relatively fast.
[0049] 5. The purpose of the attention module for sequence-segment relationship is similar to that of self-attention, but the present invention also proposes an attention mechanism for each individual sequence, which jumps the output features of the individual attention mechanism to the final features, so that the model can output features that take into account both the sequence monomer elements and the sequence segment relationship, which is more in line with the situation in which local nouns for various items in customs data occupy the main classification elements. The individual attention mechanism can be seen as assigning a weight to each 64-dimensional vector of the (3T, 64) feature. A large weight has a large impact on the classification, and a small weight has a small impact on the classification (for example, (3T, 64) represents the word vector of "cream biscuits". "Cream" and "biscuits" are both represented by a 64-dimensional vector. Now we need to classify the type of items. Only the word "biscuit" has an impact on the final classification, while "cream" has no impact or little impact. After this operation, the vector represented by "biscuits" will be multiplied by a larger weight, and vice versa). The process of generating weights is from B to B1 and then to a vector of size 3T, but the B feature does not contain the weight information. A new feature extraction part needs to be added to find the weight information, so a 4-layer convolution is added to extract the weights in the original features. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and other advantages of the present invention will become more clear.
[0051] Figure 1It is a model framework diagram of the present invention;
[0052] Table 1 shows the comparison of classification indicators of the keyword feature fusion model of customs data on the validation set with other methods. By comparing with the average accuracy, compared with BiLSTM, TextCNN and Bert methods, this method currently has the highest average accuracy on customs data. DETAILED DESCRIPTION
[0053] The present invention will be further described below with reference to the embodiments.
[0054] Example 1
[0055] Referring to the method flow of the present invention, a keyword feature fusion model and device based on customs data of the present invention specifically include the following steps:
[0056] Step 1: Merge each multi-attribute data of the customs into a sequence text of length T. For example, for the source data "Zhejiang XX Co., Ltd. | Hangzhou XX Co., Ltd. | 3301968FU0 | I | 2019 / 3 / 5 5:54 | 4612 | 4612 | 2019 / 3 / 5 0:00 | 8f55f5f0-142b-472e-b989-6804409e00f0 | 5.66041E+11 | Revlon Revlon lipstick female student genuine lipstick lasting moisturizing non-fading dog bean paste color 225 | 3304100091 | Revlon Licai lipstick 445# | 4.2g / piece; zinc citrate, mineral wax, ozokerite; polishing and brightening; Revlon; other overseas", it is processed into "Zhejiang XX Co., Ltd. Hangzhou XX Co., Ltd. 3301968FU0 I 2019 / 3 / 5 5:54 4612 4612 2019 / 3 / 5 0:00 8f55f5f0-142b-472e-b989-6804409e00f0 5.66041E+11revlon Revlon lipstick female student genuine lipstick long-lasting moisturizing non-fading dog bean paste color 225 3304100091 Revlon Li Cai lipstick 445#4.2g / piece; zinc citrate, mineral wax, ozokerite; polish and brighten; Revlon; overseas others";
[0057] Step 2: Input the sequence text of length T obtained in step 1 into the language representation layer, encode it into a word vector, and obtain a two-dimensional matrix of size (T, k);
[0058] Step 3: Convolve the two-dimensional matrix obtained in step 2 with the constant one-dimensional matrix to obtain the combined basic keyword information;
[0059] Step 4: Input the combined basic keyword features obtained in step 3 into an individual attention module and a relationship attention module respectively to extract the features of important keywords and the relationship features between discontinuous keywords;
[0060] Step 5: Combine the two features obtained in step 4;
[0061] Step 6: Input the two-dimensional matrix encoded in step 2 into the bidirectional long short-term neural network to extract the feature vector;
[0062] Step 7: Input the feature vector obtained in step 6 into the attention mechanism to extract the feature vector;
[0063] Step 8: Input the feature vector obtained in step 5 into the fully connected layer, and then go through a max pooling step;
[0064] Step 9: Input the feature vector obtained in step 7 into the fully connected layer;
[0065] Step 10: The feature vectors obtained in step 8 and step 9 are merged and then input into a fully connected softmax classification layer to obtain the conditional probability of the gate prediction.
[0066] In step 1, each customs e-commerce data contains multiple attributes, and all attributes are merged into a sequence text of length T. After processing, the data set is divided into a training set and a validation set.
[0067] In step 2, word2vec is used as a language representation model, and the sequence text obtained in step 1 is input into word2vec, so that each word of the text is converted into a word vector of length k, and finally the sequence text of length T is converted into a two-dimensional matrix of size (T, k). In one embodiment, k is set to 32.
[0068] In step 3, three constant convolution kernels are used to process the two-dimensional word vector. Considering that the length of a Chinese phrase generally does not exceed 4 Chinese characters, the convolution kernel sizes are set to (1,2), (1,3), and (1,4). Assuming that the information carried by different Chinese characters in the vocabulary is of equal weight, the constant values of the three convolution kernels are given in imitation of the Box blur kernel. Take the (1,3) convolution kernel as an example:
[0069]
[0070] Through this process, the input information is transformed into (T, 32, 3). This process completely saves the input information and evenly extracts all possible keyword vectors contained in the present invention. The features of step 3 are different from those of traditional TextCNN. The main differences include: (1). TextCNN uses (k, 2), (k, 3), (k, 4) convolution kernel sizes to mix the k-dimensional features of each character in the input information. (2). TextCNN uses learnable convolution to process the input text, destroying the original information of the input text. (3). The method of the present invention does not include a normalization layer and an activation layer. In summary, the method of the present invention integrates keyword attributes while retaining the original text information as much as possible. The feature extraction is then processed by the subsequent model, and its purpose and construction method are essentially different from TextCNN.
[0071] In step 4, the combined (3T, 32) keyword feature A obtained in step 3 is input into the feature attention layer. The purpose of this layer is to find and emphasize keyword features or keyword relationship features that are useful for a given task. It is divided into two modules to process the basic keyword feature A:
[0072] (1). The first module is an individual attention module. The present invention uses a convolution layer with an output channel of 64 (the convolution kernel size is 1 times 32) and a BN layer and an activation layer to process the input feature A into a feature B of (3T, 64). Thereafter, feature B is passed through four similar "convolution-BN layer-activation layer" structures in succession. The final output is still (3T, 64). Feature B is subjected to a one-step max pooling in 64 dimensions to become a 3T-dimensional vector. The processed vector is passed through two fully connected layers (including a BN layer and an activation layer), and the output is still a T-dimensional vector. The maximum signal value of the corresponding feature of each word is obtained through max pooling, and the importance of each data in the sequence to a given task is determined through a fully connected layer. In order to normalize the importance parameter, the activation function of the last layer is represented by an exp function, and the function form is expressed as:
[0073] y=βe -α||x||2
[0074] Among them, α and β are empirical constants. Compared with the Sigmoid function, the above function provides a smoother normalized interval. The importance parameter will not simply fall to 0 or 1. Finally, the importance parameter (3T dimension) is dot-multiplied with the 3T dimension of the (3T, 64)-dimensional feature before maxpooling to obtain the (3T, 64) output feature C.
[0075] C i =y i *B i
[0076] In the above formula, C i represents the i-th 64-dimensional feature, and C is the feature with local important information determined by the individual attention mechanism.
[0077] (2) The second module is a relational attention module, which is used to extract features of relations between words. The feature C of (3T, 64) dimension contains all features of 2-4 words, but the final classification target is not necessarily related to only individual word features, and is likely to be related to multiple word features that are far apart. Therefore, the present invention designs an attention mechanism for long-range relations to further extract features. First, the i-th 64-dimensional vector and the j-th 64-dimensional vector in the feature C of dimension (3T, 64) are concatenated to obtain a 128-dimensional vector, and i and j are traversed in the 3T dimension to construct a feature D of (3T, 3T, 128):
[0078] D ij =[C i ,C j ]
[0079] After that, let this feature D pass through four "1x1 convolution, BN, ReLU activation" layers, and the activation function of the last layer is the tanh function, which becomes a (3T, 3T, 1)-dimensional matrix D'. This matrix is used to represent the relationship features between words. For example, the relationship between the i-th and j-th vectors on the 3T dimension is reflected in the i-th row and j-th column of the matrix (3T, 64). After obtaining the relationship matrix, perform matrix multiplication on the feature C feature of dimension (3T, 64) to obtain a feature E of size (3T, 64) containing the relationship between words:
[0080] E=D′C
[0081] It is emphasized here that although the purpose of the present invention is similar to self-attention, the process is completely different from self-attention. In self-attention, the relationship matrix consists of two sets of linear plus nonlinear transformations of (3T, 64). In this process, the original (3T, 64) is transformed, and the correctness of the training is ensured only by end-to-end training. The method of the present invention directly retains the original 64-dimensional features of each dimension, and directly predicts the relationship based on the splicing of two features that may be related. Compared with self-attention, the attention mechanism of the present invention has stronger interpretability and more direct relationship parameters. The present invention uses tanh to indicate whether the relationship between two word variables is positively correlated or negatively correlated to the final classification. Compared with self-attention, the attention mechanism of the present invention is more suitable for the relationship extraction of text vocabulary in this task, and has achieved better results in experiments. Finally, the present invention connects the output C jump layer of the individual attention mechanism to the output of the relationship attention to emphasize the result of individual attention. The final output result is C+E.
[0082] In this process, the present invention uses constant convolution to extract the basic information of keywords, and uses the attention mechanism to learn useful keyword features or relationship features between keywords. Different from TextCNN, the method of the present invention efficiently finds useful keyword features in the original sequence that has not been abstracted. At the same time, the method of the present invention takes into account the long-range keyword relationship, while TextCNN can only connect long-range relationships by continuously deepening the depth. When the sequence is long or the sequence length is not fixed, TextCNN is not easy to build and has low efficiency.
[0083] In step 6, the two-dimensional matrix encoded in step 2 is input into the bidirectional long-term short-term neural network to extract the time series feature vector. In addition to the keyword features, the present invention extracts the time series features of the original text, which contain different information from the features obtained in step 5 and are used together for the final classification task;
[0084] In step 7, the feature vector obtained in step 6 is input into the traditional attention mechanism (self-attention) module, the weight is initialized, and then three sets of weight matrices K, Q, and V are derived to calculate the attention score of the input object, and then the softmax is calculated, the score is multiplied by V, and the total weighted value is obtained to obtain the output. The specific formula is expressed as:
[0085]
[0086] where d kis an empirical constant, which is set to 4.0 in the experiment. After the output two-dimensional matrix is input into the fully connected layer, it is subjected to Batch Normalization and then dropout processing. The output result is then input into the dropout, fully connected layer, Batch Normalization, and ReLU in sequence to obtain the extracted feature vector. The features of this branch contain the temporal information of the original text. The branch represented by the CNN plus the new attention mechanism module is responsible for extracting keywords and keyword relationship features that do not have temporal information. The two parts of the features are fused to obtain the final classification features.
[0087] Step 8: Input the feature vector obtained in step 5 into the fully connected layer;
[0088] Step 9: Input the feature vector obtained in step 7 into the fully connected layer;
[0089] Step 10: The feature vectors obtained in step 8 and step 9 are merged and then input into a fully connected softmax classification layer to obtain the conditional probability of the gate prediction.
[0090] Step 8: Input the feature vector obtained in step 5 into the fully connected layer, the input feature size is (3T, 64), connect each 64-dimensional vector to the fully connected layer, map it to 128 dimensions, output (3T, 128) features, perform max pooling on 128 dimensions, and output a 128-dimensional feature vector. The formula of the fully connected layer is:
[0091] y=f(wx+b)
[0092] w is the weight matrix, b is the bias vector, x is the 64-dimensional vector, and f is the ReLU function. This operation must be performed on each 64-dimensional vector, and the process is equivalent to a convolution operation with a one-dimensional convolution kernel of 1.
[0093] Step 9: Input the feature vector obtained in step 7 into the fully connected layer. The fully connected formula is as above. The branch in step 7 finally outputs a 128-dimensional feature vector by self attention, and the fully connected layer continues to map it to 128 dimensions.
[0094] Step 10: Fuse the feature vectors obtained in step 8 and step 9, and then input them into a fully connected softmax classification layer to obtain the conditional probability of the gate prediction. After fusion, the feature size is 256 dimensions, and a linear transformation is used to map it to a 96-dimensional vector, 96 being the number of gate categories, and then the probability is calculated through softmax. The softmax formula is as follows:
[0095]
[0096] where y irepresents the i-th dimension in the 96-dimensional vector, s i Represents the value of the i-th dimension after softmax.
[0097] Training hyperparameter settings, all methods use a set of training parameters: the data uses the import and export data of the customs e-commerce platform, and the training set and test set are divided according to the ratio of 7:3. The optimizer during training is Adam, where the weight_decay parameter is set to 0.001, the batch size is set to 16, and the number of training rounds is 10. In the comparative method of the embodiment, the bidirectional long short-term memory model uses 256 hidden neurons. In addition to the convolutional layer (first layer) with multiple convolution kernels, the text convolutional neural network is subsequently connected to 5 convolutional layers to extract features.
[0098] Table 1 shows the test results in this embodiment. In addition to the core checkpoint, the present invention is also applicable to another declaration element "country of origin":
[0099] Table 1
[0100] method Gate prediction accuracy Country of Origin Forecast Method of the present invention 0.85 0.62 BiLSTM 0.77 0.46 TextCNN 0.78 0.48 BERT 0.80 0.58
[0101] The experimental results are evaluated by average accuracy. By comparing the average accuracy with the BiLSTM, TextCNN and BERT methods, this method currently has the highest average accuracy on customs data. This can also be proved for the prediction of the country of origin. In the comparative experiment, all methods are run on the same equipment and framework, and the word segmentation and embedding layer methods are consistent.
[0102] An embodiment of the present invention also provides a device, such as a computer, a mobile phone or other electronic device, comprising a processor and a memory; the memory stores a program or instruction, and the program or instruction is loaded and executed by the processor to implement the above-mentioned customs data customs data prediction method.
[0103] The present invention also provides a computer-readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the above-mentioned customs data checkpoint prediction method is implemented.
Claims
1. A customs data checkpoint prediction method, characterized in that: The steps include: Step 1: Merge each multi-attribute data of the customs into a sequence text of length T; Step 2: Input the sequence text of length T obtained in step 1 into the language representation layer, encode it into a word vector, and obtain a two-dimensional matrix of size (T, k); Step 3: Convolve the two-dimensional matrix obtained in step 2 with the constant one-dimensional matrix to obtain the combined basic keyword information; Step 4: Input the combined basic keyword features obtained in step 3 into an individual attention module and a relationship attention module in turn to extract the features of important keywords and the relationship features between discontinuous keywords; Step 5: The features of the important keywords obtained in step 4 are combined with the relationship features between the discontinuous keywords based on the leap-level connection method; Step 6: Input the two-dimensional matrix encoded in step 2 into the bidirectional long short-term neural network to extract the feature vector; Step 7: Input the feature vector obtained in step 6 into the self-attention mechanism to extract the feature vector; Step 8: Input the feature vector obtained in step 5 into the fully connected layer, and then go through a max pooling step; Step 9: Input the feature vector obtained in step 7 into the fully connected layer; Step 10: The feature vectors obtained in step 8 and step 9 are merged and then input into a fully connected softmax classification layer to obtain the conditional probability of the gate prediction; In step 3, three constant convolution kernels are used to process the two-dimensional word vector, and the convolution kernel sizes are (1,2), (1,3), and (1,4) respectively; through convolution kernel processing, the input information is transformed into the keyword feature A of (3T, 32); The combined (3T, 32) keyword feature A obtained in step 3 is input into the created feature attention layer to find and emphasize keyword features or relationship features between keywords that are useful for a given task; The feature attention layer contains two modules for processing keyword feature A, one module is the body attention module, and the other module is the relationship attention module: The individual attention module uses a convolutional layer with an output channel of 64 and a BN layer and activation layer to process the input feature A into feature B of (3T, 64); Pass feature B through a four-layer convolutional network. Each layer of the convolutional network contains "convolution-BN layer-activation". The number of input and output channels of the four-layer convolution is 64, and the final output size is (3T, 64) feature B1; Perform one-step max pooling on feature B1 in 64 dimensions to convert it into a 3T-dimensional vector, and pass the 3T-dimensional vector through two fully connected layers including a BN layer and an activation layer; The activation function of the last layer is represented by the exp function, and the function form is expressed as: Among them, α and β are empirical constants, y is the output of the activation function, and x is the input of the activation function; Multiply the 3T-dimensional vector by the 3T-dimensional feature of the (3T, 64)-dimensional feature before max pooling to obtain the (3T, 64) output feature C = (C1, C2, ..., C T ), where the i-th 64-dimensional vector C i It is expressed as: C i =y i *B i B i Indicates that feature B is divided into 3T 64-dimensional vectors, and the i-th vector among them; The relational attention module concatenates the i-th 64-dimensional vector and the j-th 64-dimensional vector in the (3T, 64)-dimensional feature C to obtain a 128-dimensional vector, traverses i and j in the 3T dimension, and constructs a (3T, 3T, 128) feature D: D ij =[C i ,C j ] Let this feature D pass through four "1x1 convolution, BN, ReLU activation" layers, with the activation function of the last layer being the tanh function, and become a (3T, 3T, 1)-dimensional matrix D'. The matrix D' is used to represent the relationship features between words. Perform matrix multiplication with the feature C of dimension (3T, 64) using matrix D′ to obtain feature E of size (3T, 64) containing the relationship between words: E=D′C The output feature C of the individual attention mechanism is connected to the output feature E of the relational attention mechanism, and the output result is C+E.
2. The gateway prediction method according to claim 1, characterized in that: In step 1, each customs e-commerce data contains multiple attributes, and all attributes are merged into a sequence text of length T; after processing, the data set is divided into a training set and a validation set.
3. The gateway prediction method according to claim 1, characterized in that: In step 2, word2vec is used as the language representation model, and the sequence text obtained in step 1 is input into word2vec, so that each word of the text is converted into a word vector of length k, and finally the sequence text of length T is converted into a two-dimensional matrix of size (T, k).
4. The gateway prediction method according to claim 1, characterized in that: In step 6, the two-dimensional matrix encoded in step 2 is input into a bidirectional long short-term neural network to extract the time series feature vector.
5. The gateway prediction method according to claim 1, characterized in that: In step 7, the time series feature vector obtained in step 6 is input into the traditional attention mechanism module, the weight is initialized, and then three sets of weight matrices K, Q, and V are derived to calculate the attention score of the input object, and then the softmax is calculated, the score is multiplied by V, and the total weighted value is obtained to obtain the output. The specific formula is expressed as: Among them, d k is an empirical constant, K T is the transpose of K; After the output two-dimensional matrix is input into the fully connected layer, it goes through Batch Normalization and then performs dropout processing; The output result is then input into the dropout, fully connected layer, Batch Normalization, and ReLU in sequence to obtain the extracted feature vector.
6. A device, characterized in that: It comprises a processor and a memory; the memory stores a program or instruction, and the program or instruction is loaded and executed by the processor to implement the customs data checkpoint prediction method as described in any one of claims 1 to 5.
7. A computer-readable storage medium storing a program or instruction, wherein the program or instruction, when executed by a processor, implements the customs data checkpoint prediction method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Text emotion analysis method based on Chinese data set
CN108763216A
Chinese word segmentation method, electronic device and readable storage medium
CN110287961A