A Customs Import and Export Commodity Classification Method Based on a Multi-Model Text Intelligent Coding Algorithm

Through the multi-module text intelligent coding algorithm, the customs import and export commodity declaration text is standardized, which solves the classification difficulties caused by irregular product attribute description, and realizes efficient product classification and abnormal detection, which improves customs work efficiency.

CN113947061BActive Publication Date: 2025-07-18DALIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111235112.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-07-18
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively classify the customs import and export commodity declaration text, especially when the declaration elements are discrete and the text description of commodity attributes is not standardized, resulting in inefficient classification of customs commodity and poor generalization.

Method used

Multi-module text intelligent coding algorithm is adopted to realize the standardized processing of customs import and export commodity texts through data cleaning, splitting, sorting, keyword searching, independent word merging, synonym replacement and random code mapping, and synonym replacement is used for synonym replacement by using the BERT model to generate random codes for classification abnormality detection.

Benefits of technology

It improves the efficiency of customs product inspection, reduces the storage scale of massive data, and improves the accuracy of product classification and abnormal detection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113947061B_ABST
    Figure CN113947061B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for classifying customs import and export commodities based on a multi-module text intelligent coding algorithm. This multi-module text intelligent coding algorithm uses a customs knowledge base and standardizes the customs import and export commodity declaration text through multiple groups of intelligent processing modules to reduce the information entropy of the commodity declaration text. Subsequently, the text is converted into random codes for storage using coding logic, which not only reduces the information storage space but also enables the use of the "same code - different classification" logic to find commodities with classification anomalies, and its inspection results have a very high confidence level. Using the multi-module text-random code conversion logic, the classification of customs import and export commodity text is achieved under the premise that the declaration element content is discrete and the literal description of commodity attributes is not standardized. While improving the efficiency and effect of customs commodity inspection, the storage scale of massive data is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a method for classifying customs import and export commodities based on a multi-module text intelligent coding algorithm. Background Art

[0002] The main objects of customs supervision are import and export commodities. With the economic globalization, the throughput of customs import and export commodities has been increasing continuously, and levying taxes on import and export commodities has become a heavy task for the customs department. The tax rate of commodities depends on the classification of commodities. At present, Chinese customs mainly uses manual classification for import and export commodities. Customs officers classify commodities according to the declared text information of commodities through the customs system, and then calculate and levy taxes. This relatively traditional method is time-consuming and laborious, and can only cover a very small part of the huge amount of import and export commodities. Natural language processing is an artificial intelligence technology that specifically studies text representation. It can model the text information of commodities and construct high-dimensional spatial feature vectors of the text. These feature vectors composed of numbers carry information such as the semantics and word order of the text. Therefore, the computer can use these feature vectors to perform text task calculations and provide computing power support for the task of classifying customs import and export commodities.

[0003] The existing methods for assisting in predicting the classification of customs import and export commodities are basically based on database search. In recent years, there have also been cases of directly classifying import and export commodities using machine learning classification algorithms. However, due to the high professionalism of customs business and the non-standardization of declaration form data in the customs import and export declaration text compared with ordinary Chinese texts, simply transplanting and using the algorithms in simple natural language processing technology cannot achieve good classification effects. At the same time, using traditional rule bases to formulate rules for classifying customs import and export commodities can build the underlying logic according to business logic, but the generalization ability is weak, and it is extremely difficult to formulate rules under a large amount of data. Summary of the Invention

[0004] The purpose of this application is to provide a method for classifying customs import and export commodities based on a multi-module text intelligent coding algorithm. This method realizes the classification of customs import and export commodity texts under the premise of discrete declaration element content and non-standard literal descriptions of commodity attributes, and improves the effect of abnormal inspection of customs commodity classification.

[0005] The main object of customs commodity inspection is the declaration text of the commodity, and the judgment target is whether the commodity number in the commodity declaration text is correct. The commodity number is a 10-digit number, representing the commodity category of the commodity under the customs system. The declaration text is a text set describing the various attributes of the commodity, and the collection of attribute names is called the "Catalogue of Declaration Elements for Import and Export Commodities of the Customs". This "catalogue of elements" corresponds one-to-one with the commodity declaration text (element content) filled in by the merchant. The first 4 digits of the commodity number can be used to locate the "catalogue of elements" for which the specific content of the commodity needs to be filled in.

[0006] To achieve the above object, the technical solution of this application is: a method for classifying import and export commodities of the customs based on a multi-module text intelligent coding algorithm, specifically including:

[0007] Step 1: Clean the data of the import and export commodity declaration text, and locate the "Catalogue of Declaration Elements for Import and Export Commodities of the Customs" corresponding to the import and export commodities according to the first 4 digits of the commodity number;

[0008] Step 2: According to the "Catalogue of Declaration Elements for Import and Export Commodities of the Customs", split the import and export commodity declaration text to form element content, which corresponds one-to-one with the "catalogue of elements", and sort it;

[0009] Step 3: For the element content, perform modular data processing through keyword search, independent word merging, and synonym replacement to obtain element text;

[0010] Step 4: Obtain a random code generated by letters and numbers, establish a one-to-one mapping relationship between the element text and the random code, and convert the whole text into coded structure information;

[0011] Step 5: For the coded structure information, find out the import and export commodity declaration texts with different commodity numbers by merging the commodity declaration texts with the same code, and consider that there is a risk of abnormal commodity classification.

[0012] Further, in Step 1, a regular expression is used to clean the data of the import and export commodity declaration text.

[0013] Further, the specific implementation method of Step 2 is:

[0014] Step 21. Split the import and export commodity declaration text according to the "Catalogue of Declaration Elements for Import and Export Commodities of the Customs", and then establish a one-to-one correspondence;

[0015] Step 22. Assign a 4-digit code according to the "chapter" and the "appearance order" in a single chapter. The "chapter" code ranges from "01" to "98", and the appearance order code ranges from the numerical representation of "01" to the mixed representation of "a0", and finally to the alphabetical representation of "zz". An example of a "declaration element" code is "05b7";

[0016] Step 23. According to the rules that the order of elements is different for different chapters, with smaller sub-chapters sorted first and larger chapters sorted later, and elements that appear first in the same chapter are sorted in front, sort the "declaration elements" and "element contents" in the same order.

[0017] Furthermore, the specific implementation method of step 3 is as follows:

[0018] Step 31. Send the element content to the keyword search module for keyword replacement: when there is a keyword in the current element, the entire element content will be replaced with the keyword;

[0019] Step 32. Send the element content to the independent word search module for independent word merging: after the current element undergoes word segmentation, if there are independent words in the element, the independent words that may be split into several parts will be merged;

[0020] Step 33. Send the element content to the synonym replacement module: this module will merge all the "element contents" under a single "declaration element", represent the "element contents" as word vectors through the BERT pre-trained language model, and then calculate the cosine value between the word vectors. The larger this value is, the closer the semantics of the two words are; set words with a cosine value greater than the threshold as synonyms, and then perform synonym replacement.

[0021] Furthermore, the specific implementation method of step 4 is as follows:

[0022] Step 41. Set the number of digits of the random code identifier;

[0023] Step 42. Randomly generate a random code with a fixed number of digits from the library of numbers + capital letters + lowercase letters;

[0024] Step 43. Determine whether there is a corresponding random code for the element text in the database. If there is, convert this segment of element text into a random code; if not, select a random code to establish a mapping relationship with this segment of text and store it in the database, and this segment of text will also be replaced with the corresponding random code.

[0025] Furthermore, in step 5, integrate all the encoded element texts, merge elements with the same code one by one. When the random codes of the entire text are the same, judge whether their product numbers are consistent. If there is an inconsistent situation, it is determined that there is an incorrect declaration for the text other than the correctly declared product number.

[0026] Due to the adoption of the above technical solutions, the present invention can achieve the following technical effects: By using the multi-module text-random code conversion logic, the present invention realizes the classification of customs import and export commodity texts on the premise that the declared element contents are discrete and the literal descriptions of commodity attributes are not standardized. While improving the efficiency and effect of customs commodity inspection, the storage scale of massive data is reduced. Brief Description of the Drawings

[0027] Figure 1 It is a schematic flowchart of a method for classifying customs import and export commodities. Detailed Embodiments

[0028] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments: Taking this as an example, the present application will be further described and explained.

[0029] Embodiment 1

[0030] In the process of classifying customs import and export commodities, the role of the knowledge base in the customs professional field in the text processing process should be utilized well. Data processing based on the rule engine can maintain the focus of data to the greatest extent and reduce the noise and semantic pollution generated in the data processing process. Based on the characteristics of customs texts and the problems in the task of classifying customs import and export commodities, see Figure 1 , the present application provides a method for classifying customs import and export commodities: First, according to the first four digits of the commodity number, find the corresponding "Customs Declaration Element Catalog" for the declared commodity. The customs import and export commodity declaration text is split and sorted according to the element catalog. Then, through the knowledge acquisition module: keyword search, independent word merging, and synonym replacement for modular data processing. Among them, for the synonym replacement module, the BERT model is used to judge the cosine angle between words. Then, through the random code generated by letters and numbers, a mapping relationship is established between the element text and the random code, and the whole text is converted into a coding structure. Finally, through the "same code" merging operation, the import and export commodities with "different classifications" are found. The present application uses the multi-module text-random code conversion logic to effectively solve the problem of inaccurate classification caused by the discrete content of declared elements and the non-standard literal descriptions of commodity attributes in the problem of classifying customs import and export commodity texts, and improves the efficiency of customs commodity inspection.

[0031] The following will describe the present invention in detail in conjunction with the embodiments and drawings, so that those of ordinary skill in the art can implement it with reference to this description.

[0032] This embodiment uses Pycharm as the development platform and Python as the development language. It is carried out on 2,891,525 sentences of customs real data. The following is the specific process:

[0033] Step 1: Clean the data of the import and export commodity declaration text, and then locate the "Catalogue of Declaration Elements for Import and Export Commodities of the Customs" corresponding to the declared commodity according to the first four digits of the commodity code.

[0034] Step 2: Through the "Catalogue of Declaration Elements for Import and Export Commodities of the Customs" obtained in Step 1, split the import and export commodity declaration text to form element contents, which correspond one by one to the "Element Catalogue", and sort them; specifically:

[0035] Step 21. Split the import and export commodity declaration text according to the "Catalogue of Declaration Elements for Import and Export Commodities of the Customs" obtained in Step 1, and then establish a one-to-one correspondence; for example, data:

[0036] Data A: "9001100001|Non-dispersion shifted single-mode optical fiber|0|3|Single mode|For optical signal transmission, used for producing optical cables|No brand|G652D" After splitting, find its corresponding declaration elements:

[0037] {

[0038] "0": "Brand type"

[0039] "3": "Preferential treatment for exports"

[0040] "Single mode": "Structure"

[0041] "For optical signal transmission, used for producing optical cables": "Use

[0042] "No brand": "Brand

[0043] "G652D": "Model

[0044] }

[0045] Step 22. Assign a 4-digit code according to the "appearance order" of the elements in the "chapter" and a single chapter. The "chapter" code ranges from "01" to "98", and the appearance order code ranges from "01" represented by numbers to "a0" represented by a mixture of numbers and letters, and finally to "zz" represented by letters. For example, data:

[0046] {

[0047] "0": "Brand type": [00N3]

[0048] "3": "Preferential treatment for exports":

[0049]

[0049] "Single mode": "Structure": [00H1]

[0050] "For optical signal transmission, used for producing optical cables": "Use: [00T1]

[0051] "No brand": "Brand:

[0001]

[0052] "G652D": "Model:

[0002]

[0053] }

[0054] Step 23. Specify the element sorting rule. For the "declaration elements" and "element content", sort them in the same order according to the rule that for different chapters, the sub-chapters are sorted first; for the same chapter, the elements that appear first are sorted in the front. For example, for the data:

[0055] {

[0056] "No brand": "Brand: [0001,1]

[0057] "G652D": "Model: [0002,2]

[0058] "3": "Export preferential treatment situation": [0049,49]

[0059] "Single mode": "Structure": [00H1,76]

[0060] "0": "Brand type": [00N3,94]

[0061] "For optical signal transmission, used for producing optical cables": "Usage: [00T1,131]

[0062] }

[0063] Step 3: For the element content obtained in Step 2, through the knowledge acquisition module: keyword search, independent word merging, and synonym replacement, perform modular data processing to obtain the element text.

[0064] Step 31. Send the element content into the keyword search module for keyword replacement. When there is a keyword for the element, the entire element content will be replaced by the keyword. For example, for the data: "For optical signal transmission, used for producing optical cables", after passing through the keyword module, it becomes "Optical signal, optical cable".

[0065] Step 32. Send the element content into the independent word search module for independent word merging. When the element is split through word segmentation operation, if there are independent words for the element, the independent words that may be split into several parts will be merged.

[0066] Step 33. Send the element content to the synonym replacement module. This module merges all the "element content" under a single "declaration element", represents the "element content" as word vectors through the BERT pre-trained language model. Then calculate the cosine value between the word vectors. The larger this value, the closer the semantics of the two words. Set words with a cosine value greater than the threshold (0.87) as synonyms and perform synonym replacement. For example, for the data: the cosine value of the word vectors of "single mode" and "monomode" is 0.93, so replace "single mode" with "monomode".

[0067] Step 4: Obtain a random code generated from letters and numbers, establish a one-to-one mapping relationship between the element text obtained in Step 3 and the random code, and convert the entire text into a coding structure.

[0068] Step 41. Set the number of digits for the random code identifier;

[0069] Step 42. Randomly generate a random code with a fixed number of digits from the library of numbers + uppercase letters + lowercase letters;

[0070] Step 43. Determine whether there is a corresponding random code for the element text in the database. If there is, convert this section of text into the random code. If not, select a random code to establish a mapping relationship with this section of text and store it in the database, and this section of text will also be replaced by the corresponding random code. For example, for the data:

[0071] Text A: "9001100001|non-dispersion shifted single-mode optical fiber|0|3|single mode|optical signal, optical cable|no brand|G652D"

[0072] Code A: "9001100001,0000TfVze,0001u8Ue90,0002n9p09I,0049pm46pa,00H10Ybv5P,00N3mrHw9A,00T1XUkowA".

[0073] Step 5: For the coding information obtained in Step 4, find the imported and exported goods with "different classifications" through the "same code" merging operation.

[0074] Specifically, integrate all the coded texts, perform the same code merging element by element. When the random codes of the entire text are the same, judge whether their commodity numbers are consistent. If there is an inconsistent situation, it is considered that all the texts other than the correctly declared commodity number declare incorrect information.

[0075] Examples of classification differences:

[0076] The commodity of "dilution refrigerator". "XX Company" imported 3 batches of "dilution refrigerators" which are classified under the tariff number 8419899022 with a customs duty rate of 0. After reviewing the commodity description, they should be classified under 8418699090 with a customs duty rate of 9%. The supplementary tax for these three records is estimated to be 1.789 million yuan.

[0077] As mentioned above, it is only the preferred specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the technical field, within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent replacements or changes, and should be covered by the protection scope of the present invention.

Claims

1. A customs import and export commodity classification method based on a multi-module text intelligent coding algorithm, characterized in that, Specifically, it includes: Step 1: Clean the data of the import and export commodity declaration text, and locate the "Catalogue of Declaration Elements for Import and Export Commodities of the Customs" corresponding to the import and export commodities according to the first 4 digits of the commodity number; Step 2: Split the import and export commodity declaration text into element contents according to the "Catalogue of Declaration Elements for Import and Export Commodities of the Customs", and the element contents correspond one by one to the "Element Catalogue", and sort them; The specific implementation method is: Step 21. Split the import and export commodity declaration text according to the "Catalogue of Declaration Elements for Import and Export Commodities of the Customs", and then establish a one-to-one correspondence; Step 22. Assign a 4-digit code according to the "appearance order" of the elements in the "chapter" and a single chapter. The "chapter" code ranges from "01" to "98", and the appearance order code ranges from "01" represented by numbers to "a0" represented by a mixture of numbers and letters, and finally to "zz" represented by letters; Step 23. Sort the "declaration elements" and "element contents" in the same order according to the rules that elements in different sub-chapters of the element order are sorted first, elements in larger chapters are sorted later, and elements that appear first in the same chapter are sorted first; Step 3: Perform modular data processing on the element content through keyword search, independent word merging, and synonym replacement to obtain the element text; Step 4: Obtain a random code generated by letters and numbers, establish a one-to-one mapping relationship between the element text and the random code, and convert the whole text into coded structure information; Step 5: For the coded structure information, find out the import and export commodity declaration texts with different commodity numbers by merging the commodity declaration texts with the same code, and consider that there is a risk of abnormal commodity classification.

2. The customs import and export commodity classification method based on the multi-module text intelligent coding algorithm according to claim 1, wherein In Step 1, a regular expression is used to clean the data of the import and export commodity declaration text.

3. The customs import and export commodity classification method based on the multi-module text intelligent coding algorithm according to claim 1 is characterized in that, The specific implementation method of Step 3 is: Step 31. Send the element content to the keyword search module for keyword replacement: when there is a keyword in the current element, the whole element content will be replaced by the keyword; Step 32. Send the element content to the independent word search module for independent word merging: after the current element is segmented, if there is an independent word in the element, the independent word that may be split into several parts will be merged; Step 33. Send the element content to the synonym replacement module: This module will merge all the "element contents" under a single "declaration element", represent the "element contents" as word vectors through the BERT pre-trained language model, and then calculate the cosine value between the word vectors. The larger this value is, the closer the semantics of the two words are; Set the words with a cosine value greater than the threshold as synonyms, and then perform synonym replacement.

4. The customs import and export commodity classification method based on the multi-module text intelligent coding algorithm according to claim 1, wherein The specific implementation method of Step 4 is: Step 41. Set the number of digits of the random code identifier; Step 42. Randomly generate a random code with a fixed number of digits from the library of numbers + capital letters + lowercase letters; Step 43. Judge whether there is a corresponding random code for the element text in the database. If it exists, convert the element text into a random code; if it does not exist, select a random code to establish a mapping relationship with the element text and store it in the database, and the element text will also be replaced by the corresponding random code.

5. The customs import and export commodity classification method based on the multi-module text intelligent coding algorithm according to claim 1, characterized in that, Step 5 integrates all the encoded element texts, merges the texts with the same codes element by element. When the random codes of the whole text are the same, it is judged whether their product numbers are consistent. If there is an inconsistent situation, it is determined that there is an incorrect declaration for the text other than the correctly declared product number.

Citation Information

Patent Citations

  • Method for determining customs code and method and system for determining type information

    CN108334522A

  • Customs declaration commodity intelligent classification method based on historical data mining

    CN110471948A