Method for finding a unique Harmonized System code from a given text using an artificial neural network and system for implementing the method
Through a neural network model based on supervised learning, the allocation of customs tariff numbers is automated, which solves the problems of large human errors and complex processing in existing technologies and achieves more efficient and accurate number allocation.
Patent Information
- Application Number
- CN201880095360.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-07-04
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2038-07-04
AI Technical Summary
Existing technologies suffer from large human errors, complex and time-consuming processing in the allocation of customs tariff numbers. Existing methods rely on explicit user input and manual inspection, resulting in low efficiency.
A neural network model with five clustering concepts based on supervised learning is used to automatically classify items and assign unique numbers through preprocessing, vectorization, and training, reducing user intervention, simulating human inspection behavior, and gradually refining the number assignment process.
It improves the accuracy and efficiency of customs tariff number allocation, reduces human errors, simplifies operational processes and reduces processing time.
Smart Images

Figure CN112513901B_ABST
Abstract
Description
Technical Field
[0001] The invention disclosed herein generally relates to a method for assigning a specific / unique number to any proposed article according to an established set of standards such as the Harmonized System (HS), which is an internationally valid common commodity nomenclature standard, and a system implementing the method so as to assign the specific / unique identification number to any proposed article or goods. Background Art
[0002] The Harmonized Commodity Description and Coding System, commonly referred to as the "Harmonized System" or simply "HS," is a multi-purpose international product nomenclature developed by the World Customs Organization (WCO). During the import and export process, purchasing organizations act accordingly to this system, applying different tariff schemes for virtually every country and region. Therefore, finding a method for accurately and precisely calculating and assigning customs tariff numbers to any goods, articles, or merchandise is crucial. While the data processing load stemming from the large number of commodity groups (approximately 5,000) has been logically and legally structured, the technology at this stage still harbors considerable susceptibility to human error.
[0003] To address these concerns, the present invention focuses on a supervised learning-based approach with technical rigor; namely, a decision-making mechanism based on a neural network trained with data from five clusters capable of confidently classifying concepts.
[0004] Regarding the prior art, document KR 101571041 (B1) discloses a system for the Harmonized System (HS) classification of goods, comprising an interface processing unit for selecting an interface for receiving input; a database having HS code corresponding information; and an HS code determination unit. The disclosure is subject to a set of system-wise rules and constraints such as inclusion / exclusion and requires explicit user input. US 2016275446 (A1) discloses an apparatus and method for determining an HS code, receiving selectable determining factors from a user, and ascertaining similarity by comparing with a memory storage.
[0005] WO 2016057000 (A1), which is identified as one of the disclosures in this technical field, defines a learning-oriented classification method that states ad hoc learning rules that are subsequently adopted to improve the overall processing capabilities for a variety of goods. Other prior art documents such as EP 2704066 (A1) are related to transaction classification based on analysis of multiple variables, which is considered to facilitate the identification and classification of items. US 20050222883 (A1) discloses a general system and method for brokerage operation support, which follows steps including receiving information related to freight rates for a specific country and is characterized by an architecture including attended and unattended servers and different workstations.
[0006] Purpose of the Invention
[0007] The main purpose of the present invention is to provide a method and system for assigning customs tariff numbers to articles and goods according to the Harmonized System (HS), characterized by the application of supervised learning based on artificial neural networks trained specifically on conceptualized categories of 5 clusters, which in turn proposes a powerful and more accurate and more precise alternative when considering closely following the prior art. Summary of the Invention
[0008] In the proposed invention, an item classification and unique identifier assignment scheme is implemented that is primarily centered around an artificial neural network, or supervised learning algorithm; however, while reducing errors and improving integrity, it minimizes user intervention and the need for explicit input. In contrast to the standard method of manual inspection, which typically prioritizes manual inspection, a processing unit and database work in concert with the input methods for the traded items to be verified at customs. The machine learning method, which forms the core of the disclosed invention, is inspired by and modeled after the aforementioned standard method of manual inspection performed by actual humans, whose uncertainty leads to time-consuming binding reports to be issued and processed. In this regard, a drawback of machine learning methods using full-text search in situations with sparse pattern visibility and similar parameter ranges is that they develop models that produce outputs with considerable error. To address this, a conceptually defined parameter set is outlined in a tree-like format, which is used to increase training rigor by establishing dedicated training sets for each specific level and step, as described below.
[0009] The present invention proposes a series of processing events that are co-linear with the learning application and that make it possible to approach the problem of classification and assignment of items in a conceptual way: any ranking or possible state that an item can assume is divided into smaller parts, the logic behind which is to reduce the uncertainty and errors that may be caused entirely by the apparent size of the data set, which is related to the technical characteristics and specifications of any of the components of the item in question. Starting from the beginning, this means that a lot of data may seem unrelated and cannot be presented. As will be explained in detail in the following sections, the number to be assigned to a certain item can be divided into four pairs and four-bit bytes, which when respectively combined together constitute the category, chapter, position, sub-position and finally the complete number. This five-fold structure is maintained in an artificial neural network structure, each of which itself contains an input layer, a hidden layer and an output layer.
[0010] Initially, a query is formed through a series of preprocessing steps. The text is converted to all lowercase and parsed, then converted to all UTF-8 characters. After that, they are broken down into syllables and reorganized into phonemes and morphemes, and later divided into two- and three-word phrases. Next, all combinations of these phonemes, morphemes, and phrases are converted to vectors and then normalized. The next step is the training of the artificial neural network, which is combined with the output layer function to minimize error and reduce over / underfitting. Once the model is implemented, it is saved as a file on the local disk.
[0011] The general approach to number assignment leads to the step-by-step modeling of the number suggestion / prediction structure: the model to be trained on HS-based data is designed based on the previously mentioned grouping of numbers, which leads to the creation of different partitions of the training. Thus, the output of the category suggestion neural network is fed into the section suggestion / prediction neural network, whose output is in turn fed into the position suggestion / prediction neural network, and so on. The final set of suggestions / predictions is then displayed on the screen to the human user. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are provided merely for the purpose of illustrating the unique identification number allocation method and the system for implementing the method, the advantages of which over the prior art are summarized above and will be briefly described below.
[0013] The drawings are not intended to define the scope of protection identified in the claims, nor should the drawings alone be referenced in an attempt to interpret the scope identified in the claims without relying on the technical disclosure in the description of the present invention.
[0014] Figure 1 A possible relationship diagram regarding the customs number allocation method according to the present invention is demonstrated.
[0015] Figure 2The pre-processing steps performed before the training phase according to the present invention are demonstrated.
[0016] Figure 3 The file hierarchy for a training set of an artificial neural network for a hypothetical object according to the present invention is demonstrated.
[0017] Figure 4 An exemplary diagram illustrating the alleged information flow regarding different artificial neural networks when suggesting numbers to items according to the present invention is presented.
[0018] Figure 5 The important preprocessing steps and training process of the artificial neural network at each level according to the present invention are demonstrated.
[0019] Figure 6 The entire artificial intelligence model according to the present invention, which generally includes pre-processing, training and prediction / suggestion, is demonstrated. DETAILED DESCRIPTION
[0020] The present invention discloses a method for automatically assigning unique numbers to trade items or goods based on a query, and a system implementing the same. Artificial neural networks are used as a key element of the method to achieve increased flexibility in an otherwise rigid and complex process, as well as reduced error in a more general and straightforward approach. The unique number to be attributed to the item in question can have different components in digital format, with different components carrying information related to different aspects or characteristics of the item. Such a process has many factors to consider and numerous characteristics and constraints to carefully consider, making it laborious and time-consuming. To this end, the disclosed method and system include data processing features that are specifically designed to utilize machine learning concepts in processing mechanisms to handle various information, thereby replacing actual humans performing the aforementioned tasks. Different items present different characteristics and areas of use, which become issues when assigning unique identifiers to them, often requiring numerous lookup tables (LUTs) and accompanying laboratory testing. The disclosed invention addresses all of these issues, while also providing a systematic and procedural number / identifier assignment scheme that utilizes conceptual classification features suitable for the system and problem in question.
[0021] The strength of the technology present in the disclosed invention is built on the premise that the conceptual taxonomy of the nomenclature process is used to model the solution to generate reliable and accurate output: in areas where a large number of items need to be classified and identified, the system and method are able to assign identification numbers to items / goods with a time / accuracy trade-off that cannot be undertaken by real humans such as officials and inspectors. The technical specifications of non-generic items can vary greatly, and as such recent decades have seen a proliferation of categories, the resulting expertise burden has become cumbersome but time expensive. In various embodiments, legal liability is also taken into account, as are legal precedents, all of which is made possible with optional connectivity to ERP and web and FTP. Utilizing available web services, it becomes possible to collaboratively enhance the information available on multiple platforms, including legislative and item databases, and optionally also cases and lists covering different countries / states.
[0022] The method and system disclosed in the present invention accept two types of indirect input and one type of direct input: direct input is manual text entry, which involves a verbal definition of the article / goods in the desired language; whereas indirect input refers to catalogs and technical specifications used when externally evaluating the article / goods. Direct input is raw input and therefore needs to be processed according to customs guidelines before being used for text searching and training purposes. Once the text search is completed and the output is obtained, the legislative and article databases begin to be delivered as the evaluation process progresses. The legislative database can include information related to countries / states in multiple languages, relevant treasury information such as taxes and funds, legislation related to debts and sanctions (if any), and comparative legal status between two countries / states or similar articles / goods selected by the user.
[0023] As mentioned earlier, human text entry is raw and contains snippets of information that may be irrelevant, useless, or confusing. To circumvent this and generate data in the form of numerical vectors suitable for classification purposes, a series of steps are performed in the preprocessing process. When text of any length arrives, the lowercase processing step begins, during which the text is converted to all lowercase letters, while repeated words are removed for simplicity and size reduction. After lowercase processing, parsing is performed, where each word in the text is separated, converted to UTF-8 format, punctuation and non-characters are discarded, and then the word is deconstructed into syllables and morphemes. The final aspect of parsing is pair and triple word alignment, which can also be called tokenization. Vector conversion or vectorization is a step that numerically quantizes all tokenized and non-tokenized instances of the data input or constructed in the previous step, such as morphemes, phrases, tokens, and temporary pairs, and then forms a vector containing the numerical representation according to the following formula, where x represents the vector value of the word, c represents the character in the word, h1 and h2 represent two different experimental constants, j represents the character index in the word, and k represents the number of characters in the word:
[0024] x i =h1,x i =(x j íx i )B h2,j=1,2,...,k
[0025] Normalization, the final step of preprocessing, confines the set of vectors to a smaller range, since words can include a variety of ASCII characters, resulting in an unrealistic and therefore undesirable level of variation. This is done according to the formula given below, where x represents the word vector, y represents the sequence of text vectors, and j indicates the word count in the sequence:
[0026] y i =x i B(1·ij j 1)
[0027] The hidden layer is essentially an m by n matrix and refers to the training pool where all text entries are associated with their associated labels. Once the text is preprocessed and a vector is generated, the vector is multiplied by the hidden layer matrix and thus incorporated into the neural network, the details of which are shown below, where y corresponds to the vector of size n, B corresponds to the hidden layer matrix, b corresponds to the elements of the matrix B, and z corresponds to the output vector of size m.
[0028]
[0029] The output layer function delivers the proposed result to the final state of the output layer, which is used for both training and classifying instances. This function is given below, where z represents the proposed result, i represents the index of the proposed label, and j represents the total number of labels:
[0030] S(z i )=exp(z i )·ij j exp(z j )
[0031] The artificial neural network in the disclosed invention is generally designed to include fully connected hidden layers and outer layers. Immediately following the final output layer, there is a cost function whose purpose is to minimize the error after training for a certain number of epochs. C is the cost of the function, where lr is the learning rate, / is the actual value of the label and z i is a suggested value for the trade-off between expected error and training cost in the function outlined below:
[0032] C i =Ir B(Iz i )
[0033] Since the problem in question is modeled after the classification / attribution behavior of actual humans who inspect any item / goods, the different parts should present themselves as inspection levels so that they can be imagined as constituting steps to be taken during a routine inspection. This is possible by dividing the protocol number into smaller, but still meaningful parts and thus bringing the solution to a step-by-step, incremental quality. For example, the Harmonized System (HS) protocol assumes that a 12-digit code is valid in its entirety, which assumes the format XXXX.XX.XX.XX.XX, which by default provides a template for its subdivision. The most significant digits are the leftmost six to be predicted in groups of two, and the remainder of the code will be generated as a complete whole in the final stage of classification; including categories that proved particularly advantageous during the initial sorting of the items.
[0034] The disclosed invention differs from other full-text search and machine learning approaches in that it categorizes the classification of topics into conceptual components, thereby significantly better mimicking implicit decision-making mechanisms performed in real time. The same data is processed in batches of five, thereby marginalizing the error at each level by ensuring a four-fold error rate relaxation between them, as opposed to the number of epochs required for training to reduce the error rate to the desired level. In doing so, an indirect conceptualization of the data is ensured, similar to the conceptualization of an actual human dividing the unique identification number into smaller parts to incrementally examine the item / goods. Based on the Harmonized System (HS), one model for 21 categories, 21 models for 98 chapters, 98 models for 1239 locations, 1239 models for 5407 sub-locations, and finally a total of 5407 models for the full range of unique identification numbers were formed for training of corresponding artificial neural networks with fully connected hidden layers. Between each level of the artificial neural network a tree-like classification hierarchy is ensured, the first four of which are each associated with the most significant consecutive digit pairs; this is combined with the legislative database and regional information that can be accessed and utilized according to at least one embodiment.
[0035] After training, the system initially executes at the first level, suggesting categories for items / goods. The suggested categories are then prepared as input to the next level, the second level. The output of this next level generates suggestions for sections and is passed on to serve as input to the next level, the third level, and so on. Once all levels have executed their suggestions, the final result is presented in a form that, according to at least one embodiment, can be used in conjunction with other modules of the system, such as the legislative database and the parallelized item / goods list, to contain closely overlapping instances of unique identification numbers. In this context, it is reasonable to reflect that this system and method does not strictly limit its operation to the exact classification behavior of actual humans, namely, sequentially determining category, section, location, and sub-location; in alternative embodiments, it may also be the case that alternative correlations are also utilized. However, it should be noted that the need for categories by actual humans is reduced, and it is sometimes even possible to directly suggest unique identification numbers, which is not only inappropriate but also impossible given the system's architectural approach.
[0036] In various embodiments, the use of the methods and systems is further facilitated through the integration of various other components, such as legislative and article databases, and similar article modules, all of which serve to expand practical understanding of the disclosed invention. For example, one embodiment can request alternative suggestions based on legislation and article databases in different states. Another embodiment can retrieve matching tariffs for comparatively proposing unique identification solutions for consideration by the end user. These features / modules are not mutually exclusive, as one embodiment can include all of them simultaneously, thereby ensuring coordination with ERP systems.
[0037] In short, the disclosed invention proposes a method for assigning unique identification numbers, such as those implemented by the Harmonized System (HS), to articles and goods of any nature or area of use, by means of an artificial neural network chain capable of implicitly conceptualizing data within categories, chapters, locations, and sub-locations through a dedicated neural network trained for the purpose; the dedicated neural network includes modules and processing units that undertake the classification and assignment practices. The different modules interact with each other according to different embodiments; all of the different embodiments are optionally and cooperatively employable.
[0038] In one aspect of the present invention, a method for assigning unique identification numbers to items and goods in a systematic and conceptual manner is presented.
[0039] In one aspect of the present invention, the unique identification number assignment method includes designing at least one artificial neural network having a fully connected hidden layer.
[0040] In another aspect of the present invention, the unique identification number assignment method includes training the artificial neural network with files in the form of a training set and generating a model resulting therefrom.
[0041] In another aspect of the present invention, the unique identification number assignment method includes forming an importance-based hierarchical sequence of the artificial neural networks, ie, the output of one of the neural networks is the input of the next neural network.
[0042] In another aspect of the present invention, the unique identification number assignment method includes introducing pre-processed text vectors as input to the first level of the artificial neural network sequence, and.
[0043] In another aspect of the present invention, the unique identification number assignment method includes obtaining a unique identification number for the item / goods in question.
[0044] In another aspect of the present invention, the unique identification number assignment method includes establishing communication with at least one external module to accept raw text input.
[0045] In another aspect of the present invention, the unique identification number assignment method includes a pre-processing step for generating an input compatibility vector from an original text query.
[0046] In another aspect of the present invention, the unique identification number assignment method includes lower case processing, wherein each character in the original text query is converted to lower case letters and repeated words are eliminated.
[0047] In another aspect of the present invention, the unique identification number assignment method includes parsing, wherein words are parsed and converted into UTF8 format, punctuation marks and non-characters are discarded, and words are broken down into syllables and morphemes and arranged in groups of two and three to form tokens.
[0048] In another aspect of the present invention, the unique identification number assignment method includes vector conversion, wherein the tokenized instances and the non-tokenized instances of the data in the previous step are converted into digital vectors.
[0049] In another aspect of the present invention, the unique identification number allocation method includes normalization, in which the vector formed so far is normalized so as to be confined to a specific interval.
[0050] In another aspect of the present invention, the unique identification number assignment method includes the artificial neural network training step, further including a first-level training scheme for categories; a second-level training scheme for chapters; a third-level training scheme for locations; a fourth-level training scheme for sub-locations; and a fifth-level training scheme for identification numbers as a whole.
[0051] In another aspect of the present invention, the unique identification number assignment method includes the first, second, third and fourth level training schemes for category, chapter, location and sub-location, respectively including the division of number pairs starting from the leftmost and descending in order of importance.
[0052] In another aspect of the present invention, the unique identification number allocation method includes the message verification step, wherein the authentication of the packet is always confirmed only when all steps are performed correctly, otherwise it is discarded.
[0053] In one aspect of the present invention, a system for assigning unique identification numbers to items and goods is proposed, comprising at least one processing unit and a database.
[0054] In one aspect of the invention, the processing unit comprises modules capable of executing parallelized instances of an artificial neural network.
[0055] In one aspect of the invention, the processing unit includes manual, web service and FTP input capabilities.
[0056] In one aspect of the invention, the database includes a list relating to a catalog of items and a list of unique identification numbers associated therewith.
[0057] In one aspect of the invention, the database and processing unit comprise connections to additional modules selected from the group consisting of: legislation comparison module, similar items comparison module, matching tariff module.
[0058] In one aspect of the invention, the legislation comparison module comprises at least one piece of legislation information relating to at least one country other than the user's own country.
[0059] In one aspect of the invention, the processing unit includes an ERP compatibility feature.
Claims
1. A method for assigning a unique identification number to an item, comprising the following steps: Designing a plurality of artificial neural networks, each of the artificial neural networks comprising a fully connected hidden layer; training the plurality of artificial neural networks using files in the form of training sets to generate a plurality of artificial neural network models; forming an importance-based hierarchical artificial neural network sequence of the plurality of artificial neural network models, the sequence being organized into: a first level neural network model for suggesting categories; a second level neural network model for suggesting chapters; a third level neural network model for suggesting locations; a fourth level neural network model for suggesting sub-locations, and a fifth level neural network model for suggesting at least one identification number overall, in, using the output of the first level neural network model as the input of the second level neural network model, using the output of the second level neural network model as the input of the third level neural network model, using the output of the third level neural network model as the input of the fourth level neural network model, using the output of the fourth level neural network model as input to the fifth level neural network model, and The identification number suggested by the fifth-level neural network model includes the category suggested by the first-level neural network model, the chapter suggested by the second-level neural network model, the location suggested by the third-level neural network model, and the sub-location suggested by the fourth-level neural network model; Introducing pre-processed text vectors about products as inputs to the first level neural network model of the artificial neural network sequence; and A prediction of a unique identification number for the item is obtained from the artificial neural network sequence.
2. The method for assigning a unique identification number to a product according to claim 1, further comprising: Communication is established with at least one external module for accepting raw text input.
3. The method for assigning a unique identification number to a product according to claim 1, further comprising: A preprocessing step for generating input compatibility vectors from raw text queries.
4. The method for assigning a unique identification number to a product according to claim 3, wherein: The pre-processing step comprises the following further distinct steps: Lowercase processing, in which each character in the original text query is converted to lowercase and repeated words are eliminated; Parsing, where words are parsed and converted to UTF8 format, punctuation and non-characters are removed, words are broken into syllables and morphemes, elements are arranged in groups of two and three, and tokens are formed; Vector transformations, where tokenized and untokenized instances of data are converted to numeric vectors; and Normalization, where a vector is normalized so as to be confined to a specific interval.
5. The method for assigning a unique identification number to a product according to claim 1, wherein: The categories suggested by the first-level neural network, the chapters suggested by the second-level neural network, the positions suggested by the third-level neural network, and the sub-positions suggested by the fourth-level neural network respectively represent the divisions of the digit pairs of the identification numbers of the goods, starting from the leftmost digit pair of the identification numbers and arranged in descending order of importance.
6. A system for assigning unique identification numbers to goods, comprising at least one processing unit and a database, characterized in that: The processing unit includes a module capable of executing a parallelized instance of an artificial neural network, wherein the parallelized instance of the artificial neural network includes: a first level neural network instance for suggesting categories, a second level neural network instance for suggesting sections, a third level neural network instance for suggesting locations, a fourth level neural network instance for suggesting sub-locations, and a fifth level neural network instance for suggesting at least one identification number in general, using the output of the first level neural network instance as the input of the second level neural network instance, using the output of the second level neural network instance as input to the third level neural network instance, using the output of the third level neural network instance as input to the fourth level neural network instance, using the output of the fourth level neural network instance as input to the fifth level neural network instance, and The identification number suggested by the fifth-level neural network instance includes the category suggested by the first-level neural network instance, the chapter suggested by the second-level neural network instance, the location suggested by the third-level neural network instance, and the sub-location suggested by the fourth-level neural network instance; The processing unit includes manual, web service and FTP input capabilities; The database includes a list of items related to the catalog and a list of unique identification numbers associated with the list; and The database and processing unit include a connection to at least one of a legislation comparison module, a similar items comparison module, or a matching tariff module.
7. The system for assigning unique identification numbers to goods according to claim 6, characterized in that: The legislation comparison module also includes at least one legislation information portion relating to at least one country other than the user's own country.
8. The system for assigning unique identification numbers to goods according to claim 6, characterized in that: The processing unit includes an ERP compatibility feature.
Citation Information
Patent Citations
Multivariate transaction classification
EP2704066A1
System for determining Harmonized System(HS) classification
KR101571041B1
Market expansion through optimized resource placement
US20050222883A1
Apparatus and method for determining HS code
US20160275446A1
Relevance score assignment for artificial neural network
CN107636693A