An abnormal inspection method for customs import and export commodity texts based on keyword algorithms
By applying the TF-IDF-M algorithm and mutually exclusive vocabulary in the customs import and export commodity declaration text, the inline logical abnormality judgment of text is realized, solving the efficiency and accuracy of customs when checking the declaration text, and improving the inspection efficiency and accuracy.
Patent Information
- Application Number
- CN202111233369.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-22
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-10-22
AI Technical Summary
When checking the declaration text of import and export goods, the customs lacks effective methods to make abnormal judgments inline text logically abnormalities, resulting in frequent abnormal taxation. The existing methods rely on manual judgment and are inefficient.
The keyword scoring method based on the TF-IDF-M algorithm is used to score and standardize the declaration text of customs import and export goods, and the mutually exclusive logic judgment between each element is carried out by loading the factor directory and mutex database to realize the inline logical abnormality judgment of text.
It improves the efficiency of customs import and export goods inspection, can identify abnormal product text information in the first stage, reduces the burden of manual review, and enhances the accuracy of judgment.
Smart Images

Figure CN113946656B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a method for abnormally checking customs import and export commodity texts based on a keyword algorithm. Background Art
[0002] One of the most important functions of the customs is to supervise import and export commodities and levy tariffs. When levying taxes on commodities, the main judgment object of the customs is the declaration text of the commodity. The declaration text is a text set describing various attributes of the commodity, and the collection of attribute names is called the "catalogue of declaration elements for customs import and export commodities". The "catalogue of elements" corresponds one-to-one with the commodity declaration text (element content) filled in by the merchant. The commodity category is marked by a 10-digit commodity number used by the customs. The first 4 digits of the commodity number can be used to locate the "catalogue of elements" for which specific content needs to be filled in for the commodity. During the process of levying commodity taxes, abnormal tax levies often occur due to incorrect commodity declaration texts. At present, during the process of checking commodity declaration texts, customs officers need to have extremely strong customs business capabilities and accurate judgment capabilities, which require years of accumulated learning experience. In recent years, with the continuous increase in the throughput of import and export commodities, the workload of auditing customs officers has increased sharply. How to efficiently and accurately judge the abnormality of import and export commodity declaration texts has become a key concern of the customs. Natural language processing is a study that can explore the hidden relationships existing in text data. It can mine valuable information from historical data in customs commodity declaration texts, simplify the complex, and quickly judge the abnormality of the text while maintaining its judgment accuracy within an acceptable range.
[0003] Most of the existing methods for assisting in the abnormal inspection of customs import and export commodities are based on physical objects or mainly check commodity categories with abnormal records. However, a method for directly performing logical abnormality inspection on the internal of the import and export commodity declaration text has not been realized in the customs field. By judging the abnormality of the inline logic in the text, the abnormal commodity text information can be identified in the first stage of commodity inspection, improving the speed of commodity review and filling the gap in the method for assisting in the abnormal inspection of customs import and export commodities, which is the error recognition of the text itself. However, due to the lack of a unified filling standard for customs import and export commodity declaration texts, it is difficult for unprocessed texts to be judged abnormally by a computer.
[0004] The Term Frequency-Inverse Document Frequency (TF-IDF) algorithm is a classic algorithm for extracting keywords from text. It can filter out words with high term frequency but low importance in the text, such as prepositions, adverbs, modal particles, etc., and then score the importance of other words. Using this algorithm, corresponding keywords can be extracted from the customs import and export commodity declaration text, and then the internal logic of the text can be judged by using the keywords supplemented by the knowledge base. However, when this algorithm is used in the text of the customs field, it often filters out some proper nouns and common nouns with strong semantic features, resulting in a decrease in the accuracy of subsequent tasks. Summary of the Invention
[0005] Aiming at the above problems in the prior art, the purpose of this application is to provide a method for abnormal inspection of customs import and export commodity texts based on keyword algorithms, which realizes the determination of abnormal declaration content in customs commodity declaration texts and effectively improves the inspection efficiency of customs import and export commodities.
[0006] To achieve the above purpose, the technical solution of this application is: A method for abnormal inspection of customs import and export commodity texts based on keyword algorithms, specifically including:
[0007] Step 1: Perform data preprocessing operations on the customs import and export commodity declaration text;
[0008] Step 2: Through the first 4 digits of the commodity number, locate the customs import and export commodity declaration element catalog corresponding to the declared commodity, and split the preprocessed customs import and export commodity declaration text according to the element catalog to form element contents, and the element contents correspond to the element catalog one by one;
[0009] Step 3: Use the TF-IDF-M algorithm to score the keywords of the customs import and export commodity declaration text, and standardize the declaration text through the keywords;
[0010] Step 4: Load the element catalog mutually exclusive word library, and perform mutual exclusion logic determination between elements for a single customs import and export commodity declaration text.
[0011] Furthermore, in step 1, through regular expressions, all lowercase letters in the customs import and export commodity declaration text are converted into uppercase letters; Chinese spaces are deleted, and English spaces are retained; all full-width characters are converted into half-width characters.
[0012] Furthermore, the specific implementation method of step 2 is: Locate the customs import and export commodity declaration element catalog of the current commodity in the database through the first 4 digits of the commodity number, and then split the preprocessed declaration text into separate element contents according to the element catalog, and establish a one-to-one correspondence between the element contents and the element catalog.
[0013] Furthermore, the specific implementation of step 3 is as follows:
[0014] Step 31. For the obtained element catalog and element content, aggregate them under different commodity declaration texts. For the element content of the same declaration element, use the TF-IDF-M algorithm to obtain the word importance of each declaration element, and select the top 4-6% of the words as the keywords under this declaration element;
[0015] Step 32. Traverse the customs import and export commodity declaration texts by element. If the keyword obtained by the TF-IDF-M algorithm is mentioned in the current element content, use the keyword to replace the entire element content.
[0016] Furthermore, the formula involved in the TF-IDF-M algorithm is:
[0017]
[0018]
[0019]
[0020] TF-IDF-M = TF * IDF / M (4)
[0021] Among them, TF W is the term frequency, IDF W is the inverse document frequency, and M is the larger value between the square root of the current term frequency and the square root of the maximum term frequency. This value is used as the penalty term of the TF-IDF-M algorithm, which further improves the algorithm score of some words that frequently appear in a certain category in the customs, reduces the scores of words such as prepositions and adverbs, and at the same time improves the importance scoring of proper nouns and common nouns, improves the confidence of keyword sorting, and provides more accurate information for subsequent abnormal judgment of the inline logic in the commodity declaration text. At the same time, by performing abnormal judgment on the inline logic of the customs import and export commodity declaration text, it realizes real-time detection of abnormal commodity declaration text information, improves the speed of the customs inspection of import and export commodity declaration information, and reduces the manual burden.
[0022] Furthermore, the specific implementation of step 4 is as follows:
[0023] Step 41. Load the mutual exclusion word library of each element. First, judge whether the number of elements with mutual exclusion words in a commodity declaration text exceeds 1; if there is only one element with mutual exclusion words in this commodity declaration text, then this commodity has no possibility of mutual exclusion between words, and it can be directly skipped;
[0024] Step 42. For a commodity declaration record with more than one element having mutually exclusive words, determine whether there is a mutually exclusive relationship between the contents of each element. If so, it is determined that there is an abnormal declaration for this commodity.
[0025] Due to the adoption of the above technical solutions, the present invention can achieve the following technical effects: By using the Term Frequency-Inverse Document Frequency-Max penalty (TF-IDF-M) keyword scoring algorithm and the underlying mutual exclusion algorithm between words, the determination of abnormal declaration content in customs commodity declaration texts is realized, supplementing the deficiencies in the existing inspection of import and export commodity texts by customs, and effectively improving the inspection efficiency of customs import and export commodities. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic flow chart of a method for abnormally inspecting customs import and export commodities. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The present invention will be further described in detail below with reference to the drawings and specific embodiments: This is taken as an example to further describe and explain the present application.
[0028] Embodiment 1
[0029] During the inspection process of customs import and export commodity declaration information, by using the TF-IDF-M keyword query algorithm, various declaration contents can be standardized, reducing the overhead of the subsequent mutual exclusion algorithm between words and improving the inspection efficiency of customs import and export commodity texts. Based on the characteristics of customs texts and the problems in the task of abnormally inspecting customs import and export commodities, refer to Figure 1 , the present application provides a method for abnormally inspecting customs import and export commodities: First, perform data preprocessing operations on the import and export commodity declaration text, including converting lowercase letters in the import and export commodity declaration text to uppercase; deleting Chinese spaces and retaining English spaces; converting full-width characters to half-width characters. Then, according to the first four digits of the commodity number, locate the corresponding "Catalogue of Declaration Elements for Customs Import and Export Commodities" for the declared commodity, and split the customs import and export commodity declaration text according to the element catalogue, so that the element contents correspond one by one to the "element catalogue". Then, for the element contents, use the TF-IDF-M algorithm to score keywords for the customs import and export commodity declaration text, and standardize the text using the keywords. Finally, by loading the mutual exclusion word library of the "element catalogue", for a single customs import and export commodity declaration text, perform the mutual exclusion logic determination between each element.
[0030] This application uses the TF-IDF-M keyword algorithm and the underlying mutual exclusion algorithm between words, effectively complementing the problem of the single inspection method for the existing import and export commodity texts of the customs, and improving the inspection efficiency of the customs' import and export commodities.
[0031] The following will combine examples and drawings to make a detailed description of the present invention, so that those of ordinary skill in the art can implement it with reference to this description.
[0032] This embodiment uses Pycharm as the development platform and Python as the development language. It is carried out on a total of 131,586 sentences of customs real data. The following is the specific process:
[0033] Step 1: Perform data preprocessing operations on the declaration texts of import and export commodities. For example, the data:
[0034] Data A: "6202910090|Cashmere coat|Woven|Cold-proof top|Ladies' style|(less than 6mm)|Pure cotton|aNet".
[0035] Specifically, through regular expressions, all lowercase letters in the declaration texts of import and export commodities are converted into uppercase; Chinese spaces are deleted, and English spaces are retained; all full-width characters are converted into half-width characters.
[0036] Processed Data A: "6202910090|Cashmere coat|Woven|Cold-proof top|Ladies' style|(less than 6MM)|Pure cotton|ANET".
[0037] Step 2: Locate the "element catalog" of the current commodity in the database through the first 4 digits of the commodity number, and then split the declaration text into separate "element contents" according to the elements, and establish a one-to-one correspondence between the "element contents" and the "element catalog". For example, the data:
[0038] Data A: "6202910090|Cashmere coat|Woven|Cold-proof top|Ladies' style|(less than 6MM)|c Pure cotton|ANet". After processing, it is transformed into:
[0039] "Element content" - "Element name"
[0040] {"Cashmere coat": "Commodity name"
[0041] "Woven": "Weaving method"
[0042] "Cold-proof top": "Type (windproof coat, short coat, cloak, etc.)"
[0043] "Ladies' style": "Category"
[0044] "(less than 6MM)": "Fiber length"
[0045] "Pure cotton": "Component content"
[0046] “ANET”: "Brand"
[0047] }
[0048] Step 3: For the element content obtained in Step 2, use the TF-IDF-M algorithm to score the keywords in the customs import and export commodity declaration text, and standardize the text using the keywords. Specifically:
[0049] Step 31: For the "element catalog" and "element content", aggregate them under different commodity declaration texts. For the declaration content of the same declaration element, use the TF-IDF-M algorithm to obtain the word importance under each declaration element, and select the top 4-6% of the words as the keywords under this element. For example, the data:
[0050] {
[0051] "Cashmere coat": Keywords ["cashmere"]
[0052] "Pure cotton": Keywords ["pure cotton"]
[0053] }
[0054] Step 32. Traverse the import and export commodity declaration text by element. If the keyword obtained by the TF-IDF-M algorithm is mentioned in the element content, use the keyword to replace the entire element content.
[0055] "Cashmere coat": After replacement, it becomes: ["cashmere"].
[0056] "Pure cotton": After replacement, it becomes: ["pure cotton"].
[0057] Step 4: Load the mutually exclusive word library of the "element catalog", and perform the mutual exclusion logic determination between elements for a single customs import and export commodity declaration text. Specifically:
[0058] Step 41. Load the mutually exclusive word library of each element. First, judge whether the number of elements with mutually exclusive words in a commodity declaration text exceeds 1. If there is only one element with mutually exclusive words in this declaration, then this commodity has no possibility of mutual exclusion between words and can be skipped directly.
[0059] Step 42. For the commodity declaration records with more than 1 element with mutually exclusive words, judge whether there is a mutually exclusive relationship between the contents of each element. If so, it is determined that there is an abnormal declaration situation for this commodity. Classification difference examples:
[0060] Data A: "6202910090|Cashmere coat|Woven|Cold-proof top|Ladies' style|(<6MM)|Pure cotton|ANet". Its content is an objective and neutral record based on historical data, and there is no right or wrong in a single paragraph. However, when combined, the "wool" and "pure cotton" are contradictory, that is, the word-by-word verification shows a declaration error. It is very difficult to detect such declaration errors of goods through the existing customs inspection process. The TF-IDF-M keyword algorithm and the underlying word mutual exclusion algorithm can effectively supplement the problem of the single inspection method for existing import and export commodity texts of the customs, and improve the inspection efficiency of customs import and export goods.
[0061] As mentioned above, the above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered within the protection scope of the present invention.
Claims
1. A method for abnormal inspection of customs import and export commodity texts based on keyword algorithms, characterized in that, it specifically includes: Step 1: Perform data preprocessing operations on the customs import and export commodity declaration texts; Step 2: Through the first 4 digits of the commodity number, locate the customs import and export commodity declaration element catalog corresponding to the declared commodity, and split the preprocessed customs import and export commodity declaration text according to the element catalog to form element contents, and the element contents correspond one-to-one to the element catalog; Step 3: Use the TF-IDF-M algorithm to score keywords for the customs import and export commodity declaration text, and standardize the declaration text through the keywords; the specific implementation method is: Step 31. For the obtained element catalog and element contents, aggregate them under different commodity declaration texts. For the element contents with the same declaration elements, use the TF-IDF-M algorithm to obtain the word importance under each declaration element, and select the top 4-6% of the words as the keywords under the declaration element; Step 32. Traverse the customs import and export commodity declaration text by element. If the keyword obtained by the TF-IDF-M algorithm is mentioned in the current element content, use the keyword to replace the entire element content; The formula involved in the TF-IDF-M algorithm is: TF-IDF-M = TF * IDF / M (4) Among them, TF W is the term frequency, IDF W is the inverse document frequency, and M is the larger value between the square root of the current term frequency and the square root of the maximum term frequency. This value serves as the penalty term for the TF-IDF-M algorithm; Step 4: Load the element catalog mutual exclusion word library, and perform mutual exclusion logic determination between elements for a single customs import and export commodity declaration text.
2. The method for abnormal inspection of customs import and export commodity texts based on keyword algorithms according to claim 1, characterized in that, in Step 1, through regular expressions, convert all lowercase letters in the customs import and export commodity declaration text into uppercase letters; delete Chinese spaces and retain English spaces; convert all full-width characters into half-width characters.
3. The method for abnormal inspection of customs import and export commodity texts based on keyword algorithms according to claim 1, characterized in that, the specific implementation method of Step 2 is: Locate the customs import and export commodity declaration element catalog of the current commodity in the database through the first 4 digits of the commodity number, and then split the preprocessed declaration text into separate element contents according to the element catalog, and establish a one-to-one correspondence between the element contents and the element catalog.
4. The method for abnormal inspection of customs import and export commodity texts based on keyword algorithms according to claim 1, characterized in that, the specific implementation method of Step 4 is: Step 41. Load the mutual exclusion word library of each element. First, judge whether the number of elements with mutual exclusion words in a commodity declaration text exceeds 1; if there is only one element with mutual exclusion words in this commodity declaration text, then this commodity has no possibility of mutual exclusion between words, and it can be skipped directly; Step 42. For the commodity declaration records with the number of elements with mutual exclusion words exceeding 1, judge whether there is a mutual exclusion relationship between the element contents. If there is, it is determined that there is an abnormal declaration situation for this commodity.
Citation Information
Patent Citations
Text-subject-model-based data processing method for commodity classification
CN102929937A
TFIDF-based medical symptom keyword extraction optimization and recovery method and system
CN108133752A