Product Classification Coding Determination Method, System and Related Devices

By using a pre-trained target classification model, its classification code is automatically determined based on the product name, solving the problem of inefficiency in manually finding tax classification codes, and achieving fast and accurate automatic classification.

CN114722198BActive Publication Date: 2025-07-01KINGDEE SOFTWARE(CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210333381.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-07-01
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

During the production and operation process, it is a huge and repetitive work to manually find the corresponding tax classification code through the product name, and it is prone to errors, resulting in inefficiency.

Method used

By obtaining the pre-trained target classification model, the model is obtained by machine learning training of the initial classification model by training samples. The model preserves the target correspondence between product name features and product category, and then determines its classification code based on the full name of the target product to be classified.

Benefits of technology

It realizes the rapid and accurate automatic determination of the corresponding product classification code through the product name, which reduces manual workload, improves efficiency and reduces error rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114722198B_ABST
    Figure CN114722198B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a method, a system and related devices for determining product classification codes. The target classification model of this method is obtained by performing machine learning training on an initial classification model using training samples. The training samples include classified product name features, product classification codes corresponding to the classified product names, and unclassified product name features. Therefore, the target classification model stores the target correspondence relationship between product name features and product categories, so that when the target product name features of the full name of the target product are input into the target classification model, the product category to which the target product belongs and its corresponding target classification code can be predicted. Therefore, the embodiments of the present application can quickly and accurately automatically identify the product classification code corresponding to a product through the product name, without the need to configure a large amount of human resources and wait for a long time for the result, thereby improving the work process of product classification, sorting, or accounting, etc., and enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of Internet technologies, and in particular, to a method, a system, and related devices for determining product classification codes. Background Art

[0002] During the production and operation process, it is often necessary to classify and identify various products to improve the efficiency during classification, sorting, or statistics. For example, the finance department needs to issue value-added tax invoices for different commodities. Among them, the commodities on the invoice when issuing the invoice should be associated with the tax classification codes approved by the State Administration of Taxation. Only in this way can the invoice be correctly issued according to the tax rates and collection rates indicated on the classification codes, thereby improving the process of the tax authorities' statistics and analysis of data and enhancing the collection management efficiency.

[0003] In practical applications, if each value-added tax invoice needs to manually find its corresponding tax classification code through the commodity name, the manual workload is huge and repetitive. Moreover, with a rich sampling of commodities, manual classification is also prone to errors and poor timeliness. Therefore, it is necessary to provide an efficient method for determining product classification codes to solve the above problems. Summary of the Invention

[0004] The embodiments of the present application provide a method, a system, and related devices for determining product classification codes, which are used to automatically and accurately determine the corresponding product classification code according to the product name.

[0005] The first aspect of the embodiments of the present application provides a method for determining a product classification code, including:

[0006] Obtain a pre-trained target classification model, where the target classification model is obtained by performing machine learning training on an initial classification model using training samples. The training samples include classified product name features, classification codes corresponding to the classified product names, and unclassified product name features. The target classification model stores the target correspondence between product name features and product categories;

[0007] Determine the target product full name feature according to the full name of the target product to be classified;

[0008] Input the target product full name feature into the target classification model to output the target classification code determined by the target classification model according to the target correspondence. The target classification code is used to represent the product category to which the target product belongs.

[0009] Optionally, the weighted fusion of the full-text probability value and the local probability value corresponding to the same product category includes:

[0010] Sum the full-text probability value and the local probability value corresponding to the same product category to obtain a probability subtotal value for the current product category;

[0011] Calculate the maximum difference corresponding to each product category according to each probability subtotal value, and sum all the full-text probability values and local probability values to obtain the total probability value;

[0012] Calculate the fusion result of each product category according to the probability subtotal value, the maximum difference, and the total probability value;

[0013] Filter the fusion results that meet the fusion result conditions according to the numerical sizes of the respective fusion results, where the respective probability subtotal values and / or the respective fusion results can be sorted by size respectively.

[0014] The second aspect of the embodiments of the present application provides a product classification code determination system, including:

[0015] An acquisition unit, configured to acquire a pre-trained target classification model, where the target classification model is obtained by performing machine learning training on an initial classification model using training samples, the target classification model stores a target correspondence between product names and product categories, and the training samples include classified product name features, classification codes corresponding to the classified product name features, and unclassified product name features;

[0016] A processing unit, configured to determine a target product full name feature according to the full name of a target product to be classified;

[0017] The processing unit is further configured to input the target product full name feature into the target classification model to output a target classification code determined by the target classification model according to the target correspondence, where the target classification code is used to represent the product category to which the target product belongs.

[0018] Optionally, the acquisition unit is further configured to acquire training samples, where the training samples include classified product name features, classification codes corresponding to the classified product name features, and unclassified product name features;

[0019] The processing unit is further configured to perform machine learning training on the initial classification model using the training samples to obtain a target classification model that stores the target correspondence, where the model parameters of the initial classification model are determined by a loss function defined using a semi-supervised learning algorithm.

[0020] Optionally, the acquisition unit is specifically configured to:

[0021] Acquire classified product full names and unclassified product full names, and perform text preprocessing on the classified product full names and the unclassified product full names to obtain name segmentations for forming each product full name;

[0022] Determine the initial value of the position index of each name segmentation represented in a sparse vector form in the product word library;

[0023] The initial value of the position index in the form of a sparse vector is converted into a position index mapping value in the form of a dense vector through a word embedding generation model, and the position index mapping value is used as the product name feature in the training sample to train the initial classification model.

[0024] Optionally, the obtaining unit is specifically configured to:

[0025] Uniformly convert the letters in the full names of the classified products and the unclassified products into letters in only uppercase or only lowercase forms, and / or remove duplicate words, parentheses and the text content in the parentheses in the product full names;

[0026] Divide the cleaned product name into at least one name token.

[0027] Optionally, the processing unit is specifically configured to:

[0028] Obtain the full name of the target product, and perform text preprocessing including text cleaning and tokenization on the full name of the target product to obtain target name tokens for composing the full name of the target product;

[0029] According to the target name tokens, convert to obtain a target product full name feature with the same vector dimension as the classified product name feature;

[0030] The processing unit is specifically configured to: input the target product full name feature with the same vector dimension as the classified product name feature into the target classification model to determine the target classification code.

[0031] Optionally, the processing unit is specifically configured to:

[0032] When the target product full name feature input into the target classification model is obtained from target name concatenation words, determine whether the maximum probability value among the K full-text probability values output by the target correspondence relationship satisfies the full-text classification probability threshold, where K is a non-zero integer set by the user, and the target name concatenation words are obtained by concatenating multiple target name tokens for composing the full name of the target product, and each full-text probability value corresponds to a product category;

[0033] If the judgment result is satisfied, determine the product category corresponding to the maximum probability value as the target product category to which the target product belongs, and output the classification code corresponding to the target product category as the target classification code;

[0034] If the judgment result is not satisfied, input the target product full name features obtained from the tokens with the part-of-speech of nouns in each target name token into the target classification model respectively to obtain local probability values of the target product belonging to different product categories;

[0035] Perform weighted fusion on the full-text probability value and the local probability value corresponding to the same product category to determine that the classification code matched by the product category with the maximum fusion result is the target classification code.

[0036] Optionally, the processing unit is further configured to:

[0037] Determine whether the full name of the target product is recorded in the dictionary file, where the dictionary file pre-stores the full names of classified products and the classification codes correctly corresponding to the full names of the classified products;

[0038] If the judgment result is yes, directly output the classification code corresponding to the recorded full name of the target product in the dictionary file as the target classification code;

[0039] If the judgment result is no, determine the target classification code through the target classification model.

[0040] Optionally, the processing unit is specifically configured to:

[0041] Sum the full-text probability value and the local probability value corresponding to the same product category to obtain a probability subtotal value for the current product category;

[0042] Calculate the maximum difference corresponding to each product category according to each probability subtotal value, and sum all the full-text probability values and local probability values to obtain a total probability value;

[0043] Calculate the fusion result of each product category according to the probability subtotal value, the maximum difference, and the total probability value;

[0044] Filter the fusion results that meet the fusion result conditions according to the numerical sizes of the fusion results, where the probability subtotal values and / or the fusion results can be sorted by size respectively.

[0045] The third aspect of the embodiments of the present application provides a device for determining a product classification code, including:

[0046] A central processing unit, a memory, and an input / output interface;

[0047] The memory is a transient storage memory or a persistent storage memory;

[0048] The central processing unit is configured to communicate with the memory and execute the instruction operations in the memory to execute the method described in the first aspect or any specific implementation manner of the first aspect of the embodiments of the present application.

[0049] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium, including instructions, which when run on a computer, cause the computer to execute the method described in the first aspect or any specific implementation manner of the first aspect of the embodiments of the present application.

[0050] A computer program product is provided in the fifth aspect of the embodiments of the present application. When the computer program product runs on a computer, it causes the computer to execute the method described in the first aspect of the embodiments of the present application or any specific implementation manner of the first aspect.

[0051] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:

[0052] The target classification model of the embodiments of the present application is obtained by performing machine learning training on the initial classification model with training samples. The training samples include the classified product name features, the product classification codes corresponding to the classified product names, and the unclassified product name features. Therefore, the target classification model stores the target correspondence relationship between the product name features and the product categories, so that when the target product name features of the full name of the target product are input into the target classification model, the product category to which the target product belongs and its corresponding target classification code can be predicted. Therefore, the embodiments of the present application can quickly and accurately automatically determine the product classification code corresponding to the product through the product name, without configuring a large amount of human resources and waiting for a long time for the result, thereby improving the work process of product classification, sorting, or accounting, etc., and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0054] Figure 1 It is a schematic flowchart of a method for determining a product classification code in an embodiment of the present application;

[0055] Figure 2 It is another schematic flowchart of a method for determining a product classification code in an embodiment of the present application;

[0056] Figure 3 It is a schematic structural diagram of a system for determining a product classification code in an embodiment of the present application;

[0057] Figure 4 It is a schematic structural diagram of a device for determining a product classification code in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] For ease of explanation and understanding, the nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0059] (1) Product Lexical Database (which can be referred to as a corpus or dataset): A natural language text database for analysis and learning, containing language materials that have actually appeared in real use. These materials are usually organized with established formats and tags; Exemplarily, a dictionary file can be saved in the product lexical database.

[0060] (2) Text preprocessing: Refers to some processing carried out on text data before modeling, such as word segmentation, cleaning, normalization, and feature extraction.

[0061] (3) Word embedding: Generally speaking, it refers to converting a word into a vector representation, so word embedding is sometimes also called "word2vec"; It is a general term for language model and representation learning technologies in natural language processing. Conceptually, it means embedding a high-dimensional space with the dimension of the number of all words into a much lower-dimensional continuous vector space, and each non-phrase or phrase is mapped to a vector in the real number domain.

[0062] (4) LSTM network: Long Short-Term Memory network (LSTM) is a type of recurrent neural network, which is specifically designed to solve the long-term dependence problem existing in general recurrent neural networks (RNNs). All RNNs have a chain form of repeating neural network modules.

[0063] (5) Docker service: Docker is an open-source application container engine, open-sourced based on the Go language and compliant with the Apache 2.0 protocol.

[0064] (6) Tax classification code: Refers to the identity number of goods, commodities, taxable services, and service categories, generally determined by the State Taxation Administration.

[0065] The following introduces several traditional methods for obtaining the corresponding tax classification code according to the commodity name currently:

[0066] 1) Manual search method. First, screen out the keywords of the commodity for search. If it cannot be directly searched, first divide the industry into major categories according to the policy, and then when making a detailed sub-category division, if the classification cannot be clearly defined, for goods, the most similar code can be selected according to the material or use of the commodity, and for labor services or services, the most similar code can be selected according to the essence of the transaction, so as to finally determine the commodity name and tax rate according to the selected code. The limitation of the method based on manual search is that there will be problems such as huge manual workload, easy classification errors, and low efficiency.

[0067] 2) Semi-automatic tax coding method and system. First, extract a large number of commodity keywords, and then store the mapping relationship between keywords and tax classification codes in the database. When a new invoice needs to be issued, you only need to get the keywords of the commodity, and use the keywords to search the database to obtain the tax classification code and other information of the corresponding commodity. Among them, the first problem of the semi-automatic tax coding method and system is that it is necessary to manually screen the commodity keywords in advance, but now many commodities will add a large number of modifiers to the commodity names in order to increase the search volume and promotion. Therefore, the manual screening of keywords will consume a lot of manual work; the second problem is that the keyword data recorded in the database is limited, and there is no output result for some keywords that are not in the database, resulting in a low recognition rate; the third problem is that in most cases, the search of the database has no results, so it is necessary to use manual search methods to select the corresponding tax classification code, which cannot be fully automated.

[0068] The following will further explain this application in detail using an application example of how to obtain the tax classification code corresponding to a product based on the product name.

[0069] See also Figure 1 In a first aspect, the present application provides an embodiment of a method for determining a product classification code, comprising:

[0070] 101. Obtain a pre-trained target classification model.

[0071] Extending classification technology to text classification and recognition tasks is conducive to liberating manual classification labor and improving the output timeliness and accuracy of product classification results. The target classification model is obtained by machine learning training of the initial classification model with training samples. The target classification model stores the target correspondence between product names and product categories. The training samples include classified product name features, classification codes corresponding to classified product name features, and unclassified product name features. For example, in actual operation, the target correspondence can be specifically presented in the form of a probability formula to indicate the mapping relationship between product name features and the product category to which they belong, and then determine the product classification code corresponding to the product to be classified (target product). The product name features in the embodiment of the present application can be understood as data information used to characterize the content of the name text, which can affect the prediction accuracy of product categories and classification codes.

[0072] 102. Determine the full name characteristics of the target product based on the full name of the target product to be classified.

[0073] Since the target classification model stores the target correspondence between product names and product categories, in order to ensure that the target classification model can better utilize the input quantity and thus more optimally output the classification prediction result, it is necessary to adaptively first determine the corresponding target product full name feature based on the full name of the target product to be classified, so that this target product full name feature can be used as the input quantity to be input into the target classification model for classification prediction later.

[0074] 103. Input the target product full name feature into the target classification model to output the target product classification code.

[0075] Inputting the target product full name feature of the target product full name into the target classification model can predict the product category to which the target product belongs and its corresponding target classification code through the target correspondence therein, thereby realizing the fast and accurate automatic identification of the product classification code corresponding to the product from the product name, without the need to configure a large amount of human resources and long waiting time for results, and is conducive to the development of processes such as product classification and accounting, improving the user experience.

[0076] Please refer to Figure 2 , a second aspect of the present application provides an embodiment of a method for determining a product classification code, including: before putting the product name into the target classification model, rule matching processing can be pre-attempted to determine whether the target classification code can be obtained;

[0077] 200. Rule matching processing.

[0078] Judge whether the full name of the target product is recorded in the dictionary file, and the dictionary file pre-stores the full names of classified products and the classification codes correctly corresponding to the full names of classified products;

[0079] If the judgment result is yes, directly output the classification code corresponding to the record of the full name of the target product in the dictionary file as the target classification code, so as to quickly and accurately obtain the classification result; if the judgment result is no, determine the target classification code through the target classification model, that is, enter the model prediction processing stage (specifically, it can be divided into two model prediction processing stages: full-text classification type and partial classification type).

[0080] On the other hand, the benefit of rule matching processing is that in the face of situations that require quick response and adjustment of the model algorithm that cannot be correctly classified. For example, for a product name that has not appeared in the training data and word embedding training, it is difficult for the classification model to correctly identify its tax classification code. Therefore, in this case, it is necessary to (manually) add this data to the dictionary file so that this data can be added to the training data when the model is trained next time, thereby training the classification model iteratively.

[0081] 210. Train the initial classification model to obtain the target classification model.

[0082] If no classification result can be obtained in the rule matching and processing stage, the classification model can be used to determine the classification code of the target product. The process of machine learning training of the initial classification model with training samples specifically may include:

[0083] 2101. Obtain training samples (or construct training data).

[0084] Due to the rich variety of products and their names and the possibility of continuous new arrivals or updates, in the actual scenario, the number of products without a determined classification code (referred to as unlabeled products in the text, where the product name is known but the corresponding classification code is unknown) is much larger than that of products with a determined classification code (referred to as labeled products in the text). Therefore, it is necessary to use the data of unlabeled products, that is, the data with only the product name and the corresponding tax classification code, as part of the training samples for model training. Exemplarily, the data information of commodity invoices that have been opened in the database of the big data platform can be utilized, and two fields, namely the commodity name and the tax classification code, can be extracted from it. At the same time, the data with empty fields or obvious errors are excluded. Finally, 100,000 pieces of labeled product (classified product) data are screened out, 90,000 of which are used as training data and 10,000 are used as test data. The data is saved in the comma-separated value csv text format or the txt format. An example of the labeled product data is shown in the following table; in addition, 1 million pieces of unlabeled commodity (unclassified product) data are obtained. Here, the process of obtaining the product name can be regarded as the process of constructing the model corpus.

[0085] Serial number Commodity name Commodity tax classification code 1 Xiaochaihu Granules (Bud Culture) 1070304020000000000 2 USB White Mono Earphone Cable 1090519060000000000

[0086] In a specific embodiment, after obtaining the full names of the aforementioned classified products and unclassified products, the following feature extraction process needs to be performed on the product names that are pre-used as training sample data:

[0087] (1) Perform text preprocessing on the full names of the classified products and unclassified products to obtain the name word segmentation for composing each product full name. The text preprocessing includes text cleaning and word segmentation processing on the product full name.

[0088] The text cleaning process includes uniformly converting all the letters in the full names of classified products and unclassified products into either only uppercase or only lowercase letters, and / or removing duplicate words, parentheses, and the text content within the parentheses from the full product names; the purpose is that the content of the letters Ab and ab is essentially the same, so uniformly converting the case of the letters helps reduce data storage redundancy and resource occupancy. On the other hand, the tax code file table formulated by the tax bureau mainly corresponds to the product categories and their classification codes statistically obtained with high probability, and does not refine to minor difference categories such as product models or flavors. Moreover, the text content within the parentheses in the product name is mostly a modifying word used to distinguish the product batch number and does not affect the actual distinction of the product category attribution. For example: for the text data "Xiaojiawei Chewable Tablets (Sweet Lemon Flavor) (Gift)", first, the data after case conversion remains unchanged. Then, using regular matching to remove the parentheses and the content within the parentheses from the text, the text becomes "Xiaojiawei Chewable Tablets". After word segmentation processing, the text can finally be transformed into a list data structure form of "Xiaojiawei" and "Chewable Tablets".

[0089] The word segmentation process includes dividing the cleaned product name into at least one name segment. For example, word segmentation tools such as Jieba Segmentation or Chinese Lexical Analysis (Lac) can be used for word segmentation processing.

[0090] In practical applications, the text cleaning process may specifically further include removing function words (such as auxiliary words or conjunctions) in the full product name, which have little significance for the classification result, to avoid subsequent occupation of system resources to process these words and affecting the prediction duration.

[0091] The following steps (2) and (3) can be regarded as the operation processes included in the word embedding generation step.

[0092] (2) Determine the initial value of the position index represented in the form of a sparse vector in the product word library for each name segment.

[0093] In a specific embodiment, the initial value of the position index is a V-dimensional sparse vector in one-hot form (V represents the size of the product word library), that is, setting the index position of each segment in the product word library to 1 and setting other positions to 0. For example, the initial value of the position index of the nth segment in the vocabulary is x n =[1, 0, …, 0], X = {x1, …, x n} represents the set of index value vectors corresponding to the n text segments of a certain full product name.

[0094] (3) The position index initial value in the form of a sparse vector is converted into a position index mapping value in the form of a dense vector (300 - dimensional) through a word embedding generation model (specifically, the Embedding matrix from the word embedding generation model), and the position index mapping value is used as the product name feature in the training sample to train the initial classification model. The word embedding generation model can be specifically obtained by iteratively training the skip - gram model with multiple groups of position index initial values.

[0095] 2102. Use the training sample to perform machine learning training on the initial classification model to obtain a target classification model that stores the target correspondence.

[0096] The commodity tax classification code can essentially be regarded as a classification problem, that is, given the input features X = {x1, x2, …, x n}, model the conditional probability P(y|X) of the commodity tax classification code y. In a specific implementation, the process of training the initial classification model includes:

[0097] The aforementioned feature extraction process can be regarded as the training process of the representation layer.

[0098] In terms of the encoding layer, the bidirectional LSTM neural network algorithm can be adopted. The forward network The backward network where t represents the time step, v t is the word vector output by the representation layer, h t-1 is the output of the network unit at time step t - 1, and finally, the information of the context and the above - context at each time step is encoded to obtain the output For example, the BERT model algorithm can also be used.

[0099] In terms of the output layer, first perform max - pooling on the output of the encoding layer to obtain more useful encoding features, then perform a fully - connected mapping, and finally use Softmax normalization to obtain the probability that each product name feature belongs to each product category. The Softmax formula is:

[0100] In the formula, θ is the model parameter, K is the number of categories, and d j is the output of the fully - connected layer.

[0101] Among them, the model parameters of the initial classification model (such as the above-mentioned θ) are determined by the loss function calculated using a semi-supervised learning algorithm; semi-supervised learning is a learning method that combines supervised learning and unsupervised learning, which uses both labeled data and a large amount of unlabeled data for pattern recognition work. The supervised training algorithms (such as maximum likelihood and adversarial training algorithms) use labeled data, and the unsupervised training algorithms (minimum entropy and virtual adversarial training) use unlabeled data. Since unlabeled data is easy to obtain and in large quantities, the semi-supervised training algorithm uses these unlabeled data to learn the implicit expressions between data, which can make the classification model have better robustness and generalization in text classification tasks. Therefore, the final hybrid objective loss function used for model parameter tuning can be generated by mixing the loss function of any of the following supervised training algorithms with the loss function of any unsupervised training algorithm:

[0102] 1. Maximum likelihood. It is the most commonly used method for learning model parameters in text classification. This function is a convex function, and it is easy to find the optimal solution of the parameters. The specific algorithm loss function is as follows:

[0103]

[0104] Among them, m l is the number of labeled data, θ is the function parameter, y (i) is the predicted label of the i-th sample, and Ⅱ(;) is an indicator function.

[0105] 2. Adversarial training. This algorithm will mix some small perturbations into the sample input, and then make the model adapt to this change, so as to be robust to adversarial samples and improve the generalization ability of the model. The specific algorithm loss function is as follows:

[0106]

[0107] Among them, v *(i) is the input after adding perturbations, is the algorithm parameter, and the other parameters are the same as those of the maximum likelihood function.

[0108] 3. Minimum entropy. Information entropy is used to measure the certainty of an event occurring. The smaller the information entropy, the greater the probability of the event occurring, and the larger the information entropy, the smaller the probability of the event occurring. Here, the conditional entropy of class probabilities is used as the model loss. The smaller the conditional entropy, the more stable the model and the better the effect. Entropy minimization loss is applied to labeled and unlabeled data in an unsupervised manner. The specific algorithm loss function is as follows:

[0109]

[0110] Among them, m = m l + m u , m lis the labeled data quantity, m u is the unlabeled data quantity, and other parameters are the same as those of the maximum likelihood function.

[0111] 4. Virtual adversarial training. Virtual adversarial training can extend the supervised learning algorithm to the semi-supervised environment. This algorithm can be applied to the training of unlabeled data. By making a small perturbation to the input text representation vector, the Kullback-Leibler Divergence (KL divergence) of the class probabilities of the model output before and after the perturbation is calculated.

[0112]

[0113] where v *(i) is the input after perturbation, and other parameters are the same as those of the maximum likelihood function.

[0114] Exemplarily, the mixed objective loss function (produced by mixing the above four learning algorithms) finally used for model tuning is as follows:

[0115] L MIXED = λ ML L ML + λ AT L AT + λ EM L EM + λ VAT L VAT ,

[0116] where λ ML , λ AT , λ EM , λ VAT are the hyperparameters of the mixed objective loss function.

[0117] 220. Obtain a pre-trained target classification model.

[0118] 230. Determine the target product full name feature according to the full name of the target product to be classified.

[0119] Step 230 specifically includes: obtaining the full name of the target product, and performing text preprocessing on the full name of the target product including text cleaning and word segmentation to obtain the target name word segmentation for composing the full name of the target product; according to the target name word segmentation (specifically, it can be the noun part of speech), converting it into a target product full name feature with the same vector dimension as the classified product name feature, such as a name feature mapped into a 300-dimensional dense vector form. The specific implementation process of step 230 is similar to the feature extraction process in step 2101, and will not be elaborated here.

[0120] In practical applications, the target name word segmentation can be presented in the form of a single word or in the form of a sentence spliced by multiple individual words (also referred to as the target name spliced words in the text, such as spliced with "\t"). Exemplarily, for the full name of a commodity "Apple mobile phone", it can be first divided into two individual noun-type target name word segmentations of "Apple" and "mobile phone", and then the feature representation and encoding conversion are performed to obtain the product name features corresponding to them respectively. Finally, the classification code abbreviations (classification labels) they represent are predicted by the model as fruits and mobile communication devices. Correspondingly, the model prediction performed in this case can be called local classification prediction; or, it can be divided into the target name word segmentation of "Apple\tmobile phone" (in the form of a spliced sentence) to obtain its corresponding product name feature, and finally input into the classification model for prediction. Correspondingly, the model prediction performed in this case can be called full-text classification prediction. On the other hand, in terms of practical experience, in the case of local classification, the model labels (classification results) corresponding to the multiple individual words obtained by splitting the full name of a certain product are likely to have large type differences. For example, the above-mentioned fruits and mobile communication devices have essential differences in the functional categories, or the combination order of multiple words before and after may be different, which will affect the prediction effect of the model; while in the case of full-text classification, the target name spliced words obtained by dividing the full name of the same product (the part-of-speech of the word segmentation included is not limited to the noun part-of-speech) are in the form of a spliced sentence, so their vocabulary and vocabulary meaning are closer to the original meaning of the full name of the product, and thus the model classification result obtained is more in line with the correct label of the product, such as the above-mentioned mobile communication device; therefore, in practical applications, the priority of the full-text classification model prediction is higher than that of the local classification model prediction, which is equivalent to giving priority to considering the prediction result of the full-text classification situation. When the full-text prediction result does not meet the preset conditions (such as the probability result does not reach the threshold), then consider using the local classification prediction and its prediction result. It can be seen that the classification model of the embodiments of the present application can parse the text at the semantic level and at the same time provide users with better and faster recognition effects, thereby reducing the manual operation cost and improving user satisfaction.

[0121] It should be noted that in practical applications, a single non-noun part-of-speech word such as a verb or an adjective has little significance for identifying the product category. Therefore, to avoid occupying system resources and delaying the response time, in the case of local classification, only the product name word segmentation of the noun part-of-speech should be subjected to feature extraction processing to obtain the product name feature input to the classification model, so as to efficiently output the local classification prediction result.

[0122] 240. Input the full name feature of the target product into the target classification model to output the target product classification code.

[0123] Input the full name feature of the target product with the same vector dimension as the classified product name feature (in the form of a 300-dimensional dense vector) into the target classification model to determine the target classification code.

[0124] The process of outputting the classification code of the target product can specifically include:

[0125] The full-text classification processing stage includes: when the full name feature of the target product input to the target classification model is obtained by concatenating target name concatenation words, determining whether the maximum probability value among the K full-text probability values (top K classification label results) output by the target correspondence satisfies the full-text classification probability threshold, where K is a non-zero integer set by the user, and the target name concatenation words are obtained by concatenating multiple target name segmentation words used to form the full name of the target product (specifically, it can be concatenated with the tab character "\t"), and each full-text probability value corresponds to a product category; in practical applications, providing the aforementioned top K classification label results is beneficial to improving the recall rate.

[0126] If the judgment result is satisfied, determine the product category corresponding to the maximum probability value as the target product category to which the target product belongs, and output the classification code corresponding to the target product category as the target classification code;

[0127] The local classification processing stage includes: if the judgment result is not satisfied, respectively input the full name features of the target product obtained from the segmentation words with the part-of-speech of nouns in each target name segmentation word into the target classification model to obtain the local probability values of the target product belonging to different product categories;

[0128] Weightedly fuse the full-text probability values and local probability values corresponding to the same product category to determine the classification code matching the product category corresponding to the maximum fusion result as the target classification code.

[0129] In a specific implementation, weightedly fusing the full-text probability values and local probability values corresponding to the same product category (the same label) includes:

[0130] Sum the full-text probability values and local probability values corresponding to the same product category to obtain the probability subtotal value (which can be represented by p) for the current product category; calculate the maximum difference corresponding to each product category according to each probability subtotal value, and sum all the full-text probability values and local probability values to obtain the total probability value (which can be represented by p 总 ); calculate the fusion result of each product category according to the probability subtotal value, the maximum difference, and the total probability value; filter the fusion results that meet the fusion result conditions according to the numerical magnitudes of the fusion results, where each probability subtotal value and / or each fusion result can be sorted by magnitude respectively. Exemplarily, the formula for the maximum difference of the same product category is as follows:

[0131] round represents the "rounding function", Indicates rounding down a numerical value (floor); fusion result = p / p 总 + sub, where the purpose of sub is to ensure that the final fusion result is less than 1.

[0132] In a specific embodiment, the formula for weighted fusion of the full-text probability value and the local probability value corresponding to the same product category can also be: full-text probability value of category a × n% + local probability value of category a × (1 - n%) = fusion result for the same product category a, where n% is the weight coefficient.

[0133] In practical applications, after model training and code development are completed, for the convenience of deployment, the entire system service can be packaged into a Docker container, so that this service can be quickly and conveniently deployed on a machine with a Docker environment.

[0134] Steps 220 to 240 are similar to steps 101 to 103 and will not be elaborated here; in practical applications, the order of execution of steps 200 and 210 is not limited, that is, as long as it is ensured that the model has been trained when the classification model is enabled for prediction.

[0135] In summary, the embodiment of the present application proposes the application of the semi-supervised text classification algorithm based on the LSTM network in the commodity tax classification code. This method first uses a large amount of unlabeled data to obtain the implicit expressions between data, and then uses labeled data for adversarial training to improve the robustness and generalization ability of the model, so that the model can obtain better classification results; at the same time, this model uses dense vectors with semantic information to represent commodity texts, which can further improve the prediction effect of the model. In terms of the performance of the entire system, the time for a single call is stably maintained at about dozens of milliseconds. Finally, the method of this system can be deployed using Docker services to achieve the effect of quick deployment and use. Therefore, this technical solution can comprehensively improve the recognition effect of tax classification codes.

[0136] Please refer to Figure 3 , an embodiment of a product classification code determination system provided in the second aspect of the present application includes:

[0137] An acquisition unit 301, configured to acquire a pre-trained target classification model, where the target classification model is obtained by performing machine learning training on an initial classification model using training samples. The target classification model stores the target correspondence between product names and product categories, and the training samples include classified product name features, classification codes corresponding to the classified product name features, and unclassified product name features;

[0138] A processing unit 302, configured to determine the target product full name feature according to the target product full name to be classified;

[0139] The processing unit 302 is further configured to input the full name feature of the target product into the target classification model, so as to output a target classification code determined by the target classification model according to the target correspondence relationship, and the target classification code is used to represent the product category to which the target product belongs.

[0140] In the embodiments of the present application, the operations performed by each unit of the product classification code determination system are similar to those described in the foregoing first aspect or any specific method embodiment of the first aspect, and will not be elaborated herein specifically.

[0141] Please refer to Figure 4 , the product classification code determination device 400 of the embodiments of the present application may include one or more central processing units CPU (CPU, central processing units) 401 and a memory 405, and one or more application programs or data are stored in the memory 405.

[0142] Among them, the memory 405 may be volatile storage or persistent storage. The program stored in the memory 405 may include one or more modules, and each module may include a series of instruction operations in the product classification code determination device. Further, the central processing unit 401 may be configured to communicate with the memory 405 and execute a series of instruction operations in the memory 405 on the product classification code determination device 400.

[0143] The product classification code determination device 400 may further include one or more power supplies 402, one or more wired or wireless network interfaces 403, one or more input / output interfaces 404, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0144] The central processing unit 401 may perform the operations performed in the foregoing first aspect or any specific method embodiment of the first aspect, which will not be elaborated herein specifically.

[0145] It can be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the steps do not mean the order of execution. The execution order of each step should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0146] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0147] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system or device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0148] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0149] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0150] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product (computer program product) is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a business server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, and other various media that can store program codes.

Claims

1. A method for determining product classification codes, characterized in that, Including: Obtain a pre-trained target classification model, which is obtained by performing machine learning training on an initial classification model using training samples. The training samples include classified product name features, classification codes corresponding to the classified product names, and unclassified product name features. The target classification model stores the target correspondence between product name features and product categories; Determine the target product full name feature according to the full name of the target product to be classified; Input the target product full name feature into the target classification model to output the target classification code determined by the target classification model based on the target correspondence. The target classification code is used to represent the product category to which the target product belongs; The output of the target classification code determined by the target classification model based on the target correspondence includes: when the target product full name feature input into the target classification model is obtained by concatenating target name splicing words, determine whether the maximum probability value among the K full-text probability values output by the target correspondence meets the full-text classification probability threshold, where K is a non-zero integer set by the user. The target name splicing words are obtained by splicing multiple target name tokens used to form the full name of the target product, and each full-text probability value corresponds to a product category; if the judgment result is satisfied, determine the product category corresponding to the maximum probability value as the target product category to which the target product belongs, and output the classification code corresponding to the target product category as the target classification code.

2. The method according to claim 1, wherein Before obtaining the pre-trained target classification model, the method further includes: Obtain training samples, which include classified product name features, classification codes corresponding to the classified product name features, and unclassified product name features; Use the training samples to perform machine learning training on the initial classification model to obtain a target classification model that stores the target correspondence, where the model parameters of the initial classification model are determined by using a loss function defined by a semi-supervised learning algorithm.

3. The method according to claim 2, wherein The obtaining of the training samples includes: Obtain classified product full names and unclassified product full names, and perform text preprocessing on the classified product full names and the unclassified product full names to obtain name tokens used to form each product full name; Determine the initial value of the position index of each name token represented in the form of a sparse vector in the product word library; Convert the initial value of the position index in the form of a sparse vector into a position index mapping value in the form of a dense vector through a word embedding generation model. The position index mapping value is used as the product name feature in the training sample to train the initial classification model.

4. The method according to claim 3, wherein The text preprocessing includes text cleaning and word segmentation processing on the product full name; The text cleaning includes: uniformly converting the letters in the classified product full name and the unclassified product full name into letters in only uppercase or only lowercase form, and / or removing duplicate words, parentheses, and the text content in the parentheses in the product full name; The word segmentation processing includes dividing the cleaned product name into at least one name token.

5. The method according to claim 1, wherein The determining of the target product full name feature according to the full name of the target product to be classified includes: Obtain the full name of the target product, and perform text preprocessing on the full name of the target product, including text cleaning and word segmentation, to obtain target name word segments for composing the full name of the target product; According to the target name word segments, convert to obtain target product full name features with the same vector dimension as the classified product name features; The step of inputting the target product full name features into the target classification model includes: Input the target product full name features with the same vector dimension as the classified product name features into the target classification model to determine the target classification code.

6. The method according to claim 1, wherein Output the target classification code determined by the target classification model according to the target correspondence, including: If the judgment result is not satisfied, input the target product full name features obtained from the word segments with the part-of-speech of nouns in each target name word segment into the target classification model respectively to obtain local probability values of the target product belonging to different product categories; Perform weighted fusion on the full-text probability values and local probability values corresponding to the same product category to determine the classification code matching the product category corresponding to the maximum fusion result as the target classification code.

7. The method according to claim 1, characterized in that, Before obtaining the pre-trained target classification model, the method further includes: Judge whether the full name of the target product is recorded in the dictionary file, which pre-saves the full names of classified products and the classification codes correctly corresponding to the full names of classified products; If the judgment result is yes, directly output the classification code recorded corresponding to the full name of the target product in the dictionary file as the target classification code; If the judgment result is no, determine the target classification code through the target classification model.

8. A product classification coding determination system, characterized in that, Include: An acquisition unit for acquiring a pre-trained target classification model, which is obtained by performing machine learning training on an initial classification model with training samples. The target classification model stores a target correspondence between product names and product categories. The training samples include classified product name features, classification codes corresponding to the classified product name features, and unclassified product name features; A processing unit for determining target product full name features according to the full name of the target product to be classified; The processing unit is further configured to input the target product full name features into the target classification model to output the target classification code determined by the target classification model according to the target correspondence. The target classification code is used to represent the product category to which the target product belongs; Specifically, when the target product full name features input into the target classification model are obtained from target name concatenation words, the processing unit judges whether the maximum probability value among the K full-text probability values output by the target correspondence satisfies the full-text classification probability threshold, where K is a non-zero integer set by the user. The target name concatenation words are obtained by concatenating multiple target name word segments for composing the full name of the target product, and each full-text probability value corresponds to a product category; if the judgment result is satisfied, determine the product category corresponding to the maximum probability value as the target product category to which the target product belongs, and output the classification code corresponding to the target product category as the target classification code.

9. A device for determining product classification codes, characterized in that, Include: A central processing unit, a memory, and an input / output interface; The memory is a transient memory or a persistent memory; The central processing unit is configured to communicate with the memory and execute instruction operations in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Comprising instructions which, when run on a computer, cause the computer to execute the method according to any one of claims 1 to 7.

11. A computer program product, characterized in that, When the computer program product runs on a computer, it causes the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Invoice commodity name classification method and device

    CN114219038A