Data processing method, commodity classification method, system and device, storage medium and program product

By extracting and analyzing the key element information of the product, combining object pictures and encoding description information, the problem of low product classification accuracy in the prior art is solved, and higher classification accuracy and consistency are achieved.

CN120067908APending Publication Date: 2025-05-30ZHEJIANG WIDEWAY DIGITAL TECH SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510090313.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When using HSCode to classify products, there is a problem of low classification accuracy, especially when identifying subcategory classification and processing multiple product names, it is easy to lead to manual dismantling and misleading, affecting subsequent tax rates and export management.

Method used

By extracting the key element information of the target object, determining the classification chapter code, and selecting the adapted name value based on the object image to determine the candidate classification code. For missing key elements, the encoding description information of candidate classification codes is used for reasoning and supplementation, and the confidence is updated to determine the most accurate classification code.

Benefits of technology

It improves the accuracy of product classification, reduces manual intervention, enhances the ability to handle multi-product names, and ensures the accuracy and consistency of classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067908A_ABST
    Figure CN120067908A_ABST
Patent Text Reader

Abstract

According to the data processing scheme provided by the embodiment of the invention, key elements required by classification are extracted based on object information (such as text description information and object pictures) of a target object, and after key element information is obtained, classification chapter codes corresponding to the target object are determined according to the key element information; and determining the classification of the target object according to the classification chapter code and the key element information. According to the scheme, the object classification implementation mode of firstly determining the classification chapter code is adopted, so that the classification code matching range can be greatly reduced in the subsequent classification step of determining the target object, and the classification matching precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computing technologies, and in particular, to a data processing method, a commodity classification method, a system, a device, a storage medium, and a program product. Background Art

[0002] The commodity-related public service platform provides a pre-classification solution for the global HSCode (Harmonized System Code, also known as the tariff number or tax number) of commodities. However, in actual applications, the classification accuracy of HSCode is closely related to various factors. Once the classification is incorrect, it will affect the subsequent commodity tax rate, export (for example, it may be misidentified as a smuggled commodity when exporting), etc. Currently, for the pre-classification solution provided by the commodity-related public service platform, when using models such as named entity recognition (NER) models to identify the elements required for commodity classification, only some major classifications can be identified, and other minor classifications require manual tariff disassembling. Moreover, in the case of identifying multiple different element values belonging to the same type, it is also easy to affect the commodity classification accuracy. In addition, when performing similarity matching, the cosine similarity matching algorithm is often used, and due to some defects of the cosine similarity, there will also be a problem of low classification accuracy. Summary of the Invention

[0003] Multiple aspects of the present application provide a data processing method, a commodity classification method, a device, a storage medium, and a program product to improve the accuracy of commodity classification.

[0004] In a first aspect, the present application provides a data processing method. The method includes:

[0005] Extracting key elements required for classifying the target object based on the object information of the target object to obtain key element information;

[0006] Determining the classification chapter code corresponding to the target object based on the key element information;

[0007] Determining the classification to which the target object belongs according to the classification chapter code and the key element information.

[0008] In a second aspect, the present application further provides a data processing method. The method includes:

[0009] Obtaining the object information of the target object, where the object information includes the text description information and the object picture of the target object;

[0010] Extracting key elements required for classifying the target object based on the text description information to obtain key element information;

[0011] If there are multiple name values of the target object in the key element information, based on the object picture, select an appropriate name value from the multiple name values;

[0012] Based on the selected name value, determine multiple candidate classification codes for the target object;

[0013] Based on one of the multiple candidate classification codes, determine the classification to which the target object belongs.

[0014] In a third aspect, the present application also provides a data processing method. The method includes:

[0015] Obtain the key element information required for classifying the target object;

[0016] Based on the key element information, determine multiple candidate classification codes for the target object and the confidence level of each candidate classification code;

[0017] When the key element information meets the missing key element trigger condition, according to the code description information of the multiple candidate classification codes, determine the type of the missing key element in the key element information; where, meeting the missing key element trigger condition includes: determining that there is a missing key element in the key element information for at least one of the multiple candidate classification codes;

[0018] Based on the type of the missing key element, the code description of the at least one candidate classification code, and the text description information of the target object, determine the key element value corresponding to the type of the missing key element and the matching degree corresponding to the key element value for each candidate classification code in the at least one candidate classification code;

[0019] Based on the matching degree, update the confidence level of the at least one candidate classification code;

[0020] After the update, based on the confidence levels of the multiple candidate classification codes, select one classification code from the multiple candidate classification codes to be used to determine the classification to which the target object belongs.

[0021] In a fourth aspect, the present application provides a commodity classification method. The method includes:

[0022] Obtain the commodity information of the target commodity;

[0023] Based on the commodity information, extract the key elements required for classifying the target commodity to obtain key element information;

[0024] Based on the key element information, determine the classification chapter code corresponding to the target commodity;

[0025] Determine the classification to which the target commodity belongs according to the classification chapter code and the key element information.

[0026] In a fifth aspect, the present application provides a data processing system. The system includes:

[0027] A server for implementing the steps in each of the method embodiments provided in the present application above;

[0028] A client for displaying at least one of the following information: object information of a target object and classification information to which the target object belongs; the target object includes a target commodity, and the classification information includes a classification code to which the target object belongs and / or a confidence level of the classification code.

[0029] In a sixth aspect, the present application provides an electronic device. The electronic device includes: a memory and a processor, wherein the memory is used for storing a program; the processor is coupled to the memory and is used for executing the program stored in the memory to implement the method described in any one of the above.

[0030] In a seventh aspect, the present application provides a computer-readable storage medium storing a computer program, and the computer program can implement the method described in any one of the above when executed by a computer.

[0031] In an eighth aspect, the present application provides a computer program product including a computer program, and the computer program implements the method described in any one of the above when executed by a processor.

[0032] In the technical solution provided by the embodiments of the present application, when extracting the key elements required for classification based on the object information of the target object (such as including text description information, object pictures), after obtaining the key element information, the classification chapter code corresponding to the target object will be determined first according to the key element information, and then the classification to which the target object belongs will be determined according to the classification chapter code and the key element information. This solution adopts this implementation method of object classification that determines the classification chapter code first, so that the classification code matching range can be greatly reduced in the subsequent steps of determining the classification to which the target object belongs, and the classification matching accuracy can be improved. In addition, when multiple name values (product names) of the target object are included in the extracted key element information, an appropriate name value will be selected from the multiple name values based on the object picture of the target object, and then multiple candidate classification codes will be determined for the target object based on the selected name value. Based on one of the multiple candidate classification codes, the classification to which the target object belongs is determined. This solution integrates the object picture of the target object to denoise the multiple name values, so as to select an appropriate name value for the subsequent determination of the classification code of the target object, which is conducive to improving the accuracy of the target object classification. Further, when at least one of the multiple candidate classification codes is determined to have a missing key element in the key element information, the type of the missing key element can also be determined based on the code description of the multiple candidate classification codes. This method of supplementing the missing elements by using the multiple candidate classification codes as the judgment interval for the missing key elements can avoid the situation that the missing key elements corresponding to the correct classification code may not be included when using at least one candidate classification code as the judgment interval; and after determining the type of the missing key element, according to the type of the missing key element, the code description information of at least one candidate classification code, and the text description information of the target object (such as the object title), the key element value corresponding to the type of the missing key element and the matching degree corresponding to the key element value are determined for each candidate classification code in at least one candidate classification code, and then the confidence of at least one candidate classification code is updated based on the matching degree. After the update, based on the confidence of the multiple candidate classification codes, a classification code is selected from the multiple candidate classification codes to be used to determine the classification to which the target object belongs, so that the certainty and accuracy of the target object classification result can be effectively improved again. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0034] Figure 1 It is a schematic flow chart of a data processing method provided by an exemplary embodiment of the present application;

[0035] Figure 2 Schematic diagram of the principle for determining the classification chapter code (such as the chapter in the HS code) provided by an exemplary embodiment of the present application;

[0036] Figure 3a Schematic diagram of the principle for the classification code matching provided by an exemplary embodiment of the present application;

[0037] Figure 3b Schematic diagram of the structure of the noise reduction model provided by an exemplary embodiment of the present application;

[0038] Figure 3c Schematic diagram of the structure of the similarity analysis model provided by an exemplary embodiment of the present application;

[0039] Figure 4a Schematic diagram of the structure of the missing key element reasoning model provided by an exemplary embodiment of the present application;

[0040] Figure 4b Schematic diagram of the process of missing key element reasoning provided by an exemplary embodiment of the present application;

[0041] Figure 5 Schematic diagram of the overall process of target object classification provided by an exemplary embodiment of the present application;

[0042] Figure 6 For an exemplary embodiment of the present application, the use of Figure 5 Example of implementing the steps shown to determine the classification code of a target object (a commodity);

[0043] Figure 7 Schematic diagram of a commodity picture provided by an exemplary embodiment of the present application;

[0044] Figure 8 Schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0045] The commodity-related public service platform provides a pre-classification scheme for global HS Codes of commodities. However, in actual application, the classification accuracy of HS Codes is closely related to various factors. Once the classification is incorrect, it will affect the subsequent commodity tax rate, export (such as being misidentified as smuggled goods during export), etc. Although the classification accuracy can be improved to a certain extent through clear commodity descriptions, understanding of relevant regulations, and professional knowledge, there are still some problems with the pre-classification scheme provided by the current commodity-related public service platform.

[0046] For example, the specific implementation of the global HS Code pre-classification solution provided by the current commodity-related public service platform includes the following steps:

[0047] Stp1. Customize the NER model in the field of tariff rules to identify the elements and corresponding values required for classification from commodity information;

[0048] Stp2. Based on the classification elements and values obtained from the structured decomposition of the globally common 6-digit tariff rules, perform similarity matching with the elements and corresponding values extracted from the commodity information, and output the globally common 6-digit HS Code.

[0049] Stp3. Based on the classification elements and values obtained from the structured decomposition of the specific country's customs tariff rules, perform similarity matching with the elements and corresponding values extracted from the commodity information, and output the customs HS Code of the specific country.

[0050] Among them, the algorithm used for the above similarity matching is the common cosine similarity matching algorithm.

[0051] In the actual application process of the above solution, it is found that there are some deficiencies, which are mainly manifested in the following aspects:

[0052] (1) By using the NER model, only 13 major categories of classification elements can be identified from the commodity information, while about 330 minor categories of classification elements need to be disassembled manually from the tariff rules. Therefore, it often leads to the situation of missing classification elements in the matching results output by the model, and further leads to the uncertainty of the output results.

[0053] (2) By using the NER model to identify the elements required for classification from the commodity information, there are usually multiple element values, especially in the case of multiple product names, which will mislead the classification results of the model and further make it impossible to classify to the accurate answer.

[0054] (3) Using the cosine similarity matching algorithm to perform the corresponding similarity matching process, due to some defects of the cosine similarity, it will affect the improvement of the classification accuracy, which is mainly manifested as follows:

[0055] (31) Insensitive to the grammar and semantics of the text: The cosine similarity only considers the occurrence of words, but does not consider the grammar and semantics of words, so some important information may be ignored.

[0056] (32) Since the cosine similarity only considers the number of times a word appears in the text, the cosine similarity cannot solve the problem of semantic similarity.

[0057] In response to the above problems, this application provides a data processing solution to improve the accuracy of commodity classification.

[0058] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of them. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0059] It should be noted that in the case where the embodiments of this application involve user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject. Additionally, various models (including but not limited to language models or large models) involved in this application comply with relevant laws and standards.

[0060] The following introduces and explains each embodiment provided by this application.

[0061] First, the vocabulary involved in the embodiments of this application will be explained. It can be understood that this explanation is for a clearer understanding of the embodiments of this application and does not necessarily constitute a limitation on the embodiments of this application.

[0062] The commodity-related public service platform is provided by the corresponding e-trade platform (such as the international e-trade platform). Among them, the e-trade platform is usually an initiative led by the private sector and jointly launched by all stakeholders, aiming to cultivate e-trade rules through public-private cooperation and create a more effective and efficient policy and business environment for the development of cross-border e-trade (eTrade). That is, the main initiative of the e-trade platform is to create a more free, innovative, and inclusive international trade environment by promoting public-private dialogue and sharing practical experiences. The commodity-related public service platform is one of the important pillars of the e-trade platform and is a government-enterprise cooperation service platform built to integrate government affairs and business capabilities. On the one hand, it provides a one-stop digital solution for small and medium-sized cross-border e-commerce enterprises, covering services such as customs clearance, foreign exchange settlement, and tax rebates for organizations, as well as capabilities such as transactions, finance, payment, logistics, and settlement required by cross-border trade enterprises. On the other hand, it helps organizations address the governance challenges brought about by the rapid rise of cross-border e-commerce, win-win and co-build with organizations, and serve global trade facilitation.

[0063] HSCODE (Harmonization System Code), known as the HS code, also called the tariff number or tax number, is short for the International Convention on the Harmonized Commodity Description and Coding System. It is an internationally recognized commodity coding and classification system, also known as the commodity code. The Harmonized System is developed by the World Customs Organization and is a system for quantifying the applicable / refundable tariff rates for various products during import and export. The HS code, which is the common identity document for imported and exported goods, is the basic element for customs and commodity import / export management agencies in various countries to confirm the commodity category, conduct commodity classification management, review tariff standards, and inspect commodity quality indicators. In e-commerce terms, the HSCODE is the category system for customs to manage commodities, and each commodity is classified under only one HS. Each country can define its own HS code, but the first six digits are internationally standardized, and the subsequent digits are defined according to the actual situation of each country. Thus, the HS code (tariff number), understandably, is the commodity code number listed in the Tariff, which is used to identify specific categories of commodities and helps determine the applicable tariff rates, VAT rates, and other possible regulatory requirements for the commodity during import and export, affecting aspects such as commodity cost calculation, customs declaration procedures, and trade statistics.

[0064] Specifically, the HS code structure usually consists of six digits, which divides all international trade commodities into 22 classes and 98 chapters, further divided into headings and subheadings. Among them, the first and second digits in the HS code (commodity code) represent "chapters", the third and fourth digits represent "headings" (or called titles), and the fifth and sixth digits represent "subheadings" (or called sub-titles). The first six digits are the HS international standard codes. There are 1,241 four-digit tariff items and 5,113 six-digit subheadings in HS. Thus, the HS code is mainly divided into the following three levels: the first two digits represent "chapters" (Section), the middle two digits represent "headings" (Heading), and the last two digits represent "subheadings" (Subheading). For example, the HS code 01.01.01.00 represents horses among live animals. Also, in the HS code, "classes" are basically divided by economic sectors; the "chapters" are classified in basically two ways: one is to classify by the attributes of the commodity raw materials, and products with the same raw materials are generally classified into the same chapter, and the products within the chapter are arranged in the order from raw materials to finished products according to the processing degree; or they are classified into different chapters according to their functions or uses, regardless of the raw materials used, and the headings or subheadings are arranged according to the raw materials or processing procedures within the chapter.

[0065] The World Customs Organization (WCO) is an independent body whose mission is to enhance the effectiveness and efficiency of customs administration. Today, the WCO represents 184 customs administrations globally, which together handle approximately 98% of world trade.

[0066] Classification code matching: For tariff code matching, it refers to the process of accurately classifying goods into the corresponding HS codes according to information such as their characteristics and uses.

[0067] Product Name (or Commodity Name) of a commodity: It refers to the name used to identify and describe a specific commodity or service. It is usually one of the main ways for consumers to identify products and is also important information in commercial transactions, marketing, customs declarations, and other links.

[0068] NER (Named Entity Recognition) model, an entity extraction model, is mainly used to process text data. Specifically, it is mainly used to automatically identify and classify entities with specific meanings from unstructured text. These entities can be names of people, places, organizations, dates, amounts, commodity names, titles, etc. Usually, when it comes to pictures, since pictures are not text-format data, directly applying traditional NER models cannot process pictures. However, in this case, by combining other image recognition technologies, the NER model can be indirectly involved in understanding the content of pictures to achieve an auxiliary NER task. For example, through OCR (Optical Character Recognition) technology, the text in the picture can be extracted, and then the extracted text can be input into the NER model to identify and classify the named entities therein.

[0069] Bert (Bidirectional Encoder Representations from Transformers) is a pre-trained model for natural language processing (NLP).

[0070] Key factor: It refers to the factor that plays a decisive role in a specific task.

[0071] Multimodal model refers to a model that can process multiple data types including images, text, etc. For example, the model type of the multimodal model can be, but is not limited to, DNN (Deep Neural Networks).

[0072] The following method embodiments provided in this application are applied to corresponding electronic devices with logical operation functions. The electronic device can be a client or a server. The client can be any terminal device such as a mobile phone, a tablet computer, a smart wearable device, etc. The server can be a common server, a cloud, a virtual server, etc. In specific implementation, it can be mainly applied to the server, and specifically, the server can refer to the server corresponding to the commodity-related service platform.

[0073] Figure 1The flowchart of a data processing method provided by an embodiment of the present application is shown. As Figure 1 shown, the data processing method includes the following steps:

[0074] 101. Extract the key elements required for classifying the target object based on the object information of the target object to obtain key element information;

[0075] 102. Determine the classification chapter code corresponding to the target object based on the key element information;

[0076] 103. Determine the category to which the target object belongs based on the classification chapter code and the key element information.

[0077] In this embodiment, the target object is a commodity to be traded and classified, such as various tradable commodities like clothing, animals, electronic devices, electronic equipment, food, dairy products, cookware, etc. And the object information of the target object can be, but is not limited to, multimodal information including text description information of the target object (such as object title, object attributes, object categories, etc.), object pictures of the target object (such as main pictures and / or sub-pictures), etc. Among them, the object information of the above target object can be uploaded by the user, and / or can also be automatically generated by the corresponding device end according to some information input by the user. For example, the object title can be automatically generated according to some keywords uploaded by the user for the target object, and no specific limitation is made here.

[0078] Taking the above object information as the input of the corresponding information extraction model can extract the key element information required for classifying the target object (which is text key element information).

[0079] Based on this, in an implementable technical solution, the above step 101 "Extract the key elements required for classifying the target object based on the object information of the target object to obtain key element information" may include:

[0080] 1011. Invoke the pre-trained information extraction model;

[0081] 1012. Input the object information into the information extraction model and output the key element information.

[0082] In the above 1011, the information extraction model can be, but is not limited to, an NER model. The NER model can be trained based on a Transformer binary classification model. The Transformer binary classification model refers to a binary classification model based on the Transformer structure (a deep learning architecture), such as the BERT model. Among them, the characteristics of the Transformer binary classification model include: self-attention mechanism, multi-head attention mechanism (convolution runs multiple attention functions in parallel), feed-forward neural network, etc. Figure 2 The structural example of the NER model is shown in Figure 2 . The NER model can specifically be, for example, trained based on the BERT model. BERT is the abbreviation of Bidirectional Encoder Representations from Transformers (bidirectional encoding representation), and it is a pre-trained language representation model.

[0083] In the above 1012, before inputting the object information of the target object into the information extraction model, preprocessing can be performed first to uniformly process it into text data as the model input.

[0084] For example, the object information contains text description information such as the object title, object detail description, object attributes, object categories, etc. of the target object. Continuing to refer to Figure 2 , for content such as the object title and object detail description that are originally in text format, they can be directly used, as shown by X1~Xn; for object attributes and other formatted key-value data (Key / Value data), they can be processed into text format data of "Key:Value", as shown by P1~Pn. Finally, by concatenating the respective text format data obtained from the above processing with the corresponding delimiters, the construction of the input data can be completed.

[0085] Among them, if the object information also contains the object picture of the target object, technologies such as OCR can be used to perform text recognition and extraction on the object picture. If text information is extracted from the object picture, the text extracted from the object picture information can also be processed into corresponding text format data as the input of the NER model.

[0086] Furthermore, the NER model extracts key elements from the input, and outputs the key element information required for classifying the target object. The key element information is in text form and includes one or more key elements. Each key element includes a key element type and at least one key element value corresponding to the key element type. That is, the information text format included in the key element information is "key element type: key element value", where the key element types include but are not limited to: the name of the target object (also known as the product name), material, use, category, specification, etc.

[0087] Exemplarily, using the NER model, the object information of the target object is identified, and the key element information required for classification extracted therefrom includes the following key elements:

[0088] Name (product name): bookmark, tassel pendant, pendant, tassel, jewelry;

[0089] Material: 100% polyester fiber;

[0090] Specification: 120 pieces.

[0091] It should be supplemented here that: as seen in Figure 7 Step 1 "Classification Key Element Identification (i.e., Key Element Identification and Extraction Required for Target Object Classification)" given in, in the process of the NER model extracting the key elements required for classifying the target object based on the object information of the input target object, specifically: it will first determine whether the name (product name) of the target object is included. If it is included, it will identify the specific name according to whether the score of the name (similarity score) is higher than the set threshold. For example, the name with a score higher than the threshold can be determined as the specific name of the target object. If the scores of all names are lower than the threshold and / or it is initially determined that the input object information does not contain a name, the NER model will infer the name. The inference methods can include but are not limited to: using the context clues of the input text data, or introducing relevant external knowledge bases. For example, the vocabulary and sentence structure in the context can be used to reasonably predict the name of the target object. Another example is that resources such as industry term dictionaries can also be input into the model as additional information to achieve the identification of potential names of the target object, and so on.

[0092] The key element information output by the above NER model (such as BERT) will be applied to downstream tasks (including tasks such as determining the classification code of the target object).

[0093] Specifically in implementation, the last few layers of the NER model (such as BERT) are used to make the output applied to downstream tasks. When the output is applied to downstream tasks, a fully connected compression matrix will be performed, and feature extraction will be carried out through convolutional layers using convolutional kernels of different window sizes.

[0094] Continuing to refer to Figure 2 , the key element information output by the NER model can be used as the input of a downstream deep neural network model (DNN), so as to predict the classification chapter code (the first two digits of the classification code) of the target object through the DNN model. Specifically, in the solution of this application, two convolutional models (text convolutional models) are used to determine the first 2-digit classification chapter code in the corresponding classification code (such as the HS code). Among them, these two convolutional models are two convolutional layers in the DNN, as shown in Figure 2 the first convolutional model (the first convolutional layer Layer1) and the second convolutional model (the second convolutional layer Layer2) shown in. And when determining the second digit in the classification chapter code, the feature information of the first digit is used, which refers to the calculation of fusing the information of the second digit itself through the forget gate in the LSTM (Long Short-Term Memory Network). For example, assuming that the output result of the first convolutional model is L1 and the output result of the second convolutional model is L2, the corresponding formula here is f = σ(w f [L1, L2])·L1 + L2, where σ represents the activation function (such as the sigmoid function), w f represents the weight matrix, and [L1, L2] represents a vector matrix formed by connecting and combining L1 and L2.

[0095] Based on the above combination Figure 2 described content, in an implementable technical solution, the above step 102 "determine the classification chapter code corresponding to the target object based on the key element information" may include:

[0096] 1021. Output the key element information to the first convolutional model and the second convolutional model respectively, and obtain the first output result of the first convolutional model and the second output result of the second convolutional model;

[0097] 1022. Determine the first digit in the classification chapter code based on the first output result;

[0098] 1023. Determine the second digit in the classification chapter code based on the first output result and the second output result.

[0099] Further, the implementation of the above step 1023 may include the following steps:

[0100] 10231. Perform activation processing on a vector formed by combining the first output result and the second output result to obtain an activation processing result;

[0101] 10232. Determine the second digit according to the product value of the activation processing result and the first output result and the second output result.

[0102] Here, the first output result is the aforementioned L1, and the second output result is the aforementioned L2. The first output result can be directly determined as the first digit. For the specific determination implementation of the second digit, reference can be made to the formula f given above.

[0103] In the solution of this application, by using the key element information such as the name (product name), use, and material of the target object, through combining the semantic understanding of a large model (such as Figure 2 the DNN model shown in that contains two convolutional layers), classification precedent data, etc., to comprehensively determine the classification chapter code in the internationally common classification code (tariff number) corresponding to the target object, so as to greatly narrow the classification code matching range and improve the matching accuracy in subsequent steps.

[0104] For example, after determining the classification chapter code of the target object, further, based on the extracted key element information, the item code (leaf category code) in the classification code will be identified under this classification chapter code, so as to achieve the recognition of not only the first two digits in the classification code of the target object, but even further recognition of the first four digits or the first six digits, etc., so as to recall several candidate classification codes for the target object.

[0105] Among them, in the process of item code recognition to determine the internationally common first six - digit classification code for classification, the following solution is adopted to improve the classification accuracy and certainty of the target object: denoise the key element information, and the purpose of denoising is to perform outlier detection on various key element values included therein to remove outliers. Specifically,

[0106] 1) If there are multiple name values (product name values) in the key element information identified for the target object, then in this embodiment, the main name value will be selected through a multi - modal model that fuses image information, so that the multi - modal model selects a more suitable name among the multiple name values extracted, and then performs similarity matching to recall several candidate classification codes (tariff numbers).

[0107] 2) For other element information (such as material, use, etc.) in the key element information except the name element, noise will also be removed to purify other information. After purification, for some vertical fields and proper nouns, a key element similarity scoring model in the vertical field of this task obtained by training a Transformer binary classification model will be used for similarity matching analysis, so as to improve the matching accuracy and further improve the classification accuracy.

[0108] Based on this content, in an implementation technical solution, the above - mentioned 103 "determine the classification to which the target object belongs based on the classification chapter code and the key element information" may include:

[0109] 1031. Perform noise reduction processing on the key element information to obtain the key element information after noise reduction;

[0110] 1032. Based on the classification chapter code and the key element information after noise reduction, determine multiple candidate classification codes and use a similarity analysis model to determine the confidence of each candidate classification code; the similarity analysis model is trained based on a binary classification model relying on the self-attention mechanism;

[0111] 1033. Based on the confidence of each candidate classification code, determine a classification code from the multiple candidate classification codes for determining the classification to which the target object belongs.

[0112] In the above 1031, it can be understood that the noise reduction processing refers to performing outlier detection to remove outliers in the information, and the outliers are the noise. Specifically, in implementation, this noise reduction processing is implemented using a corresponding pre-trained noise reduction model. Figure 3b The structural example diagram of the noise reduction model is shown in. The noise reduction model uses an algorithm model including BERT (for the Transformer architecture), DNN (deep neural network), and LOF (Local Outlier Factor). Among them, LOF is an unsupervised machine learning algorithm for identifying outliers in a dataset.

[0113] Among them, when there are multiple name values (multiple product name values) in the key element information, when performing noise reduction processing on the multiple name values, a more suitable name selection is performed through a multi-modal model that fuses image information (such as Figure 3b the DNN model included in the noise reduction model shown in). For example, referring to Figure 3b , the multiple name values of the target object and the object picture will be input as prompt words into the multi-modal model (such as DNN) in the noise reduction model. Then, the multi-modal model will output a more suitable name value, and the output name value is selected from the multiple name values based on the understanding of the object picture; among them, before inputting the object picture and multiple name values into the multi-modal model, it will first be processed by the Transformer architecture in the noise reduction model, and the output of the Transformer architecture is used as the input of the multi-modal model (such as DNN).

[0114] With the above content, when there are multiple name values of the target object in the key element information, "performing noise reduction processing on the key element information" in this step 1031 may include:

[0115] 10311. Obtain the object picture of the target object from the object information of the target object;

[0116] 10312. Identify the object picture of the target object, and select a suitable name value from the multiple name values based on the recognition result.

[0117] In the above 1032, considering that when using a general coding model and cosine similarity for similarity matching analysis, there are often cases where the intended meanings are similar but the distances in the embedding space are far. Therefore, the similarity analysis model in this step is specifically a key element similarity scoring model trained by a binary classification model with a Transformer structure. Figure 3c The structural schematic diagram of the similarity analysis model is shown.

[0118] As shown in Figure 3a , first, based on the name (specifically the name value) determined after noise reduction, similarity judgments can be made on each classification code related to the classification chapter code to recall a set number of candidate classification codes with a high similarity ranking, such as recalling the top 5 candidate classification codes with a high ranking (denoted as the Top5 candidate classification codes and related to the previously determined classification chapter code). Then, using the similarity analysis model, each key element (material, product name, use, etc.) is respectively subjected to similarity matching analysis with the code description information of each recalled candidate classification code to obtain the confidence levels of each recalled candidate classification code.

[0119] For example, based on the name determined after noise reduction, similarity analysis and judgment can be carried out with the description information corresponding to each sub - code under the classification chapter code, so as to recall, for example, the top 5 sub - codes with a high similarity ranking from them, and thus the top 5 candidate classification codes can be determined. A candidate classification code includes a classification chapter code and one of the top 5 sub - codes. Suppose a recalled candidate classification code is 180610, where the chapter code 18 represents "Cocoa products; chocolate and other cocoa - containing foods", the first - level sub - code 06 represents chocolate and chocolate confectionery products, the second - level sub - code 10 represents solid or filled chocolate containing cocoa powder, and the code description information corresponding to this classification code 180610 (which can also be understood as the description information of the sub - code) is: product name (name) "chocolate candy", production materials (also production materials) "cocoa powder, cocoa butter, milk, sugar", use "for sale in the retail market, and can also be given as a snack or a gift", function "edible". Then, the materials, names, and other key elements included in the key element information after noise reduction can be respectively subjected to text similarity matching analysis with the code description information of the classification code "180610" to obtain the corresponding similarity scores P1, P2,..., Pn, and then the similarity scores P1, P2,..., Pn are weighted and calculated, so as to obtain the confidence level of this candidate classification code "180610". By analogy, the confidence levels of other recalled candidate classification codes can also be obtained.

[0120] In the above-mentioned 1033, the candidate classification codes recalled can be sorted in descending order of confidence, and the classification to which the target object belongs can be determined based on the candidate classification code ranked first. In this embodiment, this candidate classification code ranked first is referred to as the first candidate classification code. Thus, in this step 1033, the classification code determined from multiple candidate classification codes is the first candidate classification code.

[0121] Furthermore, considering that after sorting multiple candidate classification codes in descending order of confidence, for the key element information extracted previously, there may be a lack of key elements for the candidate classification code ranked first (the first candidate classification code, briefly denoted as the Top1 classification code). This is because the previously used key element extraction methods (such as the NER model and other key element extraction means (such as OCR)) have limitations that prevent the corresponding key elements from being extracted. In this case, it may lead to the true correct classification code not being ranked first. For this situation, the following solution will be adopted in this embodiment to supplement the missing key elements: taking the interval of the previously recalled multiple candidate classification codes (such as the Top5 candidate classification codes) as the range for judging missing elements, fusing the coding description information of these multiple candidate classification codes, and semantically reasoning based on the text description information of the target object (such as the object title information) to determine whether there is corresponding missing key element content. The reason for taking the interval of the recalled multiple candidate classification codes as the range for judging missing elements is as follows: Since the key elements required for the current Top1 classification code may be different from those required for the true correct classification code, in order to avoid the missing key elements corresponding to the Top1 classification code not including the missing key elements of the correct classification code, this embodiment uses the interval of multiple candidate classification codes with higher accuracy as the range for judging missing elements. After supplementing the missing key elements, the correct classification code generally receives a reward, while the classification code corresponding to the wrong same key element type generally receives a penalty.

[0122] Based on the above content, the method provided in this embodiment may further include the following steps:

[0123] S11. When the key element information meets the trigger condition for missing key elements, determine the type of missing key elements in the key element information according to the coding description information of the multiple candidate classification codes; where meeting the trigger condition for missing key elements includes: determining that there are missing key elements in the key element information for at least one candidate classification code among the multiple candidate classification codes, and the at least one candidate classification code includes at least the first candidate classification code;

[0124] S12. Based on the missing key element type, the coding descriptions of the at least one candidate classification code, and the text description information of the target object, determine the key element value corresponding to the missing key element type and the matching degree corresponding to the key element value for each candidate classification code in the at least one candidate classification code; wherein, the text description information includes the object title information of the target object.

[0125] S13. Based on the matching degree, update the confidence levels of the at least one candidate classification code, so as to re-trigger the execution of sorting the multiple candidate classification codes from largest to smallest according to the confidence levels after the update.

[0126] In the above S11, it is possible but not limited to determine whether there are missing key elements in the previously extracted key element information for each candidate classification code based on the coding description information of the multiple candidate classification codes.

[0127] For example, taking the first candidate classification code ranked first currently as an example, assuming the first candidate classification code is a classification code "180610" given in the previous steps 10311 - 10312, combining the coding description information of this classification code "180610", it can be determined that when determining whether the classification code of the target object is this "180610", the information required should include: the name of the target object, the production material, the use, and the function. However, among the previously extracted key element information for the target object, only the name, the material, and the function are included, and the use is not included. Then it can be determined that there is a missing key element in the previously extracted key element information for this first candidate classification code, and the missing key element type is the use, and the type description information corresponding to this "use" is "for sale in the retail market and can also be given as snacks or gifts". In this case, it can trigger the execution to determine the set of missing key element types within the range of the previously recalled multiple candidate classification codes. Specifically, for example, according to the coding description information of each candidate classification code, the missing key element types of the key element information for each candidate classification code can be determined, and then the determined missing key element types can be merged and de-duplicated to obtain the set of missing key element types.

[0128] For the specific implementation of whether there are missing key elements in the previously extracted key element information for other candidate classification elements, as well as the missing key element types and the corresponding type description information, reference can be made to the relevant content described above for the first candidate classification code.

[0129] In the above S12, the text description information of the target object (such as the object title of the target object), the determined missing key element type, and the coding description information of each classification code in at least one candidate classification code can be input into a preset missing key element inference model. Then, the following operations will be performed inside the missing key element inference model: Based on the missing key element type and the coding description information of each classification code in at least one candidate classification code, semantic analysis is performed on the object title information of the target object to infer the key element value corresponding to the missing key element type and the matching degree (similarity score) corresponding to the key element value for each classification code in at least one candidate classification code.

[0130] Figure 4a Specifically shows the structural schematic diagram of the above missing key element inference model. And, Figure 4b Shows the specific implementation process schematic diagram of the above steps S11 - S12. As shown in Figure 4b As shown, the missing key element inference model described in the above step S12 includes a condition - satisfaction judgment model and an inference key element information delineation model. The condition - satisfaction judgment model is used to determine the corresponding score (i.e., the matching degree) for the inference result. The condition - satisfaction judgment model can output the following three scores: -1, 0, 1. -1 indicates that there is no content required by the element, that is, it means that no key element value content that meets the missing key element type is inferred based on the object title. 0 indicates that there is content required by the element but it does not meet the requirements, that is: it means that a key element value content that meets the missing key requirement type is inferred based on the object title, but the score value corresponding to this key element value content is less than the set threshold, so it is normalized to 0. 1 indicates that there is content required by the element and it meets the requirements, that is: it means that a key element value content that meets the missing key requirement type is inferred based on the object title, and the score value corresponding to this key element value content is greater than or equal to the set threshold, so it is normalized to 1. The inference key element information delineation model uses the NER (Transformer structure) task to screen and label the data in the corresponding information.

[0131] Specifically, as shown in Figure 4a the structural schematic diagram of the missing key element inference model specifically shown in, the missing key element inference model is a multi - task algorithm model. This multi - task algorithm model combines a classification task (used to satisfy the condition judgment to output the corresponding score) and a named entity recognition (NER) task into one model, and the loss function of this model usually uses a weighted total loss to optimize these two tasks simultaneously. The loss functions of the two tasks are respectively:

[0132] The loss function of the classification task is:

[0133] The loss function for the NER task is as follows:

[0134] Among them,

[0135] Therefore, the total loss function is: TotalLoss = α × Loss cls + β × Loss ner , where α and β are hyperparameters used to control the importance of each task.

[0136] It should be supplemented and explained here that: the input data of the missing key element inference model shown in 5b: X1~Xn represent the object titles of the target objects, M1~Mn respectively represent the encoding description information of the corresponding candidate classification codes, and C represents the type of the missing key element. And, B, I, O, etc. represent the key element values of the output missing key element type.

[0137] The implementation process of the technical solution provided in this embodiment above can be simply described as Figure 5 the process shown. As Figure 5 shown, the implementation of this case mainly includes the following steps:

[0138] Step 1: Identify the key elements required for classification. Specifically, use the corresponding information extraction model (such as NER) to identify and extract the key elements required for classifying the object information of the target object to obtain key element information, such as including product name elements, material elements, and other key element values, etc.

[0139] Step 2: Recall the range of classification codes. Specifically, based on the key element information obtained in Step 1, first determine the corresponding classification chapter code, and further, predict the item code under the classification chapter code, so as to identify the first 2 digits, or even the first 4 digits, six digits, etc. of the possible classification codes, thereby realizing the recall of multiple candidate classification codes with higher rankings (such as the top 5 candidate classification codes).

[0140] Step 3: Determine the confidence level of each candidate classification code, and perform inference on the missing key elements for the previously extracted key element information.

[0141] Step 4: Based on the inference results, update the confidence levels of the corresponding candidate classification codes, and sort each candidate classification code according to the confidence level to obtain a sorting result, so as to determine the classification code of the target object based on the sorting result and further use it to determine the classification to which the target object belongs.

[0142] For the specific implementation descriptions of the above steps 1 to 4, reference can be made to the relevant content in other embodiments, and no specific elaboration will be made here.

[0143] Figure 6 shows an example of implementing the determination of the classification code of a target object (a commodity) using the steps as shown Figure 5 shown. Among them, Figure 7 the COT (Chain-of-Thought) shown in is a technology applied in large language models. It improves the quality and relevance of the content generated by the model by simulating the thinking mode of humans when solving problems. Specifically, CoT requires the model to show its thinking process before generating an answer, which not only directly gives the answer, but also includes steps of reasoning, analysis, and explanation. Large language models usually refer to those deep learning models with a huge number of parameters, rich training data, and excellent performance in natural language processing tasks. These models can understand and generate natural language texts, such as the BERT model (more specifically, the NER model), multi-modal models, etc. described in this application. And, Figure 7 in the tools shown, the information processing system includes a NER model for information recognition and extraction of the text description information, pictures and other multi-modal information of the target object (such as a commodity) to obtain key element information required for classification, etc.; the classification code promotion system includes a coding matching model (such as a multi-modal model) for narrowing down and defining the possible classification code range.

[0144] Specifically in implementation, in the overall pre-classification process of the target object provided by this solution, the following technical means are mainly used to improve the classification accuracy:

[0145] 1. Introduce the pre-classification logic of the first 2 internationally common digits. Specifically, by using the key element information such as the name, use, and material of the target object, and combining the semantic understanding of the corresponding model and the classification precedent data, first comprehensively determine the chapter of the internationally common tariff number (that is, the chapter code of the classification code described in other embodiments), which can greatly narrow the range of tariff number matching in subsequent steps and improve the accuracy.

[0146] 2. In the internationally common 6-digit pre-classification and specific international pre-classification (that is, predicting the classification code of the target object (such as a tariff number)), the following methods are used to improve the classification accuracy and certainty:

[0147] 2.1. For the situation where the existing NER model recognizes multiple product names from the target object information (commodity information), use the multi-modal ability of the corresponding model to recognize the object picture of the target object, so as to comprehensively reason and select a suitable product name to improve the classification accuracy.

[0148] Exemplarily, the target object is a commodity for trade, and the multi-modal information of this commodity is as follows:

[0149] Product Title: 120Pcs Variety colors bookmark Thread Tassel Charm PendantTassels Jewelry (120 pieces of various colored bookmark thread tassel pendants tassel jewelry), 100% Polyester;

[0150] Product Picture: As Figure 7 shown.

[0151] Using the NER model, identify the product title and extract the key elements required for classification as follows:

[0152] Product Name: 'bookmark', 'tassel charm', 'pendant', 'tassels', 'jewelry'

[0153] Material: '100% Polyester'

[0154] Specification: 120pcs (120 pieces)

[0155] It should be noted here that: OCR technology can also be used to recognize and extract text from product pictures. If text information is extracted from the product pictures, the NER model can also be used to recognize and extract the text from the product pictures to extract the corresponding key elements.

[0156] For the multiple product names recognized and extracted from the product title by the above NER model, the product names 'bookmark', 'pendant', and 'jewelry' are not accurate descriptions of the product itself. If directly used for recall classification coding (or called tariff code, which is the tariff number), incorrect recall information will be generated. For example, it may return classification codes (tariff numbers) related to Chapter 71, where Chapter 71 corresponds to natural or cultured pearls, precious or semi-precious stones, precious metals, articles of precious metal clad with other metals. Therefore, a multi-modal model is introduced to recognize the product picture (such as the main product picture), and based on the recognition result of the picture content of the product picture, the correct product name is selected from multiple product names and the selected correct product name is output, such as output: tassels, for subsequent use in determining the classification code of the product.

[0157] 2.2. Meanwhile, in view of some defects of the existing related technologies using the cosine similarity matching algorithm, noise is first removed to purify the extracted key element values. After purification, for some vertical fields and proper nouns, this solution uses a Transformer binary classification model to train the key element similarity scoring model for this task's vertical field, so as to improve the matching accuracy.

[0158] Exemplarily, the target object is a commodity for trade, and the text description information of this commodity is as follows:

[0159] Commodity Title: Outdoor Portable Silicone Water Bottle Leak-proof Fashionable Sports Cup Water Drinking Bottle For Travel Camping Hiking (Outdoor Portable Silicone Water Bottle, Leak-proof, Fashionable Sports Cup, Drinking Bottle for Travel, Camping, Hiking)

[0160] Using the NER model, identify this commodity title and extract the following key elements required for classification:

[0161] Product Name: "water bottle", "cup", "water drinking bottle", "sports bottles";

[0162] Material: "silicone";

[0163] Uses: "travel", "camping", "hiking".

[0164] By using the Transformer binary classification model to train the key element similarity scoring model for this task's vertical field, it can be judged that the uses of camping and hiking are relatively similar to the uses of tableware and kitchenware described in the tariff.

[0165] 2.3. In the case of missing key elements in the output of the top N candidate classification codes extracted by the NER model before, the missing key elements will be supplemented. For the candidate classification codes (tariff numbers) with missing key elements, the coding description information of the candidate classification code (tariff description information) and the target object information (commodity information, such as commodity title) will be used again to determine whether they match through the semantic understanding ability of the corresponding model, so as to further improve the certainty and accuracy of the results;

[0166] Exemplarily, the target object is a commodity for trade, and the text description information of the commodity is as follows:

[0167] Commodity Title: 1set Multifunction Storage Box Rotating Jars Spice RackStorage Organizer Slide Cabinet Decorative Shelves Kitchen Supplies(1 set multifunctional storage box rotating jars spice rack storage organizer slide cabinet decorative shelves kitchen supplies)

[0168] Using the NER model, identify this commodity title and extract the following key elements required for classification:

[0169] Product Name: "storage box", "jars", "spice rack", "storage organizer", "decorative shelves", "kitchen supplies", "racks", "holders"

[0170] Material: "plastic"

[0171] Function: "multifunction", "rotating"

[0172] Before using the corresponding model semantic understanding, the key element type of use is missing. According to the above commodity title (or directly based on the key element values extracted from the commodity title before) and the coding description information of the corresponding candidate classification code (which is the tariff description information), using the reasoning ability of the corresponding model, it can be inferred that it is used in the kitchen, so the key element value of the missing key element type of "use" is: kitchen.

[0173] In addition to the method embodiments described above, the present application also provides several other method embodiments. Specifically as follows:

[0174] A data processing method provided by the present application includes the following steps:

[0175] 201. Obtain the object information of the target object, where the object information includes the text description information and object picture of the target object;

[0176] 202. Based on the text description information, extract the key elements required for classifying the target object to obtain key element information;

[0177] 203. If there are multiple name values of the target object in the key element information, based on the object picture, select a suitable name value from the multiple name values;

[0178] 204. Based on the selected name value, determine multiple candidate classification codes for the target object;

[0179] 205. Based on one of the multiple candidate classification codes, determine the classification to which the target object belongs.

[0180] For the specific implementation of the above steps 201 to 203, refer to the relevant content in other embodiments.

[0181] In the above 204 to 205, first, based on the obtained key element information, determine the corresponding classification chapter code for the target object; then, as Figure 3a shown, recall multiple candidate classification codes related to this classification chapter code based on the selected name value, and determine the confidence level of each candidate classification code based on the key element information, so as to sort the multiple candidates in order according to the confidence level, and select a classification code from the multiple candidate classification codes according to the sorting result for determining the classification to which the target object belongs.

[0182] For the specific implementation of the above steps 204 to 205, refer to the relevant content in other embodiments.

[0183] It should be noted here that: in the method provided in the embodiment of the present application, in addition to the above steps, it may also include other parts or all of the steps in the above embodiments. For details, refer to the corresponding content in the above embodiments.

[0184] A data processing method provided by the present application includes the following steps:

[0185] 301. Obtain the key element information required for classifying the target object;

[0186] 302. Based on the key element information, determine multiple candidate classification codes for the target object and the confidence level of each candidate classification code;

[0187] 303. When the key element information meets the missing key element trigger condition, according to the coding description information of the multiple candidate classification codes, determine the type of the missing key element in the key element information; where, meeting the missing key element trigger condition includes: determining that there is a missing key element in the key element information for at least one of the multiple candidate classification codes;

[0188] 304. Based on the missing key element type, the coding description of the at least one candidate classification code, and the text description information of the target object, determine the key element value corresponding to the missing key element type and the matching degree corresponding to the key element value for each candidate classification code in the at least one candidate classification code;

[0189] 305. Update the confidence level of the at least one candidate classification code based on the matching degree;

[0190] 306. After the update, based on the confidence levels of the multiple candidate classification codes, select one classification code from the multiple candidate classification codes to be used to determine the classification to which the target object belongs.

[0191] In the above 303, satisfying the missing key element trigger condition specifically includes: determining that there is a missing key element in the key element information for at least one candidate classification code among the multiple candidate classification codes, and the at least one candidate classification code at least includes the candidate classification code ranked first (the first candidate classification code) when the multiple candidate classification codes are sorted from largest to smallest according to the confidence level.

[0192] Also, in the above 304, the text description information of the target object includes: the object title of the target object.

[0193] Regarding the specific implementation of the above steps 301 to 304, reference can be made to the relevant content in other embodiments, and specific details will not be elaborated here. In addition, in the method provided in the embodiments of the present application, in addition to the above steps, it may also include other parts or all of the steps in the above embodiments. For specific details, reference can be made to the corresponding content of the above embodiments.

[0194] The target object described in the foregoing embodiments is a commodity for trade. Accordingly, the present application also provides an embodiment of a commodity classification method. Specifically, the method includes the following steps:

[0195] 401. Obtain the commodity information of the target commodity;

[0196] 402. Based on the commodity information, extract the key elements required for classifying the target commodity to obtain key element information;

[0197] 403. Based on the key element information, determine the classification chapter code corresponding to the target commodity;

[0198] 404. According to the classification chapter code and the key element information, determine the classification to which the target commodity belongs.

[0199] Among the above, the product information of the target product includes, but is not limited to: the text description information of the target product (such as product title, product attributes, product categories, product detail descriptions, etc.), product pictures (such as main pictures and / or sub-pictures, etc.).

[0200] For the specific implementation of the above steps 401 to 404, reference can be made to the relevant content in other embodiments, and specific details will not be elaborated here. In addition, in the method provided in the embodiments of the present application, in addition to the above steps, it may also include other parts or all of the steps in the above embodiments. For specific details, reference can be made to the corresponding content in the above embodiments.

[0201] The present application also provides a data processing system, which includes: a server and a client. Among them, the server is used to implement the steps in the method embodiments provided in the present application above. And the client can display some relevant information of the target object on the client interface, such as the object information of the target object and / or the classification information to which the target object belongs. The target object includes the target product. And the classification information includes the classification code to which the target object belongs and / or the confidence level of the classification code. In addition, the classification information may also include other information, such as the scores (for similarity) corresponding to each key element. Each key element is identified and extracted from the object information of the target object, and the score is obtained by performing a similarity matching analysis on the value (for the key element value) in the key element and the code description information of the corresponding classification code. For the classification information of the target object to be displayed, reference can be made to Figure 7 the information content shown in Box1 in the figure.

[0202] It should be added that in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The operation numbers such as 401 and 402 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in order or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0203] The present application also provides corresponding apparatus embodiments respectively corresponding to the above methods. Specifically as follows:

[0204] The present application provides a data processing device, which includes an extraction module and a determination module. Among them, the extraction module is used to extract key elements required for classifying the target object based on the object information of the target object, so as to obtain key element information. The determination module is used to determine the classification chapter code corresponding to the target object based on the key element information; and is also used to determine the classification to which the target object belongs according to the classification chapter code and the key element information.

[0205] Further, when the determination module is used to determine the classification chapter code corresponding to the target object based on the key element information, it is specifically used to: input the key element information into the first convolutional model and the second convolutional model respectively, and obtain the first output result of the first convolutional model and the second output result of the second convolutional model; determine the first digit in the classification chapter code based on the first output result; determine the second digit in the classification chapter code based on the first output result and the second output result.

[0206] And when the determination module is used to determine the second digit in the classification chapter code based on the first output result and the second output result, it can be specifically used to: perform activation processing on a vector formed by combining the first output result and the second output result to obtain an activation processing result; determine the second digit according to the product value of the activation processing result and the first output result and the second output result.

[0207] Further, when the determination module is used to determine the classification to which the target object belongs according to the classification chapter code and the key element information, it is specifically used to: perform noise reduction processing on the key element information to obtain the key element information after noise reduction; determine a plurality of candidate classification codes based on the classification chapter code and the key element information after noise reduction, and use a similarity analysis model to determine the confidence of each candidate classification code; wherein, the similarity analysis model is trained based on a binary classification model relying on a self-attention mechanism; determine a classification code from the plurality of candidate classification codes based on the confidence of each candidate classification code to be used to determine the classification to which the target object belongs.

[0208] Further, the object information includes the text description information and object picture of the target object; the key element information includes the name value of the target object. And when there are multiple name values in the key element information, when the determination module is used to perform noise reduction processing on the key element information, it can be specifically used to: identify the object picture of the target object, and select a suitable name value from the multiple name values based on the identification result.

[0209] Further, one of the determined classification codes is a first candidate classification code, and the first candidate classification code is the candidate classification code ranked first among the multiple candidate classification codes in descending order of confidence.

[0210] In addition, the above-mentioned determination module is further configured to: when the key element information meets the missing key element trigger condition, determine the type of the missing key element in the key element information according to the code description information of the multiple candidate classification codes; where, meeting the missing key element trigger condition includes: determining that there is a missing key element in the key element information for at least one of the multiple candidate classification codes, and the at least one candidate classification code includes at least the first candidate classification code; based on the type of the missing key element, the code description of the at least one candidate classification code, and the text description information of the target object, determine the key element value corresponding to the type of the missing key element and the matching degree corresponding to the key element value for each candidate classification code in the at least one candidate classification code; where, the text description information includes the object title information of the target object. In addition, the apparatus further includes: an update module, configured to update the confidence of the at least one candidate classification code based on the matching degree, so as to re-trigger the execution of sorting the multiple candidate classification codes in descending order of confidence after the update.

[0211] The present application also provides a data processing apparatus, which includes: an acquisition module, an extraction module, a selection module, and a determination module. The above-mentioned acquisition module is configured to acquire the object information of the target object, and the object information includes the text description information and the object picture of the target object. The above-mentioned extraction module is configured to extract the key elements required for classifying the target object based on the text description information, so as to obtain key element information. The above-mentioned selection module is configured to, if the key element information contains multiple name values of the target object, select a suitable name value from the multiple name values based on the object picture. The above-mentioned determination module is configured to determine multiple candidate classification codes for the target object based on the selected name value; and is further configured to determine the classification to which the target object belongs based on one of the multiple candidate classification codes.

[0212] The present application also provides a data processing device, which includes: an acquisition module, a determination module, and an update module. The above-mentioned acquisition module is used to acquire the key element information required for classifying the target object. And, the above-mentioned determination module is used for: based on the key element information, determining multiple candidate classification codes and the confidence levels of each candidate classification code for the target object; when the key element information meets the missing key element trigger condition, determining the type of the missing key element in the key element information according to the code description information of the multiple candidate classification codes; where meeting the missing key element trigger condition includes: determining that there is a missing key element in the key element information for at least one candidate classification code among the multiple candidate classification codes; based on the type of the missing key element, the code description of the at least one candidate classification code, and the text description information of the target object, determining the key element value corresponding to the type of the missing key element and the matching degree corresponding to the key element value for each candidate classification code in the at least one candidate classification code. The above-mentioned update module is used to update the confidence level of the at least one candidate classification code based on the matching degree. The above-mentioned determination module is further used, after the update, to select one classification code from the multiple candidate classification codes based on the confidence levels of the multiple candidate classification codes for determining the classification to which the target object belongs.

[0213] The present application also provides a commodity classification device, which includes: an acquisition module, an extraction module, and a determination module. The above-mentioned acquisition module is used to acquire the commodity information of the target commodity. The above-mentioned extraction module is used to extract the key elements required for classifying the target commodity based on the commodity information to obtain key element information. The above-mentioned determination module is used to determine the classification chapter code corresponding to the target commodity based on the key element information; and is further used to determine the classification to which the target commodity belongs according to the classification chapter code and the key element information.

[0214] It should be noted for the above-mentioned devices: the above-mentioned provided devices can implement the technical solutions described in the corresponding method embodiments above. The specific implementation principles of the above-mentioned modules or units can refer to the relevant content in the corresponding method embodiments above, and will not be specifically elaborated here.

[0215] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, in practice, this electronic device includes: a memory 54 and a processor 55.

[0216] The memory 54 is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of these data include instructions for any application program or method for operating on the electronic device, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0217] A processor 55, coupled to a memory 54, is configured to execute a computer program in the memory 54 to implement the steps in the method embodiments adopted in this application.

[0218] Further, as Figure 8 shown, the electronic device further includes: other components such as a communication component 56, a display 57, a power supply component 58, an audio component 59, etc. Figure 8 Only some components are schematically shown, and it does not mean that the electronic device only includes Figure 8 the components shown. Additionally, Figure 8 the components within the dashed box are optional components, rather than essential components, and can be determined according to the product form of the electronic device. The working node of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT device, or can also be a server device such as a conventional server, a cloud server, or a server array. If the electronic device of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, etc., it may include Figure 8 the components within the dashed box; if the working node of this embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it may not include Figure 8 the components within the dashed box.

[0219] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a Static Random-Access Memory (SRAM), an Electrically Erasable Programmable Read Only Memory (EEPROM), an Erasable Programmable Read Only Memory (EPROM), a Programmable Read-Only Memory (PROM), a Read-Only Memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disc.

[0220] The above-mentioned communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as a 2G, 3G, 4G / LTE, 5G, etc. mobile communication network, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.

[0221] The above-mentioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from users. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with touch or swipe operations.

[0222] The above-mentioned power supply component provides power for various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.

[0223] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC). When the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or sent via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.

[0224] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it causes the processor to be able to implement the steps in the above method embodiments. Among them, the computer-readable storage medium can be implemented by volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, Phase-change RandomAccess Memory (PRAM), Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), other types of Random-Access Memory (RAM), Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), flash memory or other memory technologies, Compact Disc Read-Only Memory (CD-ROM), Digital Video Disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium

[0225] Accordingly, an embodiment of the present application further provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is enabled to implement each step in the above method embodiment. It should be understood that each process or the combination of multiple processes in the above method flow can be implemented by the computer program or instructions. In addition, these computer programs or instructions can be applied to the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices, so that the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices can be used as devices to implement the corresponding functions in the above method embodiment.

[0226] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.

[0227] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A data processing method, characterized in that: include: Based on the object information of the target object, key elements required for classifying the target object are extracted to obtain key element information; Based on the key element information, determining the classification chapter code corresponding to the target object; The category to which the target object belongs is determined according to the category chapter code and the key element information.

2. The method according to claim 1, characterized in that Based on the key element information, determining the classification chapter code corresponding to the target object includes: Inputting the key element information into a first convolution model and a second convolution model respectively, to obtain a first output result of the first convolution model and a second output result of the second convolution model; Based on the first output result, determining a first digit in the classification chapter code; Based on the first output result and the second output result, a second digit in the classification chapter code is determined.

3. The method according to claim 2, characterized in that Determining a second digit in the classification chapter code based on the first output result and the second output result includes: Performing activation processing on a vector formed by combining the first output result and the second output result to obtain an activation processing result; The second bit number is determined according to a product value of the activation processing result and the first output result and the second output result.

4. The method according to any one of claims 1 to 3, characterized in that Determining the classification to which the target object belongs according to the classification chapter code and the key element information includes: Performing noise reduction processing on the key element information to obtain the key element information after noise reduction; Based on the classification chapter code and the denoised key element information, multiple candidate classification codes are determined and the confidence of each candidate classification code is determined using a similarity analysis model; wherein the similarity analysis model is obtained by training a binary classification model based on a self-attention mechanism; Based on the confidence of each candidate classification code, a classification code is determined from the multiple candidate classification codes to determine the classification to which the target object belongs.

5. The method according to claim 4, characterized in that The object information includes text description information and an object picture of the target object; the key element information includes the name value of the target object; And, when the key element information includes multiple name values, the denoising process for the key element information includes: The object image of the target object is recognized to select an adapted name value from the multiple name values ​​based on the recognition result.

6. The method according to claim 5, characterized in that The determined classification code is a first candidate classification code, wherein the first candidate classification code is a candidate classification code ranked first among the plurality of candidate classification codes in descending order of confidence; And, the method further comprises: When the key element information satisfies a missing key element trigger condition, determining the type of missing key elements in the key element information according to the coding description information of the multiple candidate classification codes; wherein satisfying the missing key element trigger condition comprises: determining that there is a missing key element in the key element information for at least one candidate classification code among the multiple candidate classification codes, the at least one candidate classification code at least including the first candidate classification code; Based on the missing key element type, the coding description of the at least one candidate classification code and the text description information of the target object, determining the key element value corresponding to the missing key element type and the matching degree corresponding to the key element value for each candidate classification code in the at least one candidate classification code; wherein the text description information includes the object title information of the target object; Based on the matching degree, the confidence of the at least one candidate classification code is updated to re-trigger the sorting of the plurality of candidate classification codes from large to small according to the confidence after the updating.

7. A data processing method, characterized in that: include: Acquire object information of a target object, wherein the object information includes text description information and an object picture of the target object; Based on the text description information, extract key elements required for classifying the target object to obtain key element information; If the key element information includes multiple name values ​​of the target object, then based on the object image, select an adapted name value from the multiple name values; Based on the selected name value, determining a plurality of candidate classification codes for the target object; Based on a classification code among the multiple candidate classification codes, the classification to which the target object belongs is determined.

8. A data processing method, characterized in that: include: Obtain key element information required to classify the target object; Based on the key element information, determining a plurality of candidate classification codes and the confidence level of each candidate classification code for the target object; When the key element information meets the missing key element trigger condition, the missing key element type in the key element information is determined according to the coding description information of the multiple candidate classification codes; wherein satisfying the missing key element trigger condition includes: determining that the key element information has a missing key element for at least one candidate classification code among the multiple candidate classification codes; Based on the missing key element type, the coding description of the at least one candidate classification code and the text description information of the target object, determining the key element value corresponding to the missing key element type and the matching degree corresponding to the key element value for each candidate classification code in the at least one candidate classification code; Based on the matching degree, updating the confidence of the at least one candidate classification code; After the update, based on the confidences of the multiple candidate classification codes, a classification code is selected from the multiple candidate classification codes to determine the category to which the target object belongs.

9. A commodity classification method, characterized in that: include: Get product information of the target product; Based on the commodity information, extract key elements required for classifying the target commodity to obtain key element information; Based on the key element information, determining the classification chapter code corresponding to the target product; The category to which the target product belongs is determined according to the category chapter code and the key element information.

10. A data processing system, characterized in that: include: A server, used to implement the steps in any one of the methods of claims 1 to 9; The client is used to display at least one of the following information: object information of a target object and classification information to which the target object belongs; the target object includes a target commodity, and the classification information includes the classification code to which the target object belongs and / or the confidence of the classification code.

11. An electronic device, characterized in that: include: A memory and a processor, wherein The memory is used to store programs; The processor is coupled to the memory, and is configured to execute the program stored in the memory to implement the method according to any one of claims 1 to 9.

12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a computer, the method according to any one of claims 1 to 9 can be implemented.

13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.