Code determination method and apparatus, device, and storage medium

By acquiring target object information, dynamically selecting coding methods, and utilizing network models, the problem of low efficiency in manual coding has been solved, achieving automation and improved accuracy in import and export commodity coding, and adapting to the requirements of multiple countries' tariff codes.

WO2025246438A1PCT designated stage Publication Date: 2025-12-04BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/076481
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-30
Filing Date
2025-02-08
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

In the existing technology, the determination of import and export commodity codes relies on manual methods, which leads to low efficiency and insufficient accuracy.

Method used

By acquiring target object information, dynamically selecting an appropriate encoding method based on its category, and using the target network model to determine the encoding, including classification model and attribute prediction model, the system achieves automation and improved accuracy in encoding.

Benefits of technology

It improves the efficiency and accuracy of coding determination, ensures the consistency and interpretability of coding results, and adapts to the tariff requirements of different countries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025076481_04122025_PF_FP_ABST
    Figure CN2025076481_04122025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a code determination method and apparatus, a device, and a storage medium. The code determination method comprises: acquiring target object information corresponding to a target object to be coded (S110); on the basis of a target category to which the target object belongs, determining a target coding approach corresponding to the target object (S120); and inputting the target object information into a target network model in the target coding approach, and on the basis of an output result of the target network model, determining a target code corresponding to the target object, wherein network models in different coding approaches correspond to different model classification granularities (S130). By means of the technical solution of the embodiments of the present disclosure, automatic code determination can be implemented, improving the efficiency and accuracy of code determination.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, apparatus and storage medium for code determination

[0001] This application claims priority to Chinese Patent Application No. 202410693507.3, filed May 30, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD

[0002] Embodiments of the present disclosure relate to a method, device, apparatus and storage medium for code determination. BACKGROUND

[0003] In the context of international trade, etc., it is necessary to code objects such as imported and exported goods. Accurate coding can ensure the smooth progress of customs clearance, and is of great significance to ensuring the fairness and compliance of international trade, and can also ensure accurate tax assessment, thereby better controlling costs. Currently, professional personnel usually manually determine the coding of objects such as goods based on import and export tariffs. However, this manual determination method is time-consuming and laborious, and reduces the coding determination efficiency. SUMMARY

[0004] The present disclosure provides a method, device, apparatus and storage medium for code determination to realize automatic coding determination, thereby improving coding determination efficiency and accuracy.

[0005] In a first aspect, embodiments of the present disclosure provide a method for code determination, comprising:

[0006] obtaining target object information corresponding to a target object to be coded;

[0007] determining a target coding mode corresponding to the target object based on a target category to which the target object belongs;

[0008] inputting the target object information into a target network model in the target coding mode, and determining a target code corresponding to the target object based on an output result of the target network model;

[0009] wherein the network models in different coding modes correspond to different model classification granularities.

[0010] In a second aspect, embodiments of the present disclosure also provide a device for code determination, comprising:

[0011] a target object information obtaining module configured to obtain target object information corresponding to a target object to be coded;

[0012] a target coding mode determining module configured to determine a target coding mode corresponding to the target object based on a target category to which the target object belongs;

[0013] The target coding determination module is configured to input the target object information into a target network model in the target coding mode, and determine a target coding corresponding to the target object based on an output result of the target network model; wherein the network models in different coding modes correspond to different model classification granularities.

[0014] In a third aspect, the present disclosure provides an electronic device, which comprises:

[0015] one or more processors;

[0016] a storage device configured to store one or more programs,

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the coding determination method according to any of the embodiments of the present disclosure.

[0018] In a fourth aspect, the present disclosure provides a storage medium containing computer executable instructions, which, when executed by a computer processor, are configured to perform the coding determination method according to any of the embodiments of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0019] The above and other features, advantages, and aspects of the present embodiments will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.

[0020] FIG. 1 is a flow diagram of a coding determination method according to an embodiment of the present disclosure;

[0021] FIG. 2 is a flow diagram of another coding determination method according to an embodiment of the present disclosure;

[0022] FIG. 3 is an example diagram of a target coding determination process according to an embodiment of the present disclosure;

[0023] FIG. 4 is an example diagram of a network architecture of a classification model according to an embodiment of the present disclosure;

[0024] FIG. 5 is a flow diagram of yet another coding determination method according to an embodiment of the present disclosure;

[0025] FIG. 6 is an example diagram of a network architecture of an attribute prediction model according to an embodiment of the present disclosure;

[0026] FIG. 7 is an example diagram of a target coding determination process according to an embodiment of the present disclosure;

[0027] FIG. 8 is a structural schematic diagram of an encoding determination apparatus according to an embodiment of the present disclosure; and

[0028] FIG. 9 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are merely for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.

[0030] It should be understood that the various steps in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0031] The term "comprising" and variations thereof as used herein are open-ended, that is, "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the description below.

[0032] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0033] It should be noted that the adjectives "one", "more than one" mentioned in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that "one or more" should be understood unless the context clearly indicates otherwise.

[0034] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes and are not intended to limit the scope of the messages or information.

[0035] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.

[0036] Figure 1 is a flowchart illustrating a coding determination method provided in an embodiment of this disclosure. This embodiment is applicable to determining the coding of objects such as import and export commodities. The method can be executed by a coding determination device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.

[0037] As shown in Figure 1, the encoding determination method specifically includes the following steps:

[0038] S110. Obtain the target object information corresponding to the target object to be encoded.

[0039] The target object can refer to any object that needs to be encoded. For example, the target object could refer to import / export commodities. Import / export commodities refer to goods that need to be imported or exported in the e-commerce field. There are 22 categories and 98 chapters of commodity coding, totaling over 13,000 entries, thus requiring the determination of the target code matching the target object based on the target object information. Target object information can be information used to describe the target object. For example, target object information can include at least one of target text information and target image information used to describe the target object. Target text information refers to information that describes the target object in text form. For example, target text information includes the title information and category information of the target object. Target image information refers to information that describes the target object in image form. For example, target image information includes images taken of the target object.

[0040] Specifically, target text information and target image information corresponding to the target object to be encoded can be obtained, so as to fully explore the features of the target object by utilizing information from both text and image modalities, and further improve the accuracy of encoding determination.

[0041] S120. Based on the target category to which the target object belongs, determine the target encoding method corresponding to the target object.

[0042] The target category can refer to the category information to which the target object belongs. The target encoding method is the encoding method used to determine the corresponding object under the target category. There can be at least two encoding methods. Different encoding methods can be used to determine the corresponding code for objects under different categories. For example, the target encoding method includes a first encoding method or a second encoding method. The first encoding method is a method that directly determines the object code based on at least two classification models, while the second encoding method is a method that indirectly determines the object code based on an attribute prediction model and the correspondence between pre-configured object codes and attribute information sets.

[0043] Specifically, based on the target category to which the target object belongs, a target encoding method that is suitable for the target object can be dynamically selected, thereby improving the accuracy of encoding determination. For example, based on the target category to which the target object belongs, a target encoding method that is suitable for the target object can be dynamically selected from the first encoding method and the second encoding method, thereby improving the accuracy of encoding determination.

[0044] For example, S120 may include: if the target category to which the target object belongs is a first preset category corresponding to the first encoding method, then determine that the target encoding method corresponding to the target object is the first encoding method; if the target category to which the target object belongs is a second preset category corresponding to the second encoding method, then determine that the target encoding method corresponding to the target object is the second encoding method.

[0045] The first preset category can be pre-set, and the first encoding method applies to the category to which the processed object belongs. The second preset category can also be pre-set, and the second encoding method applies to the category to which the processed object belongs. The first and second preset categories can be pre-determined by testing the encoding accuracy of each encoding method for each object category. For example, if the encoding accuracy using the first encoding method is greater than the encoding accuracy using the second encoding method in a certain category, it indicates that encoding objects in that category using the first encoding method is more accurate, and this category can be determined as the first preset category corresponding to the first encoding method.

[0046] Specifically, the target category to which the target object belongs can be matched with the first preset category corresponding to the first encoding method. If the match is successful, meaning the first preset category includes the target category, then the target encoding method corresponding to the target object is determined to be the first encoding method. If the first preset category does not include the target category, then the target category is matched with the second preset category corresponding to the second encoding method. If the match is successful, meaning the second preset category includes the target category, then the target encoding method corresponding to the target object is determined to be the second encoding method. By using category matching, the target encoding method can be quickly determined, allowing for the use of a more accurate encoding method to classify and encode the target object, thus improving the accuracy of encoding determination.

[0047] For example, the first preset category is a popular category, and the second preset category is a non-popular category. The popular category refers to the category to which objects belong, whose demand and / or sales volume exceeds a preset threshold. The non-popular category refers to the remaining categories other than the popular category. If the target object belongs to the first preset category (i.e., the popular category), then because the classification model in the first encoding method can be sufficiently trained, the first encoding method can be used as the target encoding method for direct classification, ensuring encoding accuracy. If the target object belongs to the second preset category (i.e., the non-popular category), then the second encoding method can be used as the target encoding method for indirect classification by predicting attribute information, ensuring encoding accuracy.

[0048] S130. Input the target object information into the target network model in the target encoding method, and determine the target encoding corresponding to the target object based on the output of the target network model; wherein, the network model in different encoding methods corresponds to different model classification granularity.

[0049] In this context, the target network model refers to the neural network model used to determine the encoding method. Network models with different encoding methods have different classification targets. Model classification granularity refers to the different levels of granularity in the information classified by the network model. In other words, network models with different encoding methods have different granularity levels of classification in their output information.

[0050] For example, in the first encoding method, the network model is at least two classification models, and different classification models are used to determine different types of encoding; in the second encoding method, the network model is an attribute prediction model, and the classification granularity of the attribute prediction model is finer than that of the classification model.

[0051] The classification model can be a network model used to achieve end-to-end object classification. Object classification refers to assigning an object to the corresponding code. Each classification model is pre-trained based on sample data to ensure classification accuracy. Different classification models can be used to predict codes for different object types. The code type can refer to the object's code in different import / export countries. Because different countries have different tariffs, the same object will have different codes in different countries. The attribute prediction model can be a network model used to predict object attribute information. The attribute prediction model is also pre-trained based on sample data to ensure the accuracy of attribute information prediction. Object attribute information can refer to the attribute values ​​corresponding to the attributes possessed by the object. For example, if the object is a men's cotton down jacket, then the attribute information of the object includes: gender "male", category "down jacket", and material "cotton". Because the predicted object attributes are more granular than the object codes, the classification granularity of the attribute prediction model is more granular than that of the classification model, and the processing power of the attribute prediction model is also stronger than that of the classification model. The first encoding method can be an end-to-end object classification approach using at least two classification models. However, the classification results lack interpretability and the ability to intervene in the classification process. The second encoding method can be an indirect object classification approach using attribute prediction models and the correspondence between object codes and attribute information sets. This approach makes the classification results interpretable and allows for adjustments to the strategy to intervene in the model's classification behavior at any time.

[0052] Specifically, the target object information is input into the target network model in the target encoding method for classification processing at an appropriate granularity, and the output results of the target network model are processed to determine the final target code of the target object. Thus, the target network model can be used to realize the automatic determination of the code, improving the efficiency and accuracy of code determination.

[0053] For example, when the target encoding method is the first encoding method, at least two classification models in the first encoding method can be used to directly classify objects and determine the final target code. When the target encoding method is the second encoding method, the attribute prediction model in the second encoding method can be used to predict attribute information, and based on the predicted attribute information set and the pre-configured correspondence between the object code and the attribute information set, the target code corresponding to the target object can be indirectly determined. By using different encoding methods to determine the object codes under different categories, the accuracy of code determination can be improved.

[0054] The technical solution of this disclosure involves obtaining target object information corresponding to the target object to be encoded, and determining a target encoding method suitable for the target object based on the target category to which the target object belongs; inputting the target object information into the target network model in the target encoding method, and determining the target code corresponding to the target object based on the output result of the target network model. In this case, the network model in different encoding methods corresponds to different model classification granularities, so that the code of the target object can be determined more accurately by using the target network model adapted to the target object, realizing automatic code determination and improving the efficiency and accuracy of code determination.

[0055] Figure 2 is a flowchart illustrating another encoding determination method provided by an embodiment of this disclosure. Based on the above-described embodiments, this disclosure provides a detailed description of the process for determining the target encoding corresponding to the target object when the target encoding method is the first encoding method. Explanations of terms that are the same as or corresponding to those in the above-described embodiments are not repeated here.

[0056] As shown in Figure 2, the encoding determination method specifically includes the following steps:

[0057] S210. Obtain the target object information corresponding to the target object to be encoded.

[0058] S220. Based on the target category to which the target object belongs, determine the target encoding method corresponding to the target object.

[0059] S230. If the target encoding method is the first encoding method, then the target object information is input into at least two classification models in the first encoding method to obtain at least two candidate codes.

[0060] The classification model can be a network model used to achieve end-to-end object classification. Different classification models are used to predict the code of an object under different types, such as predicting the code under different countries. Because different countries have different tariffs, the same object will have different codes in different countries. For example, at least two classification models can include a first-type coding classification model and a second-type coding classification model. The first-type coding classification model can be used to predict the code of an object under the first type, such as the import code in the importing country. The second-type coding classification model can be used to predict the code of an object under the second type, such as the export code in the exporting country. Accordingly, at least two candidate codes can include: a candidate first-type code and a candidate second-type code.

[0061] Specifically, when the target coding method is the first coding method, the target object information can be directly input into each classification model. Each model then performs object classification based on the input target object information, obtaining at least two candidate codes corresponding to the target object. For example, the target object information can be input into a first-type coding classification model and a second-type coding classification model. The first-type classification model performs object classification based on the input target object information, obtaining candidate first-type codes (e.g., candidate import codes). The import classification model can also perform object classification based on the input target object information, obtaining candidate second-type codes (e.g., candidate export codes) corresponding to the target object.

[0062] S240. Based on at least two candidate codes, determine the target common code corresponding to the target object.

[0063] Among these, codes of different types share the same coding portion. Target shared coding can refer to the same coding portion in both the first-type and second-type codes. For example, the first 6 digits of codes from different countries are shared; that is, the first 6 digits of the same object's code are the same in different countries. For instance, the first-type code and second-type code of the same object (such as export and import codes) have the same first 6 digits, and this shared coding portion can be considered a shared coding between the two types.

[0064] Specifically, the common parts of at least two candidate codes output by all classification models can be compared to determine the target common code corresponding to the target object. For example, the first type common code (e.g., the first 6 digits of the candidate first type code) can be determined based on the candidate first type codes output by the first type coding classification model, and the second type common code (e.g., the first 6 digits of the candidate second type code) can be determined based on the candidate second type codes output by the second type coding classification model. If the first type common code and the second type common code are found to be the same, then either the first type common code or the second type common code can be directly determined as the target common code corresponding to the target object. If the first type common code and the second type common code are found to be different, then the confidence scores corresponding to the first type common code and the second type common code can be compared, and the common code with the highest confidence score can be determined as the target common code.

[0065] S250. Adjust the candidate code based on the target common code to determine the target code corresponding to the target object.

[0066] The target code can include a first-type target code and a second-type target code. For example, the first-type target code can refer to the target import code, and the second-type target code can refer to the target export code. The target import code refers to the code of the target object in the importing country. The target export code refers to the code of the target object in the exporting country.

[0067] Specifically, candidate codes can be adjusted based on the target common code to obtain the target code with the target common code. This allows for the use of at least two classification models to improve classification accuracy and ensure the consistency of multiple types of codes, thereby ensuring the consistency of codes in multiple countries.

[0068] For example, S250 may include: adjusting the encoding of the candidate first type encoding and the candidate second type encoding based on the target common encoding, and determining the target first type encoding and the target second type encoding corresponding to the target object; wherein, the common encoding part in the target first type encoding and the target second type encoding is the target common encoding.

[0069] Specifically, the common codes in the candidate first type codes and the common codes in the candidate second type codes are all adjusted to the target common codes, thereby obtaining the adjusted target first type codes and target second type codes. At this time, the two target codes have the same common code, such as the first 6 digits of the code being the same. This ensures that the target object has the same first 6 digits in the codes declared in the importing country and the exporting country, resulting in stronger consistency and higher accuracy.

[0070] For example, target object information and at least two pre-trained classification models can be used to classify objects in exporting and importing countries. This allows for the unified determination of the target code declared by the target object in both the importing and exporting countries. This ensures that the two target codes declared by the importing and exporting countries have the same common code (i.e., the first 6 digits), guaranteeing the consistency of the code for the same object across multiple countries. It also improves the accuracy of code determination and, consequently, guarantees the ability to classify numerous objects.

[0071] The technical solution of this disclosure embodiment is to input the target object information into at least two pre-trained classification models to obtain at least two candidate codes, and to determine the target common code corresponding to the target object based on the at least two candidate codes. The candidate codes are then adjusted based on the target common code to obtain a target code with the same common code. This enables the integrated prediction of multiple types of codes, making the coding results more accurate. Furthermore, it avoids the situation where the common code of the same object is inconsistent in the import and export countries, making cross-border management more convenient.

[0072] Based on the above technical solutions, at least two classification models may include: a first-type coding classification model, a second-type coding classification model, and a shared coding classification model; at least two candidate codes include candidate first-type codes, candidate second-type codes, and candidate shared codes. The shared coding classification model can be a network model used to predict the shared coding portion of the object. For example, each code has 10 bits, and the first 6 bits of codes from different countries are shared. The shared coding classification model can be used to predict the first 6 shared bits of the object. The first-type coding classification model can be used to predict the 10-bit import code of the object in the importing country. The second-type coding classification model can be used to predict the 10-bit export code of the object in the exporting country. Candidate shared codes refer to codes corresponding to the same coding portion in the first-type and second-type codes. For example, candidate shared codes refer to the first 6 bits predicted by the shared coding classification model. The target first-type code and the target second-type code have the same shared code. For example, the 10-bit target import code and the 10-bit target export code have the same first 6 bits.

[0073] For example, step S240 may include: determining the target common code corresponding to the target object based on the candidate first type code, the candidate second type code, and the candidate common code. Step S250 may include: adjusting the candidate first type code and the candidate second type code based on the target common code to determine the target first type code and the target second type code corresponding to the target object.

[0074] Specifically, as shown in Figure 3, the target object information is input into the first type coding classification model, the second type coding classification model, and the shared coding classification model for object classification. This yields three coding classification results: candidate first-type codes output by the first type coding classification model, candidate second-type codes output by the second type coding classification model, and candidate shared codes output by the shared coding classification model. The shared codes in the candidate first-type codes, the candidate second-type codes, and the candidate shared codes are compared. A voting ensemble approach can be used to determine the shared code with the highest number of votes (i.e., the most frequent occurrences) as the target shared code. If there are shared codes with the same number of votes, the shared code with the highest prediction confidence is determined as the target shared code. By utilizing the target shared code, the shared codes in the candidate first-type codes and candidate second-type codes are adjusted to limit the prediction space of the first and second type coding classification models, resulting in target first-type codes and target second-type codes with the same shared code. In other words, the shared codes in both the target first-type codes and target second-type codes are target shared codes. By adding a common coding classification model to the first and second type coding classification models, the target common coding can be determined more accurately, thereby further improving the consistency and accuracy of multiple types of coding.

[0075] Based on the above technical solutions, when the target object information includes target text information and target image information describing the target object, each classification model can include: a text encoding sub-model, an image encoding sub-model, and a classification sub-model. The network architecture of each classification model is identical. The text encoding sub-model can be an encoding network used to extract text features of the object. For example, a text encoding sub-model can refer to an unsupervised pre-trained language model based on a Transformer structure, such as the BERT (Bidirectional Encoder Representations from Transformers) model. The image encoding sub-model can be an encoding network used to extract image features of the object. For example, an image encoding sub-model can refer to a residual network model, such as ResNet-50. The classification sub-model is a network used to classify objects by fusing textual and image feature information. For example, the classification sub-model can include a fusion layer and a classification layer. A feedforward neural network (FFN) can be used as a fusion layer for feature fusion. A fully connected layer can be used as a classification layer for classification processing.

[0076] For example, as shown in Figure 4, step S230, "inputting the target object information into at least two classification models in the first encoding method to obtain at least two candidate codes," may include: for each classification model, inputting the target text information into the text encoding sub-model of the classification model for text encoding processing to obtain the target text feature information corresponding to the target object; inputting the target image information into the image encoding sub-model of the classification model for text encoding processing to obtain the target image feature information corresponding to the target object; and inputting the target text feature information and the target image feature information into the classification sub-model of the classification model for object classification processing to obtain the candidate codes corresponding to the target object.

[0077] Among them, target text feature information is used to reflect the semantic attribute information of the descriptive text corresponding to the target object. Target image feature information is used to reflect the shape attribute information of the target object, such as color, shape, and size, so as to infer the object's category, material, and purpose.

[0078] Specifically, since the available object information is limited, each classification model can perform cross-modal interaction of the input multimodal object information (i.e., target text information and target image information) compared to using only text information, thereby mining more usable information from the limited information and improving the accuracy of object classification.

[0079] For example, during the training of each classification model, the initial network weight values ​​in the text encoding sub-model are obtained by pre-training based on the object text dataset, and the initial network weight values ​​in the image encoding sub-model are obtained by pre-training based on the object image dataset.

[0080] The object text dataset can be obtained by cleaning large-scale object text data, removing non-semantic content such as garbled characters, emojis, special symbols, and code tags. The object image dataset can be obtained by hashing and deduplicating large-scale object image data, resulting in a dataset of approximately 100 million images.

[0081] Specifically, each classification model is trained based on pre-trained text-encoded sub-models and image-encoded sub-models, using sample object information and sample encoding corresponding to the sample objects. The pre-training process for the text-encoded sub-model involves pre-training the model using an object text dataset and masked language modeling, enabling it to understand object text information. The pre-training process for the image-encoded sub-model involves pre-training the model using an object image dataset and momentum contrastive unsupervised learning, enabling it to have robust image representation capabilities. Based on object ID aggregation data, images from the same object are treated as positive samples, and images from different objects as negative samples. A contrastive learning loss function (e.g., InfoNCE loss function) and a metric learning loss function (e.g., triplet loss function) are used to jointly supervise the image-encoded sub-model, narrowing the cosine distance between the embedded features of positive samples and widening the cosine distance between the embedded features of negative samples, thus enabling the image-encoded sub-model to also represent the main object. By pre-training the image coding sub-model using a large-scale object image dataset, the image coding sub-model can be made capable of understanding image content.

[0082] After the text encoding sub-model and image encoding sub-model are pre-trained, the pre-trained text network weights of the text encoding sub-model and the pre-trained image network weights of the image encoding sub-model can be loaded into each classification model. This makes the initial network weights of the text encoding sub-model and the image encoding sub-model in each classification model the pre-trained text network weights. As a result, when training each classification model, the multimodal content understanding ability can be transferred to the encoding and classification task. Consequently, each classification model can more fully understand the object information, has better generalization ability, and classify objects not covered by the sample data more accurately.

[0083] Based on the above technical solutions, each classification model is trained by using the sample object information corresponding to the sample object as input information and the probability distribution corresponding to the sample object's sample code as output label. The probability distribution includes the probability value of each optional code being the actual code of the sample object, where the probability value of the sample code being the actual code of the sample object is the largest, and the code length of the optional code and the sample code being the same is positively correlated with the probability value.

[0084] Here, "sample code" refers to the code to which a sample object is currently classified, which can be determined manually and may contain labeling errors. "Actual code" refers to the accurate code to which the sample object should be classified, i.e., the true code. "Optional code" refers to a code to which an object can be classified. Each code in the code library can be considered an optional code. "Same code length" refers to the number of bits that an optional code shares with a sample code. For example, if the total code length is 10 bits, and the first 6 bits of an optional code are the same as the first 6 bits of a sample code, then the same code length between the optional code and the sample code is determined to be 6. The longer the same code length, the higher the probability that the optional code is the actual code. The probability distribution corresponding to the sample object is used to characterize the probability distribution of the optional code being the actual code. The probability value corresponding to each candidate code in the probability distribution is a value greater than or equal to 0 and less than 1.

[0085] Specifically, compared to one-hot encoding based on sample encoding, which uses the resulting encoding vector (the dimension of the vector equals the number of optional codes, and each element in the vector corresponds one-to-one with an optional bias code, with each element having a value of 0 or 1, where 0 represents the corresponding optional code that is not a sample code and 1 represents the corresponding optional code that is a sample code) as the output label, this embodiment uses the probability distribution corresponding to the sample encoding as the output label for model training. This allows the model to learn the similarity of object encodings, making the trained classification model more robust and improving classification accuracy.

[0086] For example, the encoding can be divided into multiple levels according to the number of bits, thus converting the encoding into a tree-like hierarchical structure. For instance, a 10-bit encoding can be divided into 5 levels, with each level corresponding to a 2-bit encoding. By determining whether the encoding of an optional encoding at each level is the actual encoding, the probability value of each optional encoding being the actual encoding is determined, thereby obtaining a probability distribution. For example, suppose the i-th level has k... i Each of the following options has an equal probability p: i Other optional codes that are classified outside of the sample code can be determined based on the probability that the sample code is the actual code at level i, which is (1-(k)). i -1)p i Considering the complete 5-level classification, the probability value for determining that a sample code is the actual code of the sample object is... Similarly, the probability values ​​of other optional codes besides the sample code being the actual coded labels can be calculated, thereby transforming the vector labels corresponding to the sample codes into a probability distribution of implicit hierarchical information.

[0087] For example, the probability that other optional codes in the first level are the actual codes is p. i After that, the probability value of the optional encoding at each level being the actual encoding is 1 / k. iConsidering the complete 5-level classification, the probability value of the optional encoding being the actual encoding in this case is... (For example, if the sample code is "6102000000", all codes that do not begin with 61 fall into this category). The probability of a sample code being the actual code at the first level (i.e., the sample code and the first two digits of the actual code being the same) and other optional codes at the second level being the actual code is [value missing]. (For example, if the sample code is "6102000000", all codes that start with 61 but are not 6102 belong to this case). The sample codes of the first x levels (x∈[1,3]) are the actual codes, and the other optional codes of the (x+1)th level are the actual codes. The probability of this case is... The probability of the first four levels of samples being encoded as the actual code, and the fifth level having other optional codes being encoded as the actual code, is as follows: (For example, if the sample code is "6102000000", all codes that start with 61020000 but do not end with 00 belong to this category).

[0088] By using backpropagation to train parameters, the sample object information corresponding to the sample object is used as input information, and the probability distribution corresponding to the sample object's sample code is used as output label for supervised model training. This allows the trained classification model to learn the hierarchical structure of the code, that is, to learn the similarity of codes at the same level. It also takes into account the noise of the sample code, thus making the model more robust and improving the classification accuracy.

[0089] Figure 5 is a flowchart illustrating another encoding determination method provided by an embodiment of this disclosure. Based on the above-described embodiments, this disclosure provides a detailed description of the process for determining the target encoding corresponding to the target object when the target encoding method is the second encoding method. Explanations of terms that are the same as or corresponding to those in the above-described embodiments are not repeated here.

[0090] As shown in Figure 5, the encoding determination method specifically includes the following steps:

[0091] S310. Obtain the target object information corresponding to the target object to be encoded.

[0092] S320. Based on the target category to which the target object belongs, determine the target encoding method corresponding to the target object.

[0093] S330. If the target encoding method is the second encoding method, then the target object information is input into the attribute prediction model in the second encoding method to obtain the first attribute information set.

[0094] The attribute prediction model can be a network model used to predict object attribute information. Object attribute information can refer to the attribute values ​​corresponding to the attributes possessed by the object.

[0095] Specifically, when the target encoding method is the second encoding method, the target object information is input into a pre-trained attribute prediction model, enabling the attribute prediction model to accurately predict the attribute values ​​corresponding to the attributes possessed by the target object based on the input target object information, and then output the predicted attribute information. The attribute prediction model outputs one or more attribute information sets. All attribute information output by the attribute prediction model is combined to obtain the first attribute information set.

[0096] S340. Based on the correspondence between the first set of attribute information and the pre-configured object code and attribute information set, determine the target code corresponding to the target object.

[0097] The attribute information set can include all attribute information possessed by the object corresponding to the object code. The correspondence between the object code and the attribute information set can be pre-configured based on the tariff codes corresponding to the target object. For example, the correspondence between the object code and the attribute information set can be pre-determined based on the tariff codes of the importing country corresponding to the target object, so as to determine the import code of the target object based on this correspondence. Alternatively, the correspondence between the object code and the attribute information set can be pre-determined based on the tariff codes of the exporting country corresponding to the target object, so as to determine the export code of the target object based on this correspondence. Different countries have different tariff codes, thus requiring different correspondences between the object code and the attribute information set for different countries. For example, the correspondence between the object code and the attribute information set can be automatically determined based on a large language model and the tariff codes of the importing and exporting countries. The large language model is a deep learning model trained on massive amounts of text data. It can not only generate natural language text, but also deeply understand the meaning of text and handle various natural language tasks, such as text summarization, question answering, and translation. By calling a large language model and inputting a description of the object code in the tariff code, along with the possible attributes and attribute values ​​of the code, the large language model can automatically output a set of attribute information corresponding to each object code. Alternatively, the attributes and attribute values ​​output by the large language model can be manually reviewed to remove duplicates and add missing items, thus obtaining a more accurate correspondence between object codes and attribute information sets. For example, the attribute information set {gender: male, category: down jacket, material: cotton} corresponds to the code 6201301000 for a men's cotton down jacket.

[0098] Specifically, based on the pre-configured correspondence between object codes and attribute information sets, the object code corresponding to the first attribute information set is determined as the target code for the target object. By indirectly determining the object code through predicting attribute information, the attribute information upon which the object code is based can be obtained. This allows for rapid problem location and correction in case of misclassification, improving the interpretability of the classification results. Furthermore, when the target object information is relatively complete, it can provide higher-quality classification results, improving the accuracy of code determination. Moreover, when tariff regulations change, only the correspondence between the object code and the attribute information set needs to be modified, making the classification logic controllable and easy to maintain.

[0099] The technical solution of this disclosure, when the target encoding method is the second encoding method, can accurately determine the target encoding corresponding to the target object based on the correspondence between the first attribute information set output by the attribute prediction model and the pre-configured object encoding and attribute information set, and make the encoding result interpretable, thereby obtaining a higher quality classification result.

[0100] Based on the above technical solutions, when the target object information includes target text information and target image information used to describe the target object, the attribute prediction model can include: a text encoding sub-model, an image encoding sub-model, and an attribute prediction sub-model. The training process and network structure of the text encoding sub-model and image encoding sub-model in the attribute prediction model are the same as those in the classification model, and can be referred to the above description, so they will not be repeated here. Referring to Figure 6, the attribute prediction sub-model can include an attention processing layer and an attribute value classification layer. The attention processing layer is an attribute prediction processing network for multi-head attention. For example, the attention processing layer can include multiple serially connected Transformer modules, such as four identical Transformer modules stacked together. The attribute value classification layer can be a fully connected layer used to determine the specific attribute value corresponding to the attribute. The number of attribute value classification layers is the same as the number of candidate attributes to be identified.

[0101] For example, as shown in Figure 6, step S330, "inputting the target object information into the attribute prediction model in the second encoding method to obtain the first attribute information set", may include: inputting the target text information and target image information into the text encoding sub-model for text encoding processing to obtain the target text feature information corresponding to the target object; inputting the target image information into the image encoding sub-model for text encoding processing to obtain the target image feature information corresponding to the target object; inputting the target text feature information and target image feature information into the attribute prediction sub-model for multi-head attention attribute prediction processing to obtain the first attribute information set corresponding to the target object, such as the set composed of attribute value 1, attribute value 2 and attribute value 3 in Figure 6.

[0102] Among them, target text feature information is used to reflect the semantic attribute information of the descriptive text corresponding to the target object. Target image feature information is used to reflect the shape attribute information of the target object, such as color, shape, and size, so as to infer the object's category, material, and purpose.

[0103] Specifically, since the available object information is limited, the attribute prediction model can interact across modalities with the input multimodal object information (i.e., target text information and target image information) compared to using only text information, thereby extracting more usable information from the limited information and improving the accuracy of attribute prediction.

[0104] For example, the attribute prediction sub-model can implement the attribute prediction processing functionality of multi-head attention through the following steps:

[0105] Self-attention processing is performed on the first attribute feature vectors corresponding to all candidate attributes to obtain the second attribute feature vector; cross-attention processing is performed on the second attribute feature vector, the input target text feature information, and the target image feature information to obtain the third attribute feature vector; attribute prediction processing is performed based on the third attribute feature vector to obtain the first attribute information set corresponding to the target object.

[0106] Here, the candidate attributes can refer to any attributes an object might possess. For example, each attribute in the attribute library can be considered as a candidate attribute. The first attribute feature vector is obtained by training the attribute prediction model based on a preset attribute feature vector. The preset attribute feature vector can refer to the initial values ​​set before training the attribute prediction model. During model training, the preset attribute feature vector is continuously adjusted. The first attribute feature vector is the attribute feature vector obtained after the model training is completed. There is a one-to-one correspondence between the first attribute feature vector and the candidate attributes, and the number of first attribute feature vectors equals the number of candidate attributes.

[0107] Specifically, referring to Figure 6, the first attribute feature vectors corresponding to all candidate attributes, the target text feature information, and the target image feature information are input into the attention processing layer in the attribute prediction sub-model. In the attention processing layer, self-attention processing is performed on all first attribute feature vectors to learn the interaction between attribute vectors, thereby obtaining the second attribute feature vector corresponding to each candidate attribute. The second attribute feature vector is used as the query, and the input target text feature information and target image feature information are used as the key and value for cross-attention processing to fuse more valuable multi-modal information into the attribute vectors. A gating network activation function can be used to obtain the third attribute feature vector corresponding to each candidate attribute. The third attribute feature vector corresponding to each candidate attribute is input into the corresponding attribute classification layer, where fully connected processing is performed to obtain the attribute value and corresponding prediction probability value for each candidate attribute. Attribute values ​​with prediction probability values ​​greater than a preset threshold are output. The output attribute values ​​are combined to obtain the first attribute information set corresponding to the target object. Through self-attention processing, attribute vectors interact with each other, avoiding contradictory attribute predictions, such as the simultaneous occurrence of "Material: Steel" and "Type: Coat". Through cross-attention processing, attribute vectors can fuse the text and image information that each attribute needs to focus on. That is, the third attribute feature vector corresponding to each candidate attribute contains the information that needs to be focused on to identify the corresponding candidate attribute, thereby further improving the accuracy of attribute prediction.

[0108] For example, the attribute prediction model can be trained based on the sample object information and the sample attribute information set corresponding to the sample object, enabling the attribute prediction model to recognize attribute information. The sample attribute information set can be determined based on the sample encoding and the pre-configured correspondence between the object encoding and the attribute information set, eliminating the need for manual annotation and thus saving training costs.

[0109] Based on the above technical solutions, step S340 may include: processing the target text information based on the pre-configured correspondence between keywords and attribute information and / or the correspondence between attribute expressions and attribute information to determine the second set of attribute information corresponding to the target object; determining the target attribute information set corresponding to the target object based on the first set of attribute information and the second set of attribute information; and determining the target code corresponding to the target object based on the correspondence between the target attribute information set and the pre-configured object code and attribute information set.

[0110] Here, keywords can refer to keywords used to describe object information. Attribute expressions can refer to regular expressions that object information needs to satisfy. For example, when an object has a certain attribute, it needs to satisfy the attribute expression corresponding to that attribute. The second set of attribute information can include the third set of attribute information and / or the fourth set of attribute information. The third set of attribute information is the set of attribute information determined by keyword matching. The fourth set of attribute information is the set of attribute information determined by attribute expression matching. The target attribute information set refers to the set of attribute information that the target object ultimately possesses.

[0111] Specifically, referring to Figure 7, by inputting the target object information into the attribute prediction model, the first set of attribute information corresponding to the target object can be automatically determined. By matching keywords based on the correspondence between target keywords in the target object information and pre-configured keywords and attribute information, and / or by matching attribute expressions based on the correspondence between the target object information and pre-configured attribute expressions and attribute information, the second set of attribute information corresponding to the target object can be determined. The first and second sets of attribute information are merged and deduplicated to obtain a deduplicated set of target attribute information. Based on the correspondence between pre-configured object codes and attribute information sets, the object code corresponding to the target attribute information set is determined and used as the target code for the target object. By utilizing keyword matching and / or attribute expression matching, more and richer attribute information can be determined, thus enabling more accurate determination of the object code and further improving the accuracy of code determination.

[0112] For example, based on the pre-configured correspondence between keywords and attribute information and / or the correspondence between attribute expressions and attribute information, the target text information is processed to determine the second set of attribute information corresponding to the target object, which may include:

[0113] The system performs keyword recognition on the target text information and determines the third attribute information set corresponding to the target object based on the identified target keywords and the correspondence between pre-configured keywords and attribute information. It also checks whether the target text information satisfies each pre-configured attribute expression and determines the fourth attribute information set corresponding to the target object based on the satisfied target attribute expression and the correspondence between pre-configured attribute expressions and attribute information. Finally, it determines the second attribute information set corresponding to the target object based on the third attribute information set and / or the fourth attribute information set.

[0114] Specifically, based on the pre-configured correspondence between keywords and attribute information, the attribute information corresponding to each target keyword is determined and combined to obtain a third attribute information set. The target text information is then checked to see if it satisfies each pre-configured attribute expression. Based on the pre-configured correspondence between attribute expressions and attribute information, the attribute information corresponding to the satisfying target attribute expressions is determined and combined to obtain a fourth attribute information set. If only keyword matching exists, the third attribute information set can be directly determined as the second attribute information set. If only attribute expression matching exists, the fourth attribute information set can be directly determined as the third attribute information set. If both keyword matching and attribute expression matching exist, the third and fourth attribute information sets can be merged and deduplicated to obtain a deduplicated second attribute information set. By utilizing both keyword matching and attribute expression matching, attribute information can be accurately determined, thereby increasing the number of attribute information predictions and further improving the accuracy of encoding determination.

[0115] Figure 8 is a schematic diagram of the structure of an encoding determination device provided in an embodiment of the present disclosure. As shown in Figure 8, the device specifically includes: a target object information acquisition module 410, a target encoding method determination module 420, and a target encoding determination module 430.

[0116] The system includes a target object information acquisition module 410, which acquires target object information corresponding to the target object to be encoded; a target encoding method determination module 420, which determines the target encoding method corresponding to the target object based on the target category to which the target object belongs; and a target encoding determination module 430, which inputs the target object information into the target network model in the target encoding method and determines the target encoding corresponding to the target object based on the output of the target network model. The network models in different encoding methods correspond to different model classification granularities.

[0117] The technical solution provided in this disclosure obtains the target object information corresponding to the target object to be encoded, and determines the target encoding method adapted to the target object based on the target category to which the target object belongs; inputs the target object information into the target network model in the target encoding method, and determines the target code corresponding to the target object based on the output result of the target network model. In this way, the network model in different encoding methods corresponds to different model classification granularity, so the code of the target object can be determined more accurately by using the target network model adapted to the target object, realizing automatic code determination and improving the efficiency and accuracy of code determination.

[0118] Based on the above technical solution, the target encoding method includes a first encoding method or a second encoding method; wherein, the network model in the first encoding method is at least two classification models, and different classification models are used to determine different types of encoding; the network model in the second encoding method is an attribute prediction model, and the classification granularity of the attribute prediction model is finer than the classification granularity of the classification model.

[0119] Based on the above technical solution, when the target encoding method is the first encoding method, the target encoding determination module 430 includes:

[0120] The first information input unit is used to input the target object information into at least two classification models in the first encoding method to obtain at least two candidate codes;

[0121] A target common code determination unit is used to determine the target common code corresponding to the target object based on the at least two candidate codes;

[0122] The first encoding determination unit is used to adjust the candidate encoding based on the target common encoding to determine the target encoding corresponding to the target object.

[0123] Based on the above technical solutions, the at least two classification models include a first type coding classification model, a second type coding classification model, and a common coding classification model, and the at least two candidate codes include candidate first type codes, candidate second type codes, and candidate common codes;

[0124] The first encoding determination unit is specifically used to: adjust the encoding of the candidate first type encoding and the candidate second type encoding based on the target common encoding, and determine the target first type encoding and the target second type encoding corresponding to the target object; wherein, the common encoding part in the target first type encoding and the target second type encoding is the target common encoding.

[0125] Based on the above technical solutions, the target object information includes: target text information and target image information for describing the target object; each classification model includes: a text encoding sub-model, an image encoding sub-model, and a classification sub-model;

[0126] The first information input unit is specifically used for: for each classification model, inputting the target text information into the text encoding sub-model of the classification model for text encoding processing to obtain the target text feature information corresponding to the target object, wherein the target text feature information is used to reflect the semantic attribute information of the descriptive text corresponding to the target object; inputting the target image information into the image encoding sub-model of the classification model for text encoding processing to obtain the target image feature information corresponding to the target object, wherein the target image feature information is used to reflect the shape attribute information of the target object; and inputting the target text feature information and the target image feature information into the classification sub-model of the classification model for object classification processing to obtain the candidate code corresponding to the target object.

[0127] Based on the above technical solutions, during the training of each classification model, the initial network weight values ​​in the text encoding sub-model are obtained by pre-training based on the object text dataset, and the initial network weight values ​​in the image encoding sub-model are obtained by pre-training based on the object image dataset.

[0128] Based on the above technical solutions, each classification model is obtained by training the sample object information corresponding to the sample object as input information and the probability distribution corresponding to the sample code of the sample object as output label.

[0129] The probability distribution includes the probability value of each optional code being the actual code of the sample object, wherein the probability value of the sample code being the actual code of the sample object is the largest, and the optional code having the same code length as the sample code is positively correlated with the probability value.

[0130] Based on the above technical solutions, when the target encoding method is the second encoding method, the target encoding determination module 430 includes:

[0131] The second information input unit is used to input the target object information into the attribute prediction model in the second encoding method to obtain the first attribute information set;

[0132] The second encoding determination unit is used to determine the target encoding corresponding to the target object based on the first attribute information set and the pre-configured correspondence between the object encoding and the attribute information set.

[0133] Based on the above technical solutions, the target object information includes: target text information and target image information for describing the target object; the attribute prediction model includes: a text encoding sub-model, an image encoding sub-model, and an attribute prediction sub-model.

[0134] The second information input unit is specifically used for: inputting the target text information and the target image information into the text encoding sub-model for text encoding processing to obtain target text feature information corresponding to the target object, wherein the target text feature information is used to reflect the semantic attribute information of the descriptive text corresponding to the target object; inputting the target image information into the image encoding sub-model for text encoding processing to obtain target image feature information corresponding to the target object, wherein the target image feature information is used to reflect the shape attribute information of the target object; and inputting the target text feature information and the target image feature information into the attribute prediction sub-model for multi-head attention attribute prediction processing to obtain a first attribute information set corresponding to the target object.

[0135] Based on the above technical solutions, the attribute prediction sub-model achieves multi-head attention attribute prediction processing through the following steps:

[0136] Self-attention processing is performed on the first attribute feature vectors corresponding to all candidate attributes to obtain the second attribute feature vector, wherein the first attribute feature vector is obtained by training an attribute prediction model based on a preset attribute feature vector; cross-attention processing is performed on the second attribute feature vector, the input target text feature information, and the target image feature information to obtain the third attribute feature vector; attribute prediction processing is performed based on the third attribute feature vector to obtain the first attribute information set corresponding to the target object.

[0137] Based on the above technical solutions, the second encoding determination unit includes:

[0138] The second attribute information set determination subunit is used to process the target text information based on the pre-configured correspondence between keywords and attribute information and / or the correspondence between attribute expressions and attribute information, and to determine the second attribute information set corresponding to the target object.

[0139] The target attribute information set determination subunit is used to determine the target attribute information set corresponding to the target object based on the first attribute information set and the second attribute information set;

[0140] The target encoding determination subunit is used to determine the target encoding corresponding to the target object based on the target attribute information set and the pre-configured correspondence between the object encoding and the attribute information set.

[0141] Based on the above technical solutions, the second attribute information set determines the sub-unit, specifically for:

[0142] The target text information is subjected to keyword recognition, and based on the identified target keywords and the correspondence between pre-configured keywords and attribute information, a third attribute information set corresponding to the target object is determined; whether the target text information satisfies each pre-configured attribute expression is detected, and based on the satisfied target attribute expression and the correspondence between pre-configured attribute expression and attribute information, a fourth attribute information set corresponding to the target object is determined; based on the third attribute information set and / or the fourth attribute information set, a second attribute information set corresponding to the target object is determined.

[0143] Based on the above technical solutions, the target encoding method determination module 420 is specifically used for: if the target category to which the target object belongs is a first preset category corresponding to the first encoding method, then the target encoding method corresponding to the target object is determined to be the first encoding method; if the target category to which the target object belongs is a second preset category corresponding to the second encoding method, then the target encoding method corresponding to the target object is determined to be the second encoding method.

[0144] The encoding determination apparatus provided in this disclosure can execute the encoding determination method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the encoding determination method.

[0145] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0146] Figure 9 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Referring now to Figure 9, a schematic diagram of the structure of an electronic device (e.g., the terminal device or server in Figure 9) 500 suitable for implementing embodiments of this disclosure is shown. The terminal device in embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 9 is merely an example and should not impose any limitations on the functionality and scope of use of embodiments of this disclosure.

[0147] As shown in Figure 9, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.

[0148] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG9 shows electronic device 500 with various devices, it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0149] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0150] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0151] The electronic device provided in this embodiment and the encoding determination method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0152] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the encoding determination method provided in the above embodiments.

[0153] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0154] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0155] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0156] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire target object information corresponding to a target object to be encoded; determine the target encoding method corresponding to the target object based on the target category to which the target object belongs; input the target object information into the target network model in the target encoding method, and determine the target encoding corresponding to the target object based on the output result of the target network model; wherein, the network models in different encoding methods correspond to different model classification granularities.

[0157] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0158] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0159] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0160] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0161] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0162] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0163] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0164] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for determining an encoding, comprising: Obtain the target object information corresponding to the target object to be encoded; Based on the target category to which the target object belongs, determine the target encoding method corresponding to the target object; The target object information is input into the target network model in the target encoding method, and the target encoding corresponding to the target object is determined based on the output of the target network model. Among them, network models with different encoding methods correspond to different model classification granularities.

2. The encoding determination method according to claim 1, wherein, The target encoding method includes either a first encoding method or a second encoding method; In the first encoding method, the network model is at least two classification models, and different classification models are used to determine different types of encoding; The network model in the second encoding method is an attribute prediction model, and the classification granularity of the attribute prediction model is finer than that of the classification model.

3. The encoding determination method according to claim 2, wherein, When the target encoding method is the first encoding method, the step of inputting the target object information into the target network model in the target encoding method, and determining the target encoding corresponding to the target object based on the output of the target network model, includes: The target object information is input into at least two classification models in the first encoding method to obtain at least two candidate codes; Based on the at least two candidate codes, the target common code corresponding to the target object is determined; The candidate codes are adjusted based on the target common code to determine the target code corresponding to the target object.

4. The encoding determination method according to claim 3, wherein, The at least two classification models include a first type coding classification model, a second type coding classification model, and a common coding classification model; the at least two candidate codes include candidate first type codes, candidate second type codes, and candidate common codes. The step of adjusting the candidate codes based on the target common code to determine the target code corresponding to the target object includes: Based on the target common code, the code of the candidate first type code and the candidate second type code are adjusted to determine the target first type code and the target second type code corresponding to the target object; The common coding portion in the first type of target coding and the second type of target coding is the common coding of the target.

5. The encoding determination method according to claim 3 or 4, wherein, The target object information includes: target text information and target image information used to describe the target object; each classification model includes: a text encoding sub-model, an image encoding sub-model, and a classification sub-model. The target object information is input into at least two classification models in the first encoding method to obtain at least two candidate codes, including: For each classification model, the target text information is input into the text encoding sub-model in the classification model for text encoding processing to obtain the target text feature information corresponding to the target object. The target text feature information is used to reflect the semantic attribute information of the descriptive text corresponding to the target object. The target image information is input into the image encoding sub-model in the classification model for text encoding processing to obtain the target image feature information corresponding to the target object. The target image feature information is used to reflect the shape attribute information of the target object. The target text feature information and the target image feature information are input into the classification sub-model in the classification model for object classification processing to obtain the candidate code corresponding to the target object.

6. The encoding determination method according to claim 5, wherein, During the training of each classification model, the initial network weight values ​​in the text encoding sub-model are obtained by pre-training based on the object text dataset, and the initial network weight values ​​in the image encoding sub-model are obtained by pre-training based on the object image dataset.

7. The encoding determination method according to any one of claims 2-6, wherein, Each classification model is trained by using the sample object information corresponding to the sample object as input information and the probability distribution corresponding to the sample code of the sample object as output label. The probability distribution includes the probability value of each optional code being the actual code of the sample object, wherein the probability value of the sample code being the actual code of the sample object is the largest, and the optional code having the same code length as the sample code is positively correlated with the probability value.

8. The encoding determination method according to claim 2, wherein, When the target encoding method is the second encoding method, the step of inputting the target object information into the target network model in the target encoding method, and determining the target encoding corresponding to the target object based on the output of the target network model, includes: The target object information is input into the attribute prediction model in the second encoding method to obtain the first attribute information set; Based on the correspondence between the first set of attribute information and the pre-configured object code and attribute information set, the target code corresponding to the target object is determined.

9. The encoding determination method according to claim 8, wherein, The target object information includes: target text information and target image information used to describe the target object; the attribute prediction model includes: a text encoding sub-model, an image encoding sub-model, and an attribute prediction sub-model. The target object information is input into the attribute prediction model in the second encoding method to obtain a first attribute information set, including: The target text information and the target image information are input into the text encoding sub-model for text encoding processing to obtain the target text feature information corresponding to the target object. The target text feature information is used to reflect the semantic attribute information of the descriptive text corresponding to the target object. The target image information is input into the image encoding sub-model for text encoding processing to obtain the target image feature information corresponding to the target object. The target image feature information is used to reflect the shape attribute information of the target object. The target text feature information and the target image feature information are input into the attribute prediction sub-model for multi-head attention attribute prediction processing to obtain the first attribute information set corresponding to the target object.

10. The encoding determination method according to claim 9, wherein, The attribute prediction sub-model achieves attribute prediction processing for multi-head attention through the following steps: Self-attention processing is performed on the first attribute feature vectors corresponding to all candidate attributes to obtain the second attribute feature vectors. The first attribute feature vectors are obtained by training an attribute prediction model based on preset attribute feature vectors. The second attribute feature vector, along with the input target text feature information and target image feature information, are subjected to cross-attention processing to obtain the third attribute feature vector; Based on the third attribute feature vector, attribute prediction processing is performed to obtain the first attribute information set corresponding to the target object.

11. The encoding determination method according to any one of claims 8-10, wherein, The step of determining the target code corresponding to the target object based on the first attribute information set and the pre-configured correspondence between the object code and the attribute information set includes: Based on the pre-configured correspondence between keywords and attribute information and / or the correspondence between attribute expressions and attribute information, the target text information is processed to determine the second set of attribute information corresponding to the target object; Based on the first set of attribute information and the second set of attribute information, determine the target attribute information set corresponding to the target object; Based on the target attribute information set and the correspondence between the pre-configured object code and attribute information set, the target code corresponding to the target object is determined.

12. The encoding determination method according to claim 11, wherein, The process of processing the target text information based on the pre-configured correspondence between keywords and attribute information and / or the correspondence between attribute expressions and attribute information to determine the second set of attribute information corresponding to the target object includes: Keyword recognition is performed on the target text information, and based on the correspondence between the identified target keywords and pre-configured keywords and attribute information, the set of third attribute information corresponding to the target object is determined; Detect whether the target text information satisfies each pre-configured attribute expression, and determine the fourth attribute information set corresponding to the target object based on the satisfied target attribute expression and the correspondence between the pre-configured attribute expression and attribute information; Based on the third attribute information set and / or the fourth attribute information set, determine the second attribute information set corresponding to the target object.

13. The encoding determination method according to any one of claims 1-12, wherein, The step of determining the target encoding method corresponding to the target object based on the target category to which the target object belongs includes: If the target category to which the target object belongs is the first preset category corresponding to the first encoding method, then the target encoding method corresponding to the target object is determined to be the first encoding method; If the target category to which the target object belongs is the second preset category corresponding to the second encoding method, then the target encoding method corresponding to the target object is determined to be the second encoding method.

14. An encoding determination device, comprising: The target object information acquisition module is configured to acquire the target object information corresponding to the target object to be encoded. The target encoding method determination module is configured to determine the target encoding method corresponding to the target object based on the target category to which the target object belongs; The target encoding determination module is configured to input the target object information into the target network model in the target encoding method, and determine the target encoding corresponding to the target object based on the output of the target network model; wherein, the network models in different encoding methods correspond to different model classification granularities.

15. An electronic device comprising: One or more processors; Storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the encoding determination method as described in any one of claims 1-13.

16. A storage medium containing computer-executable instructions, wherein, The computer-executable instructions, when executed by a computer processor, are used to perform the encoding determination method as described in any one of claims 1-13.

Citation Information

Patent Citations

  • Logistics object information processing method and device and computer system

    CN110858219A

  • Code generation method and device, storage medium and electronic equipment

    CN113031943A

  • Data coding method and device, computer equipment and storage medium

    CN113792816A

  • Method for recommending HS codes for goods, electronic equipment and medium

    CN114548041A

  • Substation BIM classification coding method and device, electronic equipment and storage medium

    CN115130603A