Data object processing method, electronic device, and storage medium

By determining the feature representation of triplet object samples in the data analyzer and updating the parameters based on the weights of intra-class and inter-class matching information, the problem of slow convergence speed of the data analyzer is solved, achieving high efficiency and accuracy of HS encoding classification and reducing training costs.

CN116881761BActive Publication Date: 2026-06-02ZHEJIANG CAINIAO SUPPLY CHAIN MANAGEMENT CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG CAINIAO SUPPLY CHAIN MANAGEMENT CO LTD
Filing Date
2022-03-28
Publication Date
2026-06-02

Smart Images

  • Figure CN116881761B_ABST
    Figure CN116881761B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data object processing method, an electronic device and a storage medium, wherein the method comprises: determining a triple object sample; determining a feature representation corresponding to the triple object sample by using a data analyzer; the data analyzer is used to represent a mapping relationship between object information and the feature representation; the feature representation is used to determine target feature information corresponding to the data object; determining first matching information between a first object sample and a second object sample, and second matching information between the first object sample and a third object sample according to the feature representation corresponding to the triple object sample; updating parameters of the data analyzer according to first weight information corresponding to the first matching information and second weight information corresponding to the second matching information. Embodiments of the present application can improve the convergence speed of the data analyzer, and thus can save the cost spent in training the data analyzer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a method for processing data objects, an electronic device, and a storage medium. Background Technology

[0002] With the development of communication technology, the classification of goods and other customs clearance objects has become increasingly important. The World Customs Organization has developed the HS (Harmonized System) coding system, which uses numerical codes to represent and identify goods in cross-border trade. HS code classification is the process of finding the HS code for a product to be classified based on its information. As a universal category identifier for goods clearing customs, the HS code is the fundamental basis for customs to conduct commodity classification management, review tax standards, and inspect commodity quality indicators. Inconsistencies between the declared HS code and the actual category of the goods can cause a series of quality issues related to commodity management models, tax collection, the application of inspection standards, billing, statistics, and other related services.

[0003] Current HS coding classification methods typically use data analyzers to determine the feature representation of the customs clearance object, and then determine the feature information corresponding to the customs clearance object based on the feature representation of the customs clearance object. Summary of the Invention

[0004] This application provides a data object processing method that can improve the convergence speed of a data analyzer, thereby saving the cost of training the data analyzer.

[0005] Correspondingly, embodiments of this application also provide a data object processing device, an electronic device, and a storage medium to implement and apply the above-described method.

[0006] To address the aforementioned problems, this application discloses a method for processing data objects, the method comprising:

[0007] Determine triplet object samples; the triplet object samples include: a first object sample, a second object sample, and a third object sample; wherein, the second object sample corresponds to the same category information as the first object sample; and the third object sample corresponds to different category information than the first object sample.

[0008] A data analyzer is used to determine the feature representation corresponding to the triplet object sample; the data analyzer is used to characterize the mapping relationship between object information and feature representation; the feature representation is used to determine the target feature information corresponding to the data object.

[0009] Based on the feature representations corresponding to the triplet object samples, determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample;

[0010] The parameters of the data analyzer are updated based on the first weight information corresponding to the first matching information and the second weight information corresponding to the second matching information; the first weight information is obtained based on the first matching information and the first preset matching information, and the second weight information is obtained based on the second matching information and the second preset matching information.

[0011] To address the aforementioned problems, this application discloses a method for processing data objects, the method comprising:

[0012] Determine triplet object samples; the triplet object samples include: a first object sample, a second object sample, and a third object sample; wherein, the category information corresponding to the triplet object samples includes: N levels of subcategory information; the N levels of subcategory information corresponding to the second object sample and the first object sample are the same; the third object sample and the first object sample have M different levels of subcategory information; N is a positive integer, and M is a positive integer not greater than N;

[0013] A data analyzer is used to determine the feature representation corresponding to the triplet object sample; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0014] Based on the feature representations corresponding to the triplet object samples, determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample;

[0015] The parameters of the data analyzer are updated based on the mapping relationship between the loss information and the first matching information, the second matching information and their corresponding category parameter information; the category parameters are obtained based on the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

[0016] To address the aforementioned problems, this application discloses a method for processing data objects, the method comprising:

[0017] A data analyzer is used to determine the first feature representation corresponding to the data object; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0018] Based on the first feature representation corresponding to the data object and the second feature representation corresponding to the object sample, a target object sample matching the data object is determined; the second feature representation is obtained from the data analyzer.

[0019] Based on the feature information corresponding to the target object sample, determine the target feature information corresponding to the data object;

[0020] The training samples corresponding to the data analyzer include: triplet object samples, which include: a first object sample, a second object sample, and a third object sample; the second object sample corresponds to the same category information as the first object sample; the third object sample corresponds to different category information than the first object sample; the parameters of the data analyzer are obtained based on the first weight information corresponding to the first matching information and the second weight information corresponding to the second matching information; the first matching information corresponds to the matching information between the first object sample and the second object sample; the second matching information corresponds to the matching information between the first object sample and the third object sample; the first weight information is obtained based on the first matching information and the first preset matching information; the second weight information is obtained based on the second matching information and the second preset matching information.

[0021] To address the aforementioned problems, this application discloses a method for processing data objects, the method comprising:

[0022] A data analyzer is used to determine the first feature representation corresponding to the data object; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0023] Based on the first feature representation corresponding to the data object and the second feature representation corresponding to the object sample, a target object sample matching the data object is determined; the second feature representation is obtained from the data analyzer.

[0024] Based on the feature information corresponding to the target object sample, determine the target feature information corresponding to the data object;

[0025] The training samples corresponding to the data analyzer include: triplet object samples, which include: a first object sample, a second object sample, and a third object sample; the category information corresponding to the triplet object samples includes: N levels of sub-category information; the N levels of sub-category information corresponding to the second object sample and the first object sample are the same; the third object sample and the first object sample have M different levels of sub-category information; N is a positive integer, and M is a positive integer not greater than N; the parameters of the data analyzer are obtained according to the mapping relationship between loss information and first matching information, and second matching information and their corresponding category parameter information; the first matching information corresponds to the matching information between the first object sample and the second object sample; the second matching information corresponds to the matching information between the first object sample and the third object sample; the category parameters are obtained according to the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

[0026] To address the aforementioned problems, this application discloses a data object processing apparatus, the apparatus comprising:

[0027] A sample determination module is used to determine triplet object samples; the triplet object samples include: a first object sample, a second object sample, and a third object sample; wherein, the category information corresponding to the triplet object samples includes: N levels of subcategory information; the N levels of subcategory information corresponding to the second object sample and the first object sample are the same; the third object sample and the first object sample have M different levels of subcategory information; N is a positive integer, and M is a positive integer not greater than N;

[0028] The feature representation determination module is used to determine the feature representation corresponding to the triplet object sample using a data analyzer; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0029] The matching information determination module is used to determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample, based on the feature representation corresponding to the triplet object sample.

[0030] The parameter update module is used to update the parameters of the data analyzer according to the mapping relationship between the loss information and the first matching information, the second matching information and their corresponding category parameter information; the category parameters are obtained according to the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

[0031] To address the aforementioned problems, this application discloses a data object processing apparatus, the apparatus comprising:

[0032] A sample determination module is used to determine triplet object samples; the triplet object samples include: a first object sample, a second object sample, and a third object sample; wherein, the category information corresponding to the triplet object samples includes: N levels of subcategory information; the N levels of subcategory information corresponding to the second object sample and the first object sample are the same; the third object sample and the first object sample have M different levels of subcategory information; N is a positive integer, and M is a positive integer not greater than N;

[0033] The feature representation determination module is used to determine the feature representation corresponding to the triplet object sample using a data analyzer; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0034] The matching information determination module is used to determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample, based on the feature representation corresponding to the triplet object sample.

[0035] The parameter update module is used to update the parameters of the data analyzer according to the mapping relationship between the loss information and the first matching information, the second matching information and their corresponding category parameter information; the category parameters are obtained according to the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

[0036] To address the aforementioned problems, this application discloses a data object processing apparatus, the apparatus comprising:

[0037] The feature representation determination module is used to determine the first feature representation corresponding to the data object using a data analyzer; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0038] The target object sample determination module is used to determine a target object sample that matches the data object based on a first feature representation corresponding to the data object and a second feature representation corresponding to the object sample; the second feature representation is obtained from the data analyzer.

[0039] The feature determination module is used to determine the target feature information corresponding to the data object based on the feature information corresponding to the target object sample.

[0040] The training samples corresponding to the data analyzer include: triplet object samples, which include: a first object sample, a second object sample, and a third object sample; the second object sample corresponds to the same category information as the first object sample; the third object sample corresponds to different category information than the first object sample; the parameters of the data analyzer are obtained based on the first weight information corresponding to the first matching information and the second weight information corresponding to the second matching information; the first matching information corresponds to the matching information between the first object sample and the second object sample; the second matching information corresponds to the matching information between the first object sample and the third object sample; the first weight information is obtained based on the first matching information and the first preset matching information; the second weight information is obtained based on the second matching information and the second preset matching information.

[0041] To address the aforementioned problems, this application discloses a data object processing apparatus, the apparatus comprising:

[0042] The feature representation determination module is used to determine the first feature representation corresponding to the data object using a data analyzer; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0043] The target object sample determination module is used to determine a target object sample that matches the data object based on a first feature representation corresponding to the data object and a second feature representation corresponding to the object sample; the second feature representation is obtained from the data analyzer.

[0044] The feature determination module is used to determine the target feature information corresponding to the data object based on the feature information corresponding to the target object sample.

[0045] The training samples corresponding to the data analyzer include: triplet object samples, which include: a first object sample, a second object sample, and a third object sample; the category information corresponding to the triplet object samples includes: N levels of sub-category information; the N levels of sub-category information corresponding to the second object sample and the first object sample are the same; the third object sample and the first object sample have M different levels of sub-category information; N is a positive integer, and M is a positive integer not greater than N; the parameters of the data analyzer are obtained according to the mapping relationship between loss information and first matching information, and second matching information and their corresponding category parameter information; the first matching information corresponds to the matching information between the first object sample and the second object sample; the second matching information corresponds to the matching information between the first object sample and the third object sample; the category parameters are obtained according to the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

[0046] To address the aforementioned problems, this application discloses an electronic device, including: a processor; and a memory storing executable code thereon, wherein when the executable code is executed, the processor performs the method as described in any of the above embodiments.

[0047] To address the aforementioned issues, embodiments of this application disclose one or more machine-readable media storing executable code thereon, which, when executed, causes a processor to perform the method as described in any of the above embodiments.

[0048] Compared with the prior art, the embodiments of this application have the following advantages:

[0049] In the technical solution of this application embodiment, the parameters of the data analyzer are updated according to the first weight information and the second weight information. Since the first weight information reflects the degree of difference between the actual intra-class matching information and the intra-class optimization target in the intra-class dimension, and the second weight information reflects the degree of difference between the actual inter-class matching information and the inter-class optimization target in the inter-class dimension, this application embodiment can perform targeted processing in the intra-class and inter-class dimensions respectively based on the first and second weight information during the parameter update process, so as to achieve the corresponding optimization targets in the intra-class and inter-class dimensions respectively. Therefore, this application embodiment can improve the convergence speed of the data analyzer, thereby saving the cost of training the data analyzer. Attached Figure Description

[0050] Figure 1 This is a schematic diagram of the data object processing flow according to an embodiment of this application;

[0051] Figure 2 This is a flowchart of the steps of a data object processing method according to an embodiment of this application;

[0052] Figure 3 This is an example of encoded information from one embodiment of this application;

[0053] Figure 4 This is a comparison of one embodiment of the present application with the case where the first weight information and the second weight information are not used, and the case where the first weight information and the second weight information are used;

[0054] Figure 5 This is a flowchart of the steps of a data object processing method according to an embodiment of this application;

[0055] Figure 6 This is a flowchart of the steps of a data object processing method according to an embodiment of this application;

[0056] Figure 7This is a flowchart of the steps of a data object processing method according to an embodiment of this application;

[0057] Figure 8 This is a flowchart of the steps of a data object processing method according to an embodiment of this application;

[0058] Figure 9 This is a schematic diagram of the structure of a data object processing apparatus according to an embodiment of this application;

[0059] Figure 10 This is a schematic diagram of the structure of a data object processing apparatus according to an embodiment of this application;

[0060] Figure 11 This is a schematic diagram of the structure of a data object processing apparatus according to an embodiment of this application;

[0061] Figure 12 This is a schematic diagram of the structure of a data object processing apparatus according to an embodiment of this application;

[0062] Figure 13 This is a schematic diagram of the structure of an exemplary device provided in one embodiment of this application. Detailed Implementation

[0063] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0064] In this embodiment, the data object can be a composite information representation understood by the software. The data object can be an entity, thing, incidental event or event, role, organizational unit, location, or structure, etc. For example, the data object may include customs clearance objects such as goods or merchandise.

[0065] Reference Figure 1 The diagram illustrates a data object processing flow according to an embodiment of this application. In this flow, object information of the data object can be input into a data analyzer, which outputs a feature representation corresponding to the object information. The feature representation can be input into a preset processor, which performs preset processing on the feature representation to obtain a corresponding processing result.

[0066] Data analyzers can possess feature extraction capabilities, which can be used to characterize the mapping relationship between object information and feature representations. Object information can be shallow, typically presented in text form. Feature representations can be deep, typically presented in vector form. Taking a product as an example, object information can include attribute information such as category, material, and content.

[0067] This application embodiment can train a mathematical model based on training samples to obtain a data analyzer. A mathematical model is a scientific or engineering model constructed using mathematical logic methods and mathematical language. It is a mathematical structure that, using mathematical language, summarizes or approximates the characteristics or quantitative dependencies of a system referring to something. This mathematical structure is a relational structure characterized by mathematical symbols. A mathematical model can be one or a set of algebraic equations, differential equations, difference equations, integral equations, or statistical equations, or combinations thereof, which quantitatively or qualitatively describe the interrelationships or causal relationships between the variables of the system. Besides mathematical models described by equations, there are also models described using other mathematical tools, such as algebra, geometry, topology, and mathematical logic. In these cases, the mathematical model describes the behavior and characteristics of the system rather than its actual structure. Among them, machine learning and deep learning methods can be used to train mathematical models. Machine learning methods can include linear regression, decision trees, random forests, etc., while deep learning methods can include CNN (Convolutional Neural Networks), LSTM (Long Short-Term Memory), GRU (Gated Recurrent Unit), etc.

[0068] The mathematical model corresponding to the data analyzer in this application embodiment may include a model with feature extraction capabilities. For example, the mathematical model corresponding to the data analyzer in this application embodiment may include, but is not limited to, a pre-trained language model. Examples of pre-trained language models may include BERT (Bidirectional Encoder Representation from Transformers), ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacement Accurately), etc. Of course, the mathematical model corresponding to the data analyzer may also include other mathematical models besides pre-trained language models, such as Transformer, RNN, CNN, etc.

[0069] A preset processor can be used to perform preset processing on feature representations. Those skilled in the art can determine the preset processing method information corresponding to the preset processor based on actual application information. For example, the preset processing method information may include: determining whether a data object is a target; the target category may include: biometric information such as face, fingerprint, or voiceprint, for example, determining whether an image contains a target face. Alternatively, the preset processing method information may include: determining the target feature information corresponding to the data object. The target feature information may include: text corresponding to an image, or feature information corresponding to a product. For example, in the field of optical character recognition, preset processing can be used to determine the text corresponding to an image. As another example, in the field of HS coding classification, preset processing can be used to determine the coding information corresponding to a product. It is understood that the embodiments of this application do not limit the specific information of the preset processing.

[0070] In practical applications, data analyzers suffer from slow convergence speed. To address this issue, this application provides a data object processing scheme, which specifically includes: determining triplet object samples; the triplet object samples may include: a first object sample, a second object sample, and a third object sample; wherein the second object sample corresponds to the same category information as the first object sample; and the third object sample corresponds to different category information than the first object sample; using a data analyzer, determining the feature representation corresponding to the triplet object samples; the feature representation is used to determine the target feature information corresponding to the data object; the data analyzer is used to characterize the mapping relationship between object information and feature representation; based on the feature representation corresponding to the triplet object samples, determining the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample; updating the parameters of the data analyzer based on the first weight information corresponding to the first matching information and the second weight information corresponding to the second matching information; the first weight information can be obtained based on the first matching information and a first preset matching information, and the second weight information can be obtained based on the second matching information and a second preset matching information.

[0071] In the triplet object samples of this application embodiment, the first object sample and the second object sample correspond to the same category information, so the first matching information can characterize intra-class matching information; the first object sample and the third object sample correspond to different category information, so the second matching information can characterize inter-class matching information.

[0072] Since the first weight information is obtained based on the first matching information and the first preset matching information, and the first preset matching information can characterize the intra-class optimization objective, the first weight information can reflect the degree of difference between the actual intra-class matching information and the intra-class optimization objective. Similarly, since the second weight information is obtained based on the second matching information and the second preset matching information, and the second preset matching information can characterize the inter-class optimization objective, the second weight information can reflect the degree of difference between the actual inter-class matching information and the inter-class optimization objective.

[0073] This application embodiment updates the parameters of the data analyzer based on first weight information and second weight information. Since the first weight information reflects the degree of difference between the actual intra-class matching information and the intra-class optimization target in the intra-class dimension, and the second weight information reflects the degree of difference between the actual inter-class matching information and the inter-class optimization target in the inter-class dimension, this application embodiment can perform targeted processing in the intra-class and inter-class dimensions respectively based on the first and second weight information during the parameter update process, so as to achieve the corresponding optimization targets in the intra-class and inter-class dimensions respectively. Therefore, this application embodiment can improve the convergence speed of the data analyzer, thereby saving the cost of training the data analyzer.

[0074] Method Example 1

[0075] Reference Figure 2 The flowchart illustrates the steps of a data object processing method according to an embodiment of this application, which may specifically include the following steps:

[0076] Step 201: Determine the triplet object samples; the triplet object samples may include: a first object sample, a second object sample, and a third object sample; wherein, the second object sample and the first object sample may correspond to the same category information; the third object sample and the first object sample may correspond to different category information;

[0077] Step 202: Using a data analyzer, determine the feature representation corresponding to the triplet object sample; the data analyzer can be used to characterize the mapping relationship between object information and feature representation; the feature representation can be used to determine the target feature information corresponding to the data object;

[0078] Step 203: Based on the feature representations corresponding to the triplet object samples, determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample;

[0079] Step 204: Update the parameters of the data analyzer according to the first weight information corresponding to the first matching information and the second weight information corresponding to the second matching information; the first weight information can be obtained according to the first matching information and the first preset matching information, and the second weight information can be obtained according to the second matching information and the second preset matching information.

[0080] The embodiments of this application can be used to update the parameters of the data analyzer during the training process of the data analyzer, so as to improve the convergence speed of the data analyzer and thus save the cost of training the data analyzer.

[0081] The training process for a data analyzer can include forward propagation and backward propagation.

[0082] The forward propagation process calculates the final output information sequentially from the input layer to the output layer based on the parameters of the data analyzer. This output information can then be used to determine the loss information.

[0083] Backpropagation calculates and updates the parameters of a data analyzer sequentially from the output layer to the input layer based on the loss information. During backpropagation, the gradient information of the data analyzer parameters is determined and used to update the parameters. For example, backpropagation can use the chain rule from calculus to calculate and store the gradient information of the parameters of the data analyzer's processing layers (including the input layer, intermediate layers, and output layer) sequentially from the output layer to the input layer.

[0084] In step 201, a first object sample can be determined first, then a second object sample of the same class as the first object sample and a third object sample of a different class from the first object sample can be determined.

[0085] In practical applications, an object sample set can be constructed, which may include multiple labeled object samples.

[0086] Taking goods and other customs clearance objects as an example, the object sample can correspond to characteristic information, which can be coding information. HS codes can include 22 categories and 98 chapters. Internationally accepted HS codes can include the first two digits, the third and fourth digits, and the fifth and sixth digits. The seventh digit and subsequent digits can be determined by the country. For example, Chinese customs uses a ten-digit HS coding system, which can include 6 digits of coding information corresponding to international standards and 4 digits of coding information corresponding to national standards.

[0087] The encoding information in this application embodiment may include: 6-bit encoding or encoding of more than 6 bits. It is understood that this application embodiment does not limit the number of digits corresponding to the encoding information.

[0088] Reference Figure 3 The illustration shows an example of encoded information according to an embodiment of this application. The encoded information, in order from front to back, may include: 6-digit encoded information corresponding to international standards and 4-digit encoded information corresponding to national standards. The 6-digit encoded information corresponding to international standards may include: 2-digit encoded information corresponding to chapters, 2-digit encoded information corresponding to tariff headings, and 2-digit encoded information corresponding to subheadings.

[0089] The encoding information corresponding to the object sample in this application embodiment may include: 6-bit encoding information; in this case, this application embodiment can determine the 6-bit encoding information corresponding to the data object under the international standard based on the data analyzer and the preset processor. Of course, the encoding information corresponding to the object sample in this application embodiment may include: 10-bit encoding information; in this case, this application embodiment can determine the 10-bit encoding information corresponding to the data object under both the international standard and the Chinese standard based on the data analyzer and the preset processor. Furthermore, under the standards of other countries besides China, the number of digits corresponding to the encoding information may not be equal to 10.

[0090] In practical applications, a corresponding 6-digit code can be marked for customs clearance goods, and / or the 6-digit code can be extracted from the customs filing data; and the customs clearance goods corresponding to the above 6-digit code can be saved to the object sample set.

[0091] In a specific implementation, a first object sample A can be obtained from the object sample set, and for the first object sample A, a second object sample B and a third object sample C can be obtained from the object sample set.

[0092] Suppose that the category information corresponding to the triplet object sample (the first object sample) includes N levels of subcategory information; the second object sample and the first object sample have the same N levels of subcategory information; the third object sample and the first object sample have M different levels of subcategory information; where N can be a positive integer, and M can be a positive integer not greater than N. The method of determining the third object sample can take into account the granularity of multiple levels of subcategories and improve the diversity of the third object sample.

[0093] Assuming N=3, the third object sample C can include at least one of the following: having one different level of subcategory information from the first object sample A, having two different levels of subcategory information from the first object sample A, and having three different levels of subcategory information from the first object sample A. The method of determining the third object sample can take into account the granularity of subcategories at multiple levels, such as chapters, tax items, and subheadings, thereby improving the diversity of the third object sample.

[0094] by Figure 3 Taking the encoded information shown as an example, assuming that the category information corresponding to the encoded information of the first object sample A includes the following three levels of sub-category information: chapter, tax item, and sub-item, then for the first object sample A, under the category information corresponding to its "chapter", "tax item", and "sub-item", P second object samples B can be randomly sampled. P can be a positive integer and can be a multiple of 3.

[0095] by Figure 3 Taking the encoded information shown as an example, assuming that the category information corresponding to the encoded information of the first object sample A includes the following three levels of sub-category information: chapter, tax item, and sub-item, then for the first object sample A, under the category information corresponding to its respective "chapter" and "tax item" and different "sub-items", Q third object samples C can be randomly sampled; and under the category information corresponding to its respective "chapter" and different "tax items", Q third object samples C can be randomly sampled; and under the category information corresponding to different "chapters", Q third object samples C can be randomly sampled; Q can be a positive integer, and P can be 3 times Q, thereby achieving the matching of the number of the second object sample B and the third object sample C.

[0096] The triplet object sample in this application embodiment can be represented as: (A, B1, C1), (A, B2, C2)...(A, B... P C P ).

[0097] In step 202, the data analyzer may have corresponding parameters, which can be updated during the training process. Assuming the i-th training process corresponds to the i-th parameter, then during the i-th training process, the feature representation corresponding to the triplet object sample can be determined based on the i-th parameter. i can be a positive integer, and the i-th parameter can be an initial parameter, which can be a preset parameter.

[0098] The input to the data analyzer may include object information of the object sample, such as attribute information of the data object like category, material, and content. The output of the data analyzer may include a feature representation of the object sample, which can be presented in vector form.

[0099] The triplet object sample can include: a first object sample, a second object sample, and a third object sample. Then, the data analyzer can be used to determine the feature representation A corresponding to the first object sample, the feature representation B corresponding to the second object sample, and the feature representation C corresponding to the third object sample.

[0100] In step 203, a metric method can be used to determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample. The metric method may include Euclidean distance, cosine similarity, or information entropy, etc. It is understood that the embodiments of this application do not limit the specific metric method.

[0101] The first object sample and the second object sample correspond to the same category information, so the first matching information can represent intra-class matching information, and the first matching information can be denoted as sp. The first object sample and the third object sample correspond to different category information, so the second matching information can represent inter-class matching information, and the second matching information can be denoted as sn.

[0102] In step 204, the parameters of the data analyzer can be updated based on the first weight information and the second weight information to optimize the parameters of the data analyzer. Optimization methods may include gradient descent, Newton's method, quasi-Newton method, conjugate gradient method, etc. It is understood that the embodiments of this application do not limit the specific optimization methods.

[0103] In practical implementation, loss parameters can be determined. Optimization refers to the process of changing the parameters of the data analyzer to achieve the limit of the loss parameters. The loss function can be seen as the mapping relationship between loss information and the parameters of the data analyzer. This loss function can characterize the difference between inter-class dimensional information and intra-class dimensional information. In practical applications, the parameters of the data analyzer can be changed to achieve the lower limit of the difference information. The inter-class dimensional information can be obtained based on the second matching information, and the intra-class dimensional information can be obtained based on the first matching information. Achieving the optimization goal of the lower limit of the difference information can be achieved by increasing the first matching information and / or decreasing the second matching information.

[0104] The first weight information can reflect the degree of difference between the actual intra-class matching information and the intra-class optimization target in the intra-class dimension, and the second weight information can reflect the degree of difference between the actual inter-class matching information and the inter-class optimization target in the inter-class dimension. Therefore, in the process of updating the parameters of the data analyzer, the embodiments of this application can perform targeted processing in the intra-class dimension and the inter-class dimension according to the first weight information and the second weight information, respectively, so as to achieve the corresponding optimization targets in the intra-class dimension and the inter-class dimension. Therefore, the embodiments of this application can improve the convergence speed of the data analyzer.

[0105] Reference Figure 4 This illustrates a comparison between an embodiment of the present application that does not use the first weight information and the second weight information, and that uses the first weight information and the second weight information.

[0106] like Figure 4 As shown in case (A), without using the first and second weighting information, both sn and sp at point A are relatively small. In this case, to achieve the optimization objective, the parameters need to be adjusted according to direction D1. The increase in sp and the decrease in sn are balanced by direction D1. Since sn is already small enough, the adjustment in case (a) may lead to over-updating, causing it to deviate far from the preset solution (the preset solution corresponding to point A in the figure is T'), resulting in oscillations around the preset solution. The preset solution can refer to the values ​​of sn and sp when the loss function reaches its limit.

[0107] like Figure 4 As shown in case (B), when using the first and second weight information, both sn and sp corresponding to point A are relatively small. In this case, the parameters are adjusted according to direction D2. Direction D2 is obtained based on the first and second weight information. The increase of direction D2 in sp is obtained based on the first weight information, and the decrease of direction D2 in sn is obtained based on the second weight information. In this way, the limitation of balancing sp and sn can be overcome.

[0108] Taking the adjustment of sn as an example, the second weight information can be obtained based on the second matching information and the second preset matching information. The second preset matching information can represent the ideal inter-class matching information, which can be preset by the user. The second matching information can represent the actual inter-class matching information, that is, the inter-class matching information obtained under the conditions of the data analyzer's parameters. The second weight information can be the difference information between the second matching information and the second preset matching information, which can represent the inter-class difference information between the actual inter-class matching information and the ideal inter-class matching information. According to the inter-class difference information, the parameters of the data analyzer are updated so that the parameter update information matches the inter-class difference information. For example, if the inter-class difference information represents a small inter-class difference, the parameter update magnitude can be small; if the inter-class difference information represents a large inter-class difference, the parameter update magnitude can be large. Therefore, the embodiments of this application can, to a certain extent, avoid over-updating to the point of deviating from the preset solution, thereby causing oscillations around the preset solution.

[0109] Taking the adjustment of sp as an example, the first weight information can be obtained based on the first matching information and the first preset matching information. The first preset matching information can represent the ideal intra-class matching information, which can be preset by the user. The first matching information can represent the actual intra-class matching information, that is, the intra-class matching information obtained under the conditions of the data analyzer's parameters. The first weight information can be the difference information between the first matching information and the first preset matching information, which can represent the intra-class difference information between the actual intra-class matching information and the ideal intra-class matching information. Based on this intra-class difference information, the parameters of the data analyzer are updated so that the parameter update information matches the intra-class difference information. For example, if the intra-class difference information represents a small intra-class difference, the parameter update magnitude can be small; conversely, if the intra-class difference information represents a large intra-class difference, the parameter update magnitude can be large. Therefore, the embodiments of this application can, to a certain extent, avoid over-updating to the point of deviating from the preset solution, thereby causing oscillations around the preset solution.

[0110] In practical implementation, the parameter update information for intra-class dimensions can be determined based on the first weight information. For example, when using gradient descent, the step size information for intra-class dimensions can be determined based on the first weight information. Similarly, the parameter update information for inter-class dimensions can be determined based on the second weight information. For example, when using gradient descent, the step size information for inter-class dimensions can be determined based on the second weight information.

[0111] The step size determines the length of each step taken along the negative gradient direction during gradient descent iterations. In practical applications, partial derivatives of the loss function's parameters can be calculated, and these partial derivatives can be expressed as a vector. This vector represents the gradient information for the parameters. Based on the gradient and step size information, the parameter update can be obtained.

[0112] When using gradient descent, various methods can be employed, such as batch gradient descent, stochastic gradient descent, or mini-batch gradient descent. In practice, iteration can be performed based on a single triplet object sample corresponding to a first object sample A, or iteratively based on multiple triplet object samples corresponding to a single first object sample A, or multiple triplet object samples corresponding to multiple first object samples A. The convergence condition for these iterations can be that the loss information corresponding to the loss function meets the convergence condition. Alternatively, the convergence condition can be that the loss value corresponding to the loss information is less than a preset loss value, or the number of iterations exceeds a threshold. In other words, the iteration can terminate when the loss information corresponding to the loss function meets the convergence condition; in this case, the target parameters of the data analyzer can be obtained, which can be used in the data object processing.

[0113] In summary, the data object processing method of this application updates the parameters of the data analyzer based on the first weight information and the second weight information. Since the first weight information reflects the degree of difference between the actual intra-class matching information and the intra-class optimization target in the intra-class dimension, and the second weight information reflects the degree of difference between the actual inter-class matching information and the inter-class optimization target in the inter-class dimension, this application can perform targeted processing in the intra-class and inter-class dimensions respectively based on the first and second weight information during the parameter update process, thereby achieving the corresponding optimization targets in the intra-class and inter-class dimensions. Therefore, this application can improve the convergence speed of the data analyzer, thereby saving the cost of training the data analyzer.

[0114] Method Example 2

[0115] Reference Figure 5 The flowchart illustrates the steps of a data object processing method according to an embodiment of this application, which may specifically include the following steps:

[0116] Step 501: Determine the triplet object samples; the triplet object samples may include: a first object sample, a second object sample, and a third object sample; wherein, the second object sample and the first object sample may correspond to the same category information; the third object sample and the first object sample may correspond to different category information;

[0117] Step 502: Using a data analyzer, determine the feature representation corresponding to the triplet object sample; the data analyzer can be used to characterize the mapping relationship between object information and feature representation; the feature representation can be used to determine the target feature information corresponding to the data object;

[0118] Step 503: Based on the feature representations corresponding to the triplet object samples, determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample;

[0119] Step 504: Update the parameters of the data analyzer according to the first mapping relationship between the loss information, the first matching information and its corresponding first weight information, and the second matching information and its corresponding second weight information.

[0120] The loss function in this application embodiment can be a first mapping relationship between loss information and first matching information and its corresponding first weight information, and second matching information and its corresponding second weight information. Since this first mapping relationship utilizes both the first and second weight information, and the first weight information reflects the degree of difference in the intra-class dimension between the actual intra-class matching information and the intra-class optimization target, and the second weight information reflects the degree of difference in the inter-class dimension between the actual inter-class matching information and the inter-class optimization target, this application embodiment can obtain more accurate loss information based on the first mapping relationship. In other words, this application embodiment can improve the accuracy of the loss information. By improving the accuracy of the loss information, this application embodiment can improve the convergence speed of the data analyzer and enhance the feature discrimination capability of the data analyzer.

[0121] In practical applications, this loss function can characterize the difference between inter-class and intra-class dimensional information. The parameters of the data analyzer can be changed to achieve a lower bound on the difference. Inter-class dimensional information can be obtained based on the second matching information and its corresponding second weight information; for example, it can be the product of the second matching information and the second weight information. Intra-class dimensional information can be obtained based on the first matching information and its corresponding first weight information; for example, it can be the product of the first matching information and the first weight information. Exponential information can be obtained based on the difference between inter-class and intra-class dimensional information; for example, it can be obtained based on the difference between inter-class and intra-class dimensional information and the first parameter. Furthermore, an exponential formula can be obtained based on the exponential information and the base information. Further, a logarithmic operation can be performed on the sum of the exponential formula and the second parameter to obtain a logarithmic formula, the result of which can be considered as loss information.

[0122] In practical implementation, both the first and second matching information can correspond to the parameters of the data analyzer. Therefore, this first mapping relationship can be considered a mapping between the loss information and the parameters of the data analyzer. In practical applications, the partial derivatives of the parameters of this first mapping relationship can be calculated, and based on the partial derivatives, the optimization objective (in this case, the parameters can be called target parameters) can be obtained while achieving the lower limit of the loss information. In practical applications, the parameters can be adjusted according to the value of the loss information until the loss information meets the convergence condition.

[0123] In summary, the data object processing method of this application embodiment can obtain more accurate loss information based on the first mapping relationship; in other words, this application embodiment can improve the accuracy of the loss information. By improving the accuracy of the loss information, this application embodiment can increase the convergence speed of the data analyzer and enhance the feature discrimination capability of the data analyzer.

[0124] Method Example 3

[0125] Reference Figure 6 The flowchart illustrates the steps of a data object processing method according to an embodiment of this application, which may specifically include the following steps:

[0126] Step 601: Determine the triplet object samples; the triplet object samples may include: a first object sample, a second object sample, and a third object sample; wherein, the second object sample and the first object sample may correspond to the same category information; the third object sample and the first object sample may correspond to different category information;

[0127] Step 602: Using a data analyzer, determine the feature representation corresponding to the triplet object sample; the data analyzer can be used to characterize the mapping relationship between object information and feature representation; the feature representation can be used to determine the target feature information corresponding to the data object;

[0128] Step 603: Based on the feature representations corresponding to the triplet object samples, determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample;

[0129] Step 604: Update the parameters of the data analyzer according to the second mapping relationship between the loss information and the first matching information, the second matching information and their corresponding category parameter information; the category parameters can be obtained based on the sub-category matching information between the first object sample and the third object sample; different sub-category matching information can correspond to different category parameters.

[0130] The loss function in this embodiment can be a second mapping relationship between loss information and first matching information, as well as second matching information and its corresponding category parameter information. Since this second mapping relationship utilizes the category parameter information corresponding to the second matching information, and different sub-category matching information can correspond to different category parameters, the data analyzer can differentiate between third object samples of different sub-categories during the learning process, thereby improving the feature discrimination and generalization capabilities of the data analyzer.

[0131] In practical applications, this loss function can characterize the difference between inter-class and intra-class dimensional information. The parameters of the data analyzer can be changed to achieve a lower bound on the difference. Inter-class dimensional information can be obtained based on the second matching information and its corresponding category parameters; for example, it can be the product of the second matching information and the category parameters. Intra-class dimensional information can be obtained based on the first matching information; for example, it can be the product of the first matching information and the first weight information. The exponential information can be obtained based on the difference between inter-class and intra-class dimensional information; for example, it can be obtained based on the difference between inter-class and intra-class dimensional information and the first parameter. Furthermore, the exponential formula can be obtained based on the exponential information and the base information. Further, a logarithmic operation can be performed on the sum of the exponential formula and the second parameter to obtain the logarithmic formula, the result of which represents the loss information.

[0132] Subcategory matching information may include: the number of identical levels between the first object sample and the third object sample, and / or the number of different levels between the first object sample and the third object sample.

[0133] In practical implementation, subcategory matching information can include: first subcategory matching information and second subcategory matching information. The number of different levels corresponding to the first subcategory matching information can be greater than the number of different levels corresponding to the second subcategory matching information, and the category parameter corresponding to the first subcategory matching information can be greater than the category parameter corresponding to the second subcategory matching information. When the number of different levels is large, the category parameter is larger, thus allowing for higher weighting of inter-class dimensional information.

[0134] by Figure 3 Taking the encoded information shown as an example, assuming that the category information corresponding to the encoded information of the first object sample A includes the following three levels of sub-category information: chapter, tax item, and sub-item, the third object sample C specifically includes: the third object sample C1 under the category information corresponding to the "chapter" and "tax item" to which the first object sample A belongs, and different "sub-items", the third object sample C2 under the category information corresponding to the "chapter" to which the first object sample A belongs and different "tax items", and the third object sample C3 under the category information corresponding to different "chapter" of the first object sample A.

[0135] The number of different levels corresponding to the third object sample C1 and the first object sample A is 1, meaning that their encoding information under the sub-item and sub-category is different. Corresponding to Figure 3 The third object sample C1 is the same as the first object sample A in the first 4 bits, but different in the 5th and 6th bits. The category parameter corresponding to the third object sample C1 is called W1.

[0136] The third object sample C2 and the first object sample A have two different levels, meaning their coding information differs under tax items and subcategories. Corresponding to... Figure 3 The third object sample C1 is the same as the first object sample A in the first two positions, but different in the third to sixth positions. The category parameter corresponding to the third object sample C2 is called W2.

[0137] The third object sample C3 and the first object sample A have three different levels, meaning their coding information differs under chapters, tax items, and subcategories. (Corresponding to...) Figure 3 The third object sample C3 differs from the first object sample A in its first 1-6 positions. The category parameter corresponding to the third object sample C3 is called W3.

[0138] This application embodiment can set different category parameters for third object samples C1, C2, and C3, so that the data analyzer can differentiate between third object samples of different subcategories during the learning process. Furthermore, the category parameter W1 corresponding to third object sample C1 can be less than the category parameter W2 corresponding to third object sample C2, and the category parameter W2 corresponding to third object sample C2 can be less than the category parameter W3 corresponding to third object sample C3. This application embodiment does not limit the specific value of the category parameters; for example, the category parameters can be positive numbers less than 5.

[0139] In summary, the data object processing method of this application embodiment utilizes the category parameter information corresponding to the second matching information in the second mapping relationship. Different sub-category matching information can correspond to different category parameters, thus enabling the data analyzer to differentiate between third object samples of different sub-categories during the learning process, thereby improving the feature identification and generalization capabilities of the data analyzer.

[0140] The loss function in this application embodiment can also be a third mapping relationship between loss information and first matching information and first weight information, as well as second matching information, second weight information and category parameter information.

[0141] The embodiments of this application can obtain more accurate loss information based on the third mapping relationship; in other words, the embodiments of this application can improve the accuracy of the loss information. By improving the accuracy of the loss information, the embodiments of this application can increase the convergence speed of the data analyzer and enhance its feature discrimination capability.

[0142] Since the third mapping relationship utilizes the category parameter information corresponding to the second matching information, and different sub-category matching information can correspond to different category parameters, the data analyzer can differentiate between third object samples of different sub-categories during the learning process, thereby improving the feature discrimination and generalization capabilities of the data analyzer.

[0143] When using the third mapping relationship, inter-class dimension information can be obtained based on the second matching information and its corresponding second weight information and category parameter information. For example, inter-class dimension information can be the product of the second matching information, the second weight information, and the category parameter information. Intra-class dimension information can be obtained based on the first matching information. For example, intra-class dimension information can be the product of the first matching information and the first weight information. Exponential information can be obtained based on the difference between inter-class and intra-class dimension information. For example, exponential information can be obtained based on the difference between inter-class and intra-class dimension information and the first parameter. Furthermore, an exponential formula can be obtained based on the exponential information and the base information. Further, a logarithmic operation can be performed on the sum of the exponential formula and the second parameter to obtain a logarithmic formula. The result of the logarithmic formula can be the loss information.

[0144] This application embodiment can update the parameters of the data analyzer during the training process using the aforementioned first, second, or third mapping relationship, so as to obtain the target parameters of the data analyzer when the data analyzer converges. The following embodiments will describe the process of processing data objects using the target parameters of the data analyzer.

[0145] Method Example 4

[0146] Reference Figure 7 The flowchart illustrates the steps of a data object processing method according to an embodiment of this application, which may specifically include the following steps:

[0147] Step 701: Using a data analyzer, determine the first feature representation corresponding to the data object; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0148] Step 702: Based on the first feature representation corresponding to the data object and the second feature representation corresponding to the object sample, determine the target object sample that matches the data object; the second feature representation can be obtained from the data analyzer.

[0149] Step 703: Determine the target feature information corresponding to the data object based on the feature information corresponding to the target object sample;

[0150] The training samples corresponding to the data analyzer may include: triplet object samples, which may include: a first object sample, a second object sample, and a third object sample; the second object sample and the first object sample may correspond to the same category information; the third object sample and the first object sample may correspond to different category information; the parameters of the data analyzer may be obtained based on the first weight information corresponding to the first matching information and the second weight information corresponding to the second matching information; the first matching information corresponds to the matching information between the first object sample and the second object sample; the second matching information may correspond to the matching information between the first object sample and the third object sample; the first weight information may be obtained based on the first matching information and the first preset matching information; the second weight information may be obtained based on the second matching information and the second preset matching information.

[0151] like Figure 1 As shown, during the processing of data objects, the first feature representation corresponding to the data object can be determined based on the target parameters of the data analyzer. For example, the object information of the data object can be input into the data analyzer to obtain the first feature representation output by the data analyzer.

[0152] During the processing of data objects, or before processing data objects, the second feature representation corresponding to the object sample can be determined based on the target parameters of the data analyzer. The object sample can originate from the aforementioned set of object samples.

[0153] In a specific implementation, the first feature representation corresponding to the data object can be matched with the second feature representation corresponding to the object sample to obtain the target object sample that matches the data object. This application embodiment can utilize a measurement method to determine the third matching information between the first and second feature representations. The object samples can be sorted according to the third matching information from high to low, and the top S object samples are selected as the target object samples, where S can be a positive integer.

[0154] In the field of HS coding classification, the feature information corresponding to a target object sample can be the coding information corresponding to the product. The feature information corresponding to a single target object sample can be used as the target feature information for a data object. Alternatively, the target feature information for a data object can be determined by fusing the feature information corresponding to multiple target object samples.

[0155] In this embodiment, the parameters of the data analyzer can be obtained based on the first weight information corresponding to the first matching information and the second weight information corresponding to the second matching information. Since the first weight information reflects the degree of difference between the actual intra-class matching information and the intra-class optimization target in the intra-class dimension, and the second weight information reflects the degree of difference between the actual inter-class matching information and the inter-class optimization target in the inter-class dimension, this embodiment can perform targeted processing in the intra-class and inter-class dimensions based on the first and second weight information respectively during the updating of the data analyzer parameters, thereby achieving the corresponding optimization targets in the intra-class and inter-class dimensions respectively. Therefore, this embodiment can obtain the target parameters of the data analyzer while improving the convergence speed of the data analyzer, and then process the data object based on the target parameters.

[0156] In one implementation of this application, the parameters of the data analyzer can also be obtained based on a first mapping relationship between loss information and first matching information and its corresponding first weight information, and second matching information and its corresponding second weight information. This embodiment of the application can obtain more accurate loss information based on the first mapping relationship; in other words, this embodiment of the application can improve the accuracy of the loss information. By improving the accuracy of the loss information, this embodiment of the application can improve the convergence speed of the data analyzer and enhance its feature discrimination capability. With the improved feature discrimination capability of the data analyzer, this embodiment of the application can improve the accuracy of the processing results of the data object by using the data analyzer to process the data object.

[0157] In another implementation of this application, the parameters of the data analyzer can also be obtained based on a second mapping relationship between the loss information and the first matching information, as well as the second matching information and its corresponding category parameter information.

[0158] The second mapping relationship utilizes the category parameter information corresponding to the second matching information. Different sub-category matching information can correspond to different category parameters, thus enabling the data analyzer to differentiate between third object samples of different sub-categories during the learning process, thereby improving the feature discrimination and generalization capabilities of the data analyzer. By improving the feature discrimination and generalization capabilities of the data analyzer, the embodiments of this application utilize the data analyzer to process data objects, thereby improving the accuracy of the data object processing results.

[0159] In another implementation of this application, the parameters of the data analyzer can also be based on a third mapping relationship between the loss information and the first matching information and the first weight information, as well as the second matching information, the second weight information, and the category parameter information.

[0160] Method Example 5

[0161] Reference Figure 8 The flowchart illustrates the steps of a data object processing method according to an embodiment of this application, which may specifically include the following steps:

[0162] Step 801: Using a data analyzer, determine the first feature representation corresponding to the data object; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0163] Step 802: Based on the first feature representation corresponding to the data object and the second feature representation corresponding to the object sample, determine the target object sample that matches the data object; the second feature representation is obtained from the data analyzer.

[0164] Step 803: Determine the target feature information corresponding to the data object based on the feature information corresponding to the target object sample;

[0165] The training samples for the data analyzer can include: triplet object samples, which include: a first object sample, a second object sample, and a third object sample; the category information corresponding to the triplet object samples includes: N levels of sub-category information; the second object sample and the first object sample have the same N levels of sub-category information; the third object sample and the first object sample have M different levels of sub-category information; N is a positive integer, and M is a positive integer not greater than N; the parameters of the data analyzer are obtained based on the mapping relationship between the loss information and the first matching information, as well as the second matching information and its corresponding category parameter information; the first matching information corresponds to the matching information between the first object sample and the second object sample; the second matching information can correspond to the matching information between the first object sample and the third object sample; the category parameters can be obtained based on the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

[0166] In this embodiment of the application, the parameters of the data analyzer can also be obtained based on a second mapping relationship between the loss information and the first matching information, as well as the second matching information and its corresponding category parameter information.

[0167] The second mapping relationship utilizes the category parameter information corresponding to the second matching information. Different sub-category matching information can correspond to different category parameters, thus enabling the data analyzer to differentiate between third object samples of different sub-categories during the learning process, thereby improving the feature discrimination and generalization capabilities of the data analyzer. By improving the feature discrimination and generalization capabilities of the data analyzer, the embodiments of this application utilize the data analyzer to process data objects, thereby improving the accuracy of the data object processing results.

[0168] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of this application.

[0169] Based on the above embodiments, this application also provides a data object processing apparatus, referring to... Figure 9 The device may include the following modules:

[0170] The sample determination module 901 is used to determine triplet object samples; the triplet object samples include: a first object sample, a second object sample, and a third object sample; wherein, the category information corresponding to the triplet object samples includes: N levels of subcategory information; the N levels of subcategory information corresponding to the second object sample and the first object sample are the same; the third object sample and the first object sample have M different levels of subcategory information; N is a positive integer, and M is a positive integer not greater than N;

[0171] The feature representation determination module 902 is used to determine the feature representation corresponding to the triplet object sample using a data analyzer; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0172] The matching information determination module 903 is used to determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample, based on the feature representation corresponding to the triplet object sample.

[0173] The parameter update module 904 is used to update the parameters of the data analyzer according to the mapping relationship between the loss information and the first matching information, the second matching information and their corresponding category parameter information; the category parameters are obtained according to the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

[0174] To address the aforementioned problems, this application discloses a data object processing apparatus, the apparatus comprising:

[0175] The sample determination module 1001 is used to determine triplet object samples; the triplet object samples include: a first object sample, a second object sample, and a third object sample; wherein, the category information corresponding to the triplet object samples includes: N levels of subcategory information; the N levels of subcategory information corresponding to the second object sample and the first object sample are the same; the third object sample and the first object sample have M different levels of subcategory information; N is a positive integer, and M is a positive integer not greater than N;

[0176] The feature representation determination module 1002 is used to determine the feature representation corresponding to the triplet object sample using a data analyzer; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0177] The matching information determination module 1003 is used to determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample, based on the feature representation corresponding to the triplet object sample.

[0178] The parameter update module 1004 is used to update the parameters of the data analyzer according to the mapping relationship between the loss information and the first matching information, the second matching information and their corresponding category parameter information; the category parameters are obtained according to the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

[0179] To address the aforementioned problems, this application discloses a data object processing apparatus, the apparatus comprising:

[0180] The feature representation determination module 1101 is used to determine the first feature representation corresponding to the data object using a data analyzer; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0181] The target object sample determination module 1102 is used to determine a target object sample that matches the data object based on a first feature representation corresponding to the data object and a second feature representation corresponding to the object sample; the second feature representation is obtained based on the data analyzer.

[0182] The feature determination module 1103 is used to determine the target feature information corresponding to the data object based on the feature information corresponding to the target object sample;

[0183] The training samples corresponding to the data analyzer include: triplet object samples, which include: a first object sample, a second object sample, and a third object sample; the second object sample corresponds to the same category information as the first object sample; the third object sample corresponds to different category information than the first object sample; the parameters of the data analyzer are obtained based on the first weight information corresponding to the first matching information and the second weight information corresponding to the second matching information; the first matching information corresponds to the matching information between the first object sample and the second object sample; the second matching information corresponds to the matching information between the first object sample and the third object sample; the first weight information is obtained based on the first matching information and the first preset matching information; the second weight information is obtained based on the second matching information and the second preset matching information.

[0184] To address the aforementioned problems, this application discloses a data object processing apparatus, the apparatus comprising:

[0185] The feature representation determination module 1201 is used to determine the first feature representation corresponding to the data object using a data analyzer; the data analyzer is used to characterize the mapping relationship between object information and feature representation.

[0186] The target object sample determination module 1202 is used to determine a target object sample that matches the data object based on a first feature representation corresponding to the data object and a second feature representation corresponding to the object sample; the second feature representation is obtained from the data analyzer.

[0187] The feature determination module 1203 is used to determine the target feature information corresponding to the data object based on the feature information corresponding to the target object sample;

[0188] The training samples corresponding to the data analyzer include: triplet object samples, which include: a first object sample, a second object sample, and a third object sample; the category information corresponding to the triplet object samples includes: N levels of sub-category information; the N levels of sub-category information corresponding to the second object sample and the first object sample are the same; the third object sample and the first object sample have M different levels of sub-category information; N is a positive integer, and M is a positive integer not greater than N; the parameters of the data analyzer are obtained according to the mapping relationship between loss information and first matching information, and second matching information and their corresponding category parameter information; the first matching information corresponds to the matching information between the first object sample and the second object sample; the second matching information corresponds to the matching information between the first object sample and the third object sample; the category parameters are obtained according to the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

[0189] This application also provides a non-volatile readable storage medium storing one or more modules (programs). When these modules are applied to a device, they enable the device to execute the instructions for the method steps in this application.

[0190] This application provides one or more machine-readable media storing instructions that, when executed by one or more processors, cause an electronic device to perform one or more of the methods described in the above embodiments. In this application, the electronic device includes devices such as servers and terminal devices.

[0191] Embodiments of this disclosure can be implemented as an apparatus with any suitable hardware, firmware, software, or any combination thereof, configured as desired, and the apparatus may include electronic devices such as servers (clusters) and terminals. Figure 13 An exemplary apparatus 1300 is schematically shown that can be used to implement the various embodiments described in this application.

[0192] In one embodiment, Figure 13 An exemplary device 1300 is shown, which includes one or more processors 1302, a control module (chipset) 1304 coupled to at least one of the processors 1302, a memory 1306 coupled to the control module 1304, a non-volatile memory (NVM) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304.

[0193] Processor 1302 may include one or more single-core or multi-core processors, and processor 1302 may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, device 1300 can serve as a server, terminal, or other device as described in the embodiments of this application.

[0194] In some embodiments, apparatus 1300 may include one or more computer-readable media (e.g., memory 1306 or NVM / storage device 1308) having instructions 1314 and one or more processors 1302 that are combined with the one or more computer-readable media and configured to execute instructions 1314 to implement modules and thus perform the actions described in this disclosure.

[0195] In one embodiment, the control module 1304 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1302 and / or any suitable device or component communicating with the control module 1304.

[0196] The control module 1304 may include a memory controller module to provide an interface to the memory 1306. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0197] Memory 1306 may be used, for example, to load and store data and / or instructions 1314 for device 1300. In one embodiment, memory 1306 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, memory 1306 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).

[0198] In one embodiment, the control module 1304 may include one or more input / output controllers to provide interfaces to the NVM / storage device 1308 and (one or more) input / output devices 1310.

[0199] For example, NVM / storage device 1308 may be used to store data and / or instructions 1314. NVM / storage device 1308 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more optical disc drives (CDs), and / or one or more digital universal optical disc (DVD) drives).

[0200] NVM / storage device 1308 may include storage resources that are part of a device on which device 1300 is mounted, or that are accessible by the device but do not necessarily have to be part of the device. For example, NVM / storage device 1308 may be accessed via a network via one or more input / output devices 1310.

[0201] One or more input / output devices 1310 may provide an interface for device 1300 to communicate with any other suitable device. Input / output devices 1310 may include communication components, audio components, sensor components, etc. Network interface 1312 may provide an interface for device 1300 to communicate via one or more networks. Device 1300 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G, 5G, etc., or combinations thereof.

[0202] In one embodiment, at least one of the processors 1302 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 1304. In one embodiment, at least one of the processors 1302 may be logically packaged with one or more controllers of the control module 1304 to form a system-in-package (SiP). In one embodiment, at least one of the processors 1302 may be integrated with the logic of one or more controllers of the control module 1304 on the same die. In one embodiment, at least one of the processors 1302 may be integrated with the logic of one or more controllers of the control module 1304 on the same die to form a system-on-a-chip (SoC).

[0203] In various embodiments, device 1300 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, device 1300 may have more or fewer components and / or different architectures. For example, in some embodiments, device 1300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0204] In the device 1300, a main control chip can be used as a processor or control module, sensor data, position information, etc. are stored in a memory or NVM / storage device, the sensor group can be used as an input / output device, and the communication interface can include a network interface.

[0205] This application also provides an electronic device, including: a processor; and a memory storing executable code thereon, which, when executed, causes the processor to perform one or more methods as described in this application.

[0206] This application also provides one or more machine-readable media having executable code stored thereon, which, when executed, causes a processor to perform one or more of the methods described in this application.

[0207] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0208] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0209] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0210] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0211] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0212] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0213] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0214] The foregoing has provided a detailed description of a data object processing method, a data object processing apparatus, an electronic device, and a storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for processing data objects, characterized in that, The method includes: Determine triplet object samples; the triplet object samples include: a first object sample, a second object sample, and a third object sample; wherein, the second object sample corresponds to the same category information as the first object sample; and the third object sample corresponds to different category information than the first object sample. A data analyzer is used to determine the feature representation corresponding to the triplet object sample; the data analyzer is used to characterize the mapping relationship between object information and feature representation; the feature representation is used to determine the target feature information corresponding to the data object; wherein, the object information of the data object is input into the data analyzer, the data analyzer outputs the feature representation corresponding to the object information, the feature representation is input into a preset processor, and the preset processor performs preset processing on the feature representation to obtain the corresponding processing result; the preset processing includes: determining whether the data object is a target; or, determining the target feature information corresponding to the data object; the object information includes: text or image; the target feature information includes: text corresponding to an image, or feature information corresponding to a product; Based on the feature representations corresponding to the triplet object samples, determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample; The parameters of the data analyzer are updated based on the first weight information corresponding to the first matching information and the second weight information corresponding to the second matching information; the first weight information is obtained based on the first matching information and the first preset matching information, and the second weight information is obtained based on the second matching information and the second preset matching information.

2. The method according to claim 1, characterized in that, The category information corresponding to the triplet object sample includes: N levels of subcategory information; the N levels of subcategory information corresponding to the second object sample and the first object sample are the same; the third object sample and the first object sample have M different levels of subcategory information; N is a positive integer, and M is a positive integer not greater than M.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Based on the first mapping relationship between the loss information and the first matching information and its corresponding first weight information, and the second matching information and its corresponding second weight information, update the parameters of the data analyzer; and / or The parameters of the data analyzer are updated based on the second mapping relationship between the loss information and the first matching information, as well as the second matching information and their corresponding category parameter information; the category parameters are obtained based on the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

4. A method for processing data objects, characterized in that, The method includes: Determine triplet object samples; the triplet object samples include: a first object sample, a second object sample, and a third object sample; wherein, the category information corresponding to the triplet object samples includes: N levels of subcategory information; the N levels of subcategory information corresponding to the second object sample and the first object sample are the same; the third object sample and the first object sample have M different levels of subcategory information; N is a positive integer, and M is a positive integer not greater than N; A data analyzer is used to determine the feature representation corresponding to the triplet object sample. The data analyzer is used to characterize the mapping relationship between object information and feature representation. Specifically, the object information of the data object is input into the data analyzer, which outputs the feature representation corresponding to the object information. The feature representation is then input into a preset processor, which performs preset processing on the feature representation to obtain a corresponding processing result. The preset processing includes: determining whether the data object is a target; or, determining the target feature information corresponding to the data object. The object information includes: text or image; the target feature information includes: text corresponding to an image, or feature information corresponding to a product. Based on the feature representations corresponding to the triplet object samples, determine the first matching information between the first object sample and the second object sample, and the second matching information between the first object sample and the third object sample; The parameters of the data analyzer are updated based on the mapping relationship between the loss information and the first matching information, the second matching information and their corresponding category parameter information; the category parameters are obtained based on the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

5. The method according to claim 4, characterized in that, The sub-category matching information includes: first sub-category matching information and second sub-category matching information; wherein, the number of different levels corresponding to the first sub-category matching information is greater than the number of different levels corresponding to the second sub-category matching information, and the category parameter corresponding to the first sub-category matching information is greater than the category parameter corresponding to the second sub-category matching information.

6. A method for processing data objects, characterized in that, The method includes: The object information of the data object is input into the data analyzer, and the data analyzer is used to determine the first feature representation corresponding to the data object; the data analyzer is used to characterize the mapping relationship between the object information and the feature representation; the data object includes a clearance object, and the object information includes: the attribute information of the clearance object; Based on the first feature representation corresponding to the data object and the second feature representation corresponding to the object sample, a target object sample matching the data object is determined; the second feature representation is obtained from the data analyzer. Based on the feature information corresponding to the target object sample, the target feature information corresponding to the data object is determined; the target feature information includes: the encoding information of the clearance object; The training samples corresponding to the data analyzer include: triplet object samples, which include: a first object sample, a second object sample, and a third object sample; the second object sample corresponds to the same category information as the first object sample; the third object sample corresponds to different category information than the first object sample; the parameters of the data analyzer are obtained based on the first weight information corresponding to the first matching information and the second weight information corresponding to the second matching information; the first matching information corresponds to the matching information between the first object sample and the second object sample; the second matching information corresponds to the matching information between the first object sample and the third object sample; the first weight information is obtained based on the first matching information and the first preset matching information; the second weight information is obtained based on the second matching information and the second preset matching information.

7. The method according to claim 6, characterized in that, The parameters of the data analyzer are also obtained based on a first mapping relationship between loss information and first matching information and its corresponding first weight information, and second matching information and its corresponding second weight information; and / or The parameters of the data analyzer are also obtained based on a second mapping relationship between the loss information and the first matching information, as well as the second matching information and its corresponding category parameter information.

8. A method for processing data objects, characterized in that, The method includes: The object information of the data object is input into the data analyzer, and the data analyzer is used to determine the first feature representation corresponding to the data object; the data analyzer is used to characterize the mapping relationship between the object information and the feature representation; the data object includes a clearance object, and the object information includes: the attribute information of the clearance object; Based on the first feature representation corresponding to the data object and the second feature representation corresponding to the object sample, a target object sample matching the data object is determined; the second feature representation is obtained from the data analyzer. Based on the feature information corresponding to the target object sample, the target feature information corresponding to the data object is determined; the target feature information includes: the encoding information of the clearance object; The training samples corresponding to the data analyzer include: triplet object samples, which include: a first object sample, a second object sample, and a third object sample; the category information corresponding to the triplet object samples includes: N levels of sub-category information; the N levels of sub-category information corresponding to the second object sample and the first object sample are the same; the third object sample and the first object sample have M different levels of sub-category information; N is a positive integer, and M is a positive integer not greater than N; the parameters of the data analyzer are obtained according to the mapping relationship between loss information and first matching information, and second matching information and their corresponding category parameter information; the first matching information corresponds to the matching information between the first object sample and the second object sample; the second matching information corresponds to the matching information between the first object sample and the third object sample; the category parameters are obtained according to the sub-category matching information between the first object sample and the third object sample; different sub-category matching information corresponds to different category parameters.

9. An electronic device, characterized in that, include: processor; and A memory having executable code stored thereon, which, when executed, causes the processor to perform the method as described in any one of claims 1-8.

10. One or more machine-readable media having executable code stored thereon, which, when executed, causes a processor to perform the method as claimed in any one of claims 1-8.