Data object processing method, electronic device, and storage medium
By matching and clustering the semantic representation of data objects with standard category information, the problem of low accuracy in item classification when category information does not appear in the description information is solved, and higher classification accuracy is achieved.
Patent Information
- Application Number
- CN202210313364.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2042-03-28
AI Technical Summary
Existing item classification methods rely on category information included in the item description, resulting in low classification accuracy, especially when the category information is not present in the description, leading to either failure to classify or incorrect classification.
By determining the first semantic representation of data objects and the second semantic representation of standard category information, and utilizing semantic-level matching and clustering techniques, classification accuracy can be improved.
It improves the accuracy of item classification without relying on category information included in item description information. Through semantic matching and clustering, it unifies similar preset category information into standard category information, thereby improving classification accuracy.
Smart Images

Figure CN116860962B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of communication, in particular to a data object processing method, an electronic device and a storage medium. BACKGROUND
[0002] With the development of communication technology, the category information of data objects such as articles and commodities is becoming more and more important. In a logistics scenario, the logistics processing mode of an article needs to be determined according to the category information of the article; for example, the category information "lithium battery" belongs to nine types of dangerous goods, which corresponds to a special packaging mode and transportation mode. In a customs declaration scenario, the code of an article is also related to the category information of the article; for example, an article of the category lipstick corresponds to the first four digits "3304" of the code information.
[0003] The current classification method usually performs named entity recognition on the description information of an article to obtain the entity contained in the description information as the category information of the article.
[0004] The current classification method requires that the category information of an article must appear in the description information, and if the category information of an article does not appear in the description information, the category information cannot be obtained, thus resulting in low classification accuracy. For example, the description information "iPhone 13 256G" does not contain category information, and therefore, classification cannot be performed on the description information "iPhone 13 256G".
[0005] The current classification method may obtain two types of category information from the description information of an article, thus resulting in low classification accuracy. For example, the current classification method extracts two types of category information "pillow" and "hot water bag" from the description information "yellow pillow hot water bag", and in this case, the true category information cannot be obtained. SUMMARY
[0006] Embodiments of the present application provide a data object processing method, which can improve the classification accuracy of data objects.
[0007] Correspondingly, embodiments of the present application also provide a data object processing apparatus, an electronic device and a storage medium to implement the implementation and application of the above method.
[0008] To solve the above problems, embodiments of the present application disclose a data object processing method, which comprises:
[0009] determining a first semantic representation corresponding to a data object and a second semantic representation corresponding to standard category information; the standard category information is obtained by clustering semantic representations corresponding to a plurality of preset category information according to a semantic representation corresponding to a preset category information;
[0010] determine matching information between the data object and the standard category information according to the first semantic representation and the second semantic representation;
[0011] determine target standard category information corresponding to the data object according to the matching information.
[0012] To solve the above problems, an embodiment of the present application discloses a data object processing method, the method comprising:
[0013] determine a first semantic representation corresponding to a data object, and a second semantic representation corresponding to preset category information;
[0014] determine weight information corresponding to the first semantic representation according to attention information of the second semantic representation to the first semantic representation;
[0015] weight the first semantic representation according to the weight information to obtain a first weighted semantic representation;
[0016] determine matching information between the data object and the standard category information according to the first weighted semantic representation and the second semantic representation;
[0017] determine target category information corresponding to the data object according to the matching information.
[0018] To solve the above problems, an embodiment of the present application discloses a data object processing method, the method comprising:
[0019] determine a first semantic representation corresponding to a data object, and a second semantic representation corresponding to preset category information according to a data analyzer; the data analyzer is used to represent a mapping relationship between description information and a semantic representation; the data analyzer is obtained by comparative learning on training data;
[0020] determine matching information between the data object and the standard category information according to the first semantic representation and the second semantic representation;
[0021] determine target category information corresponding to the data object according to the matching information.
[0022] To solve the above problems, an embodiment of the present application discloses a data object processing device, the device comprising:
[0023] a semantic representation determination module, configured to determine a first semantic representation corresponding to a data object, and a second semantic representation corresponding to standard category information; the standard category information is obtained by clustering semantic representations corresponding to a plurality of preset category information according to a semantic representation corresponding to preset category information;
[0024] a matching module, configured to determine matching information between the data object and the standard category information according to the first semantic representation and the second semantic representation;
[0025] a category determining module, configured to determine target standard category information corresponding to the data object according to the matching information.
[0026] To solve the above problems, an embodiment of the present application discloses a data object processing device, the device comprising:
[0027] a semantic representation determining module, configured to determine a first semantic representation corresponding to a data object and a second semantic representation corresponding to preset category information;
[0028] a weight determining module, configured to determine weight information corresponding to the first semantic representation according to attention information of the second semantic representation to the first semantic representation;
[0029] a weighting module, configured to weight the first semantic representation according to the weight information to obtain a first weighted semantic representation;
[0030] a matching module, configured to determine matching information between the data object and the standard category information according to the first weighted semantic representation and the second semantic representation;
[0031] a category determining module, configured to determine target category information corresponding to the data object according to the matching information.
[0032] To solve the above problems, an embodiment of the present application discloses a data object processing device, the device comprising:
[0033] a semantic representation determining module, configured to determine a first semantic representation corresponding to a data object and a second semantic representation corresponding to preset category information according to a data analyzer; the data analyzer is used to represent a mapping relationship between description information and a semantic representation; the data analyzer can be obtained by comparative learning on training data;
[0034] a matching module, configured to determine matching information between the data object and the standard category information according to the first semantic representation and the second semantic representation;
[0035] a category determining module, configured to determine target category information corresponding to the data object according to the matching information.
[0036] To solve the above problems, an embodiment of the present application discloses an electronic device, comprising: a processor; and a memory having executable code stored thereon, when the executable code is executed, causing the processor to execute the method in any one of the above embodiments.
[0037] To solve the above problems, the embodiment of the application discloses a machine readable medium, which stores executable code, and when the executable code is executed, the processor executes the method as described in any one of the above embodiments.
[0038] Compared with the prior art, the embodiment of the application has the following advantages:
[0039] In the technical solution of the embodiment of the application, the matching information between the data object and the standard category information is determined based on the first semantic representation and the second semantic representation in the semantic level, which is equivalent to performing matching processing of the data object and the standard category information in the semantic level. The above matching processing in the semantic level does not require that the target standard category information must appear in the description information of the data object, and therefore the classification accuracy of the data object can be improved.
[0040] In addition, the standard category information in the embodiment of the application can be obtained by clustering the semantic representations corresponding to the preset category information and the semantic representations corresponding to the plurality of preset category information. Since the above clustering can unify the plurality of preset category information having similarity into one standard category information, for example, unifying the plurality of preset category information having similarity such as shampoo, shampoo and shampoo paste into one standard category information "shampoo", the embodiment of the application can provide the unified target standard category information, and therefore the classification accuracy of the data object can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0041] Figure 1 is a step flowchart of the data object processing method of one embodiment of the application;
[0042] Figure 2 is a step flowchart of the data object processing method of one embodiment of the application;
[0043] Figure 3 is a schematic diagram of an application environment of the data object processing method of one embodiment of the application;
[0044] Figure 4 is a structural schematic diagram of the data object processing device of one embodiment of the application;
[0045] Figure 5 is a structural schematic diagram of the word processing module 402 of one embodiment of the application;
[0046] Figure 6 is a flowchart of the data object processing method of one embodiment of the application;
[0047] Figure 7 is a step flowchart of the data object processing method of one embodiment of the application;
[0048] Figure 8 is a step flow chart of a processing method of a data object according to an embodiment of the present application;
[0049] Figure 9 is a structural schematic diagram of a processing device of a data object according to an embodiment of the present application;
[0050] Figure 10 is a structural schematic diagram of a processing device of a data object according to an embodiment of the present application;
[0051] Figure 11 is a structural schematic diagram of a processing device of a data object according to an embodiment of the present application;
[0052] Figure 12 is a structural schematic diagram of an exemplary device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to make the above objectives, features and advantages of the present application more apparent and comprehensible, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0054] In an embodiment of the present application, a data object can be a composite information representation understood by software. The data object can be an entity, a thing, an incident or event, a role, an organizational unit, a location or a structure, etc. For example, the data object can include an article, a commodity, etc.
[0055] In an embodiment of the present application, a data object such as a commodity is classified. Specifically, according to description information of the data object, category information corresponding to the data object is determined. Commodity classification refers to a process of collecting a commodity set in a management range, selecting appropriate commodity basic characteristics as classification marks, and gradually inducing a smaller range and more consistent characteristics of a sub-set (category information) such as a large category, a medium category, a small category, a fine category, and a product, a fine item, etc., so that all commodities in the range can be clearly distinguished and systematized.
[0056] The category information in an embodiment of the present application can be applied to a logistics scenario, a customs declaration scenario, etc. It can be understood that the specific application scenario corresponding to the category information is not limited in the present application. For example, in a logistics scenario, a logistics processing mode of an article can be determined according to category information of the article; for example, category information “lithium battery” belongs to nine types of dangerous goods, and corresponds to a special packaging mode and a transportation mode. For another example, in a customs declaration scenario, an article code is also related to category information of the article; for example, an article of a lipstick category corresponds to a code information of “3304” in the first four digits.
[0057] The current classification method usually performs named entity recognition on the description information of the item to obtain the entity contained in the description information as the category information of the item.
[0058] The current classification method requires that the category information of the item must appear in the description information. If the category information of the item does not appear in the description information, the category information cannot be obtained, thereby resulting in low classification accuracy. For example, the description information "iPhone 13 256G" does not contain category information, and therefore, classification cannot be performed on the description information "iPhone 13 256G".
[0059] The current classification method can obtain two types of category information from the description information of an item, thereby resulting in low classification accuracy. For example, the current classification method extracts two types of category information "pillow" and "hot water bag" from the description information "yellow pillow hot water bag", and in this case, the true category information cannot be obtained.
[0060] To solve the technical problem of low classification accuracy, the embodiments of the present application provide a processing scheme of a data object, which specifically includes: determining a first semantic representation corresponding to the data object and a second semantic representation corresponding to standard category information; the standard category information can be obtained by clustering semantic representations corresponding to multiple preset category information according to a semantic representation corresponding to a preset category information; determining matching information between the data object and the standard category information according to the first semantic representation and the second semantic representation; and determining target standard category information corresponding to the data object according to the matching information.
[0061] The first semantic representation of the embodiments of the present application can represent information of the data object at the semantic level, and the second semantic representation can represent information of the standard category information at the semantic level. The embodiments of the present application determine the matching information between the data object and the standard category information based on the first semantic representation and the second semantic representation at the semantic level, which is equivalent to performing matching processing of the data object and the standard category information at the semantic level. The above matching processing at the semantic level does not require that the target standard category information must appear in the description information of the data object, and therefore, the classification accuracy of the data object can be improved.
[0062] Furthermore, the standard category information of the embodiments of the present application can be obtained by clustering semantic representations corresponding to multiple preset category information according to a semantic representation corresponding to a preset category information. Since the above clustering can unify multiple preset category information having similarity into one standard category information, for example, unifying multiple preset category information such as "shampoo", "shampoo", and "hair wash paste" having similarity into one standard category information "shampoo", the embodiments of the present application can provide the unified target standard category information, thereby improving the classification accuracy of the data object.
[0063] Method embodiment one
[0064] This embodiment describes the learning process of the data analyzer.
[0065] The data analyzer can be used to represent the mapping relationship between the description information and the semantic representation, and the data analyzer can use machine learning methods such as deep learning to obtain the semantic representation corresponding to the description information. Deep learning is a representation learning method with multiple levels of representation, which transforms the representation of the upper level into a higher level of representation.
[0066] The embodiments of the present application can learn the mathematical model based on the training data to obtain the data analyzer. The mathematical model is a scientific or engineering model constructed by using mathematical logic methods and mathematical language. The mathematical model is a mathematical structure that describes the characteristics or quantitative dependence relationship of a certain thing system, which is described by mathematical language, and is a relationship structure described by mathematical symbols. The mathematical model can be one or a group of algebraic equations, differential equations, difference equations, integral equations or statistical equations and their combinations, which quantitatively or qualitatively describe the mutual relationship or causal relationship between variables of the system. In addition to the mathematical model described by equations, there are models described by other mathematical tools such as algebra, geometry, topology, mathematical logic, etc. The mathematical model describes the behavior and characteristics of the system rather than the actual structure of the system. The training of the mathematical model can be performed by using machine learning methods such as deep learning methods. Machine learning methods can include linear regression, decision tree, random forest, etc. Deep learning methods can include CNN (Convolutional Neural Networks), LSTM (Long Short-Term Memory), GRU (Gated Recurrent Unit), etc.
[0067] The data analyzer of the embodiments of the present application can include pre-trained models such as BERT (Bidirectional Encoder Representation from Transformers), ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacement Accurately), and Transformer.
[0068] The training process of the pre-trained model can include two stages: a pre-training stage and a fine-tuning stage. In the pre-training stage, self-supervised learning of large-scale data can be used to obtain a pre-trained model that is independent of a specific task, for example, training the pre-trained model on a large amount of general corpus to learn general language knowledge; in the fine-tuning stage, the network parameters of the pre-trained model can be corrected for a specific task. Specifically, in the embodiments of the present application, the network parameters of the pre-trained model can be corrected for a classification task of data objects, and the corrected pre-trained model can be used to determine the first semantic representation corresponding to the data objects and the second semantic representation corresponding to the standard category information.
[0069] The learning method of the data analyzer can include a supervised learning method or an unsupervised learning method, etc. The supervised learning method corresponds to labeled training data, and the unsupervised learning method corresponds to unlabeled training data. Since labeling of the labeled training data requires a large amount of time and resources, the unsupervised learning method can automatically discover the rules in the training data without labeling the training data, thereby saving the labeling cost of the training data.
[0070] The unsupervised learning method can include generative learning or contrastive learning, etc. Generative learning is represented by methods such as self-encoder (e.g., GAN, Generative Adversarial Networks), which generates data from data, so that the data is similar to the training data in the whole or high-level semantics. Contrastive learning learns the common features between similar samples (positive samples) and distinguishes the differences between non-similar samples (negative samples). Compared with generative learning, contrastive learning distinguishes data in the abstract semantic level feature space, and has the advantages of simple model and strong generalization ability. In other words, the goal of contrastive learning is to learn an encoder that encodes similar data of the same class and makes the encoding results of data of different classes as different as possible.
[0071] In the case of using contrastive learning, the training data can include a target sample, a positive sample corresponding to the target sample, and a negative sample corresponding to the target sample; wherein the target sample includes a first sample corresponding to first preset category information; the positive sample includes a second sample corresponding to the first preset category information; the negative sample includes a third sample corresponding to second preset category information; and the first preset category information is different from the second preset category information. The target sample, the positive sample corresponding to the target sample, and the negative sample corresponding to the target sample can be represented as a triple (X, X+, X-). Wherein X represents the target sample, X+ represents the positive sample, and X- represents the negative sample. The positive sample can be one or more, the negative sample can be (N-1), and N can be a positive integer greater than 1.
[0072] Assuming that the first preset category is category A, the positive samples can include: at least one sample corresponding to category A, such as a sample corresponding to category A and having description information A1, a sample corresponding to category A and having description information A2, a sample corresponding to category A and having description information A3, and so on, and a sample corresponding to category A and having description information Am (m can be a positive integer). The negative samples can include: at least one sample corresponding to category B, at least one sample corresponding to category C, and the like. In other words, the categories corresponding to the negative samples can include any one or more categories corresponding to category A.
[0073] The training of the data analyzer can include forward propagation and backward propagation.
[0074] The forward propagation can calculate the output information in sequence according to the input information from the input layer to the output layer. The output information can be used to determine the loss information. The input information of the embodiments of the present application can include description information corresponding to the target sample or the positive sample or the negative sample or the like in the triple. The output information can include semantic representation corresponding to the target sample or the positive sample or the negative sample or the like in the triple. Assuming that the semantic representation corresponding to the target sample in the triple is semantic representation A, the semantic representation corresponding to the positive sample in the triple is semantic representation B, and the semantic representation corresponding to the negative sample in the triple is semantic representation C.
[0075] The backward propagation can calculate and update the input information in sequence according to the loss information from the output layer to the input layer. During the backward propagation, the gradient information of the input information can be determined, and the input information can be updated using the gradient information. For example, the backward propagation can calculate and store the gradient information of the input information in sequence according to the chain rule in calculus from the output layer to the input layer.
[0076] The embodiments of the present application can characterize the mapping relationship between the loss information and the semantic representation A, the semantic representation B, and the semantic representation C via a loss function. The mapping relationship can include first matching information between the semantic representation A and the semantic representation B, and second matching information between the semantic representation A and the semantic representation C.
[0077] The embodiments of the present application can determine the first matching information and the second matching information by using a measurement method. The measurement method can include Euclidean distance, or cosine of the included angle, or information entropy, and it can be understood that the embodiments of the present application do not limit the specific measurement method. The first matching information can represent the matching degree between the semantic representation A and the semantic representation B, and generally the higher the matching degree, the shorter the distance between the semantic representation A and the semantic representation B.
[0078] The loss function can aim to increase the matching degree between the semantic representation A and the semantic representation B, and decrease the matching degree between the semantic representation A and the semantic representation C. In an implementation, the numerator of the mapping relationship can be determined according to the first matching information, and the denominator of the mapping relationship can be determined according to the second matching information or the first matching information and the second matching information.
[0079] The embodiments of the present application can take the loss information as a preset value as an optimization target to update the parameters of the data analyzer. The optimization method can include gradient descent method, Newton method, quasi-Newton method, conjugate gradient method, etc. It can be understood that the embodiments of the present application do not limit the specific optimization method.
[0080] In actual application, the partial derivative of the parameters of the data analyzer can be calculated, and the calculated partial derivative of the parameters is written in the form of a vector. The vector corresponding to the partial derivative can be referred to as gradient information corresponding to the parameters. The update amount corresponding to the parameters can be obtained according to the gradient information and the step information. It can be understood that the update process of the parameters is only an example, and the embodiments of the present application do not limit the specific update process of the parameters of the data analyzer.
[0081] It should be noted that the semantic representation of the embodiments of the present application can include a word-dimension semantic representation and / or a sentence-dimension semantic representation. The data analyzer can output the word-dimension semantic representation according to the description information after tokenization. The data analyzer can output the sentence-dimension semantic representation according to the description information without tokenization.
[0082] In summary, the embodiments of the present application can enable the data analyzer to learn the common features between the semantic representations of the same category information and distinguish the differences between the semantic representations of different category information based on the contrast learning of the training data. Therefore, the embodiments of the present application can improve the accuracy of the semantic representation output by the data analyzer.
[0083] Method embodiment two
[0084] The embodiments of the present application explain the process of clustering the semantic representations corresponding to the preset category information and the semantic representations corresponding to multiple preset category information to obtain the standard category information.
[0085] The standard category information of the embodiments of the present application can be obtained by clustering the semantic representations corresponding to the preset category information and the semantic representations corresponding to multiple preset category information. Since the above clustering can unify multiple preset category information with similarity into a standard category information, for example, unifying multiple preset category information with similarity such as shampoo, shampoo and shampoo paste into a standard category information "shampoo", the embodiments of the present application can provide the unified target standard category information, thereby improving the classification accuracy of the data object.
[0086] Referring to Figure 1 , a step flow chart of a data object processing method of an embodiment of the present application is shown, which can specifically include the following steps:
[0087] Step 101, determining sample data; the sample data can include: preset category information and corresponding description information thereof;
[0088] Step 102, determining semantic representation corresponding to the preset category information according to the description information corresponding to the preset category information;
[0089] Step 103, clustering semantic representations corresponding to multiple preset category information to obtain a standard category information corresponding to the multiple preset category information.
[0090] In step 101, data objects with known category information can be collected to obtain sample data. For example, the known category information of commodity A is category A, and the sample data corresponding to commodity A can be obtained: the preset category information is category A, and the description information is the description information of commodity A. It can be understood that the specific determination method of the sample data is not limited in the embodiment of the present application.
[0091] In step 102, the description information corresponding to the preset category information can be input into a data analyzer to obtain the semantic representation corresponding to the preset category information output by the data analyzer.
[0092] In step 103, a clustering method can be used to cluster the semantic representations corresponding to the multiple preset category information.
[0093] The clustering method can divide a data set into different classes or clusters according to a preset feature (such as distance), so that the similarity of data objects in the same cluster is as large as possible, and the difference of data objects not in the same cluster is as large as possible. The clustering method can include a hierarchical clustering method, a hierarchical clustering method, a density-based clustering method, etc. It can be understood that the specific clustering method is not limited in the embodiment of the present application.
[0094] An example of clustering semantic representations corresponding to multiple preset category information is provided, in which a data set D = {d1, d2,.., dn}, d1…dn represents multiple preset categories, p and p are two arbitrary preset categories in D, and the distance between p and p can be evaluated using a Euclidean distance or other measurement method. The neighborhood of p can refer to a set of preset categories with a distance less than a threshold from p.
[0095] The flow of the example can include:
[0096] Step S1, marking the preset categories in the data set as a clustering state;
[0097] Step S2, randomly selecting a preset category p corresponding to a state to be clustered from the data set, and marking the preset category p as a clustered state;
[0098] Step S3, judging whether the neighborhood of the preset category p contains M (M is a conception) preset categories, if yes, executing step S4, otherwise executing step S;
[0099] Step S4, creating a cluster C, and putting the preset category p into C;
[0100] Step S5, setting N as a set corresponding to the neighborhood of p, for a preset category p' in N:
[0101] If the preset category p' is a state to be clustered, marking p' as a clustered state;
[0102] If the neighborhood of the preset category p' contains at least M preset categories, adding at least M preset categories in the neighborhood of p' to N;
[0103] If the preset category p' is not a member of any cluster, adding p' to C, and saving C;
[0104] Step S6, marking the preset category p as noise.
[0105] The clustering of the embodiment of the application can convert the data set {d1, d2,.., dn} corresponding to the preset category information into the data set Q (q1, q2,.. q m ).
[0106] The clustering of the embodiment of the application can unify various preset category information having similar semantic representations into one standard category information, for example, unifying various preset category information having similarities such as "shampoo", "shampoo", and "shampoo paste" into one standard category information "shampoo"; therefore, the embodiment of the application can provide the unified target standard category information, and further can improve the classification accuracy of the data object.
[0107] Method embodiment three
[0108] Referring to Figure 2 , a step flowchart of a data object processing method of one embodiment of the application is shown, which can specifically include the following steps:
[0109] Step 201, determining a first semantic representation corresponding to a data object, and a second semantic representation corresponding to standard category information; the standard category information can be obtained by clustering semantic representations corresponding to various preset category information according to a semantic representation corresponding to preset category information;
[0110] Step 202, determining matching information between the data object and the standard category information according to the first semantic representation and the second semantic representation;
[0111] Step 203, determining target standard category information corresponding to the data object according to the matching information.
[0112] Embodiments of the present application can be used for classifying data objects such as commodities, and specifically, determining category information corresponding to a data object according to description information of the data object.
[0113] The processing method of the data object provided by the embodiments of the present application can be applied to Figure 3 application environments as shown in Figure 3 The client 301 and the server 302 are located in a wired or wireless network, and the client 301 and the server 302 perform data interaction through the wired or wireless network.
[0114] In actual applications, the client 301 can run on a terminal, and the terminal specifically includes but is not limited to: a smart phone, a tablet computer, an e-book reader, a recording device, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, a vehicle-mounted computer, a desktop computer, a set-top box, a smart television, a wearable device, and the like.
[0115] The client 301 can correspond to a target APP, such as a classification APP of a data object. The client 301 or the server 302 can execute at least one step included in the method of the embodiments of the present application.
[0116] In the case where the server 302 executes the method of the embodiments of the present application, the server 302 can provide a classification service to the client 301. The client 301 can send a classification request to the server 302, and the classification request can include description information of a data object. The server 302 can execute the method of the embodiments of the present application to obtain target standard category information corresponding to the data object, and return the target standard category information corresponding to the data object to the client 301.
[0117] In the process of executing the method of the embodiments of the present application, the server 302 can use a database corresponding to the standard category information. The database can include a corresponding relationship between the second semantic representation and the standard category information.
[0118] Of course, the server 302 executes the method of the embodiments of the present application, only as an example, in fact, the client 301 can also execute the method of the embodiments of the present application.
[0119] In step 201, the data object can be a data object to be classified, and the description information of the data object can be input into the data analyzer to obtain the first semantic representation output by the data analyzer. Taking the data object as a commodity as an example, the description information of the embodiments of the present application can include commodity title information, or commodity name, or commodity search words, etc. It can be understood that the embodiments of the present application do not limit the specific description information.
[0120] The embodiments of the present application can pre-input the description information of the standard category information into the data analyzer to obtain the corresponding second semantic representation output by the data analyzer, and save the corresponding relationship between the second semantic representation and the standard category information. In this way, when the data object is received, the second semantic representation can be obtained from the corresponding relationship.
[0121] In actual application, the first semantic representation can include a word-dimension first semantic representation and / or a sentence-dimension first semantic representation. The second semantic representation can include a word-dimension second semantic representation and / or a sentence-dimension second semantic representation.
[0122] In step 202, the matching information can represent the matching degree between the data object and the standard category information. Generally, the higher the matching degree, the higher the numerical value corresponding to the matching information.
[0123] In one implementation, the matching information A between the first semantic representation and the second semantic representation can be determined by using a metric method. The matching information A can represent the matching degree between the first semantic representation and the second semantic representation. Generally, the higher the matching degree between the first semantic representation and the second semantic representation, the shorter the distance between the first semantic representation and the second semantic representation.
[0124] In another implementation, the matching information B between the first weighted semantic representation and the second semantic representation can be determined by using a metric method. The matching information B can represent the matching degree between the first weighted semantic representation and the second semantic representation.
[0125] The determining process of the first weighted semantic representation can comprise: determining weight information corresponding to the first semantic representation according to the attention information of the second semantic representation to the first semantic representation; weighting the first semantic representation according to the weight information to obtain the first weighted semantic representation; in this case, the matching information between the data object and the standard category information can be determined according to the first weighted semantic representation and the second semantic representation. Since the first weighted semantic representation is obtained according to the weight information, and the weight information is obtained according to the attention information of the second semantic representation to the first semantic representation; therefore, in actual application, a larger weight can be obtained for the part (such as the first word vector) in the first semantic representation that has a higher correlation degree with the second semantic representation, and a smaller weight can be obtained for the part (such as the second word vector) in the first semantic representation that has a lower correlation degree with the second semantic representation, thereby the first weighted semantic representation can reflect the correlation degree between the first semantic representation and the second semantic representation, so as to improve the accuracy of the first weighted semantic representation and the accuracy of the matching information between the data object and the standard category information.
[0126] In actual application, the first semantic representation can comprise a word-dimension first semantic representation, such as the first semantic representation can comprise a first word vector of a1, a2, a3, …, ai, etc. The second semantic representation can comprise a word-dimension second semantic representation, such as the second semantic representation can comprise a second word vector of b1, b2, b3, …, bi, etc. Then the embodiment of the present application can first determine an attention matrix corresponding to the first word vector and the second word vector, the size of the attention matrix can be i×i (i can be a positive integer); then, the row elements of the attention matrix are fused to obtain weight information corresponding to each single word in the first word vector; then, the part of the word vector corresponding to each single word in the first word vector can be weighted according to the weight information corresponding to each single word in the first word vector to obtain the first weighted semantic representation.
[0127] The embodiment of the present application can determine the attention matrix by using the attention mechanism. For example, the attention matrix can be determined according to the product of the first word vector and the second word vector, and a normalization function. For another example, the first word vector and the second word vector can be linearly transformed respectively according to a parameter matrix, and the attention matrix can be determined according to the product of the linearly transformed first word vector and the linearly transformed second word vector, and a normalization function. It can be understood that the specific determination manner of the attention matrix is not limited in the embodiment of the present application.
[0128] The matching information can comprise word-dimension matching information and / or sentence-dimension matching information. The word-dimension matching information can be determined according to the first semantic representation of the word dimension and the second semantic representation of the word dimension, and can be represented by a word matching matrix. The sentence-dimension matching information can be determined according to the first semantic representation of the sentence dimension and the second semantic representation of the sentence dimension, and can be represented by a sentence matching matrix.
[0129] The embodiment of the present application can fuse the word-dimension matching information and the sentence-dimension matching information to obtain the matching score between the data object and the standard category information. The fusion can comprise splicing of the word-dimension matching information and the sentence-dimension matching information. The fusion can comprise neural network processing, which can fuse the matching information of multiple dimensions to obtain the matching score.
[0130] In step 203, the standard category information can be sorted according to the matching information, and the target standard category information meeting the preset condition can be obtained according to the sorting result. For example, the standard category information can be sorted in descending order of matching degree. In this way, the preset condition can comprise that the sorting position is in the first Y positions. Of course, the preset condition can also comprise that the matching score is higher than a score threshold.
[0131] The method of the embodiment of the present application can be performed by a data object processing apparatus, which can refer to Figure 4 FIG. 1 shows a structural schematic diagram of a data object processing apparatus according to an embodiment of the present application. The data object processing apparatus can specifically comprise a data analyzer 401, a word processing module 402, a sentence processing module 403, and a fusion processing module 404.
[0132] The data analyzer 401 can be configured to determine the semantic representation corresponding to the data object or the standard category information.
[0133] The word processing module 402 can be configured to process the semantic representation of the word dimension to obtain the word-dimension matching information.
[0134] The sentence processing module 403 can be configured to process the semantic representation of the sentence dimension to obtain the sentence-dimension matching information.
[0135] The fusion processing module 404 can be configured to fuse the word-dimension matching information and the sentence-dimension matching information to obtain the matching score between the data object and the standard category information. The fusion processing module 404 can also determine the target standard category information corresponding to the data object according to the matching score. The embodiment of the present application can determine the matching score according to the word-dimension matching information and the sentence-dimension matching information, which can improve the accuracy of the matching score between the data object and the standard category information.
[0136] Reference Figure 5 The diagram shows a structural schematic of a word processing module 402 according to an embodiment of this application, which may specifically include: an attention matrix determination module 421, a weight determination module 422, a weighting module 423, and a word matching module 424.
[0137] The attention matrix determination module 421 is used to determine the attention matrix corresponding to the first word vector and the second word vector.
[0138] The weight determination module 422 is used to fuse the row elements of the attention matrix to obtain the weight information corresponding to each word in the first word vector.
[0139] The weighting module 423 is used to weight the word vector parts corresponding to each word in the first word vector according to the weight information of each word in the first word vector, so as to obtain the first weighted semantic representation.
[0140] The word matching module 424 is used to determine the matching information of the word dimension based on the first weighted semantic representation of the word dimension and the second semantic representation of the word dimension. The matching information of the word dimension can be represented by the word matching matrix.
[0141] One or more of the word processing module 402, sentence processing module 403, and fusion processing module 404 can employ a neural network structure. Therefore, embodiments of this application can utilize training data to train one or more of the word processing module 402, sentence processing module 403, and fusion processing module 404, and update the corresponding parameters during the training process to improve the processing accuracy of one or more of the word processing module 402, sentence processing module 403, and fusion processing module 404. Since it has already utilized... Figure 1 In the illustrated method embodiment, the data analyzer 401 is trained. Therefore, the training of the data analyzer 401 can be independent of the training of the target modules, such as the word processing module 402, the sentence processing module 403, and the fusion processing module 404. In other words, during the training of the target modules, the data analyzer 401 can provide semantic representations to the target modules without participating in the backpropagation of the target modules; that is, the parameters of the data analyzer 401 do not need to be updated during the training of the target modules.
[0142] The training data of the target module can be the same as or different from the training data of the data analyzer 401. The training data of the target module can include a target sample and a positive sample corresponding to the target sample (hereinafter referred to as a positive sample combination), or a target sample and a negative sample corresponding to the target sample (hereinafter referred to as a negative sample combination). For the positive sample combination, the target value of the matching score output by the fusion processing module 404 can tend to 1; for the negative sample combination, the target value of the matching score output by the fusion processing module 404 can tend to 0. Therefore, according to the error information between the target value and the actual value of the matching score output by the fusion processing module 404, the back propagation of the target module can be performed, so that the update of the parameters of the target module can be realized.
[0143] Referring to Figure 6 , a flowchart of a data object processing method according to an embodiment of the present application is shown, which can include a word dimension processing branch and a sentence dimension processing branch.
[0144] In the word dimension processing branch, first, the description information of the data object is input into the data analyzer, and the corresponding first semantic representation is output by the data analyzer; the second semantic representation corresponding to the standard category information can also be obtained according to the data analyzer or the database; the first semantic representation can include first word vectors of a1, a2, a3, etc. The second semantic representation can include second word vectors of b1, b2, b3, etc. Then, the attention matrix between the first semantic representation and the second semantic representation can be determined, and the first word vectors of a1, a2, a3, etc. are weighted according to the weight information corresponding to a1, a2, a3, etc. to obtain the first weighted semantic representation. Then, the word matching matrix between the first weighted semantic representation and the second semantic representation can be determined.
[0145] In the sentence dimension processing branch, first, the description information of the data object is input into the data analyzer, and the corresponding first semantic representation is output by the data analyzer; the second semantic representation corresponding to the standard category information can also be obtained according to the data analyzer or the database; the first semantic representation can include first word vectors of a1, a2, a3, etc. The second semantic representation can include second word vectors of b1, b2, b3, etc. Then, the first word vectors of a1, a2, a3, etc. are spliced, the second word vectors of b1, b2, b3, etc. are spliced, and the sentence matching matrix between the spliced first word vectors and the spliced second word vectors is determined.
[0146] Further, the word matching matrix and the sentence matching matrix can be input into the fusion processing module, and the fusion processing module can perform fusion processing on the word matching matrix and the sentence matching matrix to obtain the corresponding matching score.
[0147] In summary, the data object processing method of the embodiments of the present application determines the matching information between the data object and the standard category information based on the first semantic representation and the second semantic representation at the semantic level, which is equivalent to performing matching processing of the data object and the standard category information at the semantic level. The above-mentioned matching processing at the semantic level does not require that the target standard category information must appear in the description information of the data object, and thus can improve the classification accuracy of the data object.
[0148] In addition, the standard category information of the embodiments of the present application can be obtained by clustering the semantic representations corresponding to the plurality of preset category information. Since the above-mentioned clustering can unify the plurality of preset category information having similarity into one standard category information, for example, unifying the plurality of preset category information such as “shampoo”, “shampoo”, and “shampoo” having similarity into one standard category information “shampoo”, the embodiments of the present application can provide the unified target standard category information, and thus can improve the classification accuracy of the data object.
[0149] Method embodiment four
[0150] Reference Figure 7 , a step flowchart of a data object processing method of an embodiment of the present application is shown, which can specifically include the following steps:
[0151] Step 701, determining the first semantic representation corresponding to the data object, and the second semantic representation corresponding to the preset category information;
[0152] Step 702, determining the weight information corresponding to the first semantic representation according to the attention information of the second semantic representation to the first semantic representation;
[0153] Step 703, weighting the first semantic representation according to the weight information to obtain the first weighted semantic representation;
[0154] Step 704, determining the matching information between the data object and the standard category information according to the first weighted semantic representation and the second semantic representation;
[0155] Step 705, determining the target category information corresponding to the data object according to the matching information.
[0156] The data object processing method of the embodiments of the present application determines the matching information between the data object and the standard category information based on the first semantic representation and the second semantic representation at the semantic level, which is equivalent to performing matching processing of the data object and the standard category information at the semantic level. The above-mentioned matching processing at the semantic level does not require that the target standard category information must appear in the description information of the data object, and thus can improve the classification accuracy of the data object.
[0157] Further, the embodiment of the present application determines the matching information between the data object and the standard category information according to the first weighted semantic representation and the second semantic representation. Since the first weighted semantic representation is obtained according to the weight information, and the weight information is obtained according to the attention information of the second semantic representation to the first semantic representation, in actual application, a larger weight can be obtained for a part (such as the first word vector) in the first semantic representation that has a higher correlation with the second semantic representation, and a smaller weight can be obtained for a part (such as the second word vector) in the first semantic representation that has a lower correlation with the second semantic representation, so that the first weighted semantic representation can reflect the correlation degree between the first semantic representation and the second semantic representation, thereby improving the accuracy of the first weighted semantic representation and the accuracy of the matching information between the data object and the standard category information.
[0158] Method embodiment five
[0159] Reference Figure 8 The step flowchart of the data object processing method of one embodiment of the present application is shown, which can specifically include the following steps:
[0160] Step 801, determining the first semantic representation corresponding to the data object and the second semantic representation corresponding to the preset category information according to the data analyzer; the data analyzer can be used to represent the mapping relationship between the description information and the semantic representation; the data analyzer is obtained by comparative learning on training data;
[0161] Step 802, determining the matching information between the data object and the standard category information according to the first semantic representation and the second semantic representation;
[0162] Step 803, determining the target category information corresponding to the data object according to the matching information.
[0163] The data object processing method of the embodiment of the present application determines the matching information between the data object and the standard category information based on the first semantic representation and the second semantic representation at the semantic level, which is equivalent to performing semantic-level matching processing on the data object and the standard category information. The above-mentioned semantic-level matching processing does not require that the target standard category information must appear in the description information of the data object, so that the classification accuracy of the data object can be improved.
[0164] Further, the embodiment of the present application obtains the data analyzer by comparative learning on the training data. Comparative learning distinguishes data in the abstract semantic level feature space, has the advantages of simple model and strong generalization ability, so that the accuracy of the semantic representation output by the data analyzer can be improved, and in turn the classification accuracy of the data object can be improved.
[0165] It should be noted that, for the method embodiments, the series of acts combined is described for simplicity, but those skilled in the art should know that the application embodiments are not limited to the order of the acts described, because according to the application embodiments, certain steps can be performed in other orders or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the acts involved are not necessarily the application embodiments.
[0166] On the basis of the above-mentioned embodiments, the application embodiments further provide a data object processing device, which refers to Figure 9 The device can include the following modules:
[0167] The semantic representation determination module 901 is configured to determine a first semantic representation corresponding to the data object and a second semantic representation corresponding to the standard category information; the standard category information is obtained by clustering semantic representations corresponding to a plurality of preset category information according to a semantic representation corresponding to preset category information.
[0168] The matching module 902 is configured to determine matching information between the data object and the standard category information according to the first semantic representation and the second semantic representation.
[0169] The category determination module 903 is configured to determine target standard category information corresponding to the data object according to the matching information.
[0170] Optionally, the device can further include:
[0171] The sample determination module is configured to determine sample data; the sample data includes preset category information and description information corresponding thereto.
[0172] The semantic determination module is configured to determine a semantic representation corresponding to the preset category information according to the description information corresponding to the preset category information.
[0173] The clustering module is configured to cluster semantic representations corresponding to a plurality of preset category information to obtain a standard category information corresponding to the plurality of preset category information.
[0174] Optionally, the matching module 902 can specifically include:
[0175] The weight determination module is configured to determine weight information corresponding to the first semantic representation according to attention information of the second semantic representation to the first semantic representation.
[0176] The weighting module is configured to weight the first semantic representation according to the weight information to obtain a first weighted semantic representation.
[0177] The matching determining module is configured to determine matching information between the data object and the standard category information according to the first weighted semantic representation and the second semantic representation.
[0178] Optionally, the first semantic representation is obtained according to a data analyzer, and the second semantic representation is obtained according to the data analyzer; the data analyzer is configured to represent a mapping relationship between description information and a semantic representation; and the data analyzer is obtained by contrast learning on training data.
[0179] Optionally, the training data can include a target sample, a positive sample corresponding to the target sample, and a negative sample corresponding to the target sample; the target sample includes a first sample corresponding to first preset category information; the positive sample includes a second sample corresponding to the first preset category information; the negative sample includes a third sample corresponding to second preset category information; and the first preset category information is different from the second preset category information.
[0180] The embodiment of the present application further provides a data object processing device, referring to Figure 10 The device can include:
[0181] The semantic representation determining module 1001 is configured to determine a first semantic representation corresponding to a data object and a second semantic representation corresponding to preset category information.
[0182] The weight determining module 1002 is configured to determine weight information corresponding to the first semantic representation according to attention information of the second semantic representation to the first semantic representation.
[0183] The weighting module 1003 is configured to weight the first semantic representation according to the weight information to obtain a first weighted semantic representation.
[0184] The matching module 1004 is configured to determine matching information between the data object and the standard category information according to the first weighted semantic representation and the second semantic representation.
[0185] The category determining module 1005 is configured to determine target category information corresponding to the data object according to the matching information.
[0186] The embodiment of the present application further provides a data object processing device, referring to Figure 11 The device can include:
[0187] The semantic representation determination module 1101 is configured to determine a first semantic representation corresponding to a data object and a second semantic representation corresponding to preset category information according to a data analyzer. The data analyzer is used to represent a mapping relationship between description information and a semantic representation. The data analyzer can be obtained by contrastive learning on training data.
[0188] The matching module 1102 is configured to determine matching information between the data object and the standard category information according to the first semantic representation and the second semantic representation.
[0189] The category determination module 1103 is configured to determine target category information corresponding to the data object according to the matching information.
[0190] The embodiments of the present application further provide a non-volatile readable storage medium, which stores one or more programs. When the one or more programs are applied to a device, the device can execute instructions of each method step in the embodiments of the present application.
[0191] The embodiments of the present application provide one or more machine readable media, which store instructions. When executed by one or more processors, the instructions cause an electronic device to perform a method according to any one of the above embodiments. In the embodiments of the present application, the electronic device includes a server, a terminal device, and the like.
[0192] Embodiments of the present disclosure can be implemented as an apparatus configured in a desired manner using any appropriate hardware, firmware, software, or any combination thereof, which can include a server (cluster), a terminal, and the like. Figure 12 An exemplary apparatus 1300 that can be used to implement various embodiments described herein is shown schematically.
[0193] For one embodiment, Figure 12 An exemplary apparatus 1300 is shown having one or more processors 1302, a control module (chipset) 1304 coupled to at least one of the processor(s) 1302, a memory 1306 coupled to the control module 1304, a non-volatile memory (NVM) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304.
[0194] The processor 1302 can include one or more single core or multi core processors, which can include any combination of general-purpose processors or dedicated processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, the apparatus 1300 can be capable of acting as a server, a terminal, or the like, as described in embodiments herein.
[0195] In some embodiments, the apparatus 1300 can include one or more computer readable media (e.g., the memory 1306 or the NVM / storage 1308) with instructions 1314 and one or more processors 1302 incorporated with the one or more computer readable media configured to execute the instructions 1314 to implement modules to perform the actions described in the present disclosure.
[0196] For one embodiment, the control module 1304 can include any suitable interface controllers to provide for any suitable interface to at least one of the processor(s) 1302 and / or any suitable device or component in communication with the control module 1304.
[0197] The control module 1304 can include a memory controller module to provide an interface to the memory 1306. The memory controller module can be a hardware module, a software module, and / or a firmware module.
[0198] The memory 1306 can be used to, for example, load and store data and / or instructions 1314 for the apparatus 1300. For one embodiment, the memory 1306 can include any suitable volatile memory, such as suitable DRAM. In some embodiments, the memory 1306 can include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).
[0199] For one embodiment, the control module 1304 can include one or more input / output controllers to provide an interface to the NVM / storage 1308 and the input / output device(s) 1310.
[0200] The NVM / storage 1308 can be used, for example, to store data and / or instructions 1314. The NVM / storage 1308 can include any suitable non-volatile memory (e.g., flash memory) and / or can include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).
[0201] The NVM / storage 1308 can include storage resources of a device on which the apparatus 1300 is installed as part of the device, or it can be accessed by the device and not necessarily part of the device. For example, the NVM / storage 1308 can be accessed over a network via the input / output device(s) 1310.
[0202] The input / output device(s) 1310 can provide an interface between the apparatus 1300 and any other suitable device, and can include communication components, audio components, sensor components, and the like. The network interface 1312 can provide an interface between the apparatus 1300 and one or more networks, and the apparatus 1300 can communicate wirelessly with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as to access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, 5G, and the like, or combinations thereof.
[0203] For one embodiment, at least one of the processor(s) 1302 can be packaged together with logic of one or more controllers of the control module 1304, such as a memory controller module. For one embodiment, at least one of the processor(s) 1302 can be packaged together with logic of one or more controllers of the control module 1304 to form a system in a package (SiP). For one embodiment, at least one of the processor(s) 1302 can be fabricated together with logic of one or more controllers of the control module 1304 on the same die. For one embodiment, at least one of the processor(s) 1302 can be fabricated together with logic of one or more controllers of the control module 1304 on the same die to form a system on a chip (SoC).
[0204] In various embodiments, the apparatus 1300 can be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet, a netbook, etc.). In various embodiments, the apparatus 1300 can have more or less components, and / or different architectures. For example, in some embodiments, the apparatus 1300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0205] In various embodiments, the apparatus 1300 can include a master chip as a processor or control module, sensor data, location information, etc. can be stored in a memory or NVM / storage device, a sensor group can be an input / output device, and a communication interface can include a network interface.
[0206] The embodiment of the present application further provides an electronic device, comprising: a processor; and a memory having executable codes stored thereon, which, when executed, cause the processor to perform the method according to any one or more of the embodiments of the present application.
[0207] The embodiment of the present application further provides one or more machine readable medium having executable codes stored thereon, which, when executed, cause a processor to perform the method according to any one or more of the embodiments of the present application.
[0208] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts are described in the part of the method embodiment.
[0209] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same parts between the embodiments can be referred to each other.
[0210] The embodiments of the present application are described with reference to flowcharts and / or block diagrams according to the method, terminal device (system), and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the computer or other programmable data processing terminal device produce a device for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows and / or blocks
[0211] These computer program instructions can also be stored in a computer readable memory capable of guiding the computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer readable memory produce a product including instruction devices, which implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows and / or blocks
[0212] These computer program instructions can also be loaded into a computer or other programmable data processing terminal device, so that a series of operation steps are performed on the computer or other programmable terminal device to produce a computer implemented process, so that the instructions executed on the computer or other programmable terminal device provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows and / or blocks
[0213] Although preferred embodiments of the application have been described in detail, those skilled in the art will appreciate that various modifications and alterations can be made to the embodiments without departing from the scope of the application. Accordingly, the appended claims are intended to encompass all such modifications and alterations. In particular, the application is intended to cover any and all combinations of the features set forth in the claims, and any and all combinations and sub-combinations of the features set forth in the detailed description above.
[0214] Finally, it should be noted that the terms "first", "second", and the like, herein do not denote any order, quantity, combination, or importance, but rather are used to distinguish one element from another, and are not intended to denote a physical or logical relationship between such elements. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0215] The above provides a data object processing method, a data object processing device, an electronic device and a storage medium. The principles and implementation manners of the application are described by using specific examples. The above description of the embodiments is only used to help understand the method and the core idea of the application. For those skilled in the art, according to the idea of the application, the specific implementation manners and application scopes can be changed. The above description of the application should not be understood as a limitation.
Claims
1. A method of processing data objects, characterized by, The method comprises: determining a first semantic representation corresponding to a data object and a second semantic representation corresponding to standard category information; the standard category information is obtained by clustering semantic representations corresponding to a plurality of preset category information according to a semantic representation corresponding to a preset category information; determining matching information between the data object and the standard category information according to the first semantic representation and the second semantic representation; determining target standard category information corresponding to the data object according to the matching information; wherein the determination of the matching information between the data object and the standard category information comprises: determining weight information corresponding to the first semantic representation according to attention information of the second semantic representation to the first semantic representation; weighting the first semantic representation according to the weight information to obtain a first weighted semantic representation; and determining the matching information between the data object and the standard category information according to the first weighted semantic representation and the second semantic representation.
2. The method of claim 1, wherein, The method further comprises: determining sample data; the sample data comprises: preset category information and description information corresponding thereto; determining a semantic representation corresponding to a preset category information according to description information corresponding to the preset category information; clustering semantic representations corresponding to a plurality of preset category information to obtain a standard category information corresponding to the plurality of preset category information.
3. The method according to claim 1 or 2, characterized in that, The first semantic representation is obtained according to a data analyzer, and the second semantic representation is obtained according to the data analyzer; the data analyzer is used to represent a mapping relationship between description information and semantic representation; the data analyzer is obtained by comparative learning on training data.
4. The method of claim 3, wherein, The training data can include: a target sample, a positive sample corresponding to the target sample, and a negative sample corresponding to the target sample; wherein the target sample includes a first sample corresponding to a first preset category information; the positive sample includes a second sample corresponding to the first preset category information; the negative sample includes a third sample corresponding to a second preset category information; the first preset category information is different from the second preset category information.
5. A method of processing data objects, characterized by, The method comprises: determining a first semantic representation corresponding to a data object and a second semantic representation corresponding to a preset category information; determining weight information corresponding to the first semantic representation according to attention information of the second semantic representation to the first semantic representation; weighting the first semantic representation according to the weight information to obtain a first weighted semantic representation; determining matching information between the data object and standard category information according to the first weighted semantic representation and the second semantic representation; the standard category information is obtained by clustering semantic representations corresponding to a plurality of preset category information according to a semantic representation corresponding to a preset category information; determining target category information corresponding to the data object according to the matching information.
6. A method of processing data objects, characterized by, The method comprises: According to the data analyzer, determine the first semantic representation corresponding to the data object, and the second semantic representation corresponding to the preset category information; the data analyzer is used to represent the mapping relationship between the description information and the semantic representation; the data analyzer is obtained by comparing the training data; According to the first semantic representation and the second semantic representation, determine the matching information between the data object and the standard category information; the standard category information is obtained by clustering the semantic representations corresponding to a plurality of preset category information according to the semantic representation corresponding to the preset category information; According to the matching information, determine the target category information corresponding to the data object; Wherein, the determination of the matching information between the data object and the standard category information comprises: according to the attention information of the second semantic representation to the first semantic representation, determining the weight information corresponding to the first semantic representation; according to the weight information, weighting the first semantic representation to obtain the first weighted semantic representation; according to the first weighted semantic representation and the second semantic representation, determine the matching information between the data object and the standard category information.
7. An apparatus for processing of data objects, characterized by The device comprises: The semantic representation determination module is used to determine the first semantic representation corresponding to the data object, and the second semantic representation corresponding to the standard category information; the standard category information is obtained by clustering the semantic representations corresponding to a plurality of preset category information according to the semantic representation corresponding to the preset category information; The matching module is used to determine the matching information between the data object and the standard category information according to the first semantic representation and the second semantic representation; The category determination module is used to determine the target standard category information corresponding to the data object according to the matching information; Wherein, the matching module comprises: The weight determination module is used to determine the weight information corresponding to the first semantic representation according to the attention information of the second semantic representation to the first semantic representation; The weighting module is used to weight the first semantic representation according to the weight information to obtain the first weighted semantic representation; The matching determination module is used to determine the matching information between the data object and the standard category information according to the first weighted semantic representation and the second semantic representation.
8. An electronic device, comprising: Comprise: Processor; And Memory, which has stored executable code, when the executable code is executed, make the processor execute the method as claimed in any one of claims 1-6.
9. One or more machine readable media having stored thereon executable code that, when executed, cause a processor to perform the method as claimed in any one of claims 1-6.
9. One or more machine readable media having stored thereon executable code that, when executed, cause a processor to perform the method as claimed in any one of claims 1-6.
Citation Information
Patent Citations
Custom declaration commodity information processing method and system, storage medium and electronic equipment
CN111881265A
Object processing method and device, electronic equipment and storage medium
CN112434154A