A method for determining multi-level categories, a method for model training, and related devices
Through the first and second classifiers in the hierarchical classification model, the distribution and fusion vectors are generated using text encoding vectors, which solves the problem of insufficient classification accuracy of multi-category texts in the prior art, and achieves a higher classification accuracy.
Patent Information
- Application Number
- CN202111114531.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-23
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-09-23
AI Technical Summary
The prior art still has room for improvement in the classification accuracy based on keywords or neural network models in text information classification, especially when facing multi-category text.
By obtaining text encoding vectors, using the first and second classifiers in the hierarchical classification model, the distribution vector and the fusion vector are generated, and the next level of categories are predicted based on the output of the previous level, and the upper and lower levels of the category system are fully utilized.
The accuracy of category classification has been enhanced, and the classification effect of secondary categories has been improved by focusing on the prediction results of first-level categories.
Smart Images

Figure CN114328906B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method for determining multi-level categories, a method for model training, and related devices. Background Art
[0002] With the continuous development of computer technology, the amount of information faced by users is increasing day by day. When users face a large amount of information, they can obtain information related to the keyword by searching for the keyword. For example, when a user enters the keyword "football" on a video platform, based on this, the video platform can provide videos related to "football".
[0003] The premise of search and recommendation is to classify information. Therefore, in traditional solutions, the category to which the text belongs can be directly determined by extracting keywords in the text information, or a trained neural network model can be used to perform semantic recognition on the input text information, thereby realizing the classification of information.
[0004] The inventors found that there are at least the following problems in the prior art. The limitation of classification based on keywords is relatively large. If the text information involves keywords in different categories, it is easy to cause classification errors. Although information classification based on a neural network model can improve the accuracy of classification to a certain extent, there is still a large room for improvement in classification accuracy. Summary of the Invention
[0005] Embodiments of this application provide a method for determining multi-level categories, a method for model training, and related devices. This application predicts the output of the next level based on the output result of the previous level, and can fully and effectively utilize the constraint relationship between the upper and lower layers in the category system. Thereby enhancing the effect of category classification and improving the classification accuracy.
[0006] In view of this, on the one hand, this application provides a method for determining multi-level categories, including:
[0007] Obtain the text encoding vector corresponding to the target text information;
[0008] Based on the text encoding vector, obtain the first distribution vector through the first classifier included in the hierarchical classification model, where the first distribution vector includes M first element scores, and each first element score represents a probability value of a first-level category, and M is an integer greater than 1;
[0009] Generate a text fusion vector according to the first distribution vector and the text encoding vector;
[0010] Based on the text fusion vector, obtain a second distribution vector through the second classifier included in the hierarchical classification model, where the second distribution vector includes N second element scores, and each second element score represents a probability value of a secondary category, and the secondary category belongs to the next-level category of the primary category, and N is an integer greater than 1;
[0011] Determine the target primary category to which the target text information belongs according to the first distribution vector, and determine the target secondary category to which the target text information belongs according to the second distribution vector.
[0012] On the other hand, the present application provides a method for model training, including:
[0013] Obtain the predicted text encoding vector corresponding to the text information to be trained, where the text information to be trained corresponds to a primary annotation category and a secondary annotation category;
[0014] Based on the predicted text encoding vector, obtain a first predicted distribution vector through the first classifier to be trained included in the hierarchical classification model to be trained, where the first predicted distribution vector includes M first element scores, and each first element score represents a probability value of a primary category, and M is an integer greater than 1;
[0015] Generate a predicted text fusion vector according to the first predicted distribution vector and the predicted text encoding vector;
[0016] Based on the predicted text fusion vector, obtain a second predicted distribution vector through the second classifier to be trained included in the hierarchical classification model to be trained, where the second predicted distribution vector includes N second element scores, and each second element score represents a probability value of a secondary category, and the secondary category belongs to the next-level category of the primary category, and N is an integer greater than 1;
[0017] Update the model parameters of the hierarchical classification model to be trained according to the first predicted distribution vector, the second predicted distribution vector, the primary annotation category, and the secondary annotation category until the model training condition is satisfied, and output the hierarchical classification model, where the hierarchical classification model includes the first classifier and the second classifier involved in the above aspects.
[0018] On the other hand, the present application provides a multi-level category determination device, including:
[0019] An acquisition module, configured to acquire the text encoding vector corresponding to the target text information;
[0020] The acquisition module is further configured to obtain a first distribution vector through the first classifier included in the hierarchical classification model based on the text encoding vector, where the first distribution vector includes M first element scores, and each first element score represents a probability value of a primary category, and M is an integer greater than 1;
[0021] A generation module, configured to generate a text fusion vector according to a first distribution vector and a text encoding vector;
[0022] The acquisition module is further configured to obtain a second distribution vector based on the text fusion vector through a second classifier included in the hierarchical classification model, where the second distribution vector includes N second element scores, and each second element score represents a probability value of a secondary category, and the secondary category belongs to a next-level category of a primary category, and N is an integer greater than 1;
[0023] A determination module, configured to determine a target primary category to which the target text information belongs according to the first distribution vector, and determine a target secondary category to which the target text information belongs according to the second distribution vector.
[0024] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0025] The generation module is specifically configured to generate a prior semantic vector according to the first distribution vector based on a primary category vector mapping relationship, where the primary category vector mapping relationship includes a one-to-one mapping relationship between M index values and M semantic vectors, and each index value corresponds to a primary category;
[0026] Generate a text fusion vector according to the prior semantic vector and the text encoding vector.
[0027] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0028] The generation module is specifically configured to determine the top K first element scores with the largest probability values from the first distribution vector, where each first element score corresponds to an index value of a primary category, and K is an integer greater than 1 and less than M;
[0029] Obtain the index values corresponding to each of the top K first element scores to obtain K index values;
[0030] Based on the primary category vector mapping relationship, obtain the corresponding K semantic vectors according to the K index values;
[0031] For each of the K index values, perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value to obtain an updated semantic vector corresponding to the index value;
[0032] Perform a summation calculation on the updated semantic vectors corresponding to the K index values to obtain a prior semantic vector.
[0033] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0034] A generation module, specifically configured to determine a first element score with the largest probability value from a first distribution vector, where the first element score corresponds to an index value of a first-level category;
[0035] If the first element score is greater than or equal to an element score threshold, obtain the index value corresponding to the first element score;
[0036] Based on the first-level category vector mapping relationship, obtain the corresponding semantic vector according to the index value;
[0037] Perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value to obtain a prior semantic vector.
[0038] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0039] A generation module, specifically configured to obtain the index values corresponding to each first element score in the first distribution vector to obtain M index values;
[0040] Based on the first-level category vector mapping relationship, obtain M corresponding semantic vectors according to the M index values;
[0041] For each of the M index values, perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value to obtain an updated semantic vector corresponding to the index value;
[0042] Perform a summation calculation on the updated semantic vectors corresponding to the K index values to obtain a prior semantic vector.
[0043] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0044] A generation module, specifically configured to splice the prior semantic vector and the text encoding vector to obtain a text fusion vector;
[0045] Or,
[0046] A generation module, specifically configured to determine a first-order vector mapping according to the prior semantic vector, the text encoding vector, and a first parameter matrix, where the prior semantic vector is represented as a p-dimensional vector, the text encoding vector is represented as a q-dimensional vector, and the first parameter matrix is represented as a [d*(p + q)]-dimensional matrix, and p, q, and d are all integers greater than 1;
[0047] Determine a second-order vector mapping according to the prior semantic vector, the text encoding vector, and a second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix;
[0048] Generate a text fusion vector based on a first-order vector mapping, a second-order vector mapping, and a bias vector.
[0049] In a possible design, in another implementation of another aspect of the embodiments of the present application,
[0050] The generation module is specifically configured to splice the first distribution vector and the text encoding vector to obtain a text fusion vector;
[0051] Or,
[0052] The generation module is specifically configured to determine a first-order vector mapping according to the first distribution vector, the text encoding vector, and a first parameter matrix, where the first distribution vector is represented as a p-dimensional vector, the text encoding vector is represented as a q-dimensional vector, and the first parameter matrix is represented as a [d*(p + q)]-dimensional matrix, and p, q, and d are all integers greater than 1;
[0053] Determine a second-order vector mapping according to the first distribution vector, the text encoding vector, and a second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix;
[0054] Generate a text fusion vector based on the first-order vector mapping, the second-order vector mapping, and the bias vector.
[0055] In a possible design, in another implementation of another aspect of the embodiments of the present application,
[0056] Before the acquisition module is further configured to acquire the text encoding vector corresponding to the target text information, acquire the target text information corresponding to the target video, where the target text information includes at least one of the title information, summary information, subtitle information, and comment information of the target video;
[0057] Or,
[0058] Before the acquisition module is further configured to acquire the text encoding vector corresponding to the target text information, acquire the target text information corresponding to the target picture, where the target text information includes at least one of the title information, author information, optical character recognition (OCR) information, and summary information of the target picture;
[0059] Or,
[0060] Before the acquisition module is further configured to acquire the text encoding vector corresponding to the target text information, acquire the target text information corresponding to the target commodity, where the target text information includes at least one of the commodity name information, origin information, comment information, and commodity description information of the target commodity;
[0061] Or,
[0062] The obtaining module is further configured to obtain the target text information corresponding to the target text before obtaining the text encoding vector corresponding to the target text information, where the target text information includes at least one of the title information, author information, abstract information, comment information, and body information of the target text.
[0063] In a possible design, in another implementation of another aspect of the present application, the multi-level category determination device further includes a receiving module and a sending module;
[0064] The receiving module is configured to receive a category query instruction sent by the terminal device for the content to be searched;
[0065] The sending module is configured to, in response to the category query instruction, if the content to be searched is video content, send a video search result to the terminal device;
[0066] The sending module is further configured to, in response to the category query instruction, if the content to be searched is picture content, send a picture search result to the terminal device;
[0067] The sending module is further configured to, in response to the category query instruction, if the content to be searched is commodity content, send a commodity search result to the terminal device;
[0068] The sending module is further configured to, in response to the category query instruction, if the content to be searched is text content, send a text search result to the terminal device.
[0069] Another aspect of the present application provides a model training device, including:
[0070] The obtaining module is configured to obtain the predicted text encoding vector corresponding to the text information to be trained, where the text information to be trained corresponds to a first-level labeled category and a second-level labeled category;
[0071] The obtaining module is further configured to, based on the predicted text encoding vector, obtain a first predicted distribution vector through a first classifier to be trained included in the hierarchical classification model to be trained, where the first predicted distribution vector includes M first element scores, and each first element score represents a probability value of a first-level category, and M is an integer greater than 1;
[0072] The generating module is configured to generate a predicted text fusion vector according to the first predicted distribution vector and the predicted text encoding vector;
[0073] The obtaining module is further configured to, based on the predicted text fusion vector, obtain a second predicted distribution vector through a second classifier to be trained included in the hierarchical classification model to be trained, where the second predicted distribution vector includes N second element scores, and each second element score represents a probability value of a second-level category, the second-level category belongs to the next level of the first-level category, and N is an integer greater than 1;
[0074] A training module, configured to update the model parameters of the hierarchical classification model to be trained according to the first prediction distribution vector, the second prediction distribution vector, the first-level annotation categories, and the second-level annotation categories until the model training conditions are met, and output the hierarchical classification model, where the hierarchical classification model includes the first classifier and the second classifier involved in the above aspects.
[0075] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0076] The training module is specifically configured to calculate the first loss value of the text information to be trained by using a first classification loss function according to the first prediction distribution vector and the first-level annotation categories;
[0077] Calculate the second loss value of the text information to be trained by using a second classification loss function according to the second prediction distribution vector and the second-level annotation categories;
[0078] Determine the comprehensive loss value of the text information to be trained according to the first loss value and the second loss value;
[0079] Update the model parameters of the hierarchical classification model to be trained according to the comprehensive loss value.
[0080] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0081] The training module is specifically configured to calculate the first loss value of the text information to be trained by using a first classification loss function according to the first prediction distribution vector and the first-level annotation categories;
[0082] Calculate the second loss value of the text information to be trained by using a second classification loss function according to the second prediction distribution vector and the second-level annotation categories;
[0083] Determine the first element prediction score corresponding to the first-level annotation categories from the first prediction distribution vector, and determine the second element prediction score corresponding to the second-level annotation categories from the second prediction distribution vector;
[0084] Calculate the third loss value of the text information to be trained by using a hinge loss function according to the first element prediction score, the second element prediction score, and the target hyperparameter;
[0085] Determine the comprehensive loss value of the text information to be trained according to the first loss value, the second loss value, and the third loss value;
[0086] Update the model parameters of the hierarchical classification model to be trained according to the comprehensive loss value.
[0087] On the other hand, the present application provides a computer device, including: a memory, a processor, and a bus system;
[0088] The memory is used to store programs;
[0089] The processor is used to execute the programs in the memory, and the processor is used to execute the methods provided in the above aspects according to the instructions in the program code;
[0090] The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.
[0091] On the other hand, the present application provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the methods provided in the above aspects.
[0092] In another aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the methods provided in the above aspects.
[0093] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:
[0094] In the embodiments of the present application, a method for determining multi-level categories is provided. First, a text encoding vector corresponding to target text information is obtained, and then the text encoding vector is used as the input of a first classifier, thereby obtaining a first distribution vector. The first classifier is part of a hierarchical classification model, and the hierarchical classification model further includes a second classifier. Next, according to the first distribution vector and the text encoding vector, a text fusion vector is generated, and then the text fusion vector is used as the output of the second classifier, thereby obtaining a second distribution vector. Finally, based on the first distribution vector, the target first-level category to which the target text information belongs is determined, and based on the second distribution vector, the target second-level category to which the target text information belongs is determined. In the above manner, the prediction result corresponding to the first-level category is used as prior knowledge, and after fusing the text encoding vector, a text fusion vector is obtained. The text fusion vector is used as the basis for predicting the second-level category, that is, the output result of the upper level is used to predict the output of the lower level, which can fully and effectively utilize the constraint relationship between the upper and lower levels in the category system. Thus, when the second classifier makes a prediction, it can pay more attention to the second-level categories related to the prediction result of the first-level category, thereby enhancing the category classification effect and improving the classification accuracy. Description of the Drawings
[0095] Figure 1It is a schematic architecture diagram of the multi-level category determination system in the embodiment of this application;
[0096] Figure 2 It is a schematic diagram for determining multi-level categories based on a hierarchical classification model in the embodiment of this application;
[0097] Figure 3 It is a schematic flowchart of the multi-level category determination method in the embodiment of this application;
[0098] Figure 4 It is a schematic diagram for encoding text information based on an encoder in the embodiment of this application;
[0099] Figure 5 It is a schematic diagram of the video content hierarchical classification task in the embodiment of this application;
[0100] Figure 6 It is a schematic diagram of the picture content hierarchical classification task in the embodiment of this application;
[0101] Figure 7 It is a schematic diagram of the commodity content hierarchical classification task in the embodiment of this application;
[0102] Figure 8 It is a schematic diagram of the text content hierarchical classification task in the embodiment of this application;
[0103] Figure 9 It is a schematic diagram of an interface for displaying video search results in the embodiment of this application;
[0104] Figure 10 It is a schematic diagram of an interface for displaying picture search results in the embodiment of this application;
[0105] Figure 11 It is a schematic diagram of an interface for displaying commodity search results in the embodiment of this application;
[0106] Figure 12 It is a schematic diagram of an interface for displaying text search results in the embodiment of this application;
[0107] Figure 13 It is a schematic flowchart of the model training method in the embodiment of this application;
[0108] Figure 14 It is a schematic diagram of the multi-level category determination device in the embodiment of this application;
[0109] Figure 15 It is a schematic diagram of the model training device in the embodiment of this application;
[0110] Figure 16 It is a schematic structure diagram of the server in the embodiment of this application;
[0111] Figure 17 This is a schematic structural diagram of a terminal device in an embodiment of the present application. Detailed implementation manners
[0112] The embodiments of the present application provide a method for determining multi-level categories, a method for model training, and related devices. Based on the output results of the upper level, the present application predicts the output of the lower level, and can fully and effectively utilize the constraint relationship between the upper and lower levels in the category system. Thereby enhancing the effect of category classification and improving the classification accuracy.
[0113] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0114] With the continuous development of Internet technology, more and more network interaction platforms have emerged. These network interaction platforms have provided great convenience for people's daily lives. At the same time, they have also increased the difficulty of content integration on the network interaction platforms. Therefore, the content on the network interaction platforms can be classified to facilitate tasks such as content search and content recommendation. Exemplarily, in a network video platform, users can view interesting video content according to the video classification results. Exemplarily, in a network e-commerce platform, users can purchase goods according to the product classification results. Exemplarily, in a network game platform, users can play video games according to the game classification results. Exemplarily, in a network education platform, users can study courses according to the course classification results. Exemplarily, in a network e-book platform, users can read articles according to the e-book classification results.
[0115] In order to achieve more accurate content classification in the above scenarios, the present application proposes a method for determining multi-level categories, which is applied to Figure 1The multi-level category determination system shown in the figure, as shown, the multi-level category determination system includes a server and a terminal device, and the client is deployed on the terminal device. Among them, the client can run on the terminal device in the form of a browser, or can also run on the terminal device in the form of an independent application (APP), etc. For the specific display form of the client, it is not limited here. The server involved in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), as well as big data and artificial intelligence platforms. The terminal device includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, and in-vehicle terminals, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here. The number of the server and the terminal device is also not limited. The solution provided by this application can be completed independently by the terminal device, or can be completed independently by the server, or can also be completed by the cooperation of the terminal device and the server. Regarding this, this application does not make specific limitations. The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, and assisted driving, etc.
[0116] Based on Figure 1 The level category determination system shown in the figure. Specifically, a large amount of content to be classified (such as video content, picture content, commodity content, and text content, etc.) is stored in the database. These contents to be classified are used as the input of the hierarchical classification model. Thus, the multi-level categories of the content to be classified are obtained. Among them, hierarchical classification is a special type of text classification task, that is, there is a hierarchical structure relationship between categories, which can generally be represented as a tree or an undirected graph. And the multi-level category at least includes a first-level category and a second-level category. Based on this, the content and its corresponding multi-level categories are stored in the category mapping table, and the server of the network interaction platform can call the category mapping table. The user selects a multi-level category through the terminal device. Then, the server pushes relevant content to the terminal device used by the user according to the category mapping table and the multi-level category selected by the user.
[0117] The task of classifying content based on a hierarchical classification model specifically involves computer vision (CV) technology, natural language processing (NLP) technology, and machine learning (ML) technology based on artificial intelligence (AI), etc. Among them, AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, AI is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. AI technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. AI basic technologies generally include technologies such as sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. AI software technologies mainly include several major directions such as CV technology, speech processing technology, NLP technology, and ML / deep learning, autonomous driving, and intelligent transportation.
[0118] Among them, CV is a science that studies how to make machines "see". Further speaking, it refers to machine vision that uses cameras and computers to replace human eyes to identify, trace, and measure targets, and further performs graphic processing to make the images processed by the computer more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, CV studies relevant theories and technologies and attempts to build an AI system that can obtain information from images or multi-dimensional data. CV technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0119] Among them, NLP is an important direction in the fields of computer science and AI. It studies various theories and methods that can enable effective communication between humans and computers in natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. NLP technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph, and other technologies.
[0120] Among them, ML is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. ML is the core of AI and the fundamental way to make computers intelligent, and its applications cover all fields of AI. ML and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and formal teaching learning.
[0121] The following will combine Figure 2 to introduce the task execution process of the hierarchical classification model. Please refer to Figure 2 , Figure 2 This is a schematic diagram for determining multiple levels of categories based on the hierarchical classification model in the embodiments of the present application. As shown in the figure, the hierarchical classification model uses a multi-classification model (i.e., the first classifier) for both the first-level categories and the second-level categories. The target text information corresponding to the content to be classified is input into the encoder to obtain a text encoding vector. The text encoding vector is input into the first classifier, and then the first classifier outputs the first distribution vector. Based on the first distribution vector, K semantic vectors corresponding to K first element scores can be determined. After fusing the K semantic vectors with the text encoding vector, a text fusion vector is obtained. The text fusion vector is input into the second classifier, and then the second classifier outputs the second distribution vector. Combining the first distribution vector and the second distribution vector can determine the multiple levels of categories of the content to be classified. It can be seen that the categories have a hierarchical relationship, where the upper-level category is the parent of the lower-level category, and the finer the granularity is as it goes down to the lower levels.
[0122] Combined with the above introduction, the solution provided in the embodiments of the present application involves technologies such as NLP, CV, and ML of AI. The following will introduce the method for determining multiple levels of categories in the present application. Please refer to Figure 3 , an embodiment of the method for determining multiple levels of categories in the embodiments of the present application includes:
[0123] 110. Obtain the text encoding vector corresponding to the target text information;
[0124] In one or more embodiments, the multi-level category determination device obtains target text information, which may be sourced from videos, pictures, articles, etc., and the target text information may be a title sentence, paragraph, etc., without limitation here.
[0125] Specifically, assume the target text information is "Jump and learn the strategies to get 600 points", and use this target text information as the input to the encoder. Whether word segmentation is required for the input or whether it needs to be directly input at the character granularity mainly depends on the type of the encoder. For example, if the encoder uses Convolutional Neural Networks (CNN) or Long-Short Term Memory (LSTM), word segmentation is mainly used. Another example is that if the encoder uses the Bidirectional Encoder Representations from Transformers (BERT) model, the character granularity is mainly used.
[0126] Taking the encoder using the BERT model as an example for introduction, for easy understanding, please refer to Figure 4 , Figure 4 This is a schematic diagram for encoding text information based on the encoder in the embodiments of the present application. As shown in the figure, the target text information is input into the trained BERT model in the form of character granularity, and after encoding, a text encoding vector is generated. Taking the semantic vector of each character as 768 dimensions as an example, usually the output vector (768 dimensions) of the first character "CLS" is taken as the vector representation of the entire target text information. Therefore, the final text encoding vector is also 768 dimensions. That is, l1_emb = BERT(text), where l1_emb represents the text encoding vector and text represents the target text information.
[0127] It should be noted that the multi-level category determination device may be deployed on the server, or may be deployed on the terminal device, or may be deployed in a system composed of the terminal device and the server, without limitation here.
[0128] 120. Based on the text encoding vector, obtain the first distribution vector through the first classifier included in the hierarchical classification model, where the first distribution vector includes M first element scores, and each first element score represents a probability value of a first-level category, and M is an integer greater than 1;
[0129] In one or more embodiments, the multi-level category determination device inputs the text encoding vector into the first classifier (classify1) included in the hierarchical classification model, and the first classifier obtains the first distribution vector (logits1). That is, logits1 = classify1(l1_emb). Among them, the first distribution vector is the prediction result of the first-level category, and the first distribution vector includes M first element scores, where M represents the total number of first-level categories, and each first element score represents the probability value of the corresponding first-level category.
[0130] Specifically, taking M as 5 as an example, and assuming that the first distribution vector is (0, 0.1, 0.7, 0.2, 0), where the first first element score "0" represents that the probability value of belonging to "Games" is "0", the first element score "0.1" represents that the probability value of belonging to "Dance" is "0.1", the first element score "0.7" represents that the probability value of belonging to "Technology" is "0.7", the first element score "0.2" represents that the probability value of belonging to "Nature" is "0.2", and the last first element score "0" represents that the probability value of belonging to "Sports" is "0". It can be seen that the probability that the target text information belongs to "Technology" is the largest.
[0131] 130. Generate a text fusion vector according to the first distribution vector and the text encoding vector;
[0132] In one or more embodiments, the classification results obtained by the high-level model (for example, the first classifier) are generally accurate. Therefore, the first distribution vector can be used as prior knowledge and input into the next-level model (for example, the second classifier).
[0133] Specifically, the multi-level category determination device combines the first distribution vector and the text encoding vector to generate a text fusion vector. That is, the text fusion vector has feature vectors from two dimensions. One is the text encoding vector of the target text information itself, and the other is the first distribution vector output by the first classifier. Therefore, the fusion of multi-dimensional feature vectors is required here.
[0134] 140. Based on the text fusion vector, obtain the second distribution vector through the second classifier included in the hierarchical classification model, where the second distribution vector includes N second element scores, and each second element score represents the probability value of a second-level category. The second-level category belongs to the next level of the first-level category, and N is an integer greater than 1;
[0135] In one or more embodiments, the multi-level category determination device inputs the text fusion vector (L2_logits) into the second classifier (classify2) included in the hierarchical classification model, and the second classifier obtains the second distribution vector (logits2). That is, logits2 = classify2(L2_logits), where the second distribution vector is the prediction result of the secondary category, and the second distribution vector includes N second element scores, N represents the total number of secondary categories, and each second element score represents the probability value of the corresponding secondary category.
[0136] It should be noted that the secondary category belongs to the next-level category (i.e., subcategory) of the primary category. Usually, the value of M is less than the value of N. For example, there are 44 primary categories, including thematic coarse-grained categories such as "Sports", "Games", and "Entertainment", and each primary category can be further divided into multiple secondary categories, such as a total of 305 fine-grained secondary categories.
[0137] 150. Determine the target primary category to which the target text information belongs according to the first distribution vector, and determine the target secondary category to which the target text information belongs according to the second distribution vector.
[0138] In one or more embodiments, the multi-level category determination device determines the primary category corresponding to the maximum first element score according to the first distribution vector, and uses this primary category as the target primary category. Similarly, determine the secondary category corresponding to the maximum second element score according to the second distribution vector, and use this secondary category as the target secondary category.
[0139] It should be noted that this application takes the output of the target primary category and the target secondary category as an example for introduction. In practical applications, more levels of categories can also be output. Correspondingly, the hierarchical classification model also needs to include classifiers for applying and outputting different categories, which will not be elaborated here.
[0140] In the embodiments of the present application, a method for determining multi-level categories is provided. Through the above method, the prediction result corresponding to the primary category is used as prior knowledge, and after fusing the text encoding vector, the text fusion vector is obtained. The text fusion vector is used as the basis for predicting the secondary category, that is, predicting the output of the next level based on the output result of the previous level, which can fully and effectively utilize the constraint relationship between the upper and lower levels in the category system. Thus, when the second classifier makes a prediction, it can pay more attention to the secondary categories related to the prediction result of the primary category, thereby enhancing the effect of category classification and improving the classification accuracy.
[0141] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding various embodiments, a text fusion vector is generated according to the first distribution vector and the text encoding vector, specifically including:
[0142] Based on the first-level category vector mapping relationship, a prior semantic vector is generated according to the first distribution vector, where the first-level category vector mapping relationship includes a one-to-one mapping relationship between M index values and M semantic vectors, and each index value corresponds to a first-level category;
[0143] According to the prior semantic vector and the text encoding vector, a text fusion vector is generated.
[0144] In one or more embodiments, a method for generating a text fusion vector based on a prior semantic vector and a text encoding vector is introduced. As can be seen from the foregoing embodiments, first, the top K first element scores need to be selected from the first distribution vector, where K can be an integer greater than or equal to 1 and less than or equal to M. Then, a prior semantic vector is generated based on the first-level category vector mapping relationship. Finally, the prior semantic vector and the text encoding vector are fused to obtain a text fusion vector.
[0145] Specifically, the first-level category vector mapping relationship includes a one-to-one mapping relationship between M index values and M semantic vectors, and each index value corresponds to a first-level category. Among them, the semantic vector can be specifically represented as a word vector, and the word vector can be trained on a general corpus or a specific corpus. The training methods generally include word2vec, one-hot, BERT, or matrix factorization. Exemplarily, taking word2vec as an example, please refer to Table 1, which is an illustration of the first-level category vector mapping relationship.
[0146] Table 1
[0147] First-level category Index value Semantic vector Game 1 (0.1,-0.7,0.03,…,-0.9) Movie 2 (0.3,-0.2,0.03,…,-0.8) Sports 3 (-0.1,-0.9,0.23,…,0.1) Live broadcast 4 (0.2,-0.1,0.13,…,-0.7) Variety show 5 (0.3,-0.4,0.22,…,-0.1) Anime 6 (0.3,-0.6,0.04,…,-0.9)
[0148] Among them, the index value has a unique mapping relationship with the first-level category, and the semantic vector also has a unique mapping relationship with the first-level category. Based on this, according to the top K first element scores in the first distribution vector, the corresponding K index values can be found, and then the K semantic vectors corresponding to these K index values can be determined. Thus, the K semantic vectors can be used as the basis for calculating the prior semantic vector.
[0149] Secondly, in the embodiments of the present application, a method for generating a text fusion vector based on a prior semantic vector and a text encoding vector is provided. Through the above method, the first-level category vector mapping relationship is introduced as the basis for generating the prior semantic vector, strengthening the feature expression of the corresponding first-level category in the first distribution vector, thereby facilitating the improvement of the accuracy of category classification.
[0150] Optionally, in the above Figure 3Based on the corresponding various embodiments, in another alternative embodiment provided by the embodiments of the present application, based on the first-level category vector mapping relationship, a prior semantic vector is generated according to the first distribution vector, specifically including:
[0151] Determine the top K first element scores with the largest probability values from the first distribution vector, where each first element score corresponds to an index value of a first-level category, and K is an integer greater than 1 and less than M;
[0152] Obtain the index values corresponding to each of the top K first element scores to obtain K index values;
[0153] Based on the first-level category vector mapping relationship, obtain the corresponding K semantic vectors according to the K index values;
[0154] For each of the K index values, perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value to obtain the updated semantic vector corresponding to the index value;
[0155] Perform a summation calculation on the updated semantic vectors corresponding to the K index values to obtain the prior semantic vector.
[0156] In one or more embodiments, a method for calculating the prior semantic vector is introduced. As can be seen from the foregoing embodiments, assuming that K is an integer greater than 1 and less than M, that is, the top K first element scores with the largest probability values are determined from the first distribution vector, which will be described below with examples.
[0157] Specifically, assume that M is 6 and the first distribution vector is (0.4, 0.2, 0.05, 0.05, 0, 0.3). For the convenience of understanding, please refer to Table 2, which is a schematic diagram of the correspondence between the first element scores and the index values in the first distribution vector.
[0158] Table 2
[0159] First element score Index value 0.4 1 0.2 2 0.05 3 0.05 4 0 5 0.3 6
[0160] Based on this, assume that K is 3. The top K first element scores with the largest probability values determined from the first distribution vector are "0.4", "0.2", and "0.3" respectively, and the corresponding K index values are "1", "2", and "6" respectively. Combining the first-level category vector mapping relationship provided in Table 1 in the foregoing embodiments, three corresponding semantic vectors can be obtained, that is:
[0161] The semantic vector corresponding to the index value "1" is (0.1, -0.7, 0.03, …, -0.9);
[0162] The semantic vector corresponding to the index value "2" is (0.3, -0.2, 0.03, …, -0.8);
[0163] The semantic vector corresponding to the index value "6" is (0.3, -0.6, 0.04, …, -0.9);
[0164] Thus, the semantic vector corresponding to the same index value and the first element score are weighted and calculated to obtain the updated semantic vector corresponding to the index value, that is:
[0165] The updated semantic vector corresponding to the index value "1" is (0.04, -0.28, 0.012, …, -0.36);
[0166] The updated semantic vector corresponding to the index value "2" is (0.06, -0.04, 0.006, …, -0.16);
[0167] The updated semantic vector corresponding to the index value "6" is (0.09, -0.18, 0.012, …, -0.27);
[0168] Finally, the updated semantic vectors corresponding to the above three index values are summed, that is, the element values at the corresponding positions in the updated semantic vectors are added to obtain the prior semantic vector (0.19, -0.5, 0.03, …, -0.79).
[0169] Again, in the embodiments of the present application, a method for calculating the prior semantic vector is provided. Through the above method, the semantic vector corresponding to the index value can be queried using the first-level category vector mapping relationship. Thus, the feature expression of the corresponding first-level category in the first distribution vector is strengthened, which is beneficial to improving the accuracy of category classification.
[0170] Optionally, based on the above Figure 3 On the basis of the corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, a prior semantic vector is generated according to the first distribution vector based on the first-level category vector mapping relationship, specifically including:
[0171] Determine the first element score with the largest probability value from the first distribution vector, where the first element score corresponds to the index value of a first-level category;
[0172] If the first element score is greater than or equal to the element score threshold, obtain the index value corresponding to the first element score;
[0173] Based on the first-level category vector mapping relationship, obtain the corresponding semantic vector according to the index value;
[0174] Weight and calculate the semantic vector corresponding to the index value and the first element score corresponding to the index value to obtain the prior semantic vector.
[0175] In one or more embodiments, a method for calculating a prior semantic vector is introduced. As can be known from the foregoing embodiments, assuming K is 1, that is, the first element score with the largest probability value is determined from the first distribution vector. The following will be described in conjunction with examples.
[0176] Specifically, assume M is 6 and the first distribution vector is (0.4, 0.2, 0.05, 0.05, 0, 0.3). For ease of understanding, please refer to Table 2 again. Based on this, assuming K is 1, the first element score with the largest probability value determined from the first distribution vector is "0.4", and its corresponding K index values are "1". Combining with the first-level category vector mapping relationship provided in Table 1 in the foregoing embodiments, the corresponding semantic vector can be obtained, that is:
[0177] The semantic vector corresponding to the index value "1" is (0.1, -0.7, 0.03, …, -0.9);
[0178] Thus, the semantic vector corresponding to the same index value and the first element score are weighted and calculated to obtain the prior semantic vector corresponding to the index value, that is:
[0179] The prior semantic vector corresponding to the index value "1" is (0.04, -0.28, 0.012, …, -0.36).
[0180] Furthermore, in the embodiments of the present application, another method for calculating the prior semantic vector is provided. Through the above method, the semantic vector corresponding to an index value can be queried by using the first-level category vector mapping relationship. Thus, the feature expression of the most likely first-level category in the first distribution vector is strengthened, which is beneficial to improving the accuracy of category classification.
[0181] Optionally, on the basis of the corresponding embodiments above, Figure 3 In another optional embodiment provided by the embodiments of the present application, based on the first-level category vector mapping relationship, a prior semantic vector is generated according to the first distribution vector, which specifically includes:
[0182] Obtain the index values corresponding to each first element score in the first distribution vector to obtain M index values;
[0183] Based on the first-level category vector mapping relationship, obtain the corresponding M semantic vectors according to the M index values;
[0184] For each index value among the M index values, the semantic vector corresponding to the index value and the first element score corresponding to the index value are weighted and calculated to obtain the updated semantic vector corresponding to the index value;
[0185] Sum the updated semantic vectors corresponding to the K index values to obtain a prior semantic vector.
[0186] In one or more embodiments, a method for calculating a prior semantic vector is introduced. As can be seen from the foregoing embodiments, assuming that K is an integer equal to M, that is, each first element score is determined from the first distribution vector. This will be described below with examples.
[0187] Specifically, assume that M is 6 and the first distribution vector is (0.4, 0.2, 0.05, 0.05, 0, 0.3). For ease of understanding, please refer to Table 2 again. Based on this, assume that K is equal to M. The respective first element scores determined from the first distribution vector are "0.4", "0.2", "0.05", "0.05", "0", and "0.3", and the corresponding K index values are "1", "2", "3", "4", "5", and "6". Combining the first-level category vector mapping relationship provided in Table 1 in the foregoing embodiments, the corresponding six semantic vectors can be obtained, namely:
[0188] The semantic vector corresponding to the index value "1" is (0.1, -0.7, 0.03, …, -0.9);
[0189] The semantic vector corresponding to the index value "2" is (0.3, -0.2, 0.03, …, -0.8);
[0190] The semantic vector corresponding to the index value "3" is (-0.1, -0.9, 0.23, …, 0.1);
[0191] The semantic vector corresponding to the index value "4" is (0.2, -0.1, 0.13, …, -0.7);
[0192] The semantic vector corresponding to the index value "5" is (0.3, -0.4, 0.22, …, -0.1);
[0193] The semantic vector corresponding to the index value "6" is (0.3, -0.6, 0.04, …, -0.9);
[0194] Thus, perform a weighted calculation on the semantic vectors and the first element scores corresponding to the same index value to obtain the updated semantic vector corresponding to the index value, that is:
[0195] The updated semantic vector corresponding to the index value "1" is (0.04, -0.28, 0.012, …, -0.36);
[0196] The updated semantic vector corresponding to the index value "2" is (0.06, -0.04, 0.006, …, -0.16);
[0197] The updated semantic vector corresponding to the index value "3" is (-0.005, -0.045, 0.0115, …, 0.005);
[0198] The updated semantic vector corresponding to the index value "4" is (0.01, -0.005, 0.0065, …, -0.035);
[0199] The updated semantic vector corresponding to the index value "5" is (0, 0, 0, …, 0);
[0200] The updated semantic vector corresponding to the index value "6" is (0.09, -0.18, 0.012, …, -0.27);
[0201] Finally, perform a summation calculation on the updated semantic vectors corresponding to the above six index values, that is, add the element values at the corresponding positions in the updated semantic vectors to obtain the prior semantic vector (0.195, -0.55, 0.048, …, -0.83).
[0202] Again, in the embodiments of the present application, another way to calculate the prior semantic vector is provided. Through the above method, the semantic vectors corresponding to all index values can be queried using the first-level category vector mapping relationship. Thus, the feature expressions of each first-level category in the first distribution vector are strengthened, which is beneficial to improving the accuracy of category classification.
[0203] Optionally, based on the corresponding embodiments above, Figure 3 In another optional embodiment provided by the embodiments of the present application, a text fusion vector is generated according to the prior semantic vector and the text encoding vector, specifically including:
[0204] Perform a splicing process on the prior semantic vector and the text encoding vector to obtain the text fusion vector;
[0205] Or,
[0206] Generate a text fusion vector according to the prior semantic vector and the text encoding vector, specifically including:
[0207] Determine a first-order vector mapping according to the prior semantic vector, the text encoding vector, and the first parameter matrix, where the prior semantic vector is represented as a p-dimensional vector, the text encoding vector is represented as a q-dimensional vector, and the first parameter matrix is represented as a [d*(p + q)]-dimensional matrix, and p, q, and d are all integers greater than 1;
[0208] Determine a second-order vector mapping according to the prior semantic vector, the text encoding vector, and the second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix;
[0209] Generate a text fusion vector based on a first-order vector mapping, a second-order vector mapping, and a bias vector.
[0210] In one or more embodiments, two ways of generating a text fusion vector by using a prior semantic vector and a text encoding vector are introduced. One way is direct concatenation (concat), and the other way is to introduce a first-order vector mapping and a second-order vector mapping to strengthen the interaction of multi-dimensional vector features. The following will introduce them separately.
[0211] Exemplarily, assume that the prior semantic vector (e_emb) is 300-dimensional and the text encoding vector (l1_emb) is 768-dimensional. After concatenation, a 1068-dimensional text fusion vector (L2_logits) is obtained.
[0212] Exemplarily, a bilinear fusion method is adopted, that is, a hybrid first-order vector mapping and a second-order vector mapping are introduced, which can effectively strengthen the interaction of two-dimensional vector features, thereby improving the effect. The interaction process of the two-dimensional features is as follows:
[0213]
[0214] Among them, L2_logits represents the text fusion vector. represents the first-order vector mapping. l1_emb represents the text encoding vector, and l1_emb ∈ R q , that is, the text encoding vector is represented as a q-dimensional vector. e_emb represents the prior semantic vector, and e_emb ∈ R p , that is, the prior semantic vector is represented as a p-dimensional vector. V represents the first parameter matrix, and V ∈ R d*(p+q) , that is, the first parameter matrix is represented as a [d*(p + q)]-dimensional matrix. [] represents the concatenation process. b represents the bias vector. l1_emb * W [1:d] * e_emb represents the second-order vector mapping, that is, bilinear interaction (i.e., second-order interaction). W [1:d] represents the second parameter matrix, and W [1:d] ∈ R p *q*d , that is, the second parameter matrix is represented as a (d * p * q)-dimensional matrix. d represents the output dimension and belongs to a hyperparameter.
[0215] Again, in the embodiments of the present application, two ways of generating a text fusion vector by using a prior semantic vector and a text encoding vector are provided. Through the above methods, in one implementation, the vectors can be directly concatenated. Thus, while achieving feature fusion, the operation difficulty can be reduced. In another implementation, a method of introducing a hybrid first-order vector mapping and a second-order vector mapping is adopted, which can effectively strengthen the interaction between two-dimensional vector features, thereby improving the effect.
[0216] Optionally, based on the above Figure 3 corresponding respective embodiments, in another alternative embodiment provided by the embodiments of the present application, a text fusion vector is generated according to the first distribution vector and the text encoding vector, specifically including:
[0217] Concatenate the first distribution vector and the text encoding vector to obtain a text fusion vector;
[0218] Or,
[0219] Generate a text fusion vector according to the first distribution vector and the text encoding vector, specifically including:
[0220] Determine a first-order vector mapping according to the first distribution vector, the text encoding vector, and the first parameter matrix, where the first distribution vector is represented as a p-dimensional vector, the text encoding vector is represented as a q-dimensional vector, and the first parameter matrix is represented as a [d*(p + q)]-dimensional matrix, and p, q, and d are all integers greater than 1;
[0221] Determine a second-order vector mapping according to the first distribution vector, the text encoding vector, and the second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix;
[0222] Generate a text fusion vector according to the first-order vector mapping, the second-order vector mapping, and the bias vector.
[0223] In one or more embodiments, two ways of generating a text fusion vector using the first distribution vector and the text encoding vector are introduced. One way is direct concatenation (concat), and the other way is to introduce a first-order vector mapping and a second-order vector mapping method to strengthen the interaction of multi-dimensional vector features. The following will be introduced separately.
[0224] Exemplarily, assume that the first distribution vector (logits1) is 44-dimensional and the text encoding vector (l1_emb) is 768-dimensional. After concatenation, an 812-dimensional text fusion vector (L2_logits) is obtained.
[0225] Exemplarily, the bilinear fusion method is adopted, that is, a hybrid first-order vector mapping and a second-order vector mapping are introduced, which can effectively strengthen the interaction of the two-dimensional vector features, thereby improving the effect. The interaction process of the two-dimensional features is as follows:
[0226]
[0227] Among them, L2_logits represents the text fusion vector. represents the first-order vector mapping. l1_emb represents the text encoding vector, and l1_emb ∈ Rq , that is, the text encoding vector is represented as a q-dimensional vector. logits1 represents the first distribution vector, and logits1 ∈ R p , that is, the first distribution vector is represented as a p-dimensional vector. V represents the first parameter matrix, and V ∈ R d*(p+q) , that is, the first parameter matrix is represented as a [d * (p + q)]-dimensional matrix. [] represents the concatenation process. b represents the bias vector. l1_emb * W [1:d] * logits1 represents a second-order vector mapping, that is, a bilinear interaction (i.e., a second-order interaction). W [1:d] represents the second parameter matrix, and W [1:d] ∈ R p*q*d , that is, the second parameter matrix is represented as a (d * p * q)-dimensional matrix. d represents the output dimension and belongs to a hyperparameter.
[0228] Secondly, in the embodiments of the present application, two methods for generating a text fusion vector by using the first distribution vector and the text encoding vector are provided. Through the above methods, in one implementation, vectors can be directly concatenated. Thus, while achieving feature fusion, the operation difficulty can also be reduced. In another implementation, a method that combines first-order vector mapping and second-order vector mapping is introduced, which can effectively strengthen the interaction between the features of the two-dimensional vectors, thereby improving the effect.
[0229] Optionally, on the basis of the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, before obtaining the text encoding vector corresponding to the target text information, it may further include:
[0230] Obtain the target text information corresponding to the target video, where the target text information includes at least one of the title information, abstract information, subtitle information, and comment information of the target video;
[0231] Or,
[0232] Obtain the target text information corresponding to the target picture, where the target text information includes at least one of the title information, author information, optical character recognition (OCR) information, and abstract information of the target picture;
[0233] Or,
[0234] Obtain the target text information corresponding to the target commodity, where the target text information includes at least one of the commodity name information, origin information, comment information, and commodity description information of the target commodity;
[0235] Or,
[0236] Obtain the target text information corresponding to the target text, where the target text information includes at least one of the title information, author information, abstract information, comment information, and body information of the target text.
[0237] In one or more embodiments, various ways of extracting target text information from multiple types of information are introduced. As can be seen from the foregoing embodiments, the target text information can be sourced from video platforms, official accounts, e-commerce platforms, or Weibo, etc. Below, the hierarchical classification tasks will be introduced respectively using target videos, target pictures, target products, and target texts as examples.
[0238] Task 1: Perform hierarchical classification on the target video;
[0239] Exemplarily, the video title, abstract information, subtitle information, and comment information are the main components of the video content. Combining algorithms such as NLP to complete the parsing of the text, thereby strengthening the understanding of the video semantic information, is one of the important tasks of the video search system. Among them, the abstract information can be a brief introduction to the video, the subtitle information is the subtitle extracted from the video, and the comment information refers to the information of users' comments on the video.
[0240] For ease of understanding, please refer to Figure 5 , Figure 5 , which is a schematic diagram of the hierarchical classification task of the video content in the embodiment of the present application. As shown in the figure, assume that the target text information extracted from the target video is "This hero didn't develop well, was economically suppressed, and couldn't continue fighting at all". Input the target text information into the hierarchical classification model. Thus, the target first-level category is "Games", and the target second-level category is "Mobile Games".
[0241] Task 2: Perform hierarchical classification on the target picture;
[0242] Exemplarily, the title information, author information, OCR information, and abstract information are the main components of the picture content. Combining algorithms such as NLP to complete the parsing of the text, thereby strengthening the understanding of the picture semantic information, is one of the important tasks of the picture search system. Among them, the author information represents the name of the photographer or the information of the production party of the picture, the OCR information represents the text information recognized from the picture, and the abstract information can be a brief introduction to the picture.
[0243] For ease of understanding, please refer to Figure 6 , Figure 6 , which is a schematic diagram of the hierarchical classification task of the picture content in the embodiment of the present application. As shown in the figure, assume that the target text information extracted from the target picture is "The figure in the painting is proud and self-respecting. She is wearing luxurious clothes and sitting in a magnificent open carriage". Input the target text information into the hierarchical classification model. Thus, the target first-level category is "Paintings", and the target second-level category is "Figures".
[0244] Task 3: Hierarchically classify the target product;
[0245] Exemplarily, product name information, origin information, review information, and product description information are the main components of product content. Combining algorithms such as NLP to complete the parsing of the text, thereby strengthening the understanding of product semantic information, is one of the important tasks of the product search system. Among them, review information represents the reviews of the buyer on the product, and product description information represents the brief introduction of the product by the merchant.
[0246] For ease of understanding, please refer to Figure 7 , Figure 7 which is a schematic diagram of the hierarchical classification task of product content in the embodiments of the present application. As shown in the figure, assume that the target text information extracted from the target picture is "Product type: Calendar; Origin: Zhejiang; Sales volume: 5000 / month; Price: 18 yuan". Input the target text information into the hierarchical classification model. Thus, the target first-level category is "Electrical appliances", and the target second-level category is "Telephone".
[0247] Task 4: Hierarchically classify the target text;
[0248] Exemplarily, title information, author information, abstract information, review information, and body information are the main components of text content. Combining algorithms such as NLP to complete the parsing of the text, thereby strengthening the understanding of text semantic information, is one of the important tasks of the text search system. Among them, review information represents the reviews of the buyer on the product, and product description information represents the brief introduction of the product by the merchant.
[0249] For ease of understanding, please refer to Figure 8 , Figure 8 which is a schematic diagram of the hierarchical classification task of text content in the embodiments of the present application. As shown in the figure, assume that the target text information extracted from the target picture is "Roses are in the Rosales order,..., Roses have always been highly regarded". Input the target text information into the hierarchical classification model. Thus, the target first-level category is "Popular science", and the target second-level category is "Plants".
[0250] It should be noted that in practical applications, there may also be more types of tasks. The four types of hierarchical classification tasks introduced in the present application are only for illustration and should not be construed as a limitation of the present application.
[0251] Secondly, in the embodiments of the present application, multiple ways of extracting target text information from multiple types of information are provided. Through the above methods, it can be applied to different category classification scenarios. Whether it is video or picture, whether it is product or text, the methods provided in the present application can be used to extract the corresponding target text information for further prediction, thereby improving the flexibility and diversity of the solution.
[0252] Optionally, based on each of the foregoing Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, it may further include:
[0253] Receiving a category query instruction sent by a receiving terminal device for content to be searched;
[0254] In response to the category query instruction, if the content to be searched is video content, sending a video search result to the terminal device;
[0255] In response to the category query instruction, if the content to be searched is picture content, sending a picture search result to the terminal device;
[0256] In response to the category query instruction, if the content to be searched is product content, sending a product search result to the terminal device;
[0257] In response to the category query instruction, if the content to be searched is text content, sending a text search result to the terminal device.
[0258] In one or more embodiments, a variety of ways to push corresponding content in the search scenario are introduced. As can be seen from the foregoing embodiments, the user can also send a category query instruction for the content to be searched to the server through the terminal device. The server responds to the category query instruction and determines the search result that needs to be pushed to the terminal device according to the content to be searched. The following will be described in combination with four types of application scenarios.
[0259] I. Video search scenario;
[0260] The user can search for videos on a video platform, that is, the content to be searched is video content. For ease of understanding, please refer to Figure 9 , Figure 9 which is a schematic diagram of an interface showing video search results in the embodiments of the present application. As shown in FIG. (A) in Figure 9 , multiple first-level categories are displayed on the video platform. Assuming that the user triggers a category query instruction for "variety shows", based on this, second-level categories related to the first-level category "variety shows" can be displayed. For example, "reality shows", "dancing", and "emotions", etc. Assuming that the user triggers a category query instruction for "reality shows", as shown in FIG. (B) in Figure 9 , content related to the third-level category is displayed in the video search results. For example, "Super Script Murder" and "Escape Maze", etc.
[0261] II. Picture search scenario;
[0262] The user can search for wallpapers on a wallpaper platform, that is, the content to be searched is picture content. For ease of understanding, please refer to Figure 10 ,Figure 10 This is a schematic diagram of an interface for displaying image search results in an embodiment of the present application. As shown in Figure 10 (A) of the figure, multiple first-level categories are displayed on the wallpaper platform. Assume that the user triggers a category query instruction for "constellation wallpapers". Based on this, second-level categories related to the first-level category "constellation wallpapers" can be displayed. As shown in Figure 10 (B) of the figure, content related to the second-level category is displayed in the image search results. For example, "Aries" and "Taurus", etc.
[0263] III. Commodity search scenario;
[0264] The user can search for commodities on the e-commerce platform, that is, the content to be searched is commodity content. For ease of understanding, please refer to Figure 11 Figure 11 This is a schematic diagram of an interface for displaying commodity search results in an embodiment of the present application. As shown in Figure 11 (A) of the figure, multiple first-level categories are displayed on the e-commerce platform. Assume that the user triggers a category query instruction for "household appliances". Based on this, second-level categories related to the first-level category "household appliances" can be displayed. As shown in Figure 11 (B) of the figure, content related to the second-level category is displayed in the commodity search results. For example, "TV", "air conditioner", and "refrigerator", etc.
[0265] IV. Text search scenario;
[0266] The user can search for novels on the e-book platform, that is, the content to be searched is text content. For ease of understanding, please refer to Figure 12 Figure 12 This is a schematic diagram of an interface for displaying text search results in an embodiment of the present application. As shown in Figure 12 (A) of the figure, multiple first-level categories are displayed on the e-book platform. Assume that the user triggers a category query instruction for "science fiction". Based on this, second-level categories related to the first-level category "science fiction" can be displayed. As shown in Figure 12 (B) of the figure, content related to the second-level category is displayed in the text search results. For example, "Machine Age" and "Science Fiction World", etc.
[0267] It should be noted that in actual applications, more application scenarios may be involved. The four application scenarios introduced in this application are only for illustration and should not be construed as a limitation of this application.
[0268] Again, in the embodiments of the present application, a variety of ways to push corresponding content of corresponding categories in the search scenario are provided. Through the above methods, the background can determine the search object (for example, video content, picture content, commodity content, text content, etc.) according to the content to be searched input by the user. Based on this, the background can efficiently search for the content of interest to the user in combination with the pre-determined multi-level categories, and push it to the terminal device used by the user, thereby improving the search efficiency.
[0269] Combined with the above introduction, the method for model training in the present application will be introduced below. Please refer to Figure 13 , an embodiment of the model training method in the embodiments of the present application includes:
[0270] 210. Obtain a predicted text encoding vector corresponding to the text information to be trained, where the text information to be trained corresponds to a first-level labeled category and a second-level labeled category;
[0271] In one or more embodiments, the model training device obtains the text information to be trained, where the text information to be trained has been pre-determined with its corresponding first-level labeled category and second-level labeled category in a manual annotation manner. For the convenience of description, the following will take the training of a piece of text information to be trained as an example for introduction. In the actual training process, a larger number of text information to be trained is required, which will not be elaborated here.
[0272] Specifically, the text information to be trained is used as the input of the encoder. Whether word segmentation is required for the input or whether it needs to be directly input at the character granularity mainly depends on the type of the encoder. Taking the encoder using the BERT model as an example for introduction, the text information to be trained is input into the trained BERT model in the form of character granularity, and after encoding, a predicted text encoding vector is generated. Taking the semantic vector of each character as 768 dimensions as an example, usually the output vector (768 dimensions) of the first character "CLS" is taken as the vector representation of the entire text information to be trained. Therefore, the final predicted text encoding vector is also 768 dimensions.
[0273] It should be noted that the model training device can be deployed on the server, or can be deployed on the terminal device, or can be deployed on a system composed of the terminal device and the server, which is not limited here.
[0274] 220. Based on the predicted text encoding vector, obtain a first predicted distribution vector through the first classifier to be trained included in the hierarchical classifier to be trained, where the first predicted distribution vector includes M first element scores, and each first element score represents a probability value of a first-level category, and M is an integer greater than 1;
[0275] In one or more embodiments, the model training device inputs the predicted text encoding vector into the first classifier to be trained included in the hierarchical classifier to be trained, and the first classifier to be trained obtains the first predicted distribution vector. Among them, the first predicted distribution vector is the prediction result of the first-level category, and the first predicted distribution vector includes M first element scores, where M represents the total number of first-level categories, and each first element score represents the probability value of the corresponding first-level category.
[0276] It should be noted that the content described in step 220 is similar to that in step 120 in the embodiment, so it will not be elaborated here.
[0277] 230. Generate a predicted text fusion vector according to the first predicted distribution vector and the predicted text encoding vector;
[0278] In one or more embodiments, the classification result obtained by the first classifier to be trained is generally accurate. Therefore, the first predicted distribution vector can be used as prior knowledge and input into the second classifier to be trained.
[0279] Specifically, the model training device combines the first predicted distribution vector and the predicted text encoding vector to generate a predicted text fusion vector. That is, the predicted text fusion vector has feature vectors from two dimensions. One is the predicted text encoding vector of the text to be trained itself, and the other is the first predicted distribution vector. Therefore, the fusion of multi-dimensional feature vectors is required here.
[0280] It should be noted that step 230 is similar to the content described in step 130 in the embodiment and Figure 3 the corresponding embodiments, so it will not be elaborated here.
[0281] 240. Based on the predicted text fusion vector, obtain a second predicted distribution vector through the second classifier to be trained included in the hierarchical classifier to be trained. Among them, the second predicted distribution vector includes N second element scores, and each second element score represents the probability value of a second-level category. The second-level category belongs to the next level of the first-level category, and N is an integer greater than 1;
[0282] In one or more embodiments, the model training device inputs the predicted text fusion vector into the second classifier to be trained included in the hierarchical classifier to be trained, and the second classifier to be trained obtains the second predicted distribution vector. Among them, the second predicted distribution vector is the prediction result of the second-level category, and the second predicted distribution vector includes N second element scores, where N represents the total number of second-level categories, and each second element score represents the probability value of the corresponding second-level category.
[0283] It should be noted that the content described in step 240 is similar to that in step 140 in the embodiment, so it will not be elaborated here.
[0284] 250. Update the model parameters of the hierarchical classification model to be trained according to the first prediction distribution vector, the second prediction distribution vector, the first-level labeled category, and the second-level labeled category until the model training conditions are met, and output the hierarchical classification model, where the hierarchical classification model includes the first classifier and the second classifier involved in the above embodiments.
[0285] In one or more embodiments, the model training device uses a loss function to calculate a comprehensive loss value according to the first prediction distribution vector, the second prediction distribution vector, the first-level labeled category, and the second-level labeled category. Then, the stochastic gradient descent method can be used to calculate the gradient of the comprehensive loss value, thereby updating the model parameters of the hierarchical classification model to be trained. When the number of iterations reaches the threshold, or the comprehensive loss value has converged, it means that the model training conditions are met. Thus, a hierarchical classification model including the trained first classifier and second classifier is output.
[0286] In the embodiments of the present application, a method for model training is provided. Through the above method, a hierarchical classification model for implementing multi-level category classification can be trained. Based on this, the prediction result corresponding to the first-level category is used as prior knowledge, and after fusing the text encoding vector, a text fusion vector is obtained. The text fusion vector is used as the basis for predicting the second-level category, that is, the output of the next level is predicted based on the output result of the previous level, which can fully and effectively utilize the constraint relationship between the upper and lower levels in the category system. Thus, when the second classifier makes a prediction, it can pay more attention to the second-level categories related to the prediction result of the first-level category, thereby enhancing the effect of category classification and improving the classification accuracy.
[0287] Optionally, based on the corresponding embodiments above Figure 13 In another optional embodiment provided by the embodiments of the present application, updating the model parameters of the hierarchical classification model to be trained according to the first prediction distribution vector, the second prediction distribution vector, the first-level labeled category, and the second-level labeled category specifically includes:
[0288] Calculate the first loss value of the text information to be trained by using the first classification loss function according to the first prediction distribution vector and the first-level labeled category;
[0289] Calculate the second loss value of the text information to be trained by using the second classification loss function according to the second prediction distribution vector and the second-level labeled category;
[0290] Determine the comprehensive loss value of the text information to be trained according to the first loss value and the second loss value;
[0291] Update the model parameters of the hierarchical classification model to be trained according to the comprehensive loss value.
[0292] In one or more embodiments, a method for updating model parameters using a cross - entropy loss function is introduced. As can be seen from the foregoing embodiments, the static weight sum method can be used as the final loss function for the task during training. In addition, the dynamic weight sum method can also be used as the final loss function for the task. For example, introducing Dynamic Task Priority (DTP) or Focal Loss, such losses can adjust the relevant weight values according to the magnitude of the loss values of the actual loss modules, making the final loss value pay more attention to the modules with large loss values.
[0293] Specifically, the loss function used in this application can be composed of two parts. Among them, the first prediction distribution vector and the first - level annotation category use the first classification loss function (i.e., the negative logarithm loss function), and the second prediction distribution vector and the second - level annotation category use the second classification loss function (i.e., the negative logarithm loss function).
[0294] Based on this, the total loss function can be expressed as the following formula:
[0295] Loss=λ1loss cls1 +λ2loss cls2 ; Equation (3)
[0296] Where Loss represents the comprehensive loss value of the text information to be trained. λ1 represents the first weight value (i.e., the hyperparameter used to adjust the first - level category classification task). λ2 represents the second weight value (i.e., the hyperparameter used to adjust the second - level category classification task). loss cls1 represents the first loss value of the text information to be trained. loss cls2 represents the second loss value of the text information to be trained.
[0297] Since this application is introduced by taking one text information to be trained as an example, if there are multiple text information to be trained, the comprehensive loss values of these text information to be trained need to be accumulated.
[0298] Based on this, it is also necessary to calculate the first loss value and the second loss value respectively. The calculation method of the first loss value is as follows:
[0299]
[0300] Where loss cls1 represents the first loss value of the text information to be trained. M represents the total number of first - level categories. i represents the i - th first - level category. y i represents the first - level annotation category for the i - th first - level category (i.e., indicating whether it belongs to the i - th first - level category, y i is 1 when it belongs to the i - th first - level category, y iWhen it is 0, it means it does not belong to the i-th first-level category). a i Represents the i-th first element score in the first prediction distribution vector (i.e., the probability value of predicting the i-th first-level category).
[0301] The calculation method of the second loss value is as follows:
[0302]
[0303] Among them, loss cls2 Represents the second loss value of the text information to be trained. N represents the total number of second-level categories. j represents the j-th second-level category. y j Represents the second-level annotation category for the j-th second-level category (i.e., annotating whether it belongs to the j-th second-level category, y j When it is 1, it means it belongs to the j-th second-level category, y j When it is 0, it means it does not belong to the j-th second-level category). a j Represents the j-th second element score in the second prediction distribution vector (i.e., the probability value of predicting the j-th second-level category).
[0304] Secondly, in the embodiments of the present application, a method for updating model parameters using a cross-entropy loss function is provided. Through the above method, the model is trained using the cross-entropy loss values corresponding to multiple classifiers, which can effectively improve the classification effect of the classifier, thereby improving the accuracy of multi-level category classification.
[0305] Optionally, on the basis of the above Figure 13 In another optional embodiment provided by the embodiments of the present application, based on the corresponding embodiments, the model parameters of the hierarchical classification model to be trained are updated according to the first prediction distribution vector, the second prediction distribution vector, the first-level annotation category, and the second-level annotation category, specifically including:
[0306] According to the first prediction distribution vector and the first-level annotation category, the first loss value of the text information to be trained is calculated using the first classification loss function;
[0307] According to the second prediction distribution vector and the second-level annotation category, the second loss value of the text information to be trained is calculated using the second classification loss function;
[0308] Determine the first element prediction score corresponding to the first-level annotation category from the first prediction distribution vector, and determine the second element prediction score corresponding to the second-level annotation category from the second prediction distribution vector;
[0309] According to the first element prediction score, the second element prediction score, and the target hyperparameter, the third loss value of the text information to be trained is calculated using the hinge loss function;
[0310] Determine the comprehensive loss value of the text information to be trained according to the first loss value, the second loss value, and the third loss value;
[0311] Update the model parameters of the hierarchical classification model to be trained according to the comprehensive loss value.
[0312] In one or more embodiments, a method for updating model parameters using a cross-entropy loss function and a hinge loss function is introduced. As can be seen from the foregoing embodiments, the static weight sum and / or the dynamic weight sum can be used as the final loss function of the task during training.
[0313] Specifically, the loss function used in this application can be composed of three parts. Among them, the first prediction distribution vector and the first-level annotation category use the first classification loss function (i.e., the negative logarithm loss function), and the second prediction distribution vector and the second-level annotation category use the second classification loss function (i.e., the negative logarithm loss function). To ensure the consistency of the two-level classification results, a hinge loss function is added. Assuming that the classification of the upper-level category is always easier than that of the lower-level category, therefore, adding the hinge loss function can make the probability of the first-level category always greater than the corresponding second-level category.
[0314] Based on this, the total loss function can be expressed as the following formula:
[0315] Loss=λ1loss cls1 +λ2loss cls2 +λ3loss h ; Equation (6)
[0316] Where Loss represents the comprehensive loss value of the text information to be trained. λ1 represents the first weight value (i.e., the hyperparameter used to adjust the first-level category classification task). λ2 represents the second weight value (i.e., the hyperparameter used to adjust the second-level category classification task). λ3 represents the third weight value. loss cls1 represents the first loss value of the text information to be trained. loss cls2 represents the second loss value of the text information to be trained. loss h represents the third loss value of the text information to be trained.
[0317] Since this application introduces an example of a text information to be trained, if multiple text information to be trained are involved, the comprehensive loss values of these text information to be trained need to be accumulated.
[0318] Based on this, it is also necessary to calculate the first loss value, the second loss value, and the third loss value respectively. It should be noted that the calculation method of the first loss value is as described in Equation (4) in the foregoing embodiments, and the calculation method of the second loss value is as described in Equation (5) in the foregoing embodiments. The calculation method of the third loss value (i.e., the hinge loss function) is as follows:
[0319] loss h = max(0, λ + l2_score - l1_score); Equation (7)
[0320] where loss h represents the third loss value of the text information to be trained. λ represents the target hyperparameter. l2_score represents the second element prediction score determined in the second prediction distribution vector corresponding to the secondary annotation category. For example, if the secondary annotation category is "mobile game", "mobile game" corresponds to the 5th index value in the second prediction distribution vector. Therefore, the second element score corresponding to the 5th index value in the second prediction distribution vector is used as the second element prediction score. Similarly, l1_score represents the first element prediction score determined in the first prediction distribution vector corresponding to the primary annotation category. For example, if the primary annotation category is "game", "game" corresponds to the 30th index value in the first prediction distribution vector. Therefore, the first element score corresponding to the 30th index value in the first prediction distribution vector is used as the first element prediction score. max(·) represents taking the maximum value.
[0321] Secondly, in the embodiments of the present application, a method for updating model parameters using the cross-entropy loss function and the hinge loss function is provided. By the above method, adding the hinge loss function can ensure the consistency of the two-level categories, that is, assuming that the upper-level category is always easier than the fine-grained lower-level classification result, that is, the fine-grained classification is more difficult. Therefore, adding the hinge loss function can be used to ensure that the probability of the primary category should always be greater than the corresponding secondary category.
[0322] The following will describe in detail the multi-level category determination device in the present application. Please refer to Figure 14 , Figure 14 which is a schematic diagram of an embodiment of the multi-level category determination device in the embodiments of the present application. The multi-level category determination device 30 includes:
[0323] An acquisition module 310, configured to acquire a text encoding vector corresponding to the target text information;
[0324] The acquisition module 310 is further configured to, based on the text encoding vector, obtain a first distribution vector through a first classifier included in the hierarchical classification model, where the first distribution vector includes M first element scores, and each first element score represents a probability value of a primary category, and M is an integer greater than 1;
[0325] A generation module 320, configured to generate a text fusion vector according to the first distribution vector and the text encoding vector;
[0326] The obtaining module 310 is further configured to obtain a second distribution vector through a second classifier included in the hierarchical classification model based on the text fusion vector, where the second distribution vector includes N second element scores, and each second element score represents a probability value of a secondary category, and the secondary category belongs to a next-level category of a primary category, and N is an integer greater than 1;
[0327] The determining module 330 is configured to determine a target primary category to which the target text information belongs according to the first distribution vector, and determine a target secondary category to which the target text information belongs according to the second distribution vector.
[0328] In an embodiment of the present application, a multi-level category determination device is provided. By using the above device, taking the prediction result corresponding to the primary category as prior knowledge, and fusing the text encoding vector to obtain a text fusion vector, and using the text fusion vector as the basis for predicting the secondary category, that is, predicting the output of the next level based on the output result of the upper level, can fully and effectively utilize the constraint relationship between the upper and lower levels in the category system. Therefore, when the second classifier makes a prediction, it can pay more attention to the secondary categories related to the prediction result of the primary category, thereby enhancing the effect of category classification and improving the classification accuracy.
[0329] Optionally, based on the corresponding embodiment above, Figure 14 In another embodiment of the multi-level category determination device 30 provided in the embodiment of the present application,
[0330] The generating module 320 is specifically configured to generate a prior semantic vector according to the first distribution vector based on the primary category vector mapping relationship, where the primary category vector mapping relationship includes a one-to-one mapping relationship between M index values and M semantic vectors, and each index value corresponds to a primary category;
[0331] Generate a text fusion vector according to the prior semantic vector and the text encoding vector.
[0332] In an embodiment of the present application, a multi-level category determination device is provided. By using the above device, introducing the primary category vector mapping relationship as the basis for generating the prior semantic vector, strengthening the feature expression of the corresponding primary category in the first distribution vector, which is beneficial to improving the accuracy of category classification.
[0333] Optionally, based on the corresponding embodiment above, Figure 14 In another embodiment of the multi-level category determination device 30 provided in the embodiment of the present application,
[0334] The generating module 320 is specifically configured to determine the top K first element scores with the largest probability values from the first distribution vector, where each first element score corresponds to an index value of a primary category, and K is an integer greater than 1 and less than M;
[0335] Obtain the index values corresponding to each of the top K first element scores, resulting in K index values;
[0336] Based on the first-level category vector mapping relationship, obtain the corresponding K semantic vectors according to the K index values;
[0337] For each of the K index values, perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value to obtain the updated semantic vector corresponding to the index value;
[0338] Perform a summation calculation on the updated semantic vectors corresponding to the K index values to obtain the prior semantic vector.
[0339] In an embodiment of the present application, a multi-level category determination device is provided. By using the above device and utilizing the first-level category vector mapping relationship, the semantic vector corresponding to the index value can be queried. Thus, the feature expression of the corresponding first-level category in the first distribution vector is strengthened, which is beneficial to improving the accuracy of category classification.
[0340] Optionally, based on the corresponding embodiment above, in another embodiment of the multi-level category determination device 30 provided in the embodiment of the present application, Figure 14 a generation module 320, specifically configured to determine the first element score with the largest probability value from the first distribution vector, where the first element score corresponds to the index value of a first-level category;
[0341] If the first element score is greater than or equal to the element score threshold, obtain the index value corresponding to the first element score;
[0342] Based on the first-level category vector mapping relationship, obtain the corresponding semantic vector according to the index value;
[0343] Perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value to obtain the prior semantic vector.
[0344] In an embodiment of the present application, a multi-level category determination device is provided. By using the above device and utilizing the first-level category vector mapping relationship, the semantic vector corresponding to one index value can be queried. Thus, the feature expression of the most likely first-level category in the first distribution vector is strengthened, which is beneficial to improving the accuracy of category classification.
[0345] Optionally, based on the corresponding embodiment above, in another embodiment of the multi-level category determination device 30 provided in the embodiment of the present application,
[0346] Optionally, based on the corresponding embodiment above, in another embodiment of the multi-level category determination device 30 provided in the embodiment of the present application, Figure 14 On the basis of the corresponding embodiment above, in another embodiment of the multi-level category determination device 30 provided in the embodiment of the present application,
[0347] A generation module 320, specifically configured to obtain the index values corresponding to each first element score in the first distribution vector, obtaining M index values;
[0348] Based on the first-level category vector mapping relationship, obtain the corresponding M semantic vectors according to the M index values;
[0349] For each index value among the M index values, perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value, obtaining the updated semantic vector corresponding to the index value;
[0350] Perform a summation calculation on the updated semantic vectors corresponding to the K index values, obtaining a prior semantic vector.
[0351] In an embodiment of the present application, a multi-level category determination device is provided. By using the above device and utilizing the first-level category vector mapping relationship, the semantic vectors corresponding to all index values can be queried. Thus, the feature expressions of each first-level category in the first distribution vector are strengthened, which is beneficial to improving the accuracy of category classification.
[0352] Optionally, based on the corresponding embodiment above, Figure 14 In another embodiment of the multi-level category determination device 30 provided in the embodiment of the present application,
[0353] The generation module 320 is specifically configured to splice the prior semantic vector and the text encoding vector to obtain a text fusion vector;
[0354] Or,
[0355] The generation module 320 is specifically configured to determine a first-order vector mapping according to the prior semantic vector, the text encoding vector, and the first parameter matrix, where the prior semantic vector is represented as a p-dimensional vector, the text encoding vector is represented as a q-dimensional vector, and the first parameter matrix is represented as a [d*(p + q)]-dimensional matrix, and p, q, and d are all integers greater than 1;
[0356] Determine a second-order vector mapping according to the prior semantic vector, the text encoding vector, and the second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix;
[0357] Generate a text fusion vector according to the first-order vector mapping, the second-order vector mapping, and the bias vector.
[0358] In an embodiment of the present application, a multi-level category determination device is provided. By using the above device, in one implementation, vectors can be directly spliced. Thus, while achieving feature fusion, the operation difficulty can also be reduced. In another implementation, a method of introducing a hybrid first-order vector mapping and second-order vector mapping is adopted, which can effectively strengthen the interaction between the features of two-dimensional vectors, thereby improving the effect.
[0359] Optionally, based on the above Figure 14 corresponding embodiment, in another embodiment of the multi-level category determination device 30 provided by the embodiments of the present application,
[0360] The generation module 320 is specifically configured to splice the first distribution vector and the text encoding vector to obtain a text fusion vector;
[0361] Or,
[0362] The generation module 320 is specifically configured to determine a first-order vector mapping according to the first distribution vector, the text encoding vector, and the first parameter matrix, where the first distribution vector is represented as a p-dimensional vector, the text encoding vector is represented as a q-dimensional vector, and the first parameter matrix is represented as a [d*(p + q)]-dimensional matrix, and p, q, and d are all integers greater than 1;
[0363] Determine a second-order vector mapping according to the first distribution vector, the text encoding vector, and the second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix;
[0364] Generate a text fusion vector according to the first-order vector mapping, the second-order vector mapping, and the bias vector.
[0365] In the embodiments of the present application, a multi-level category determination device is provided. Using the above device, in one implementation, vectors can be directly spliced. Thus, while achieving feature fusion, the operation difficulty can be reduced. In another implementation, a method of introducing a hybrid first-order vector mapping and second-order vector mapping can effectively strengthen the interaction between the two-dimensional vector features, thereby improving the effect.
[0366] Optionally, based on the above Figure 14 corresponding embodiment, in another embodiment of the multi-level category determination device 30 provided by the embodiments of the present application,
[0367] Before the acquisition module 310 further acquires the text encoding vector corresponding to the target text information, it acquires the target text information corresponding to the target video, where the target text information includes at least one of the title information, abstract information, subtitle information, and comment information of the target video;
[0368] Or,
[0369] Before the acquisition module 310 further acquires the text encoding vector corresponding to the target text information, it acquires the target text information corresponding to the target picture, where the target text information includes at least one of the title information, author information, optical character recognition (OCR) information, and abstract information of the target picture;
[0370] Or,
[0371] The obtaining module 310 is further configured to obtain the target text information corresponding to the target product before obtaining the text encoding vector corresponding to the target text information, where the target text information includes at least one of the product name information, origin information, review information, and product description information of the target product;
[0372] Or,
[0373] The obtaining module 310 is further configured to obtain the target text information corresponding to the target text before obtaining the text encoding vector corresponding to the target text information, where the target text information includes at least one of the title information, author information, abstract information, review information, and body information of the target text.
[0374] In the embodiments of the present application, a multi-level category determination device is provided. By using the above device, it can be applied to different category classification scenarios. Whether it is a video or a picture, whether it is a product or a text, the method provided by the present application can be used to extract the corresponding target text information for further prediction, thereby improving the flexibility and diversity of the solution.
[0375] Optionally, based on the corresponding embodiments above, Figure 14 In another embodiment of the multi-level category determination device 30 provided in the embodiments of the present application, the multi-level category determination device 30 further includes a receiving module 340 and a sending module 350;
[0376] The receiving module 340 is configured to receive a category query instruction sent by a terminal device for the content to be searched;
[0377] The sending module 350 is configured to, in response to the category query instruction, if the content to be searched is video content, send a video search result to the terminal device;
[0378] The sending module 350 is further configured to, in response to the category query instruction, if the content to be searched is picture content, send a picture search result to the terminal device;
[0379] The sending module 350 is further configured to, in response to the category query instruction, if the content to be searched is product content, send a product search result to the terminal device;
[0380] The sending module 350 is further configured to, in response to the category query instruction, if the content to be searched is text content, send a text search result to the terminal device.
[0381] In an embodiment of the present application, a multi-level category determination device is provided. By using the above device, the background can determine a search object (such as video content, picture content, commodity content, or text content, etc.) according to the content to be searched input by the user. Based on this, the background can efficiently search for the content of interest to the user in combination with the pre-determined multi-level categories, and push it to the terminal device used by the user, thereby improving the search efficiency.
[0382] The multi-level category determination device in the present application will be described in detail below. Please refer to Figure 15 , Figure 15 which is a schematic diagram of an embodiment of a model training device in an embodiment of the present application. The model training device 40 includes:
[0383] An acquisition module 410, configured to acquire a predicted text encoding vector corresponding to the text information to be trained, where the text information to be trained corresponds to a first-level labeled category and a second-level labeled category;
[0384] The acquisition module 410 is further configured to, based on the predicted text encoding vector, obtain a first predicted distribution vector through a first classifier to be trained included in the hierarchical classification model to be trained, where the first predicted distribution vector includes M first element scores, and each first element score represents a probability value of a first-level category, and M is an integer greater than 1;
[0385] A generation module 420, configured to generate a predicted text fusion vector according to the first predicted distribution vector and the predicted text encoding vector;
[0386] The acquisition module 410 is further configured to, based on the predicted text fusion vector, obtain a second predicted distribution vector through a second classifier to be trained included in the hierarchical classification model to be trained, where the second predicted distribution vector includes N second element scores, and each second element score represents a probability value of a second-level category, and the second-level category belongs to the next level of the first-level category, and N is an integer greater than 1;
[0387] A training module 430, configured to update the model parameters of the hierarchical classification model to be trained according to the first predicted distribution vector, the second predicted distribution vector, the first-level labeled category, and the second-level labeled category until the model training condition is met, and output a hierarchical classification model, where the hierarchical classification model includes the first classifier and the second classifier involved in the above aspects.
[0388] In an embodiment of the present application, a model training device is provided. By using the above device, a hierarchical classification model for implementing multi-level category classification can be trained. Based on this, the prediction result corresponding to the first-level category is used as prior knowledge, and after fusing the text encoding vector, a text fusion vector is obtained. The text fusion vector is used as the basis for predicting the second-level category, that is, the output of the next level is predicted based on the output result of the previous level, and the constraint relationship between the upper and lower levels in the category system can be fully and effectively utilized. Thus, when the second classifier makes a prediction, it can pay more attention to the second-level categories related to the prediction result of the first-level category, thereby enhancing the effect of category classification and improving the classification accuracy.
[0389] Optionally, based on the corresponding embodiment above, in another embodiment of the model training device 40 provided in the embodiment of the present application, Figure 15 the training module 430 is specifically configured to calculate the first loss value of the text information to be trained by using the first classification loss function according to the first prediction distribution vector and the first-level labeled category;
[0390] calculate the second loss value of the text information to be trained by using the second classification loss function according to the second prediction distribution vector and the second-level labeled category;
[0391] determine the comprehensive loss value of the text information to be trained according to the first loss value and the second loss value;
[0392] update the model parameters of the hierarchical classification model to be trained according to the comprehensive loss value.
[0393]
[0394] In an embodiment of the present application, a model training device is provided. By using the above device, the model is trained by using the cross-entropy loss values corresponding to multiple classifiers, which can effectively improve the classification effect of the classifier, thereby improving the accuracy of multi-level category classification.
[0395] Optionally, based on the corresponding embodiment above, in another embodiment of the model training device 40 provided in the embodiment of the present application, Figure 15 the training module is specifically configured to calculate the first loss value of the text information to be trained by using the first classification loss function according to the first prediction distribution vector and the first-level labeled category;
[0396] calculate the second loss value of the text information to be trained by using the second classification loss function according to the second prediction distribution vector and the second-level labeled category;
[0397] calculate the second loss value of the text information to be trained by using the second classification loss function according to the second prediction distribution vector and the second-level labeled category;
[0398] Determine the first element prediction score corresponding to the first-level annotation category from the first prediction distribution vector, and determine the second element prediction score corresponding to the second-level annotation category from the second prediction distribution vector;
[0399] According to the first element prediction score, the second element prediction score, and the target hyperparameter, use the hinge loss function to calculate the third loss value of the text information to be trained;
[0400] Determine the comprehensive loss value of the text information to be trained according to the first loss value, the second loss value, and the third loss value;
[0401] Update the model parameters of the hierarchical classification model to be trained according to the comprehensive loss value.
[0402] In the embodiments of the present application, a model training device is provided. By using the above device, adding the hinge loss function can ensure the consistency of the two-level categories, that is, it is assumed that the upper-level category is always easier than the fine-grained lower-level classification result, that is, the fine-grained classification is more difficult. Therefore, adding the hinge loss function can be used to ensure that the probability of the first-level category should always be greater than the corresponding second-level category.
[0403] The embodiments of the present application also provide a multi-level category determination device and a model training device that can be deployed on a server. Please refer to Figure 16 ., Figure 16 FIG. is a schematic structural diagram of a server provided by the embodiments of the present application. The server 500 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 522 (for example, one or more processors) and a memory 532, and one or more storage media 530 (for example, one or more mass storage devices) for storing application programs 542 or data 544. Among them, the memory 532 and the storage media 530 may be transient storage or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 522 may be configured to communicate with the storage media 530 and execute a series of instruction operations in the storage media 530 on the server 500.
[0404] The server 500 may further include one or more power supplies 526, one or more wired or wireless network interfaces 550, one or more input / output interfaces 558, and / or one or more operating systems 541, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM, FreeBSD TM and so on.
[0405] In the above embodiments, the steps executed by the server may be based on the Figure 16 server structure shown.
[0406] The embodiments of the present application also provide a multi-level category determination device and a model training device that can be deployed on a terminal device. As Figure 17 shown, for the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal device may be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS) device, an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example:
[0407] Figure 17 Shown is a block diagram of a part of the structure of a mobile phone related to the terminal device provided by the embodiments of the present application. Referring to Figure 17 , the mobile phone includes: a radio frequency (RF) circuit 610, a memory 620, an input unit 630, a display unit 640, a sensor 650, an audio circuit 660, a wireless fidelity (WiFi) module 670, a processor 680, and a power supply 690, etc. Those skilled in the art can understand that Figure 17 the mobile phone structure shown in
[0408] does not constitute a limitation on the mobile phone, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Figure 17 The following specifically introduces each component of the mobile phone in conjunction with
[0409] The RF circuit 610 can be used for receiving and sending information or signals during communication. Specifically, it receives the downlink information from the base station and processes it with the processor 680. Additionally, it sends the uplink data designed to the base station. Generally, the RF circuit 610 includes but is not limited to antennas, at least one amplifier, transceivers, couplers, low noise amplifiers (LNAs), duplexers, etc. In addition, the RF circuit 610 can also communicate with networks and other devices via wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0410] The memory 620 can be used to store software programs and modules. The processor 680 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 620. The memory 620 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.), etc.; the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 620 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices.
[0411] The input unit 630 can be used to receive input numeric or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 630 can include a touch panel 631 and other input devices 632. The touch panel 631, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch panel 631), and drive corresponding connection devices according to a pre-set program. Optionally, the touch panel 631 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 680, and can receive and execute commands sent by the processor 680. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 631. In addition to the touch panel 631, the input unit 630 can also include other input devices 632. Specifically, the other input devices 632 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0412] The display unit 640 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 640 can include a display panel 641. Optionally, the display panel 641 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 631 can cover the display panel 641. When the touch panel 631 detects a touch operation on or near it, it is transmitted to the processor 680 to determine the type of touch event. Subsequently, the processor 680 provides corresponding visual output on the display panel 641 according to the type of touch event. Although in Figure 17 the touch panel 631 and the display panel 641 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 631 and the display panel 641 can be integrated to realize the input and output functions of the mobile phone.
[0413] The mobile phone may further include at least one sensor 650, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 641 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 641 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the mobile phone can also be configured with, they will not be elaborated here.
[0414] The audio circuit 660, the speaker 661, and the microphone 662 can provide an audio interface between the user and the mobile phone. The audio circuit 660 can transmit the electrical signal converted from the received audio data to the speaker 661, and the speaker 661 converts it into a sound signal for output; on the other hand, the microphone 662 converts the collected sound signal into an electrical signal, which is received by the audio circuit 660 and then converted into audio data. After the audio data is output to the processor 680 for processing, it is sent through the RF circuit 610 to, for example, another mobile phone, or the audio data is output to the memory 620 for further processing.
[0415] WiFi belongs to short - range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 670, which provides users with wireless broadband Internet access. Although Figure 17 the WiFi module 670 is shown, it can be understood that it does not belong to an essential component of the mobile phone and can be omitted completely within the scope of not changing the essence of the invention according to needs.
[0416] The processor 680 is the control center of the mobile phone. It connects various parts of the entire mobile phone through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 620, and by calling data stored in the memory 620, it executes various functions of the mobile phone and processes data. Optionally, the processor 680 may include one or more processing units; optionally, the processor 680 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 680 either.
[0417] The mobile phone further includes a power source 690 (such as a battery) for supplying power to each component. Optionally, the power source can be logically connected to the processor 680 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system.
[0418] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated herein.
[0419] In the above embodiments, the steps performed by the terminal device may be based on the Figure 17 shown terminal device structure.
[0420] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. When it runs on a computer, it enables the computer to execute the methods described in the foregoing embodiments.
[0421] An embodiment of the present application also provides a computer program product including a program. When it runs on a computer, it enables the computer to execute the methods described in the foregoing embodiments.
[0422] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.
[0423] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.
[0424] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0425] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0426] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0427] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A method for determining multi-level categories, characterized in that, Including: Obtaining a text encoding vector corresponding to target text information; Based on the text encoding vector, obtaining a first distribution vector through a first classifier included in a hierarchical classification model, where the first distribution vector includes M first element scores, and each first element score represents a probability value of a first-level category, and M is an integer greater than 1; Generating a text fusion vector according to the first distribution vector and the text encoding vector; Based on the text fusion vector, obtaining a second distribution vector through a second classifier included in the hierarchical classification model, where the second distribution vector includes N second element scores, and each second element score represents a probability value of a second-level category, and the second-level category belongs to a next-level category of the first-level category, and N is an integer greater than 1; Determining a target first-level category to which the target text information belongs according to the first distribution vector, and determining a target second-level category to which the target text information belongs according to the second distribution vector; Wherein, the generating a text fusion vector according to the first distribution vector and the text encoding vector includes: Based on a first-level category vector mapping relationship, generating a prior semantic vector according to the first distribution vector, where the first-level category vector mapping relationship includes a one-to-one mapping relationship between M index values and M semantic vectors, and each index value corresponds to a first-level category; Determining a first-order vector mapping according to the prior semantic vector, the text encoding vector, and a first parameter matrix, where the prior semantic vector is represented as a p-dimensional vector, the text encoding vector is represented as a q-dimensional vector, the first parameter matrix is represented as a [d*(p+q)]-dimensional matrix, and p, q, and d are all integers greater than 1; Determining a second-order vector mapping according to the prior semantic vector, the text encoding vector, and a second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix; Generating the text fusion vector according to the first-order vector mapping, the second-order vector mapping, and a bias vector; Or, the generating a text fusion vector according to the first distribution vector and the text encoding vector includes: Determining a first-order vector mapping according to the first distribution vector, the text encoding vector, and a first parameter matrix, where the first distribution vector is represented as a p-dimensional vector, the text encoding vector is represented as a q-dimensional vector, the first parameter matrix is represented as a [d*(p+q)]-dimensional matrix, and p, q, and d are all integers greater than 1; Determining a second-order vector mapping according to the first distribution vector, the text encoding vector, and a second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix; Generating the text fusion vector according to the first-order vector mapping, the second-order vector mapping, and a bias vector.
2. The determination method according to claim 1, wherein The generating a prior semantic vector according to the first distribution vector based on the first-level category vector mapping relationship includes: Determine the top K first element scores with the largest probability values from the first distribution vector, where each of the first element scores corresponds to an index value of a first-level category, and K is an integer greater than 1 and less than M; Obtain the index values corresponding to each of the top K first element scores among the top K first element scores, obtaining K index values; Based on the first-level category vector mapping relationship, obtain the corresponding K semantic vectors according to the K index values; For each of the K index values, perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value to obtain the updated semantic vector corresponding to the index value; Perform a summation calculation on the updated semantic vectors corresponding to the K index values to obtain the prior semantic vector.
3. The determination method according to claim 1, wherein The generating the prior semantic vector according to the first distribution vector based on the first-level category vector mapping relationship includes: Determine the first element score with the largest probability value from the first distribution vector, where the first element score corresponds to an index value of a first-level category; If the first element score is greater than or equal to the element score threshold, obtain the index value corresponding to the first element score; Based on the first-level category vector mapping relationship, obtain the corresponding semantic vector according to the index value; Perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value to obtain the prior semantic vector.
4. The determination method according to claim 1, wherein The generating the prior semantic vector according to the first distribution vector based on the first-level category vector mapping relationship includes: Obtain the index values corresponding to each of the first element scores in the first distribution vector, obtaining the M index values; Based on the first-level category vector mapping relationship, obtain the corresponding M semantic vectors according to the M index values; For each of the M index values, perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value to obtain the updated semantic vector corresponding to the index value; Perform a summation calculation on the updated semantic vectors corresponding to the M index values to obtain the prior semantic vector.
5. The determination method according to claim 1, wherein Before obtaining the text encoding vector corresponding to the target text information, the method further includes: Obtain the target text information corresponding to the target video, where the target text information includes at least one of the title information, abstract information, subtitle information, and comment information of the target video; Or, Obtain the target text information corresponding to the target picture, where the target text information includes at least one of the title information, author information, optical character recognition (OCR) information, and abstract information of the target picture; Or, Obtain the target text information corresponding to the target commodity, where the target text information includes at least one of the commodity name information, origin information, comment information, and commodity description information of the target commodity; Or, Obtain the target text information corresponding to the target text, where the target text information includes at least one of the title information, author information, abstract information, comment information, and body information of the target text.
6. The determination method according to any one of claims 1 to 5, characterized in that The method further includes: Receiving a category query instruction sent by a terminal device for content to be searched; In response to the category query instruction, if the content to be searched is video content, sending a video search result to the terminal device; In response to the category query instruction, if the content to be searched is picture content, sending a picture search result to the terminal device; In response to the category query instruction, if the content to be searched is commodity content, sending a commodity search result to the terminal device; In response to the category query instruction, if the content to be searched is text content, sending a text search result to the terminal device.
7. A method for model training, characterized in that, It includes: Obtaining a predicted text encoding vector corresponding to the text information to be trained, where the text information to be trained corresponds to a first-level annotation category and a second-level annotation category; Based on the predicted text encoding vector, obtaining a first predicted distribution vector through a first classifier to be trained included in the hierarchical classification model to be trained, where the first predicted distribution vector includes M first element scores, and each first element score represents a probability value of a first-level category, and M is an integer greater than 1; Generating a predicted text fusion vector according to the first predicted distribution vector and the predicted text encoding vector; Based on the predicted text fusion vector, obtaining a second predicted distribution vector through a second classifier to be trained included in the hierarchical classification model to be trained, where the second predicted distribution vector includes N second element scores, and each second element score represents a probability value of a second-level category, and the second-level category belongs to the next level of the first-level category, and N is an integer greater than 1; Updating the model parameters of the hierarchical classification model to be trained according to the first predicted distribution vector, the second predicted distribution vector, the first-level annotation category, and the second-level annotation category until the model training condition is satisfied, and outputting a hierarchical classification model, where the hierarchical classification model includes the first classifier and the second classifier as described in any one of claims 1 to 6 above; Wherein, the generating a predicted text fusion vector according to the first predicted distribution vector and the predicted text encoding vector includes: Based on a first-level category vector mapping relationship, generating a prior semantic vector according to the first predicted distribution vector, where the first-level category vector mapping relationship includes a one-to-one mapping relationship between M index values and M semantic vectors, and each index value corresponds to a first-level category; Determining a first-order vector mapping according to the prior semantic vector, the predicted text encoding vector, and a first parameter matrix, where the prior semantic vector is represented as a p-dimensional vector, the predicted text encoding vector is represented as a q-dimensional vector, and the first parameter matrix is represented as a [d*(p+q)]-dimensional matrix, and p, q, and d are all integers greater than 1; Determining a second-order vector mapping according to the prior semantic vector, the predicted text encoding vector, and a second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix; Generate the predicted text fusion vector according to the first-order vector mapping, the second-order vector mapping, and the bias vector; Alternatively, generating the predicted text fusion vector according to the first predicted distribution vector and the predicted text encoding vector includes: Determine a first-order vector mapping according to the first predicted distribution vector, the predicted text encoding vector, and a first parameter matrix, where the first predicted distribution vector is represented as a p-dimensional vector, the predicted text encoding vector is represented as a q-dimensional vector, the first parameter matrix is represented as a [d*(p + q)]-dimensional matrix, and p, q, and d are all integers greater than 1; Determine a second-order vector mapping according to the first predicted distribution vector, the predicted text encoding vector, and a second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix; Generate the predicted text fusion vector according to the first-order vector mapping, the second-order vector mapping, and the bias vector.
8. The method according to claim 7, wherein Updating the model parameters of the to-be-trained hierarchical classification model according to the first predicted distribution vector, the second predicted distribution vector, the first-level annotation category, and the second-level annotation category includes: Calculate a first loss value of the to-be-trained text information by using a first classification loss function according to the first predicted distribution vector and the first-level annotation category; Calculate a second loss value of the to-be-trained text information by using a second classification loss function according to the second predicted distribution vector and the second-level annotation category; Determine a comprehensive loss value of the to-be-trained text information according to the first loss value and the second loss value; Update the model parameters of the to-be-trained hierarchical classification model according to the comprehensive loss value.
9. The method according to claim 7, wherein Updating the model parameters of the to-be-trained hierarchical classification model according to the first predicted distribution vector, the second predicted distribution vector, the first-level annotation category, and the second-level annotation category includes: Calculate a first loss value of the to-be-trained text information by using a first classification loss function according to the first predicted distribution vector and the first-level annotation category; Calculate a second loss value of the to-be-trained text information by using a second classification loss function according to the second predicted distribution vector and the second-level annotation category; Determine a first element prediction score corresponding to the first-level annotation category from the first predicted distribution vector, and determine a second element prediction score corresponding to the second-level annotation category from the second predicted distribution vector; Calculate a third loss value of the to-be-trained text information by using a hinge loss function according to the first element prediction score, the second element prediction score, and a target hyperparameter; Determine a comprehensive loss value of the to-be-trained text information according to the first loss value, the second loss value, and the third loss value; Update the model parameters of the to-be-trained hierarchical classification model according to the comprehensive loss value.
10. A multi-level category determination device, characterized in that, including: An acquisition module for acquiring a text encoding vector corresponding to target text information; The obtaining module is further configured to obtain a first distribution vector based on the text encoding vector through a first classifier included in the hierarchical classification model, where the first distribution vector includes M first element scores, and each first element score represents a probability value of a first-level category, and M is an integer greater than 1; The generating module is configured to generate a text fusion vector according to the first distribution vector and the text encoding vector; The obtaining module is further configured to obtain a second distribution vector based on the text fusion vector through a second classifier included in the hierarchical classification model, where the second distribution vector includes N second element scores, and each second element score represents a probability value of a second-level category, and the second-level category belongs to the next level of the first-level category, and N is an integer greater than 1; The determining module is configured to determine a target first-level category to which the target text information belongs according to the first distribution vector, and determine a target second-level category to which the target text information belongs according to the second distribution vector; Wherein, the generating module is specifically configured to: Generate a prior semantic vector according to the first distribution vector based on a first-level category vector mapping relationship, where the first-level category vector mapping relationship includes a one-to-one mapping relationship between M index values and M semantic vectors, and each index value corresponds to a first-level category; Determine a first-order vector mapping according to the prior semantic vector, the text encoding vector, and a first parameter matrix, where the prior semantic vector is represented as a p-dimensional vector, the text encoding vector is represented as a q-dimensional vector, and the first parameter matrix is represented as a [d*(p+q)]-dimensional matrix, and p, q, and d are all integers greater than 1; Determine a second-order vector mapping according to the prior semantic vector, the text encoding vector, and a second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix; Generate the text fusion vector according to the first-order vector mapping, the second-order vector mapping, and a bias vector; Alternatively, the generating module is specifically configured to: Determine a first-order vector mapping according to the first distribution vector, the text encoding vector, and a first parameter matrix, where the first distribution vector is represented as a p-dimensional vector, the text encoding vector is represented as a q-dimensional vector, and the first parameter matrix is represented as a [d*(p+q)]-dimensional matrix, and p, q, and d are all integers greater than 1; Determine a second-order vector mapping according to the first distribution vector, the text encoding vector, and a second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix; Generate the text fusion vector according to the first-order vector mapping, the second-order vector mapping, and a bias vector.
11. The determination device according to claim 10, characterized in that, The generating module is specifically configured to: Determine the top K first element scores with the largest probability values from the first distribution vector, where each first element score corresponds to an index value of a first-level category, and K is an integer greater than 1 and less than M; Obtain the index values corresponding to each of the top K first element scores, obtaining K index values; Based on the first-level category vector mapping relationship, obtain the corresponding K semantic vectors according to the K index values; For each of the K index values, perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value, obtaining the updated semantic vector corresponding to the index value; Perform a summation calculation on the updated semantic vectors corresponding to the K index values, obtaining the prior semantic vector.
12. The determination device according to claim 10, characterized in that, The generating module is specifically configured to: Determine the first element score with the largest probability value from the first distribution vector, where the first element score corresponds to an index value of a first-level category; If the first element score is greater than or equal to the element score threshold, obtain the index value corresponding to the first element score; Based on the first-level category vector mapping relationship, obtain the corresponding semantic vector according to the index value; Perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value, obtaining the prior semantic vector.
13. The determining device according to claim 10, characterized in that, The generating module is specifically configured to: Obtain the index values corresponding to each of the first element scores in the first distribution vector, obtaining the M index values; Based on the first-level category vector mapping relationship, obtain the corresponding M semantic vectors according to the M index values; For each of the M index values, perform a weighted calculation on the semantic vector corresponding to the index value and the first element score corresponding to the index value, obtaining the updated semantic vector corresponding to the index value; Perform a summation calculation on the updated semantic vectors corresponding to the M index values, obtaining the prior semantic vector.
14. The determination device according to claim 10, wherein The obtaining module is further configured to obtain the target text information corresponding to the target video before obtaining the text encoding vector corresponding to the target text information, where the target text information includes at least one of title information, abstract information, subtitle information, and comment information of the target video; Or, The obtaining module is further configured to obtain the target text information corresponding to the target picture before obtaining the text encoding vector corresponding to the target text information, where the target text information includes at least one of title information, author information, optical character recognition (OCR) information, and abstract information of the target picture; Or, The obtaining module is further configured to obtain the target text information corresponding to the target commodity before obtaining the text encoding vector corresponding to the target text information, where the target text information includes at least one of commodity name information, origin information, comment information, and commodity description information of the target commodity; Or, The obtaining module is further configured to obtain the target text information corresponding to the target text before obtaining the text encoding vector corresponding to the target text information, where the target text information includes at least one of title information, author information, abstract information, comment information, and body information of the target text.
15. The determination device according to any one of claims 10 to 14, characterized in that The device further includes: a receiving module and a sending module; The receiving module is configured to receive a category query instruction sent by a terminal device for content to be searched; The sending module is configured to, in response to the category query instruction, if the content to be searched is video content, send a video search result to the terminal device; The sending module is further configured to, in response to the category query instruction, if the content to be searched is picture content, send a picture search result to the terminal device; The sending module is further configured to, in response to the category query instruction, if the content to be searched is commodity content, send a commodity search result to the terminal device; The sending module is further configured to, in response to the category query instruction, if the content to be searched is text content, send a text search result to the terminal device.
16. A model training device, characterized in that, Including: An obtaining module, configured to obtain a predicted text encoding vector corresponding to text information to be trained, where the text information to be trained corresponds to a first-level labeled category and a second-level labeled category; The obtaining module is further configured to, based on the predicted text encoding vector, obtain a first predicted distribution vector through a first classifier to be trained included in the hierarchical classification model to be trained, where the first predicted distribution vector includes M first element scores, and each first element score represents a probability value of a first-level category, and M is an integer greater than 1; A generating module, configured to generate a predicted text fusion vector according to the first predicted distribution vector and the predicted text encoding vector; The obtaining module is further configured to, based on the predicted text fusion vector, obtain a second predicted distribution vector through a second classifier to be trained included in the hierarchical classification model to be trained, where the second predicted distribution vector includes N second element scores, and each second element score represents a probability value of a second-level category, and the second-level category belongs to a next-level category of the first-level category, and N is an integer greater than 1; A training module, configured to update model parameters of the hierarchical classification model to be trained according to the first predicted distribution vector, the second predicted distribution vector, the first-level labeled category, and the second-level labeled category until a model training condition is met, and output a hierarchical classification model, where the hierarchical classification model includes the first classifier and the second classifier as described in any one of claims 1 to 6 above; Wherein, the generating module is specifically configured to: Generate a prior semantic vector according to the first predicted distribution vector based on a first-level category vector mapping relationship, where the first-level category vector mapping relationship includes a one-to-one mapping relationship between M index values and M semantic vectors, and each index value corresponds to a first-level category; Determine a first-order vector mapping according to the prior semantic vector, the predicted text encoding vector, and the first parameter matrix, where the prior semantic vector is represented as a p-dimensional vector, the predicted text encoding vector is represented as a q-dimensional vector, the first parameter matrix is represented as a [d*(p+q)]-dimensional matrix, and p, q, and d are all integers greater than 1; Determine a second-order vector mapping according to the prior semantic vector, the predicted text encoding vector, and the second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix; Generate the predicted text fusion vector according to the first-order vector mapping, the second-order vector mapping, and the bias vector; Alternatively, the generating module is specifically configured to: Determine a first-order vector mapping according to the first predicted distribution vector, the predicted text encoding vector, and the first parameter matrix, where the first predicted distribution vector is represented as a p-dimensional vector, the predicted text encoding vector is represented as a q-dimensional vector, the first parameter matrix is represented as a [d*(p+q)]-dimensional matrix, and p, q, and d are all integers greater than 1; Determine a second-order vector mapping according to the first predicted distribution vector, the predicted text encoding vector, and the second parameter matrix, where the second parameter matrix is represented as a (d*p*q)-dimensional matrix; Generate the predicted text fusion vector according to the first-order vector mapping, the second-order vector mapping, and the bias vector.
17. The device according to claim 16, characterized in that The training module is specifically configured to: Calculate a first loss value of the text information to be trained by using a first classification loss function according to the first predicted distribution vector and the first-level labeled category; Calculate a second loss value of the text information to be trained by using a second classification loss function according to the second predicted distribution vector and the second-level labeled category; Determine a comprehensive loss value of the text information to be trained according to the first loss value and the second loss value; Update the model parameters of the hierarchical classification model to be trained according to the comprehensive loss value.
18. The device according to claim 16, characterized in that, The training module is specifically configured to: Calculate a first loss value of the text information to be trained by using a first classification loss function according to the first predicted distribution vector and the first-level labeled category; Calculate a second loss value of the text information to be trained by using a second classification loss function according to the second predicted distribution vector and the second-level labeled category; Determine a first element prediction score corresponding to the first-level labeled category from the first predicted distribution vector, and determine a second element prediction score corresponding to the second-level labeled category from the second predicted distribution vector; Calculate a third loss value of the text information to be trained by using a hinge loss function according to the first element prediction score, the second element prediction score, and the target hyperparameter; Determine a comprehensive loss value of the text information to be trained according to the first loss value, the second loss value, and the third loss value; Update the model parameters of the hierarchical classification model to be trained according to the comprehensive loss value.
19. A computer device, characterized in that, Comprising: A memory, a processor, and a bus system; Wherein, the memory is used for storing programs; The processor is used for executing the programs in the memory, and the processor is used for executing the multi-level category determination method according to the instructions in the program code as described in any one of claims 1 to 6, or executing the model training method as described in any one of claims 7 to 9; The bus system is used for connecting the memory and the processor so that the memory and the processor can communicate with each other.
20. A computer-readable storage medium, including instructions, when running on a computer, enabling the computer to execute the multi-level category determination method as described in any one of claims 1 to 6, or execute the model training method as described in any one of claims 7 to 9.
21. A computer program product, comprising a computer program and instructions, characterized in that, When the computer program / instructions are executed by the processor, the multi-level category determination method as described in any one of claims 1 to 6 is implemented, or the model training method as described in any one of claims 7 to 9 is executed.
Citation Information
Patent Citations
Content classification method and device, computer equipment and storage medium
CN110737801A
System of text classification model and training method thereof
CN111309919A