Method and apparatus for processing label classification
By constructing a tag semantic relationship diagram and using classifier parameters of related tags, the problem of high data processing cost in multi-label classification is solved, and efficient tag classification training is achieved.
Patent Information
- Application Number
- CN202110461702.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-04-27
AI Technical Summary
In the multi-label classification, the prior art requires continuous collection and labeling of large amounts of data as the tag set is expanded in multi-label classification, resulting in high data processing costs and increasing difficulty in algorithm iteration.
By constructing a tag semantic relationship diagram, determine the initial tag that meets the conditions for the correlation degree of the tag to be added as the relevant tag, and use its classifier parameters to obtain the classifier of the tag to be added, and only a small amount of data is required to train it to achieve tag classification.
This greatly saves data collection and data labeling costs, improves the training efficiency of label classification, and reduces the training steps for adding new labels.
Smart Images

Figure CN113761291B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of machine learning, and particularly to a method and device for processing label classification. Background Art
[0002] Label classification refers to tagging content. Herein, the content can be images, videos, news, music, and so on. Label classification can be used in application scenarios such as content understanding and content review.
[0003] In the practical application scenario of multi-label classification, the label set may continuously expand as the task progresses. For example, originally, 5 different categories were defined for annotation in a multi-label task, and then as the business requirements evolved, 5 new categories needed to be added, expanding from 5 categories to 10 categories. In this case, a direct and commonly used strategy in the industry is to collect sufficient data for annotation for the newly added 5 categories. As the label set continues to expand, using this direct method, it is necessary to continuously collect samples belonging to the new categories and perform a large amount of annotation at the same time. Consequently, the costs of data collection and data annotation continue to increase. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, and storage medium for processing label classification that can reduce the data processing cost.
[0005] A method for processing label classification, the method comprising:
[0006] Obtaining a to-be-added label;
[0007] Based on the constructed label semantic relationship graph, determining an initial label whose relevance to the to-be-added label meets the relevance condition as a relevant label; wherein, the label semantic relationship graph is constructed based on a plurality of initial labels and the relevance between the plurality of initial labels;
[0008] Determining the classifier parameters corresponding to the relevant label;
[0009] Based on the classifier parameters corresponding to the relevant label, obtaining a classifier for the to-be-added label, the classifier for the to-be-added label being used to identify whether the input to-be-identified content belongs to the to-be-added label.
[0010] A device for processing label classification, the device comprising:
[0011] A to-be-added label obtaining module, configured to obtain a to-be-added label;
[0012] A related tag acquisition module, configured to determine, based on a constructed tag semantic relationship graph, an initial tag whose relevance to the to-be-added tag meets the relevance condition as a related tag; wherein, the tag semantic relationship graph is constructed based on a plurality of initial tags and the relevance between the plurality of initial tags;
[0013] A parameter determination module, configured to determine classifier parameters corresponding to the related tags;
[0014] A classifier acquisition module, configured to obtain a classifier for the to-be-added tag based on the classifier parameters corresponding to the related tags, and the classifier for the to-be-added tag is used to identify whether the input content to be recognized belongs to the to-be-added tag.
[0015] A computer device, comprising a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0016] Obtain a to-be-added tag;
[0017] Based on a constructed tag semantic relationship graph, determine an initial tag whose relevance to the to-be-added tag meets the relevance condition as a related tag; wherein, the tag semantic relationship graph is constructed based on a plurality of initial tags and the relevance between the plurality of initial tags;
[0018] Determine classifier parameters corresponding to the related tags;
[0019] Based on the classifier parameters corresponding to the related tags, obtain a classifier for the to-be-added tag, and the classifier for the to-be-added tag is used to identify whether the input content to be recognized belongs to the to-be-added tag.
[0020] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0021] Obtain a to-be-added tag;
[0022] Based on a constructed tag semantic relationship graph, determine an initial tag whose relevance to the to-be-added tag meets the relevance condition as a related tag; wherein, the tag semantic relationship graph is constructed based on a plurality of initial tags and the relevance between the plurality of initial tags;
[0023] Determine classifier parameters corresponding to the related tags;
[0024] Based on the classifier parameters corresponding to the related tags, obtain a classifier for the to-be-added tag, and the classifier for the to-be-added tag is used to identify whether the input content to be recognized belongs to the to-be-added tag.
[0025] The above-mentioned method, device, computer equipment and storage medium for processing label classification, when a new label needs to be added, obtain the label to be newly added, and based on the constructed label semantic relationship graph, determine the initial label whose relevance to the label to be newly added meets the relevance condition as the relevant label, and obtain the classifier parameter of the label to be newly added based on the classifier parameter corresponding to the relevant label. By using the label semantic relationship graph constructed in advance based on multiple initial labels and their relevance relationships, when adding a new label, first determine the initial label related to the label to be newly added, and then based on the classifier parameter of this initial label, a classifier that can be used to classify the label to be newly added can be obtained. Of course, in order to obtain a classifier with better performance, only a small amount of data of the label to be newly added needs to be collected for training, so there is no need to use a large amount of data for algorithm iteration, greatly saving the cost of data collection and data annotation, reducing the training steps of the classifier for the label to be newly added, and improving the training efficiency. Brief Description of the Drawings
[0026] Figure 1 It is an application environment diagram of the method for processing label classification in an embodiment;
[0027] Figure 2 It is a schematic flowchart of the method for processing label classification in an embodiment;
[0028] Figure 3 It is a label semantic relationship graph in an embodiment;
[0029] Figure 4 It is a schematic flowchart of the steps for constructing a classifier for an initial label in an embodiment;
[0030] Figure 5 It is a schematic diagram of the process of processing label classification in another embodiment;
[0031] Figure 6 It is a structural block diagram of the device for processing label classification in an embodiment;
[0032] Figure 7 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments
[0033] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0034] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including the theory, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0035] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0036] The solution provided in the embodiments of this application relates to technologies such as machine learning in artificial intelligence, and is specifically described through the following embodiments:
[0037] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0038] The processing method for label classification provided in this application can be applied to the application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through the network. The terminal 102 sends content to the server 104 by interacting with the server 104. Among them, based on different application scenarios, the interaction methods between the terminal 102 and the server 104 are also different. The server has an identification task for the content. For example, when the terminal 102 uploads a video and the server has an audit task, the server 104 can identify the content through the processing method for label classification of this application.
[0039] Specifically, the server obtains the label to be newly added; based on the constructed label semantic relationship graph, determines the initial labels whose relevance to the label to be newly added meets the relevance condition as relevant labels; wherein, the label semantic relationship graph is constructed based on multiple initial labels and the relevance between multiple initial labels; determines the classifier parameters corresponding to the relevant labels; based on the classifier parameters corresponding to the relevant labels, obtains the classifier for the label to be newly added, and the classifier for the label to be newly added is used to identify whether the input content to be recognized belongs to the label to be newly added. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers. Among them, multiple servers can form a blockchain, and the server is a node on the blockchain.
[0040] In one embodiment, as Figure 2 shown, a processing method for label classification is provided. Taking the server in Figure 1 as an example, the method includes the following steps:
[0041] Step 202, obtain the label to be newly added.
[0042] The label to be newly added refers to a label that is not pre-included in the initial label set in the target business scenario and is the label to be added. The initial label set in the target business scenario is a set of labels determined according to the target application scenario for implementing label classification, which has initial labels and can add new labels on the basis of the initial labels.
[0043] In the practical application scenario of multi-label classification, the initial label set of a target business scenario may continuously expand as the task progresses. For example, originally, 5 different categories were defined for annotation in a multi-label task, and five initial labels were trained. Later, as the business requirements evolve, 5 new categories need to be added, expanding from 5 categories to 10 categories. The five newly added labels are the labels to be newly added. Among them, the initial label set is determined according to the application scenario of label classification. The initial label sets for different application scenarios and application objects are different. For example, in the application scenario of classifying music, the labels in its initial label set are usually the names of singers and the types of songs.
[0044] Step 204, based on the constructed label semantic relationship graph, determine the initial labels whose relevance to the label to be newly added meets the relevance condition as relevant labels; wherein, the label semantic relationship graph is constructed based on multiple initial labels and the relevance between multiple initial labels.
[0045] In some embodiments, each initial label in the label semantic relationship graph may already have classifier parameters obtained through pre-training. Among them, the classifier of an initial label can identify whether the input content to be recognized belongs to the initial label. The label semantic relationship graph is constructed based on multiple initial labels and the relevance between multiple initial labels.
[0046] In one embodiment, the label semantic relationship graph is constructed based on multiple initial labels. Among them, the label semantic relationship graph can be entirely constructed based on the initial labels, so that each initial label in the label semantic relationship graph already has pre-trained classifier parameters. In this embodiment, the initial labels can be common categories in daily life and belong to the common label set. That is to say, the common labels can be used as the initial labels, the classifier parameters of all the initial labels are pre-trained, and the label semantic relationship graph is constructed based on multiple initial labels and the relevance between the initial labels. In this embodiment, the more initial labels there are, the larger the constructed label semantic relationship graph is, the more it can reflect the relevant relationships between the common labels, and the more types of new labels to be added can be added.
[0047] In one embodiment, the label semantic relationship graph can be constructed based on the initial labels and the common labels. Among them, the initial labels are the initial labels of the initial label set of the target business scenario, that is, the existing labels. That is to say, the label semantic relationship graph is constructed based on the existing labels and other common labels. The classifier parameters of the existing labels have been trained. The common labels are common categories in daily life. It can be understood that the initial labels can be common labels. In this embodiment, the classifier parameters of the initial labels of the initial label set in the target business scenario are pre-trained, and the label semantic relationship graph is constructed based on the initial labels, the common labels, and the relevance between each label.
[0048] Among them, the relevance condition is predefined and used to screen out the relevant labels of the new label to be added from the label semantic relationship graph. That is to say, the relevant labels of the new label to be added are the initial labels whose relevance with the new label to be added in the label semantic relationship graph meets the relevance condition.
[0049] In one implementation, whether the relevance meets the relevance condition can be determined by the result obtained by comparing the relevance with a relevance threshold. For example, the initial labels with a relevance greater than the relevance threshold to the label to be newly added are used as the relevant labels of the label to be newly added. In another implementation, whether the relevance meets the relevance condition can also be determined by sorting the relevance to the label to be newly added and judging whether the relevance is among the top N highest relevances. For example, if N is 1, according to the sorting result, the initial label corresponding to the highest relevance is used as the relevant label of the label to be newly added, that is, the initial label most relevant to the label to be newly added is used as the relevant label of the label to be newly added. Another example is that if N is 3, according to the sorting result, the initial labels corresponding to the top 3 highest relevances are used as the relevant labels of the label to be newly added.
[0050] In one implementation, if the label semantic relationship graph is constructed based on multiple initial labels, then based on the semantic information of the label to be newly added, the initial labels whose relevance to the label to be newly added meets the relevance condition are determined in the label semantic relationship graph, that is, according to the semantic information of the label to be newly added, the initial labels with the semantic information relevance meeting the relevance condition are searched in the label semantic relationship graph and used as the relevant labels of the label to be newly added. In this implementation, there are two cases: One case is that there is an initial label identical to the label to be newly added among the multiple initial labels used to construct the label semantic relationship graph (it can be understood that at this time the label to be newly added is in the label semantic relationship graph), then the relevant labels found at this time include the label to be newly added itself. Another case is that there is no initial label identical to the label to be newly added among the multiple initial labels used to construct the label semantic relationship graph (it can be understood that at this time the label to be newly added is not in the label semantic relationship graph), then according to the semantic information of the label to be newly added, the initial labels with the semantic information relevance meeting the relevance condition are searched in the label semantic relationship graph and used as the relevant labels.
[0051] In one implementation, if the label semantic relationship graph is constructed based on the initial labels and common labels of the initial set of the target business scenario, there are two cases: One case is that the label to be newly added is not in the label semantic relationship graph, then similar labels are first found according to the semantic information of the label to be newly added, and the similar labels are the labels with the most matching semantic information to the label to be newly added, and then the initial labels whose relevance to the similar labels meets the relevance condition are found as the relevant labels. Another case is that the label to be newly added is in the label semantic relationship graph, then the initial labels whose relevance to it meets the relevance condition are directly searched in the label semantic relationship graph and used as the relevant labels.
[0052] In one embodiment, the method for processing label classification further includes: obtaining a plurality of initial labels, using the semantic information of each initial label as vertices, and using the relevance between each initial label to represent the connection relationship between each initial label, to construct a label semantic relationship graph.
[0053] Among them, the initial tags can be common tags. The common tags can be determined according to the usage frequency of the tags, and usually include common categories in daily life.
[0054] For the tag classification task, in different application scenarios, there is an initial tag set corresponding to the target application scenario, and the content to be recognized can be attributed to the initial tags in the initial tag set corresponding to the target application scenario. For example, in a certain agricultural platform, the initial tags in the initial tag set are mainly plant names, and in a certain music platform, the initial tags in the initial tag set are mainly music types and singer names. Therefore, for the initial tag set of a certain application scenario, usually only a small part of the common tags are used. Assume that the common tag set is C all , and the initial tag set C current corresponding to a certain target application scenario only uses a small part of the common tag set. Generally, the number of tags in the initial tag set C current corresponding to a certain target application scenario is much smaller than the number of tags in the common tag set C all .
[0055] The semantic information is the semantic information of the name of the tag, which can be obtained by performing word vector conversion on the name of the tag. Each tag in the common tag set is a category, and the semantic information of the tag is obtained according to the word vector of the tag name. The semantic information is the category expression of this category. For example, the category expression refers to the word2vec vector pre-trained on a large-scale corpus. The name of each category corresponds to a unique word2vec vector, and the vector dimension is fixed.
[0056] The word vector can reflect the characteristics of the tag in general semantics. For example, for the word vectors corresponding to the three tags of cat, dog, and bus, the cosine similarity between the word vectors of the two tags of cat and dog is greater than the cosine similarity between the word vectors of cat and bus. This reflects that in general semantic concepts, cat and dog are more similar categories and both belong to pets; while bus belongs to transportation tools and is farther away from these two categories.
[0057] In this embodiment, a tag semantic relationship graph is constructed. The tag semantic relationship graph is a graph structure (Graph), and the expression of the semantic relationship graph is G = {V, A}. Among them, V = {v0, v1,..., v C-1} represents C vertices, and each vertex corresponds to the category expression of a tag.
[0058] A = {a 00 , a 01 , …, a (c-1)(c-1)}\ is the connection matrix of the label semantic relationship graph, representing the connection relationships between C vertices. Among them, a ij characterizes the correlation between two category expressions v i and v j . Among them, the connection relationships between labels can be characterized by the correlation between labels. Specifically, the connection relationship is positively correlated with the correlation between labels. The higher the correlation between labels, the closer the connection relationship between labels.
[0059] Among them, when the correlation is the semantic similarity, the connection relationship can be reflected as the distance between labels. The higher the correlation, the closer the connection relationship, and the shorter the distance between labels in the label semantic relationship graph. In other embodiments, when the correlation is the co-occurrence relationship degree, or the weighted number of the co-occurrence relationship degree and the semantic similarity, the connection relationship can be reflected as the connection weight between labels. The higher the correlation, the greater the connection weight between labels. As Figure 3 shown, the initial labels include: pig, cat, dog, bus, tour bus, dance,..., and the constructed label semantic relationship graph takes the semantic information V of each label as vertices, and the connections between vertices represent the correlation between initial labels.
[0060] Step 206, determine the classifier parameters corresponding to the relevant labels.
[0061] As mentioned before, the classifier parameters of the initial labels can be pre-trained. The relevant labels are the initial labels, and the pre-trained classifier parameters of the initial labels are stored in the memory, and the corresponding classifier parameters can be found by according to the name of the relevant labels.
[0062] Step 208, based on the classifier parameters corresponding to the relevant labels, obtain the classifier of the to-be-added label, and the classifier of the to-be-added label is used to identify whether the input content to be recognized belongs to the to-be-added label.
[0063] In one embodiment, if the found relevant label is itself, obtain the trained classifier parameters of this label to get the classifier parameters of the to-be-added label. In this embodiment, the classifier parameters of common labels can be pre-trained. When it is necessary to add a to-be-added label to the initial label set of the target business scenario, the classifier parameters of the to-be-added label can be directly obtained, improving the addition efficiency.
[0064] In one embodiment, if the retrieved relevant label is an initial label whose relevance meets the relevance condition, obtain the classifier parameters of the relevant label, use them as the initial parameters of the label to be newly added, and combine the existing label samples of the label to be newly added to train the classifier of the label to be newly added, so as to obtain the classifier of the label to be newly added. In this embodiment, by using the classifier parameters of the relevant label of the label to be newly added and migrating them to the label to be newly added, since the relevance between the relevant label and the label to be newly added meets the relevance condition, when migrating to the label to be newly added, it is also relatively reliable, and a relatively reliable mapping relationship of the classifier parameters of the label to be newly added can be obtained, and then the classifier parameters of the label to be newly added can be trained.
[0065] The classifier of the label to be newly added is used to identify whether the input content to be recognized belongs to the label to be newly added, that is, for label classification.
[0066] For the above-mentioned label classification processing method, when a new label needs to be added, obtain the label to be newly added, and based on the pre-constructed label semantic relationship graph, determine the initial label whose relevance to the label to be newly added meets the relevance condition as the relevant label, and obtain the classifier parameters of the label to be newly added based on the classifier parameters corresponding to the relevant label. By using the label semantic relationship graph pre-constructed based on multiple initial labels and their relevance relationships, when adding a new label, first determine the initial label related to the label to be newly added, and then based on the classifier parameters of this initial label, a classifier that can be used to classify the label to be newly added can be obtained. Of course, in order to obtain a classifier with better performance, only a small amount of data of the label to be newly added needs to be collected for training, so there is no need to use a large amount of data for algorithm iteration, greatly saving the costs of data collection and data annotation, reducing the training steps of the classifier of the label to be newly added, and improving the training efficiency.
[0067] In another embodiment, obtaining the classifier of the label to be newly added based on the classifier parameters corresponding to the relevant label includes: when the relevant label includes the label to be newly added, use the classifier parameters corresponding to the relevant label as the classifier parameters of the classifier of the label to be newly added to obtain the classifier of the label to be newly added.
[0068] Specifically, in one implementation manner, if the label semantic relationship graph is constructed based on all initial labels, based on the semantic information of the label to be newly added, determine the initial label whose relevance to the label to be newly added meets the relevance condition in the label semantic relationship graph, that is, according to the semantic information of the label to be newly added, search for the initial label whose semantic information relevance meets the relevance condition in the label semantic relationship graph as the relevant label of the label to be newly added. In this implementation manner, there are two cases. One case is that the relevant label found for the label to be newly added in the label semantic relationship graph includes itself.
[0069] If the retrieved related tags include itself, that is, the related tags include the to-be-added tag, then obtain the classifier parameters of this tag that have been trained to obtain the classifier parameters of the to-be-added tag. In this embodiment, the classifier parameters of common tags can be pre-stored. When it is necessary to add a to-be-added tag to the initial tag set of the target business scenario, the classifier parameters of the to-be-added tag can be directly obtained, improving the addition efficiency.
[0070] In another embodiment, obtaining the classifier of the to-be-added tag based on the classifier parameters corresponding to the related tags includes: when the related tags do not include the to-be-added tag, using the classifier parameters corresponding to the related tags as the classifier parameters of the classifier of the to-be-added tag, and training the classifier of the to-be-added tag based on the existing labeled samples labeled with the to-be-added tag to obtain the classifier of the to-be-added tag.
[0071] Specifically, when the related tags do not include the to-be-added tag, obtain the classifier parameters of the related tags and use them as the initial parameters of the to-be-added tag. Combine the existing labeled samples of the to-be-added tag to train the classifier of the to-be-added tag to obtain the classifier of the to-be-added tag. In this embodiment, the classifier parameters of the related tags of the to-be-added tag are utilized and transferred to the to-be-added tag. Since the relevance between the related tags and the to-be-added tag meets the requirements, when transferred to the to-be-added tag, it is also relatively reliable, and relatively reliable initial parameters of the classifier of the to-be-added tag can be obtained, and then the classifier parameters of the to-be-added tag are trained.
[0072] Taking the most relevant tag as the related tag as an example, the most relevant tag is the initial tag with the highest relevance to the to-be-added tag, and obtain the semantic information of the most relevant initial tag. Taking Figure 3 the shown tag semantic relationship graph as an example, if the to-be-added tag is cat, and the initial tags include pig and dog, and the initial tag with the highest relevance to the tag cat obtained from the tag semantic relationship graph is dog, then the tag dog is the most relevant initial tag of the tag cat, and obtain the semantic information of dog.
[0073] Among them, the classifier parameters of each initial tag have been pre-trained. Specifically, the adjacency relationship graph of the semantic information of the initial tags is processed by a graph convolutional neural network to obtain the mapping relationship between the semantic information of the initial tags and the classifier parameters of the initial tags, and then the parameters of the classifier are learned from the tag features based on the mapping relationship.
[0074] In this embodiment, only a small number of training samples of the to-be-added tag need to be collected. Use the classifier parameters corresponding to the related tags as the classifier parameters of the classifier of the to-be-added tag, and train the classifier of the to-be-added tag based on the existing labeled samples labeled with the to-be-added tag to obtain the classifier of the to-be-added tag.
[0075] The classifier parameters of the initial tags have been trained using a graph convolutional neural network. In the constructed label semantic relationship graph, the correlation degrees between each pair of labels are identified. Based on the correlation degrees, the correlation levels between each pair of labels can be determined. This leads to the fact that when the relatively reliable classifier parameters learned by the GCN from the most relevant labels are transferred to the newly added label, they are also relatively reliable, and relatively reliable classifier parameters for the newly added label can be obtained, that is, the classifier parameters of the relevant labels of the newly added label are used as the classifier parameters of the classifier for the newly added label.
[0076] Furthermore, the transferred classifier parameters can be directly used to train the classifier for the newly added label based on the existing labeled samples marked with the newly added label, so as to obtain the classifier for the newly added label. The classifier for the newly added label is constructed to identify whether the input content to be recognized belongs to the label according to the classifier for the newly added label.
[0077] When the initial label set needs to be extended in the target business scenario of multi-label classification, the traditional strategy can only collect a large amount of data for annotation again for the newly added categories. As the label set continues to expand, the data collection and annotation costs of this technical solution will also continue to rise, increasing the difficulty of algorithm update and iteration. In the label classification processing method of this embodiment, when expanding the label set, since the classifier parameters of the relevant labels can be directly used as the initial parameters of the classifier for the newly added label, that is, the initial parameters are transferred from the classifier parameters of the relevant labels of the newly added label, with the initial parameters, only a small number of samples are needed for training, without using a large amount of data for algorithm iteration, greatly saving the costs of data collection and data annotation, reducing the training steps of the classifier for the newly added label, and improving the training efficiency.
[0078] In one embodiment, the method for determining the correlation degree between each pair of initial labels includes: for any two initial labels among multiple initial labels, obtaining the semantic similarity between the two initial labels; and determining the correlation degree between the two initial labels according to the semantic similarity between the two initial labels.
[0079] Among them, the semantic similarity is specifically the degree of similarity of the semantic information of the initial label. The cosine similarity between the semantic information v i 、v j of two initial labels represents the semantic similarity degree between the two labels. The more similar two labels i,j are, the more similar the semantic information (word vectors) between the two categories are, and the closer the value of the correlation degree a ij is to 1, indicating that the connection relationship between the two categories is stronger.
[0080] In another embodiment, determining the relevance between two initial tags according to the semantic similarity between the two initial tags includes: when there are corresponding existing labeled samples for both of the two initial tags, obtaining the co-occurrence relationship degree between the two initial tags; determining the relevance between the two initial tags based on the larger value between the semantic similarity between the two initial tags and the co-occurrence relationship degree between the two initial tags.
[0081] Specifically, the relevance between tags can be considered from two aspects. One is the semantic similarity between the two tags, and the other is the co-occurrence relationship degree between the two tags. Among them, the co-occurrence relationship degree of the initial tags can refer to the probability that when there are corresponding existing labeled samples for both of the two initial tags, the existing labeled samples are labeled with both of the two initial tags at the same time. Specifically, it is the probability that an existing labeled sample (such as a piece of content) is also labeled with the j-th category when it is labeled as the i-th category. For example, when determining the co-occurrence relationship degree between the two initial tags of the i-th category and the j-th category, all the corresponding existing labeled samples of the two initial tags can be obtained, and the target existing labeled samples that are labeled with both the j-th category and the j-th category are determined from them. The ratio of the target existing labeled samples to all the existing labeled samples corresponding to the two initial tags is used as the co-occurrence relationship degree between the two initial tags of the i-th category and the j-th category.
[0082] In some embodiments, when there are corresponding existing labeled samples for both of the two initial tags, the co-occurrence relationship degree between the two initial tags can be obtained, and the relevance between the two initial tags is determined based on the larger value between the semantic similarity and the co-occurrence relationship degree between the two initial tags.
[0083] Specifically, the two initial tags include a first initial tag and a second initial tag. When there are corresponding existing labeled samples for both of the two initial tags, obtaining the co-occurrence relationship degree between the two initial tags includes: in the set of existing labeled samples corresponding to the two initial tags, determining the target existing labeled samples that are labeled with both of the two initial tags; among them, the set of existing labeled samples corresponding to the two initial tags includes at least the existing labeled samples labeled with the first initial tag and the existing labeled samples labeled with the second initial tag; according to the ratio of the target existing labeled samples to all the existing labeled samples in the set of existing labeled samples, the co-occurrence relationship degree between the two initial tags is obtained.
[0084] Specifically, the set of existing labeled samples corresponding to the two initial tags is the set of existing labeled samples including these two initial tags, and can be the set of existing labeled samples of the initial tags for which the classifier parameters have been trained.
[0085] The target existing labeled samples are the existing labeled samples in the set of labeled samples that are labeled with both the first initial tag and the second initial tag, that is, the target existing labeled samples are labeled as belonging to both the first initial tag and the second initial tag at the same time.
[0086] According to the proportion of the target existing labeled samples in all the existing labeled samples in the existing labeled sample set, the co-occurrence relationship degree between two initial labels is obtained. For example, if the number of samples in the existing labeled sample set is 100 and the number of target existing labeled samples of two initial labels is 17, then the co-occurrence relationship degree between the two is 17 / 100. The co-occurrence relationship degree can measure the co-occurrence degree of two labels in multi-label classification and better reflect the relevance of multi-labels in terms of content. If two initial labels are not very similar but have a strong co-occurrence relationship and always appear in the same content, the relationship between such two labels is also very close.
[0087] In practical applications, since the common label set has a relatively large number of labels, it takes a lot of time to collect and label data for preparing the training samples of each initial label. For the initial labels in the initial label set under the target application scenario, the training samples of the initial labels must be prepared, that is, for the calculation of the co-occurrence relationship degree, there is no need to deliberately prepare. Therefore, to improve the processing efficiency, the training samples of the initial labels can be used to calculate the co-occurrence relationship degree only for the initial labels. That is to say, in this embodiment, only two initial labels that appear simultaneously in the same training sample in the training sample set of the initial labels have a co-occurrence relationship. In this embodiment, by using the training samples of the initial labels to calculate the co-occurrence relationship degree between the initial labels, the relationship of the initial labels can be mined from the existing labeled data of the training samples, saving data processing time and cost.
[0088] The correlation degree is the larger value of the similarity degree and the co-occurrence relationship degree between labels. If there is no co-occurrence relationship degree between labels, the similarity degree is used as the correlation degree. The correlation degree a ij Take the maximum result of the two values of the similarity degree and the co-occurrence relationship degree, and the value range is between 0 and 1.
[0089] In another embodiment, multiple initial labels correspond to at least one business scenario, and the to-be-added label belongs to the target business scenario in at least one business scenario. The label classification processing method further includes the step of constructing a classifier for the initial labels. It can be understood that the step of constructing a classifier for the initial labels can be obtained by pre-processing or determined when determining the classifier parameters corresponding to the relevant labels. As Figure 4 shown, this step includes:
[0090] Step 402, obtaining a target existing labeled sample set corresponding to the initial label set under the target business scenario; wherein, the target existing labeled sample set includes target existing labeled samples, and the target existing labeled samples are labeled with the initial labels in the initial label set under the target business scenario.
[0091] Specifically, the target existing labeled sample is the labeled sample corresponding to the initial label. It can be understood that for each initial label in the initial label set under the target business scenario, different initial labels correspond to different target existing labeled samples. The existing labeled samples of each initial label constitute the target existing labeled sample set.
[0092] Step 404, for the target initial label under the target business scenario, determine the classifier parameters of the target initial label through a graph network based on the label semantic relationship graph; wherein, the target initial label is an initial label in the initial label set under the target business scenario.
[0093] The target initial label is an initial label in the initial label set under the target business scenario. The purpose of this embodiment is to train a classifier for the target initial label. Specifically, a graph network is used to determine the classifier parameters of the target initial label based on the label semantic relationship graph.
[0094] Graph Convolutional Network, abbreviated as GCN. Each node of the label semantic relationship graph is the semantic information of the initial label. GCN directly maps the semantic information to a set of mutually dependent classifiers, and these classifiers can further be directly applied to the classification of the content to be recognized.
[0095] Specifically, by using a graph convolutional network in advance, learn a mapping relationship from the semantic information v i of the label to the classifier parameters:
[0096] w i = GCN(v i )
[0097] Specifically, determining the classifier parameters of the target initial label through a graph network based on the label semantic relationship graph includes: through the graph network, updating the node features of the target initial label based on the relevance between the target initial label and the adjacent initial labels in the label semantic relationship graph, and obtaining the classifier parameters of the target initial label according to the node features of the target initial label; wherein, the label semantic relationship graph is constructed with the semantic information of multiple initial labels as nodes and the relevance between each initial label as the connection relationship; wherein, the adjacent initial label is an initial label having a connection relationship with the target initial label.
[0098] Specifically, the label semantic relationship graph is a graph structure (Graph), and the expression of the semantic relationship graph is G = {V, A}. Wherein V = {v0, v1,..., v C-1} represents C vertices, and each vertex corresponds to a category expression of a label. A = {a 00 , a 01 ,..., a (c-1)(c-1)} is the connection matrix of the label semantic relationship graph, representing the connection relationships between C vertices. Among them, the connection relationship is related to the correlation between labels.
[0099] When training the classifier of the target initial label based on the graph network, the adjacency relationship graph of the target initial label can be obtained based on the label semantic relationship graph, and training can be performed based on the semantic adjacency relationship graph of the target initial label to improve the training efficiency.
[0100] Among them, based on the adjacency relationship graph of an initial label, the semantic information including this initial label, the semantic information of other labels adjacent to the semantic information of this initial label, and the correlation between this initial label and the adjacent other labels can be determined. Based on the label semantic relationship graph, the semantic information of other labels having a connection relationship with a label can be determined. For example, Figure 3 There is a connection relationship between the semantic information of the pig in [] and the semantic information of the dog. According to the connection relationship of the semantic information of the initial label, the semantic information of other labels adjacent to the semantic information of the initial label can be obtained in the label semantic relationship graph.
[0101] In one implementation, the adjacency relationship graph of an initial label is composed of other labels (which can be denoted as the adjacent labels or adjacent nodes of this initial label) in the label semantic relationship graph with the semantic information of the initial label as the center point and having an adjacency relationship with the semantic information of the initial label. Among them, only one adjacency level can be extracted according to needs, or more adjacency levels can be extracted to obtain the adjacency relationship graph. Among them, when extracting the adjacency relationship graph, only the adjacent nodes with a correlation greater than the threshold can be considered. Taking the threshold as 0.6 as an example, taking the initial label as a cat, in the label semantic relationship graph, only the adjacent labels with a correlation greater than 0.6 with the semantic information of the cat are extracted. If the correlation between an adjacent label and the semantic information of the cat is 0.5, it will not appear in the adjacency relationship graph of the cat.
[0102] When training through the graph network based on the label relationship graph, the final output of each GCN node is designed as a classifier related to the label. A classifier is learned from the label features through a mapping function based on GCN, where the input of each GCN layer I takes the node features of the previous layer as the input and then outputs new node features. The input of the first layer is the semantic information of the initial label, and the output of the last layer of the matrix is the classifier.
[0103] The advantage of using GCN is that when GCN performs forward propagation, it will fuse the information of a node and its adjacent nodes in the semantic information adjacency relationship graph of the initial label, and the degree of information fusion depends on the connection relationship a between two nodes in the connection matrix ijIf the connection relationship is large, the degree of considering the information of this adjacent node is greater; conversely, the information of this adjacent node is hardly considered. Such a strategy conforms to human cognitive intuition. For the recognition task, it can consider more features of similar categories, that is, the input of each GCN layer I takes the node features of the previous layer as the input, and then outputs new node features. The goal of GCN is to learn a function of graph G. The input of this function is the feature description and the correlation coefficient matrix, so as to update the node features according to the feature description and the correlation coefficient matrix.
[0104] Step 406: Obtain the predicted labels of the target existing labeled samples based on the classifier parameters of the target initial labels and the target existing labeled samples.
[0105] Among them, the target existing labeled samples can be samples only labeled with the target initial labels, or all target existing labeled samples.
[0106] Specifically, as Figure 5 shown, after obtaining the mapping relationship w i and the features x of the target existing labeled samples, the classification score can be obtained through the inner product of the classifier parameters and the sample features:
[0107]
[0108] Then, using the sigmoid function, the probability that the sample belongs to the c-th label can be obtained:
[0109]
[0110] Take the label with the highest probability as the predicted label of the existing labeled sample.
[0111] Step 408: Train the graph network based on the difference between the predicted labels of the target existing labeled samples and the target initial labels to adjust the classifier parameters of the target initial labels.
[0112] After obtaining the predicted labels of the existing labeled samples, in the training stage, the standard cross-entropy loss function can be used for model training. Specifically, backpropagation is performed according to the difference between the predicted labels and the labeled labels to adjust the classifier parameters of the target initial labels.
[0113] Step 410: Determine the classifier parameters of the target initial labels based on the adjusted classifier parameters of the target initial labels.
[0114] Specifically, after the iterative training is completed, determine the classifier parameters of the target initial labels based on the adjusted classifier parameters of the target initial labels.
[0115] The above-mentioned label classification method learns the mapping relationship between the classifier parameters and the semantic information of the labels. Using this mapping relationship, training is performed based on the sample data of the initial labels, and the resulting classifier takes into account the semantic information of the labels. The mapping relationship is obtained by processing the label semantic relationship graph through a graph convolutional network, which can fuse the semantic information of the labels and adjacent labels, enabling the finally obtained classifier to consider the semantic information of other labels with similar semantics to the label. Multiple label classifiers can be obtained using this method, so that each label classifier takes into account the semantic information of other labels with similar semantics to the label, enabling each label classifier to consider the semantic features of other similar labels. When performing multi-label classification on the content, it can better reflect the relevance of multiple labels in the content, thereby improving the accuracy of multi-label classification.
[0116] In another embodiment, multiple initial labels correspond to at least one business scenario, and the to-be-added label belongs to the target business scenario in at least one business scenario. Based on the constructed semantic relationship graph, determining the initial label whose relevance to the to-be-added label meets the relevance condition as the relevant label includes: based on the constructed semantic relationship graph, determining, from the set of initial labels under the target business scenario, the initial label whose relevance to the to-be-added label meets the relevance condition as the relevant label; wherein, when the relevant label does not include the to-be-added label, after obtaining the classifier of the to-be-added label based on the classifier parameters corresponding to the relevant label, the method further includes: taking the to-be-added label as an initial label under the target business scenario, and updating the set of initial labels under the target business scenario so that the updated set of initial labels under the target business scenario includes the to-be-added label.
[0117] Specifically, after obtaining the classifier of the to-be-added label, the to-be-added label is used as an extended label of the set of initial labels under the target business scenario, and the initial labels in the set of initial labels are updated, thereby updating the set of initial labels under the target business scenario. That is to say, the updated set of initial labels includes the to-be-added label and the initial labels. Through this step, the update of the set of initial labels is realized.
[0118] In another embodiment, the label classification processing method further includes: obtaining the content to be recognized under the target business scenario; based on the classifiers of the initial labels in the set of initial labels under the target business scenario, identifying whether the content to be recognized belongs to the initial labels in the set of initial labels under the target business scenario, and obtaining the multi-label classification result of the content to be recognized.
[0119] Specifically, according to the classifiers of the set of initial labels under the target business scenario, identifying whether the input content to be recognized belongs to the initial labels in the set of initial labels under the target business scenario, and obtaining the multi-label classification result of the content to be recognized.
[0120] Among them, the feature vector of the content to be recognized can be extracted by using a pre-trained feature extraction model. Different types of content to be recognized have different applicable feature extraction models. For example, for text-type content to be recognized, common feature extraction models include LSTM, etc. For image-type content to be recognized, common feature extraction models include CNN, etc. For video-type content to be recognized, common feature extraction networks include TSN, TSM, SlowFast, etc. Taking the content to be recognized as a video as an example, obtaining the feature vector of the content to be recognized includes: obtaining the video, dividing the video into N segments, randomly extracting one frame of picture from each segment, and combining them to obtain a video sequence; using the pre-trained feature extraction network to extract the video features of the video sequence.
[0121] Specifically, what the terminal uploads is a video. The length of the video is not fixed. For the convenience of subsequent model processing, the video sequence is evenly divided into N segments, and then one frame of picture is randomly extracted from each segment, and combining them together obtains a video sequence with a fixed length of N.
[0122] After that, use the feature extraction network to extract the features of the video sequence. There is no restriction on the structure of the feature extraction network itself, as long as it can effectively extract the spatio-temporal information of the video. For the video sequence, assume that the extracted feature is x ∈ R D , where D represents the dimension of the video feature, and the specific value of D varies according to different feature extraction networks.
[0123] Then, according to the feature vector x of the content to be recognized and the classifier parameter w i obtain the score of the content to be recognized under this label. Specifically, the score can be obtained by calculating the inner product of the trained classifier parameter and the features of the content to be recognized:
[0124]
[0125] Then use the sigmoid function to obtain the probability that the content belongs to the c-th label:
[0126]
[0127] According to the magnitude of the probability value, determine the specific label to which the content to be recognized belongs. According to the probability values of the content to be recognized belonging to each label in the application label set, obtain the multi-label classification result of the object, so as to realize the multi-label classification of the content.
[0128] The advantage of the label classification processing method of this application lies in its scalability. When the label set is expanded, only a small amount of new category data needs to be collected and only this part of new category data is used, and a classifier with good performance can be obtained, avoiding the processes of a large amount of data collection, annotation, and re-training.
[0129] It should be understood that although Figure 2 and Figure 4 each step in the flowchart is shown in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 and Figure 4 at least a part of the steps in can include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0130] In one embodiment, as Figure 6 shown, a processing device 600 for label classification is provided. This device can be a software module, a hardware module, or a combination of both to become a part of a computer device. Specifically, this device includes:
[0131] A to-be-added label acquisition module 602, configured to acquire a to-be-added label.
[0132] A related label acquisition module 604, configured to determine, based on a pre-constructed label semantic relationship graph, an initial label whose relevance to the to-be-added label meets a relevance condition as a related label; wherein, the label semantic relationship graph is constructed based on multiple initial labels and the relevance between the multiple initial labels.
[0133] A parameter determination module 606, configured to determine classifier parameters corresponding to the related label.
[0134] A classifier acquisition module 608, configured to obtain a classifier for the to-be-added label based on the classifier parameters corresponding to the related label, where the classifier for the to-be-added label is used to identify whether the input content to be recognized belongs to the to-be-added label.
[0135] The processing device for the above-mentioned tag classification, when a new tag needs to be added, obtains the to-be-added tag, determines an initial tag that meets the relevance condition with the to-be-added tag as the relevant tag based on the constructed tag semantic relationship graph, and obtains the classifier parameter of the to-be-added tag based on the classifier parameter corresponding to the relevant tag. This method uses the tag semantic relationship graph pre-constructed based on multiple initial tags and their relevance relationships. When adding a new tag, it first determines the initial tag related to the to-be-added tag, and then based on the classifier parameter of this initial tag, a classifier that can be used to classify the to-be-added tag can be obtained. Of course, in order to obtain a classifier with better performance, only a small amount of data of the to-be-added tag needs to be collected for training, thus eliminating the need to use a large amount of data for algorithm iteration, greatly saving the costs of data collection and data annotation, reducing the training steps of the classifier for the to-be-added tag, and improving the training efficiency.
[0136] In another embodiment, the classifier acquisition module is configured to, when the relevant tag does not include the to-be-added tag, use the classifier parameter corresponding to the relevant tag as the classifier parameter of the classifier for the to-be-added tag, and train the classifier for the to-be-added tag based on the existing annotated samples labeled with the to-be-added tag to obtain the classifier for the to-be-added tag.
[0137] In another embodiment, the processing device for tag classification further includes:
[0138] The similarity acquisition module is configured to obtain the semantic similarity between any two of the multiple initial tags.
[0139] The relevance acquisition module determines the relevance between the two initial tags according to the semantic similarity between the two initial tags.
[0140] Among them, the relevance acquisition module is configured to, when both of the two initial tags have corresponding existing annotated samples, obtain the co-occurrence relationship degree between the two initial tags; and determine the relevance between the two initial tags based on the larger value of the semantic similarity between the two initial tags and the co-occurrence relationship degree between the two initial tags.
[0141] Among them, the two initial tags include a first initial tag and a second initial tag. In the set of existing annotated samples corresponding to the two initial tags, determine the target existing annotated samples that are simultaneously labeled with the two initial tags; where the set of existing annotated samples corresponding to the two initial tags at least includes the existing annotated samples labeled with the first initial tag and the existing annotated samples labeled with the second initial tag; and obtain the co-occurrence relationship degree between the two initial tags according to the proportion of the target existing annotated samples in all the existing annotated samples in the set of existing annotated samples.
[0142] In another embodiment, the multiple initial tags correspond to at least one business scenario, and the processing device for tag classification further includes:
[0143] A sample acquisition module, configured to acquire a target existing labeled sample set corresponding to the initial tag set in the target business scenario; wherein, the target existing labeled sample set includes target existing labeled samples, and the target existing labeled samples are labeled with the initial tags in the initial tag set in the target business scenario.
[0144] A training module, configured to determine classifier parameters of the target initial tag based on the tag semantic relationship graph through a graph network for the target initial tag in the target business scenario; wherein, the target initial tag is an initial tag in the initial tag set in the target business scenario;
[0145] A prediction module, configured to obtain predicted tags of the target existing labeled samples based on the target existing labeled samples and the classifier parameters of the target initial tag;
[0146] An adjustment module, configured to train the graph network based on the difference between the predicted tags of the target existing labeled samples and the target initial tag, so as to adjust the classifier parameters of the target initial tag;
[0147] A classifier determination module, configured to determine classifier parameters of the target initial tag based on the adjusted classifier parameters of the target initial tag.
[0148] In another embodiment, the training module is configured to update the node features of the target initial tag based on the relevance between the target initial tag and adjacent tags in the tag semantic relationship graph, and obtain classifier parameters of the target initial tag according to the node features of the target initial tag; the tag semantic relationship graph is constructed with the semantic information of multiple initial tags as nodes and the relevance between each initial tag as the connection relationship; the adjacent initial tag is an initial tag having a connection relationship with the target initial tag.
[0149] In another embodiment, the multiple initial tags correspond to at least one business scenario, and a relevant tag acquisition module is configured to determine, based on the constructed semantic relationship graph, an initial tag whose relevance to the to-be-added tag meets the relevance condition as a relevant tag from the initial tag set in the target business scenario; wherein,
[0150] It further includes an update module, which is used to take the to-be-added tag as an initial tag in the target business scenario and update the initial tag set in the target business scenario, so that the updated initial tag set in the target business scenario includes the to-be-added tag.
[0151] In another embodiment, the processing device for tag classification further includes:
[0152] A content acquisition module, which is used to acquire the content to be recognized in the target business scenario;
[0153] A classification module, which is used to identify whether the content to be recognized belongs to each initial tag in the initial tag set in the target business scenario based on the classifiers of each initial tag in the initial tag set in the target business scenario, and obtain a multi-tag classification result of the content to be recognized.
[0154] For the specific limitations of the processing device for tag classification, reference can be made to the limitations of the processing method for tag classification in the above text, which will not be elaborated here. Each module in the above processing device for tag classification can be implemented in whole or in part through software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or be independent of it, or be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0155] In one embodiment, a computer device is provided. This computer device can be a server, and its internal structure diagram can be as Figure 7 shown. The computer device 700 includes a processor 702, a memory, and a network interface 704 connected through a system bus. Among them, the processor 702 of the computer device 700 is used to provide computing and control capabilities. The memory of the computer device 700 includes a non-volatile storage medium 706 and an internal memory 708. The non-volatile storage medium 706 stores an operating system, a computer program, and a database. The internal memory 708 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium 706. The database of the computer device 700 is used to store tag data. The network interface 704 of the computer device 700 is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a processing method for tag classification.
[0156] Those skilled in the art can understand that Figure 7 the structure shown in
[0157] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented. The computer device may be Figure 7 the computer device shown in the figure.
[0158] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0159] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.
[0160] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0161] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0162] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for processing label classification, characterized in that, The method includes: Obtain the label to be newly added; Based on the constructed label semantic relationship graph, determine the initial labels whose relevance to the label to be newly added meets the relevance condition as relevant labels; wherein, the label semantic relationship graph is constructed based on multiple initial labels and the relevance between the multiple initial labels; the relevance between two initial labels is determined based on the larger value of the semantic similarity and the co-occurrence relationship degree between the two initial labels; the co-occurrence relationship degree between the two initial labels is obtained when there are corresponding existing annotation samples for both of the two initial labels; Determine the classifier parameters corresponding to the relevant labels; Based on the classifier parameters corresponding to the relevant labels, obtain the classifier for the label to be newly added, and the classifier for the label to be newly added is used to identify whether the input content to be recognized belongs to the label to be newly added.
2. The method according to claim 1, wherein The obtaining the classifier for the label to be newly added based on the classifier parameters corresponding to the relevant labels includes: When the relevant labels do not include the label to be newly added, use the classifier parameters corresponding to the relevant labels as the classifier parameters of the classifier for the label to be newly added, and train the classifier for the label to be newly added based on the existing annotation samples labeled with the label to be newly added, so as to obtain the classifier for the label to be newly added.
3. The method according to claim 2, wherein The determination method of the relevance between the multiple initial labels includes: For any two initial labels among the multiple initial labels, obtain the semantic similarity between the two initial labels; When there are corresponding existing annotation samples for both of the two initial labels, obtain the co-occurrence relationship degree between the two initial labels; Based on the larger value of the semantic similarity between the two initial labels and the co-occurrence relationship degree between the two initial labels, determine the relevance between the two initial labels.
4. The method according to claim 3, wherein The two initial labels include a first initial label and a second initial label. The obtaining the co-occurrence relationship degree between the two initial labels when there are corresponding existing annotation samples for both of the two initial labels includes: In the set of existing annotation samples corresponding to the two initial labels, determine the target existing annotation samples that are simultaneously labeled with the two initial labels; wherein, the set of existing annotation samples corresponding to the two initial labels at least includes the existing annotation samples labeled with the first initial label and the existing annotation samples labeled with the second initial label; Obtain the co-occurrence relationship degree between the two initial labels according to the proportion of the target existing annotation samples in all the existing annotation samples in the set of existing annotation samples.
5. The method according to claim 1, characterized in that, The multiple initial labels correspond to at least one business scenario, and the label to be newly added belongs to the target business scenario among the at least one business scenario. Before determining the classifier parameters corresponding to the relevant labels, the method further includes: Obtain the set of target existing annotation samples corresponding to the set of initial labels in the target business scenario; wherein, the set of target existing annotation samples includes target existing annotation samples, and the target existing annotation samples are labeled with the initial labels in the set of initial labels in the target business scenario. For the target initial label in the target business scenario, determine the classifier parameters of the target initial label through a graph network based on the label semantic relationship graph; wherein, the target initial label is an initial label in the set of initial labels in the target business scenario. Based on the target existing labeled samples and the classifier parameters of the target initial label, obtain the predicted labels of the target existing labeled samples. Based on the difference between the predicted labels of the target existing labeled samples and the target initial label, train the graph network to adjust the classifier parameters of the target initial label. Based on the adjusted classifier parameters of the target initial label, determine the classifier parameters of the target initial label.
6. The method according to claim 5, characterized in that, Determining the classifier parameters of the target initial label through a graph network based on the label semantic relationship graph includes: Through the graph network, based on the relevance between the target initial label and adjacent initial labels in the label semantic relationship graph, update the node features of the target initial label, and according to the node features of the target initial label, obtain the classifier parameters of the target initial label; the label semantic relationship graph is constructed with the semantic information of multiple initial labels as nodes and the relevance between each initial label as the connection relationship; the adjacent initial label is an initial label that has a connection relationship with the target initial label.
7. The method according to claim 1 or 2, characterized in that, The multiple initial labels correspond to at least one business scenario, and the to-be-added label belongs to the target business scenario in the at least one business scenario. Based on the already constructed semantic relationship graph, determining the initial label whose relevance with the to-be-added label meets the relevance condition as the relevant label includes: Based on the already constructed semantic relationship graph, from the set of initial labels in the target business scenario, determine the initial label whose relevance with the to-be-added label meets the relevance condition as the relevant label; wherein, When the relevant label does not include the to-be-added label, after obtaining the classifier of the to-be-added label based on the classifier parameters corresponding to the relevant label, the method further includes: Taking the to-be-added label as an initial label in the target business scenario, and updating the set of initial labels in the target business scenario so that the updated set of initial labels in the target business scenario includes the to-be-added label.
8. The method according to claim 7, wherein The method further includes: Obtain the content to be recognized in the target business scenario. Based on the classifiers of each initial label in the set of initial labels in the target business scenario, identify whether the content to be recognized belongs to each initial label in the set of initial labels in the target business scenario, and obtain the multi-label classification result of the content to be recognized.
9. A processing device for label classification, characterized in that, Includes: A to-be-added label acquisition module, configured to acquire a to-be-added label. A related tag acquisition module, configured to determine, based on the constructed tag semantic relationship graph, an initial tag whose relevance to the to-be-added tag meets the relevance condition as a related tag; wherein, the tag semantic relationship graph is constructed based on a plurality of initial tags and the relevance between the plurality of initial tags; the relevance between two initial tags is determined based on the larger value of the semantic similarity and the co-occurrence relationship degree between the two initial tags; the co-occurrence relationship degree between the two initial tags is obtained when both of the two initial tags have corresponding existing annotation samples; A parameter determination module, configured to determine the classifier parameters corresponding to the related tags; A classifier acquisition module, configured to obtain a classifier for the to-be-added tag based on the classifier parameters corresponding to the related tags, and the classifier for the to-be-added tag is used to identify whether the input content to be recognized belongs to the to-be-added tag.
10. The device according to claim 9, characterized in that, The classifier acquisition module is specifically configured to: When the related tags do not include the to-be-added tag, use the classifier parameters corresponding to the related tags as the classifier parameters of the classifier for the to-be-added tag, and train the classifier for the to-be-added tag based on the existing annotation samples labeled with the to-be-added tag to obtain the classifier for the to-be-added tag.
11. The device according to claim 10, wherein The apparatus further includes: A similarity acquisition module, configured to obtain the semantic similarity between any two of the plurality of initial tags; A relevance acquisition module, configured to obtain the co-occurrence relationship degree between two initial tags when both of the two initial tags have corresponding existing annotation samples; and determine the relevance between the two initial tags based on the larger value of the semantic similarity between the two initial tags and the co-occurrence relationship degree between the two initial tags.
12. The device according to claim 11, wherein The two initial tags include a first initial tag and a second initial tag, and the relevance acquisition module is further configured to: Determine a target existing annotation sample that is simultaneously labeled with the two initial tags in the set of existing annotation samples corresponding to the two initial tags; wherein, the set of existing annotation samples corresponding to the two initial tags at least includes the existing annotation samples labeled with the first initial tag and the existing annotation samples labeled with the second initial tag; Obtain the co-occurrence relationship degree between the two initial tags according to the proportion of the target existing annotation sample in all the existing annotation samples in the set of existing annotation samples.
13. The device according to claim 9, wherein The plurality of initial tags correspond to at least one business scenario, and the tag classification processing apparatus further includes: A sample acquisition module, configured to obtain a set of target existing annotation samples corresponding to the set of initial tags in the target business scenario; wherein, the set of target existing annotation samples includes target existing annotation samples, and the target existing annotation samples are labeled with the initial tags in the set of initial tags in the target business scenario; A training module, which is used to determine the classifier parameters of the target initial label for the target business scenario through a graph network based on the label semantic relationship graph; wherein, the target initial label is an initial label in the initial label set for the target business scenario. A prediction module, which is used to obtain the predicted label of the target existing labeled sample based on the target existing labeled sample and the classifier parameters of the target initial label. An adjustment module, which is used to train the graph network based on the difference between the predicted label of the target existing labeled sample and the target initial label, so as to adjust the classifier parameters of the target initial label. A classifier determination module, which is used to determine the classifier parameters of the target initial label based on the adjusted classifier parameters of the target initial label.
14. The device according to claim 13, characterized in that Specifically, the training module is used for: Through the graph network, based on the relevance between the target initial label and the adjacent initial labels in the label semantic relationship graph, update the node features of the target initial label, and obtain the classifier parameters of the target initial label according to the node features of the target initial label; the label semantic relationship graph is constructed with the semantic information of multiple initial labels as nodes and the relevance between each initial label as the connection relationship. The adjacent initial label is an initial label that has a connection relationship with the target initial label.
15. The device according to claim 9 or 10, characterized in that, The multiple initial labels correspond to at least one business scenario, and the to-be-added label belongs to the target business scenario among the at least one business scenario. Specifically, the relevant label acquisition module is used for: based on the constructed semantic relationship graph, determine, from the initial label set for the target business scenario, the initial labels whose relevance to the to-be-added label meets the relevance condition as relevant labels. The apparatus further includes: An update module, which is used to use the to-be-added label as an initial label for the target business scenario and update the initial label set for the target business scenario, so that the updated initial label set for the target business scenario includes the to-be-added label.
16. The device according to claim 15, characterized in that, The apparatus further includes: A content acquisition module, which is used to acquire the content to be recognized for the target business scenario. A classification module, which is used to identify whether the content to be recognized belongs to each initial label in the initial label set for the target business scenario based on the classifiers of each initial label in the initial label set for the target business scenario, and obtain the multi-label classification result of the content to be recognized.
17. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
18. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 8.
19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Development method and system of small sample classification model based on graph convolutional neural network
CN112183620A