Label identification method and device, equipment and storage medium
By extracting and fusing the text semantic features and label semantic features of the object to be identified, the sub-labels of the target label are identified, which solves the problem of low label recognition accuracy in the prior art and achieves a higher recognition accuracy.
Patent Information
- Application Number
- CN202311546172.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
In the prior art, multi-classification tasks based on the tag system are difficult to improve the accuracy of the identification of the tags to be identified, resulting in inaccurate identification results.
By obtaining text information of the object to be identified, text semantic features are extracted, and classification processing is performed based on these features, the initial target tag and credibility are obtained. Then, the tag semantic features of the target tag are extracted separately, and weighted splicing is performed to form comprehensive tag features. Finally, based on the comprehensive label features and text semantic features, the sub-labels of the target label are identified to improve the recognition accuracy.
By utilizing the fusion of text semantic features and tag semantic features, the accuracy of the tag identification of objects to be identified can be improved, the label hierarchical relationship can be fully utilized, and the recognition effect can be enhanced.
Smart Images

Figure CN120020771A_ABST
Abstract
Description
Background Art
[0002] In order to more quickly implement business operations such as searching, querying, and recommending the object to be recognized, label recognition is performed on the object to be recognized according to the text content of the object to be recognized. During label recognition, multiple labels can be recognized for the same object to be recognized, and there is a certain correlation between the labels.
[0003] Currently, the correlation between labels is determined based on hierarchical classification; among them, hierarchical classification is an important task in multi-classification in fields such as natural language processing and computer vision. Its main feature is that the labels in hierarchical classification have a superior-subordinate relationship, the superior label is the parent of the subordinate label, and the finer the granularity of the hierarchical classification is as it goes down.
[0004] In the related art, for multi-classification tasks based on a label system, most of them regard the classification task as several basic multi-classification tasks, directly predict the secondary labels of the object to be recognized, and directly backtrack the primary labels from the prediction results. This method results in a low accuracy of label recognition for the object to be recognized.
[0005] Therefore, how to improve the accuracy of label recognition for the object to be recognized is a technical problem that needs to be solved currently. Summary of the Invention
[0006] Embodiments of the present application provide a label recognition method, device, equipment, and storage medium to improve the accuracy of label recognition for the object to be recognized.
[0007] In a first aspect, embodiments of the present application provide a label recognition method, which includes:
[0008] Obtain the text information of the object to be recognized, and extract the text semantic features of the text information;
[0009] Perform classification processing on the object to be recognized based on the text semantic features to obtain at least one first target label of the object to be recognized and the corresponding first credibility;
[0010] Extract the label semantic features of at least one first target label respectively, and perform splicing processing on the corresponding label semantic features based on the obtained at least one first credibility to obtain comprehensive label features;
[0011] Based on the comprehensive label features and the text semantic features, obtain at least one second target label of the object to be recognized; where the second target label is a sub-label of the first target label.
[0012] In a second aspect, embodiments of the present application provide a label recognition device, which includes:
[0013] An acquisition unit, configured to acquire text information of an object to be recognized, and extract text semantic features of the text information;
[0014] A classification unit, configured to classify the object to be recognized based on the text semantic features, and obtain at least one first target label of the object to be recognized and a corresponding first confidence level;
[0015] A splicing unit, configured to respectively extract label semantic features of at least one first target label, and perform splicing processing on the corresponding label semantic features based on the obtained at least one first confidence level, to obtain comprehensive label features;
[0016] An obtaining unit, configured to obtain at least one second target label of the object to be recognized based on the comprehensive label features and the text semantic features; wherein, the second target label is a sub-label of the first target label.
[0017] In a possible implementation manner, the obtaining unit is specifically configured to:
[0018] Perform fusion processing on the comprehensive label features and the text semantic features to obtain fused target fusion semantic features;
[0019] Perform classification processing on the object to be recognized based on the target fusion semantic features, and obtain at least one second candidate label to which the object to be recognized belongs and a corresponding second confidence level;
[0020] Select, from the at least one second candidate label, at least one second target label whose second confidence level meets the screening condition.
[0021] In a possible implementation manner, the obtaining unit is specifically configured to:
[0022] Obtain a label parameter matrix and a text parameter matrix for feature fusion;
[0023] Perform second-order linear fusion processing on the comprehensive label features based on the label parameter matrix to obtain second-order fusion label features;
[0024] Perform second-order linear fusion processing on the text semantic features based on the text parameter matrix to obtain second-order fusion semantic features;
[0025] Perform fusion processing on the second-order fusion label features and the second-order fusion semantic features to obtain fused target fusion semantic features.
[0026] In a possible implementation manner, the acquisition unit is specifically configured to:
[0027] Perform word segmentation processing on the text information to obtain at least one corresponding text word segment;
[0028] Extract features from at least one text tokenization respectively to obtain corresponding token semantic features.
[0029] Perform splicing processing on the obtained at least one token semantic feature to obtain a text semantic feature.
[0030] In a possible implementation manner, the classification unit is specifically used for:
[0031] Classify the object to be recognized based on the text semantic feature, and obtain at least one first candidate label to which the object to be recognized belongs and the corresponding first credibility.
[0032] Select at least one first target label from the at least one first candidate label, where the first credibility meets the credibility condition.
[0033] In a possible implementation manner, the splicing unit is specifically used for:
[0034] Use at least one first credibility to perform weighted processing on the corresponding label semantic features respectively to obtain at least one weighted label feature.
[0035] Perform concatenation splicing processing on the at least one weighted label feature to obtain a comprehensive label feature.
[0036] In a possible implementation manner, the label recognition method involved in the embodiments of the present application is executed by a label recognition model; wherein, the label recognition device further includes a training unit, and the label recognition model is obtained by the training unit through the following manner:
[0037] Perform cyclic iterative training on the label recognition model to be trained according to the sample training set to obtain the label recognition model; perform the following operations in one cyclic iteration:
[0038] Select a training sample from the sample training set; wherein, the training sample is a historical object including a first labeled label and a second labeled label, and the second labeled label is a sub-label of the first labeled label.
[0039] Input the training sample into the label recognition model to be trained, and predict the first predicted label and the second predicted label associated with the historical object.
[0040] Based on the first predicted label and the second predicted label, construct a loss function, and use the loss function to adjust the parameters of the label recognition model to be trained.
[0041] In a possible implementation manner, the training unit is specifically used for:
[0042] Determine a first loss function based on the main label difference between the first predicted label and the first annotated label, determine a second loss function based on the sub-label difference between the second predicted label and the second annotated label, and determine a constraint function based on the relationship difference between the first predicted label and the second predicted label;
[0043] Construct a loss function based on the first loss function, the second loss function, and the constraint function; wherein, the constraint function is used to ensure that the second predicted label is a sub-label of the first predicted label.
[0044] In a possible implementation manner, the label recognition model to be trained includes: a first classification layer, a fusion layer, and a second classification layer; the training unit is specifically configured to:
[0045] Through the first classification layer, classify the historical object based on the historical text semantic features extracted from the historical text information of the historical object, and obtain at least one first predicted label;
[0046] Perform splicing processing on the historical label semantic features of at least one first predicted label to obtain corresponding historical comprehensive label features;
[0047] Through the fusion layer, perform fusion processing on the historical comprehensive label features and the historical text semantic features to obtain the fused historical fusion semantic features;
[0048] Through the second classification layer, classify the historical object based on the historical fusion semantic features to obtain at least one second predicted label.
[0049] In a third aspect, an embodiment of the present application provides a computing device, including: a memory and a processor, wherein the memory is used to store a computer program; the processor is used to execute the computer program to implement the steps of the label recognition method provided by the embodiment of the present application.
[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of the label recognition method provided by the embodiment of the present application are implemented.
[0051] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer-readable storage medium; when the processor of the computing device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the computing device executes the steps of the label recognition method provided by the embodiment of the present application.
[0052] The beneficial effects of the present application are as follows:
[0053] The embodiments of the present application provide a method, device, equipment and storage medium for label recognition, which relate to technical fields such as text processing and artificial intelligence; and can be applied to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, assisted driving, information recommendation, etc.
[0054] In the embodiments of the present application, when performing label recognition on an object to be recognized, first obtain the text information of the object to be recognized, and extract the text semantic features of the text information; then classify the object to be recognized based on the text semantic features to obtain at least one first target label of the object to be recognized and the corresponding first confidence information. For the label system, the higher the level, the coarser the granularity, the shorter the recognition difficulty, and at the same time, the label recognition accuracy is relatively high, that is, the accuracy of the recognized first target label is relatively high.
[0055] Then, extract the label semantic features of at least one first target label respectively, and perform splicing processing on the corresponding label semantic features based on the obtained at least one first confidence to obtain comprehensive label features; finally, based on the comprehensive label features and the text semantic features, obtain at least one second target label of the object to be recognized, and the second target label is a sub-label of the first target label. It can be seen that when determining the second target label, the upper-layer label recognition result with higher accuracy is used as prior knowledge, and together with the text semantic features, it is used as the input of the lower-layer label recognition process. More reference information is fused in the lower-layer label recognition process, and the label hierarchy relationship is fully utilized, thereby enhancing the recognition effect and improving the accuracy of the label recognition of the object to be recognized. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0057] Figure 1 It is a schematic diagram of an application scenario provided by the embodiments of the present application;
[0058] Figure 2 It is a schematic diagram of a specific application scenario provided by the embodiments of the present application;
[0059] Figure 3 It is a schematic diagram of the structure of a label recognition model provided by the embodiments of the present application;
[0060] Figure 4 It is a schematic diagram of extracting text semantic features through a BERT model provided by the embodiments of the present application;
[0061] Figure 5Schematic diagram for identifying tags and credibility through a first classification layer provided by an embodiment of the present application;
[0062] Figure 6 Schematic diagram of the specific structure of a tag recognition model provided by an embodiment of the present application;
[0063] Figure 7 Schematic diagram of a video tag hierarchy provided by an embodiment of the present application;
[0064] Figure 8 Structural diagram of a sample training set provided by an embodiment of the present application;
[0065] Figure 9 Schematic diagram of a tag system provided by an embodiment of the present application;
[0066] Figure 10 Flowchart of a method for training a tag recognition model provided by an embodiment of the present application;
[0067] Figure 11 Training schematic diagram of a tag recognition model provided by an embodiment of the present application;
[0068] Figure 12 Flowchart of a tag recognition method provided by an embodiment of the present application;
[0069] Figure 13 Schematic diagram of a specific implementation scenario for tag recognition provided by an embodiment of the present application;
[0070] Figure 14 Structural diagram of a tag recognition device provided by an embodiment of the present application;
[0071] Figure 15 Structural diagram of a computing device provided by an embodiment of the present application;
[0072] Figure 16 Another structural diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0073] In order to make the objectives, technical solutions and beneficial effects of the present application more clear and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0074] The following explains some terms in the embodiments of the present application to facilitate the understanding of those skilled in the art.
[0075] 1. The object is the object that needs to be label-identified. In the embodiments of the present application, the object can be a multimedia file. A multimedia file refers to media information in the forms of text, image, video, sound, animation, etc., such as videos, short videos, audios, graphic and text content, etc. The object can also be an Internet entity, such as a product in an e-commerce system. The object can also be the content in an Internet search system, etc. In the embodiments of the present application, the object to be identified is the object involved in label identification using the label identification model, and the historical object is the object involved in training the label identification model.
[0076] 2. Hierarchical Multi-Label Classification is an important task in multi-classification in fields such as Natural Language Processing (NLP) and Computer Vision (CV). Its main feature is that the labels in hierarchical classification have a superior-subordinate relationship. The superior label is the parent of the subordinate label, and the finer the granularity of hierarchical classification is as it goes down to the lower levels.
[0077] 3. The constraint function is a loss function used to ensure the accurate classification of lower-level labels, that is, to ensure that the lower-level labels are subclasses of the upper-level labels. In this field, the hinge loss function can also be used as the constraint function. In a support vector machine, when constructing the objective function, the hinge loss function can be selected as the loss function. Moreover, the hinge loss function has a loss of 0 only when the classification is correct and the confidence is high enough. That is to say, the hinge loss function has higher requirements for learning.
[0078] As used hereinafter, the word "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" does not have to be construed as superior to or better than other embodiments.
[0079] The terms "first" and "second" in the text are only used for descriptive purposes and cannot be construed as explicitly or implicitly indicating relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0080] The design concept of the embodiments of the present application is briefly introduced below.
[0081] With the development of technology, the number of objects to be recognized is increasing. In order to more quickly implement business operations such as searching, querying, and recommending objects to be recognized, label recognition is performed on the objects to be recognized according to the text content of the objects to be recognized. During label recognition, multiple labels can be recognized for the same object to be recognized, and there is a certain correlation between the labels.
[0082] Currently, the correlation between labels is determined based on hierarchical classification. Among them, hierarchical classification is an important task in multi-classification in fields such as natural language processing and computer vision. Its main feature is that the labels in hierarchical classification have a superior-subordinate relationship, the superior label is the parent of the subordinate label, and the finer the granularity of the hierarchical classification is towards the lower level.
[0083] In related technologies, for multi-classification tasks based on a label system, most of them regard the classification task as several basic multi-classification tasks, directly predict the secondary labels of the objects to be recognized, and directly backtrack the primary labels from the prediction results. This method results in a low accuracy of label recognition for the objects to be recognized.
[0084] Therefore, how to improve the accuracy of label recognition for objects to be recognized is a technical problem that needs to be solved currently.
[0085] In view of this, the embodiments of the present application provide a label recognition method, device, equipment, and storage medium, which relate to fields such as computer technology and artificial intelligence technology, and can be applied to various scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.
[0086] In the embodiments of the present application, when performing label recognition on an object to be recognized, first, the text information of the object to be recognized is obtained, and the text semantic features of the text information are extracted; then, classification processing is performed on the object to be recognized based on the text semantic features to obtain at least one first target label of the object to be recognized and the corresponding first confidence information. For the label system, the higher the level, the coarser the granularity, the shorter the recognition difficulty, and at the same time, the label recognition accuracy is relatively high, that is, the accuracy of the recognized first target label is relatively high; then, the label semantic features of at least one first target label are respectively extracted, and based on the obtained at least one confidence, the corresponding label semantic features are concatenated to obtain a comprehensive label feature; finally, based on the comprehensive label feature and the text semantic features, at least one second target label of the object to be recognized is obtained, and the second target label is a sub-label of the first target label. It can be seen that when determining the second target label, the recognition result of the upper-level label with higher accuracy is used as prior knowledge, and together with the text semantic features, it is used as the input of the lower-level label recognition process. More reference information is fused in the lower-level label recognition process, and the label hierarchical relationship is fully utilized, thereby enhancing the recognition effect and improving the accuracy of label recognition for the object to be recognized.
[0087] The process of performing label recognition on the object to be recognized is mainly executed based on a label recognition model. Therefore, in order to ensure the accuracy of label recognition, an embodiment of this application also proposes a method for training a label recognition model. During the training of the label model, the relationship difference between the upper-layer label recognition result (i.e., the first predicted label in the embodiment of this application) and the lower-layer label recognition result (i.e., the second predicted label) is introduced, and the label hierarchy relationship is used to constrain the consistency of multi-level label recognition, that is, to ensure that the lower-layer label recognition result is a sub-label of the upper-layer label recognition result.
[0088] The embodiments of this application relate to artificial intelligence (AI) and machine learning technologies, and are designed based on speech technology, natural language processing technology, and machine learning (ML) in artificial intelligence.
[0089] Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence.
[0090] Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technologies mainly include several major directions such as computer vision technology, natural language processing technology, and machine learning / deep learning. With the research and progress of artificial intelligence technology, artificial intelligence has been studied and applied in multiple fields. For example, common applications include smart home, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, driverless, autonomous driving, robots, intelligent healthcare, etc. It is believed that with the development of technology, artificial intelligence will be applied in more fields and play an increasingly important role.
[0091] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0092] A brief description of the application scenarios set in this application is provided below. It should be noted that the following scenarios are only for illustrative purposes and not for limitation. In specific implementations, the technical solutions provided in the embodiments of this application can be flexibly applied according to actual needs.
[0093] See Figure 1 , Figure 1 which is a schematic diagram of an application scenario provided by an embodiment of this application. This application scenario includes a terminal device 110 and a server 120, and the terminal device 110 and the server 120 can communicate through a communication network.
[0094] In an alternative embodiment, the communication network can be a wired network or a wireless network. Therefore, the terminal device 110 and the server 120 can be directly or indirectly connected through wired or wireless communication methods. For example, the terminal device 110 can be indirectly connected to the server 120 through a wireless access point, or the terminal device 110 can be directly connected to the server 120 through the Internet. This application does not make any restrictions here.
[0095] Among them, the terminal device 110 includes, but is not limited to, devices such as mobile phones, tablet computers, laptop computers, desktop computers, e-book readers, intelligent voice interaction devices, smart home appliances, and in-vehicle terminals.
[0096] The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0097] Figure 1 The above is only an example. In fact, the number of the terminal device 110 and the server 120 is not limited and is not specifically defined in the embodiments of this application. In the embodiments of this application, when the number of the server 120 is multiple, multiple servers 120 can form a blockchain, and the server 120 is a node on the blockchain.
[0098] In a possible application scenario, the tag recognition method proposed in the embodiments of this application can be applied to the business scenario of multimedia file recommendation.
[0099] In the business scenario of multimedia file recommendation, a client supporting multimedia file playback function can be installed on the terminal device 110, and the server 120 is a background server corresponding to the client installed in the terminal device 110.
[0100] SeeFigure 2 , Figure 2 This is a schematic diagram of a specific application scenario provided by an embodiment of the present application, where the terminal device 110 is connected to the server 120 through a network.
[0101] Play a certain multimedia file in the client of the terminal device 110, and when receiving a multimedia file recommendation request, send a recommendation request carrying the multimedia file to the server through the network; wherein, the multimedia file recommendation request can be determined based on operations such as refreshing and searching triggered by the viewer;
[0102] After receiving the recommendation request, the server 120 first obtains the multimedia file in the recommendation request, identifies the text information of the multimedia file, and extracts the text semantic features of the text information; secondly, classifies the multimedia file based on the text semantic features to obtain at least one first target label of the multimedia file and the corresponding first credibility; then, extracts the label semantic features of at least one first target label respectively, and based on the obtained at least one first credibility, performs splicing processing on the corresponding label semantic features to obtain comprehensive label features; then, based on the comprehensive label features and the text semantic features, obtains at least one second target label of the object to be recognized; wherein, the second target label is a sub-label of the first target label; finally, obtains the multimedia file to be recommended from the multimedia resource pool based on the first target label and the associated second target label, sends the multimedia file to be recommended to the terminal device 110, and displays it on the client interface of the terminal device 110.
[0103] It should be noted that in addition to the business scenario of multimedia file recommendation, the label recognition method proposed in the embodiments of the present application can also be applied to all business scenarios that require text label extraction, such as content classification in search and product title classification in e-commerce systems.
[0104] The embodiments of the present application can also be implemented with the help of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing.
[0105] Among them, the above application scenario is only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.
[0106] To further illustrate the technical solutions provided by the embodiments of the present application, the following takes the server executing alone as an example and describes the label recognition model training method (i.e., the label recognition model training process) and the label recognition method (i.e., the label recognition model application process) provided by the exemplary embodiments of the present application in combination with the accompanying drawings.
[0107] Training process of the label recognition model:
[0108] In the embodiment of the present application, the training process of the label recognition model is a process of multi-round iterative training using the first sample set and the second sample set, which mainly includes a model design stage, a data preparation stage, and an iterative training stage. Each stage will be introduced separately below.
[0109] I. Model design stage
[0110] Refer to Figure 3 , Figure 3 which is a schematic structural diagram of a label recognition model provided in the embodiment of the present application. It can be seen from Figure 3 that the label recognition model includes an encoding layer, a first classification layer, a fusion layer, and a second classification layer. Among them, the encoding layer is used to extract features from the text information of the object to obtain corresponding text semantic features; the first classification layer is used to classify the object based on the text semantic features to obtain at least one first label (first predicted label or first target label) to which the object belongs and the corresponding first credibility; the fusion layer is used to fuse the comprehensive label features obtained by splicing the label semantic features of at least one first label based on the first credibility and the text semantic features to obtain the fused fused semantic features (historical fused semantic features or target fused semantic features); the second classification layer is used to classify the object based on the fused semantic features to obtain at least one second label (second predicted label or second target label) to which the object belongs; where the second label is a sub-label of the first label. Thus, the first label corresponding to the object and the corresponding second label can be determined.
[0111] In a possible implementation manner, the encoding layer can be implemented by a pre-trained semantic feature extraction model, that is, a pre-trained semantic feature extraction model is set in the encoding layer, or the pre-trained semantic feature extraction model is determined as the encoding layer; at this time, the pre-trained semantic feature extraction model is used to extract features from the text information to obtain the text semantic features of the text information.
[0112] Among them, the pre-trained semantic feature extraction model includes, but is not limited to: Bidirectional Encoder Representations from Transformers (BERT) model, Convolutional Neural Network (CNN) model, Long Short-Term Memory (LSTM), LSTM++Attention.
[0113] Exemplarily, taking the BERT model as the encoding layer in the label recognition model and extracting the features of a certain text information to obtain the text semantic features of the text information as an example, the implementation manner of obtaining the text semantic features will be described in detail. Refer to Figure 4 , Figure 4 FIG. Figure 4 is a schematic diagram of extracting text semantic features by the BERT model provided in an embodiment of the present application.
[0114] As can be seen from Figure 4 : According to the character granularity of the text information, the text information is input into the BERT model. In the BERT model, each character in the text information is first vector-encoded to obtain the corresponding encoded information. The obtained encoded information passes through multiple Transformers to obtain the semantic features of each character, and the semantic features of all characters are concatenated to obtain the text semantic features of the text information.
[0115] In a possible implementation manner, the first classification layer and the second classification layer can be implemented by a fully connected layer followed by a softmax layer; refer to Figure 5 , Figure 5 FIG. Figure 4 is a schematic diagram of identifying labels and credibility through the first classification layer provided in an embodiment of the present application.
[0116] As can be seen from Figure 5 : After obtaining the text semantic features, the text semantic features are input into the fully connected layer, that is, the input in the fully connected layer is the text semantic features. Then, after the text semantic features are mapped through at least one hidden layer, the mapping result is input into the softmax layer through the output layer, and the softmax layer determines the credibility of the mapping result; wherein, the mapping result represents at least one label to which the text information belongs.
[0117] In a possible implementation manner, the fusion layer introduces a fusion method of bilinear tensor decomposition vector mapping to effectively strengthen the interaction between the two-dimensional vector features of the comprehensive label features and the text semantic features, thereby improving the feature fusion effect.
[0118] Taking the label recognition of a video as an example, a specific structural schematic diagram of a label recognition model is provided. Refer to Figure 6 . As can be seen from Figure 6It can be seen from the following that the text information of the video "T-T, teach you the game strategy to reach 600" is input into the encoding layer; the text semantic features of the text information are output through the encoding layer, and the text semantic features are input into the first classification layer; the first labels "game", "teaching", and "strategy" to which the text information belongs are obtained through the first classification layer; the label semantic features of "game", "teaching", and "strategy" are obtained, and the label semantic features are concatenated to obtain comprehensive label features; the comprehensive label features and the text semantic features are input into the fusion layer, and the fusion layer performs fusion processing on the comprehensive label features and the text semantic features to obtain the fused semantic features after fusion; the fused semantic features are input into the second classification layer, and the second labels "single-player game", "mini-game", and "strategy" to which the text information belongs are obtained through the second classification layer.
[0119] The execution process of the above model will be introduced in detail in the subsequent content, so it will not be elaborated here too much.
[0120] II. Data Preparation Stage
[0121] Data preparation is of utmost importance in machine learning and can be said to be the most important link. The data preparation stage of the embodiments of the present application mainly includes: the training sample preparation process and the label system construction process.
[0122] 1. Training Sample Preparation Process
[0123] In the embodiments of the present application, the training samples include: an object, and a first annotation label and a second annotation label set for the object, and the second annotation label is a sub-label of the first annotation label.
[0124] In the training sample preparation process, for the target application scenario, the objects in the target application scenario can be obtained, and the corresponding first annotation label and second annotation label can be determined based on the text information of the object. The target application scenario can be business scenarios such as recommendation, storage, and search involving label recognition. The object can be a multimedia file, such as a video, audio, article, etc., or an Internet entity, such as a commodity in an e-commerce system, or the content in an Internet search system, etc. In addition, when the object is an Internet entity, the text information of the object is a descriptive text for explaining the content, features (such as the shape features and usage purposes of the commodity) of the object; when the object is a multimedia file, the text information of the object can be the title text, text summary of the multimedia file, or at least part of the text content obtained by content recognition of the multimedia file; for example, when the multimedia file is a video, the text information associated with the multimedia file can be obtained by recognizing content such as actor roles and subtitle texts; when the multimedia file is an audio, the text information can be obtained through the operation of converting the audio to text.
[0125] Taking an object as a video as an example, considering that the video title is one of the main components of the video, based on the title text of the video and combining basic elements such as natural language, the parsing of the text can be completed, thereby strengthening the understanding of the video semantic information and determining the tags to which the video belongs. This is one of the core tasks of a video search system and is helpful for video search; see Figure 7 , which is a schematic diagram of the video tag hierarchy provided by an embodiment of the present application. Figure 7 In , for the video content, the first annotation tag and the second annotation tag of the video content are determined according to the video title.
[0126] Therefore, when the object is a video, the text information is the video title, and the first annotation tag and the second annotation tag are determined based on the video title. See Figure 8 , Figure 8 , which is a structural diagram of a sample training set provided by an embodiment of the present application. The sample training set includes the video titles of multiple videos, as well as the first annotation tag and the second annotation tag corresponding to each video title.
[0127] 2. Label system construction process
[0128] Considering that the embodiment of the present application is a technical solution related to label recognition, and in the label recognition process, it is based on the text information of the object to match with multiple preset labels, and the label to which the object belongs is determined according to the matching result. At the same time, considering that there is more than one label to which the object belongs, and in the case of multiple labels, there will be an association relationship between the labels. Therefore, the embodiment of the present application pre-constructs a label system for a label recognition model corresponding to the object. The label system includes two levels of labels with a hierarchical relationship; that is, the label system includes a first label and an associated second label.
[0129] Exemplarily, taking the construction of a label system for a video-related business as an example, that is, when the object is a video, a label system for a label recognition model corresponding to the video can be constructed; see Figure 9 , Figure 9 is a schematic diagram of a label system provided by an embodiment of the present application. The label system includes two levels of labels with a hierarchical relationship. The first label includes thematic coarse-grained labels such as sports, games, entertainment, and dance. The first label is further divided into multiple second labels. For example, under sports, it is divided into basketball, football, sports meetings, etc.; under games, it is divided into mini-games, mobile games, PC games, etc.; under entertainment, it is divided into movies, variety shows, stars, etc.; under dance, it is divided into square dance, hip-hop, ballroom dance, etc. Based on this label system, the first label to which the video belongs and the second label associated with the first label are obtained; for example, for the video information with the video title "Square dance makes you healthier", it can be classified into the first label corresponding to "dance" and the second label corresponding to "square dance".
[0130] In the embodiments of the present application, after the training samples for training are prepared, these training samples can be used to train the constructed model.
[0131] III. Iterative training stage
[0132] In the embodiments of the present application, through iterative training of the label recognition model, a trained label recognition model is obtained. During the model training process, a total of N rounds (such as 100) of iteration are performed on all sample pairs. Among them, one round of iteration means that all sample pairs are trained once in the label recognition model. In each round of iteration, due to the limited video memory resources of the training machine, all sample pairs cannot be input into the model for training at one time. Therefore, all sample pairs need to be trained in batches. Each batch of samples is generated by means such as random division, and each batch of samples is respectively input into the model for forward calculation, backward calculation, model parameter update and other trainings.
[0133] Before the first round of training, the parameters of the label recognition model are initialized. Further, after setting hyperparameters such as batch size, number of epochs, and learning rate respectively, the training starts, and finally a target label recognition model is obtained.
[0134] See Figure 10 , which is a flowchart of a method for training a label recognition model provided by the embodiments of the present application, including the following steps:
[0135] Step S100, obtain a sample training set; each training sample is a historical object including a first labeled label and a second labeled label, and the second labeled label is a sub-label of the first labeled label.
[0136] Step S101, select a training sample from the sample training set.
[0137] Step S102, for the historical object in the training sample, obtain the historical text information of the historical object.
[0138] In a possible implementation manner, for different historical objects, the method for determining the historical text information is different; for example: when the historical object is a video, the video title of the video is used as the historical text information; when the historical object is an audio, the text information converted from the audio is used as the text information; when the historical object is an article, the title of the article is used as the historical file information; when the historical object is a commodity, the description information of the commodity is used as the text information.
[0139] Step S103, input the historical text information, the first labeled label, and the second labeled label of the historical object into the label recognition model to be trained, and the label recognition model to be trained executes the following steps S1030 to S1035:
[0140] Step S1030: Extract the historical text semantic features of the historical text information through the encoding layer.
[0141] When extracting the historical text semantic features of the historical text information through the encoding layer, the text information input into the encoding layer has been subjected to word segmentation processing; that is, before inputting the text information into the encoding layer, the text information should be first subjected to word segmentation processing to obtain at least one historical text word segment.
[0142] In a possible implementation, the way to perform word segmentation processing on the text information is: perform word cutting according to semantic information, and stop words can be removed during the word cutting process, and the data after word cutting is used as the historical text word segment; for example, if the text information is "Game strategy to teach you to get 600 points", perform word cutting on the text information according to semantic information, and the at least one historical text word segment obtained is "teach", "you", "get", "600 points", "game", "strategy".
[0143] In a possible implementation, the way to perform word segmentation processing on the text information is: keyword extraction, and the extracted keywords are used as at least one historical text word segment; for example, if the text information is "Game strategy to teach you to get 600 points", perform keyword extraction, and the at least one historical text word segment obtained is "get 600 points", "game", "strategy";
[0144] In a possible implementation, the way to perform word segmentation processing on the text information is: perform word segmentation according to character granularity, and each character in the text information is a historical text word segment; for example, if the text information is "Game strategy to teach you to get 600 points", perform word segmentation according to character granularity, and the at least one historical text word segment obtained is "teach", "you", "get", "600", "points", "game", "strategy", "attack", "strategy".
[0145] It should be noted that whether word segmentation processing is required and in which way to perform word segmentation to obtain the historical text word segment input into the encoding layer is determined by the encoding layer.
[0146] In a possible implementation, when extracting the historical text semantic features of the historical text information through the encoding layer: the encoding layer respectively extracts features from at least one historical text word segment of the historical text information to obtain the corresponding historical text semantic features, and performs splicing processing on the obtained at least one historical text semantic feature to obtain the corresponding historical text semantic feature.
[0147] Step S1031: Through the first classification layer, classify the historical object based on the historical text semantic features to obtain at least one first prediction label and the corresponding first credibility.
[0148] In a possible implementation, when classifying historical objects based on historical text semantic features to obtain at least one first prediction label and the corresponding first confidence level: determining the scores of the historical text belonging to each first label in the label system based on the historical text semantic features, obtaining at least one first label from the label system as the first prediction label based on the scores of each first label in the label system, and determining the corresponding first confidence level by normalization based on the scores associated with the first prediction labels. Exemplarily, sorting the first labels according to the scores and selecting the top Top-K first labels as the first prediction labels, where K is a positive integer greater than or equal to 1; or selecting the first labels with scores greater than the first score threshold as the first prediction labels.
[0149] In another possible implementation, when classifying historical objects based on historical text semantic features to obtain at least one first prediction label and the corresponding first confidence level: determining the scores of the historical text belonging to each first label in the label system based on the historical text semantic features, performing normalization processing on the scores to obtain the corresponding first confidence level, and selecting the first prediction labels from the label system based on the first confidence levels of each first label; Exemplarily, sorting the first labels according to the first confidence levels and selecting the top Top-K first labels as the first prediction labels, where K is a positive integer greater than or equal to 1; or selecting the first labels with first confidence levels greater than the first confidence level threshold as the first prediction labels.
[0150] Step S1032: Obtain the label semantic features of at least one first prediction label.
[0151] It should be noted that the label semantic features of the first prediction labels are trained based on general text, such as based on word2vec. Exemplarily, the label semantic features of games = [0.1, -0.7, 0.03… -0.9], the label semantic features of movies = [0.3, -0.2, 0.03… -0.8], and the label semantic features of sports = [-0.1, -0.9, 0.23… 0.1].
[0152] Step S1033: Based on the at least one first confidence level obtained, perform splicing processing on the corresponding label semantic features to obtain historical comprehensive label features.
[0153] In a possible implementation, when obtaining the historical comprehensive label feature by performing splicing processing on the corresponding label semantic features based on at least one obtained first credibility: the label semantic features of at least one first predicted label are weighted respectively by using the first credibility of at least one first predicted label to obtain at least one weighted label feature; the at least one weighted label feature is concatenated to obtain the historical comprehensive label feature. That is, a weighted sum of the label semantic features of the first predicted label is performed, and the weight of each label semantic feature is the first credibility of the corresponding first predicted label.
[0154] Step S1034, through a fusion layer, fuse the historical comprehensive label feature and the historical text semantic feature to obtain the fused historical fused semantic feature.
[0155] In a possible implementation, when fusing the historical comprehensive label feature and the historical text semantic feature to obtain the fused historical fused semantic feature: First, obtain a label parameter matrix and a text parameter matrix for feature fusion; then, perform second-order linear fusion processing on the historical comprehensive label feature based on the label parameter matrix to obtain a second-order fused label feature; and perform second-order linear fusion processing on the historical text semantic feature based on the text parameter matrix to obtain a second-order fused semantic feature; finally, fuse the second-order fused label feature and the second-order fused semantic feature to obtain the fused historical fused semantic feature.
[0156] In the embodiments of the present application, when fusing the label feature and the text semantic feature, a fusion method of bilinear tensor decomposition is introduced, which can effectively strengthen the fusion interaction of the two-dimensional features, thereby improving the effect.
[0157] The fusion interaction process of the two-dimensional features of the label feature and the semantic feature is specifically as follows:
[0158] l2 logits = l1 emb *(U * V T ) * E emb T + b
[0159] After transformation, it is:
[0160] l2 logits = (U * l1 emb ) ο (V * E emb ) T + b
[0161] Among them, l2 logits represents the fused historical fused semantic feature, U represents the label parameter matrix, l1 emb represents the historical comprehensive label feature, V represents the text parameter matrix, Eemb a represents the semantic features of historical texts, b represents the bias vector, "*" represents matrix-vector multiplication, corresponding to second-order linear fusion processing, and "ο" represents vector dot product, corresponding to the fusion processing of second-order fusion label features and second-order fusion semantic features.
[0162] The feature fusion method proposed in the embodiment of the present application reduces the complexity of the feature matrix. Compared with the fusion method involving the introduction of the W tensor in the related art, the number of parameters is reduced. Exemplarily, the number of parameters of the introduced W tensor is m*n*k; while in the present application, if U*V is used to fit W, where U is k*m and V is n*k, the number of parameters here is k(m + n), and the number of parameters is less than m*n*k. Additionally, the accuracy effect of the finally recognized label is also better here.
[0163] Step S1035, through the second classification layer, classify the historical object based on the historical fusion semantic features to obtain at least one second predicted label.
[0164] In a possible implementation manner, when classifying the historical object based on the historical fusion semantic features to obtain at least one second predicted label: determine the scores of the historical text belonging to each second label in the label system based on the historical fusion semantic features, and based on the scores of each second label in the label system, obtain at least one second label from the label system as the second predicted label, and determine the corresponding second credibility through a normalization method based on the scores associated with the second predicted label. Exemplarily, sort the second labels according to the scores, and select the top Top-K second labels as the second predicted labels, where K is a positive integer greater than or equal to 1; or select the second labels with scores greater than the first score threshold as the second predicted labels.
[0165] In another possible implementation manner, when classifying the historical object based on the historical fusion semantic features to obtain at least one second predicted label: determine the scores of the historical text belonging to each second label in the label system based on the historical fusion semantic features, perform normalization processing on the scores to obtain the corresponding second credibility, and select the second predicted label from the label system based on the second credibility of each second label; Exemplarily, sort the second labels according to the second credibility, and select the top Top-K second labels as the second predicted labels, where K is a positive integer greater than or equal to 1; or select the second labels with second credibility greater than the second credibility threshold as the second predicted labels.
[0166] Step S104, based on the first predicted label and the second predicted label, construct a loss function, and use the loss function to adjust the parameters of the label recognition model to be trained.
[0167] In a possible implementation, when constructing a loss function based on a first prediction label and a second prediction label:
[0168] First, determine a first loss function based on the main label difference between the first prediction label and the first annotation label;
[0169] And determine a second loss function based on the sub-label difference between the second prediction label and the second annotation label;
[0170] Meanwhile, determine a constraint function based on the relationship difference between the first prediction label and the second prediction label; wherein, the constraint function is used to ensure that the second prediction label is a sub-label of the first prediction label;
[0171] Finally, construct a loss function based on the first loss function, the second loss function, and the constraint function.
[0172] Exemplarily, the first loss function can adopt a negative log-likelihood loss function, and the specific formula is as follows:
[0173]
[0174] wherein, cls1 corresponds to the first classification layer, n is the number of first labels in the label system, y i takes 0 or 1, representing whether the object belongs to the first label, whether the first label is the i-th class, and a i is the first confidence of the first prediction label.
[0175] Exemplarily, the second loss function can adopt a negative log-likelihood loss function, and the specific formula is as follows:
[0176]
[0177] wherein, cls2 corresponds to the second classification layer, m is the number of second labels in the label system, y j takes 0 or 1, representing whether the object belongs to the second label, whether the second label is the j-th class, and a j is the second confidence of the second prediction label.
[0178] Exemplarily, the constraint function can adopt a hinge loss function. The specific formula is as follows:
[0179]
[0180] Among them, the constraint function is used to constrain the consistency between the first predicted label and the second predicted label, that is, to ensure that the predicted second predicted label belongs to the first predicted label, which means that the second predicted label is a sub-label of the first predicted label. λ is a hyperparameter greater than 0. The meaning of λ + l2_score - l1_score is that it is expected that the score l1_score of the first predicted label obtained by prediction is always greater than the corresponding second predicted label l2_score. If λ + l2_score - l1_score <= 0, that is, l1_score is greater than l2_score by λ, then lossh = 0 and no loss is generated.
[0181] Based on the first loss function, the second loss function, and the constraint function, the constructed loss function is:
[0182] loss = λ1loss cls1 + λ2loss cls2 + λ3loss h
[0183] Among them, λ1, λ2, and λ3 are the weights of each loss function.
[0184] It should be noted that in addition to using the weighted sum method as the final loss function of the task, other loss function reconciliation methods can also be replaced. In addition to the negative log loss function, it can also be set as the cross-entropy loss function.
[0185] In this application, during the model training process, during the label recognition process, the upper-layer label prediction result will be input into the lower-layer label prediction model as prior knowledge, so as to realize that multi-dimensional features are used as the input of the second predicted label prediction process. The recognition process of the second predicted label can refer to more features to ensure accuracy. And in terms of multi-dimensional feature fusion, a second-order tensor mapping method based on bilinear tensor multiplication is proposed to strengthen the interaction of multi-dimensional vector features, ensure that the input information of the second predicted label is more accurate, and introduce the relationship difference between the first predicted label and the second predicted label into the loss function to constrain the consistency of the multi-level label results, effectively utilizing the upper and lower layer constraint relationships of the label system, so as to ensure the accuracy of the label recognition model.
[0186] For the sake of easy understanding, taking the label recognition of a video as an example, the training process of the label recognition model is described; among them, the encoding layer of the label recognition model is the BERT model, and the first classification layer and the second classification layer are fully connected layers followed by softmax layers; see Figure 11 This application embodiment provides a training schematic diagram of a label recognition model. From Figure 11 it can be seen that:
[0187] The video title of the video is "T-T, Teach You the Game Strategy to Score 600 Points", and the video title carries two-level annotation tags. The first annotation tag is "Game", and the second annotation tag is "Mini Game"; taking the video title as historical text information, inputting the historical text information into the BERT model at the character granularity, and obtaining the historical text semantic features through the BERT model; inputting the historical text semantic features into the first classification layer, in the first classification layer, after the historical text semantic features are processed by the fully connected layer mapping and the softmax layer normalization, determining each first candidate label and the corresponding first credibility of the historical text information associated with the historical text semantic features; selecting the first candidate labels with the top K first credibilities from each first candidate label as the first prediction labels, and determining the first loss function based on the first prediction labels and the first annotation labels; determining the historical label semantic features of each first prediction label, and splicing all the determined historical label semantic features to obtain the corresponding historical comprehensive label features; inputting the historical comprehensive label features into the fusion layer, and obtaining the historical fusion semantic features by fusing the historical comprehensive label features with the historical text semantic features; inputting the historical fusion semantic features into the second classification layer, in the second classification layer, after the historical fusion semantic features are processed by the fully connected layer mapping and the softmax layer normalization, determining each second candidate label and the corresponding second credibility of the historical text information associated with the historical fusion semantic features; selecting the second candidate labels with the top K second credibilities from each second candidate label as the second prediction labels, and determining the second loss function based on the second prediction labels and the second prediction labels; determining the constraint function based on the first prediction labels and the second prediction labels; determining the loss function based on the first loss function, the second loss function and the constraint function, and adjusting the parameters based on the loss parameter until the loss function value converges or the number of iterations reaches the upper limit.
[0188] Application process of the label recognition model:
[0189] See Figure 12 , which is a flowchart of a label recognition method provided by an embodiment of the present application, including the following steps:
[0190] Step S1200, obtaining the text information of the object to be recognized, and extracting the text semantic features of the text information.
[0191] In a possible implementation manner, when obtaining the text information of the object to be recognized and extracting the text semantic features of the text information: First, perform word segmentation on the text information to obtain at least one corresponding text segment; then, perform feature extraction on each of the at least one text segment to obtain the corresponding word semantic features; finally, splice the at least one obtained word semantic features to obtain the text semantic features.
[0192] Step S1201: Classify the object to be recognized based on the text semantic features to obtain at least one first target label of the object to be recognized and the corresponding first confidence level.
[0193] In a possible implementation manner, when classifying the object to be recognized based on the text semantic features to obtain at least one first target label of the object to be recognized and the corresponding first confidence level: classify the object to be recognized based on the text semantic features to obtain at least one first candidate label to which the object to be recognized belongs and the corresponding first confidence level; select at least one first target label whose first confidence level meets the confidence level condition from the at least one first candidate label.
[0194] Step S1202: Extract the label semantic features of at least one first target label respectively, and perform splicing processing on the corresponding label semantic features based on the obtained at least one first confidence level to obtain comprehensive label features.
[0195] In a possible implementation manner, when performing splicing processing on the corresponding label semantic features based on the obtained at least one first confidence level to obtain comprehensive label features: use at least one first confidence level to perform weighted processing on the corresponding label semantic features respectively to obtain at least one weighted label feature; perform serial splicing processing on the at least one weighted label feature to obtain comprehensive label features.
[0196] Step S1203: Obtain at least one second target label of the object to be recognized based on the comprehensive label features and the text semantic features; wherein, the second target label is a sub-label of the first target label.
[0197] In a possible implementation manner, when obtaining at least one second target label of the multimedia text based on the comprehensive label features and the text semantic features: First, obtain a label parameter matrix and a text parameter matrix for feature fusion, and perform second-order linear fusion processing on the comprehensive label features based on the label parameter matrix to obtain second-order fusion label features, and perform second-order linear fusion processing on the text semantic features based on the text parameter matrix to obtain second-order fusion semantic features; then, perform fusion processing on the second-order fusion label features and the second-order fusion semantic features to obtain the fused target fusion semantic features; then, classify the object to be recognized based on the target fusion semantic features to obtain at least one second candidate label to which the object to be recognized belongs and the corresponding second confidence level; finally, select at least one second target label whose second confidence level meets the screening condition from the at least one second candidate label.
[0198] It should be noted that for the training process and application process of the label recognition model, the operations performed inside the label recognition model are the same. Therefore, the content discussed in detail during the training process will not be repeated during the application process.
[0199] For ease of understanding, taking the label recognition in the video recommendation process as an example, a schematic diagram of a specific implementation scenario of label recognition is provided. Refer to Figure 13 , from Figure 13 it can be seen that the terminal device sends a recommendation request to the server, and the recommendation request carries a video. After receiving the recommendation request, the server obtains the video title of the video in the recommendation request, uses the video title as text information, determines the first target label and the second target label of the video, and based on the first target label and the second target label, obtains the videos to be recommended associated with the first target label and the second target label from the video repository, and feeds back the videos to be recommended to the terminal device, thus completing the video recommendation process.
[0200] In this application, when recognizing the label of the object to be recognized based on the text information of the object to be recognized, first, the text information of the object to be recognized is obtained, and the text semantic features of the text information are extracted; then, based on the text semantic features, the object to be recognized is classified to obtain at least one first target label of the object to be recognized and the corresponding first confidence information. For the label system, the higher the level, the coarser the granularity, the shorter the recognition difficulty, and at the same time, the label recognition accuracy is relatively high, that is, the accuracy of the recognized first target label is relatively high; then, the label semantic features of at least one first target label are respectively extracted, and based on the obtained at least one first confidence, the corresponding label semantic features are spliced to obtain a comprehensive label feature; finally, based on the comprehensive label feature and the text semantic features, at least one second target label of the object to be recognized is obtained, and the second target label is a sub-label of the first target label. It can be seen that when determining the second target label, the upper-layer label recognition result with higher accuracy is used as prior knowledge, and together with the text semantic features, it is used as the input of the lower-layer label recognition process. More reference information is fused in the lower-layer label recognition process, and the label hierarchy relationship is fully utilized, thereby enhancing the recognition effect and improving the accuracy of the label recognition of the object to be recognized.
[0201] It should be noted that although the operations of the method of this application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0202] In addition, it should be noted that in the specific implementation of this application, when it comes to data related to users, when the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.
[0203] Based on the same inventive concept, an embodiment of this application also provides a label recognition device. Refer to Figure 14 , which is a structural diagram of a label recognition device provided by an embodiment of this application. The label recognition device 1400 includes: an acquisition unit 1401, a classification unit 1402, a splicing unit 1403, and an obtaining unit 1404; where:
[0204] The acquisition unit 1401 is configured to acquire the text information of the object to be recognized and extract the text semantic features of the text information.
[0205] The classification unit 1402 is configured to perform classification processing on the object to be recognized based on the text semantic features to obtain at least one first target label of the object to be recognized and the corresponding first credibility.
[0206] The splicing unit 1403 is configured to respectively extract the label semantic features of at least one first target label and perform splicing processing on the corresponding label semantic features based on the obtained at least one first credibility to obtain comprehensive label features.
[0207] The obtaining unit 1404 is configured to obtain at least one second target label of the object to be recognized based on the comprehensive label features and the text semantic features; where the second target label is a sub-label of the first target label.
[0208] In a possible implementation manner, the obtaining unit 1404 is specifically configured to:
[0209] Perform fusion processing on the comprehensive label features and the text semantic features to obtain the fused target fusion semantic features;
[0210] Perform classification processing on the object to be recognized based on the target fusion semantic features to obtain at least one second candidate label to which the object to be recognized belongs and the corresponding second credibility;
[0211] Select at least one second target label from the at least one second candidate label whose second credibility meets the screening conditions.
[0212] In a possible implementation manner, the obtaining unit 1404 is specifically configured to:
[0213] Obtain a label parameter matrix and a text parameter matrix for feature fusion;
[0214] Perform a second-order linear fusion process on the comprehensive label features based on the label parameter matrix to obtain second-order fused label features;
[0215] Perform a second-order linear fusion process on the text semantic features based on the text parameter matrix to obtain second-order fused semantic features;
[0216] Fuse the second-order fused label features and the second-order fused semantic features to obtain the fused target fused semantic features.
[0217] In a possible implementation manner, the obtaining unit 1401 is specifically configured to:
[0218] Perform word segmentation on the text information to obtain at least one corresponding text word segment;
[0219] Extract features from at least one text word segment respectively to obtain corresponding word segment semantic features,
[0220] Perform splicing processing on the obtained at least one word segment semantic feature to obtain text semantic features.
[0221] In a possible implementation manner, the classification unit 1402 is specifically configured to:
[0222] Perform classification processing on the object to be recognized based on the text semantic features to obtain at least one first candidate label to which the object to be recognized belongs and corresponding first credibility;
[0223] Select at least one first target label from the at least one first candidate label whose first credibility meets the credibility condition.
[0224] In a possible implementation manner, the splicing unit 1403 is specifically configured to:
[0225] Use at least one first credibility to perform weighted processing on the corresponding label semantic features respectively to obtain at least one weighted label feature;
[0226] Perform concatenation splicing processing on the at least one weighted label feature to obtain comprehensive label features.
[0227] In a possible implementation manner, the label recognition method involved in the embodiments of the present application is executed by a label recognition model; wherein, the label recognition device further includes a training unit 1405, and the label recognition model is obtained by the training unit 1405 through the following manner:
[0228] Perform cyclic iterative training on the label recognition model to be trained according to the sample training set to obtain the label recognition model; perform the following operations in one cyclic iteration:
[0229] Select training samples from the sample training set; where the training samples are historical objects that include a first annotation label and a second annotation label, and the second annotation label is a sub-label of the first annotation label.
[0230] Input the training samples into the label recognition model to be trained, and predict the first predicted label and the second predicted label associated with the historical object.
[0231] Based on the first predicted label and the second predicted label, construct a loss function, and use the loss function to adjust the parameters of the label recognition model to be trained.
[0232] In a possible implementation, the training unit 1405 is specifically configured to:
[0233] Determine a first loss function based on the main label difference between the first predicted label and the first annotation label, determine a second loss function based on the sub-label difference between the second predicted label and the second annotation label, and determine a constraint function based on the relationship difference between the first predicted label and the second predicted label.
[0234] Construct a loss function based on the first loss function, the second loss function, and the constraint function; where the constraint function is used to ensure that the second predicted label is a sub-label of the first predicted label.
[0235] In a possible implementation, the label recognition model to be trained includes: a first classification layer, a fusion layer, and a second classification layer; the training unit 1405 is specifically configured to:
[0236] Through the first classification layer, classify the historical object based on the historical text semantic features extracted from the historical text information of the historical object, and obtain at least one first predicted label.
[0237] Perform splicing processing on the historical label semantic features of at least one first predicted label to obtain corresponding historical comprehensive label features.
[0238] Through the fusion layer, perform fusion processing on the historical comprehensive label features and the historical text semantic features to obtain the fused historical fusion semantic features.
[0239] Through the second classification layer, classify the historical object based on the historical fusion semantic features to obtain at least one second predicted label.
[0240] It should be noted that although several units (or modules) of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more units (or modules) described above can be embodied in one unit (or module). Conversely, the features and functions of one unit (or module) described above can be further divided and embodied by multiple units (or modules). Of course, when implementing the present application, the functions of each unit (or module) can also be implemented in the same or multiple software or hardware.
[0241] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0242] After introducing the label recognition method and device of the exemplary embodiments of the present application, the following introduces another exemplary embodiment of the present application, the computing device.
[0243] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0244] In a possible implementation manner, the computing device provided by the embodiments of the present application can at least include a processor and a memory. Among them, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes any step in the label recognition methods of various exemplary embodiments of the present application.
[0245] In another possible implementation manner, the computing device can be a terminal device, and the structure of the computing device can be as Figure 15 shown, including: communication component 1510, memory 1520, display unit 1530, camera 1540, sensor 1550, audio circuit 1560, Bluetooth module 1570, processor 1580 and other components.
[0246] The communication component 1510 is used to communicate with a server. In some embodiments, it may include a Wi-Fi (Wireless Fidelity) module. The Wi-Fi module belongs to short-range wireless transmission technology. Through the Wi-Fi module, the computing device can help users send and receive information.
[0247] The memory 1520 can be used to store software programs and data. By running the software programs or data stored in the memory 1520, the processor 1580 executes various functions of the terminal device and data processing. The memory 1520 may include high-speed random access memory, and may also include non-volatile memory, such as at least one type of disk storage device, flash memory device, or other volatile solid-state storage devices. The memory 1520 stores an operating system that enables the terminal device to run. In this application, the memory 1520 can store the operating system and various application programs, and can also store the code for implementing the label recognition method of the embodiments of this application.
[0248] The display unit 1530 can also be used to display information input by the user or information provided to the user, as well as the graphical user interface (GUI) of various menus of the terminal device. Specifically, the display unit 1530 may include a display screen 1532 disposed on the front of the terminal device. Among them, the display screen 1532 can be configured in the form of a liquid crystal display, a light-emitting diode, etc.
[0249] The display unit 1530 can also be used to receive input numerical or character information, and generate signal inputs related to the user settings and function control of the terminal device. Specifically, the display unit 1530 may include a touch screen 1531 disposed on the front of the terminal device, which can collect touch operations of the user thereon or nearby, such as clicking buttons, dragging scroll boxes, etc.
[0250] Among them, the touch screen 1531 can cover the display screen 1532, or the touch screen 1531 and the display screen 1532 can be integrated to implement the input and output functions of the terminal device. After integration, it can be simply referred to as a touch display screen.
[0251] The camera 1540 can be used to capture static images. The camera 1540 can be one or multiple. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 1580 to be converted into a digital image signal.
[0252] The terminal device may further include at least one sensor 1550, such as an acceleration sensor 1551, a distance sensor 1552, a fingerprint sensor 1553, and a temperature sensor 1554. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.
[0253] The audio circuit 1560, the speaker 1561, and the microphone 1562 can provide an audio interface between the user and the terminal device. The audio circuit 1560 can transmit the electrical signal converted from the received audio data to the speaker 1561, and the speaker 1561 converts it into a sound signal for output. The terminal device may also be configured with volume buttons for adjusting the volume of the sound signal. On the other hand, the microphone 1562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1560, converted into audio data, and then the audio data is output to the communication component 1510 to be sent to, for example, another terminal device, or the audio data is output to the memory 1520 for further processing.
[0254] The Bluetooth module 1570 is used to interact with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable computing device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 1570 to perform data interaction.
[0255] The processor 1580 is the control center of the terminal device, connecting various parts of the entire terminal using various interfaces and lines. By running or executing software programs stored in the memory 1520 and calling data stored in the memory 1520, it executes various functions of the terminal device and processes data. In some embodiments, the processor 1580 may include one or more processing units; the processor 1580 may also integrate an application processor and a baseband processor, where the application processor mainly processes the operating system, user interface, and application programs, etc., and the baseband processor mainly processes wireless communication. It can be understood that the above baseband processor may not be integrated into the processor 1580. In this application, the processor 1580 can run the operating system, application programs, user interface display, and touch response, as well as the label recognition method of the embodiments of this application; in addition, the processor 1580 is coupled to the display unit 1530.
[0256] In another possible implementation, the computing device may be a server. The structure of the computing device may be as Figure 16 shown. The components of the computing device 1600 may include but are not limited to: at least one processor 1601, at least one memory 1602, and a bus 1603 connecting different system components (including the memory 1602 and the processor 1601).
[0257] The memory 1602 can be used to store software programs and data. The processor 1601 executes various functions and data processing by running the software programs or data stored in the memory 1602, so as to implement the label recognition method of the embodiments of the present application.
[0258] The bus 1603 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a processor, or a local area bus using any bus structure in a variety of bus structures.
[0259] The memory 1602 may include a readable medium in the form of volatile memory, such as a random access memory (RAM) 16021 and / or a cache memory 16022, and may further include a read-only memory (ROM) 16023.
[0260] The memory 1602 may further include a program / utilities 16025 having a set (at least one) of program modules 16024. Such program modules 16024 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0261] The computing device 1600 can also communicate with one or more external devices 1604 (such as a keyboard, a pointing device, etc.), can also communicate with one or more devices that enable a user to interact with the computing device 1600, and / or communicate with any device that enables the computing device 1600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 1605. And, the computing device 1600 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1606. As Figure 16 shown, the network adapter 1606 communicates with other modules for the computing device 1600 through the bus 1603. It should be understood that although Figure 16 not shown in the figure, other hardware and / or software modules can be used in combination with the computing device 1600, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0262] In some possible implementation manners, aspects of the label recognition method provided in the present application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on a computing device, the computer program is used to enable the computing device to execute the steps in the label recognition method according to various exemplary embodiments of the present application described above in this specification.
[0263] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0264] The program product of the embodiments of the present application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may be run on a computing device. However, the program product of the present application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0265] The readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a readable computer program is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0266] The computer program contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0267] The computer program for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The computer program may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0268] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable computer programs.
[0269] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable devices generate means for implementing the specified functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0270] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0271] These computer program instructions can also be loaded onto a computer or other programmable device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0272] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0273] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to cover these changes and modifications.
Claims
1. A tag recognition method, characterized in that: The method comprises: Acquire text information of the object to be identified, and extract text semantic features of the text information; Classify the object to be identified based on the text semantic features to obtain at least one first target label and a corresponding first credibility of the object to be identified; Extracting label semantic features of the at least one first target label respectively, and concatenating corresponding label semantic features based on the obtained at least one first credibility to obtain a comprehensive label feature; Based on the comprehensive label feature and the text semantic feature, at least one second target label of the object to be identified is obtained; wherein the second target label is a sub-label of the first target label.
2. The method according to claim 1, characterized in that The obtaining, based on the comprehensive tag feature and the text semantic feature, at least one second target tag of the multimedia text comprises: Fusing the comprehensive label feature and the text semantic feature to obtain a fused target fused semantic feature; Classify the object to be identified based on the target fusion semantic feature to obtain at least one second candidate label to which the object to be identified belongs and a corresponding second credibility; The at least one second target tag whose second credibility meets the filtering condition is screened out from the at least one second candidate tag.
3. The method according to claim 2, characterized in that The fusing process of the comprehensive label feature and the text semantic feature to obtain the fused target fused semantic feature includes: Get the label parameter matrix and text parameter matrix for feature fusion; Performing a second-order linear fusion process on the comprehensive label features based on the label parameter matrix to obtain a second-order fusion label feature; Performing second-order linear fusion processing on the text semantic features based on the text parameter matrix to obtain second-order fused semantic features; The second-order fusion label feature and the second-order fusion semantic feature are fused to obtain a fused target fusion semantic feature.
4. The method according to claim 1, characterized in that The step of obtaining text information of an object to be identified and extracting text semantic features of the text information includes: Performing word segmentation processing on the text information to obtain at least one corresponding text word segmentation; Feature extraction is performed on the at least one text segmentation to obtain corresponding segmentation semantic features, and concatenation processing is performed on the obtained at least one segmentation semantic feature to obtain the text semantic feature.
5. The method according to claim 1, characterized in that The classifying process of the object to be identified based on the text semantic feature to obtain at least one first target label of the object to be identified and a corresponding first credibility includes: Classify the object to be identified based on the text semantic features to obtain at least one first candidate label to which the object to be identified belongs and a corresponding first credibility; At least one first target tag whose first credibility meets the credibility condition is screened out from the at least one first candidate tag.
6. The method according to claim 1, characterized in that The step of performing splicing processing on corresponding label semantic features based on the at least one first credibility obtained to obtain a comprehensive label feature includes: Using at least one first credibility, weighting the corresponding label semantic features respectively to obtain at least one weighted label feature; The at least one weighted label feature is serially concatenated to obtain the comprehensive label feature.
7. The method according to any one of claims 1 to 6, characterized in that: The tag recognition method is performed by a tag recognition model, and the tag recognition model is obtained by training in the following manner: According to the sample training set, a loop iteration training is performed on the label recognition model to be trained to obtain the label recognition model; the following operations are performed in one loop iteration: Selecting training samples from the sample training set; wherein the training samples are: historical objects including a first annotation label and a second annotation label, wherein the second annotation label is a sub-label of the first annotation label; Inputting the training sample into the to-be-trained label recognition model to predict a first predicted label and a second predicted label associated with the historical object; A loss function is constructed based on the first predicted label and the second predicted label, and the loss function is used to adjust parameters of the label recognition model to be trained.
8. The method according to claim 7, characterized in that The constructing a loss function based on the first predicted label and the second predicted label includes: Determine a first loss function based on a main label difference between the first predicted label and the first labeled label, determine a second loss function based on a sub-label difference between the second predicted label and the second labeled label, and determine a constraint function based on a relationship difference between the first predicted label and the second predicted label; The loss function is constructed based on the first loss function, the second loss function and the constraint function; wherein the constraint function is used to ensure that the second predicted label is a sublabel of the first predicted label.
9. The method according to claim 7, characterized in that The label recognition model to be trained includes: a first classification layer, a fusion layer and a second classification layer; The step of inputting the training sample into the to-be-trained label recognition model to predict the first predicted label and the second predicted label associated with the historical object comprises: By means of the first classification layer, based on the historical text semantic features extracted from the historical text information of the historical object, the historical object is classified to obtain at least one first prediction label; Performing concatenation processing on the historical tag semantic features of the at least one first predicted tag to obtain corresponding historical comprehensive tag features; Through the fusion layer, the historical comprehensive label feature and the historical text semantic feature are fused to obtain a fused historical fusion semantic feature; The historical object is classified based on the historical fusion semantic feature through the second classification layer to obtain at least one second prediction label.
10. A tag recognition device, characterized in that: The device comprises: An acquisition unit, used to acquire text information of an object to be identified and extract text semantic features of the text information; A classification unit, configured to classify the object to be identified based on the text semantic features, and obtain at least one first target label and a corresponding first credibility of the object to be identified; A splicing unit, configured to respectively extract the label semantic features of the at least one first target label, and based on the obtained at least one first credibility, splice the corresponding label semantic features to obtain a comprehensive label feature; An obtaining unit is used to obtain at least one second target label of the object to be identified based on the comprehensive label feature and the text semantic feature; wherein the second target label is a sub-label of the first target label.
11. A computing device, characterized in that: The computing device comprises: a processor and a memory, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to implement the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
13. A computer program product, characterized in that The invention comprises a computer program, which implements the steps of any one of the methods of claims 1 to 9 when executed by a processor.