Label recognition method and device
By extracting multimodal features and category recognition of videos and combining with new hot tag collections, the problem of difficult to identify refined video tags and maintain label timeliness in the prior art is solved, and higher label recognition accuracy and timeliness are achieved.
Patent Information
- Application Number
- CN202110184917.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-02-10
AI Technical Summary
The prior art is difficult to identify refined and specific video tags, and the tags predicted by the model are difficult to maintain timeliness.
The video tag is determined by extracting the multimodal features of the video to be identified, combining category recognition and a new hot tag collection. The specific steps include: extracting multimodal features to determine the first tag, identifying the second tag according to the video category, and obtaining the third tag through the new hot tag collection updated in real time, and finally determining the video tag in combination with the three.
It improves the accuracy and timeliness of label identification, can be flexibly updated and migrated, improves reusability and reduces maintenance difficulty.
Smart Images

Figure CN113407778B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of big data technology, and in particular, to a method and device for label recognition. Background Art
[0002] Video labels play a very important role in services such as video recommendation and video search. Video labels can not only accurately depict the characteristics of a video, but also assist in depicting the interests and habits of users, and can provide a comprehensive and accurate basis for services such as video recommendation and video search.
[0003] In the related art, mainly a training model is used to predict the labels of a video. However, this solution usually can only achieve good recognition effects on general and abstract labels, and cannot recognize refined and specific labels, making it difficult to meet the actual needs of video labels.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present application is to provide a method, device, electronic device and computer-readable storage medium for label recognition, so as to improve the accuracy of label recognition and the timeliness of labels to a certain extent.
[0006] According to the first aspect of the present application, a method for label recognition is provided. The method includes: extracting multi-modal features of the video to be recognized, and using the multi-modal features to determine the first label of the video to be recognized; classifying and recognizing the video to be recognized to obtain the category of the video to be recognized, and recognizing the second label of the video to be recognized based on the category of the video to be recognized; obtaining the third label of the video to be recognized through a newly updated popular label set, and combining the first label, the second label and the third label to determine the video label of the video to be recognized.
[0007] In an exemplary embodiment of the present application, based on the foregoing embodiment, the recognizing the second label of the video to be recognized based on the category of the video to be recognized includes:
[0008] When the category of the video to be recognized is a film, television and variety category, recognizing the target object in the video to be recognized;
[0009] Determining the second label of the video to be recognized based on the target object.
[0010] In an exemplary embodiment of the present application, based on the foregoing embodiments, the target object includes human features, and determining the second label of the video to be recognized based on the target object includes:
[0011] Recognize the human features in the video to be recognized, and determine the human label of the video to be recognized;
[0012] Obtain the names of movies, TV shows, and variety shows related to the human label through the knowledge graph model;
[0013] Calculate the similarity between the video to be recognized and the video of the movie, TV show, or variety show corresponding to the name, and determine the movie, TV show, and variety show label of the video to be recognized;
[0014] Determine the second label according to the human label and the movie, TV show, and variety show label.
[0015] In an exemplary embodiment of the present application, based on the foregoing embodiments, recognizing the second label of the video to be recognized based on the category of the video to be recognized includes:
[0016] When the category of the video to be recognized is a game category, match the game template data with the video to be recognized, and determine the second label of the video to be recognized according to the matched game template data. In an exemplary embodiment of the present application, based on the foregoing embodiments,
[0017] The game template data includes a skill frame template, and matching the game template data with the video to be recognized and determining the second label of the video to be recognized according to the matching result includes:
[0018] Match the video to be recognized with multiple skill frame templates to determine the target skill frame that matches the video to be recognized;
[0019] Obtain the target game character associated with the target skill frame;
[0020] Determine the second label according to the target game character.
[0021] In an exemplary embodiment of the present application, based on the foregoing embodiments, obtaining the third label of the video to be recognized through the newly popular label set updated in real time includes:
[0022] Obtain the newly popular label set updated in real time, and obtain the video data corresponding to the newly popular label set;
[0023] Calculate the similarity between the title of the video to be recognized and the title of the video data, and screen out multiple target videos from the video data whose similarity meets the preset threshold;
[0024] Determine the third tag of the video to be recognized through the new hot tags corresponding to the multiple target videos.
[0025] In an exemplary embodiment of the present application, based on the foregoing embodiment, the determining the video tag of the video to be recognized by combining the first tag, the second tag, and the third tag includes:
[0026] According to the credibility strategies corresponding to the first tag and the second tag respectively, determine the tags that conform to the credibility strategies among the first tag and the second tag as the fourth tag;
[0027] Determine the video tag of the video to be recognized according to the fourth tag and the third tag.
[0028] In an exemplary embodiment of the present application, based on the foregoing embodiment, the determining the tags that conform to the credibility strategies among the first tag and the second tag as the fourth tag according to the credibility strategies corresponding to the first tag and the second tag respectively includes:
[0029] For each first tag, if the model that outputs the first tag is the confidence model corresponding to the first tag, then determine the first tag as a fourth tag; wherein, the confidence degree output by the confidence model corresponding to the first tag conforms to the credibility strategy corresponding to the first tag;
[0030] For each second tag, if the model that outputs the second tag is the confidence model corresponding to the second tag, then determine the second tag as a fourth tag; wherein, the confidence degree output by the confidence model corresponding to the second tag conforms to the credibility strategy corresponding to the second tag.
[0031] In an exemplary embodiment of the present application, based on the foregoing embodiment, the using the multimodal features to determine the first tag of the video to be recognized includes: extracting the multimodal features of the video to be recognized through a first model, and determining the first tag of the video to be recognized according to the multimodal features;
[0032] The recognizing the second tag of the video to be recognized based on the category of the video to be recognized includes: recognizing the video to be recognized through a second model corresponding to the category of the video to be recognized to obtain the second tag of the video to be recognized;
[0033] Before determining the tags that conform to the credibility strategies among the first tag and the second tag as the target tag according to the credibility strategies corresponding to the first tag and the second tag respectively, the method further includes:
[0034] Identify first labels of multiple video samples by using the first model as a first test result; and identify second labels of the multiple video samples by using the second model as a second test result;
[0035] Calculating a first confidence of the first model for each label included in the first test result, and calculating a second confidence of the second model for each label included in the second test result;
[0036] For target labels that overlap in the first test result and the second test result, determine the confidence model corresponding to each target label from the first model and the second model based on the first confidence level and the second confidence level of each target label, and use the correspondence between the target label and the confidence model as the confidence strategy corresponding to the target label.
[0037] According to a second aspect of the present application, a tag recognition device is provided, the device comprising: a multimodal feature extraction module, a category recognition module and a new hot tag recognition module.
[0038] Among them, the multimodal feature extraction module is used to extract the multimodal features of the video to be identified, and use the multimodal features to determine the first label of the video to be identified.
[0039] The category recognition module is used to classify and recognize the video to be recognized, obtain the category of the video to be recognized, and identify the second label of the video to be recognized based on the category of the video to be recognized.
[0040] The new hot tag identification module is used to obtain the third tag of the video to be identified through the real-time updated new hot tag set, and determine the video tag of the video to be identified by combining the first tag, the second tag and the third tag.
[0041] In an exemplary embodiment of the present application, based on the aforementioned embodiment, the category identification module includes: a film and television category identification module, which is used to identify the target object in the video to be identified when the category of the video to be identified is the film and television category; and a film and television label determination module, which is used to determine the second label of the video to be identified based on the target object.
[0042] In an exemplary embodiment of the present application, based on the aforementioned embodiment, the target object includes character features, and the film and television comprehensive label determination module of the present application may include a character recognition module, a knowledge graph module, a similarity calculation module and a label output module.
[0043] The character recognition module is used to recognize the features of the characters in the video to be recognized and determine the character tags of the video to be recognized.
[0044] A knowledge graph module for obtaining the names of film, television, and variety shows related to the character tags through a knowledge graph model.
[0045] A similarity calculation module for calculating the similarity between the video to be recognized and the video of the film, television, or variety show corresponding to the name of the film, television, or variety show, and determining the film, television, and variety show tags of the video to be recognized.
[0046] A tag output module for determining the second tag based on the character tags and the film, television, and variety show tags.
[0047] In an exemplary embodiment of the present application, based on the foregoing embodiment, the category recognition module is configured to: when the category of the video to be recognized is a game category, match the game template data with the video to be recognized, and determine the second tag of the video to be recognized according to the matched game template data.
[0048] In an exemplary embodiment of the present application, based on the foregoing embodiment, the game template data includes a skill frame template, and the category recognition module may include a skill frame matching module, a game character determination module, and a game tag determination module.
[0049] Among them, the skill frame matching module is used to match the video to be recognized with multiple skill frame templates to determine the target skill frame that matches the video to be recognized.
[0050] The game character determination module is used to obtain the target game character associated with the target skill frame.
[0051] The game tag determination module is used to determine the second tag according to the target game character.
[0052] In an exemplary embodiment of the present application, based on the foregoing embodiment, the new and popular tag recognition module may include a video data acquisition module, a video similarity calculation module, and a new and popular tag determination module.
[0053] Among them, the video data acquisition module is used to acquire a set of real-time updated new and popular tags and acquire the video data corresponding to the set of new and popular tags.
[0054] The video similarity calculation module is used to calculate the similarity between the title of the video to be recognized and the title of the video data, and screen out multiple target videos from the video data whose similarity meets a preset threshold.
[0055] The new and popular tag determination module is used to determine the third tag of the video to be recognized through the new and popular tags corresponding to the multiple target videos.
[0056] In an exemplary embodiment of the present application, based on the foregoing embodiment, the new hot tag recognition module may be configured to: determine, according to the credibility strategies corresponding to the first tag and the second tag respectively, the tags that conform to the credibility strategies among the first tag and the second tag as the fourth tag; and determine the video tag of the video to be recognized according to the fourth tag and the third tag.
[0057] In an exemplary embodiment of the present application, the new hot tag recognition module may be configured to: for each first tag, if the model that outputs the first tag is the confidence model corresponding to the first tag, then determine the first tag as a fourth tag; wherein, the confidence model corresponding to the first tag outputs a confidence level of the first tag that conforms to the credibility strategy corresponding to the first tag; for each second tag, if the model that outputs the second tag is the confidence model corresponding to the second tag, then determine the second tag as a fourth tag; wherein, the confidence model corresponding to the second tag outputs a confidence level of the second tag that conforms to the credibility strategy corresponding to the second tag.
[0058] In an exemplary embodiment of the present application, based on the foregoing embodiment, the determining the first tag of the video to be recognized by using the multi-modal features includes: extracting the multi-modal features of the video to be recognized through a first model, and determining the first tag of the video to be recognized according to the multi-modal features; the recognizing the second tag of the video to be recognized based on the category of the video to be recognized includes: recognizing the video to be recognized through a second model corresponding to the category of the video to be recognized to obtain the second tag of the video to be recognized; the apparatus further includes a sample acquisition module, a confidence level test module, and a credibility strategy determination module.
[0059] Among them, the first test result module is used to recognize the first tags of multiple video samples through the first model as the first test result.
[0060] The second test result module recognizes the second tags of the multiple video samples through the second model as the second test result.
[0061] The confidence level test module is used to calculate the first confidence level of each tag included in the first test result by the first model, and calculate the second confidence level of each tag included in the second test result by the second model.
[0062] A confidence strategy determination module is used to determine, for target labels that overlap in the first test result and the second test result, a confidence model corresponding to each target label from the first model and the second model according to the first confidence level and the second confidence level of each target label, and use the correspondence between the target label and the confidence model as the confidence strategy corresponding to the target label.
[0063] According to a third aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the tag identification method described in any embodiment of the first aspect is implemented.
[0064] According to the fourth aspect of the present application, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the tag identification method described in any embodiment of the above-mentioned first aspect by executing the executable instructions.
[0065] According to a fifth aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the tag recognition method provided in each of the above embodiments.
[0066] The exemplary embodiments of the present application may have some or all of the following beneficial effects:
[0067] In the tag identification scheme provided in an example embodiment of the present application, the multimodal features of the video to be identified are extracted, the first tag of the video to be identified is determined by using the multimodal features, and the video to be identified is classified and identified, and the second tag is further identified according to the category of the video to be identified, and the third tag is also identified by a new hot tag set, and finally the video tag of the video to be identified is finally determined by combining the first tag, the second tag and the third tag. It can be seen that in this technical scheme, on the one hand, the first tag of the video is determined as a whole by multimodal features, so as to ensure the universality of the tag; on the other hand, the video is classified and refined, and the second tag is determined in a targeted manner according to the category of the video, which can improve the accuracy of the tag; on the other hand, the real-time hot content is identified by a new hot tag set, which can improve the timeliness of the tag. In addition, the first tag, the second tag and the third tag have less data dependence on each other, can be flexibly updated and migrated, can improve reusability, and reduce the difficulty of maintenance.
[0068] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. Brief Description of the Drawings
[0069] The drawings herein are incorporated into and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0070] Figure 1 The schematic diagram of the system architecture showing the application scenario of a label recognition method to which the embodiments of the present application can be applied.
[0071] Figure 2 The schematic diagram of the structure of the computer system of the electronic device suitable for implementing the embodiments of the present application.
[0072] Figure 3 The schematic flow chart showing the label recognition method according to an embodiment of the present application.
[0073] Figure 4 The schematic flow chart showing the process of recognizing a second label according to an embodiment of the present application.
[0074] Figure 5 The schematic flow chart showing the process of recognizing a second label according to another embodiment of the present application.
[0075] Figure 6 The schematic flow chart showing the process of recognizing a third label according to an embodiment of the present application.
[0076] Figure 7 The schematic flow chart showing the process of recognizing a third label according to another embodiment of the present application.
[0077] Figure 8 The schematic flow chart showing the process of determining an acceptance strategy according to an embodiment of the present application.
[0078] Figure 9 The schematic diagram of the system architecture of the label recognition method according to an embodiment of the present application.
[0079] Figure 10 The schematic diagram of the label system showing the label recognition method according to an embodiment of the present application.
[0080] Figure 11 The schematic diagram of the video label display effect showing the label recognition method according to an embodiment of the present application.
[0081] Figure 12The structural schematic diagram of a tag recognition device according to an embodiment of the present application is shown. Detailed implementation manners
[0082] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will recognize that one or more of the specific details may be omitted, or other methods, components, devices, steps, etc. may be used. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of this application.
[0083] In addition, the accompanying drawings are only schematic illustrations of this application and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0084] In the related art of recognizing video tags, multi-modal features can be used to predict the tags of a video. The entire model is divided into two stages. First, single-modal features are extracted, and then multi-modal feature fusion prediction is performed. It has high flexibility and can meet the requirements of big data scenarios. For example, a pre-trained model is used to extract the word vectors of a video to obtain the text modality, a convolutional network is used to extract the image features of the video to obtain the image modality, and a convolutional network is used to extract the audio features to obtain the audio modality. Then, the multi-modal information is interactively fused and integrated into a unified representation vector to predict the tags of the video. However, this solution can only achieve results on conceptual tags and a small number of entity tags, and the recognition effect is difficult to meet the actual requirements for tags in expert fields that require refined understanding. Moreover, this solution requires a large amount of training data to be accumulated, and the actual scene has a very fast video update speed, resulting in the tags predicted by the model being difficult to maintain timeliness.
[0085] Based on this, this exemplary embodiment provides a tag recognition method that can overcome one or more of the above problems. Figure 1The figure shows a schematic diagram of a system architecture of an exemplary application environment to which a label recognition method according to an embodiment of the present application can be applied.
[0086] As Figure 1 shown, the system architecture 100 may include one or more of terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The terminal devices 101, 102, 103 may be various electronic devices with a display screen, including but not limited to desktop computers, portable computers, smartphones, and tablet computers, etc. It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0087] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. For example, the server 105 may be a server cluster composed of multiple servers, etc.
[0088] Figure 2 The figure shows a schematic diagram of a computer system of an electronic device suitable for implementing an embodiment of the present application.
[0089] It should be noted that Figure 2 the computer system 200 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0090] As Figure 2 shown, the computer system 200 includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 202 or the program loaded from the storage section 208 into the random access memory (RAM) 203. In the RAM 203, various programs and data required for system operations are also stored. The CPU 201, ROM 202, and RAM 203 are connected to each other through a bus 204. The input / output (I / O) interface 205 is also connected to the bus 204.
[0091] The following components are connected to the I / O interface 205: an input section 206 including a keyboard, a mouse, etc.; an output section 207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 208 including a hard disk, etc.; and a communication section 209 including a network interface card such as a LAN card, a modem, etc. The communication section 209 performs communication processing via a network such as the Internet. A drive 210 is also connected to the I / O interface 205 as needed. A removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 210 as needed so that a computer program read from it can be installed into the storage section 208 as needed.
[0092] Specifically, according to an embodiment of the present application, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 209, and / or installed from the removable medium 211. When the computer program is executed by a central processing unit (CPU) 201, various functions defined in the methods and apparatuses of the present application are executed. In some embodiments, the computer system 200 may further include an AI (Artificial Intelligence) processor for processing computational operations related to machine learning.
[0093] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have functions of perception, reasoning, and decision-making.
[0094] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0095] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement, and further performs image processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0096] The key technologies of Speech Technology include Automatic Speech Recognition (ASR), Text To Speech (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.
[0097] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph, etc. technologies.
[0098] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0099] The technical solutions of the embodiments of the present application will be elaborated in detail as follows:
[0100] Figure 3 A flowchart of a tag recognition method according to an embodiment of the present application is schematically shown. This tag recognition method can be applied to the above-mentioned server 105, or can be applied to one or more of the above-mentioned terminal devices 101, 102, and 103. No special limitation is made in this exemplary embodiment. Refer to Figure 3 As shown, this tag recognition method may include the following steps:
[0101] Step S310. Extract multi-modal features of the video to be recognized, and use the multi-modal features to determine the first tag of the video to be recognized.
[0102] Step S320. Classify and recognize the video to be recognized, obtain the category of the video to be recognized, and recognize the second tag of the video to be recognized based on the category of the video to be recognized.
[0103] Step S330. Obtain the third tag of the video to be recognized through a newly updated set of popular tags, and combine the first tag, the second tag, and the third tag to determine the video tag of the video to be recognized.
[0104] In the tag recognition method provided in this exemplary embodiment, on the one hand, the first tag of the video is determined as a whole through multi-modal features, thereby ensuring the generality of the tag; on the other hand, the video is classified and refined, and the second tag is determined specifically according to the category of the video, which can improve the accuracy of the tag; on the other hand, through the set of newly popular tags, real-time hot content is recognized, which can improve the timeliness of the tag. Moreover, there is less data dependence among the first tag, the second tag, and the third tag, and they can be flexibly updated and migrated, which can improve the reusability and reduce the maintenance difficulty at the same time.
[0105] Next, the above steps in this exemplary embodiment will be described in more detail.
[0106] In step S310, multi-modal features of the video to be recognized are extracted, and the first tag of the video to be recognized is determined using the multi-modal features.
[0107] Among them, modality refers to the form of information. Each source or form of information can be called a modality. For example, information can be in the form of speech, video, text, etc., and a person can have vision, hearing, smell, etc. Each of these can be called a modality. In this embodiment, the multi-modal features can include audio features, image features, and text features. To extract the multi-modal features of the video to be recognized, the image features, audio features, and text features of the video to be recognized can be extracted separately; then the image features, audio features, and text features are fused to obtain the multi-modal features.
[0108] Specifically, first, the features of each modality are extracted separately. The video to be recognized includes multiple frames of images, and a certain number of frames of images can be extracted from them for feature extraction. For example, one frame of image can be extracted per second for extraction, or one frame of image can be extracted every 2 seconds for extraction, etc. There are various ways to extract image features. For example, the image features can be extracted through EfficientNet, which is a convolutional neural network; the image features can also be extracted through other convolutional neural networks, such as VGGNet, ResNet, etc. This embodiment is not limited to this. The video to be recognized also includes audio. The audio is sampled, and then the audio features are extracted. For example, a 0.96-second audio segment is sampled, the Mel spectrogram is extracted, and then the audio features are obtained by using VGGish. VGGish is a VGG model based on TensorFlow, and this model can extract semantic embedding features from the audio spectrum. The video to be recognized usually also includes a title or name, and the text features in the title or name of the video to be recognized can be extracted by using the BERT model.
[0109] After the features of each modality are extracted, the features of each modality can be aggregated to obtain multi-modal features, that is, multiple single-modal features are fused into one multi-modal feature. Exemplarily, the image features, audio features, and text features can be added to obtain the multi-modal features, or the image features, audio features, and text features can be multiplied to obtain the multi-modal features, or the method of a linear function can be used for fusion, etc.
[0110] After the modal features are extracted, the multi-modal features can be classified to obtain the first label. Here, classification refers to classifying the field corresponding to the video content. Therefore, the first label can include labels for all fields. For example, the first label can include TV dramas, movies, games, food, variety shows, etc., and can also include types in other field dimensions, such as live broadcasts, news, etc. Correspondingly, the second label includes more refined labels for each field, including but not limited to TV drama names, actors in TV dramas, events corresponding to the video, etc. For example, the first label can be "TV drama", and the second label can be "My First Half of Life". Another example is that the first label can be "game", and the second label can be "Honor of Kings", etc. By pre-determining the label system, all labels can be divided into the first label and the second label. Both the first label and the second label can include multiple labels, and in some field dimensions, the labels included in the first label and the second label can overlap. For example, in the dimension of a person, the first label can be Person A, and the second label can also be Person A. However, the second label is more detailed than the first label in the dimension of a person. Therefore, in addition to including Person A, the second label can also include Persons B, C, etc.
[0111] For example, the video to be recognized can be input into the NeXtVLAD model, and the recognition result of the model can be used as the first label. Using NeXtVLAD can extract the features of video frames and compress them into feature vectors to achieve video classification, and can better fuse multi-modal information. At the same time, text information such as the title or name of the video to be recognized is input into the TextCNN and BI-LSTM models. Using TextCNN or BI-LSTM can generate the representation vectors of the text information to achieve text classification, and the obtained classification result can be used as the first label. For example, the first label can be: TV drama, Mainland drama, TV drama teaser, etc. Among them, the NeXtVLAD model introduces an attention mechanism to aggregate the features of each video frame, and can be applied to video classification methods with any number of frames, and can better aggregate image features and audio features; TextCNN is a text classification model based on a convolutional neural network, which can encode text information to achieve text classification; BI-LSTM is a bidirectional LSTM model, which can consider the context order of the text and has a more accurate classification effect. To improve the classification effect of multi-modal features, the NeXtVLAD model can be used to recognize the video to be recognized to obtain the recognition label of the video to be recognized. Using the TextCNN or BI-LSTM model can also obtain the recognition label of the video to be recognized. The union of the above two types of recognition labels can be used as the first label to ensure the generality of the first label.
[0112] In step S320, the video to be recognized is classified and recognized to obtain the category of the video to be recognized, and the second label of the video to be recognized is recognized based on the category of the video to be recognized.
[0113] In this exemplary embodiment, the classification and recognition of the video to be recognized can be performed by using a machine learning model. For example, the above NeXtVLAD algorithm is used to train a classification model, so as to classify the video to be recognized by using multi-modal features and obtain its category. Exemplarily, the categories of the video to be recognized may include film, television, variety shows, games, food, and news. In advance, corresponding video sample data can be obtained and marked for each category, and the convolutional neural network can be trained by using the marked video sample data to obtain a video classification model. Then, the trained video classification model can be used to classify and recognize the video to be recognized, and the category of the video to be recognized is output. In addition, the video can be classified in a more fine-grained manner according to actual needs. For example, the categories may include TV dramas, movies, entertainment news, current affairs news, domestic cuisine, foreign food, etc. This embodiment is not limited thereto.
[0114] For each category of the video, a corresponding label prediction model can be trained. After determining the category of the video to be recognized, the video to be recognized is input into the label prediction model corresponding to the category for more fine-grained classification to determine the second label of the video to be recognized. In an exemplary embodiment, when the category of the video to be recognized is film, television, and variety shows, the target object in the recognized video is recognized; and then the second label of the video to be recognized is determined based on the target object. Among them, the target object may refer to the object included in the video, including but not limited to people, animals, objects, scenes, etc. Exemplarily, the target object may be the scene in the video, and the corresponding label prediction model can be used to recognize the scene in the video to obtain the location feature corresponding to the video, and then the location feature is used as the second label of the video. Exemplarily, the target object may also be a person in the video, then the person in the video is recognized to determine the person feature, and the obtained person feature is used as the second label. In other embodiments, the target object may also be a vehicle, a room, a road, etc. in the video, and then the features such as the vehicle, road, and room in the video can be recognized by the model as the second label, etc., which also belongs to the protection scope of this application.
[0115] In an exemplary embodiment, when the target object is a person feature, the method for recognizing the second label of the video to be recognized may include the following steps S410, step S420, step S430, and step S440, as Figure 4 shown.
[0116] In step S410, the human features in the video to be recognized are recognized to determine the human label of the video to be recognized. Among them, the human features may include face features, body features, or features of some body parts, etc. In this exemplary embodiment, the face is taken as an example to illustrate the recognition process. Specifically, the video to be recognized is input into a face recognition model. Through this face recognition model, multiple image frames can be sampled from the video, and face detection and face alignment are performed on the image frames. Then, face embedding features are extracted and retrieved in the face database, and the face information with the highest confidence is output. This face information is used as the human label. For example, if the confidence of retrieving actor A is 0.7 and the confidence of retrieving actor B is 0.9, then actor B is output as the retrieval result, and "actor B" is the human label of the video to be recognized.
[0117] In step S420, the names of film and television variety shows related to the human label are obtained through the knowledge graph model. Among them, the names of film and television variety shows may include movie names, TV drama names, variety show names, or any combination of two or three of them.
[0118] A knowledge graph is a particularly large semantic network system. The main purpose is to describe the association relationships between entities or concepts in the real world. Through a large amount of data collection, it is organized into a knowledge base that can be processed by machines to achieve visual display.
[0119] The knowledge graph model in this embodiment can determine the relationship between a person and a film and television drama by pre-collecting relevant data of film and television dramas, such as video data of film and television dramas, information of cast and crew, plot descriptions, etc., and construct a data model that can be recognized and processed by a computer. By querying the knowledge graph model for the names of film and television variety shows associated with the human label, for example, if the human label is "actor A", then all the names of the film and television dramas in which "actor A" participated can be output through this knowledge graph model. It can be understood that one human label can be associated with multiple names of film and television variety shows, that is, one person can star in multiple film and television dramas; multiple human labels can also be associated with the same name of a film and television variety show, that is, multiple actors can star in the same film and television drama. If multiple human labels are recognized in step S410, then the names of film and television variety shows associated with all these multiple human labels can be obtained.
[0120] In step S430, calculate the similarity between the video to be recognized and the video of the film, television, or variety show corresponding to the name of the film, television, or variety show to determine the film, television, or variety show label of the video to be recognized. In this exemplary embodiment, the name of the film, television, or variety show related to the person label may include one or more. Obtain the videos corresponding to each name of the film, television, or variety show respectively, calculate the similarity between the video to be recognized and the video of the film, television, or variety show, and use the name of the film, television, or variety show with the highest similarity as the film, television, or variety show label of the video to be recognized. Specifically, for each video of the film, television, or variety show, multiple frames of pictures can be extracted in advance to construct a picture library. For example, extract pictures of key plot segments in the video of the film, television, or variety show, extract manually marked pictures, etc.; then extract the key frames from the video to be recognized, calculate the similarity between each key frame and each picture in the picture library respectively, and sum up the obtained similarities as the similarity calculation result between the picture library and the video to be recognized; finally, select the picture library with the highest similarity to the video to be recognized according to the similarity calculation result, and use the film corresponding to this picture library as the final film, television, or variety show label. For example, assume that there are film A, film B, and film C related to the person label. Calculate the similarity between the video of film A and the video to be recognized, the similarity between film B and the video to be recognized, and the similarity between film C and the video to be recognized respectively, and use the film with the highest similarity as the final film, television, or variety show label.
[0121] In step S440, the second label is determined according to the character label and the film and television variety label. The character label and the film and television variety label can be combined as the second label. For example, if the character label includes "Actor A, Actor B" and the film and television variety label includes "TV drama A", then the second label can include "Actor A, Actor B, TV drama A". Moreover, according to the film and television variety label, other information related to the film and television variety label can be obtained through the above knowledge graph model, and then the obtained information, together with the character label and the film and television variety label, is added to the second label to more comprehensively determine the label of the video to be recognized. For example, information about all actors, directors, etc. related to the film and television variety label, or information with the strongest association with the film and television variety label, such as the leading actors of a TV drama. In the knowledge graph model, different weights can be assigned to the association relationships between different entities, and the weights are used to represent the strength of the association relationships. If the association relationship between two entities is stronger, the weight is higher. Therefore, after obtaining the film and television variety information most similar to the video to be recognized, the entity with the highest association relationship weight with the film and television variety information can be searched from the knowledge graph model and output as the second label, or a certain number of entities related to the film and television variety can be output as the second label according to the weights of the association relationships, etc. For example, if the character label includes "Actor A, Actor B" and the film and television variety label includes "TV drama A", according to the knowledge graph model, the entity with the highest association relationship weight with "TV drama A" is "Actor C", and their association relationship can be "lead actor", then "Actor A, Actor B, Actor C, TV drama A" can be used as the second label of the video to be recognized.
[0122] In an exemplary embodiment, the category of the video to be recognized can also be a game category. When the category of the video to be recognized is a game category, the game template data can be matched with the video to be recognized, and the second label of the video to be recognized can be determined according to the matched game template data. Among them, the game template data can include various template information, such as game character templates, game scene templates, models of objects in the game, skill frame templates, etc., and can also include other templates, such as game map templates, etc. This embodiment does not make special limitations on this. In advance, the game characters, skill frames, and game scenes included in each game can be collected as templates to construct game template data. By matching each template included in the game template data with the video to be recognized one by one, the template that matches the video to be recognized can be determined, and thus the information of the template can be used as the second label. For example, if the video to be recognized matches the template of game scene A, then the second label can be game scene A.
[0123] In an exemplary embodiment, when the game template data includes a skill frame template, the method for determining the second label can include the following steps S510, step S520, and step S530, as Figure 5 shown.
[0124] In step S510, the video to be recognized is matched with multiple skill box templates to determine a target skill box that matches the video to be recognized. In a game application, a skill box refers to the area where the control for triggering a skill is located in the game interface. Different game characters have different skills. In advance, the skill boxes of the game characters in each game application can be collected and saved as skill box templates, and the number of skill box templates can be multiple. Then, a certain number of frame pictures are extracted from the video to be recognized, and the extracted pictures are matched with multiple skill box templates to determine the skill box template that matches the picture as the target skill box. In one example, there are 100 skill box templates. The 100 skill box templates are matched one by one with the pictures included in the video to be recognized. If the one that matches the pictures included in the video to be recognized is "Skill Box Template A", then Skill Box Template A is the target skill box.
[0125] In step S520, the target game character associated with the target skill box is obtained. A skill box can be associated with a game character. The association relationship between the skill box and the game character can also be constructed through a knowledge graph or through a relational database key-value. After determining the target skill box, the knowledge graph or the corresponding database can be queried to determine the target game character associated with the target skill box.
[0126] In step S530, the second label is determined according to the target game character. The target game character can be used as the second label of the video to be recognized. For example, if Skill Box A is associated with Game Character B, then the second label of the video to be recognized can be "Game Character B". Moreover, after determining the target game character corresponding to the video to be recognized, other relevant information such as the game name and game type can also be retrieved according to the target game character, and the retrieved information can also be used as the second label.
[0127] In this exemplary embodiment, the categories of the video to be recognized may also include various other categories, such as food, news, etc. For each category, a label prediction model can be pre-trained to be responsible for predicting the fine-grained second label under that category, making the label more accurate. For example, for the food category, a large number of food pictures of different cuisines can be collected to train a classification model. Then, the classification model is used to classify the video to be recognized in the food category, and the cuisine corresponding to the video to be recognized is output, and this output cuisine is used as the second label of the video to be recognized. Moreover, for each cuisine, a more fine-grained model can be trained separately to recognize more specific dish names or food names. For example, a classification model is trained for the dessert category, and through this classification model, it is recognized which kind of dessert is in the video. For example, the picture frames in the video to be recognized are input into the classification model corresponding to the dessert category, and this classification model can output results such as "cake" or "egg tart", and this output result is used as the second label, so as to improve the accuracy of the label.
[0128] Continue to refer to Figure 3 , in step S330, the third label of the video to be recognized is obtained through the newly popular label set updated in real time, and the video label of the video to be recognized is determined by combining the first label, the second label, and the third label.
[0129] Among them, the newly popular label set includes multiple newly popular labels. The newly popular labels can include labels related to current hot events. In an actual scenario, the content of short videos updates very quickly and is easily a focus and hot spot of people's attention. And this kind of hot content usually lasts for a short time and may soon fade, with strong timeliness. Therefore, the newly popular labels can be updated in real time, such as once a day, once every 8 hours, etc. The update frequency can be determined according to actual needs, and this embodiment does not limit this. For example, through manual operation and marking, the newly popular labels found on the same day can be marked and saved into the newly popular label set to realize the real-time update of the newly popular label set. People can predict hot and explosive events, so as to speed up the discovery of newly popular labels and make the newly popular labels maintain strong timeliness.
[0130] After obtaining the updated newly popular label set, the video to be recognized can be matched with the newly popular labels in the newly popular label set. If the title information of the video to be recognized can be accurately matched with a newly popular label, then this newly popular label is used as the third label of the video to be recognized. For example, the newly popular label set includes four hot events A, B, C, and D. The title of the video to be recognized is matched with these four hot events respectively. If A matches the field in the title of the video to be recognized, then A is used as the third label of the video to be recognized.
[0131] Since the new hot tags have short timeliness, and model prediction requires long-term data accumulation with slow response speed, the new hot tag set in this exemplary embodiment can meet the actual need for rapid response of tags, improving the timeliness and accuracy of the tags.
[0132] In an exemplary embodiment, the method for obtaining the third tag of the video to be recognized may include the following steps S610, step S620, and step S630, as Figure 6 shown.
[0133] In step S610, obtain a newly updated set of hot tags and obtain the video data corresponding to the set of hot tags. The set of hot tags can be updated once a day, once every 8 hours, once every 4 hours, etc. Each time it is updated, the newly discovered hot tags can be manually updated into the set of hot tags, and the set of hot tags can be saved in a specific directory. Then, according to the pre-determined update time, the updated set of hot tags can be obtained from the predetermined directory. According to the updated set of hot tags, the video data with the hot tags can be pulled from the content library. The same video data can carry two or more hot tags.
[0134] In step S620, calculate the similarity between the title of the video to be recognized and the title of the video data, and filter out multiple target videos from the video data whose similarity meets the preset threshold. The set of hot tags can include multiple hot tags, and each hot tag can correspond to multiple video data. For example, calculate the similarity between the title of video data N and the title of the video to be recognized. If the similarity meets the preset threshold, then this video data N is used as the target video. By analogy, calculate the similarity between the title of each video data and the title of the video to be recognized to obtain all the target videos whose similarity meets the preset threshold. Among them, the preset threshold can be determined according to actual needs, such as 0.6, 0.7, etc., or it can be 0.65, 0.8, etc. This embodiment does not make special limitations on this.
[0135] In step S630, the third label of the video to be recognized is determined through the new hot labels corresponding to the multiple target videos. Specifically, according to the number of times each new hot label is matched in the target video, a vote can be conducted on the new hot labels, and the new hot label with the highest score is used as the third label of the video to be recognized. Each time a new hot label is matched, it can get one point. For example, the new hot labels corresponding to target video a are A and B, the new hot labels corresponding to target video b are B and C, the new hot labels corresponding to target video c are C, and the new hot labels corresponding to target video d are D and B. Then the score of A is 1, the score of B is 3, the score of C is 2, and the score of D is 1. The one with the highest score is B, so the new hot label B can be used as the third label of the video to be recognized. In addition, one or more can be selected from the new hot labels corresponding to the target video as the third label. For example, the new hot label B with the highest score and the new hot label C with the second highest score are used as the third label, and so on.
[0136] In this exemplary embodiment, a new hot label prediction model can also be trained according to the new hot label set and its corresponding video data, and the new hot labels of the video to be recognized are predicted through this model. Or the new hot label set is determined in combination with this new hot label prediction model. Specifically, verification data is obtained, and the verification data is respectively recognized by using the above new hot label set and the new hot label prediction model to determine the new hot labels of the verification data, and the precision-recall rate of the two methods for the verification data is calculated. If the precision-recall rate of the new hot label set for the verification data is higher, the current new hot label set is warehoused and used to recognize the video to be recognized; if the precision-recall rate of the new hot label prediction model is higher, the new hot labels adopted by this model are merged into the new hot label set, and the merged new hot label set is warehoused.
[0137] Figure 7 Schematically shows a flowchart of a method for obtaining the third label of a video to be recognized through a new hot label set. As Figure 7As shown, in step S710, the new hot tags of the operation tags are stored in the database to obtain an updated set of new hot tags. In step S720, a new hot pool is constructed. The new hot pool may include video data with new hot tags; after the operation tags the videos related to the new hot tags and uploads them online for everyone to view, the video data with new hot tags is pulled from the online database according to the updated set of new hot tags. In step S730, the titles included in the new hot pool are matched with the title of the video to be recognized. If the similarity between a title in the new hot pool and the title of the video to be recognized meets the threshold, the video data corresponding to this title can be output as the matching result. After matching each title in the new hot pool one by one, all target videos with similarity meeting the threshold can be obtained. For example, if the new hot pool contains 100 videos in total, and the titles of these 100 videos are respectively matched with the title of the video to be recognized, and there are 10 videos whose title similarity meets the threshold, then these 10 videos with title similarity meeting the threshold are output as the matching result, or the new hot tags corresponding to these 10 videos are output as the matching result. In step S740, the title of the video to be recognized is matched with the new hot tags. The matching result is the new hot tags whose similarity with the title of the video to be recognized meets the threshold. In step S750, the matching result is obtained, and this matching result includes the new hot tags of the target videos matched in step S730 and the new hot tags matched in step S740. In step S760, the scores of each new hot tag are calculated according to the matching result. Exemplarily, the scores of each new hot tag can be the number of times each new hot tag is matched. For example, if new hot tag A is matched 4 times, then its score is 4. In step S770, the prediction result of the new hot tags of the video to be recognized is determined according to the scores. Exemplarily, the new hot tag with the highest score can be used as the prediction result of the video to be recognized, or a certain number of new hot tags can be used as the prediction result according to the scores, etc. The prediction result of the new hot tags in this step is the third tag of the video to be recognized.
[0138] In step S330, after obtaining the third label of the video to be identified, the final video label of the video to be identified can be determined by combining the first label, the second label and the third label. Exemplarily, the first label, the second label and the third label can be subjected to a union operation, and their union can be used as the video label of the video to be identified, and then the video to be identified can be marked. Taking the first label, the second label and the third label as the final video label of the video to be identified can maximize the recall rate of the video label. After determining the video label of the video to be identified, other videos with similar labels can be recommended to users in video recommendation scenarios and video search scenarios based on the video label of the video to be identified. In an exemplary embodiment, according to the acceptance strategies corresponding to the first label and the second label, the labels in the first label and the second label that meet the acceptance strategy can be determined as the fourth label; and then the video label of the video to be identified can be determined based on the fourth label and the third label. Among them, the acceptance strategy can refer to the conditions for whether the label is accepted, for example, the acceptance strategy is to accept labels with recognition probabilities exceeding a threshold. Both the first label and the second label can be multiple labels, and the acceptance strategies corresponding to each label can be different. For example, the acceptance strategy of label A is that the recognition probability exceeds 80%. If the probability of label A being recognized does not exceed 80%, it can be determined that label A does not meet the acceptance strategy, and label A is deleted from the final result.
[0139] Before identifying the video to be identified, the credibility strategy for each tag can be determined in advance. Specifically, first, multiple tag samples can be obtained. In an implementation manner of the present application, the first tag and the second tag can be identified through a model. That is, the above steps S310 and S320 are executed through the model to output the first tag and the second tag. It can be understood that the model responsible for identifying the first tag and the model responsible for identifying the second tag can be different models. For example, the multi-modal features of the video to be identified are extracted through the first model, and the first tag of the video to be identified is determined by using the multi-modal features. The second tag of the video to be identified is identified through the second model. Among them, the first model can be used to extract the multi-modal features of the video, so as to classify the video by using the multi-modal features. For example, the first model can be the NeXtVLAD model, etc. The second model can be determined according to the category of the video to be identified. For example, if the category of the video to be identified is the category of film, television and variety shows, any one or more of the models related to film, television and variety shows, such as the face recognition model, the knowledge graph association model of film and television dramas, and the title retrieval model of film, television and variety shows, can be used as the second model to perform a more fine-grained understanding of the video to be identified. Another example is that if the category of the video to be identified is the game category, any one or more of the game category models, such as the game character name tag model for identifying game character names and the game name tag model for identifying game names, can be used as the second model. Thus, a customized tag model (such as the aforementioned models related to film, television and variety shows, game category models, etc.) can be fused on the basis of the general video tag model to improve the recognition ability on some entity tags. After the first model and the second model are trained, the confidence levels of the first model and the second model can be determined, and then the credibility strategy can be determined according to the confidence levels. Specifically, for each first tag, if the model that outputs the first tag is the confidence model corresponding to the first tag, the first tag is determined as a fourth tag. Among them, the confidence level output by the confidence model corresponding to the first tag conforms to the credibility strategy corresponding to the first tag. For each second tag, if the model that outputs the second tag is the confidence model corresponding to the second tag, the second tag is determined as a fourth tag. Among them, the confidence level output by the confidence model corresponding to the second tag conforms to the credibility strategy corresponding to the second tag.
[0140] Determining the fourth label from the first label and the second label depends on whether the model for the output label (including the first label and the second label) is the confidence model corresponding to the label, and the confidence model corresponding to each label can be determined in advance. For the first label, its confidence model is a model whose confidence level for outputting the first label conforms to the adoption strategy. For example, if the first model identifies a video as the first label A, and if according to the adoption strategy, the confidence model for the label A is the first model, then the label A can be determined as a fourth label; on the contrary, if the confidence model for the label A in the adoption strategy is the second model, then the model outputting the label A is not its corresponding confidence model, and the label A cannot be used as the fourth label. The same applies to the second label. Among them, the confidence level is the accuracy rate of the model's identification of the label. After training the first model and the second model, the accuracy rate of their identification of each label can be calculated by testing the first model and the second model.
[0141] Before determining the fourth label among the first label and the second label of the video to be recognized, it is necessary to first determine the confidence model corresponding to the label, that is, determine the adoption strategy. In an exemplary embodiment, the specific process of determining the adoption strategy may include steps S810 to S840, as Figure 8 shown, where:
[0142] Step S810. Identify the first label of multiple video samples through the first model as the first test result.
[0143] Step S820. Identify the second label of the multiple video samples through the second model as the second test result.
[0144] Step S830. Calculate the first confidence level of the first model for each label included in the first test result, and calculate the second confidence level of the second model for each label included in the second test result.
[0145] Step S840. For the target labels that overlap in the first test result and the second test result, determine the confidence model corresponding to each target label from the first model and the second model according to the first confidence level and the second confidence level of each target label, and use the corresponding relationship between the target label and the confidence model as the adoption strategy corresponding to the target label.
[0146] In step S810, the video sample can be understood as a video for which the video label has been determined. Exemplarily, a certain number of videos can be obtained in advance, and the video label of each video can be determined manually, and the corresponding video label can be marked on the video to obtain the video sample. Inputting the video sample into the first model can obtain the first test result of the first model.
[0147] In step S820, the video sample is input into the second model to obtain a second test result. The first test result may include multiple first labels, and the second test result may include multiple second labels.
[0148] In step S830, it can be verified whether the first label is correctly classified based on the video label of the video sample, so as to calculate the first confidence of the first model for each label in the test result. Confidence refers to the accuracy of model recognition, and the higher the accuracy, the higher the confidence. For example, there are 10 samples in the video sample with the video label A. If the first model recognizes the label A of these 10 samples, the confidence of the first model for label A is 1. If the first model recognizes the label A of 9 samples among the 10 samples, the confidence for label A is 0.9. Similarly, the confidence of the second model for each label in the second test result is calculated as the second confidence.
[0149] In step S840, the intersection operation is first performed on the first test result and the second test result to determine the target label that overlaps in the first test result and the second test result. Then, for the target label, the first confidence of the first model is compared with the second confidence of the second model, and the higher of the first confidence and the second confidence is used to determine the confidence model of the target label. For example, the target label may include multiple, such as A, B, C, and D. The confidences of the first model for the four target labels are: 0.75, 0.7, 0.7, and 0.8, respectively, and the confidences of the second model for the target label are: 0.8, 0.75, 0.6, and 0.75, respectively. Then, it can be determined that the confidence model for label A is the second model, the confidence model for label B is the second model, the confidence model for label C is the first model, and the confidence model for label D is the first model. After determining the corresponding confidence model for each target label, the corresponding relationship between the target label and the corresponding confidence model is saved as a confidence adoption strategy. For labels with the same first confidence and second confidence, any one of the first model and the second model can be selected as the confidence model, for example, the first model can be selected as the confidence model.
[0150] After obtaining the first label and the second label of the video to be identified, the saved acceptance strategy can be queried to determine the confidence model corresponding to each first label and the second label. If the model output of the first label is different from the confidence model, it can be determined that the first label does not comply with the acceptance strategy. If the model output of the first label is the same as the confidence model, it can be determined that the first label complies with the acceptance strategy. For example, the first label of the video to be identified is: A, B, C, and the second label is A, C, E, F, G. The confidence models corresponding to A, B, C, E, F, and G can be determined through the acceptance strategy. For example, if the confidence model of label A is the second model, it can be determined that label A does not comply with the acceptance strategy. For another example, if the confidence model of label C is the second model, it can be determined that C complies with the acceptance strategy.
[0151] The labels included in the first test result and the second test result may not cover all labels in the label system. For example, if the label system is designed with 1,000 labels, the video sample may include 100 labels or 50 of them. For labels not covered by the first test result and the second test result, the probability threshold may be used as a trust strategy. For example, a probability threshold is determined, and labels with recognition probabilities exceeding the probability threshold are determined as labels that meet the trust strategy, and labels that do not exceed the probability threshold are determined as labels that do not meet the trust strategy.
[0152] After taking the first label of the video to be identified and the label in the second label that meets the adoption strategy as the fourth label, the union of the fourth label and the third label can be output as the final video label of the video to be identified; the third label can also be screened, and the third label and the fourth label that meet the conditions are taken as the final video label; for example, a threshold is set, and the third label and the fourth label that meet the threshold are output as the video label of the video to be identified, etc.
[0153] In the above method, the method of selecting the union of the first label, the second label and the third label can maximize the recall rate of the label. In this embodiment, the first label and the second label are screened, and the labels with higher confidence are filtered out as the video labels of the video to be identified through the adoption strategy, which can avoid the labels with lower confidence from affecting the accuracy of label identification, thereby improving the accuracy of video labels while ensuring the recall rate as much as possible.
[0154] Figure 9 A system architecture diagram of the tag identification method of this example embodiment is schematically shown. Figure 9As shown, the system architecture 900 may include a first label prediction model 901, a second label prediction model 902, a third label prediction model 903, and a fusion module 904. The video to be recognized can be input into models 901, 902, and 903 simultaneously. The first label prediction model 901 extracts multi-modal features of the video to be recognized and predicts the first label of the video to be recognized using the multi-modal features; the second label prediction model 902 predicts the second label of the video to be recognized; the third label prediction model 903 can predict the third label of the video to be recognized; and then the fusion module 904 fuses the results predicted by models 901, 902, and 903 to obtain the final video label of the video to be recognized.
[0155] Exemplarily, the second label prediction model 902 may include a classification model 9021 and multiple vertical models. Exemplarily, the vertical models may be a movie and TV drama recognition model 9022, a game recognition model 9023, a news label recognition model 9024, and a food label recognition model 9025. Among them, the classification model 9021 can perform classification recognition on the video to be recognized, predict the category of the video to be recognized, and perform gating control according to the recognized category, and distribute the video to be recognized to the corresponding vertical model. For example, if the category is "game category", the video to be recognized is sent to the game recognition model 9025, and if the category is "movie, TV drama, and variety show category", the video to be recognized is sent to the movie and TV drama recognition model 9022, etc.
[0156] The vertical model can perform label recognition for a specific category and has a more fine-grained understanding. Multiple vertical models can be designed according to the label classification system. Exemplarily, the label system can be as Figure 10 shown. According to the Figure 10 shown label system, four types of vertical models can be designed, such as a movie and TV drama recognition model, a game recognition model, a news label recognition model, and a food label recognition model.
[0157] For example, the film and television drama recognition model 9022 can be responsible for performing face recognition on the video to be recognized, determining that there are characters in it, predicting the character tags of the video to be recognized, and obtaining the names of related film and television dramas in combination with the knowledge graph. Then, by retrieving the video of the related film and television drama in the film and television drama database and comparing it with the video to be recognized, the corresponding film and television variety label of the video to be recognized is determined. Moreover, using this film and television variety label, other relevant information can be obtained again through the knowledge graph, such as the lead actor, other actors besides the above-recognized characters, etc. The relevant information, together with the film and television variety label, character tags, etc., is output as the second label to the fusion module 904. The game recognition model 9023 can use skill box template matching to recognize information such as the game name, game character, game skills, etc. of the video to be recognized, and output the recognized results as the second label to the fusion module 904. The news label recognition model 9024 can be trained based on pre-collected news events, so as to use this model to recognize the news events of the video to be recognized and output them as the second label to the fusion module. The food label recognition model 9025 can be trained by collecting pictures of various foods; using this model, the foods included in the video to be recognized can be recognized, and then the food names are output as the second label to the fusion module.
[0158] The third label prediction model 903 can be used to predict the new hot tags of the video to be recognized through the newly updated hot tag set in real time, and output them as the third label to the fusion module 904.
[0159] The fusion module 904 can calculate the union of the first label, second label, and third label output by all models, and output the union as the video label. Figure 11 Schematically shows the display effect of the video label. As Figure 11 shown, the recognized video labels of the video to be recognized A can include "drama name, type, region, highlights, person name", etc., and the recognition confidence can also be displayed in the recognition result. For example, the drama name label 1101 can indicate that the confidence in recognizing the drama name as "A" is "1.0", and the type label 1102 can indicate that the confidence in recognizing the type as "drama trailer" is "1.0", etc.
[0160] Exemplarily, when fusing the output results of the above-mentioned first label prediction model 901 and the second label prediction model 902, the precision-recall rates of the models 901 and 902, as well as the precision-recall rate corresponding to the merged output results of the models 901 and 902, can be verified first, and the best fusion method can be selected by comparing these three precision-recall rates for fusion. Specifically, after obtaining the models 901 and 902, a validation dataset D can be obtained, and the samples in the validation dataset can be recognized by using the models 901 and 902 respectively to obtain recognition results E and T; and the label U after merging the recognition results of these two models, that is, the first label in the recognition result E of the model 901 and the second label in the recognition result F of the model 902 are included in U. Then, for the label X, calculate the precision-recall rate e of the model 901, the precision-recall rate f of the model 902, and the precision-recall rate u of U obtained by merging the two on the validation dataset D, compare the magnitudes of these three types of precision-recall rates. If e is the largest, the first label output by the model 901 and the third label recognized by the model 903 can be merged in the fusion module 904 as the final video label.
[0161] Moreover, each module in this embodiment can be split. For example, the first label prediction model 901 and the second label prediction model 902 mentioned above can be combined to predict the labels of the video to obtain the final video label. For another example, the first label prediction model 901 and the third label prediction model 903 mentioned above can be combined to recognize the video labels; for another example, the third label prediction model 903 can be combined with other label prediction models to jointly predict the video labels and so on. It can be seen that the label recognition method provided in this embodiment has strong reusability and flexibility.
[0162] To further verify the effectiveness of the present application, the inventor compared the recognition effect of the label recognition method implemented according to the above system architecture 900 with that of the video label models in other methods. Specifically, it was found through experiments that the accuracy rate of other video label models was 82.2% and the recall rate was 65.3%. When using the first label recognition model 901 and the second label recognition model 902 in the above system architecture of this embodiment to recognize video labels, the accuracy rate increased by 0.7% to reach 82.9%, and the recall rate increased by 2.4% to reach 67.7%; when using the above system architecture 900 as a whole to recognize video labels, the accuracy rate increased by 1% to reach 82.2%; the recall rate increased by 20% to reach 85.3%, and the response speed to new and popular labels was greatly improved, ensuring the timeliness of the labels.
[0163] From the comparison results, using the label recognition method of the present application to recognize video labels can make the recognized video labels more accurate and have higher accuracy.
[0164] Those skilled in the art can understand that all or part of the steps for implementing the above embodiments are realized as a computer program executed by a processor (including a CPU and a GPU). When the computer program is executed by the processor, the above functions defined by the above method provided in this application are executed. The program can be stored in a computer-readable storage medium, which can be a read-only memory, a magnetic disk, an optical disc, etc.
[0165] In addition, it should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0166] The following introduces the tag recognition device provided by this technical solution:
[0167] A tag recognition device provided in this exemplary embodiment. Refer to Figure 12 As shown, the tag recognition device 1200 includes: a multi-modal feature extraction module 1201, a category recognition module 1202, and a new and hot tag recognition module 1203.
[0168] Among them, the multi-modal feature extraction module 1201 is used to extract the multi-modal features of the video to be recognized, and determine the first tag of the video to be recognized by using the multi-modal features.
[0169] The category recognition module 1202 is used to classify and recognize the video to be recognized to obtain the category of the video to be recognized, and recognize the second tag of the video to be recognized based on the category of the video to be recognized.
[0170] The new and hot tag recognition module 1203 is used to obtain the third tag of the video to be recognized through a real-time updated new and hot tag set, and determine the video tag of the video to be recognized by combining the first tag, the second tag, and the third tag.
[0171] In an exemplary embodiment of the present application, based on the foregoing embodiment, the category recognition module 1202 includes: a film, television, and variety show category recognition module, which is used to recognize the target object in the video to be recognized when the category of the video to be recognized is a film, television, and variety show category; and a film, television, and variety show tag determination module, which is used to determine the second tag of the video to be recognized based on the target object.
[0172] In an exemplary embodiment of the present application, based on the foregoing embodiment, the target object includes human features, and the film, television, and variety show tag determination module may include a human recognition module, a knowledge graph module, a similarity calculation module, and a tag output module.
[0173] Among them, the person recognition module is used to recognize the person features in the video to be recognized and determine the person label of the video to be recognized.
[0174] The knowledge graph module is used to obtain the names of movies, TV shows, and variety shows related to the person label through the knowledge graph model.
[0175] The similarity calculation module is used to calculate the similarity between the video to be recognized and the video of the movie, TV show, or variety show corresponding to the name, and determine the movie, TV show, and variety show label of the video to be recognized.
[0176] The label output module is used to determine the second label according to the person label and the movie, TV show, and variety show label.
[0177] In an exemplary embodiment of the present application, based on the foregoing embodiment, the category recognition module 1202 is configured to: when the category of the video to be recognized is a game category, match the game template data with the video to be recognized, and determine the second label of the video to be recognized according to the matched game template data.
[0178] In an exemplary embodiment of the present application, based on the foregoing embodiment, the game template data includes a skill frame template, and the category recognition module may include a skill frame matching module, a game character determination module, and a game label determination module.
[0179] Among them, the skill frame matching module is used to match the video to be recognized with multiple skill frame templates to determine the target skill frame that matches the video to be recognized.
[0180] The game character determination module is used to obtain the target game character associated with the target skill frame.
[0181] The game label determination module is used to determine the second label according to the target game character.
[0182] In an exemplary embodiment of the present application, based on the foregoing embodiment, the new and popular label recognition module 1203 may include a video data acquisition module, a video similarity calculation module, and a new and popular label determination module.
[0183] Among them, the video data acquisition module is used to obtain a set of real-time updated new and popular labels and obtain the video data corresponding to the set of new and popular labels.
[0184] The video similarity calculation module is used to calculate the similarity between the title of the video to be recognized and the title of the video data, and screen out multiple target videos from the video data whose similarity meets a preset threshold.
[0185] A new hot tag determination module, configured to determine a third tag of the video to be recognized through new hot tags corresponding to the multiple target videos.
[0186] In an exemplary embodiment of the present application, based on the foregoing embodiment, the new hot tag recognition module may be configured to: determine tags that conform to the adoption strategy in the first tag and the second tag as a fourth tag according to the adoption strategies corresponding to the first tag and the second tag respectively; determine the video tag of the video to be recognized according to the fourth tag and the third tag.
[0187] In an exemplary embodiment of the present application, the new hot tag recognition module may be configured to: for each first tag, if the model outputting the first tag is the confidence model corresponding to the first tag, determine the first tag as a fourth tag; wherein, the confidence level output by the confidence model corresponding to the first tag conforms to the adoption strategy corresponding to the first tag; for each second tag, if the model outputting the second tag is the confidence model corresponding to the second tag, determine the second tag as a fourth tag; wherein, the confidence level output by the confidence model corresponding to the second tag conforms to the adoption strategy corresponding to the second tag.
[0188] In an exemplary embodiment of the present application, based on the foregoing embodiment, the determining the first tag of the video to be recognized by using the multi-modal features includes: extracting the multi-modal features of the video to be recognized through a first model, and determining the first tag of the video to be recognized according to the multi-modal features; the recognizing the second tag of the video to be recognized based on the category of the video to be recognized includes: recognizing the video to be recognized through a second model corresponding to the category of the video to be recognized to obtain the second tag of the video to be recognized; the apparatus further includes a sample acquisition module, a confidence level test module, and an adoption strategy determination module.
[0189] Among them, the first test result module is configured to recognize the first tags of multiple video samples through the first model as the first test result.
[0190] The second test result module recognizes the second tags of the multiple video samples through the second model as the second test result.
[0191] The confidence level test module is configured to calculate a first confidence level of the first model for each tag included in the first test result, and calculate a second confidence level of the second model for each tag included in the second test result.
[0192] The credibility strategy determination module is configured to, for the target tags that overlap between the first test result and the second test result, determine the credibility model corresponding to each target tag from the first model and the second model according to the first confidence level and the second confidence level of each target tag, and use the correspondence between the target tag and the credibility model as the credibility strategy corresponding to the target tag.
[0193] The specific details of each module or unit in the above label recognition device have been described in detail in the corresponding label recognition method, so they will not be elaborated here.
[0194] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0195] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the above-mentioned module, segment of a program, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0196] The units involved in the embodiments described in the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation to the unit itself.
[0197] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device is caused to implement the label recognition method described in the above embodiments.
[0198] For example, the electronic device may implement as Figure 3 shown: Step S310. Extract multi-modal features of the video to be recognized, and use the multi-modal features to determine the first label of the video to be recognized; Step S320. Classify and recognize the video to be recognized to obtain the category of the video to be recognized, and recognize the second label of the video to be recognized based on the category of the video to be recognized; and Step S330. Obtain the third label of the video to be recognized through a newly updated set of hot labels, and combine the first label, the second label, and the third label to determine the video label of the video to be recognized.
[0199] For another example, the electronic device may implement each step as Figures 4 to 7 shown.
[0200] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0201] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0202] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0203] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A label recognition method, characterized in that, it includes: extracting multi-modal features of the video to be recognized, extracting the multi-modal features of the video to be recognized through a first model, and determining a first label of the video to be recognized according to the multi-modal features; classifying and recognizing the video to be recognized to obtain the category of the video to be recognized, and identifying a second label of the video to be recognized based on the category of the video to be recognized; wherein, the identifying the second label of the video to be recognized based on the category of the video to be recognized includes: performing gating control based on the category of the video to be recognized, distributing the video to be recognized to a second model corresponding to the category, and performing recognition on the video to be recognized based on the second model to obtain the second label of the video to be recognized, and the second model includes one or more models related to the category; obtaining a third label of the video to be recognized through a newly updated popular label set, and determining a fourth label by taking the labels that meet the credibility strategy from the first label and the second label according to the credibility strategies corresponding to the first label and the second label respectively; wherein, the credibility strategy corresponding to the target label is obtained through the following process: identifying the first labels of multiple video samples through the first model as the first test result; identifying the second labels of the multiple video samples through the second model as the second test result; calculating the first confidence of the first model for each label included in the first test result, and calculating the second confidence of the second model for each label included in the second test result; for the target labels that overlap between the first test result and the second test result, determining the confidence model corresponding to each target label from the first model and the second model according to the first confidence and the second confidence of each target label, and taking the corresponding relationship between the target label and the confidence model as the credibility strategy corresponding to the target label; determining a video label of the video to be recognized according to the fourth label and the third label.
2. The label recognition method according to claim 1, characterized in that, the identifying the second label of the video to be recognized based on the category of the video to be recognized includes: when the category of the video to be recognized is the category of film, television, and variety shows, identifying the target object in the video to be recognized; determining the second label of the video to be recognized based on the target object.
3. The label recognition method according to claim 2, characterized in that, the target object includes human features, and the determining the second label of the video to be recognized based on the target object includes: recognizing the human features in the video to be recognized to determine a human label of the video to be recognized; acquiring the film, television, and variety show names related to the human label through a knowledge graph model; calculating the similarity between the video to be recognized and the film, television, and variety show video corresponding to the film, television, and variety show name, and determining a film, television, and variety show label of the video to be recognized; determining the second label according to the human label and the film, television, and variety show label.
4. The label recognition method according to claim 1, It is characterized in that The second label of the video to be recognized recognized based on the category of the video to be recognized includes When the category of the video to be recognized is a game category, match the game template data with the video to be recognized, and determine the second label of the video to be recognized according to the matched game template data 5. The label recognition method according to claim 4 It is characterized in that The game template data includes a skill box template, and the matching of the game template data with the video to be recognized and the determination of the second label of the video to be recognized according to the matching result include Match the video to be recognized with a plurality of skill box templates to determine a target skill box that matches the video to be recognized Obtain a target game character associated with the target skill box Determine the second label according to the target game character 6. The label recognition method according to claim 1 It is characterized in that The obtaining of the third label of the video to be recognized through the newly popular label set updated in real time includes Obtain the newly popular label set updated in real time, and obtain the video data corresponding to the newly popular label set Calculate the similarity between the title of the video to be recognized and the title of the video data, and screen out a plurality of target videos from the video data whose similarity meets a preset threshold Determine the third label of the video to be recognized through the newly popular labels corresponding to the plurality of target videos 7. The label recognition method according to claim 1 It is characterized in that The determining the first label and the second label that meet the adoption strategy as the fourth label according to the adoption strategies corresponding to the first label and the second label respectively includes For each first label, if the model outputting the first label is the confidence model corresponding to the first label, then determine the first label as a fourth label; wherein, the confidence model corresponding to the first label outputs a confidence level of the first label that meets the adoption strategy corresponding to the first label For each second label, if the model outputting the second label is the confidence model corresponding to the second label, then determine the second label as a fourth label; wherein, the confidence model corresponding to the second label outputs a confidence level of the second label that meets the adoption strategy corresponding to the second label 8. A label recognition device It is characterized in that Comprising A multimodal feature extraction module, configured to extract multimodal features of a video to be recognized, extract multimodal features of the video to be recognized through a first model, and determine a first label of the video to be recognized according to the multimodal features A category recognition module, used for classifying and identifying the video to be identified, obtaining the category of the video to be identified, and identifying a second label of the video to be identified based on the category of the video to be identified; wherein the identifying the second label of the video to be identified based on the category of the video to be identified includes: performing gate control based on the category of the video to be identified, distributing the video to be identified to a second model corresponding to the category, identifying the video to be identified based on the second model, and obtaining the second label of the video to be identified, wherein the second model includes one or more models related to the category; A new hot tag identification module, used to obtain the third tag of the video to be identified through the real-time updated new hot tag set, and determine the video tag of the video to be identified by combining the first tag, the second tag and the third tag; The new hot tag identification module is configured to: determine the tag that complies with the acceptance strategy among the first tag and the second tag as a fourth tag according to the acceptance strategy respectively corresponding to the first tag and the second tag; determine the video tag of the video to be identified according to the fourth tag and the third tag; The device also includes: A first test result module, used for identifying first labels of multiple video samples by using the first model as a first test result; A second test result module, configured to identify second labels of the plurality of video samples by using the second model as a second test result; A confidence testing module, used to calculate a first confidence of the first model for each label included in the first test result, and to calculate a second confidence of the second model for each label included in the second test result; A confidence strategy determination module is used to determine, for target labels that overlap in the first test result and the second test result, a confidence model corresponding to each target label from the first model and the second model according to the first confidence level and the second confidence level of each target label, and use the correspondence between the target label and the confidence model as the confidence strategy corresponding to the target label.
9. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the tag identification method according to any one of claims 1 to 7 is implemented.
10. An electronic device, It is characterized in that include: processor; as well as A memory for storing executable instructions of the processor; wherein the processor is configured to perform the tag identification method according to any one of claims 1 to 7 by executing the executable instructions.
11. A computer program product, It is characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the tag identification method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for determining video label and computer equipment
CN111125435A