Video processing method and device, computer device and storage medium
By performing entity recognition on the target video and disambiguation processing using the video feature information of the publisher, the problem of low accuracy of video tags in traditional methods is solved, and more accurate video tag settings are achieved.
Patent Information
- Application Number
- CN202111439163.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2041-11-29
AI Technical Summary
Traditional entity disambiguation methods rely on the completeness of video content. When the video content is not clear or complete enough, the accuracy of video tag setting decreases.
By performing entity recognition on the target video, the video feature information of the object that published the video is obtained and associated with the features of its historical videos. Based on this information, multiple candidate entities are disambiguated to determine the target entity and set video tags.
It enhances the entity disambiguation effect, improves the accuracy of video tag settings, and makes the disambiguated candidate entities more closely match the historical video feature information of the publishing user.
Smart Images

Figure CN114329064B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to a video processing method and device, a computer device and a storage medium. BACKGROUND
[0002] With the rapid development of network technology and the popularization and application of multimedia, various video sources are continuously generated, and video and other media data have become the main body of big data. In order to facilitate the recommendation and distribution of a large number of videos, a video tag related to the video content can be added to the video, such as displaying the title or abstract of the video in the user interface.
[0003] For the video tag, most of the video tags are entity tags, but there can be multiple candidate entities for the same entity tag, such as an entity tag with the name "Zhang Fei" corresponding to multiple candidate entities, such as historical figures and game characters. Therefore, when setting the video tag, the multiple candidate entities need to be disambiguated, so that a reasonable tag can be accurately set for the video. At present, the traditional entity disambiguation method depends on the completeness of the video content, and when the video content is not clear and complete enough, it is not sufficient to support entity disambiguation, thereby reducing the accuracy of video tag setting. SUMMARY
[0004] The embodiments of the present application provide a video processing method and device, a computer device and a storage medium, which can enhance the entity disambiguation effect and improve the accuracy of video tag setting.
[0005] In one aspect, the embodiments of the present application provide a video processing method, which comprises:
[0006] performing entity recognition on a target video to determine at least one entity tag corresponding to the target video;
[0007] determining a target entity tag with multiple candidate entities in the at least one entity tag;
[0008] obtaining video feature information of an object publishing the target video, the video feature information being associated with the video features of historical videos published by the object;
[0009] performing disambiguation processing on the multiple candidate entities corresponding to the target entity tag based on the video feature information, to obtain a target entity corresponding to the target entity tag, the target entity including one or more of the multiple candidate entities;
[0010] determining the target entity corresponding to each entity tag in the at least one entity tag as a video tag of the target video.
[0011] In one aspect, the embodiments of the present application provide a video processing device, which comprises:
[0012] An identifying unit is configured to perform entity identification on a target video to determine at least one entity label corresponding to the target video.
[0013] A determining unit is configured to determine, among the at least one entity label, a target entity label of which there are multiple candidate entities.
[0014] An obtaining unit is configured to obtain video feature information of an object that publishes the target video, the video feature information being associated with video features of historical videos published by the object.
[0015] A disambiguating unit is configured to perform disambiguation processing on the multiple candidate entities corresponding to the target entity label based on the video feature information, to obtain a target entity corresponding to the target entity label, the target entity including one or more of the multiple candidate entities.
[0016] The determining unit is further configured to determine, as a video label of the target video, a target entity corresponding to each entity label among the at least one entity label.
[0017] In an aspect, an embodiment of the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program. The computer program is executed by the processor to cause the processor to perform the video processing method.
[0018] In an aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program. The computer program is read and executed by a processor of a computer device to cause the computer device to perform the video processing method.
[0019] In an aspect, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and executes the computer instructions to cause the computer device to perform the video processing method.
[0020] By the embodiment of the present application, entity recognition is performed on a target video to determine at least one entity label corresponding to the target video; a target entity label in which multiple candidate entities exist is determined in the at least one entity label; video feature information of an object publishing the target video is obtained, the video feature information being associated with video features of historical videos published by the object; multiple candidate entities corresponding to the target entity label are disambiguated based on the video feature information to obtain a target entity corresponding to the target entity label, the target entity including one or more of the multiple candidate entities; and target entities corresponding to each entity label in the at least one entity label are determined as video labels of the target video. It should be understood that entity disambiguation is performed using video feature information of an object publishing the target video, so that the disambiguated candidate entity is more matched with feature information of historical videos of the publishing user, thereby enhancing the entity disambiguation effect and improving the accuracy of video label setting. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0022] Figure 1 is a schematic diagram of a video processing system architecture provided by an embodiment of the present application;
[0023] Figure 2 is a schematic diagram of a video processing method provided by an embodiment of the present application;
[0024] Figure 3 is a schematic diagram of a target video and corresponding video labels provided by an embodiment of the present application;
[0025] Figure 4 is a structural schematic diagram of an entity recognition model provided by an embodiment of the present application;
[0026] Figure 5 is a video processing flowchart provided by an embodiment of the present application;
[0027] Figure 6 is a schematic diagram of another video processing method provided by an embodiment of the present application;
[0028] Figure 7 is a schematic diagram of another video processing method provided by an embodiment of the present application;
[0029] Figure 8 is a structural schematic diagram of a deep matching model provided by an embodiment of the present application;
[0030] Figure 9 is a flowchart of another video processing method provided by an embodiment of the present application;
[0031] Figure 10 is a flowchart of another video processing method provided by an embodiment of the present application;
[0032] Figure 11 is a structural diagram of a video processing device provided by an embodiment of the present application;
[0033] Figure 12 is a structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0034] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative work fall within the scope of protection of the present application.
[0035] It should be noted that the terms "first", "second", etc. involved in the embodiments of the present application are only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the technical features indicated or the number of the technical features indicated. Therefore, the technical features limited by "first" and "second" can explicitly or implicitly include at least one of the features.
[0036] First, some nouns involved in the embodiments of the present application are explained to facilitate the understanding of those skilled in the art.
[0037] SPERT model: SPERT model is a joint entity and relationship extraction model taking transformer network BERT as the core, which realizes joint extraction by adopting the idea of classification. The entity extraction and relationship extraction models are both classification models, which predict the entity types of all possible text segments in a given text by adopting the idea of exhaustion. Relationship extraction depends on the extracted entities, and predicts the relationship types of all combinations of the extracted entities. The text feature information between entities is considered in relationship extraction.
[0038] BERT model: full name Bidirectional Encoder Representations from Transformers, is a new language model proposed by Google, which pre-trains bidirectional deep representation (Embedding) by jointly adjusting all layers of bidirectional transformer (Transformer).
[0039] Transformer model: The Transformer model is a classical model of natural language processing (NLP). The Transformer model encodes input and calculates output completely based on attention, without relying on sequence-aligned recurrent neural networks or convolutional neural networks. The Transformer model uses a self-attention mechanism, without the sequential structure of recurrent neural networks (RNNs), so that the model can be trained in parallel and can have global information. The structure of the Transformer model consists of an encoder layer and a decoder layer.
[0040] Optical character recognition (OCR) technology: refers to the process of analyzing and recognizing image files of text materials to obtain text and layout information. That is, the text in the image is recognized and returned in the form of text.
[0041] Automatic speech recognition (ASR) technology: a technology that converts human speech into text. Through speech signal processing and pattern recognition, the machine automatically recognizes and understands the speech signal and converts it into corresponding text or commands. The main process includes: speech input, encoding (feature extraction), decoding and text output.
[0042] Self-attention (Self-Attention) model: the attention model simulates the internal process of biological observation behavior, that is, a mechanism that aligns internal experience and external feeling to increase the observation accuracy of certain areas. The attention model can quickly extract important features of sparse data, so it is widely used in natural language processing tasks, especially machine translation. The self-attention mechanism is an improvement of the attention model, which reduces the dependence on external information and is better at capturing the internal correlation of data or features.
[0043] The video processing method provided in the embodiments of the present application can be implemented based on an artificial intelligence technology. The artificial intelligence (AI) is a theory, method, technology and application system for using a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive an environment, acquire knowledge and use the knowledge to obtain optimal results. In other words, the artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. The artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.
[0044] The artificial intelligence technology is a comprehensive discipline, involves a wide range of fields, and has both hardware level technology and software level technology. The artificial intelligence basic technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. The artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0045] Among them, the computer vision technology (CV) is a science that studies how to make a machine "see", and further refers to using a camera and a computer to replace human eyes to identify, track and measure a target and other machine vision, and further to perform image processing, so that the computer processing becomes an image more suitable for human eye observation or transmission to an instrument detection. As a scientific discipline, the computer vision researches related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. The computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. It also includes common face recognition, fingerprint recognition and other biometric identification technologies.
[0046] The key technologies of the speech technology include an automatic speech recognition technology (ASR), a speech synthesis technology (TTS) and a voiceprint recognition technology. Letting a computer can hear, can see, can speak, can feel is the development direction of future human-computer interaction, and the speech becomes one of the most promising human-computer interaction modes in the future.
[0047] Natural language processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, i.e., the language used in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph, etc.
[0048] With the research and progress of artificial intelligence technology, artificial intelligence technology is being researched and applied in many fields, such as common smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned vehicles, autonomous vehicles, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0049] Currently, the traditional entity disambiguation method relies on the completeness of the video content. When the video content is not clear and complete enough, it is not sufficient to support entity disambiguation, thereby reducing the accuracy of video tag setting. The embodiments of the present application consider the video feature information of the object publishing the target video, which is used to indicate the video features of the historical videos published by the object, such as the video type of the historical video and the video tag of the historical video. To some extent, disambiguating the candidate entities based on the video feature information of the object publishing the target video can enhance the entity disambiguation effect and improve the accuracy of video tag setting. Therefore, the embodiments of the present application propose a video processing scheme, specifically, performing entity recognition on a target video to obtain at least one entity tag of the target video, for the entity tag with multiple candidate entities, disambiguating the multiple candidate entities corresponding to the entity tag based on the video feature information to obtain a target entity corresponding to the entity tag, and further determining the target entity corresponding to each entity tag in the at least one entity tag as a video tag of the target video.
[0050] The video processing method proposed in the present application is executed by a computer device, which can be a terminal device such as a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart car, etc., but is not limited thereto; the computer device can also be a server, such as a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and basic cloud computing services such as big data and artificial intelligence platforms.
[0051] Alternatively, it can be performed jointly by computer equipment and video processing equipment, such as the computer equipment being the terminal device and the video processing equipment being the server; or the computer equipment being the server and the video processing equipment being the terminal device.
[0052] For example, assuming the computer device is the terminal device and the video processing device is the server, the video processing solution proposed in this application can be implemented using the following video processing system architecture. Please see [link to relevant documentation]. Figure 1 , Figure 1 This is a schematic diagram of the architecture of a video processing system according to an embodiment of this application, such as... Figure 1 As shown, the video processing system 100 may include one or more terminal devices 101 and a server 102. Of course, the video processing system 100 may also include one or more terminal devices 101 and multiple servers 102; this embodiment does not limit this. The terminal device 101 is mainly used to send one or more target videos to the server 102 and receive video tags of the target videos sent by the server 102. The server 102 is mainly used to execute the relevant steps of the video processing method, obtain the video tags of the target videos, and send the video tags of the target videos to the terminal device 101. The terminal device 101 and the server 102 can be connected, and the connection method may include wired connection and wireless connection, which is not limited here.
[0053] Based on the video processing system described above, the video processing method of this application embodiment can generally include:
[0054] Terminal device 101 sends the target video to server 102. After receiving the target video from terminal device 101, server 102 performs entity recognition on the target video to determine at least one entity tag corresponding to the target video. For entity tags with multiple candidate entities, server 102 further obtains the video feature information of the object that published the target video, uses this video feature information to disambiguate the multiple candidate entities, and determines the target entity corresponding to each entity tag as the video tag of the target video, then sends the video tag of the target video to terminal device 101. Alternatively, terminal device 101 can also perform entity recognition on the target video, obtain at least one entity tag corresponding to the target video, and then send the at least one entity tag to server 102. By determining the video tag of the target video through this method, and using the video feature information of the object that published the target video for entity disambiguation, the disambiguated candidate entities are more closely matched with the video feature information of the video publishing object, thereby enhancing the entity disambiguation effect and improving the accuracy of video tag setting.
[0055] In an embodiment, there can be multiple candidate entities corresponding to an entity label, or there can be only one candidate entity. For an entity label with only one candidate entity, entity disambiguation is not needed, and the application can directly determine the candidate entity as the target entity corresponding to the entity label. For example, assuming that the server 102 performs entity recognition on a target video and determines three entity labels corresponding to the target video, namely entity label 1, entity label 2, and entity label 3, where entity label 1 has multiple candidate entities, entity label 2 has only one candidate entity, and entity label 3 has only one candidate entity. Then the server 102 can disambiguate the multiple candidate entities corresponding to entity label 1 based on the video feature information described above to obtain the target entity corresponding to entity label 1. The server 102 can also determine the candidate entity corresponding to entity label 2 as the target entity corresponding to entity label 2, and determine the candidate entity corresponding to entity label 3 as the target entity corresponding to entity label 3. Further, the server 102 can determine the target entity corresponding to entity label 1, the target entity corresponding to entity label 2, and the target entity corresponding to entity label 3 as the video label of the target video, and send the video label to the terminal device 101. Subsequent embodiments are described in the scenario where an entity label has multiple candidate entities, but the entity labels corresponding to the target video are not limited to only entity labels with multiple candidate entities.
[0056] It can be understood that the system architecture diagram described in the embodiments of the application is used to more clearly illustrate the technical solutions of the embodiments of the application, and does not constitute a limitation on the technical solutions provided by the embodiments of the application. Those skilled in the art can know that, as the system architecture evolves and new business scenarios appear, the technical solutions provided by the embodiments of the application are also applicable to similar technical problems.
[0057] Based on the above description of the video processing system architecture, the embodiments of the application disclose a video processing method, please see Figure 2 , a flowchart of a video processing method disclosed by the embodiments of the application. The video processing method can be executed by a computer device, which can be specifically a server 102 in a video processing system. The video processing method can specifically include steps S201-S205:
[0058] S201, performing entity recognition on a target video to determine at least one entity label corresponding to the target video.
[0059] In the embodiments of the present application, the target video can be any video published by a user, which is not limited herein. In order to facilitate the recommendation and distribution of a large number of videos, a video tag related to the content of the video can be added to the video, for example, the title or abstract of the video is displayed in the user interface. For the video tag, most of the video tags include entity tags. By performing entity recognition on the target video, at least one entity tag corresponding to the target video can be determined.
[0060] As shown in Figure 3 Figure 3 is a schematic diagram of a target video and a corresponding video tag provided by the embodiments of the present application, wherein the video serial number of the left video is 843017, the entity tag corresponding to the video is "Zhang Fei", the entity type corresponding to the entity tag is "culture-person name", the entity corresponding to the entity tag is "a famous general of Shu Han during the Three Kingdoms Period", and the video tag of the video is "Romance of the Three Kingdoms: Zhang Fei is brave and overbearing, and is a famous general of Shu Han during the Three Kingdoms Period". The video serial number of the right video is 55752408, the entity tag corresponding to the video is "Zhang Fei", the entity type corresponding to the entity tag is "game character", the entity corresponding to the entity tag is "a character in XX game", and the video tag of the video is "XX game: Zhang Fei's attack skill is really strong". As can be seen, the entity tags of the left video and the right video are both "Zhang Fei".
[0061] In a possible implementation, the entity recognition on the target video to determine at least one candidate entity tag corresponding to the target video comprises: obtaining text information corresponding to the target video; calling an entity recognition model to perform recognition processing on the text information to obtain intermediate result information corresponding to each text segment of the text information, the intermediate result information including one or more of entity region information, interval length feature information, interval context feature information and global context feature information; calling the entity recognition model to perform fusion processing on each information included in the intermediate result information to obtain at least one candidate entity tag corresponding to the target video. That is, the text information corresponding to the target video can be divided into multiple text segments for recognition processing to obtain the intermediate result information of each text segment, and further fusion processing is performed according to each information included in the intermediate result information of each text segment, so as to determine at least one candidate entity tag corresponding to the target video. The fusion processing here can be an average pooling operation on each information included in the intermediate result information, so as to realize the prediction of the entity tag. In addition, the length of each text segment can be pre-set, and the set value is not limited herein. The text information corresponding to the target video can be obtained by performing OCR recognition and ASR recognition on the target video to recognize the subtitle text, voice dialogue text and the like in the target video, which is not limited herein.
[0062] For example, the length of the text segment is set to 10 words in advance, the text information corresponding to the target video includes 30 words, and therefore the text information can be divided into 3 text segments. The entity recognition model is called to process the 3 text segments corresponding to the text information, to obtain intermediate result information of the 3 text segments. The information included in the intermediate result information of the 3 text segments is further fused to obtain at least one candidate entity label corresponding to the target video.
[0063] It should be noted that the fusion processing of the plurality of feature information can comprehensively utilize a plurality of feature information, realize the complementary advantages of multiple features, reduce the influence of single feature limitations, and thus improve the accuracy of information recognition. The fusion processing mode of the plurality of feature information can be feature fusion based on the Bayesian theory, feature fusion based on the sparse representation theory, or feature fusion based on the deep learning theory, which is not limited herein. For example, the fusion processing mode of the information included in the intermediate result information can be splicing, i.e., summation calculation, or maximum or average calculation, or network layer fusion.
[0064] Please refer to Figure 4 , Figure 4 is a structural schematic diagram of an entity recognition model provided in the embodiments of the present application. As shown in Figure 4 , the entity recognition model includes a SPERT layer composed of a SPERT network. The text information corresponding to the target video is input into the SPERT layer in the entity recognition model for processing, to output intermediate result information corresponding to each text segment of the text information. The intermediate result information includes entity region information, interval length feature information, interval context feature information, and global context feature information. Then, the information included in the intermediate result information is fused to obtain entity prediction results of each text segment, so as to determine at least one candidate entity label corresponding to the target video. The entity region information is used to indicate the entity region of the text segment.
[0065] S202, determining a target entity label of the plurality of candidate entities in the at least one entity label.
[0066] In the embodiments of the present application, the candidate entity corresponding to the entity label can exist only one or multiple. For example, please refer to Figure 3 , Figure 3The entity label in the video on the left is "Zhang Fei," corresponding to the entity "a famous general of Shu Han during the Three Kingdoms period." The entity label in the video on the right is also "Zhang Fei," corresponding to the entity "a character in XX game." Therefore, the entity label "Zhang Fei" corresponds to multiple candidate entities, including a famous general of Shu Han during the Three Kingdoms period and a character in XX game; thus, this entity label has multiple candidate entities. For example, if the entity label is "hydrogen," this entity label has only one candidate entity, namely "gas," therefore, this entity label has only one candidate entity. After identifying a target entity label with multiple candidate entities among at least one entity label, entity disambiguation processing is subsequently performed on this target entity label.
[0067] S203. Obtain the video feature information of the object that published the target video, and associate the video feature information with the video features of the historical videos published by the object.
[0068] In this embodiment, the object that publishes the target video can refer to the user who published the target video, and the video feature information is used to indicate the video features of the historical videos published by the object. This video feature information can be obtained through processing and analysis of historical videos, or it can be directly obtained from other devices; no limitation is made here. The purpose of obtaining the video feature information of the object who published the target video is to facilitate subsequent disambiguation processing of multiple candidate entities corresponding to the target entity tag using this video feature information.
[0069] In one possible implementation, obtaining the video feature information of the object that published the target video includes: obtaining historical videos published by the object within a preset time period; analyzing and processing the historical videos to obtain historical video data information, which includes the video type and video tag of the historical videos; and determining the historical video data information as the video feature information of the object. It should be understood that the video feature information of the object includes the video type and video tag of the historical videos, and may also include other data information, which is not limited here.
[0070] For example, please see again Figure 3 ,Will Figure 3 The video on the left is considered a historical video posted by the entity that posted the target video. This historical video is categorized as a film or television work, and its corresponding video tag is "Romance of the Three Kingdoms: Zhang Fei was exceptionally brave and a famous general of Shu Han during the Three Kingdoms period." This video tag is set based on the entity "Famous General of Shu Han during the Three Kingdoms period." Therefore, the historical video data information includes: the video type is a film or television work, and the corresponding video tag is "Romance of the Three Kingdoms: Zhang Fei was exceptionally brave and a famous general of Shu Han during the Three Kingdoms period."
[0071] S204, perform disambiguation processing on the multiple candidate entities corresponding to the target entity label based on the video feature information, to obtain a target entity corresponding to the target entity label, the target entity including one or more of the multiple candidate entities.
[0072] In the embodiments of the present application, the disambiguation processing refers to determining one candidate entity that best matches the target entity label from the multiple candidate entities, i.e., the target entity corresponding to the target entity label. Based on the analysis of a large number of historical videos published by the video publishing object, it is found that the video content published by the video publishing object tends to be stable, for example, a creator who publishes historical dramas will not suddenly switch to publishing game videos. Therefore, using the video feature information of the object publishing the target video to perform entity disambiguation makes the target entity corresponding to the target entity label more matched with the feature information of the historical videos of the publishing user, thereby enhancing the entity disambiguation effect and improving the accuracy of video label setting.
[0073] S205, determining the target entity corresponding to each entity label in the at least one entity label as a video label of the target video.
[0074] In the embodiments of the present application, for the target entity label with multiple candidate entities, the disambiguation processing is performed on the multiple candidate entities corresponding to the target entity label based on the video feature information, to obtain a target entity corresponding to the target entity label. For the entity label with only one candidate entity, the candidate entity is directly determined as the target entity corresponding to the entity label. Further, the target entity corresponding to each entity label in the at least one entity label is determined as a video label of the target video. It should be noted that the video label can include the target entity corresponding to each entity label, or the video label of the target video can be generated according to the brief information of the target entity corresponding to each entity label, which is not limited herein. Based on this way, a more accurate video label can be constructed for the target video, which is beneficial to the subsequent recommendation and distribution of the target video.
[0075] For example, the entity labels of the target video include entity label M and entity label N, wherein the entity label M has candidate entity A and candidate entity B, and the entity label N has one candidate entity C. Based on the video feature information, the disambiguation processing is performed on the two candidate entities corresponding to the entity label M, to obtain the target entity corresponding to the entity label M as the candidate entity A; the target entity corresponding to the entity label N is the candidate entity C. Therefore, the candidate entity A and the candidate entity C are determined as the video labels of the target video.
[0076] In summary, this video processing method can be divided into three steps: entity recognition of the video content; enhanced disambiguation of video entities based on publisher user characteristics; and the use of the disambiguated video entities as the basis for video tagging for video distribution. Please refer to [link / reference]. Figure 5 , Figure 5 This is a video processing flowchart provided in an embodiment of this application. The specific implementation of step S201 can be summarized as "entity recognition of video content". The specific implementation of steps S202 to S204 can be summarized as "strengthening video entity disambiguation through publisher user features". The specific implementation of step S205 can be summarized as "using the disambiguated video entities as basic data for video tagging and distributing the video". Here, publisher user features correspond to the video feature information of the object that published the target video, video entities correspond to the candidate entities, and disambiguated video entities correspond to the target entities.
[0077] The following specific examples illustrate the determination of video tags:
[0078] Please see again. Figure 3 ,Will Figure 3 The video on the left is considered the target video for which we need to set video tags. First, we perform entity recognition on this target video, identifying the corresponding entity tag as "Zhang Fei." This entity tag has multiple candidate entities, including famous generals of Shu Han during the Three Kingdoms period and characters from games. Next, we obtain the video feature information of the user who posted the target video. This video feature information indicates the video characteristics of the user's historical videos, specifically that the video type is film / TV and the video tag is "Romance of the Three Kingdoms." Then, based on this video feature information, we perform disambiguation processing on the multiple candidate entities corresponding to the target entity tag, obtaining the target entity corresponding to the target entity tag, which is "famous general of Shu Han during the Three Kingdoms period." Finally, we can determine this target entity as the video tag for the target video, i.e., the video tag for the target video is determined as "Zhang Fei: Famous general of Shu Han during the Three Kingdoms period." Alternatively, we can generate the video tag for the target video based on the brief information of the target entity, i.e., the video tag for the target video is determined as "Romance of the Three Kingdoms: Zhang Fei was exceptionally brave and was a famous general of Shu Han during the Three Kingdoms period."
[0079] To sum up, in the embodiment of the present application, entity recognition is performed on a target video to determine at least one entity label corresponding to the target video; a target entity label in which multiple candidate entities exist is determined in the at least one entity label; video feature information of an object publishing the target video is obtained, the video feature information being associated with video features of historical videos published by the object; multiple candidate entities corresponding to the target entity label are disambiguated based on the video feature information to obtain a target entity corresponding to the target entity label, the target entity including one or more of the multiple candidate entities; and target entities corresponding to each entity label in the at least one entity label are determined as video labels of the target video. It should be understood that entity disambiguation is performed using video feature information of an object publishing the target video, so that the disambiguated candidate entity is more matched with feature information of historical videos of the publishing user, thereby enhancing the entity disambiguation effect and improving the accuracy of video label setting.
[0080] Based on the above description of the video processing system architecture, the embodiment of the present application discloses a video processing method, please see Figure 6 The flowchart of another video processing method disclosed by the embodiment of the present application is shown in the figure, the video processing method can be executed by a computer device, and the computer device can be specifically a server 102 in a video processing system. The video processing method can specifically include steps S601-S606. Step S604 and step S605 are a specific implementation of step S204. Wherein:
[0081] S601, entity recognition is performed on a target video to determine at least one entity label corresponding to the target video.
[0082] S602, a target entity label in which multiple candidate entities exist is determined in the at least one entity label.
[0083] S603, video feature information of an object publishing the target video is obtained, the video feature information being associated with video features of historical videos published by the object.
[0084] Wherein, the specific implementation of steps S601-S603 is the same as the specific implementation of steps S201-S203, which will not be repeated here.
[0085] S604, the correlation between each candidate entity in the multiple candidate entities and the video label of the historical video is determined.
[0086] In the embodiment of the present application, the correlation between each candidate entity in the plurality of candidate entities and the video label of the historical video is calculated in a statistical manner. Based on this manner, the disambiguation of the plurality of candidate entities can be performed by using the video feature information of the object publishing the target video, the entity disambiguation effect is enhanced, and the accuracy of video label setting is improved.
[0087] In a possible implementation, the determining the correlation between each candidate entity in the plurality of candidate entities and the video label of the historical video comprises: obtaining a first matching degree of an entity type corresponding to each candidate entity in the plurality of candidate entities and the video type, and a first association degree between the each candidate entity and the video label of the historical video; and determining the correlation between the each candidate entity and the video label of the historical video based on the first matching degree and the first association degree. It should be understood that the correlation between each candidate entity and the video label of the historical video is calculated by using the first matching degree of the entity type corresponding to each candidate entity and the video type and the first association degree between each candidate entity and the video label of the historical video. Specifically, it can be calculated by using formula (1) as follows:
[0088] Ps=Ps_class*Ps_en_asso (1)
[0089] Wherein, Ps represents the correlation between each candidate entity and the video label of the historical video, Ps_class represents the first matching degree of the entity type corresponding to each candidate entity and the video type, and Ps_en_asso represents the first association degree between each candidate entity and the video label of the historical video.
[0090] It should be noted that the entity type refers to the category to which the entity belongs, and the video type refers to the category to which the video belongs. For example, for the entity label "Zhang Fei", the entity label corresponds to two candidate entities, wherein the first candidate entity is "a general of Shu Han in the Three Kingdoms Period", and the entity type corresponding to the candidate entity is "culture-person name"; the second candidate entity is "a character in XX game", and the entity type corresponding to the candidate entity is "game-character". Assuming that the video type of the historical video published by a user is "game video", the entity type corresponding to the second candidate entity is more matched with the video type of the historical video.
[0091] In a possible implementation, the first matching degree of the entity type corresponding to each candidate entity in the plurality of candidate entities and the video type is obtained by: obtaining the number of plays of the video in the historical video whose video type matches the entity type corresponding to each candidate entity, and the total number of plays of the historical video; and determining the first matching degree of the entity type corresponding to each candidate entity meeting the type corresponding to the video label of the historical video based on the number of plays and the total number of plays. It should be understood that the first matching degree of the entity type corresponding to each candidate entity in the plurality of candidate entities and the video type is calculated by using the number of plays of the video in the historical video whose video type is the same as the entity type corresponding to each candidate entity and the total number of plays of the historical video. Specifically, the first matching degree can be calculated by using formula (2) as follows:
[0092]
[0093] wherein Ps_class represents the first matching degree of the entity type corresponding to each candidate entity and the video type, Ps_class_m represents the number of plays of the video in the historical video whose video type is the same as the entity type corresponding to each candidate entity, and Ps_class_n represents the total number of plays of the historical video.
[0094] For example, in the historical video published by the object publishing the target video, the number of plays of the video whose entity type is the same as the candidate entity P is 80, and the total number of plays of the historical video is 100, then the first matching degree of the entity type corresponding to the candidate entity P and the video type is 80%.
[0095] In a possible implementation, the first association degree between each candidate entity and the video label of the historical video is obtained by: obtaining a second association degree between the entity label corresponding to each candidate entity and the video label of the historical video, and a use degree of the video label of the historical video; determining the first association degree between each candidate entity and the video label of the historical video based on the two association degrees and the use degree; wherein the use degree of the video label of the historical video is determined by the number of times the video label of the historical video is marked. It should be understood that the first association degree between each candidate entity and the video label of the historical video is calculated by using the second association degree between the entity label corresponding to each candidate entity and the video label of the historical video and the use degree of the video label of the historical video. Specifically, the first association degree can be calculated by using formula (3) as follows:
[0096] Ps_en_asso=sum_k(P_Use_k*Asso_kx) (3)
[0097] wherein, Ps_en_asso represents the first association degree between each candidate entity and the video tag of the historical video, k represents the video tag of the historical video, x represents the entity tag corresponding to each candidate entity, P_Use_k represents the use degree of the video tag of the historical video, Asso_kx represents the second association degree between the entity tag corresponding to each candidate entity and the video tag of the historical video, sum_k represents summation operation, since the video tag of the historical video can be more than one, summation operation needs to be performed on the parameters (here, the parameter refers to the product of the use degree and the second association degree) corresponding to all video tags of the historical video, so as to obtain the first association degree between each candidate entity and the video tag of the historical video.
[0098] For example, the use degree of the video tag 1 of the historical video is P_Use_1, the use degree of the video tag 2 of the historical video is P_Use_2, the second association degree between the entity tag corresponding to the candidate entity a and the video tag 1 of the historical video is Asso_1a, the second association degree between the entity tag corresponding to the candidate entity a and the video tag 2 of the historical video is Asso_2a, the second association degree between the entity tag corresponding to the candidate entity b and the video tag 1 of the historical video is Asso_1b, and the second association degree between the entity tag corresponding to the candidate entity b and the video tag 2 of the historical video is Asso_2b. Therefore, the first association degree between the candidate entity a and the video tag of the historical video is the product of P_Use_1 and Asso_1a plus the product of P_Use_2 and Asso_2a, and the first association degree between the candidate entity b and the video tag of the historical video is the product of P_Use_1 and Asso_1b plus the product of P_Use_2 and Asso_2b.
[0099] In a possible implementation, the second association degree between the entity tag corresponding to each candidate entity and the video tag of the historical video is obtained by: obtaining the first marking times of the entity tag corresponding to each candidate entity and the video tag of the historical video being marked on the same video, and the total first marking times of the entity tag corresponding to each candidate entity and the video tag of the historical video being marked; and determining the second association degree between the entity tag corresponding to each candidate entity and the video tag of the historical video based on the first marking times and the total first marking times. It should be understood that the second association degree between the entity tag corresponding to each candidate entity and the video tag of the historical video is calculated by using the first marking times of the entity tag corresponding to each candidate entity and the video tag of the historical video being marked on the same video, and the total first marking times of the entity tag corresponding to each candidate entity and the video tag of the historical video being marked. Specifically, it can be calculated by using formula (4) as follows:
[0100]
[0101] wherein, Asso_ij represents the second association degree between the entity label corresponding to each candidate entity and the video label of the historical video, i represents the entity label corresponding to the candidate entity, j represents the video label of the historical video, Asso_ij_m represents the first marking times that the entity label corresponding to each candidate entity and the video label of the historical video are marked on the same video, and Asso_ij_n represents the first marking total times that the entity label corresponding to each candidate entity and the video label of the historical video are marked.
[0102] For example, the first marking times that the entity label i corresponding to the candidate entity and the video label j of the historical video are marked on the same video are 80 times, the marking times of the entity label i corresponding to the candidate entity are 100 times, and the marking times of the video label j of the historical video are 100 times. Therefore, the first marking total times that the entity label i corresponding to the candidate entity and the video label j of the historical video are marked are 200 times, and the second association degree between the entity label i corresponding to the candidate entity and the video label j of the historical video is 40% according to formula (4).
[0103] In a possible implementation, the obtaining the use degree of the video label of the historical video comprises: obtaining the second marking times that the video label of the historical video is marked in the historical video and the second marking total times that the video label of the historical video is marked; and determining the use degree of each candidate entity in the historical video based on the second marking times and the second marking total times. It should be understood that the use degree of the video label of the historical video is calculated by using the second marking times that the video label of the historical video is marked in all historical videos and the second marking total times that the video label of all historical videos is marked. Specifically, the use degree of the video label of the historical video can be calculated by using formula (5) as shown below:
[0104]
[0105] wherein, P_Use_k represents the use degree of the video label of the historical video, P_Use_k_m represents the second marking times that the video label of the historical video is marked in the historical video, P_Use_k_n represents the second marking total times that the video label of the historical video is marked, and k represents the video label of the historical video.
[0106] For example, the second marking times that the video label 1 of the historical video is marked in all historical videos are 80 times, and the second marking total times that the video label of all historical videos is marked are 200 times. Therefore, the use degree of the use degree 1 of the video label of the historical video is 40% according to formula (5).
[0107] The determination of the relevance between the candidate entity and the video label of the historical video is described below by using specific examples:
[0108] Suppose there is a candidate entity a, an entity label i corresponding to the candidate entity a, and a video label j of a historical video.
[0109] 1. Calculation of the second correlation degree between the entity label i corresponding to the candidate entity a and the video label j of the historical video:
[0110] The first marking times of the entity label i corresponding to the candidate entity and the video label j of the historical video are marked on the same video, which is 80 times. The marking times of the entity label i corresponding to the candidate entity are 100 times, and the marking times of the video label j of the historical video are 100 times. Therefore, the first marking total times of the entity label i corresponding to the candidate entity and the video label j of the historical video are 200 times, and the second correlation degree between the entity label i corresponding to the candidate entity and the video label j of the historical video is calculated according to formula (4) as 40%.
[0111] 2. Calculation of the usage degree of the video label j of the historical video:
[0112] The second marking times of the video label j of the historical video are marked in all historical videos, which is 80 times, and the total second marking times of the video labels of all historical videos are 200 times. Therefore, the usage degree of the video label j of the historical video is calculated according to formula (5) as 40%.
[0113] 3. Calculation of the first correlation degree between the candidate entity a and the video label j of the historical video:
[0114] According to formula (3), the first correlation degree between the candidate entity a and the video label j of the historical video is calculated as 16% by using the second correlation degree between the entity label i corresponding to the candidate entity a and the video label j of the historical video and the usage degree of the video label j of the historical video.
[0115] 4. Calculation of the first matching degree of the entity type corresponding to the candidate entity a and the video type:
[0116] In the historical video published by the object publishing the target video, the play times of the video with the same entity type as the candidate entity a are 80, and the total play times of the historical video are 100. Therefore, the first matching degree of the entity type corresponding to the candidate entity a and the video type is calculated according to formula (2) as 80%.
[0117] 5. Calculation of the relevance between the candidate entity a and the video label j of the historical video:
[0118] According to formula (3), the relevance between the candidate entity a and the video label j of the historical video is 12.8%, which is calculated by using the first matching degree of the entity type corresponding to the candidate entity a and the video type and the first association degree between the candidate entity a and the video label j of the historical video.
[0119] S605, disambiguating the multiple candidate entities corresponding to the target entity label based on the relevance of each candidate entity, to obtain a target entity corresponding to the target entity label, the relevance between the target entity and the video label of the historical video being greater than the relevance between each entity other than the target entity in the multiple candidate entities and the video label of the historical video.
[0120] In the embodiment of the present application, by comparing the relevance between each candidate entity and the video label of the historical video, the higher the relevance is, the more matched the candidate entity is with the feature information of the historical video of the publishing user. The candidate entity with the highest relevance is taken as the target entity corresponding to the target entity label, to realize disambiguation of the multiple candidate entities.
[0121] For example, the target entity label corresponds to two candidate entities, which are candidate entity a and candidate entity b. The relevance between the candidate entity a and the video label of the historical video is 12.8%, and the relevance between the candidate entity b and the video label of the historical video is 30%, so the relevance between the candidate entity b and the video label of the historical video is the highest, and the candidate entity b is taken as the target entity corresponding to the target entity label.
[0122] S606, determining the target entity corresponding to each entity label in the at least one entity label as the video label of the target video.
[0123] The specific implementation of step S606 is the same as that of step S205 described above, and will not be repeated here.
[0124] To sum up, in the embodiment of the present application, entity recognition is performed on a target video to determine at least one entity label corresponding to the target video; a target entity label in which multiple candidate entities exist is determined in the at least one entity label; video feature information of an object publishing the target video is obtained, the video feature information being associated with video features of historical videos published by the object; a correlation between each candidate entity in the multiple candidate entities and a video label of the historical videos is determined; the multiple candidate entities corresponding to the target entity label are disambiguated based on the correlation of each candidate entity to obtain a target entity corresponding to the target entity label, the correlation between the target entity and the video label of the historical videos being greater than the correlation between entities other than the target entity in the multiple candidate entities and the video label of the historical videos; and a target entity corresponding to each entity label in the at least one entity label is determined as a video label of the target video. It should be understood that, based on a statistical manner, entity disambiguation is performed using the video feature information of the object publishing the target video, so that the disambiguated candidate entity is more matched with the feature information of the historical videos of the publishing user, thereby enhancing the entity disambiguation effect and improving the accuracy of video label setting.
[0125] Based on the above description of the video processing system architecture, the embodiment of the present application discloses a video processing method, please see Figure 7 , the flowchart of another video processing method disclosed by the embodiment of the present application, which can be executed by a computer device, specifically a server 102 in a video processing system. The video processing method specifically can include steps S701-S707. Steps S704-S706 are another specific implementation of step S204 described above. Wherein:
[0126] S701, entity recognition is performed on a target video to determine at least one entity label corresponding to the target video.
[0127] S702, a target entity label in which multiple candidate entities exist is determined in the at least one entity label.
[0128] S703, video feature information of an object publishing the target video is obtained, the video feature information being associated with video features of historical videos published by the object.
[0129] Wherein, the specific implementation of steps S701-S703 is the same as the specific implementation of steps S201-S203 described above, which will not be repeated here.
[0130] S704, text information corresponding to the target video is obtained.
[0131] In the embodiments of the present application, the text information corresponding to the target video can be obtained by performing OCR recognition and ASR recognition on the target video to recognize the subtitle text, voice dialogue text, etc. in the target video, which is not limited herein. The purpose of obtaining the text information corresponding to the target video is to facilitate the use of the text information corresponding to the target video in the subsequent entity disambiguation process.
[0132] In the embodiments of the present application, the second entity information corresponding to the candidate entity can be obtained through existing knowledge base data, or can be obtained in other ways, which is not limited herein. For example, the existing knowledge base data stores entity labels, candidate entities corresponding to the entity labels, and entity information corresponding to the candidate entities, wherein the entity information corresponding to the candidate entities includes the entity type corresponding to the candidate entity and the description information corresponding to the candidate entity. In the knowledge base data, the entity information corresponding to the candidate entity can be queried through the name of the candidate entity. As shown in Table 1, the entity label "Zhang Fei" is stored in Table 1, the candidate entities corresponding to the entity label include "a famous general of Shu Han during the Three Kingdoms Period", the entity type corresponding to the candidate entity is "culture-person name", and the description information corresponding to the candidate entity is "Zhang Fei is a brave general of Shu Han during the Three Kingdoms Period"; the candidate entities corresponding to the entity label also include "a character in XX game", the entity type corresponding to the candidate entity is "game-character", and the description information corresponding to the candidate entity is "Zhang Fei has strong attack skills in XX game".
[0133] In the embodiments of the present application, the second entity information corresponding to the candidate entity can be obtained through existing knowledge base data, or can be obtained in other ways, which is not limited herein. For example, the existing knowledge base data stores entity labels, candidate entities corresponding to the entity labels, and entity information corresponding to the candidate entities, wherein the entity information corresponding to the candidate entities includes the entity type corresponding to the candidate entity and the description information corresponding to the candidate entity. In the knowledge base data, the entity information corresponding to the candidate entity can be queried through the name of the candidate entity. As shown in Table 1, the entity label "Zhang Fei" is stored in Table 1, the candidate entities corresponding to the entity label include "a famous general of Shu Han during the Three Kingdoms Period", the entity type corresponding to the candidate entity is "culture-person name", and the description information corresponding to the candidate entity is "Zhang Fei is a brave general of Shu Han during the Three Kingdoms Period"; the candidate entities corresponding to the entity label also include "a character in XX game", the entity type corresponding to the candidate entity is "game-character", and the description information corresponding to the candidate entity is "Zhang Fei has strong attack skills in XX game".
[0134] Table 1
[0135]
[0136]
[0137] In the embodiments of the present application, the second entity information corresponding to the candidate entity can be obtained through existing knowledge base data, or can be obtained in other ways, which is not limited herein. For example, the existing knowledge base data stores entity labels, candidate entities corresponding to the entity labels, and entity information corresponding to the candidate entities, wherein the entity information corresponding to the candidate entities includes the entity type corresponding to the candidate entity and the description information corresponding to the candidate entity. In the knowledge base data, the entity information corresponding to the candidate entity can be queried through the name of the candidate entity. As shown in Table 1, the entity label "Zhang Fei" is stored in Table 1, the candidate entities corresponding to the entity label include "a famous general of Shu Han during the Three Kingdoms Period", the entity type corresponding to the candidate entity is "culture-person name", and the description information corresponding to the candidate entity is "Zhang Fei is a brave general of Shu Han during the Three Kingdoms Period"; the candidate entities corresponding to the entity label also include "a character in XX game", the entity type corresponding to the candidate entity is "game-character", and the description information corresponding to the candidate entity is "Zhang Fei has strong attack skills in XX game".
[0138] In the embodiments of the present application, the second entity information corresponding to the candidate entity can be obtained through existing knowledge base data, or can be obtained in other ways, which is not limited herein. For example, the existing knowledge base data stores entity labels, candidate entities corresponding to the entity labels, and entity information corresponding to the candidate entities, wherein the entity information corresponding to the candidate entities includes the entity type corresponding to the candidate entity and the description information corresponding to the candidate entity. In the knowledge base data, the entity information corresponding to the candidate entity can be queried through the name of the candidate entity. As shown in Table 1, the entity label "Zhang Fei" is stored in Table 1, the candidate entities corresponding to the entity label include "a famous general of Shu Han during the Three Kingdoms Period", the entity type corresponding to the candidate entity is "culture-person name", and the description information corresponding to the candidate entity is "Zhang Fei is a brave general of Shu Han during the Three Kingdoms Period"; the candidate entities corresponding to the entity label also include "a character in XX game", the entity type corresponding to the candidate entity is "game-character", and the description information corresponding to the candidate entity is "Zhang Fei has strong attack skills in XX game".
[0139] In a possible implementation, the method for performing disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the text information, the second entity information corresponding to the plurality of candidate entities, and the video feature information of the target video comprises: calling a deep matching model to process the text information, the second entity information corresponding to the plurality of candidate entities, and the video feature information, to obtain a context feature vector of the text information, a feature vector of the second entity information corresponding to the plurality of candidate entities, and a feature vector of the video feature information; performing self-attention calculation on the context feature vector of the text information, the feature vector of the second entity information corresponding to the plurality of candidate entities, and the feature vector of the video feature information, to obtain a second matching degree of each candidate entity in the plurality of candidate entities with the historical video; and performing disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the second matching degree of each candidate entity, to obtain the target entity corresponding to the target entity label, wherein the second matching degree of the target entity with the historical video is greater than the second matching degree of each entity other than the target entity in the plurality of candidate entities with the historical video.
[0140] It should be understood that, based on the entity disambiguation manner of the deep matching model, the text information of the target video, the second entity information corresponding to the plurality of candidate entities, and the video feature information of the historical video are modeled in depth, the deep matching model is trained on the entity label data, and the second matching degree of the candidate entity with the context of the target video and the video feature information of the historical video of the target video publisher can be output. The second entity information corresponding to the plurality of candidate entities includes the type, description information, and other text features of the candidate entity, and the video feature information of the historical video of the target video publisher includes the type, description information, and other text features of the object, and the types of the top k entities with large quantities are taken as additional feature information of the object, to enhance the modeling of the video feature information of the target video publisher. After each candidate entity is processed by the deep matching model, the candidate entity with the highest matching degree is selected as the target entity corresponding to the target entity label.
[0141] See Figure 8 , Figure 8 is a structural schematic diagram of a deep matching model provided by an embodiment of the present application. As Figure 8As shown, the deep matching model includes a SPERT layer composed of a SPERT network, a BERT layer composed of a BERT network, and a self-attention layer. The text information of the target video is input into the SPERT layer for processing, the second entity information corresponding to the plurality of candidate entities is input into the BERT layer for processing, and the video feature information is also input into the BERT layer for processing, to obtain a context feature vector (Tag-Context Attention) of the text information, a feature vector (Mention representation) of the second entity information corresponding to the plurality of candidate entities, and a feature vector (Tag-User Attention) of the video feature information. Then, the self-attention layer is used to perform self-attention calculation on the context feature vector of the text information and the feature vector of the second entity information corresponding to the plurality of candidate entities, and to perform self-attention calculation on the feature vector of the second entity information corresponding to the plurality of candidate entities and the feature vector of the video feature information, to obtain the second matching degree of each candidate entity in the plurality of candidate entities with the historical video. It should be noted that, Figure 8 In the above method, the "label candidate disambiguation probability" corresponds to the "second matching degree of each candidate entity with the historical video", the "video text content" corresponds to the "text information of the target video", the "entity candidate content (type, introduction)" corresponds to the "second entity information corresponding to the plurality of candidate entities", and the "user-type, introduction, historical entity type" corresponds to the "video feature information".
[0142] S707, determining the target entity corresponding to each entity label in the at least one entity label as the video label of the target video.
[0143] In the above method, the "label candidate disambiguation probability" corresponds to the "second matching degree of each candidate entity with the historical video", the "video text content" corresponds to the "text information of the target video", the "entity candidate content (type, introduction)" corresponds to the "second entity information corresponding to the plurality of candidate entities", and the "user-type, introduction, historical entity type" corresponds to the "video feature information".
[0144] To sum up, in the embodiment of the present application, entity recognition is performed on a target video to determine at least one entity label corresponding to the target video; a target entity label in which multiple candidate entities exist is determined in the at least one entity label; video feature information of an object publishing the target video is obtained, the video feature information being associated with video features of historical videos published by the object; text information corresponding to the target video is obtained; second entity information corresponding to the multiple candidate entities is obtained, the second entity information including one or more of an entity type corresponding to the multiple candidate entities and description information corresponding to the multiple candidate entities; multiple candidate entities corresponding to the target entity label are disambiguated based on the text information, the second entity information corresponding to the multiple candidate entities, and the video feature information, to obtain a target entity corresponding to the target entity label; and target entities corresponding to each entity label in the at least one entity label are determined as video labels of the target video. It should be understood that, based on a deep matching model, entity disambiguation is performed by using video feature information of an object publishing a target video and text information of the target video, so that the disambiguated candidate entities are more matched with feature information and video content information of historical videos of a publishing user, the entity disambiguation effect is further enhanced, and the accuracy of video label setting is improved.
[0145] Based on the above description of the video processing system architecture, the embodiment of the present application discloses a video processing method, please refer to Figure 9 For another flowchart of the video processing method disclosed by the embodiment of the present application, the video processing method can be executed by a computer device, which can be a server 102 in a video processing system. The video processing method can specifically include steps S901-S908. Steps S904-S907 are another specific implementation of step S204. Wherein:
[0146] S901, entity recognition is performed on a target video to determine at least one entity label corresponding to the target video.
[0147] S902, a target entity label in which multiple candidate entities exist is determined in the at least one entity label.
[0148] S903, video feature information of an object publishing the target video is obtained, the video feature information being associated with video features of historical videos published by the object.
[0149] S904, a correlation between each candidate entity in the multiple candidate entities and a video label of the historical video is determined.
[0150] S905, acquire text information corresponding to the target video and second entity information corresponding to each candidate entity, the second entity information including one or more of an entity type corresponding to each candidate entity and description information corresponding to each candidate entity.
[0151] S906, determine a second matching degree of each candidate entity based on the text information, the second entity information corresponding to each candidate entity, and the video feature information.
[0152] The specific implementation of step S904 is the same as that of step S604 described above; the specific implementation of step S905 is the same as that of steps S704 and S705 described above; and the specific implementation of step S906 is the same as that of step S706 described above, which will not be repeated here.
[0153] S907, perform disambiguation processing on the multiple candidate entities corresponding to the target entity label based on the relevance of each candidate entity and the second matching degree of each candidate entity, to obtain a target entity corresponding to the target entity label.
[0154] In the embodiments of the present application, the relevance between each candidate entity in the multiple candidate entities and the video label of the historical video is calculated based on a statistical method, and the second matching degree between each candidate entity and the historical video is obtained based on a deep matching model. The two methods are combined to realize disambiguation of the multiple candidate entities, so that the disambiguated candidate entities are more matched with the feature information and video content information of the historical video of the publishing object, greatly enhancing the entity disambiguation effect and improving the accuracy of video label setting.
[0155] It should be noted that the two parts of values (i.e., each relevance and each matching degree) can be fused by linear difference, or other fusion methods can be used for processing, which is not limited here. For the fusion method of linear difference, the key is to determine the weight corresponding to the two parts of values. The weight here can be a weight preset according to the experience of related technical personnel, or a best weight value trained based on a training model through network search and post-evaluation, which is not limited here.
[0156] For example, the candidate entities corresponding to the target entity tag include candidate entity a and candidate entity b. The relevance between candidate entity a and the video tag of the historical video is 12.8%, and the relevance between candidate entity b and the video tag of the historical video is 30%. The second matching degree between candidate entity a and the historical video is 20%, and the second matching degree between candidate entity b and the historical video is 35%. The weight for relevance is 0.4, and the weight for matching degree is 0.6. By adding the weights, the fused matching degree corresponding to candidate entity a is 17.12%, and the fused matching degree corresponding to candidate entity b is 33%. Therefore, the candidate entity with the highest fused matching degree is taken as the target entity corresponding to the target entity tag, that is, candidate entity b is the target entity.
[0157] S908. Determine the target entity corresponding to each entity label in the at least one entity label as the video label of the target video.
[0158] The specific implementation method of step S908 is the same as that of step S707 above, and will not be described in detail here.
[0159] To sum up, in the embodiment of the present application, entity recognition is performed on a target video to determine at least one entity label corresponding to the target video; a target entity label in which multiple candidate entities exist is determined in the at least one entity label; video feature information of an object publishing the target video is obtained, the video feature information being associated with video features of historical videos published by the object; text information corresponding to the target video is obtained; second entity information corresponding to the multiple candidate entities is obtained, the second entity information including one or more of an entity type corresponding to the multiple candidate entities and description information corresponding to the multiple candidate entities; a deep matching model is called to process the text information, the second entity information corresponding to the multiple candidate entities, and the video feature information to obtain a context feature vector of the text information, a feature vector of the second entity information corresponding to the multiple candidate entities, and a feature vector of the video feature information; self-attention calculation is performed on the context feature vector of the text information, the feature vector of the second entity information corresponding to the multiple candidate entities, and the feature vector of the video feature information to obtain a second matching degree of each candidate entity in the multiple candidate entities with the historical video; a correlation degree between each candidate entity in the multiple candidate entities and a video label of the historical video is determined; the multiple candidate entities corresponding to the target entity label are disambiguated based on the correlation degree of each candidate entity and the second matching degree of each candidate entity to obtain a target entity corresponding to the target entity label; and the target entity corresponding to each entity label in the at least one entity label is determined as a video label of the target video. It should be understood that the entity disambiguation scheme based on statistics and the entity disambiguation scheme based on the deep matching model are combined to jointly perform entity disambiguation, so that the disambiguated candidate entity is more matched with the feature information and the video content information of the historical videos of the publishing user, thereby greatly enhancing the entity disambiguation effect and improving the accuracy of video label setting.
[0160] Based on the above description of the video processing system architecture, the embodiment of the present application discloses a video processing method, please see Figure 10 , the flowchart of another video processing method disclosed by the embodiment of the present application, which can be executed by a computer device, specifically a server 102 in a video processing system. The video processing method specifically can include steps S1001-S1008. Steps S1004-S1007 are another specific implementation of step S204. Wherein:
[0161] S1001, entity recognition is performed on a target video to determine at least one entity label corresponding to the target video.
[0162] S1002, a target entity label in which multiple candidate entities exist is determined in the at least one entity label.
[0163] S1003, obtain video feature information of an object publishing the target video, the video feature information being associated with video features of historical videos published by the object.
[0164] S1004, determine a relevance between each candidate entity in the plurality of candidate entities and a video label of the historical video.
[0165] The specific implementation of steps S1001-S1004 is the same as that of steps S601-S604, and thus is not repeated here.
[0166] S1005, determine a plurality of target candidate entities based on the relevance of each candidate entity.
[0167] In the embodiment of the present application, the relevance of each candidate entity in the plurality of candidate entities is determined in a statistical manner, and t target candidate entities are selected according to the relevance of each candidate entity, where t is a positive integer greater than 1. For example, the relevance of candidate entity A is 10%, the relevance of candidate entity B is 20%, and the relevance of candidate entity C is 30%. The relevance of each candidate entity is sorted, and the top 2 candidate entities with the largest relevance are selected as target candidate entities. Therefore, candidate entity B and candidate entity C are the two target candidate entities determined.
[0168] S1006, obtain text information corresponding to the target video and first entity information corresponding to the plurality of target candidate entities, the first entity information including one or more of an entity type corresponding to the plurality of target candidate entities and description information corresponding to the plurality of target candidate entities.
[0169] S1007, perform disambiguation processing on the plurality of target candidate entities corresponding to the target entity label based on the text information, the first entity information corresponding to the plurality of target candidate entities, and the video feature information, to obtain a target entity corresponding to the target entity label.
[0170] In the embodiment of the present application, the text information corresponding to the target video, the first entity information corresponding to the plurality of target candidate entities, and the video feature information are further used in a deep matching model to perform entity disambiguation, to obtain a target entity corresponding to the target entity label. It should be understood that the embodiment of the present application first uses a statistical method to process and determine a plurality of target candidate entities with higher relevance, and then uses a deep matching model to further disambiguate the plurality of target candidate entities, thereby enhancing the effect of entity disambiguation and improving the accuracy of video label setting.
[0171] S1008, determine the target entity corresponding to each entity label in the at least one entity label as a video label of the target video.
[0172] The specific implementation of step S1006 is the same as that of step S905 described above; the specific implementation of step S1007 is the same as that of step S706 described above; the specific implementation of step S1008 is the same as that of step S707 described above, and thus is not described herein again.
[0173] To sum up, in the embodiment of the present application, the target video is subjected to entity recognition to determine at least one entity label corresponding to the target video; a target entity label in which multiple candidate entities exist is determined in the at least one entity label; video feature information of an object publishing the target video is acquired, the video feature information being associated with a video feature of a historical video published by the object; a correlation between each candidate entity in the multiple candidate entities and a video label of the historical video is determined; multiple target candidate entities are determined based on the correlation of each candidate entity; text information corresponding to the target video and first entity information corresponding to the multiple target candidate entities are acquired, the first entity information including one or more of an entity type corresponding to the multiple target candidate entities and description information corresponding to the multiple target candidate entities; the multiple target candidate entities corresponding to the target entity label are subjected to disambiguation processing based on the text information, the first entity information corresponding to the multiple target candidate entities and the video feature information, to obtain a target entity corresponding to the target entity label; and the target entity corresponding to each entity label in the at least one entity label is determined as a video label of the target video. It should be understood that the multiple target candidate entities are first determined by using a statistical method, and then the entity disambiguation is performed based on a deep matching model, so that the disambiguated candidate entities are more matched with the feature information and the video content information of the historical video published by the user, thereby greatly enhancing the entity disambiguation effect and improving the accuracy of video label setting.
[0174] Based on the video processing method described above, an embodiment of the present application provides a video processing device. Please refer to Figure 11 FIG. 1 is a structural schematic diagram of a video processing device provided by an embodiment of the present application. The video processing device 1100 can run the following units:
[0175] The recognition unit 1101 is configured to perform entity recognition on a target video to determine at least one entity label corresponding to the target video.
[0176] The determination unit 1102 is configured to determine a target entity label in which multiple candidate entities exist in the at least one entity label.
[0177] The acquisition unit 1103 is configured to acquire video feature information of an object publishing the target video, the video feature information being associated with a video feature of a historical video published by the object.
[0178] The disambiguation unit 1104 is configured to perform disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the video feature information, to obtain a target entity corresponding to the target entity label, the target entity including one or more of the plurality of candidate entities.
[0179] The determination unit 1102 is further configured to determine the target entity corresponding to each entity label in the at least one entity label as a video label of the target video.
[0180] In an embodiment, when the disambiguation unit 1104 performs disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the video feature information, to obtain a target entity corresponding to the target entity label, the disambiguation unit 1104 is specifically configured to: determine a relevance between each candidate entity in the plurality of candidate entities and the video label of the historical video; determine a plurality of target candidate entities based on the relevance of each candidate entity; obtain text information corresponding to the target video and first entity information corresponding to the plurality of target candidate entities, the first entity information including one or more of an entity type corresponding to the plurality of target candidate entities and description information corresponding to the plurality of target candidate entities; and perform disambiguation processing on the plurality of target candidate entities corresponding to the target entity label based on the text information, the first entity information corresponding to the plurality of target candidate entities, and the video feature information, to obtain the target entity corresponding to the target entity label.
[0181] In an embodiment, when the disambiguation unit 1104 performs disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the video feature information, to obtain a target entity corresponding to the target entity label, the disambiguation unit 1104 is specifically configured to: determine a relevance between each candidate entity in the plurality of candidate entities and the video label of the historical video; and perform disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the relevance of each candidate entity, to obtain the target entity corresponding to the target entity label, the relevance between the target entity and the video label of the historical video being greater than the relevance between entities other than the target entity in the plurality of candidate entities and the video label of the historical video.
[0182] In an embodiment, when the disambiguation unit 1104 determines the relevance between each candidate entity in the plurality of candidate entities and the video label of the historical video, the disambiguation unit 1104 is specifically configured to: obtain a first matching degree of an entity type corresponding to each candidate entity in the plurality of candidate entities and the video type, and a first association degree between the each candidate entity and the video label of the historical video; and determine the relevance between the each candidate entity and the video label of the historical video based on the first matching degree and the first association degree.
[0183] In an embodiment, the disambiguation unit 1104, when obtaining the first matching degree of the entity type corresponding to each candidate entity in the plurality of candidate entities and the video type, is specifically configured to: obtain a play number of a video in the historical video whose video type matches the entity type corresponding to the each candidate entity, and a total play number of the historical video; and determine the first matching degree of the entity type corresponding to the each candidate entity conforming to the type corresponding to the video label of the historical video based on the play number and the total play number.
[0184] In an embodiment, the disambiguation unit 1104, when obtaining the first association degree between the each candidate entity and the video label of the historical video, is specifically configured to: obtain a second association degree between the entity label corresponding to the each candidate entity and the video label of the historical video, and a usage degree of the video label of the historical video; wherein the usage degree of the video label of the historical video is determined by a number of times that the video label of the historical video is labeled; and determine the first association degree between the each candidate entity and the video label of the historical video based on the two association degrees and the usage degree.
[0185] In an embodiment, the disambiguation unit 1104, when obtaining the second association degree between the entity label corresponding to the each candidate entity and the video label of the historical video, is specifically configured to: obtain a first labeling number of the entity label corresponding to the each candidate entity and the video label of the historical video being labeled on the same video, and a first total labeling number of the entity label corresponding to the each candidate entity and the video label of the historical video being labeled; and determine the second association degree between the entity label corresponding to the each candidate entity and the video label of the historical video based on the first labeling number and the first total labeling number.
[0186] In an embodiment, the disambiguation unit 1104, when obtaining the usage degree of the video label of the historical video, is specifically configured to: obtain a second labeling number of the video label of the historical video being labeled in the historical video, and a second total labeling number of the video label of the historical video being labeled; and determine the usage degree of the video label of the historical video based on the second labeling number and the second total labeling number.
[0187] In an embodiment, the disambiguation unit 1104, when disambiguating the plurality of candidate entities corresponding to the target entity label based on the video feature information to obtain the target entity corresponding to the target entity label, is specifically configured to: obtain text information corresponding to the target video; obtain second entity information corresponding to the plurality of candidate entities, the second entity information including one or more of entity types corresponding to the plurality of candidate entities and description information corresponding to the plurality of candidate entities; disambiguate the plurality of candidate entities corresponding to the target entity label based on the text information, the second entity information corresponding to the plurality of candidate entities, and the video feature information to obtain the target entity corresponding to the target entity label.
[0188] In an embodiment, the disambiguation unit 1104, when disambiguating the plurality of candidate entities corresponding to the target entity label based on the text information, the second entity information corresponding to the plurality of candidate entities, and the video feature information to obtain the target entity corresponding to the target entity label, is specifically configured to: call a deep matching model to process the text information, the second entity information corresponding to the plurality of candidate entities, and the video feature information to obtain a context feature vector of the text information, a feature vector of the second entity information corresponding to the plurality of candidate entities, and a feature vector of the video feature information; perform self-attention calculation on the context feature vector of the text information, the feature vector of the second entity information corresponding to the plurality of candidate entities, and the feature vector of the video feature information to obtain a second matching degree of each candidate entity in the plurality of candidate entities with the historical video; and disambiguate the plurality of candidate entities corresponding to the target entity label based on the second matching degree of each candidate entity to obtain the target entity corresponding to the target entity label, the second matching degree of the target entity with the historical video being greater than the second matching degree of each entity other than the target entity in the plurality of candidate entities with the historical video.
[0189] In an embodiment, the disambiguation unit 1104, when disambiguating the plurality of candidate entities corresponding to the target entity label based on the video feature information to obtain the target entity corresponding to the target entity label, is specifically configured to: determine a relevance between each candidate entity in the plurality of candidate entities and a video label of the historical video; obtain text information corresponding to the target video, and second entity information corresponding to each candidate entity, the second entity information including one or more of entity types corresponding to each candidate entity and description information corresponding to each candidate entity; determine a second matching degree of each candidate entity based on the text information, the second entity information corresponding to each candidate entity, and the video feature information; and disambiguate the plurality of candidate entities corresponding to the target entity label based on the relevance of each candidate entity and the second matching degree of each candidate entity to obtain the target entity corresponding to the target entity label.
[0190] In an embodiment, the acquisition unit 1103, in acquiring the video feature information of the object publishing the target video, is specifically configured to: acquire historical videos published by the object within a preset time period; perform analysis and processing on the historical videos to obtain historical video data information, the historical video data information including a video type of the historical videos and video labels of the historical videos; and determine the historical video data information as the video feature information of the object.
[0191] In summary, the target video is subjected to entity recognition to determine at least one entity label corresponding to the target video; a target entity label with multiple candidate entities is determined in the at least one entity label; video feature information of an object publishing the target video is acquired, the video feature information being associated with video features of historical videos published by the object; the multiple candidate entities corresponding to the target entity label are subjected to disambiguation processing based on the video feature information to obtain a target entity corresponding to the target entity label, the target entity including one or more of the multiple candidate entities; and the target entity corresponding to each entity label in the at least one entity label is determined as a video label of the target video. It should be understood that the video feature information of the object publishing the target video is used for entity disambiguation, so that the disambiguated candidate entity is more matched with the feature information of the historical videos of the publishing user, thereby enhancing the entity disambiguation effect and improving the accuracy of video label setting.
[0192] Based on the above-mentioned embodiments of the video processing method and the video processing device, the present embodiment provides a computer device, which corresponds to the aforementioned server. Please refer to Figure 12 is a structural schematic diagram of a computer device provided by the present embodiment, which can at least include a processor 1201, a communication interface 1202, and a computer storage medium 1203. The processor 1201, the communication interface 1202, and the computer storage medium 1203 can be connected through a bus or other means.
[0193] The computer storage medium 1203 can be stored in the memory 1204 of the computer device 1200, and is used to store a computer program including program instructions, and the processor 1201 is configured to execute the program instructions stored in the computer storage medium 1203. The processor 1201 (also called CPU (Central Processing Unit, Central Processing Unit)) is the computing core and control core of the computer device 1200, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing:
[0194] The target video is subjected to entity recognition to determine at least one entity label corresponding to the target video; a target entity label in which a plurality of candidate entities exist is determined in the at least one entity label; video feature information of an object publishing the target video is acquired, the video feature information being associated with a video feature of a historical video published by the object; the plurality of candidate entities corresponding to the target entity label are disambiguated based on the video feature information to obtain a target entity corresponding to the target entity label, the target entity including one or more of the plurality of candidate entities; and the target entity corresponding to each entity label in the at least one entity label is determined as a video label of the target video.
[0195] In one embodiment, when the plurality of candidate entities corresponding to the target entity label are disambiguated based on the video feature information to obtain the target entity corresponding to the target entity label, the processor 1201 is specifically configured to: determine a relevance between each candidate entity in the plurality of candidate entities and a video label of the historical video; and determine a plurality of target candidate entities based on the relevance of each candidate entity.
[0196] In one embodiment, when the plurality of candidate entities corresponding to the target entity label are disambiguated based on the video feature information to obtain the target entity corresponding to the target entity label, the processor 1201 is specifically configured to: determine a relevance between each candidate entity in the plurality of candidate entities and a video label of the historical video; and disambiguate the plurality of candidate entities corresponding to the target entity label based on the relevance of each candidate entity to obtain the target entity corresponding to the target entity label, the relevance between the target entity and the video label of the historical video being greater than the relevance between entities other than the target entity in the plurality of candidate entities and the video label of the historical video.
[0197] In one embodiment, when the relevance between each candidate entity in the plurality of candidate entities and the video label of the historical video is determined, the processor 1201 is specifically configured to: acquire a first matching degree of an entity type corresponding to each candidate entity in the plurality of candidate entities and the video type, and a first association degree between the each candidate entity and the video label of the historical video; and determine the relevance between the each candidate entity and the video label of the historical video based on the first matching degree and the first association degree.
[0198] In an embodiment, the processor 1201, in obtaining the first matching degree of the entity type corresponding to each candidate entity in the plurality of candidate entities and the video type, is specifically configured to: obtain a play number of a video in the historical video whose video type matches the entity type corresponding to each candidate entity, and a total play number of the historical video; and determine the first matching degree of the entity type corresponding to each candidate entity meeting the type corresponding to the video label of the historical video based on the play number and the total play number.
[0199] In an embodiment, the processor 1201, in obtaining the first association degree between each candidate entity and the video label of the historical video, is specifically configured to: obtain a second association degree between the entity label corresponding to each candidate entity and the video label of the historical video, and a usage degree of the video label of the historical video; wherein the usage degree of the video label of the historical video is determined by a number of times that the video label of the historical video is labeled; and determine the first association degree between each candidate entity and the video label of the historical video based on the two association degrees and the usage degree.
[0200] In an embodiment, the processor 1201, in obtaining the second association degree between the entity label corresponding to each candidate entity and the video label of the historical video, is specifically configured to: obtain a first labeling number of the entity label corresponding to each candidate entity and the video label of the historical video being labeled on the same video, and a first total labeling number of the entity label corresponding to each candidate entity and the video label of the historical video being labeled; and determine the second association degree between the entity label corresponding to each candidate entity and the video label of the historical video based on the first labeling number and the first total labeling number.
[0201] In an embodiment, the processor 1201, in obtaining the usage degree of the video label of the historical video, is specifically configured to: obtain a second labeling number of the video label of the historical video being labeled in the historical video, and a second total labeling number of the video label of the historical video being labeled; and determine the usage degree of the video label of the historical video based on the second labeling number and the second total labeling number.
[0202] In an embodiment, the processor 1201, when performing disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the video feature information to obtain the target entity corresponding to the target entity label, is specifically configured to: obtain text information corresponding to the target video; obtain second entity information corresponding to the plurality of candidate entities, the second entity information including one or more of entity types corresponding to the plurality of candidate entities and description information corresponding to the plurality of candidate entities; perform disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the text information, the second entity information corresponding to the plurality of candidate entities, and the video feature information to obtain the target entity corresponding to the target entity label.
[0203] In an embodiment, the processor 1201, when performing disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the text information, the second entity information corresponding to the plurality of candidate entities, and the video feature information to obtain the target entity corresponding to the target entity label, is specifically configured to: call a deep matching model to process the text information, the second entity information corresponding to the plurality of candidate entities, and the video feature information to obtain a context feature vector of the text information, a feature vector of the second entity information corresponding to the plurality of candidate entities, and a feature vector of the video feature information; perform self-attention calculation on the context feature vector of the text information, the feature vector of the second entity information corresponding to the plurality of candidate entities, and the feature vector of the video feature information to obtain a second matching degree of each candidate entity in the plurality of candidate entities with the historical video; and perform disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the second matching degree of each candidate entity to obtain the target entity corresponding to the target entity label, the second matching degree of the target entity with the historical video being greater than the second matching degree of each entity in the plurality of candidate entities other than the target entity with the historical video.
[0204] In an embodiment, the processor 1201, when performing disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the video feature information to obtain the target entity corresponding to the target entity label, is specifically configured to: determine a relevance between each candidate entity in the plurality of candidate entities and a video label of the historical video; obtain text information corresponding to the target video, and second entity information corresponding to each candidate entity, the second entity information including one or more of entity types corresponding to each candidate entity and description information corresponding to each candidate entity; determine a second matching degree of each candidate entity based on the text information, the second entity information corresponding to each candidate entity, and the video feature information; and perform disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the relevance of each candidate entity and the second matching degree of each candidate entity to obtain the target entity corresponding to the target entity label.
[0205] In an embodiment, the processor 1201, when acquiring the video feature information of the object publishing the target video, specifically is configured to: acquire historical videos published by the object within a preset time period; analyze and process the historical videos to obtain historical video data information, the historical video data information including a video type of the historical videos and a video label of the historical videos; and determine the historical video data information as the video feature information of the object.
[0206] In summary, the target video is subjected to entity recognition to determine at least one entity label corresponding to the target video; a target entity label with multiple candidate entities is determined in the at least one entity label; video feature information of an object publishing the target video is acquired, the video feature information being associated with a video feature of a historical video published by the object; the multiple candidate entities corresponding to the target entity label are subjected to disambiguation processing based on the video feature information to obtain a target entity corresponding to the target entity label, the target entity including one or more of the multiple candidate entities; and the target entity corresponding to each entity label in the at least one entity label is determined as a video label of the target video. It should be understood that the video feature information of the object publishing the target video is used for entity disambiguation, so that the disambiguated candidate entity is more matched with the feature information of the historical video of the publishing user, thereby enhancing the entity disambiguation effect and improving the accuracy of video label setting.
[0207] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments. The technical solutions of the present application essentially or say the parts that make contributions to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes instructions for causing a computer device (which can be a personal computer, a server or a network device, etc., specifically a processor in the computer device) to execute all or part of the steps of the above-mentioned method of each embodiment of the present application. Among them, the storage medium can include: U disk, mobile hard disk, magnetic disk, optical disk, read-only memory (English: Read-Only Memory, abbreviated: ROM) or random access memory (English: Random Access Memory, abbreviated: RAM) and various program code storage media.
[0208] Those skilled in the art can be aware that units and steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0209] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable devices. Computer instructions can be stored in a computer storage medium or transmitted through a computer storage medium. Computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner. The computer storage medium can be any available medium that the computer can access or a data storage device such as a server, data center, etc. containing one or more available media sets. The available media can be magnetic media (for example, floppy disk, hard disk, magnetic tape), optical media (for example, DVD), or semiconductor media (for example, solid state disk (SSD)) and the like.
[0210] The above description is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of video processing, the method comprising: The method comprises: performing entity recognition on a target video to determine at least one entity label corresponding to the target video; the determination manner of the at least one entity label comprises: obtaining text information corresponding to the target video; calling an entity recognition model to perform recognition processing on the text information to obtain intermediate result information corresponding to each text segment in the text information, the intermediate result information comprising multiple types of information including entity region information, interval length feature information, interval context feature information and global context feature information; and calling the entity recognition model to perform fusion processing on each information included in the intermediate result information to obtain the at least one entity label corresponding to the target video; determining, in the at least one entity label, a target entity label in which multiple candidate entities exist; obtaining video feature information of an object publishing the target video, the video feature information being associated with video features of historical videos published by the object; performing disambiguation processing on the multiple candidate entities corresponding to the target entity label based on the video feature information to obtain a target entity corresponding to the target entity label, the target entity comprising one or more of the multiple candidate entities; the multiple candidate entities corresponding to the target entity label have the same label name, and different candidate entities in the multiple candidate entities correspond to different entity types; the disambiguation processing refers to determining, in the multiple candidate entities, a candidate entity that is most matched with the target entity label for the target video; determining the target entity corresponding to each entity label in the at least one entity label as a video label of the target video.
2. The method of claim 1, wherein, The disambiguation processing on the multiple candidate entities corresponding to the target entity label based on the video feature information to obtain a target entity corresponding to the target entity label comprises: determining a relevance between each candidate entity in the multiple candidate entities and a video label of the historical video; determining multiple target candidate entities based on the relevance of each candidate entity; obtaining text information corresponding to the target video and first entity information corresponding to the multiple target candidate entities, the first entity information comprising one or more of an entity type corresponding to the multiple target candidate entities and description information corresponding to the multiple target candidate entities; performing disambiguation processing on the multiple target candidate entities corresponding to the target entity label based on the text information, the first entity information corresponding to the multiple target candidate entities and the video feature information to obtain a target entity corresponding to the target entity label.
3. The method according to claim 1 or 2, characterized in that, The disambiguation processing on the multiple candidate entities corresponding to the target entity label based on the video feature information to obtain a target entity corresponding to the target entity label comprises: determining a relevance between each candidate entity in the multiple candidate entities and a video label of the historical video; The target entity corresponding to the target entity label is obtained by disambiguating the multiple candidate entities corresponding to the target entity label based on the relevance of each candidate entity, and the relevance between the target entity and the video label of the historical video is greater than the relevance between each entity in the multiple candidate entities except the target entity and the video label of the historical video.
4. The method of claim 3, wherein, The relevance between each candidate entity in the multiple candidate entities and the video label of the historical video is determined, including: A first matching degree of an entity type corresponding to each candidate entity in the multiple candidate entities and a video type, and a first association degree between the each candidate entity and the video label of the historical video are obtained; The relevance between the each candidate entity and the video label of the historical video is determined based on the first matching degree and the first association degree.
5. The method of claim 4, wherein, The first matching degree of the entity type corresponding to each candidate entity in the multiple candidate entities and the video type is obtained, including: A play frequency of a video in the historical video whose video type matches the entity type corresponding to the each candidate entity, and a total play frequency of the historical video are obtained; The first matching degree of the entity type corresponding to each candidate entity in the multiple candidate entities and the video type is determined based on the play frequency and the total play frequency.
6. The method of claim 4, wherein, The first association degree between the each candidate entity and the video label of the historical video is obtained, including: A second association degree between an entity label corresponding to the each candidate entity and the video label of the historical video, and a usage degree of the video label of the historical video are obtained; wherein the usage degree of the video label of the historical video is determined by a frequency of the video label of the historical video being labeled; The first association degree between the each candidate entity and the video label of the historical video is determined based on the second association degree and the usage degree.
7. The method of claim 6, wherein, The second association degree between the entity label corresponding to the each candidate entity and the video label of the historical video is obtained, including: A first labeling frequency of the entity label corresponding to the each candidate entity and the video label of the historical video being labeled on the same video, and a first total labeling frequency of the entity label corresponding to the each candidate entity and the video label of the historical video being labeled are obtained; The second association degree between the entity label corresponding to the each candidate entity and the video label of the historical video is determined based on the first labeling frequency and the first total labeling frequency.
8. The method of claim 6, wherein, The usage degree of the video label of the historical video is obtained, including: A second labeling frequency of the video label of the historical video being labeled in the historical video, and a second total labeling frequency of the video label of the historical video being labeled are obtained; The usage degree of the video label of the historical video is determined based on the second labeling frequency and the second total labeling frequency.
9. The method of claim 1 or 2, wherein, The target entity corresponding to the target entity label is obtained by disambiguating the multiple candidate entities corresponding to the target entity label based on the video feature information, including: Text information corresponding to the target video is obtained; obtaining second entity information corresponding to the plurality of candidate entities, the second entity information including one or more of entity types corresponding to the plurality of candidate entities and description information corresponding to the plurality of candidate entities; performing disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the text information, the second entity information corresponding to the plurality of candidate entities, and the video feature information, to obtain a target entity corresponding to the target entity label.
10. The method of claim 9, wherein, The disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the text information, the second entity information corresponding to the plurality of candidate entities, and the video feature information, to obtain a target entity corresponding to the target entity label, includes: calling a deep matching model to process the text information, the second entity information corresponding to the plurality of candidate entities, and the video feature information, to obtain a context feature vector of the text information, a feature vector of the second entity information corresponding to the plurality of candidate entities, and a feature vector of the video feature information; performing self-attention calculation on the context feature vector of the text information, the feature vector of the second entity information corresponding to the plurality of candidate entities, and the feature vector of the video feature information, to obtain a second matching degree of each candidate entity in the plurality of candidate entities with the historical video; performing disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the second matching degree of each candidate entity, to obtain a target entity corresponding to the target entity label, the target entity having a second matching degree with the historical video greater than a second matching degree of entities other than the target entity in the plurality of candidate entities with the historical video.
11. The method of claim 1 or 2, wherein, The disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the video feature information, to obtain a target entity corresponding to the target entity label, includes: determining a relevance between each candidate entity in the plurality of candidate entities and a video label of the historical video; obtaining text information corresponding to the target video, and second entity information corresponding to each candidate entity, the second entity information including one or more of entity types corresponding to each candidate entity and description information corresponding to each candidate entity; determining a second matching degree of each candidate entity based on the text information, the second entity information corresponding to each candidate entity, and the video feature information; performing disambiguation processing on the plurality of candidate entities corresponding to the target entity label based on the relevance of each candidate entity and the second matching degree of each candidate entity, to obtain a target entity corresponding to the target entity label.
12. The method of claim 1 or 2, wherein, The obtaining of the video feature information of the object publishing the target video includes: obtaining historical videos published by the object within a preset time period; performing analysis and processing on the historical videos to obtain historical video data information, the historical video data information including a video type of the historical videos and video labels of the historical videos; determining the historical video data information as the video feature information of the object.
13. A video processing apparatus, comprising: The apparatus includes: The recognition unit is configured to perform entity recognition on a target video to determine at least one entity label corresponding to the target video. The at least one entity label is determined in the following manner: obtaining text information corresponding to the target video; calling an entity recognition model to perform recognition processing on the text information to obtain intermediate result information corresponding to each text segment of the text information. The intermediate result information includes multiple types of information, such as entity region information, interval length feature information, interval context feature information, and global context feature information. The entity recognition model is called to perform fusion processing on each piece of information included in the intermediate result information to obtain the at least one entity label corresponding to the target video. The determination unit is configured to determine, among the at least one entity label, a target entity label that includes multiple candidate entities. The obtaining unit is configured to obtain video feature information of an object that publishes the target video. The video feature information is associated with video features of historical videos published by the object. The disambiguation unit is configured to perform disambiguation processing on the multiple candidate entities corresponding to the target entity label based on the video feature information to obtain a target entity corresponding to the target entity label. The target entity includes one or more of the multiple candidate entities. The multiple candidate entities corresponding to the target entity label have the same label name, and different candidate entities in the multiple candidate entities correspond to different entity types. The disambiguation processing refers to determining, among the multiple candidate entities, a candidate entity that is most matched with the target entity label for the target video. The determination unit is further configured to determine, as a video label of the target video, a target entity corresponding to each entity label in the at least one entity label.
14. A computer device, comprising: The computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the video processing method of any one of claims 1-12.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more computer programs. The one or more computer programs are adapted to be loaded and executed by the processor to perform the video processing method of any one of claims 1-12.
16. A computer program product, characterised in that, The computer program product includes computer instructions executed by the processor to implement the video processing method of any one of claims 1-12.
Citation Information
Patent Citations
Method and device for processing short video data, computer equipment and storage medium
CN111010619A
Method for determining video label, server and storage medium
CN111274442A
Video tag extension method and device, computer equipment and storage medium
CN111368141A