Category determination method, apparatus, device, storage medium, and product

By identifying entities in video description information and introducing knowledge extension information, and combining knowledge graphs and entity recognition models, the problem of low accuracy in video category identification is solved, and high accuracy in determining video categories is achieved.

CN114676285BActive Publication Date: 2025-12-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210297726.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-24
Publication Date
2025-12-30
Estimated Expiration
2042-03-24

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of category identification based on video title information is low, mainly because the title information is short and limited.

Method used

By acquiring entities from the video description information, expanding entity knowledge information, and combining knowledge graphs and entity recognition models, the video category can be determined.

Benefits of technology

It improves the accuracy of video categorization, especially in fine-grained classification, enabling more accurate determination of secondary video categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114676285B_ABST
    Figure CN114676285B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a category determination method, device, equipment, storage medium and product, belongs to the computer technology field, and can be applied to video classification, artificial intelligence and vehicle-mounted scenes in computer technology. The method comprises the following steps: obtaining first video description information of a target video, wherein the first video description information is used for describing video content of the target video; performing entity identification on the first video description information to obtain a first entity in the first video description information; obtaining first knowledge expansion information of the first entity; and determining a category of the target video based on the first video description information and the first knowledge expansion information. According to the scheme, the information of the entity is more abundant by introducing the knowledge expansion information, and the entity can be more accurately understood, so that the video category can be more accurately determined based on the knowledge expansion information and the video description information, and the accuracy of the video category is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, storage medium, and product for classifying data. Background Technology

[0002] With the continuous development of computer and internet technologies, the amount of video data disseminated online is increasing. To facilitate users' retrieval of desired video data, retrieval systems categorize videos based on their titles to determine video types, and then provide users with search data based on these categories.

[0003] Currently, most methods use category recognition models to classify video content based on title information. However, title information is generally short and contains very limited information, resulting in low accuracy in category classification obtained in this way. Summary of the Invention

[0004] This application provides a method, apparatus, device, storage medium, and product for determining video categories, which can improve the accuracy of video categorization. The technical solution is as follows:

[0005] On the one hand, a method for determining a category is provided, the method comprising:

[0006] Obtain first video description information of the target video, wherein the first video description information is used to describe the video content of the target video;

[0007] Entity recognition is performed on the first video description information to obtain the first entity in the first video description information;

[0008] Obtain the first knowledge extension information of the first entity;

[0009] Based on the first video description information and the first knowledge extension information, the category of the target video is determined.

[0010] On the other hand, a category determination apparatus is provided, the apparatus comprising:

[0011] The first acquisition module is used to acquire first video description information of the target video, wherein the first video description information is used to describe the video content of the target video.

[0012] The recognition module is used to perform entity recognition on the first video description information to obtain the first entity in the first video description information;

[0013] The second acquisition module is used to acquire the first knowledge extension information of the first entity;

[0014] The determination module is used to determine the category of the target video based on the first video description information and the first knowledge extension information.

[0015] In some embodiments, the category of the target video includes a primary category and a secondary category; the determining module includes:

[0016] The first determining unit is configured to determine the primary category of the target video based on the first video description information;

[0017] The second determining unit is used to determine the secondary category of the target video based on the first video description information and the first knowledge extension information.

[0018] In some embodiments, the first determining unit is configured to determine the descriptive semantic features of the first video description information, and determine the primary category of the target video based on the descriptive semantic features;

[0019] The second determining unit is used to determine the extended semantic features of the first knowledge extended information, fuse the descriptive semantic features and the extended semantic features to obtain fused features, and determine the secondary category of the target video based on the fused features.

[0020] In some embodiments, the category of the target video includes a primary category and a secondary category; the determining module is configured to input the first video description information and the first knowledge extension information into the category recognition model, and output the primary category and secondary category of the target video, wherein the category recognition model is configured to determine the primary category of the target video based on the first video description information, and determine the secondary category of the target video based on the first video description information and the first knowledge extension information.

[0021] In some embodiments, the apparatus further includes:

[0022] The third acquisition module is used to acquire sample data, which includes the second video description information of the sample video, the primary category of the sample video, and the secondary category of the sample video.

[0023] The second acquisition module is further configured to acquire second knowledge extension information of the second entity, wherein the second entity is an entity in the second video description information;

[0024] The determining module is further configured to perform category recognition on the second video description information and the second knowledge extension information through the category recognition model before training, so as to obtain the predicted primary category and predicted secondary category of the sample video;

[0025] The training module is used to train the pre-training category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category, and the sample secondary category, to obtain the category recognition model.

[0026] In some embodiments, the training module includes:

[0027] The first determining unit is configured to determine a first loss value based on the predicted primary category and the sample primary category;

[0028] The second determining unit is used to determine a second loss value based on the predicted secondary category and the sample secondary category;

[0029] A fusion unit is used to fuse the first loss value and the second loss value to obtain a third loss value;

[0030] The training unit is used to train the pre-training category recognition model based on the third loss value to obtain the category recognition model.

[0031] In some embodiments, the training module is further configured to determine a first probability and a second probability, and the category recognition model is configured to determine the probability that the target video belongs to multiple primary categories and the probability that the target video belongs to multiple secondary categories, wherein the first probability is the probability of at least one primary category determined by the category recognition model, and the second probability is the probability of at least one secondary category determined by the category recognition model.

[0032] The training module is used to train the pre-training category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category, the sample secondary category, the first probability, and the second probability, to obtain the category recognition model.

[0033] In some embodiments, the training module is configured to perform any of the following:

[0034] The probability of the correct primary category determined by the category recognition model is defined as the first probability, and the probability of the correct secondary category determined by the category recognition model is defined as the second probability. The correct primary category is the same as the primary category of the sample, and the correct secondary category is the same as the secondary category of the sample.

[0035] The probability of predicting the first-level category is determined as the first probability, and the probability of predicting the second-level category is determined as the second probability;

[0036] The probability of predicting the first-level category is determined as the first probability, and the probabilities of the multiple second-level categories determined by the category recognition model are determined as the second probability;

[0037] The probabilities of multiple primary categories determined by the category recognition model are defined as the first probability, and the probabilities of multiple secondary categories determined by the category recognition model are defined as the second probability.

[0038] In some embodiments, the training module includes:

[0039] The first determining unit is configured to determine a first loss value based on the predicted primary category and the sample primary category;

[0040] The second determining unit is used to determine a second loss value based on the predicted secondary category and the sample secondary category;

[0041] The third determining unit is used to determine a fourth loss value based on the first probability and the second probability;

[0042] A fusion unit is used to fuse the first loss value, the second loss value, and the fourth loss value to obtain a fifth loss value;

[0043] The training unit is used to train the category recognition model based on the fifth loss value.

[0044] In some embodiments, the third determining unit is configured to determine the fourth loss value based on the difference between the first probability and the second probability when the first probability is less than the second probability, wherein the fourth loss value is positively correlated with the difference; and to determine the fourth loss value as 0 when the first probability is not less than the second probability.

[0045] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory being used to store at least one computer program, the at least one computer program being loaded and executed by the processor to implement the category determination method in the embodiments of this application.

[0046] On the other hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to implement the category determination method as described in the embodiments of this application.

[0047] On the other hand, a computer program product or computer program is provided, which includes computer program code stored in a computer-readable storage medium, wherein a processor of a computer device reads the computer program code from the computer-readable storage medium, and executes the computer program code, causing the computer device to perform the category determination method provided in various alternative implementations of the above aspects.

[0048] This application provides a method, apparatus, device, storage medium, and product for classifying videos. When classifying videos based on video description information, it introduces extended knowledge information about the entities within the video description information by identifying those entities. Since additional information about the entity can be obtained based on this extended knowledge information, the entity's information is enriched, and it can be understood more accurately. Therefore, when determining the video category based on this extended knowledge information and the video description information, the video category can be determined more accurately, improving the accuracy of video category classification. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a schematic diagram of the implementation environment of a category determination method provided in an embodiment of this application;

[0051] Figure 2 This is a flowchart of a category determination method provided in an embodiment of this application;

[0052] Figure 3 This is a flowchart of a category determination method provided in an embodiment of this application;

[0053] Figure 4 This is a schematic diagram illustrating the process of obtaining extended knowledge information from a knowledge graph, as provided in an embodiment of this application.

[0054] Figure 5 This is a flowchart of a category determination method provided in an embodiment of this application;

[0055] Figure 6 This is a schematic diagram of a category recognition model provided in an embodiment of this application;

[0056] Figure 7 This is a flowchart illustrating a training method for a category recognition model provided in an embodiment of this application;

[0057] Figure 8 This is a block diagram of a category determination device provided in an embodiment of this application;

[0058] Figure 9 This is a block diagram of another category determination device provided in the embodiments of this application;

[0059] Figure 10This is a structural block diagram of a terminal provided in an embodiment of this application;

[0060] Figure 11 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0062] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with essentially the same function. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on the quantity or execution order.

[0063] In this application, the term "at least one" means one or more, and "multiple" means two or more.

[0064] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the video description information, videos, sample data, etc. involved in this application were obtained with full authorization.

[0065] The following is an explanation of the terms used in this application.

[0066] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0067] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0068] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0069] The solutions provided in this application involve technologies such as natural language processing in artificial intelligence, and are specifically illustrated through the following embodiments.

[0070] The following describes the implementation environment involved in this application:

[0071] The category determination method provided in this application can be executed by a computer device. In some embodiments, the computer device is a terminal. The terminal is a smartphone, tablet, laptop, desktop computer, intelligent voice interaction device, smart home appliance, vehicle terminal, etc., but is not limited to these. In some embodiments, the computer device is a server. The server can be an independent server, a server cluster of multiple physical servers, or a distributed system. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0072] In some embodiments, the computer device includes a terminal and a server. The following is a schematic diagram illustrating the implementation environment of the category determination method provided in this application, using a computer device including terminal 101 and server 102 as an example. See also... Figure 1 The implementation environment includes terminal 101 and server 102. Terminal 101 and server 102 can be connected directly or indirectly via wired or wireless communication, which is not limited herein.

[0073] In some embodiments, server 102 primarily undertakes computing tasks, while terminal 101 undertakes secondary computing tasks; or, server 102 provides secondary computing services, while terminal 101 undertakes primary computing tasks; or, server 102 and terminal 101 collaborate on computing using a distributed computing architecture.

[0074] In some embodiments, terminal 101 uploads a video and its title information. Server 102 receives the video and title information uploaded by terminal 101 and determines the video's category based on the title information. Subsequently, server 102 can manage the video based on its category. For example, if the video is shared by terminal 101, server 102 can recommend the video to other terminals 101 based on its category; or, server 102 can provide accurate search results to terminal 101 when it searches for videos, based on the video's category.

[0075] The following describes the application scenarios of this application:

[0076] The category determination method provided in this application can be applied to video processing scenarios such as video classification and video retrieval.

[0077] For example, it can be applied to video classification scenarios.

[0078] When users upload videos, they need to set the category of the video. If the category determination method provided in this application is used, the category of the video can be automatically determined for the user without the need for manual setting, which reduces user operations and improves video upload efficiency.

[0079] For example, it can be applied to video retrieval scenarios.

[0080] The video retrieval system categorizes multiple videos in a video library and then provides video search results to users based on the video categories. If the solution provided in this application is adopted, the introduction of extended knowledge information enriches the reference information for video classification, enabling more accurate determination of video categories, improving the accuracy of video category identification, and providing users with more accurate video search results.

[0081] It should be noted that the category determination method provided in this application embodiment can also be applied to other video processing scenarios, and this application embodiment does not limit it.

[0082] Figure 2 This is a flowchart of a category determination method provided according to an embodiment of this application. See also... Figure 2 In this embodiment of the application, a computer device is used as the execution subject for illustrative purposes. The method for determining the category includes the following steps:

[0083] 201. The computer device acquires first video description information of the target video, the first video description information being used to describe the video content of the target video.

[0084] In this embodiment of the application, the target video can be any video. For example, the target video can be a video locally on the computer device; or it can be a video obtained by the computer device from another device; or it can be a video shot by the computer device; or it can be a video uploaded by any user; or it can be any video in the video library, etc. This embodiment of the application does not limit the target video.

[0085] The first video description information describes the video content of the target video. This video content can be the main theme of the target video, such as the video title; it can also be the content of any video frame in the target video, such as the first frame; it can also be the audio content of the target video; for example, it can be subtitle information. It should be noted that the first video description information can be any one or a combination of several video description information shown in the embodiments of this application, or it can be other video description information. The embodiments of this application do not limit the first video description information.

[0086] 202. The computer device performs entity recognition on the first video description information to obtain the first entity in the first video description information.

[0087] An entity can be considered an instance of a concept. For example, "time" is a concept, and "Mid-Autumn Festival" is an instance of "time," therefore "Mid-Autumn Festival" is a time entity. Similarly, "location" is a concept, and "Scenic Spot A" is an instance of "location," therefore "Scenic Spot A" is an instance of "location."

[0088] The computer device performs entity recognition on the first video description information, which means identifying the entities in the first video description information.

[0089] 203. Computer devices acquire the first knowledge extension information of the first entity.

[0090] In this embodiment, the first knowledge extension information is knowledge information related to the first entity. This knowledge information includes information other than the information related to the first entity in the first video description information, thus extending the knowledge information of the first entity. Therefore, it is called the knowledge extension information of the first entity. Taking the first video description information as title information as an example, title information is generally relatively short. Therefore, there is little information about the first entity in the title information. By obtaining knowledge information related to the first entity, the information of the first entity can be extended.

[0091] 204. The computer device determines the category of the target video based on the first video description information and the first knowledge extension information.

[0092] In this embodiment of the application, when determining the category of a video, not only the information of the video itself—video description information—is considered, but additional information—knowledge extension information—can also be introduced, making the information referenced in the category determination process richer, thereby determining a more accurate video category.

[0093] The category determination method provided in this application, when classifying videos based on video description information, introduces extended knowledge information about the entities in the video description information by identifying those entities. Since additional information about the entity can be obtained based on this extended knowledge information, the entity's information is enriched, and the entity can be understood more accurately. Therefore, when determining the video category based on this extended knowledge information and the video description information, the video category can be determined more accurately, improving the accuracy of video category determination.

[0094] Figure 3 This is a flowchart illustrating a category determination method provided in an embodiment of this application. The category determination method includes the following steps:

[0095] 301. The computer device acquires first video description information of the target video, the first video description information being used to describe the video content of the target video.

[0096] In some embodiments, the first video description information may be existing information that the computer device can directly obtain. For example, the first video description information may be the title information of a target video. The video library stores multiple videos and their title information, and the computer device can directly obtain the title information of a specific video from the video library.

[0097] In some embodiments, the first video description information may be information that does not currently exist and requires processing to obtain. In such cases, the computer device needs to obtain the first video description information through data processing. For example, the first video description information may be the subtitle information of the target video. The computer device obtains the audio data of the target video, performs speech recognition on the audio data, and obtains the subtitle information of the target video.

[0098] 302. The computer device performs entity recognition on the first video description information to obtain the first entity in the first video description information.

[0099] In this embodiment, any entity recognition method in the field of NLP can be used when performing entity recognition on the first video description information, and this embodiment does not limit the method. In some embodiments, the computer device can perform the task based on the NLP tool TextSmart (an entity recognition tool). In some embodiments, the computer device can also perform the task using mature NLP toolkits such as hanNLP (an NLP toolkit). Mature toolkits integrate a series of basic NLP functions such as word segmentation, part-of-speech tagging, and entity recognition. Therefore, based on mature NLP toolkits, entity recognition can be performed on the first video description information, and accurate recognition results can be obtained, that is, accurate first entities can be obtained.

[0100] 303. The computer device acquires the first knowledge extension information of the first entity.

[0101] In some embodiments, a computer device obtains first knowledge extension information of a first entity from a knowledge graph. The knowledge graph records multiple entities, types of multiple entities, attributes of multiple entities, and other objects associated with multiple entities. Optionally, the knowledge graph includes multiple nodes and edges between the nodes, where nodes represent entities, entity types, entity attributes, or other objects associated with entities, and edges between two nodes represent the relationship between those two nodes. The first knowledge extension information consists of nodes and edges in the knowledge graph related to the node representing the first entity.

[0102] Optionally, the computer device obtains the first extended information of the first entity from the knowledge graph, including: the computer device determines the first node representing the first entity from the knowledge graph, and extracts a subgraph of the target jump from the knowledge graph with the first node as the center, and determines the subgraph as the first extended information of the first entity.

[0103] The target jump can be any number of jumps, such as 2, 3, or 5. This application embodiment does not limit the target jump and can set it according to the data processing capability and accuracy requirements of the actual application scenario.

[0104] For example, if the target jump is 2 hops, in the knowledge graph, with the first node as the center, we extract multiple second nodes connected to the first node, as well as third nodes connected to these second nodes. It should be noted that when extracting multiple second nodes connected to the first node, the edges connecting the first and second nodes are also extracted.

[0105] For example, knowledge graphs Figure 4 As shown, it should be noted that Figure 4 Only a portion of the knowledge graph is shown. This knowledge graph includes the node "Game Name A," as well as type information for node "Game Name A," namely "Mini Program" and "Game," the platform "Application B" on which "Game Name A" resides, the node "Game Name C" connected to node "Game," and the nodes "Mobile Phone" and "Company D" connected to node "Game Name C." If the first entity is "Game Name A," extracting the first knowledge extension information from the knowledge graph yields the type information for node "Game Name A," namely "Mini Program" and "Game," and also the platform "Application B" on which "Game Name A" resides.

[0106] When a computer device determines the first node representing the first entity from the knowledge graph, it can use a keyword matching method or an entity chaining method in the field of NLP. This application does not limit the specific method used.

[0107] In some embodiments, the computer device may obtain first knowledge extension information of the first entity from the retrieval platform. For example, the computer device may use the first entity as retrieval data, perform a retrieval in the retrieval platform, and use the retrieval results as the first knowledge extension information of the first entity.

[0108] It should be noted that the embodiments of this application are only used as examples of the first knowledge extension information as a sub-graph of the knowledge graph and the retrieval results of the first entity, and are not limited to the first knowledge extension information. The first knowledge extension information can also be other information that can extend the knowledge information of the entities in the first video description information.

[0109] 304. The computer device determines the primary category of the target video based on the first video description information.

[0110] In this embodiment, the target video is categorized hierarchically into primary and secondary categories. Primary categories are coarse-grained, while secondary categories are more detailed. For example, primary categories include 44 subcategories such as sports, games, and entertainment. Secondary categories include 305 subcategories such as mini-games, mobile games, square dancing, street dance, and Latin dance. Since secondary categories are derived from each primary category, there is a hierarchical relationship between them. For instance, a primary category might be "dance," and its secondary categories could include various dance styles such as square dancing, street dance, and Latin dance.

[0111] Since the primary category is a coarse-grained category, it is not necessary to use a lot of information to determine the primary category of the target video. To improve the efficiency of category determination and reduce the amount of computation required for category determination, only the first video description information can be used to determine the primary category of the target video. In some embodiments, the computer device determines the primary category of the target video based on the first video description information, including: determining the descriptive semantic features of the first video description information, and determining the primary category of the target video based on the descriptive semantic features.

[0112] The semantic features described are used to describe the semantics of the first video description information. The semantic features can be represented by vectors or the like, but this application does not limit this.

[0113] In some embodiments, the computer device determines the descriptive semantic features of the first video description information, including: the computer device extracts features from the first video description information to obtain the descriptive semantic features. The computer device may employ any feature extraction method to extract features from the first video description information. For example, it may use any model for processing text data, such as the BERT model, to extract features from the first video description information; or it may use methods such as word vector conversion to extract features from the first video description information. This application embodiment does not limit the feature extraction method.

[0114] 305. The computer device determines the secondary category of the target video based on the first video description information and the first knowledge extension information.

[0115] Since secondary categories are more detailed categories, more information can be input when determining the secondary category of a target video to obtain more accurate results.

[0116] In this embodiment, first knowledge extension information is added to the input information to enrich the input information. In some embodiments, the computer device fuses multiple pieces of input information before classifying them. Optionally, the computer device determines the secondary category of the target video based on the first video description information and the first knowledge extension information, including: determining the extended semantic features of the first knowledge extension information, fusing the description semantic features and the extended semantic features to obtain fused features, and determining the secondary category of the target video based on the fused features.

[0117] The descriptive semantic features are calculated when the computer device determines the primary category of the target video. When determining the secondary category of the target video, the descriptive semantic features can be used directly without calculating the first video description information again.

[0118] The extended semantic features can be obtained using any feature extraction method. It should be noted that different feature extraction methods are used for knowledge extension information of different data types. For example, if the knowledge extension information is an image, any image feature extraction method is used; if the knowledge extension information is text, any text feature extraction method is used. This application does not limit this.

[0119] It should be noted that any fusion method can be used to fuse descriptive semantic features and extended semantic features in the embodiments of this application. The embodiments of this application are only illustrated by the following two methods. In some embodiments, the computer device fuses descriptive semantic features and extended semantic features to obtain fused features, including: the computer device extracts features from the descriptive semantic features and extended semantic features to obtain fused features. In some embodiments, the computer device fuses descriptive semantic features and extended semantic features to obtain fused features, including: the computer device concatenates the descriptive semantic features and extended semantic features to obtain fused features.

[0120] It should be noted that the embodiments of this application are merely illustrative examples illustrating the inclusion of primary and secondary categories in video classification, and do not limit the types of videos. In some embodiments, video classification only has secondary categories and no primary categories; that is, the computer device only performs a relatively detailed classification of videos, without performing a coarse-grained classification. In some embodiments, video classification includes primary, secondary, and tertiary categories, etc., and the method provided in the embodiments of this application can also be applied to determine the tertiary category of a video.

[0121] It should be noted that the method provided in this application embodiment can also be applied to determine the primary category of a video, and this application embodiment does not limit this application.

[0122] The category determination method provided in this application, when classifying videos based on video description information, introduces extended knowledge information about the entities in the video description information by identifying those entities. Since additional information about the entity can be obtained based on this extended knowledge information, the entity's information is enriched, and the entity can be understood more accurately. Therefore, when determining the video category based on this extended knowledge information and the video description information, the video category can be determined more accurately, improving the accuracy of video category determination.

[0123] Furthermore, this embodiment of the application only considers video description information when determining the primary category of a video. While ensuring the accuracy of the primary category, it reduces the amount of computation and improves the efficiency of category determination. When determining the secondary category of a video, it considers not only the video description information but also the knowledge extension information of the entities within the video description information. This allows the computer device to obtain more information when performing more detailed classification, thereby enabling it to determine the secondary category of the video more accurately.

[0124] It should be noted that, in the embodiments of this application, when the computer device determines the category of the target video based on the first video description information and the first knowledge extension information, this can be accomplished by a category recognition model. In some embodiments, the computer device determines the category of the target video based on the first video description information and the first knowledge extension information by: inputting the first video description information and the first knowledge extension information into the category recognition model, and outputting the primary category and secondary category of the target video. The category recognition model is used to determine the primary category of the target video based on the first video description information, and to determine the secondary category of the target video based on the first video description information and the first knowledge extension information. The following example illustrates the determination of the target video category using a category recognition model.

[0125] Figure 5 This is a flowchart illustrating a category determination method provided in an embodiment of this application. The category determination method includes the following steps:

[0126] 501. The computer device inputs the first video description information of the target video and the first knowledge extension information of the first entity in the first video description information into the category recognition model.

[0127] 502. The computer device extracts features from the first video description information using a category recognition model to obtain the descriptive semantic features of the first video description information.

[0128] In some embodiments, the category recognition model includes a first feature extraction layer. The computer device inputs first video description information into the first feature extraction layer and processes the first video description information through the first feature extraction layer. For example, the computer device extracts features from the first video description information through the category recognition model to obtain the descriptive semantic features of the first video description information, including: the computer device extracts features from the first video description information through the first feature extraction layer to obtain the descriptive semantic features of the first video description information.

[0129] The first feature extraction layer can be a network structure used to extract text features, such as BERT (a language representation model) or LSTM (Long Short-Term Memory). This application does not limit the first feature extraction layer.

[0130] For example, the first feature extraction layer is a BERT sub-model. After inputting the video title information into the BERT sub-model, a CLS_token is obtained. This CLS_token is the output of the BERT sub-model, and CLS_token = BERT(title), where BERT represents the BERT sub-model and title is the video title information. In this embodiment, the embedding information corresponding to the CLS_token in the BERT sub-model output information is used as a descriptive semantic feature.

[0131] In some embodiments, when the category recognition model extracts features from the first video description information, it can perform word segmentation or character segmentation on the first video description information. When the category recognition model extracts features from the word segmentation results or character segmentation results, it will extract features from the word segmentation results or character segmentation results based on the context information of the word segmentation results or character segmentation results, so as to obtain more accurate semantic features.

[0132] 503. The computer device extracts features from the first knowledge extension information of the first entity in the first video description information using a category recognition model to obtain extended semantic features.

[0133] In some embodiments, the category recognition model includes a second feature extraction layer. The computer device inputs first knowledge extension information into the second feature extraction layer and processes the first knowledge extension information through the second feature extraction. For example, the computer device extracts features from the first knowledge extension information of a first entity in the first video description information using the category recognition model to obtain extended semantic features, including: the computer device extracts features from the first knowledge extension information through the second feature extraction layer to obtain extended semantic features.

[0134] In some embodiments, the first knowledge extension information is graph information, and the second feature extraction layer is a network for processing images, such as GCN (Graph Convolutional Neural Networks), node2vec, etc.

[0135] In some embodiments, a Graph Convolutional Network (GCN) is used to extract features from the first knowledge extension information. This feature extraction process is essentially the encoding process of the first knowledge extension information. Taking the first knowledge extension information as a subgraph as an example, using a GCN to encode the subgraph can effectively embed the topological information of the graph in the semantic features of nodes and edges. The semantic features of the nodes and edges are Nodei|Sidej = GCN(sub_graph). Here, Nodei represents the i-th node, Sidej represents the j-th node, GCN is a graph convolutional network, and sub_graph is the subgraph of the first entity.

[0136] In some embodiments, the semantic features of each node and edge in the graph information can be averaged using pooling to obtain extended semantic features. For example, the extended semantic features are G_emb = avgPooling(Nodei|Sidej), where avgPooling represents the pooling averaging process.

[0137] In some embodiments, other methods may be used to fuse the semantic features of each node and edge to obtain extended semantic features, and this application embodiment does not limit this.

[0138] 504. Computer devices fuse descriptive semantic features and extended semantic features through a category recognition model to obtain fused features.

[0139] In this embodiment, the computer device can employ any fusion method to fuse descriptive semantic features and extended semantic features. For example, the computer device can concatenate descriptive semantic features and extended semantic features using a category recognition model to obtain the fused feature, such as Fusion = [CLS_token:G_emb]. Alternatively, the category recognition model may include a third feature extraction layer, through which the computer device extracts features from the descriptive semantic features and extended semantic features to obtain the fused feature.

[0140] 505. The computer device uses a category recognition model to identify the category of the descriptive semantic features and obtains the first category of the target video.

[0141] In some embodiments, the category recognition model includes a first recognition layer for identifying the primary category of the video. Optionally, the category recognition model is used to determine the primary category to which the video belongs from a plurality of primary categories. Optionally, the category recognition model is used to determine the probability that the video belongs to a plurality of primary categories, and determine the primary category corresponding to the highest probability as the primary category to which the video belongs.

[0142] 506. The computer device uses a category recognition model to identify the category of the fused features and obtains the second category of the target video.

[0143] In some embodiments, the category recognition model includes a second recognition layer for identifying the secondary category of the video. Optionally, the category recognition model is used to determine the secondary category to which the video belongs from a plurality of secondary categories. Optionally, the category recognition model is used to determine the probability that the video belongs to a plurality of secondary categories, and determine the secondary category corresponding to the highest probability as the secondary category to which the video belongs.

[0144] It should be noted that, in this embodiment of the application, the target video may have multiple secondary categories. In some embodiments, the category recognition model is used to determine the probability that the video belongs to multiple secondary categories, and the secondary category whose probability meets the recognition condition is determined as the secondary category to which the video belongs.

[0145] Optionally, determining the secondary category to which the video belongs based on the probability satisfying the recognition condition includes: determining the secondary category to which the video belongs based on the probability greater than the target probability. The target probability can be an empirical value or a value set by a technician, etc., and this application embodiment does not limit the target probability.

[0146] It should be noted that the first recognition layer and the second recognition layer in the embodiments of this application can be a softmax classifier or other classifiers, and the embodiments of this application do not limit this.

[0147] In some embodiments, such as Figure 6 As shown, the category recognition model includes a first feature extraction layer, a second feature extraction layer, a first recognition layer, and a second recognition layer; wherein the first feature extraction layer is connected to the first and second recognition layers, and the second feature extraction layer is connected to the second recognition layer. First video description information is input to the first feature extraction layer, and first knowledge extension information is input to the second feature extraction layer. The first feature extraction layer obtains the descriptive semantic features of the first video description information, and the second feature extraction layer obtains the extended semantic features of the first knowledge extension information. The descriptive semantic features are input to the first recognition layer, which outputs the primary category of the target video. The descriptive semantic features and extended semantic features are concatenated and input to the second recognition layer, which outputs the secondary category of the target video.

[0148] The category determination method provided in this application's embodiments processes the first video description information and the first knowledge extension information using a category recognition model to obtain the primary and secondary categories of the target video, ensuring the accuracy of the determined primary and secondary categories.

[0149] It should be noted that the category recognition model provided in this application embodiment is a trained model that has achieved the target accuracy. The target accuracy is an empirical value or is set by a technician. This application embodiment does not limit the target accuracy.

[0150] Figure 7 This is a flowchart illustrating a training method for a category recognition model provided in an embodiment of this application. The training method for the category recognition model includes the following steps:

[0151] 701. The computer device acquires sample data, which includes second video description information of the sample video, and the primary and secondary categories of the sample video.

[0152] In this embodiment, the sample data can be one or multiple data points; this embodiment does not limit the number. The primary and secondary categories of the sample are the correct primary and secondary categories of the sample video. The sample data can be manually labeled or obtained from the internet; this embodiment does not limit the number of data points.

[0153] The embodiments of this application illustrate the sample data through Table 1.

[0154] Table 1

[0155]

[0156] 702. The computer device acquires second knowledge extension information of the second entity, wherein the second entity is an entity in the second video description information.

[0157] Step 702 is the same as step 303 above, and will not be described in detail here.

[0158] 703. The computer device uses the pre-trained category recognition model to perform category recognition on the second video description information and the second knowledge extension information, and obtains the predicted primary category and predicted secondary category of the sample video.

[0159] The category recognition model before training refers only to the current training and is the category recognition model prior to this training. Therefore, the category recognition model can be an untrained category recognition model; it can also be a trained model that is not fully trained; or it can be a trained model that has already been put into use and is being trained again to improve accuracy. This application does not limit the category recognition model before training.

[0160] It should be noted that the process in step 703 above, which involves "using the pre-trained category recognition model to perform category recognition on the second video description information and the second knowledge extension information to obtain the predicted primary category and predicted secondary category of the sample video," is the same as described above. Figure 5 The process of "determining the primary and secondary categories of the target video based on the first video description information and the first knowledge extension information using the category recognition model" is similar and will not be described in detail here.

[0161] 704. The computer equipment trains the category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category, and the sample secondary category to obtain the category recognition model.

[0162] The category recognition model is trained in this embodiment to ensure its output is accurate. Specifically, the predicted primary category should match the sample's primary category, and the predicted secondary category should also match the sample's secondary category. If the predicted primary category differs from the sample's primary category, the model's output is incorrect, and its parameters need adjustment to improve accuracy. Similarly, if the predicted secondary category differs from the sample's secondary category, the model's output is also incorrect, requiring further adjustment to improve accuracy.

[0163] In this embodiment of the application, when the computer device trains the category recognition model before training based on the predicted primary category, the sample primary category, the predicted secondary category, and the sample secondary category, it can train the entire category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category, and the sample secondary category, or it can train different parts of the category recognition model separately.

[0164] In some embodiments, the computer device trains different parts of the category recognition model separately. Optionally, the category recognition model includes a first feature extraction layer, a second feature extraction layer, a first recognition layer, and a second recognition layer. The computer device trains the untrained category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category, and the sample secondary category to obtain the category recognition model, including: the computer device trains the first feature extraction layer and the first recognition layer based on the predicted primary category and the sample primary category; and trains the second feature extraction layer and the second recognition layer based on the predicted secondary category and the sample secondary category to obtain the category recognition model.

[0165] In some embodiments, the computer device trains the entire category recognition model together. The computer device trains the untrained category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category, and the sample secondary category to obtain the category recognition model, including: determining a first loss value based on the predicted primary category and the sample primary category; determining a second loss value based on the predicted secondary category and the sample secondary category; fusing the first loss value and the second loss value to obtain a third loss value; and training the untrained category recognition model based on the third loss value to obtain the category recognition model.

[0166] Among them, the first loss value is positively correlated with the difference between the predicted first-level category and the sample first-level category, and the second loss value is positively correlated with the difference between the predicted second-level category and the sample second-level category.

[0167] In some embodiments, the computer device processes the predicted primary category and the sample primary category based on a first loss function to obtain a first loss value; the computer device also processes the predicted secondary category and the sample secondary category based on a second loss function to obtain a second loss value. Optionally, the first loss function and the second loss function are negative logarithmic loss functions. Optionally, the first loss function and the second loss function are other types of loss functions.

[0168] In some embodiments, the computer device fuses a first loss value and a second loss value to obtain a third loss value, including: the computer device weighting the first loss value and the second loss value to obtain the third loss value. The weights of the first loss value and the second loss value can be empirical values, i.e., set by a technician. Optionally, the sum of the weights of the first loss value and the second loss value is 1. Optionally, the computer device fuses the first loss value and the second loss value to obtain the third loss value, including: the computer device using a loss harmonic function to harmonicize the first loss value and the second loss function to obtain the third loss value. The loss harmonic function can be any loss harmonic function in the related art; this application embodiment does not limit the loss harmonic function.

[0169] It should be noted that, since the primary category in this embodiment is a coarse-grained classification and the secondary category is a more refined classification, determining the secondary category is more difficult than determining the primary category. Therefore, to better ensure the accuracy of the category recognition model, this embodiment trains the category recognition model to determine a primary category with a higher probability than the determined secondary category. Thus, when training the category recognition model, both the probability of the primary category and the probability of the secondary category are considered.

[0170] In some embodiments, the computer device trains a pre-training category recognition model based on a predicted primary category, a sample primary category, a predicted secondary category, and a sample secondary category to obtain a category recognition model, including: determining a first probability and a second probability, wherein the category recognition model is used to determine the probability that a target video belongs to multiple primary categories and the probability that a target video belongs to multiple secondary categories, the first probability being the probability of at least one primary category determined by the category recognition model, and the second probability being the probability of at least one secondary category determined by the category recognition model; and training the pre-training category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category, the sample secondary category, the first probability, and the second probability to obtain a category recognition model.

[0171] The first probability is the probability of at least one primary category determined by the category recognition model, which is also the probability that the sample video belongs to the at least one primary category determined by the category recognition model; the second probability is the probability of at least one secondary category determined by the category recognition model, which is also the probability that the sample video belongs to the at least one secondary category determined by the category recognition model.

[0172] Wherein, at least one primary category can be the correct primary category of the sample video, or the primary category to which the sample video belongs as determined by the category recognition model, or multiple primary categories in the primary category classification, etc. The embodiments of this application do not limit the requirement of at least one primary category. At least one secondary category is similar to at least one primary category, and will not be described in detail here.

[0173] Optionally, the computer device determines a first probability and a second probability, including any one of the following:

[0174] (1) The probability of the correct first-level category determined by the category recognition model is determined as the first probability, and the probability of the correct second-level category determined by the category recognition model is determined as the second probability. The correct first-level category is the same as the first-level category of the sample, and the correct second-level category is the same as the second-level category of the sample.

[0175] (2) The probability of predicting the first category is determined as the first probability, and the probability of predicting the second category is determined as the second probability.

[0176] (3) The probability of predicting the first-level category is determined as the first probability, and the probabilities of the multiple second-level categories determined by the category recognition model are determined as the second probability.

[0177] (4) The probabilities of multiple first-level categories determined by the category recognition model are determined as the first probability, and the probabilities of multiple second-level categories determined by the category recognition model are determined as the second probability.

[0178] In some embodiments, a computer device trains a pre-training category recognition model based on a predicted primary category, a sample primary category, a predicted secondary category, a sample secondary category, a first probability, and a second probability to obtain a category recognition model, including: determining a first loss value based on the predicted primary category and the sample primary category; determining a second loss value based on the predicted secondary category and the sample secondary category; determining a fourth loss value based on the first probability and the second probability; fusing the first loss value, the second loss value, and the fourth loss value to obtain a fifth loss value; and training the category recognition model based on the fifth loss value.

[0179] The first loss value is positively correlated with the difference between the predicted primary category and the sample primary category, and the second loss value is positively correlated with the difference between the predicted secondary category and the sample secondary category. The fourth loss value is used to penalize the category recognition model when the first probability is less than the second probability; that is, the model parameters of the category recognition model are adjusted based on the loss value.

[0180] Optionally, the computer device determines a fourth loss value based on a first probability and a second probability, including: when the first probability is less than the second probability, determining a fourth loss value based on the difference between the first probability and the second probability, wherein the fourth loss value is positively correlated with the difference; and when the first probability is not less than the second probability, determining the fourth loss value to be 0.

[0181] In some embodiments, the computer device processes the predicted primary category and the sample primary category based on a first loss function to obtain a first loss value; the computer device processes the predicted secondary category and the sample secondary category based on a second loss function to obtain a second loss value; and the computer device processes the first probability and the second probability based on a third loss function to obtain a fourth loss value.

[0182] In some embodiments, the first loss function and the second loss function are negative logarithmic loss functions. For example, the first loss function is:

[0183]

[0184] Where loss1 is the first loss value. The function is a summation function, where n is the number of first-level categories, i represents the i-th first-level category, and y... iIndicate whether the sample video belongs to the i-th primary category. If y i A value of 1 indicates that the sample data belongs to the i-th primary category; if y i A value of 0 indicates that the sample data does not belong to the i-th primary category; log is the logarithmic function, a i This represents the probability that the sample video, as predicted by the category recognition model, belongs to the i-th primary category. i is a positive integer greater than 0 and less than or equal to n, and n is a positive integer greater than 1.

[0185] The second loss function is:

[0186]

[0187] Where loss2 is the second loss value. The function is a summation function, where m is the number of second-level categories, j represents the j-th first-level category, and y... j Indicate whether the sample video belongs to the j-th primary category. If y j A value of 1 indicates that the sample data belongs to the j-th primary category; if y j A value of 0 indicates that the sample data does not belong to the j-th primary category; log is the logarithmic function, a j This represents the probability that the sample video, as predicted by the category recognition model, belongs to the j-th primary category. J is a positive integer greater than 0 and less than or equal to m, where m is a positive integer greater than 1.

[0188] The third loss function is:

[0189]

[0190] Where loss3 is the fourth loss value, Let be the summation function, k represent the k-th correct second-order category, z represent the number of correct second-order categories, and max represent the maximum value function.

[0191] When λ+I 2k When -I1 is less than or equal to 0, the output of the max function is 0; when λ+I 2k When -I1 is less than or equal to 0, the output of the max function is λ+I. 2k -I1; where λ is a hyperparameter greater than 0, determined by empirical values. 2k Let Ik represent the probability of the k-th correct secondary category determined by the category recognition model, and I1 represent the probability of the correct primary category determined by the category recognition model. Here, k is a positive integer greater than 0 and less than or equal to z, and z is a positive integer greater than 0.

[0192] The first loss value, the second loss value, and the fourth loss value are fused using the following formula:

[0193] LOSS=λ1loss1+λ2loss2+λ3loss3

[0194] LOSS is the fifth loss value, λ1 is the weight of the first loss value, λ2 is the weight of the second loss value, λ3 is the weight of the fourth loss value, loss1 is the first loss value, loss2 is the second loss value, and loss3 is the fourth loss value. Among them, λ1, λ2, and λ3 are hyperparameters greater than 0.

[0195] The training method for the category recognition model provided in this application uses the video description information of sample videos and the knowledge extension information of entities in the video description information to train the category recognition model into a model that can determine both the primary and secondary categories of the video. Compared with related technologies that only identify the secondary category of the video and then backtrack to the primary category from the secondary category, the category recognition model provided in this application can determine the category of the video more accurately.

[0196] Furthermore, the training method for the category recognition model provided in this application not only trains the category recognition model based on the prediction results and sample results, but also takes into account that the difficulty of determining the first-level category is lower than that of determining the second-level category. Therefore, by controlling the probability of the first-level category to be greater than that of the second-level category, the category recognition model can more accurately determine the category of the video.

[0197] Figure 8 This is a block diagram of a category determination apparatus according to an embodiment of this application. The apparatus is used to perform the above-described category determination method. (See also...) Figure 8 The device includes:

[0198] The first acquisition module 801 is used to acquire first video description information of the target video, which is used to describe the video content of the target video.

[0199] The recognition module 802 is used to perform entity recognition on the first video description information to obtain the first entity in the first video description information;

[0200] The second acquisition module 803 is used to acquire the first knowledge extension information of the first entity;

[0201] The determination module 804 is used to determine the category of the target video based on the first video description information and the first knowledge extension information.

[0202] like Figure 9 As shown, in some embodiments, the category of the target video includes a primary category and a secondary category; the determining module 804 includes:

[0203] The first determining unit 8041 is used to determine the primary category of the target video based on the first video description information;

[0204] The second determining unit 8042 is used to determine the secondary category of the target video based on the first video description information and the first knowledge extension information.

[0205] In some embodiments, the first determining unit 8041 is configured to determine the descriptive semantic features of the first video description information, and determine the primary category of the target video based on the descriptive semantic features;

[0206] The second determining unit 8042 is used to determine the extended semantic features of the first knowledge extended information, fuse the descriptive semantic features and the extended semantic features to obtain fused features, and determine the secondary category of the target video based on the fused features.

[0207] In some embodiments, the category of the target video includes a primary category and a secondary category; the determining module 804 is used to input the first video description information and the first knowledge extension information into the category recognition model, and output the primary category and secondary category of the target video. The category recognition model is used to determine the primary category of the target video based on the first video description information, and to determine the secondary category of the target video based on the first video description information and the first knowledge extension information.

[0208] In some embodiments, the device further includes:

[0209] The third acquisition module 805 is used to acquire sample data, which includes the second video description information of the sample video, the primary category of the sample video, and the secondary category of the sample video.

[0210] The second acquisition module 803 is further configured to acquire second knowledge extension information of the second entity, wherein the second entity is an entity in the second video description information;

[0211] The determining module 804 is also used to perform category recognition on the second video description information and the second knowledge extension information through the category recognition model before training, so as to obtain the predicted primary category and predicted secondary category of the sample video;

[0212] Training module 806 is used to train the pre-training category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category, and the sample secondary category to obtain the category recognition model.

[0213] In some embodiments, the training module 806 includes:

[0214] The first determining unit 8061 is used to determine a first loss value based on the predicted first-level category and the sample first-level category;

[0215] The second determining unit 8062 is used to determine a second loss value based on the predicted secondary category and the sample secondary category;

[0216] The fusion unit 8063 is used to fuse the first loss value and the second loss value to obtain a third loss value;

[0217] Training unit 8064 is used to train the pre-training category recognition model based on the third loss value to obtain the category recognition model.

[0218] In some embodiments, the training module 806 is further configured to determine a first probability and a second probability, and the category recognition model is configured to determine the probability that the target video belongs to multiple first-level categories and the probability that the target video belongs to multiple second-level categories, wherein the first probability is the probability of at least one first-level category determined by the category recognition model, and the second probability is the probability of at least one second-level category determined by the category recognition model.

[0219] The training module 806 is used to train the pre-training category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category, the sample secondary category, the first probability, and the second probability, to obtain the category recognition model.

[0220] In some embodiments, the training module 806 is configured to perform any of the following:

[0221] The probability of the correct primary category determined by the category recognition model is defined as the first probability, and the probability of the correct secondary category determined by the category recognition model is defined as the second probability. The correct primary category is the same as the primary category of the sample, and the correct secondary category is the same as the secondary category of the sample.

[0222] The probability of predicting the first category is defined as the first probability, and the probability of predicting the second category is defined as the second probability;

[0223] The probability of predicting the first-level category is defined as the first probability, and the probabilities of the multiple second-level categories determined by the category recognition model are defined as the second probability.

[0224] The probabilities of multiple primary categories determined by the category recognition model are defined as the first probability, and the probabilities of multiple secondary categories determined by the category recognition model are defined as the second probability.

[0225] In some embodiments, the training module 806 includes:

[0226] The first determining unit 8061 is used to determine a first loss value based on the predicted first-level category and the sample first-level category;

[0227] The second determining unit 8062 is used to determine a second loss value based on the predicted secondary category and the sample secondary category;

[0228] The third determining unit 8065 is used to determine the fourth loss value based on the first probability and the second probability;

[0229] The fusion unit 8063 is used to fuse the first loss value, the second loss value and the fourth loss value to obtain a fifth loss value;

[0230] Training unit 8064 is used to train the category recognition model based on the fifth loss value.

[0231] In some embodiments, the third determining unit 8065 is configured to determine the fourth loss value based on the difference between the first probability and the second probability when the first probability is less than the second probability, wherein the fourth loss value is positively correlated with the difference; and to determine the fourth loss value as 0 when the first probability is not less than the second probability.

[0232] It should be noted that the category determination method provided in the above embodiments is only an example of the division of the above functional modules when determining video categories. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the category determination device and the category determination method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0233] In this application embodiment, the computer device can be configured as a terminal or a server. When the computer device is configured as a terminal, the terminal acts as the execution subject to implement the technical solution provided in this application embodiment; when the computer device is configured as a server, the server acts as the execution subject to implement the technical solution provided in this application embodiment; or, the technical solution provided in this application can be implemented through the interaction between the terminal and the server. This application embodiment does not limit this.

[0234] Figure 10 This is a structural block diagram of a terminal provided in an embodiment of this application. The terminal 1000 is used to execute the steps executed by the terminal in the above embodiments, and can be a portable mobile terminal, such as a laptop computer, desktop computer, intelligent voice interaction device, smart home appliance, vehicle terminal, etc.

[0235] Typically, terminal 1000 includes a processor 1001 and a memory 1002.

[0236] Processor 1001 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. Processor 1001 may employ DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA.

[0237] The processor 1001 may be implemented using at least one hardware form of a Programmable Logic Array (PLA). The processor 1001 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 1001 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, the processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0238] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 are used to store at least one program code, which is executed by the processor 1001 to implement the category determination method provided in the method embodiments of this application.

[0239] In some embodiments, the terminal 1000 may also optionally include a peripheral device interface 1003 and at least one peripheral device. The processor 1001, memory 1002, and peripheral device interface 1003 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1003 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007, and a power supply 1008.

[0240] Peripheral device interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1001 and memory 1002. In some embodiments, processor 1001, memory 1002 and peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1001, memory 1002 and peripheral device interface 1003 can be implemented on separate chips or circuit boards, and this application embodiment does not limit this.

[0241] The radio frequency (RF) circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1004 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1004 can communicate with other terminals via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1004 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0242] Display screen 1005 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1005 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1001 for processing. In this case, display screen 1005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1005, serving as the front panel of terminal 1000; in other embodiments, there may be at least two display screens, respectively disposed on different surfaces of terminal 1000 or in a folded design; in still other embodiments, display screen 1005 may be a flexible display screen, disposed on a curved or folded surface of terminal 1000. Furthermore, display screen 1005 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1005 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0243] The camera assembly 1006 is used to acquire images or videos. Optionally, the camera assembly 1006 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1006 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0244] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1001 for processing, or input to the radio frequency circuit 1004 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 1000. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1007 may also include a headphone jack.

[0245] The power supply 1008 is used to power the various components in the terminal 1000. The power supply 1008 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When the power supply 1008 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0246] In some embodiments, the terminal 1000 further includes one or more sensors 1010. The one or more sensors 1010 include, but are not limited to: an acceleration sensor 1011, a gyroscope sensor 1012, a pressure sensor 1013, an optical sensor 1014, and a proximity sensor 1015.

[0247] Accelerometer 1011 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal 1000. For example, accelerometer 1011 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1001 can control display screen 1005 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1011. Accelerometer 1011 can also be used for games or for acquiring user motion data.

[0248] The gyroscope sensor 1012 can detect the orientation and rotation angle of the terminal 1000. The gyroscope sensor 1012, in conjunction with the accelerometer sensor 1011, can collect 3D motion data from the user on the terminal 1000. Based on the data collected by the gyroscope sensor 1012, the processor 1001 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0249] The pressure sensor 1013 can be disposed on the side bezel of the terminal 1000 and / or on the lower layer of the display screen 1005. When the pressure sensor 1013 is disposed on the side bezel of the terminal 1000, it can detect the user's grip signal on the terminal 1000, and the processor 1001 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1013. When the pressure sensor 1013 is disposed on the lower layer of the display screen 1005, the processor 1001 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1005. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0250] An optical sensor 1014 is used to collect ambient light intensity. In one embodiment, the processor 1001 can control the display brightness of the display screen 1005 based on the ambient light intensity collected by the optical sensor 1014. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1005 is increased; when the ambient light intensity is low, the display brightness of the display screen 1005 is decreased. In another embodiment, the processor 1001 can also dynamically adjust the shooting parameters of the camera assembly 1006 based on the ambient light intensity collected by the optical sensor 1014.

[0251] The proximity sensor 1015, also known as a distance sensor, is typically installed on the front panel of the terminal 1000. The proximity sensor 1015 is used to detect the distance between the user and the front of the terminal 1000. In one embodiment, when the proximity sensor 1015 detects that the distance between the user and the front of the terminal 1000 is gradually decreasing, the processor 1001 controls the display screen 1005 to switch from a screen-on state to a screen-off state; when the proximity sensor 1015 detects that the distance between the user and the front of the terminal 1000 is gradually increasing, the processor 1001 controls the display screen 1005 to switch from a screen-off state to a screen-on state.

[0252] Those skilled in the art will understand that Figure 10 The structure shown does not constitute a limitation on terminal 1000 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0253] Figure 11This is a schematic diagram of a server structure provided in an embodiment of this application. The server 1100 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1101 and one or more memories 1102. The memories 1102 store at least one line of program code, which is loaded and executed by the processor 1101 to implement the methods provided in the above-described method embodiments. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated upon here.

[0254] This application also provides a computer device, which includes a processor and a memory. The memory stores at least one piece of program code, which is loaded by the processor and executes the operations performed in the communication connection establishment of the above embodiments.

[0255] This application also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to implement the category determination method as described in the above embodiments.

[0256] This application also provides a computer program product or computer program that includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to perform the operations performed in the category determination method of the above embodiments.

[0257] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware, or by a program or program code related to hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0258] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A category determining method characterized by comprising: The method comprises: obtaining first video description information for describing the video content of a target video; performing entity recognition on the first video description information to obtain a first entity in the first video description information; determining a first node representing the first entity from a knowledge graph recording a plurality of entities, types of the plurality of entities, attributes of the plurality of entities, and other objects associated with the plurality of entities, and extracting a sub-knowledge graph of a target hop from the knowledge graph with the first node as the center as first knowledge extension information of the first entity, the sub-knowledge graph including nodes directly or indirectly connected to the first node and edges between the nodes; the first knowledge extension information includes other information in addition to information related to the first entity in the first video description information, which extends the knowledge information of the first entity; performing feature extraction on the first video description information through a first feature extraction layer of a category recognition model to obtain description semantic features for describing semantics of the first video description information; performing category recognition on the description semantic features through a first recognition layer in the category recognition model to determine a primary category of the target video; encoding the first knowledge extension information through a second feature extraction layer of the category recognition model to obtain extension semantic features of the first knowledge extension information; fusing the description semantic features and the extension semantic features through a third feature extraction layer of the category recognition model to obtain fusion features; performing category recognition on the fusion features through a second recognition layer of the category recognition model to determine a secondary category of the target video; the category recognition model is trained based on a fifth loss value, the fifth loss value is obtained by fusing a fourth loss value and other loss values, the other loss values are determined based on a correct category of a sample video and a predicted category of the sample video obtained by a category recognition model based on second video description information and second knowledge extension information of the sample video, the fourth loss value is determined based on a probability of a correct primary category identical to a sample primary category of the sample video, a probability of a correct secondary category identical to a sample secondary category of the sample video, and a number of correct secondary categories.

2. The method of claim 1, wherein, The training process of the category recognition model comprises: obtaining sample data, the sample data including second video description information of a sample video, a sample primary category and a sample secondary category of the sample video; obtaining second knowledge extension information of a second entity, the second entity being an entity in the second video description information; performing category recognition on the second video description information and the second knowledge extension information through a category recognition model before training to obtain a predicted primary category and a predicted secondary category of the sample video; training the category recognition model before training based on the predicted primary category, the sample primary category, the predicted secondary category, and the sample secondary category to obtain the category recognition model.

3. The method of claim 2, wherein, The training of the pre-training category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category and the sample secondary category comprises: determining a first loss value based on the predicted primary category and the sample primary category; determining a second loss value based on the predicted secondary category and the sample secondary category; fusing the first loss value and the second loss value to obtain a third loss value; training the pre-training category recognition model based on the third loss value to obtain the category recognition model.

4. The method of claim 2, wherein, The training of the pre-training category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category and the sample secondary category comprises: determining a first probability and a second probability, wherein the category recognition model is configured to determine a probability of the target video belonging to a plurality of primary categories and a probability of the target video belonging to a plurality of secondary categories, the first probability is a probability of at least one primary category determined by the category recognition model, and the second probability is a probability of at least one secondary category determined by the category recognition model; training the pre-training category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category, the sample secondary category, the first probability and the second probability to obtain the category recognition model.

5. The method of claim 4, wherein, The training of the pre-training category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category and the sample secondary category comprises: determining a first loss value based on the predicted primary category and the sample primary category; determining a second loss value based on the predicted secondary category and the sample secondary category; determining a fourth loss value based on the first probability and the second probability; fusing the first loss value, the second loss value and the fourth loss value to obtain a fifth loss value; training the category recognition model based on the fifth loss value.

6. The method of claim 5, wherein, The determination of the fourth loss value based on the first probability and the second probability comprises: in a case where the first probability is less than the second probability, determining the fourth loss value based on a difference between the first probability and the second probability, wherein the fourth loss value is positively correlated with the difference; in a case where the first probability is not less than the second probability, determining the fourth loss value as 0.

7. A category determining apparatus characterized by comprising: The apparatus comprises: a first obtaining module configured to obtain first video description information for describing video content of a target video; an identification module configured to perform entity recognition on the first video description information to obtain first entities in the first video description information; The second acquisition module is configured to determine a first node representing the first entity from a knowledge graph recording a plurality of entities, types of the plurality of entities, attributes of the plurality of entities, and other objects associated with the plurality of entities, and to extract, as first knowledge expansion information of the first entity, a sub-knowledge graph of a target hop from the knowledge graph with the first node as a center, the sub-knowledge graph including nodes directly or indirectly connected to the first node and edges between the nodes; the first knowledge expansion information includes other information in addition to information related to the first entity in the first video description information, and plays an expansion role on knowledge information of the first entity; The determination module is configured to perform feature extraction on the first video description information through a first feature extraction layer of the category recognition model to obtain description semantic features for describing semantics of the first video description information, perform category recognition on the description semantic features through a first recognition layer in the category recognition model to determine a primary category of the target video, encode the first knowledge expansion information through a second feature extraction layer of the category recognition model to obtain expansion semantic features of the first knowledge expansion information, fuse the description semantic features and the expansion semantic features through a third feature extraction layer of the category recognition model to obtain fused features, and perform category recognition on the fused features through a second recognition layer of the category recognition model to determine a secondary category of the target video. The category recognition model is trained based on a fifth loss value, the fifth loss value is obtained by fusing a fourth loss value and other loss values, the other loss values are determined based on a correct category of a sample video and a predicted category of the sample video obtained by the category recognition model based on second video description information and second knowledge expansion information of the sample video, and the fourth loss value is determined based on a probability of a correct primary category identical to a sample primary category of the sample video, a probability of a correct secondary category identical to a sample secondary category of the sample video, and a number of correct secondary categories.

8. The apparatus of claim 7, wherein, The apparatus further includes: The third acquisition module is configured to acquire sample data, the sample data including second video description information of a sample video, a sample primary category, and a sample secondary category of the sample video. The second acquisition module is further configured to acquire second knowledge expansion information of a second entity, the second entity being an entity in the second video description information. The determination module is further configured to perform category recognition on the second video description information and the second knowledge expansion information through the category recognition model before training to obtain a predicted primary category and a predicted secondary category of the sample video. The training module is configured to train the category recognition model before training based on the predicted primary category, the sample primary category, the predicted secondary category, and the sample secondary category to obtain the category recognition model.

9. The apparatus of claim 8, wherein, The training module includes: The first determination unit is configured to determine a first loss value based on the predicted primary category and the sample primary category. The second determining unit is configured to determine a second loss value based on the predicted secondary category and the sample secondary category; The fusion unit is configured to fuse the first loss value and the second loss value to obtain a third loss value; The training unit is configured to train the pre-training category recognition model based on the third loss value to obtain the category recognition model.

10. The apparatus of claim 8, wherein, The training module is further configured to: determine a first probability and a second probability, the category recognition model being configured to determine probabilities of the target video belonging to a plurality of primary categories and probabilities of the target video belonging to a plurality of secondary categories, the first probability being a probability of at least one primary category determined by the category recognition model, and the second probability being a probability of at least one secondary category determined by the category recognition model; train the pre-training category recognition model based on the predicted primary category, the sample primary category, the predicted secondary category, the sample secondary category, the first probability, and the second probability to obtain the category recognition model.

11. The apparatus of claim 10, wherein, The training module includes: The first determining unit is configured to determine a first loss value based on the predicted primary category and the sample primary category; The second determining unit is configured to determine a second loss value based on the predicted secondary category and the sample secondary category; The third determining unit is configured to determine the fourth loss value based on the first probability and the second probability; The fusion unit is configured to fuse the first loss value, the second loss value, and the fourth loss value to obtain a fifth loss value; The training unit is configured to train the category recognition model based on the fifth loss value.

12. The apparatus of claim 11, wherein, The third determining unit is configured to: in a case where the first probability is less than the second probability, determine the fourth loss value based on a difference between the first probability and the second probability, the fourth loss value being positively correlated with the difference; in a case where the first probability is not less than the second probability, determine the fourth loss value as 0.

13. A computer device, comprising: The computer device includes a processor and a memory, the memory being configured to store at least one piece of computer program, the at least one piece of computer program being loaded and executed by the processor to implement the category determination method in any one of claims 1 to 6.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium is configured to store at least one piece of computer program, the at least one piece of computer program being configured to implement the category determination method in any one of claims 1 to 6.

15. A computer program product comprising computer program code, characterised in that, The computer program code is stored in the computer-readable storage medium, and the processor of the computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code to cause the computer device to implement the category determination method in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multimedia resource classification method and device, electronic equipment and storage medium

    CN113239215A

  • Content classification method and device, electronic equipment and computer readable storage medium

    CN113821632A