Video classification method and device, computer device and storage medium

CN116229329BActive Publication Date: 2026-10-09DOUYIN VISION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310287019.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2026-10-09
Estimated Expiration
2043-03-22

AI Technical Summary

Technical Problem

[0003]一般的,可以利用神经网络提取短视频的内容特征,得到特征数据,神经网络根据提取到的特征数据,进行短视频的分类,但是,由于短视频天然的具有时短、强编辑等特点,使得神经网络的分类结果的准确性较低

Benefits of technology

[0015] The video classification method, apparatus, computer device, and storage medium provided in this disclosure, after acquiring a video to be detected, acquire non-content attribute information of the video to be detected, such as video duration, number of likes, etc., and perform feature extraction on the video to be detected to generate feature data of multiple modalities matching the content of the video to be detected; and based on the task content of the classification task, encode the non-content attribute information to generate attribute feature data, which matches the classification task, so that the attribute feature data can provide auxiliary information for video classification, and thus, based on the feature data of multiple modalities and attribute feature data, can more accurately determine the classification result of the video to be detected under the classification task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116229329B_ABST
    Figure CN116229329B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video classification method and device, computer equipment and a storage medium, wherein the method comprises: obtaining a to-be-detected video and non-content attribute information of the to-be-detected video; performing feature extraction on the to-be-detected video to generate feature data of multiple modalities matched with the content of the to-be-detected video; and encoding the non-content attribute information based on the task content of a classification task to generate attribute feature data; and determining a classification result of the to-be-detected video under the classification task based on the feature data of the multiple modalities and the attribute feature data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to a video classification method, apparatus, computer device, and storage medium. Background Technology

[0002] With the rapid development of science and technology, mobile devices have become indispensable in people's lives. Due to the widespread use of mobile devices, short videos have become a common form of multimedia content. Therefore, how to accurately classify short videos has become a hot research topic.

[0003] Generally, neural networks can be used to extract content features from short videos to obtain feature data. The neural network then classifies the short videos based on the extracted feature data. However, due to the inherent characteristics of short videos, such as their short duration and heavy editing, the accuracy of the classification results obtained by the neural network is relatively low. Summary of the Invention

[0004] This disclosure provides at least one video classification method, apparatus, computer device, and storage medium.

[0005] In a first aspect, embodiments of this disclosure provide a video classification method, including:

[0006] Obtain the video to be detected and its non-content attribute information;

[0007] Feature extraction is performed on the video to be detected to generate feature data of multiple modalities that match the content of the video to be detected; and based on the task content of the classification task, the non-content attribute information is encoded to generate attribute feature data.

[0008] Based on the feature data of the multiple modalities and the attribute feature data, the classification result of the video to be detected under the classification task is determined.

[0009] Secondly, embodiments of this disclosure also provide a video classification device, comprising:

[0010] The acquisition module is used to acquire the video to be detected and the non-content attribute information of the video to be detected;

[0011] The generation module is used to extract features from the video to be detected and generate feature data of multiple modalities that match the content of the video to be detected; and to encode the non-content attribute information based on the task content of the classification task to generate attribute feature data.

[0012] The determination module is used to determine the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data.

[0013] Thirdly, embodiments of this disclosure also provide a computer device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the first aspect above, or any possible implementation of the first aspect, are performed.

[0014] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the first aspect or any possible implementation thereof.

[0015] The video classification method, apparatus, computer device, and storage medium provided in this disclosure, after acquiring a video to be detected, acquire non-content attribute information of the video to be detected, such as video duration, number of likes, etc., and perform feature extraction on the video to be detected to generate feature data of multiple modalities matching the content of the video to be detected; and based on the task content of the classification task, encode the non-content attribute information to generate attribute feature data, which matches the classification task, so that the attribute feature data can provide auxiliary information for video classification, and thus, based on the feature data of multiple modalities and attribute feature data, can more accurately determine the classification result of the video to be detected under the classification task.

[0016] Considering that non-content attributes of a video can significantly impact the classification results of the video to be detected, and that different tasks focus on different attribute information, this disclosure obtains the non-content attribute information of the video to be detected, encodes the non-content attribute information based on the task content of the classification task, and generates attribute feature data. This approach not only focuses on the content-related features of the video to be detected but also on the non-content-related features, resulting in a richer set of feature information. Consequently, the classification results obtained based on feature data from multiple modalities and attribute feature data have higher accuracy.

[0017] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0019] Figure 1 A flowchart of a video classification method provided by an embodiment of this disclosure is shown;

[0020] Figure 2a A flowchart is shown illustrating a specific method for generating attribute feature data in the video classification method provided in this embodiment of the present disclosure.

[0021] Figure 2b A flowchart is shown for another specific method for generating attribute feature data in the video classification method provided in this disclosure embodiment;

[0022] Figure 2c A flowchart is shown for another specific method for generating attribute feature data in the video classification method provided in this disclosure embodiment;

[0023] Figure 3 A schematic diagram of a video classification apparatus provided in an embodiment of this disclosure is shown;

[0024] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0026] In related technologies, neural networks used for video classification can be used to extract content features from short videos to obtain feature data. The neural network then classifies the short videos based on the extracted feature data. However, the aforementioned neural networks for video classification are generally applied to long video classification scenarios. Due to the inherent characteristics of short videos, such as short duration and heavy editing, the accuracy of the classification results of the aforementioned neural networks is relatively low.

[0027] Based on the above research, this disclosure provides a video classification method. After acquiring the video to be detected, non-content attribute information of the video to be detected is obtained, such as video duration, number of likes, etc., and features are extracted from the video to be detected to generate feature data of multiple modalities that match the content of the video to be detected. Based on the task content of the classification task, the non-content attribute information is encoded to generate attribute feature data, which matches the classification task, so that the attribute feature data can provide auxiliary information for video classification. Therefore, based on the feature data of multiple modalities and the attribute feature data, the classification result of the video to be detected under the classification task can be determined more accurately.

[0028] Considering that non-content attributes of a video can significantly impact the classification results of the video to be detected, and that different tasks focus on different attribute information, this disclosure obtains the non-content attribute information of the video to be detected, encodes the non-content attribute information based on the task content of the classification task, and generates attribute feature data. This approach not only focuses on the content-related features of the video to be detected but also on the non-content-related features, resulting in a richer set of feature information. Consequently, the classification results obtained based on feature data from multiple modalities and attribute feature data have higher accuracy.

[0029] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0030] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0031] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0032] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0033] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0034] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0035] To facilitate understanding of this embodiment, a detailed description of the video classification method disclosed in this disclosure is provided first. The execution entity of the video classification method provided in this disclosure is generally a computer device with a certain computing capability. This computer device may include, for example, a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a handheld device, a computing device, etc. In some possible implementations, the video classification method can be implemented by a processor calling computer-readable instructions stored in memory.

[0036] The video classification method provided in this disclosure embodiment will be described below using a server as the execution subject as an example.

[0037] See Figure 1 The diagram shows a flowchart of a video classification method provided in an embodiment of this disclosure. The method includes steps S101 to S103, wherein:

[0038] S101, Obtain the video to be detected and the non-content attribute information of the video to be detected;

[0039] S102, extract features from the video to be detected to generate feature data of multiple modalities that match the content of the video to be detected; and based on the task content of the classification task, encode the non-content attribute information to generate attribute feature data;

[0040] S103, Based on the feature data of the multiple modalities and the attribute feature data, determine the classification result of the video to be detected under the classification task.

[0041] The following sections will provide specific explanations of S101-S103.

[0042] Regarding S101:

[0043] The video to be tested can be any short video, which refers to a short video on a multimedia platform. After obtaining the video to be tested, non-content attribute information can also be obtained. This non-content attribute information refers to other attribute information of the video besides content features. For example, non-content attribute information may include: attribute information related to duration features, attribute information related to the evaluation features of each user on the video, attribute information related to the historical operation features of each user on the video, etc. The various attribute information included in the non-content attribute information can be set as needed.

[0044] The non-content attribute information includes at least one of the following: number of likes, number of dislikes, number of comments, number of reposts, number of favorites, number of shares, video duration, number of plays, and completion rate. For example, the completion rate can be the ratio of the number of times the video has been fully played to the total number of plays. In practice, the non-content attribute information of the video to be detected can be recorded during its dissemination process so that the non-content attribute information of the video to be detected can be obtained later.

[0045] Regarding S102:

[0046] After acquiring the video to be detected, features can be extracted from it to generate feature data in multiple modalities that match the content of the video. These modalities include, but are not limited to, video feature data, text feature data, and audio feature data.

[0047] In implementation, video feature data can be extracted according to the following steps: sample the video to be detected to determine multiple keyframes contained in the video; then extract features from the multiple keyframes to obtain intermediate feature data; perform pooling operations (such as average pooling, max pooling, etc.) on the intermediate feature data to obtain the video feature data. Alternatively, a pre-trained video feature extraction sub-model can be used to extract the video feature data; or an existing image detection model can be used.

[0048] Audio feature data can be extracted according to the following steps: Obtain the audio data contained in the video to be detected; preprocess the audio data to obtain processed audio data; preprocessing includes, but is not limited to, sampling rate conversion, noise reduction, and filtering. Sampling rate conversion can improve the extraction efficiency of audio feature data; noise reduction and filtering can remove at least some useless data from the audio data, so that audio feature data can be obtained more accurately from the processed audio data; then use a pre-trained sound classification model to extract features from the processed audio data to generate audio feature data.

[0049] The text feature data can be extracted according to the following steps: Determine the first text information corresponding to the audio data contained in the video to be detected, for example, by using Automatic Speech Recognition (ASR) technology; determine the second text information contained in the image frames of the video to be detected, for example, by using Optical Character Recognition (OCR); and obtain the third text information containing the title information, description information, etc. of the video to be detected; then preprocess the text data containing the first, second, and third text information, such as word segmentation and determining word vector representations, to obtain processed text data; and finally use a pre-trained text classification model to extract features from the processed text data to generate text feature data.

[0050] Furthermore, based on the task content of the classification task, non-content attribute information can be encoded to obtain a vector corresponding to each attribute information. These vectors are then concatenated to generate attribute feature data. However, for the same video to be detected, different classification tasks will yield different attribute feature data.

[0051] During implementation, attribute feature data can be generated in the following manner:

[0052] Method 1, which involves encoding the non-content attribute information based on the task content of the classification task to generate attribute feature data, includes: determining the weights of various attribute information in the non-content attribute information based on the task content of the classification task; and encoding the non-content attribute information based on the weights of various attribute information in the non-content attribute information to generate attribute feature data.

[0053] During implementation, the weights of various attributes in the non-content attribute information can be determined based on the task content of the classification task. For example, non-content attribute information can be divided into strongly relevant primary attribute information and weakly relevant secondary attribute information according to their relevance to the task content. The weight of the primary attribute information can be set as the primary weight, and the weight of the secondary attribute information can be set as the secondary weight, where the primary weight is greater than the secondary weight. Alternatively, the attributes in the non-content attribute information can be sorted according to their relevance to the task content, and the weight of each attribute information can be determined according to the sorting result; that is, the greater the relevance, the greater the weight of the attribute information.

[0054] For example, if the task of a classification task is to classify video quality, video quality is generally more related to attributes such as video duration, completion rate, and number of favorites. Therefore, the weights of video duration, completion rate, and number of favorites can be set higher, while the weights of other attributes such as number of likes, number of dislikes, number of comments, and number of plays can be set lower.

[0055] Then, based on the weights of various attributes within the non-content attribute information, the non-content attribute information is encoded to generate attribute feature data. For example, an encoding model can be used to encode non-content attributes.

[0056] Here, by setting the weights of attribute information and encoding non-content attributes according to the weights of various attribute information, more accurate attribute feature data can be generated.

[0057] Method 2: The non-content attribute information is encoded based on the task content of the classification task to generate attribute feature data, including: selecting target attribute information from various attribute information included in the non-content attribute information based on the task content of the classification task; and encoding the selected target attribute information to generate attribute feature data.

[0058] In implementation, based on the task content of the classification task, target attribute information can be selected from various attribute information included in non-content attribute information. This target attribute information is attribute information related to the task content of the classification task. For example, if the task content of the classification task is video quality classification, then video duration, completion rate, and number of favorites can be determined as target attribute information. The selected target attribute information is encoded to generate attribute feature data. As another example, if the target content of the classification task is to predict the future number of views of a video, then the target attribute information can include information such as completion rate and number of shares.

[0059] Here, by selecting target attribute information related to the classification task from the various attribute information included in the content attribute information and encoding the selected target attribute information, the interference of other attribute information besides the target attribute information in the non-content attribute information can be mitigated, thereby improving the accuracy of attribute feature data.

[0060] Method 3: Based on the task content of the classification task, select target attribute information from various attribute information included in non-content attribute information, and determine the weight corresponding to the target attribute information; according to the weight corresponding to the target attribute information, encode the selected target attribute information to generate attribute feature data.

[0061] During implementation, the generation method can be determined according to the task content of the classification task. For example, if the classification task is to detect the quality of the video, attribute feature data can be generated according to method one; if the classification task is to predict the future number of times the video will be played, attribute feature data can be generated according to method two.

[0062] This section provides multiple methods for generating attribute feature data in a flexible manner to meet the needs of different classification tasks.

[0063] Regarding S103:

[0064] In implementation, feature data from multiple modalities and the attribute feature data can be fused. Based on the fused feature data, the classification result of the video to be detected in the classification task can be determined. For example, attribute feature data can be concatenated with video feature data from multiple modalities to obtain concatenated video feature data. Then, based on the concatenated video feature data and text and audio feature data from multiple modalities, the classification result of the video to be detected in the classification task can be determined. The classification result can include the score of the video to be detected in each predicted category. The predicted category is related to the task content of the classification task. For example, if the task content of the classification task is to detect video quality, the predicted category can be high quality, medium quality, low quality, etc.

[0065] In one optional implementation, determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data includes: concatenating the feature data of the multiple modalities and the attribute feature data to generate first concatenated feature data; and processing the first concatenated feature data sequentially using a first fully connected processing layer and a second fully connected processing layer to generate the classification result of the video to be detected under the classification task.

[0066] See Figure 2aAs shown, in implementation, feature data from multiple modalities, namely video feature data, audio feature data, text feature data, and attribute feature data, can first be concatenated to obtain the first concatenated feature data. Then, the first fully connected processing layer is used to process the first concatenated feature data to generate processed feature data, and the second fully connected processing layer is used to process the processed feature data to generate the classification result of the video to be detected under the classification task.

[0067] In another optional implementation, determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data includes: using a first fully connected processing layer to process the feature data of the multiple modalities and the attribute feature data respectively, generating processed feature data of the multiple modalities and processed attribute feature data; concatenating the processed feature data of the multiple modalities and the processed attribute feature data to generate second concatenated feature data; and using a second fully connected processing layer to process the second concatenated feature data to generate the classification result of the video to be detected under the classification task.

[0068] See Figure 2b As shown, the first fully connected processing layer can be used to process the feature data and attribute feature data of multiple modalities respectively, generating processed feature data of multiple modalities and processed attribute feature data; then the processed feature data of multiple modalities and processed attribute feature data are concatenated to generate second concatenated feature data; the second fully connected processing layer is used to process the second concatenated feature data to generate the classification result of the video to be detected under the classification task.

[0069] In another optional implementation, determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data includes: using a first fully connected processing layer to process the feature data of the multiple modalities and the attribute feature data respectively, generating processed feature data of the multiple modalities and processed attribute feature data; using a second fully connected processing layer to process the processed feature data of the multiple modalities and the processed attribute feature data respectively, generating intermediate classification results corresponding to the feature data of each modality and intermediate classification results corresponding to the attribute feature data; and weightedly fusing the intermediate classification results corresponding to the feature data of the multiple modalities and the intermediate classification results corresponding to the attribute feature data to generate the classification result of the video to be detected under the classification task.

[0070] See Figure 2cAs shown, the first fully connected processing layer is used to process the feature data and attribute feature data of multiple modalities, generating processed feature data and attribute feature data for each modality. Then, the first fully connected processing layer is used again to process the processed feature data and attribute feature data for each modality, generating intermediate classification results for each modality's feature data, such as intermediate classification results for video feature data, audio feature data, text feature data, and attribute feature data. The intermediate classification results include the video's score under multiple preset categories.

[0071] Finally, the intermediate classification results corresponding to the feature data of multiple modalities and the intermediate classification results corresponding to the attribute feature data are weighted and fused to generate the classification result of the video to be detected in the classification task (that is, the scores of each intermediate classification result are weighted and fused to obtain the final score). The weights used for weighted fusion can be obtained by training.

[0072] Here, multiple methods for fusion and determination of classification results are set up, which improves the diversity and flexibility of the determination of classification results.

[0073] In specific implementation, the classification result is generated using a trained target classification model, and the feature data of the multiple modalities includes video feature data, text feature data, and audio feature data; the method further includes: obtaining a trained feature extraction sub-model, which includes a video feature extraction sub-model, a text feature extraction sub-model, and an audio feature extraction sub-model; generating an initial classification model based on the feature extraction sub-model and the constructed fusion classification sub-model; obtaining sample data matching the classification task, and adjusting the initial classification model using the sample data to generate the target classification model.

[0074] During implementation, the trained feature extraction sub-models are obtained. The feature extraction sub-models include: video feature extraction sub-model, text feature extraction sub-model, and audio feature extraction sub-model. The feature extraction sub-models are models that have been initially trained. The model structures of the video feature extraction sub-model, text feature extraction sub-model, and audio feature extraction sub-model can be determined as needed.

[0075] The fusion classification sub-model is used to determine the classification result based on feature data and attribute feature data from multiple modalities. The model structure of the fusion classification sub-model can be determined according to the feature fusion method. For example, the fusion classification sub-model may include a first fully connected processing layer, a second fully connected processing layer, convolutional layers, etc. The fusion classification sub-model is an untrained model. Then, the feature extraction sub-model and the constructed fusion classification sub-model are combined to generate the initial classification model.

[0076] To acquire sample data matching the classification task, for example, if the task is to identify medical-related short videos, then various medical-related sample videos can be acquired and labeled. Videos containing labeled tags are then identified as matching the classification task. Alternatively, after acquiring the sample video data, it can be preprocessed to obtain preprocessed sample video data. Preprocessing may include data cleaning (removing duplicate, outlier, or missing data), data transformation (converting non-numerical data to numerical data), or data augmentation (transforming data to increase diversity, such as flipping, rotating, cropping, or deforming image data, or randomly replacing, deleting, or adding text data). Finally, the preprocessed sample video data is labeled, and those containing labeled tags are identified as matching the classification task. Subsequently, by using sample data that matches the classification task, the initial model can be trained, which enables the trained target classification model to perform the reasoning process of the classification task better, improves the reasoning accuracy, and thus obtains a higher accuracy classification result.

[0077] The sample data is then input into the initial classification model to obtain prediction results. Based on the prediction results and labeled data, the loss value is determined. The initial classification model is then adjusted based on this loss value; for example, the model parameters of the feature extraction sub-model are fine-tuned, and the network parameters of the fusion classification sub-model are adjusted to obtain an updated classification model. This model is then iteratively trained multiple times until a training cutoff condition is met. The model obtained from the last training iteration is determined as the target classification model. The training cutoff condition can be set as needed, such as the number of training iterations exceeding a threshold, or the loss value of the classification model being less than a threshold. Since the selected feature extraction sub-model is a pre-trained model with modified network parameters, only fine-tuning of the network parameters of the feature extraction sub-model based on a small amount of sample data is required, improving the efficiency of determining the target classification model.

[0078] After training the target classification model, the feature extraction sub-model can be deployed locally, while the fusion classification sub-model can be deployed online. Specifically, the feature extraction sub-model stored locally extracts features from the video to be detected, obtaining feature data in multiple modalities. Based on the task content of the classification task, the non-content attribute information of the video to be detected is encoded to obtain attribute feature data. Then, the feature data in multiple modalities and the attribute feature data are input into the online-deployed fusion classification sub-model to determine the classification result of the video to be detected. Considering that deploying the entire target classification model online requires reloading the entire model for each inference, resulting in long loading times, low detection efficiency, and a poor user experience, the above deployment method, which deploys the feature extraction sub-model separately and separates it from the subsequent fusion classification sub-model, reduces model loading time, improves model inference speed, shortens the latency when the model goes online, ensures the model's real-time performance and availability, and further improves the model's efficiency and accuracy.

[0079] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0080] Based on the same inventive concept, this disclosure also provides a video classification device corresponding to the video classification method. Since the principle of the device in this disclosure for solving the problem is similar to that of the video classification method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0081] Reference Figure 3 The diagram shown is a schematic representation of the architecture of a video classification device according to an embodiment of this disclosure. The device includes: an acquisition module 301, a generation module 302, and a determination module 303; wherein,

[0082] The acquisition module 301 is used to acquire the video to be detected and the non-content attribute information of the video to be detected;

[0083] The generation module 302 is used to extract features from the video to be detected and generate feature data of multiple modalities that match the content of the video to be detected; and to encode the non-content attribute information based on the task content of the classification task to generate attribute feature data.

[0084] The determination module 303 is used to determine the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data.

[0085] In one possible implementation, the non-content attribute information includes at least one of the following: number of likes, number of dislikes, number of comments, number of reposts, number of favorites, number of shares, video duration, number of plays, and completion rate.

[0086] In one possible implementation, the generation module 302, when encoding the non-content attribute information based on the task content of the classification task to generate attribute feature data, is used for:

[0087] Based on the task content of the classification task, determine the weights of various attribute information in the non-content attribute information;

[0088] Based on the weights of various attribute information in the non-content attribute information, the non-content attribute information is encoded to generate attribute feature data.

[0089] In one possible implementation, the generation module 302, when encoding the non-content attribute information based on the task content of the classification task to generate attribute feature data, is used for:

[0090] Based on the task content of the classification task, target attribute information is selected from the various attribute information included in the non-content attribute information;

[0091] The selected target attribute information is encoded to generate attribute feature data.

[0092] In one possible implementation, the determining module 303, when determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data, is used to:

[0093] The feature data of the multiple modalities and the attribute feature data are concatenated to generate the first concatenated feature data;

[0094] The first spliced ​​feature data is processed sequentially using the first fully connected processing layer and the second fully connected processing layer to generate the classification result of the video to be detected under the classification task.

[0095] In one possible implementation, the determining module 303, when determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data, is used to:

[0096] The first fully connected processing layer is used to process the feature data of the multiple modalities and the attribute feature data respectively, to generate processed feature data of the multiple modalities and processed attribute feature data.

[0097] The processed feature data of multiple modalities and the processed attribute feature data are concatenated to generate second concatenated feature data.

[0098] The second fully connected processing layer is used to process the second stitched feature data to generate the classification result of the video to be detected under the classification task.

[0099] In one possible implementation, the determining module 303, when determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data, is used to:

[0100] The first fully connected processing layer is used to process the feature data of the multiple modalities and the attribute feature data respectively, to generate processed feature data of the multiple modalities and processed attribute feature data.

[0101] Using the second fully connected processing layer, the processed feature data of multiple modalities and the processed attribute feature data are processed respectively to generate intermediate classification results corresponding to the feature data of each modality and intermediate classification results corresponding to the attribute feature data.

[0102] The intermediate classification results corresponding to the feature data of the multiple modalities and the intermediate classification results corresponding to the attribute feature data are weighted and fused to generate the classification result of the video to be detected under the classification task.

[0103] In one possible implementation, the classification result is generated using a trained target classification model, and the feature data of the multiple modalities includes video feature data, text feature data, and audio feature data; the device further includes: a training module 304, used for:

[0104] Obtain trained feature extraction sub-models, including: video feature extraction sub-model, text feature extraction sub-model, and audio feature extraction sub-model;

[0105] Based on the feature extraction sub-model and the constructed fusion classification sub-model, an initial classification model is generated;

[0106] Obtain sample data that matches the classification task, and use the sample data to adjust the initial classification model to generate the target classification model.

[0107] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0108] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 4 The diagram shows the structure of a computer device 400 provided in this embodiment, including a processor 401, a memory 402, and a bus 403. The memory 402 stores execution instructions and includes main memory 4021 and external memory 4022. The main memory 4021, also called internal memory, is used to temporarily store computational data in the processor 401 and data exchanged with external memory 4022 such as a hard disk. The processor 401 exchanges data with the external memory 4022 through the main memory 4021. When the computer device 400 is running, the processor 401 and the memory 402 communicate through the bus 403, causing the processor 401 to execute the following instructions:

[0109] Obtain the video to be detected and its non-content attribute information;

[0110] Feature extraction is performed on the video to be detected to generate feature data of multiple modalities that match the content of the video to be detected; and based on the task content of the classification task, the non-content attribute information is encoded to generate attribute feature data.

[0111] Based on the feature data of the multiple modalities and the attribute feature data, the classification result of the video to be detected under the classification task is determined.

[0112] In one optional implementation, the non-content attribute information in the instructions executed by the processor 401 includes at least one of the following: number of likes, number of dislikes, number of comments, number of reposts, number of favorites, number of shares, video duration, number of plays, and completion rate.

[0113] In one optional implementation, the instructions executed by processor 401, which involve encoding the non-content attribute information based on the task content of the classification task to generate attribute feature data, include:

[0114] Based on the task content of the classification task, determine the weights of various attribute information in the non-content attribute information;

[0115] Based on the weights of various attribute information in the non-content attribute information, the non-content attribute information is encoded to generate attribute feature data.

[0116] In one optional implementation, the instructions executed by processor 401, which involve encoding the non-content attribute information based on the task content of the classification task to generate attribute feature data, include:

[0117] Based on the task content of the classification task, target attribute information is selected from the various attribute information included in the non-content attribute information;

[0118] The selected target attribute information is encoded to generate attribute feature data.

[0119] In one optional implementation, the instructions executed by processor 401, wherein determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data, includes:

[0120] The feature data of the multiple modalities and the attribute feature data are concatenated to generate the first concatenated feature data;

[0121] The first spliced ​​feature data is processed sequentially using the first fully connected processing layer and the second fully connected processing layer to generate the classification result of the video to be detected under the classification task.

[0122] In one optional implementation, the instructions executed by processor 401, wherein determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data, includes:

[0123] The first fully connected processing layer is used to process the feature data of the multiple modalities and the attribute feature data respectively, to generate processed feature data of the multiple modalities and processed attribute feature data.

[0124] The processed feature data of multiple modalities and the processed attribute feature data are concatenated to generate second concatenated feature data.

[0125] The second fully connected processing layer is used to process the second stitched feature data to generate the classification result of the video to be detected under the classification task.

[0126] In one optional implementation, the instructions executed by processor 401, wherein determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data, includes:

[0127] The first fully connected processing layer is used to process the feature data of the multiple modalities and the attribute feature data respectively, to generate processed feature data of the multiple modalities and processed attribute feature data.

[0128] Using the second fully connected processing layer, the processed feature data of multiple modalities and the processed attribute feature data are processed respectively to generate intermediate classification results corresponding to the feature data of each modality and intermediate classification results corresponding to the attribute feature data.

[0129] The intermediate classification results corresponding to the feature data of the multiple modalities and the intermediate classification results corresponding to the attribute feature data are weighted and fused to generate the classification result of the video to be detected under the classification task.

[0130] In one optional implementation, the classification result in the instructions executed by the processor 401 is generated using a trained target classification model, and the feature data of the multiple modalities includes video feature data, text feature data, and audio feature data; the method further includes:

[0131] Obtain trained feature extraction sub-models, including: video feature extraction sub-model, text feature extraction sub-model, and audio feature extraction sub-model;

[0132] Based on the feature extraction sub-model and the constructed fusion classification sub-model, an initial classification model is generated;

[0133] Obtain sample data that matches the classification task, and use the sample data to adjust the initial classification model to generate the target classification model.

[0134] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the video classification method described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.

[0135] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the video classification method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0136] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0137] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0138] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0139] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0140] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0141] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A video classification method, characterized in that, include: Obtain the video to be detected and its non-content attribute information; Feature extraction is performed on the video to be detected to generate feature data of multiple modalities that match the content of the video to be detected; and based on the task content of the classification task, the non-content attribute information is encoded to generate attribute feature data, wherein the attribute feature data matches the classification task, and the encoding of the non-content attribute information is based on the correlation between the non-content attribute information and the task content. Based on the feature data of the multiple modalities and the attribute feature data, the classification result of the video to be detected under the classification task is determined.

2. The method according to claim 1, characterized in that, The non-content attribute information includes at least one of the following: number of likes, number of dislikes, number of comments, number of reposts, number of favorites, number of shares, video duration, number of plays, and completion rate.

3. The method according to claim 1 or 2, characterized in that, The task content based on the classification task encodes the non-content attribute information to generate attribute feature data, including: Based on the task content of the classification task, determine the weights of various attribute information in the non-content attribute information; Based on the weights of various attribute information in the non-content attribute information, the non-content attribute information is encoded to generate attribute feature data.

4. The method according to claim 1 or 2, characterized in that, The task content based on the classification task encodes the non-content attribute information to generate attribute feature data, including: Based on the task content of the classification task, target attribute information is selected from the various attribute information included in the non-content attribute information; The selected target attribute information is encoded to generate attribute feature data.

5. The method according to claim 1, characterized in that, The process of determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data includes: The feature data of the multiple modalities and the attribute feature data are concatenated to generate the first concatenated feature data; The first spliced ​​feature data is processed sequentially using the first fully connected processing layer and the second fully connected processing layer to generate the classification result of the video to be detected under the classification task.

6. The method according to claim 1, characterized in that, The process of determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data includes: The first fully connected processing layer is used to process the feature data of the multiple modalities and the attribute feature data respectively, to generate processed feature data of the multiple modalities and processed attribute feature data. The processed feature data of multiple modalities and the processed attribute feature data are concatenated to generate second concatenated feature data. The second fully connected processing layer is used to process the second stitched feature data to generate the classification result of the video to be detected under the classification task.

7. The method according to claim 1, characterized in that, The process of determining the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data includes: The first fully connected processing layer is used to process the feature data of the multiple modalities and the attribute feature data respectively, to generate processed feature data of the multiple modalities and processed attribute feature data. Using the second fully connected processing layer, the processed feature data of multiple modalities and the processed attribute feature data are processed respectively to generate intermediate classification results corresponding to the feature data of each modality and intermediate classification results corresponding to the attribute feature data. The intermediate classification results corresponding to the feature data of the multiple modalities and the intermediate classification results corresponding to the attribute feature data are weighted and fused to generate the classification result of the video to be detected under the classification task.

8. The method according to claim 1, characterized in that, The classification result is generated using a trained target classification model, and the feature data of the multiple modalities includes video feature data, text feature data, and audio feature data. The method further includes: Obtain trained feature extraction sub-models, including: video feature extraction sub-model, text feature extraction sub-model, and audio feature extraction sub-model; Based on the feature extraction sub-model and the constructed fusion classification sub-model, an initial classification model is generated; Obtain sample data that matches the classification task, and use the sample data to adjust the initial classification model to generate the target classification model.

9. A video classification device, characterized in that, include: The acquisition module is used to acquire the video to be detected and the non-content attribute information of the video to be detected; The generation module is used to extract features from the video to be detected and generate feature data of multiple modalities that match the content of the video to be detected; and to encode the non-content attribute information based on the task content of the classification task to generate attribute feature data, wherein the attribute feature data matches the classification task, and the encoding of the non-content attribute information is based on the correlation between the non-content attribute information and the task content. The determination module is used to determine the classification result of the video to be detected under the classification task based on the feature data of the multiple modalities and the attribute feature data.

10. A computer device, characterized in that, include: The computer device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the video classification method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the video classification method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Video classification processing method and device, computer equipment and storage medium

    CN110162669A

  • Video classification method and device

    CN110334689A

  • Video screening method and device, electronic equipment and storage medium

    CN110990631A

  • Task processing model determination method, video category determination method and device

    CN115098725A