Video classification method and device, storage medium and electronic device

By using the visual Transformer layer to extract image features and language representation models in the video classification method, the title information features are extracted and fused, and the problem of low video classification accuracy in the prior art is solved, and higher video classification accuracy is achieved.

CN114494942BActive Publication Date: 2025-05-09BEIJING OPPO TELECOMM CORP LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111612621.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-05-09
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

The video classification method in the prior art has low classification accuracy and cannot effectively utilize the multimodal features in video.

Method used

The visual Transformer layer is used to extract feature of multiple target images, and the title information is extracted in combination with the language representation model. The target feature vector is obtained by fusing the two, thereby determining the classification results of the video.

Benefits of technology

By capturing the global image features and title information features in the video, the accuracy of video classification is significantly improved, making the classification results more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494942B_ABST
    Figure CN114494942B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of image and video processing technology, and specifically to a video classification method and device, a computer-readable storage medium, and an electronic device, wherein the method comprises: obtaining a target video and title information of the target video; obtaining multiple target images in the target video; extracting features of the multiple target images using a visual Transformer layer to obtain a first feature vector corresponding to the target video; extracting features of the title information using a language representation model to obtain a second feature vector corresponding to the target video; determining a target feature vector of the target video based on the first feature vector and the second feature vector; and determining a target classification result of the target video based on the target feature vector. The technical solution of the embodiment of the present disclosure improves the accuracy of the video classification result.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the rapid development of the Internet, Internet videos have entered a new stage of explosive growth. The massive amount of video data has also put forward higher requirements for common related technologies such as video processing, classification, and recommendation.

[0003] The classification accuracy of the video classification method in the prior art is low, so it is necessary to propose a new video classification method.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0005] The purpose of the present disclosure is to provide a video classification method, a video classification device, a computer-readable medium and an electronic device, thereby improving the accuracy of video classification results at least to a certain extent.

[0006] According to a first aspect of the present disclosure, a video classification method is provided, including: obtaining a target video and title information of the target video; obtaining multiple target images in the target video; using a visual Transformer layer to perform feature extraction on the multiple target images to obtain a first feature vector corresponding to the target video; using a language representation model to perform feature extraction on the title information to obtain a second feature vector corresponding to the target video; determining a target feature vector of the target video based on the first feature vector and the second feature vector; and determining a target classification result of the target video based on the target feature vector.

[0007] According to a second aspect of the present disclosure, a video classification device is provided, including: a first acquisition module, used to acquire a target video and title information of the target video; a second acquisition module, used to acquire multiple target images in the target video; a first feature extraction module, used to perform feature extraction on the multiple target images using a visual Transformer layer to obtain a first feature vector corresponding to the target video; a second feature extraction module, used to perform feature extraction on the title information using a language representation model to obtain a second feature vector corresponding to the target video; a feature fusion module, used to determine a target feature vector of the target video based on the first feature vector and the second feature vector; and a video classification module, used to determine a classification result of the target video based on the target feature vector.

[0008] According to a third aspect of the present disclosure, a computer-readable medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented.

[0009] According to a fourth aspect of the present disclosure, an electronic device is provided, characterized in that it includes: one or more processors; and a memory for storing one or more programs, which enables the one or more processors to implement the above method when the one or more programs are executed by the one or more processors.

[0010] An embodiment of the present disclosure provides a video classification method, which obtains a target video and title information of the target video; obtains multiple target images in the target video; uses a visual Transformer layer to perform feature extraction on the multiple target images to obtain a first feature vector corresponding to the target video; uses a language representation model to perform feature extraction on the title information to obtain a second feature vector corresponding to the target video; determines a target feature vector of the target video based on the first feature vector and the second feature vector; and determines a target classification result of the target video based on the target feature vector. Compared with the prior art, using a visual Transformer layer to perform feature extraction on multiple target images can more accurately capture the global features of the images in the video and improve the accuracy of video classification. Furthermore, the second feature vector corresponding to the title information is fused with the first feature vector to obtain a target feature vector, making video classification more accurate.

[0011] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0013] Figure 1 A schematic diagram showing an exemplary system architecture to which embodiments of the present disclosure may be applied;

[0014] Figure 2 A flowchart of a video classification method in an exemplary embodiment of the present disclosure is schematically shown;

[0015] Figure 3 A flowchart for determining a target classification result in an exemplary embodiment of the present disclosure is schematically shown;

[0016] Figure 4 A data flow diagram for training a multi-head self-attention mechanism network in an exemplary embodiment of the present disclosure is schematically shown;

[0017] Figure 5 A flowchart for updating target classification results in an exemplary embodiment of the present disclosure is schematically shown;

[0018] Figure 6 A schematic diagram schematically shows the composition of a video classification device in an exemplary embodiment of the present disclosure;

[0019] Figure 7 A schematic diagram schematically shows the composition of another video classification device in an exemplary embodiment of the present disclosure;

[0020] Figure 8 A schematic diagram of an electronic device to which the embodiments of the present disclosure can be applied is shown. DETAILED DESCRIPTION

[0021] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0022] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0023] Figure 1 The schematic diagram of the system architecture is shown, and the system architecture 100 may include a terminal 110 and a server 120. The terminal 110 may be a terminal device such as a smart phone, a tablet computer, a desktop computer, a laptop computer, etc. The server 120 generally refers to a background system that provides video classification related services in this exemplary embodiment, and may be a server or a cluster formed by multiple servers. The terminal 110 and the server 120 may be connected via a wired or wireless communication link to perform data exchange.

[0024] In one implementation, the video classification method may be performed by the terminal 110. For example, a user uses the terminal 110 to shoot an image or the user selects a target video and the title information of the target video in the album of the terminal 110, and the terminal 110 classifies the image and outputs the classification result.

[0025] In one implementation, the video classification method may be performed by the server 120. For example, a user uses the terminal 110 to shoot an image or the user selects a target video and the title information of the target video in the photo album of the terminal 110, and the terminal 110 uploads the target video and the title information of the target video to the server 120, and the server 120 classifies the target video and returns the classification result to the terminal 110.

[0026] As can be seen from the above, the execution subject of the video classification method in this exemplary embodiment may be the terminal 110 or the server 120, and the present disclosure does not limit this.

[0027] The exemplary embodiment of the present disclosure also provides an electronic device for executing the above-mentioned video classification method, and the electronic device may be the above-mentioned terminal 110 or the server 120. Generally, the electronic device may include a processor and a memory, the memory is used to store executable instructions of the processor, and the processor is configured to execute the above-mentioned image video classification method by executing the executable instructions.

[0028] In related technologies, deep networks are usually used for video understanding tasks. The main solution is to use ResNet50 to extract the single-modal feature of video images for video classification. The main disadvantages are that ResNet50 is a CNN network structure. The CNN network is composed of multiple CNN blocks stacked together. The features extracted from the feature map are local features of the original image, and the connection between the local features of the original image is not fully utilized. Video is a multimodal form composed of images, texts, and voices. Only the features of video images are used for classification, and the features of other modes of the video are not fully utilized. In addition, this solution performs random data enhancement on each frame of the video to reduce the correlation between the time series of video images. Or the GAN model is used to process the video source data to obtain similar new source data to increase the training data. In addition, the single-modal image features extracted by ResNet are given to LSTM for video classification. The training of LSTM has very high hardware requirements and requires storage bandwidth binding calculation, which makes training difficult and has low applicability.

[0029] Combine the following Figure 2 The image quality evaluation method in this exemplary embodiment is described. Figure 2 An exemplary process of the image quality assessment method is shown, which may include:

[0030] Step S210, obtaining a target video and title information of the target video;

[0031] Step S220, acquiring multiple target images in the target video;

[0032] Step S230, using a visual Transformer layer to extract features from the plurality of target images to obtain a first feature vector corresponding to the target video;

[0033] Step S240, extracting features from the title information using a language representation model to obtain a second feature vector corresponding to the target video;

[0034] Step S250, determining a target feature vector of the target video according to the first feature vector and the second feature vector;

[0035] Step S260: determining a target classification result of the target video according to the target feature vector.

[0036] Based on the above method, compared with the existing technology, the use of the visual Transformer layer to extract features from multiple target images can more accurately capture the global features of the images in the video and improve the video classification accuracy. Furthermore, the second feature vector corresponding to the title information is fused with the first feature vector to obtain the target feature vector, making the video classification more accurate.

[0037] Below Figure 2 Each step in the procedure is described in detail.

[0038] refer to Figure 2 In step S210, a target video and title information of the target video are obtained.

[0039] In an exemplary implementation of the present disclosure, the processor may obtain a target video to be processed from a database, and determine title information of the target video.

[0040] In step S220, a plurality of target images are acquired in the target video;

[0041] In this example implementation, the processor can obtain multiple target images in the target video. Specifically, the processor can obtain multiple target images in the target video at intervals of a first preset time. When obtaining, the processor can first determine the first preset time and the total duration of the target video, and then determine the number of the target images. The first preset time can be 10 milliseconds, 1 second, 10 seconds, etc., and can also be customized according to user needs, which is not specifically limited in this example implementation.

[0042] In another example implementation, the number of the target images and the total duration of the target video may be determined first, and then the first preset time may be calculated, wherein the vectors of the target images may be 10, 20, etc., and may also be customized according to user needs, which is not specifically limited in this example implementation.

[0043] After acquiring the target image, data enhancement can be performed on the image, including but not limited to geometric transformation enhancement and color transformation enhancement. Geometric transformation enhancement includes: flipping, rotating, cropping, deforming and scaling each frame of the advertising video. Color transformation enhancement includes: noise transformation, blur transformation and color transformation for each frame of the advertising video.

[0044] In step S230, a visual Transformer layer is used to extract features from the plurality of target images to obtain a first feature vector corresponding to the target video.

[0045] In this example implementation, the visual Transformer layer can be used to extract features from multiple target images to obtain a first feature vector corresponding to the target video. Specifically, the Q matrix, K matrix and V matrix corresponding to the autonomous attention mechanism unit in the Transformer layer can be determined based on the target image. Then, the target image can be divided into regions. When dividing the regions, the target image can be divided into 9 regions, or it can be customized according to user needs. When dividing the regions, the target image can be divided by average division or by sliding window division, which is not specifically limited in this example implementation.

[0046] In this example implementation, after completing the region division of the target image, the target image can be processed by an autonomous attention mechanism unit, and the first feature vector is obtained through the normalization layer of the visual Transformer layer, the fusion layer, and the perception layer.

[0047] In another example implementation of the present disclosure group, when the above-mentioned visual Transformer layer is used to extract features of the target image, the time information of each of the above-mentioned target images can also be obtained, and the time information is synchronously input into the above-mentioned visual Transformer layer, so that the first feature vector obtained includes the temporal correlation between multiple images, thereby further improving the video classification accuracy.

[0048] The above steps are described in detail below through a specific embodiment.

[0049] Taking multiple images as a batch (concurrency number) as an example, the size of the target image read into VIT (Vision Transformer, visual Transformer layer) is (1, N, C, W, H), that is, N target images are input, usually C = 3, that is, RGB image, H and W are the height and width of the target image, which are fixed values, such as 256; (1, N, C, W, H) can be first transformed into a tensor of (1*N, C, W, H) size and sent to VIT to extract image features, and obtain image features of (1*N, 768) size; finally, (1*N, 768) is transformed into the first feature image of (1, N, 768) size.

[0050] In step S240, a language representation model is used to extract features from the title information to obtain a second feature vector corresponding to the target video.

[0051] In this example implementation, the language representation model can be a Chinese_base_BERT model, which is a Bert network used for feature extraction of Chinese. It can also be customized according to user needs, which is not specifically limited in this example implementation.

[0052] The processor may input the title information into the language representation model to obtain a second feature vector corresponding to the target video.

[0053] For example, the Bert network reads the text information of the target video title information (such as K characters), and now embeds it into (1, 56), where 56 is the longest length of the video title obtained by statistics, and can also be customized according to needs; it can be input into the chinese_base_bert model together with the text vector of all zeros (1, 56) and the position vector of all ones (1, 56) to obtain a second feature vector of (1, 56) size, and then the second feature vector can be dimensionally upgraded according to the above first feature vector to convert the above second feature vector into a size of (1, 1, 768) to facilitate fusion with the above first feature vector.

[0054] In step S250, a target feature vector of the target video is determined according to the first feature vector and the second feature vector.

[0055] In this example implementation, after obtaining the first feature vector and the second feature vector, the second feature vector may be dimensionally upgraded according to the first feature vector to facilitate fusion with the first feature vector. For example, when the size of the first feature vector is (1, N, 768) and the size of the second feature vector is (1, 56), the second feature vector may be dimensionally upgraded to (1, 1, 768). Then, the first feature vector and the second feature vector are fused to obtain the target feature vector.

[0056] In step S260, a target classification result of the target video is determined according to the target feature vector.

[0057] In this example implementation, determining the target classification result of the target video according to the target feature vector may include steps S310 to S320, which are described in detail below.

[0058] In step S310, a pre-trained video classification model is obtained.

[0059] In this example implementation, when obtaining a pre-trained video classification model, an initial model can be first obtained. The initial model can be a CNN model, a multi-head self-attention mechanism network, or other networks. It can also be customized according to user needs and is not specifically limited in this example implementation.

[0060] The following is an example of the above initial model being a multi-head self-attention mechanism network. After obtaining the multi-head self-attention mechanism network, multiple reference videos and real labels corresponding to the reference videos can be obtained, wherein the real labels can be action labels (such as playing tennis), scene labels (such as beaches), and object labels (such as cars), and various custom labels can also be annotated according to actual applications. The reference videos and the real labels corresponding to the reference videos are used as training data.

[0061] In this example implementation, after the training data is obtained, the title information of each reference video is obtained, and a plurality of initial images are obtained in the reference video; Figure 4As shown, the above-mentioned initial image can be first divided into regions, for example, into 9 regions, and then the divided images can be used to extract features from multiple initial images using the visual Transformer layer 410 to obtain the fourth eigenvector corresponding to the reference video; the title information 420 can be extracted using the language representation model to obtain the fifth eigenvector corresponding to the reference video, wherein the above-mentioned language representation model can be the chinese_base_bert layer 430; the final eigenvector of the reference video is determined according to the fourth eigenvector and the fifth eigenvector, that is, the final eigenvector is obtained according to the fourth eigenvector and the fifth eigenvector using the fusion module 440; the multi-head self-attention mechanism network is trained according to the final eigenvector and the true label to obtain the video classification model.

[0062] In this example implementation, refer to Figure 4 As shown, the cross entropy and softmax loss function 460 can be used to update the parameters in the multi-head self-attention mechanism network 450 to obtain the video classification model.

[0063] For example, in the above multi-head self-attention mechanism network, assuming that the above final feature vector is X, let Q = K = V = X, d k is the scaling factor. Self-Attention is the process of calculating the importance of each part of the sequence tensor X, that is, the softmax part, and then weighting it with itself V. Multi-HeadAttention combines the results of multiple self-Attentions and plays a voting-like role. In the above figure, the first eigenvector obtained by VIT and the second eigenvector obtained by Bert are combined. The use of Multi-head Attention fully reflects the importance of the N-frame target images and title information in a video, and at the same time, it can utilize the temporal relationship between the N target images in the video, which can better perform video classification tasks.

[0064] In step S320, a target classification result of the target video is determined based on the target feature vector using a trained video classification model.

[0065] After obtaining the above video classification model, the above target feature vector may be input into the above video classification model to obtain the target classification result of the above target video.

[0066] In an exemplary embodiment of the present disclosure, referring to Figure 5 As shown, the above method may further include step S510 to step S550. The above steps are described in detail below.

[0067] In step S510, at least one set of image sets is obtained in the target video, each set of the image sets including a plurality of reference images;

[0068] In this example implementation, at least one set of image sets can be acquired in the target video, and the image set includes a plurality of reference images. Specifically, a plurality of image sets can be acquired in the target video at intervals of different second preset times, wherein the second preset time is different from the first preset time.

[0069] In this example implementation, the number of the above-mentioned image sets can be two groups, three groups, etc., and can also be customized according to user needs, which is not specifically limited in this example implementation.

[0070] In step S520, a visual Transformer layer is used to perform feature extraction on each of the image sets to obtain at least one third feature vector corresponding to the target video.

[0071] In this example implementation, the visual Transformer layer can be used to perform feature extraction on multiple reference images to obtain a third feature vector corresponding to the target video. Specifically, the Q matrix, K matrix and V matrix corresponding to the autonomous attention mechanism unit in the Transformer layer can be determined based on the reference image, and then the reference image can be divided into regions. When dividing the regions, the reference image can be divided into 9 regions, or it can be customized according to user needs. When dividing the regions, the average division method can be used, or the sliding window method can be used to divide the reference image into regions, which is not specifically limited in this example implementation.

[0072] In step S530, at least one reference feature vector of the target video is determined according to the third feature vector and the second feature vector.

[0073] In this example implementation, the second eigenvector may first be subjected to a dimensionality increase operation based on the third eigenvector. The specific process of the dimensionality increase operation has been described in detail above, and therefore will not be described again here.

[0074] After the second eigenvector is dimensionally upgraded, the third eigenvector and the second eigenvector may be concatenated and fused to obtain the reference eigenvector.

[0075] In step S540, a reference classification result of the target video is determined according to at least one reference feature vector.

[0076] In this example implementation, in this example implementation, the reference feature vector may be input into the video classification model to obtain at least one reference classification result of the target video.

[0077] In step S550, the target classification result is updated using the reference classification result.

[0078] In this example implementation, after obtaining the above-mentioned multiple reference classification results, the above-mentioned target classification results are updated according to the above-mentioned multiple reference classification results. For example, assuming that the number of the above-mentioned reference classification results is 2, namely news tags and sports tags, if the above-mentioned target classification result is a news tag, the news tag can be used as the above-mentioned target classification result.

[0079] In another example implementation, the weight information of the reference classification labels and the weight of the target classification result may be determined, and the target classification result of the target video screen may be determined according to the weights.

[0080] In summary, in this exemplary embodiment, compared with the prior art, the use of a visual Transformer layer to extract features from multiple target images can more accurately capture the global features of the images in the video, improve the accuracy of video classification, and further, the second feature vector corresponding to the title information is fused with the first feature vector to obtain the target feature vector, so that the video classification is more accurate. Further, when the visual Transformer layer is used to extract features from multiple target images, the information of each target image is input, the temporal correlation between each target image is increased, and the classification accuracy is enhanced. Further, a multi-head self-attention mechanism network is used to train the video classification model. When training the model, the amount of data calculated is small, and the obtained video classification model has a high accuracy in video classification. Further, the same target image is classified multiple times, and the target classification result is determined based on the multiple classification results, which prevents misclassification, increases the fault tolerance rate, and further improves the accuracy of video classification.

[0081] It should be noted that the above figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not intended to be limiting. It is easy to understand that the processes shown in the above figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.

[0082] For further reference, Figure 6As shown, in the implementation of this example, a video classification device 600 is also provided, including a first acquisition module 610, a second acquisition module 620, a first feature extraction model 630, a second feature extraction model 640, a feature fusion module 650 and a video classification module 660. Among them:

[0083] The first acquisition module 610 can be used to acquire the target video and the title information of the target video.

[0084] The second acquisition module 620 may be configured to acquire a plurality of target images in the target video. Specifically, the plurality of target images may be acquired in the target video at intervals of a first preset time.

[0085] The first feature extraction model 630 can be used to use the visual Transformer layer to perform feature extraction on multiple target images to obtain a first feature vector corresponding to the target video. Specifically, the time information of the multiple target images is obtained; and the visual Transformer layer is used to perform feature extraction on the multiple target images according to the time information to obtain the first feature vector corresponding to the target video.

[0086] The second feature extraction model 640 may be used to extract features from the title information using a language representation model to obtain a second feature vector corresponding to the target video.

[0087] The feature fusion module 650 may be configured to determine a target feature vector of a target video according to the first feature vector and the second feature vector.

[0088] The video classification module 660 can be used to determine the target classification result of the target video according to the target feature vector. Specifically, a pre-trained video classification model can be first obtained; and then the target classification result of the target video can be determined according to the target feature vector using the trained video classification model.

[0089] In an example implementation of the present disclosure, when obtaining a pre-trained video classification model, a multi-head self-attention mechanism network can be first obtained; multiple reference videos and real labels corresponding to the reference videos are obtained; title information of the reference video is obtained, and multiple initial images are obtained in the reference video; the visual Transformer layer is used to extract features from the multiple initial images to obtain a fourth feature vector corresponding to the reference video; the language representation model is used to extract features from the title information to obtain a fifth feature vector corresponding to the reference video; the final feature vector of the reference video is determined based on the fourth feature vector and the fifth feature vector; and the multi-head self-attention mechanism network is trained based on the final feature vector and the real labels to obtain a video classification model.

[0090] In an exemplary embodiment of the present disclosure, referring to Figure 7As shown, the video classification device 600 may further include a result update module 670 for updating the target classification result. Specifically, the result update module 670 obtains at least one set of image sets in the target video, each set of image sets includes multiple reference images; uses the visual Transformer layer to extract features from each image set to obtain at least one third feature vector corresponding to the target video; determines at least one reference feature vector of the target video based on the third feature vector and the second feature vector; determines a reference classification result of the target video based on at least one reference feature vector; and updates the target classification result using the reference classification result.

[0091] The specific details of each module in the above device have been described in detail in the implementation method of the method part. The undisclosed details can be found in the implementation method of the method part, so they will not be repeated here.

[0092] Below Figure 8 The mobile terminal 800 in FIG. 1 is taken as an example to exemplify the structure of the electronic device. It should be understood by those skilled in the art that, in addition to the components specifically used for mobile purposes, Figure 8 The construction in can also be applied to fixed type equipment.

[0093] like Figure 8 As shown, the mobile terminal 800 may specifically include: a processor 801, a memory 802, a bus 803, a mobile communication module 804, an antenna 1, a wireless communication module 805, an antenna 2, a display screen 806, a camera module 807, an audio module 808, a power module 809 and a sensor module 810.

[0094] The processor 201 may include one or more processing units, for example, the processor 801 may include an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor and / or an NPU (Neural-Network Processing Unit), etc. The video classification method in this exemplary embodiment may be executed by an AP, a GPU or a DSP, and when the method involves processing related to a neural network, it may be executed by an NPU.

[0095] The encoder can encode (i.e. compress) an image or video, for example, the target image can be encoded into a specific format to reduce the data size for easy storage or transmission. The decoder can decode (i.e. decompress) the encoded data of the image or video to restore the image or video data, such as reading the encoded data of the target image, decoding it through the decoder to restore the data of the target image, and then perform relevant processing on the data for video classification. The mobile terminal 800 can support one or more encoders and decoders. In this way, the mobile terminal 800 can process images or videos in multiple encoding formats, such as: JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (Bitmap) and other image formats, MPEG (Moving Picture Experts Group) 1, MPEG2, H.263, H.264, HEVC (High Efficiency Video Coding) and other video formats.

[0096] The processor 801 may be connected to the memory 802 or other components via a bus 803 .

[0097] The memory 802 may be used to store computer executable program codes, which include instructions. The processor 801 executes various functional applications and data processing of the mobile terminal 800 by running the instructions stored in the memory 802. The memory 802 may also store application data, such as images, videos and other files.

[0098] The communication function of the mobile terminal 800 can be implemented by the mobile communication module 804, antenna 1, wireless communication module 805, antenna 2, modulation and demodulation processor and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module 804 can provide 2G, 3G, 4G, 5G and other mobile communication solutions applied to the mobile terminal 800. The wireless communication module 805 can provide wireless communication solutions such as wireless LAN, Bluetooth, near field communication, etc. applied to the mobile terminal 800.

[0099] The display screen 806 is used to implement display functions, such as displaying user interfaces, images, and videos. The camera module 807 is used to implement shooting functions, such as shooting images and videos. The audio module 808 is used to implement audio functions, such as playing audio and collecting voice. The power module 809 is used to implement power management functions, such as charging the battery, powering the device, and monitoring the battery status. The sensor module 810 may include a depth sensor 8101, a pressure sensor 8102, a gyroscope sensor 8103, an air pressure sensor 8104, etc., to implement corresponding sensing detection functions.

[0100] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as "circuits", "modules" or "systems".

[0101] The exemplary embodiments of the present disclosure also provide a computer-readable storage medium on which a program product capable of implementing the above-mentioned method of the present specification is stored. In some possible implementations, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section of the present specification, for example, it can execute Figures 3 to 5 Any one or more steps in .

[0102] It should be noted that the computer-readable medium shown in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0103] In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0104] In addition, program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0105] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and examples are to be considered as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.

[0106] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A video classification method, characterized in that: include: Obtaining a target video and title information of the target video; Acquire multiple target images in the target video; Using a visual Transformer layer to extract features from the plurality of target images to obtain a first feature vector corresponding to the target video; The step of extracting features from the plurality of target images using the visual Transformer layer to obtain a first feature vector corresponding to the target video includes: obtaining time information of the plurality of target images; extracting features from the plurality of target images using the visual Transformer layer according to the time information to obtain a first feature vector corresponding to the target video; the first feature vector includes a temporal association of the plurality of images; Using a language representation model to extract features from the title information to obtain a second feature vector corresponding to the target video; Determine the target feature vector of the target video according to the first feature vector and the second feature vector; the determining the target feature vector of the target video according to the first feature vector and the second feature vector comprises: performing dimension-upgrading processing on the second feature vector according to the first feature vector to obtain an intermediate feature vector; and concatenating and fusing the first feature vector and the intermediate feature vector to obtain a target feature vector; Determine a target classification result of the target video according to the target feature vector; The method further comprises: Acquire at least one set of image sets in the target video, each set of the image sets including a plurality of reference images; Using a visual Transformer layer to extract features from each of the image sets to obtain at least one third feature vector corresponding to the target video; Determine at least one reference feature vector of the target video according to the third feature vector and the second feature vector; the step of determining at least one reference feature vector of the target video according to the third feature vector and the second feature vector comprises: performing a dimension-up operation on the second feature vector according to the third feature vector, and concatenating and fusing the third feature vector with the dimension-upgraded second feature vector to obtain a reference feature vector; Determine a reference classification result of the target video according to at least one reference feature vector; The target classification result is updated using the reference classification result.

2. The method according to claim 1, characterized in that The acquiring of a plurality of target images in the target video comprises: A plurality of target images are acquired in the target video at intervals of a first preset time.

3. The method according to claim 1, characterized in that Determining the target classification result of the target video according to the target feature vector includes: Get a pre-trained video classification model; The target classification result of the target video is determined according to the target feature vector using a trained video classification model.

4. The method according to claim 3, characterized in that The obtaining of the pre-trained video classification model comprises: Get a multi-head self-attention mechanism network; Obtain multiple reference videos and real labels corresponding to the reference videos; Obtaining title information of a reference video, and obtaining a plurality of initial images in the reference video; Using a visual Transformer layer to extract features from the plurality of initial images to obtain a fourth feature vector corresponding to the reference video; Using a language representation model to extract features from the title information to obtain a fifth feature vector corresponding to the reference video; Determining a final feature vector of the reference video according to the fourth feature vector and the fifth feature vector; The multi-head self-attention mechanism network is trained according to the final feature vector and the true label to obtain the video classification model.

5. A video classification device, characterized in that: include: A first acquisition module, used to acquire a target video and title information of the target video; A second acquisition module is used to acquire multiple target images in the target video; A first feature extraction module is used to extract features from the plurality of target images using a visual Transformer layer to obtain a first feature vector corresponding to the target video; the first feature extraction module is configured to: obtain time information of the plurality of target images; extract features from the plurality of target images using the visual Transformer layer according to the time information to obtain a first feature vector corresponding to the target video; the first feature vector includes a temporal association of the plurality of images; A second feature extraction module, used to extract features from the title information using a language representation model to obtain a second feature vector corresponding to the target video; A feature fusion module is used to determine a target feature vector of the target video according to the first feature vector and the second feature vector; the feature fusion module is configured to: perform dimension-upgrading processing on the second feature vector according to the first feature vector to obtain an intermediate feature vector; and concatenate and fuse the first feature vector with the intermediate feature vector to obtain a target feature vector; A video classification module, used to determine a target classification result of the target video according to the target feature vector; The device is also configured to: Acquire at least one set of image sets in the target video, each set of the image sets including a plurality of reference images; Using a visual Transformer layer to extract features from each of the image sets to obtain at least one third feature vector corresponding to the target video; Determine at least one reference feature vector of the target video according to the third feature vector and the second feature vector; the step of determining at least one reference feature vector of the target video according to the third feature vector and the second feature vector comprises: performing a dimension-up operation on the second feature vector according to the third feature vector, and concatenating and fusing the third feature vector with the dimension-upgraded second feature vector to obtain a reference feature vector; Determine a reference classification result of the target video according to at least one reference feature vector; The target classification result is updated using the reference classification result.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the video classification method according to any one of claims 1 to 4 is implemented.

7. An electronic device, characterized in that: include: one or more processors; as well as A memory for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the video classification method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Label data processing method and device and computer readable storage medium

    CN111711869A

  • Method and device for detecting image by adopting target detection model, equipment and medium

    CN113222916A

  • Video processing method and device, equipment and medium

    CN113395594A