Method, apparatus, electronic device and medium for video classification
By extracting the multimodal features of video data and performing fusion processing, the problem of low video classification accuracy in the prior art is solved, and more efficient video classification is achieved.
Patent Information
- Application Number
- CN202111556380.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-17
AI Technical Summary
In the prior art, video classification has low accuracy and usually requires manual classification, which is not efficient.
By obtaining the video data to be classified, input it to the audio and video learning network and the text learning network, image features, audio features and text features are extracted, and these features are input to the fusion learning network to generate fusion feature vectors, and finally classified through the Softmax classifier.
It improves the accuracy of video classification, reduces dependence on manual classification, and improves classification efficiency.
Smart Images

Figure CN114037946B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to data processing technologies, and in particular, to a method, apparatus, electronic device, and medium for video classification. Background Art
[0002] With the rapid development of mobile Internet technologies, the continuous improvement of network transmission speeds, and the continuous progress of compression technologies, various multimedia information has emerged continuously. A large amount of video data has been generated and used in digital libraries, distance education, video on demand, digital video broadcasting, interactive television, etc.
[0003] On this basis, video classification has also become an important research topic in the field of multimedia analysis. Video classification is the basis for many video applications, and it provides convenience for the management of the increasing video data. Technologies such as content-based video retrieval, video summary, video indexing, and tagging are all driving the development of video classification technologies.
[0004] However, in the related technologies, the accuracy of the method for automatically classifying videos by a computer is relatively low, and generally, videos can only be classified manually. This also results in low classification efficiency. Summary of the Invention
[0005] Embodiments of the present application provide a method, apparatus, electronic device, and medium for video classification. It is used to solve the problem in the related technologies that videos cannot be accurately classified.
[0006] Among them, according to one aspect of the embodiments of the present application, a method for video classification is provided, including:
[0007] Obtain video data to be classified;
[0008] Input the video data to be classified into an audio-visual learning network to obtain image features and audio features corresponding to the video to be classified; and input the video data to be classified into a text learning network to obtain text features corresponding to the video to be classified;
[0009] Input the image features, the audio features, and the text features into a fusion learning network to obtain a fusion feature vector;
[0010] Input the fusion feature vector into a Softmax classifier, and use the classification result output by the classifier as the classification result of the video to be classified.
[0011] Optionally, in another embodiment based on the above method of the present application, the step of inputting the image features, the audio features, and the text features into a fusion learning network to obtain a fusion feature vector includes:
[0012] Perform vector transformation on the image features, the audio features, and the text features respectively to obtain an image feature vector, an audio feature vector, and a text feature vector;
[0013] Perform vector addition on the image feature vector, the audio feature vector, and the text feature vector to obtain a first fusion feature vector; and perform product normalization on the image feature vector, the audio feature vector, and the text feature vector to obtain a second fusion feature vector
[0014] Based on the first fusion feature vector and the second fusion feature vector, obtain the fusion feature vector.
[0015] Optionally, in another embodiment based on the above method of the present application, the obtaining the fusion feature vector based on the first fusion feature vector and the second fusion feature vector includes:
[0016] Generate a plurality of weight coefficient vectors, and perform normalization processing on the plurality of weight coefficient vectors to obtain fusion weight coefficients;
[0017] Use the fusion weight coefficients to perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain the fusion feature vector.
[0018] Optionally, in another embodiment based on the above method of the present application, the using the fusion weight coefficients to perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain the fusion feature vector includes:
[0019] Use the fusion weight coefficients multiple times to perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain a plurality of preliminary fusion feature vectors;
[0020] Use a loss function to adaptively update the fusion weight coefficients to obtain updated weight coefficients;
[0021] Use the updated weight coefficients to perform weighted summation on the plurality of preliminary fusion feature vectors to obtain the fusion feature vector.
[0022] Optionally, in another embodiment based on the above method of the present application, the inputting the video data to be classified into a text learning network to obtain the text features corresponding to the video to be classified includes:
[0023] Perform speech recognition on the video data to be classified to obtain text to be processed;
[0024] Use a preset conversion rule to convert the letter fields and emoji fields included in the text to be processed into text fields;
[0025] Convert the text to be processed containing the text field into a one-hot vector;
[0026] Input the one-hot vector into the text learning network for deep semantic feature extraction to obtain the text feature.
[0027] Optionally, in another embodiment based on the above method of the present application, the inputting the video data to be classified into the audio-visual learning network to obtain the image feature corresponding to the video data to be classified includes:
[0028] Divide the video data to be classified into multiple sub-video data according to a preset interval duration;
[0029] Extract a key frame image from each sub-video data to obtain a set of multiple key frame images;
[0030] Arrange the multiple key frame images in the set of key frame images in chronological order to obtain the image data to be input;
[0031] Input the image data to be input into the audio-visual learning network to obtain the image feature corresponding to the video data to be classified.
[0032] Optionally, in another embodiment based on the above method of the present application, the inputting the video data to be classified into the audio-visual learning network to obtain the audio feature corresponding to the video data to be classified includes:
[0033] Extract the audio data to be processed included in the video data to be classified;
[0034] Input the audio data to be processed into the audio-visual learning network to obtain the audio feature corresponding to the video data to be classified.
[0035] Wherein, according to another aspect of the embodiments of the present application, a video classification device is provided, which is characterized by including:
[0036] An acquisition module, configured to acquire video data to be classified;
[0037] A generation module, configured to input the video data to be classified into the audio-visual learning network to obtain the image feature and audio feature corresponding to the video data to be classified; and input the video data to be classified into the text learning network to obtain the text feature corresponding to the video data to be classified;
[0038] A fusion module, configured to input the image feature, the audio feature, and the text feature into the fusion learning network to obtain a fusion feature vector;
[0039] A classification module, configured to input the fused feature vector into a Softmax classifier, and use the classification result output by the classifier as the classification result of the video to be classified.
[0040] According to another aspect of the embodiments of the present application, an electronic device is provided, including:
[0041] A memory for storing executable instructions; and
[0042] A display for operating with the memory to execute the executable instructions to complete the operations of the video classification method described above in any one of the preceding items.
[0043] According to still another aspect of the embodiments of the present application, a computer-readable storage medium is provided for storing computer-readable instructions, and when the instructions are executed, the operations of the video classification method described above in any one of the preceding items are performed.
[0044] In the present application, video data to be classified can be obtained; the video data to be classified is input into an audio-visual learning network to obtain image features and audio features corresponding to the video to be classified; and the video data to be classified is input into a text learning network to obtain text features corresponding to the video to be classified; the image features, audio features, and text features are input into a fusion learning network to obtain a fused feature vector; the fused feature vector is input into a Softmax classifier, and the classification result output by the classifier is used as the classification result of the video to be classified. By applying the technical solution of the present application, after obtaining the video to be classified, the image features, audio features, and text features of the video data can be obtained by using a preset learning network model, and after fusing the three features, the classification result of the video to be classified can be determined according to the fused features. Thus, the drawback of inaccurate classification of video data in the related art is avoided.
[0045] The technical solution of the present application will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The drawings forming a part of the specification depict embodiments of the present application and, together with the description, are used to explain the principles of the present application.
[0047] Referring to the drawings, the present application can be more clearly understood according to the following detailed description, where:
[0048] Figure 1 is a schematic diagram of a video classification method proposed by the present application;
[0049] Figure 2 is an overall network architecture diagram of a video classification proposed by the present application;
[0050] Figure 3Schematic structural diagram of an electronic device for video classification proposed in this application;
[0051] Figure 4 Schematic structural diagram of an electronic device for video classification proposed in this application. Detailed implementation manners
[0052] Various exemplary embodiments of this application will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and values set forth in these embodiments do not limit the scope of this application.
[0053] Meanwhile, it should be understood that, for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships.
[0054] The following description of at least one exemplary embodiment is merely illustrative in nature and is not intended as any limitation to this application and its application or use.
[0055] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.
[0056] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0057] In addition, the technical solutions between various embodiments of this application can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.
[0058] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of this application are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will change accordingly.
[0059] The following is combined with Figure 1 - Figure 2 to describe the method for video classification according to the exemplary embodiments of this application. It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of this application, and the embodiments of this application are not limited in this regard. On the contrary, the embodiments of this application can be applied to any applicable scenario.
[0060] This application also proposes a method, device, electronic device, and medium for video classification.
[0061] Figure 1 Schematically shown is a flowchart of a method for video classification according to an embodiment of the present application. As Figure 1 shown, the method includes:
[0062] S101, obtaining video data to be classified.
[0063] S102, inputting the video data to be classified into an audio-visual learning network to obtain image features and audio features corresponding to the video to be classified; and inputting the video data to be classified into a text learning network to obtain text features corresponding to the video to be classified.
[0064] S103, inputting the image features, audio features, and text features into a fusion learning network to obtain a fusion feature vector.
[0065] S104, inputting the fusion feature vector into a Softmax classifier, and taking the classification result output by the classifier as the classification result of the video to be classified.
[0066] With the rapid development of mobile Internet technology, a large amount of video data has been generated. This includes long videos and short videos. Taking short videos as an example, nowadays, more and more mobile Internet users display their lives, express their views and emotions through short videos, which makes short videos have a strong social attribute and are extremely likely to trigger and spread public opinions.
[0067] In one approach. With the development of computer vision, especially the significant progress made by Convolutional Neural Network (CNN) in image classification tasks, related technologies have begun to attempt to use deep learning methods to solve video classification problems. However, existing methods are difficult to achieve the expected results in the classification practice of short videos. The main reasons are that existing classification methods are mostly trained using publicly available datasets based on general video data, while the data distributions in different fields vary greatly. Secondly, existing classification methods rely too much on image information, and their key features are easily masked by noise. In addition, short videos often have a large number of comments, and most existing classification models do not consider or cannot effectively utilize the semantic information contained in short video comments. In addition, in terms of short video datasets, the currently available publicly available datasets are basically for general fields, and it is difficult to find a suitable publicly available short video dataset for the university field.
[0068] To address the above problems, this application proposes a short video classification method based on a multi-network structure (Multimodal Micro-video Classification Based on Multi-network Structure, MMS). It can be understood that this application can extract multi-modal features from the multi-modal information carried by video data, such as feature extraction networks for image data, audio data, and text data, perform joint training on the multi-modal features, and finally fuse the three types of features and obtain the video classification result based on the fused features.
[0069] Furthermore, the overall structure of a video classification method proposed in this application is as Figure 2 shown. To extract the image features, audio features, and text features of the video to be classified. The overall process of this application includes an audio-visual learning network, a text learning network, and a fusion learning network. Among them, the audio-visual learning network can include two sub-networks, namely a visual network and an audio network. It can also be an audio-visual learning network that processes audio-visual data as a whole. This application does not make a limitation on this.
[0070] In one way, the audio-visual network or the visual network is responsible for capturing the behavioral information presented in the video, extracting appropriate visual behavioral features through learning a large number of behaviors, which can be used for subsequent multi-modal feature fusion or single-modal data processing.
[0071] In addition, the audio learning network in this application uses convolution to extract the waveform features of different segments of the audio, and sends them as tokens into the Transformer encoder to obtain audio feature vectors. The text network has the ability to extract the deep semantic features of text information. In the embodiments of this application, the video to be classified can be input into different sub-networks for corresponding feature extraction. Specifically, it includes inputting the video data to be classified into the audio-visual learning network to obtain the image features and audio features corresponding to the video to be classified; and, after inputting the video data to be classified into the text learning network to obtain the text features corresponding to the video to be classified, the fusion learning network fuses the three features, inputs the fused feature vector into the Softmax classifier, and uses the classification result output by the classifier as the classification result of the video to be classified.
[0072] In this application, video data to be classified can be obtained; the video data to be classified is input into an audio-visual learning network to obtain image features, audio features, and text features corresponding to the video to be classified; and the image features, audio features, and text features corresponding to the video to be classified are input into a fusion learning network to obtain a fusion feature vector; the fusion feature vector is input into a Softmax classifier, and the classification result output by the classifier is used as the classification result of the video to be classified. By applying the technical solution of this application, after obtaining the video to be classified, the image features, audio features, and text features of the video data can be obtained by using a preset learning network model, and after fusing these three features, the classification result of the video to be classified is determined according to the fused features. Thus, the drawback of inaccurate classification of video data in the related art is avoided. The navigation route is selected as the planned route with the lowest passing cost. Furthermore, the purpose of selecting the navigation route with the lowest passing cost for the user can be achieved from the standards of the passing cost and passing efficiency of the vehicle.
[0073] Optionally, in another embodiment based on the above method of this application, inputting the image features, the audio features, and the text features into a fusion learning network to obtain a fusion feature vector includes:
[0074] Vector conversions are respectively performed on the image features, the audio features, and the text features to obtain an image feature vector, an audio feature vector, and a text feature vector;
[0075] The image feature vector, the audio feature vector, and the text feature vector are added vectorially to obtain a first fusion feature vector; and the image feature vector, the audio feature vector, and the text feature vector are multiplied and normalized to obtain a second fusion feature vector
[0076] Based on the first fusion feature vector and the second fusion feature vector, the fusion feature vector is obtained.
[0077] Optionally, in another embodiment based on the above method of this application, the obtaining the fusion feature vector based on the first fusion feature vector and the second fusion feature vector includes:
[0078] Generate a plurality of weight coefficient vectors, and perform normalization processing on the plurality of weight coefficient vectors to obtain fusion weight coefficients;
[0079] Using the fusion weight coefficients, perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain the fusion feature vector.
[0080] Optionally, in another embodiment of the method based on the present application, the step of using the fusion weight coefficient to perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain the fusion feature vector includes:
[0081] Using the fusion weight coefficient multiple times to perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain multiple preliminary fusion feature vectors;
[0082] Using a loss function to adaptively update the fusion weight coefficient to obtain an updated weight coefficient;
[0083] Using the updated weight coefficient to perform weighted summation on the multiple preliminary fusion feature vectors to obtain the fusion feature vector.
[0084] Optionally, in another embodiment of the method based on the present application, the step of inputting the video data to be classified into a text learning network to obtain the text feature corresponding to the video data to be classified includes:
[0085] Performing speech recognition on the video data to be classified to obtain a text to be processed;
[0086] Using a preset conversion rule to convert the letter fields and emoji fields included in the text to be processed into text fields;
[0087] Converting the text to be processed including the text fields into a one-hot vector;
[0088] Inputting the one-hot vector into the text learning network for deep semantic feature extraction to obtain the text feature.
[0089] Optionally, in another embodiment of the method based on the present application, the step of inputting the video data to be classified into an audio-visual learning network to obtain the image feature corresponding to the video data to be classified includes:
[0090] Dividing the video data to be classified into multiple sub-video data at a preset interval duration;
[0091] Extracting a key frame image from each sub-video data respectively to obtain a plurality of key frame image sets;
[0092] Arranging the multiple key frame images in the key frame image sets in chronological order to obtain input image data;
[0093] Inputting the input image data into the audio-visual learning network to obtain the image feature corresponding to the video data to be classified.
[0094] Optionally, in another embodiment based on the above method of the present application, the step of inputting the video data to be classified into the audio-visual learning network to obtain the audio features corresponding to the video to be classified includes:
[0095] Extracting the audio data to be processed included in the video data to be classified;
[0096] Inputting the audio data to be processed into the audio-visual learning network to obtain the audio features corresponding to the video to be classified.
[0097] In one approach, the audio-visual learning network or the visual learning network is responsible for capturing the behavior information presented in the video, and extracting appropriate visual behavior features through learning a large number of behaviors, which can be used for subsequent multi-modal feature fusion or single-modal data processing. In addition, the audio learning network in the present application uses convolution to extract the waveform features of different segments of the audio, and uses them as tokens to be fed into the Transformer encoder to obtain the audio feature vectors. The text network has the ability to extract deep semantic features of text information.
[0098] Specifically, in the embodiments of the present application, the video to be classified can be input into different sub-networks for corresponding feature extraction. Specifically, it includes inputting the video data to be classified into the audio-visual learning network to obtain the image features and audio features corresponding to the video to be classified; and, after inputting the video data to be classified into the text learning network to obtain the text features corresponding to the video to be classified, the fusion learning network fuses the three features, and then inputs the fused feature vector into the Softmax classifier, and uses the classification result output by the classifier as the classification result of the video to be classified.
[0099] Specifically, for obtaining the image features corresponding to the video to be classified, in the embodiments of the present application, the video data to be classified can be first stored in the form of RGB frames. In order to reduce the amount of calculation, certain preprocessing needs to be performed on the video data to be classified. This may include evenly dividing the classified video data into a fixed number of sub-video data according to the duration, and then randomly selecting one frame from each sub-video data as the key frame of this sub-video data.
[0100] In addition, the key frames of the extracted sub-video data are arranged in order of the time dimension to form an RGB input sequence of size T×H×W to obtain the image data to be input. Then, the image data to be input is input into the audio-visual learning network to obtain the image features corresponding to the video to be classified.
[0101] In one approach, the present application can first extract an audio waveform file from a video source file, use a deep convolutional network on the waveform file to obtain a potential audio representation, and regard the audio representation as a special form of text and input it into a Transformer encoder to obtain a final feature vector. In one approach, the present application can use a standard wav2vec learning model to extract audio features and obtain the audio features corresponding to the video to be classified.
[0102] Furthermore, for obtaining the text features corresponding to the video to be classified, the present application needs to perform speech recognition on the video data to be classified to obtain the text to be processed, where the text to be processed can be converted from the audio data appearing in the video to be classified, or can be comment information, title information, tag information, etc. for the video to be classified.
[0103] Further, due to social short texts such as video titles and comments, there are a large number of expressions such as abbreviations, homophonic words, and emoticons that are completely different from traditional written languages. In this case, appropriate preprocessing operations are very crucial. In one approach, the present application can first establish an emoji-meaning mapping table to replace the special emojis contained in the text with standard text, then replace the letters and abbreviations with the first candidate in the corresponding results of Baidu Input Method, and finally convert the obtained standard text information into a one-hot vector and input it into a text learning network for deep semantic feature extraction to obtain a feature vector.
[0104] Specifically, for the construction of the text network, through in-depth research on current natural language processing models, it is found that the existing BERT model can meet the text processing requirements. Therefore, the text network draws on the improved BERT model in terms of structure. BERT uses multiple layers of Transformer bidirectional encoders to obtain the deep semantics of the text by coordinating the context representations of all layers. First, the input text is tokenized into token vectors. In classification tasks, special tokens are used to identify the first token of the sentence, and special tokens are used to separate non-contiguous token sequences, and the processed tokens are fed into the BERT model to obtain the vector representation of the sentence. Since it is difficult to train a bidirectional language model in the traditional left-to-right or right-to-left manner, a random masking method is used in the BERT model to mask some tokens in the input sequence and predict the masked parts.
[0105] Furthermore, since the three feature dimensions extracted from the video data are different, the semantic information contained in the three features is not completely consistent, which may also lead to contradictory results when directly classifying using the three feature vectors. To alleviate the problem of different modality semantic conflicts, the present application may select to input the image features, audio features, and text features into a fusion learning network to obtain a fusion feature vector, so that subsequently, the fusion feature vector is input into a Softmax classifier, and the classification result output by the classifier is used as the classification result of the video to be classified.
[0106] Specifically, for fusing the image features, audio features, and text features, the following steps may be included:
[0107] Step 1: Process the three feature vectors (image features, audio features, and text features) using a multi-layer perceptron to obtain three vectors with the same dimension.
[0108] Step 2: Normalize the three vectors;
[0109] Step 3: Add the three vectors for preliminary fusion to obtain a first fusion feature vector;
[0110] Step 4: Calculate the Hadamard product of the three vectors to reduce errors and obtain a second fusion feature vector;
[0111] Step 5: Randomly generate multiple weight coefficient vectors (for example, 4);
[0112] Step 6: Normalize the multiple weight coefficient vectors to ensure that the sum of the elements in the corresponding positions of the multiple vectors is 1 to obtain fusion weight coefficients;
[0113] Step 7: Use the fusion weight coefficients to perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain a fusion feature vector;
[0114] Step 8: Use a loss function to adaptively update the fusion weight coefficients to obtain updated weight coefficients;
[0115] Step 9: Use the updated weight coefficients to perform weighted summation on multiple preliminary fusion feature vectors to obtain a fusion feature vector
[0116] Step 10: Input the fusion feature vector into a Softmax classifier, and use the classification result output by the classifier as the classification result of the video to be classified.
[0117] By applying the technical solution of the present application, after obtaining the video to be classified, the image features, audio features, and text features of the video data can be obtained by using a preset learning network model, and after fusing these three features, the classification result of the video to be classified can be determined according to the fused features. Thereby avoiding the disadvantages of inaccurate classification of video data in the related art.
[0118] Optionally, in another implementation manner of the present application, as Figure 3 shown, the present application further provides a video classification device. It includes:
[0119] An acquisition module 201, configured to acquire video data to be classified;
[0120] A generation module 202, configured to input the video data to be classified into an audio-visual learning network to obtain the image features and audio features corresponding to the video to be classified; and input the video data to be classified into a text learning network to obtain the text features corresponding to the video to be classified;
[0121] A fusion module 203, configured to input the image features, the audio features, and the text features into a fusion learning network to obtain a fused feature vector;
[0122] A classification module 204, configured to input the fused feature vector into a Softmax classifier, and use the classification result output by the classifier as the classification result of the video to be classified.
[0123] In the present application, video data to be classified can be acquired; the video data to be classified is input into an audio-visual learning network to obtain the image features and audio features corresponding to the video to be classified; and the text features corresponding to the video to be classified; the image features, audio features, and text features are input into a fusion learning network to obtain a fused feature vector; the fused feature vector is input into a Softmax classifier, and the classification result output by the classifier is used as the classification result of the video to be classified. By applying the technical solution of the present application, after obtaining the video to be classified, the image features, audio features, and text features of the video data can be obtained by using a preset learning network model, and after fusing these three features, the classification result of the video to be classified can be determined according to the fused features. Thereby avoiding the disadvantages of inaccurate classification of video data in the related art. Select the planned route with the lowest passing cost as the navigation route. Furthermore, the purpose of selecting the navigation route with the lowest passing cost for the user can be achieved from the standards of the passing cost and passing efficiency of the vehicle.
[0124] In another implementation manner of the present application, the steps that the acquisition module 201 is configured to execute include:
[0125] Perform vector transformation on the image feature, the audio feature, and the text feature respectively to obtain an image feature vector, an audio feature vector, and a text feature vector;
[0126] Perform vector addition on the image feature vector, the audio feature vector, and the text feature vector to obtain a first fusion feature vector; and, perform product normalization on the image feature vector, the audio feature vector, and the text feature vector to obtain a second fusion feature vector
[0127] Based on the first fusion feature vector and the second fusion feature vector, obtain the fusion feature vector.
[0128] In another embodiment of the present application, the steps that the acquisition module 201 is configured to execute include:
[0129] Generate a plurality of weight coefficient vectors, and perform normalization processing on the plurality of weight coefficient vectors to obtain fusion weight coefficients;
[0130] Use the fusion weight coefficients to perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain the fusion feature vector.
[0131] In another embodiment of the present application, the steps that the acquisition module 201 is configured to execute include:
[0132] Use the fusion weight coefficients multiple times to perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain a plurality of preliminary fusion feature vectors;
[0133] Use a loss function to adaptively update the fusion weight coefficients to obtain updated weight coefficients;
[0134] Use the updated weight coefficients to perform weighted summation on the plurality of preliminary fusion feature vectors to obtain the fusion feature vector.
[0135] In another embodiment of the present application, the steps that the acquisition module 201 is configured to execute include:
[0136] Perform speech recognition on the video data to be classified to obtain a text to be processed;
[0137] Use a preset conversion rule to convert the letter fields and emoji fields included in the text to be processed into text fields;
[0138] Convert the text to be processed including the text fields into a one-hot vector;
[0139] Input the one-hot vector into the text learning network for deep semantic feature extraction to obtain the text feature.
[0140] In another embodiment of the present application, the steps that the acquisition module 201 is configured to execute include:
[0141] Dividing the video data to be classified into multiple sub-video data according to a preset interval duration;
[0142] Extracting a key frame image from each sub-video data respectively to obtain a plurality of key frame image sets;
[0143] Arranging the multiple key frame images in the key frame image set in chronological order to obtain the image data to be input;
[0144] Inputting the image data to be input into the audio-visual learning network to obtain the image features corresponding to the video data to be classified.
[0145] In another embodiment of the present application, the steps that the acquisition module 201 is configured to execute include:
[0146] Extracting the audio data to be processed included in the video data to be classified;
[0147] Inputting the audio data to be processed into the audio-visual learning network to obtain the audio features corresponding to the video data to be classified.
[0148] Figure 4 It is a logical structure block diagram of an electronic device shown according to an exemplary embodiment. For example, the electronic device 300 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0149] In an exemplary embodiment, there is also provided a non-transitory computer-readable storage medium including instructions, such as a memory including instructions. The above instructions can be executed by a processor of an electronic device to complete the above video classification method. The method includes: acquiring video data to be classified; inputting the video data to be classified into an audio-visual learning network to obtain image features and audio features corresponding to the video data to be classified; and inputting the video data to be classified into a text learning network to obtain text features corresponding to the video data to be classified; inputting the image features, the audio features, and the text features into a fusion learning network to obtain a fusion feature vector; inputting the fusion feature vector into a Softmax classifier, and using the classification result output by the classifier as the classification result of the video data to be classified. Optionally, the above instructions can also be executed by a processor of an electronic device to complete other steps involved in the above exemplary embodiment. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0150] In an exemplary embodiment, an application program / computer program product is also provided, including one or more instructions that can be executed by a processor of an electronic device to complete the above method for video classification. The method includes: obtaining video data to be classified; inputting the video data to be classified into an audio-visual learning network to obtain image features and audio features corresponding to the video data to be classified; and inputting the video data to be classified into a text learning network to obtain text features corresponding to the video data to be classified; inputting the image features, the audio features, and the text features into a fusion learning network to obtain a fusion feature vector; inputting the fusion feature vector into a Softmax classifier, and using the classification result output by the classifier as the classification result of the video data to be classified. Optionally, the above instructions can also be executed by a processor of an electronic device to complete other steps involved in the above exemplary embodiment.
[0151] Figure 4 FIG. is an example diagram of an electronic device 300. Those skilled in the art can understand that the schematic Figure 4 is merely an example of the electronic device 300, and does not constitute a limitation on the electronic device 300. It may include more or fewer components than shown, or combine certain components, or different components. For example, the electronic device 300 may further include input / output devices, network access devices, buses, etc.
[0152] The so-called processor 302 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor 302 may also be any conventional processor, etc. The processor 302 is the control center of the electronic device 300, and connects various parts of the entire electronic device 300 through various interfaces and lines.
[0153] The memory 301 can be used to store computer-readable instructions 303. By running or executing the computer-readable instructions or modules stored in the memory 301 and invoking the data stored in the memory 301, the processor 302 realizes various functions of the electronic device 300. The memory 301 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 300. In addition, the memory 301 can include a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash device, Read-Only Memory (ROM), Random Access Memory (RAM), or other non-volatile / volatile storage devices.
[0154] If the modules integrated in the electronic device 300 are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by computer-readable instructions to instruct relevant hardware. The computer-readable instructions can be stored in a computer-readable storage medium. When the computer-readable instructions are executed by the processor, the steps of the above-mentioned various method embodiments can be realized.
[0155] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application. These variations, uses, or adaptations follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0156] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for video classification, characterized in that, it includes: Obtain video data to be classified; Input the video data to be classified into an audio-visual learning network to obtain image features and audio features corresponding to the video to be classified; And input the video data to be classified into a text learning network to obtain text features corresponding to the video to be classified; Input the image features, the audio features, and the text features into a fusion learning network to obtain a fusion feature vector; Input the fusion feature vector into a Softmax classifier, and use the classification result output by the classifier as the classification result of the video to be classified; Among them, the step of inputting the video data to be classified into a text learning network to obtain text features corresponding to the video to be classified includes: performing speech recognition on the video data to be classified to obtain text to be processed; using a preset conversion rule to convert the letter fields and emoji fields included in the text to be processed into text fields; converting the text to be processed including the text fields into a one-hot vector; inputting the one-hot vector into the text learning network for deep semantic feature extraction to obtain the text features; among them, establish an emoji and meaning mapping table, and replace the emoji fields included in the text to be processed with standard text; replace letters and abbreviations with the first candidate word in the corresponding result of the input method. The step of inputting the image features, the audio features, and the text features into a fusion learning network to obtain a fusion feature vector includes: respectively performing vector conversion on the image features, the audio features, and the text features to obtain an image feature vector, an audio feature vector, and a text feature vector; performing vector addition on the image feature vector, the audio feature vector, and the text feature vector to obtain a first fusion feature vector; and performing product normalization on the image feature vector, the audio feature vector, and the text feature vector to obtain a second fusion feature vector; based on the first fusion feature vector and the second fusion feature vector, obtain the fusion feature vector; among them, the way to obtain the second fusion feature vector is to perform Hadamard product on the image feature vector, the audio feature vector, and the text feature vector.
2. The method according to claim 1, characterized in that, the step of obtaining the fusion feature vector based on the first fusion feature vector and the second fusion feature vector includes: Generate a plurality of weight coefficient vectors, and perform normalization processing on the plurality of weight coefficient vectors to obtain fusion weight coefficients; Use the fusion weight coefficients to perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain the fusion feature vector.
3. The method according to claim 2, characterized in that, the step of using the fusion weight coefficients to perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain the fusion feature vector includes: Utilize the fusion weight coefficients multiple times to perform weighted summation on the first fusion feature vector and the second fusion feature vector to obtain multiple preliminary fusion feature vectors; Utilize a loss function to adaptively update the fusion weight coefficients to obtain updated weight coefficients; Utilize the updated weight coefficients to perform weighted summation on the multiple preliminary fusion feature vectors to obtain the fusion feature vector.
4. The method according to claim 1, wherein, the inputting the video data to be classified into an audio-visual learning network to obtain the image features corresponding to the video data to be classified includes: dividing the video data to be classified into multiple sub-video data according to a preset interval duration; extracting a key frame image from each sub-video data respectively to obtain a plurality of key frame image sets; arranging the multiple key frame images in the key frame image sets in chronological order to obtain the image data to be input; inputting the image data to be input into the audio-visual learning network to obtain the image features corresponding to the video data to be classified.
5. The method according to claim 1 or 4, wherein, the inputting the video data to be classified into an audio-visual learning network to obtain the audio features corresponding to the video data to be classified includes: extracting the audio data to be processed included in the video data to be classified; inputting the audio data to be processed into the audio-visual learning network to obtain the audio features corresponding to the video data to be classified.
6. An apparatus for video classification, wherein, it includes: an acquisition module configured to acquire video data to be classified; a generation module configured to input the video data to be classified into an audio-visual learning network to obtain the image features and audio features corresponding to the video data to be classified; and input the video data to be classified into a text learning network to obtain the text features corresponding to the video data to be classified; a fusion module configured to input the image features, the audio features, and the text features into a fusion learning network to obtain a fusion feature vector; a classification module configured to input the fusion feature vector into a Softmax classifier and use the classification result output by the classifier as the classification result of the video data to be classified; the generation module is further configured to perform speech recognition on the video data to be classified to obtain the text to be processed; utilize a preset conversion rule to convert the letter fields and expression fields included in the text to be processed into text fields; convert the text to be processed including the text fields into a one-hot vector; input the one-hot vector into the text learning network to perform deep semantic feature extraction to obtain the text features; wherein, establish an expression and meaning mapping table, replace the expression fields included in the text to be processed with standard text; replace the letters and abbreviations with the first candidate word in the corresponding result of the input method; The fusion module is further configured to perform vector transformation on the image feature, the audio feature, and the text feature respectively to obtain an image feature vector, an audio feature vector, and a text feature vector; perform vector addition on the image feature vector, the audio feature vector, and the text feature vector to obtain a first fusion feature vector; and perform product normalization on the image feature vector, the audio feature vector, and the text feature vector to obtain a second fusion feature vector; obtain the fusion feature vector based on the first fusion feature vector and the second fusion feature vector; wherein the second fusion feature vector is obtained by taking the Hadamard product of the image feature vector, the audio feature vector, and the text feature vector.
7. An electronic device characterized in that it includes: a memory for storing executable instructions; and a processor for cooperating with the memory to execute the executable instructions so as to complete the operations of the video classification method according to any one of claims 1-5.
8. A computer-readable storage medium for storing computer-readable instructions characterized in that when the instructions are executed, the operations of the video classification method according to any one of claims 1-5 are executed.
Citation Information
Patent Citations
Video classification method and device, storage medium and server
CN111209970A