Short video classification method, system and device and storage medium
By decomposing short videos into video streams and audio streams, extracting the features of time synchronization and using pre-trained classification models for identification, the problem of low classification efficiency in the existing technology is solved, and fast and efficient video classification is achieved.
Patent Information
- Application Number
- CN202510071276.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-17
AI Technical Summary
Existing video classification methods based on deep learning face huge data processing volume when processing short videos, resulting in increased computational complexity and reduced recognition efficiency.
The short video is decomposed into video streams and audio streams, the local texture features of keyframes are extracted from the video stream, and the audio features with time synchronization are extracted based on these features, and the merged features are identified through pre-trained classification models.
It realizes rapid classification of short videos, reduces computational complexity, and improves recognition efficiency.
Smart Images

Figure CN120164138A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a short video classification method, system, device and storage medium. Background Art
[0002] As a new way of Internet content dissemination, short videos have quickly attracted the attention of the majority of Internet users with their unique charm. The duration of this video form is usually controlled within 5 minutes, which not only meets the fragmented entertainment needs in the fast-paced life of modern people, but also facilitates rapid dissemination on platforms such as social media. With the increasing popularity of mobile terminal devices and the rapid development of Internet technology, especially the significant improvement in network speed, short videos have gradually become the focus pursued by major platforms, and at the same time have won the love of many fans and the favor of capital.
[0003] Today, short videos have become an important carrier and dissemination method for Internet users to obtain information, enjoy services, interact and engage in cultural entertainment. It not only enriches people's spare time life, but also promotes the diversified development of Internet culture.
[0004] However, the existing video classification methods based on deep learning still face some challenges when dealing with short videos. Since in-depth analysis of video data is required, this often means a huge amount of data processing. This not only increases the computational complexity, but may also lead to a decrease in recognition efficiency. These problems are particularly prominent when dealing with large-scale short video data. Summary of the Invention
[0005] In view of the above deficiencies of the prior art, the present invention provides a short video classification method, system, device and storage medium to solve the above technical problems.
[0006] In a first aspect, the present invention provides a short video classification method, including: Decompose the short video into a video stream and an audio stream; Extract key frames from the video stream and extract local texture features from the key frames; Extract audio features from the audio stream; Align the local texture features with the audio features based on the synchronization of the timestamps of the local texture features and the audio features; Use a pre-trained classification model to identify the category of the short video based on the aligned local texture features and the audio features.
[0007] In an optional embodiment, extracting key frames from the video stream and extracting local texture features from the key frames includes: Decode the video stream, identify and extract I frames in the video stream, and the I frame is an independent frame in video coding; Select the image frame with the earliest timestamp and the image frame with the latest timestamp from all I-frames, and randomly select three image frames from the remaining I-frames; Save all the selected image frames as images in a specified format; Extract texture features from the images using a local binary pattern processing model.
[0008] In an alternative embodiment, saving all the selected image frames as images in a specified format includes: Perform a forward discrete cosine transform on the image to obtain the discrete cosine transform coefficients of the image; Quantize the discrete cosine transform coefficients using a weighting function, truncate the discrete cosine transform coefficients, and only retain the coefficient values within a certain range to reduce the data volume; Convert the quantized DCT coefficients into binary form, divide the binary data into codeword blocks of size 7, and perform BCH coding processing on each codeword block; Remove the parity bits generated during the BCH coding process; Recombine all the data blocks after BCH coding and removing the parity bits into complete image data.
[0009] In an alternative embodiment, extracting texture features from the images using a local binary pattern processing model includes: Divide the image into 3x3 independent blocks, count the histogram data of each small block, and arrange it into a vector, which is the texture feature.
[0010] In an alternative embodiment, extracting audio features from the audio stream includes: Extract the first audio feature from the audio stream using mel-frequency cepstral coefficients; Extract the second audio feature from the audio stream using constant Q transform technology.
[0011] In an alternative embodiment, aligning the local texture features with the audio features based on the synchronization of the timestamps of the local texture features and the audio features includes: Sort the local texture features in the order of their corresponding timestamps to obtain a video feature sequence; Use the timestamps of the local texture features as the standard timestamps, calculate the time difference between the timestamps of the audio features and the standard timestamps, and select the minimum time difference; If the minimum time difference does not exceed a preset time threshold, determine that the corresponding audio feature is valid; if the minimum time difference exceeds the time threshold, determine that the corresponding audio feature is invalid; Filter out the effective audio features and arrange the effective audio features in chronological order to form an audio feature sequence.
[0012] In an optional embodiment, using a pre-trained classification model to identify the category of a short video based on the aligned local texture features and the audio features, includes: Calculate the weighted sum of the video feature sequence and the audio feature sequence to obtain a feature sequence; Input the feature sequence into the pre-trained classification model to obtain the category of the short video.
[0013] In a second aspect, the present invention provides a short video classification system, including: A data decomposition module, configured to decompose the short video into a video stream and an audio stream; An image processing module, configured to extract key frames from the video stream and extract local texture features from the key frames; An audio processing module, configured to extract audio features from the audio stream; A synchronization alignment module, configured to align the local texture features with the audio features based on the synchrony of the timestamps of the local texture features and the audio features; A classification processing module, configured to use a pre-trained classification model to identify the category of a short video based on the aligned local texture features and the audio features.
[0014] In a third aspect, a device is provided, including: A memory, configured to store a short video classification program; A processor, configured to implement the steps of the short video classification method provided in the first aspect when executing the short video classification program.
[0015] In a fourth aspect, a computer-readable storage medium is provided, on which a short video classification program is stored. When the short video classification program is executed by a processor, the steps of the short video classification method provided in the first aspect are implemented.
[0016] The beneficial effects of the present invention are that the short video classification method, system, device, and storage medium provided by the present invention decompose a short video into a video stream and an audio stream, extract local texture features of key frames from the video stream, and extract audio features with time synchrony based on the local texture features of the key frames, and perform fusion classification on the two features, so as to realize the rapid classification of short videos.
[0017] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention.
[0020] Figure 2 is a schematic block diagram of the system according to an embodiment of the present invention.
[0021] Figure 3 is a schematic structural diagram of a device provided by an embodiment of the present invention. Detailed implementation manners
[0022] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0024] The short video classification method provided by the embodiment of the present invention is executed by a computer device. Correspondingly, the short video classification system runs in the computer device.
[0025] Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention. Among them, Figure 1 The execution subject can be a short video classification system. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.
[0026] Such as Figure 1 shown, the method includes: S1. Decompose the short video into a video stream and an audio stream.
[0027] Use professional video processing tools or libraries (such as FFmpeg, OpenCV, etc.) to parse short video files. Short video files usually contain a composite data stream of video and audio. Our task is to separate them into independent video streams and audio streams for subsequent individual processing and analysis.
[0028] S2. Extract key frames from the video stream and extract local texture features from the key frames.
[0029] In a video stream, key frames refer to those frames that can represent the main changes or important information in the video content. We can use video processing algorithms (such as inter-frame difference detection, motion estimation, etc.) to identify and extract these key frames. Next, we need to extract features from the extracted key frames. Here, we choose to extract local texture features because texture features are one of the important clues in video content analysis. Common local texture feature extraction methods include Local Binary Pattern (LBP), Gabor filters, Histogram of Oriented Gradients (HOG), etc.
[0030] S3. Extract audio features from the audio stream.
[0031] Audio feature extraction is an important step in audio processing. In this step, we need to use audio processing tools or libraries (such as librosa, PyDub, etc.) to parse the audio stream and extract parameters that can represent the characteristics of the audio content. Common audio features include Mel Frequency Cepstral Coefficients (MFCC), spectral centroid, pitch, rhythm, etc. These features can reflect key information such as the sound quality, pitch, and rhythm of the audio.
[0032] S4. Align the local texture features with the audio features based on the synchrony of the timestamps of the local texture features and the audio features.
[0033] Since the video and audio are recorded simultaneously, there is a temporal synchrony between them. We can utilize this characteristic to align the extracted local texture features with the audio features. The alignment method usually involves comparing and matching timestamps to ensure the temporal consistency of the video and audio features. This step is crucial for subsequent short video classification and recognition because it can ensure the effective fusion and utilization of video and audio information in the classification model.
[0034] S5. Use a pre-trained classification model to identify the category of the short video based on the aligned local texture features and the audio features.
[0035] Select a suitable classification model to classify short videos. The classification model can be based on traditional machine learning algorithms (such as support vector machines, decision trees, etc.) or based on deep learning (such as convolutional neural networks, recurrent neural networks, etc.). Before classification, we need to fuse the aligned local texture features and audio features to form a unified feature vector. Then, we input this feature vector into the classification model for training and testing. Finally, the classification model will output the category label of the short video to achieve the classification and recognition of the short video.
[0036] Based on the above embodiments, in order to further improve the processing speed of the short video classification provided by the above embodiments, as an implementable way, utilize the parallel computing power of GPGPU to accelerate the entire processing flow and reduce the data transmission and storage burden.
[0037] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.
[0038] S201. Decode the video stream, identify and extract the I-frames in the video stream. The I-frame is an independent frame in video coding.
[0039] Use video decoding technology to parse and process the video stream. The video stream is usually stored and transmitted in a compressed form. Therefore, we need to decode it to restore the original sequence of image frames. In video coding, I-frames (also known as key frames or independent frames) are those image frames that can be decoded independently without relying on other frames. They contain complete image information and are an important part of video content analysis.
[0040] In a specific example, the method for identifying independent frames is as follows: I. Video Coding Principle and Frame Types First, we need to understand the basic principle of video coding and the classification of frame types. Video coding standards (such as H.264, H.265, etc.) usually divide the video into a series of frames and encode them according to the frame types. Among them, I-frames (key frames or independent frames) are those image frames that can be decoded independently without relying on other frames and they contain complete image information. P-frames (predicted frames) rely on the previous I-frame or P-frame for decoding, while B-frames (bi-directionally predicted frames) rely on the previous and next I-frame or P-frame.
[0041] II. Video Decoder and I-frame Extraction To extract I-frames, we need to use a video decoder to decode the video stream. The video decoder can restore the compressed video stream to the original sequence of image frames according to the rules of the video coding standard. During the decoding process, the decoder will identify and extract all I-frames based on the frame type information.
[0042] III. Specific implementation steps Select a video decoding library: Select a video decoding library that supports the video coding standard you need, such as FFmpeg, libavcodec, etc. These libraries provide rich video processing functions, including decoding, encoding, transcoding, etc.
[0043] Initialize the decoder: According to the selected video decoding library, initialize the decoder and set relevant parameters. This includes specifying the video coding standard, allocating the memory space required for decoding, etc.
[0044] Read the video stream: Read the compressed video stream from the video file and pass it to the decoder for decoding. This usually involves file I / O operations and the parsing of the video stream format.
[0045] Decode and extract I-frames: During the decoding process, the decoder will classify the image frames according to the frame type information (such as I-frames, P-frames, B-frames, etc.). We can use the API or callback function provided by the decoder to extract the I-frames when they are decoded.
[0046] S202. Select the image frame with the earliest timestamp and the image frame with the latest timestamp from all I-frames, and randomly select three image frames from the remaining I-frames.
[0047] After extracting all the I-frames, we need to further screen out the image frames with representative or special significance. In this step, we adopt the following strategy: Select the image frame with the earliest timestamp: This image frame is usually the starting frame of the video, which represents the beginning of the video content. Through this image frame, we can initially understand the theme and background of the video.
[0048] Select the image frame with the latest timestamp: This image frame is the ending frame of the video, which represents the end of the video content. Through this image frame, we can generally understand the development and changes of the video.
[0049] Randomly select three image frames from the remaining I-frames: In addition to the starting frame and the ending frame, the video stream may also contain a large number of other I-frames. To select representative samples from these frames, we adopt the method of random selection. Randomly selecting three image frames can ensure that we can obtain diverse information from different parts of the video, so as to understand the content of the video more comprehensively.
[0050] S203. Save all the selected image frames as images in a specified format.
[0051] 1. Perform a forward discrete cosine transform (DCT) on the image First, divide the image into 8x8 pixel blocks (for the common JPEG compression standard). Then, use the two-dimensional DCT formula to perform a forward discrete cosine transform (DCT) on each 8x8 pixel block, converting the image from the spatial domain to the frequency domain. The purpose of this step is to concentrate the energy of the image on a few DCT coefficients, especially the low-frequency coefficients. The DCT coefficients are usually organized as an 8x8 matrix, where the coefficients in the upper left corner (low frequency) have larger values and the coefficients in the lower right corner (high frequency) have smaller values.
[0052] 2. Quantize the DCT coefficients using a weighting function Next, use a predefined quantization table (also known as a quantization matrix) to quantize the DCT coefficients. The quantization table is usually an 8x8 matrix, and the values in it determine the scaling degree of each DCT coefficient. The quantization process reduces the precision of the DCT coefficients, thereby reducing the data volume.
[0053] Divide the corresponding elements of the DCT coefficient matrix by the quantization table matrix, and then round or truncate. The values of the quantization table can be adjusted according to the image content and compression requirements.
[0054] 3. Truncate the quantized DCT coefficients To further reduce the data volume, the quantized DCT coefficients can be truncated, only retaining the coefficient values within a certain range (for example, setting a threshold, setting the coefficients with absolute values less than the threshold to zero, and only retaining the coefficients with absolute values greater than a certain threshold). This step is usually combined with the quantization process because quantization itself will cause some smaller coefficients to become zero.
[0055] This step can be regarded as a sparsification process for the quantized DCT coefficients.
[0056] 4. Convert the quantized DCT coefficients to binary form Convert the quantized and truncated DCT coefficients to binary form for subsequent encoding processing. This step usually involves encoding the sign and magnitude of each coefficient.
[0057] Encode the sign and magnitude of the coefficients using variable-length coding (VLC) or fixed-length coding (FLC).
[0058] VLC can assign codewords of different lengths according to the probability distribution of the coefficients, thereby achieving a higher compression ratio.
[0059] 5. Divide the binary data into codeword blocks of size 7 and perform BCH encoding on each codeword block Divide the binary data into codeword blocks (or data blocks) of size 7, and then perform BCH (Bose-Chaudhuri-Hocquenghem) encoding on each codeword block. BCH encoding is a cyclic code that can correct multiple errors. It improves the reliability of data by expanding the original data into a codeword containing parity bits.
[0060] Select appropriate BCH encoding parameters (such as the error correction ability t and the codeword length n) to ensure that the required number of errors can be corrected. Perform BCH encoding on each 7-bit codeword block to generate an extended codeword containing parity bits.
[0061] 6. Remove the parity bits generated during the BCH encoding process Remove the parity bits generated during the BCH encoding process to further reduce the data volume.
[0062] Remove the parity bits from the extended codeword and only retain the original data bits. Record or transmit information related to the removed parity bits (such as BCH encoding parameters and error correction ability) so that the parity bits can be reconstructed or error detection can be performed during decoding.
[0063] 7. Recombine all the data blocks after BCH encoding and removal of parity bits into the complete image data Finally, recombine all the data blocks after BCH encoding and removal of parity bits into the complete image data. This step usually involves parsing and reorganizing the binary data to restore the original image structure.
[0064] According to the block division method of the original image and the block size of BCH encoding, recombine the encoded data blocks into an 8x8 DCT coefficient matrix. If the parity bits are removed, additional measures need to be taken during decoding to detect or correct errors. Perform inverse quantization and inverse DCT transformation on the recombined DCT coefficient matrix to restore the original image pixel values.
[0065] S204. Extract texture features from the image using the local binary pattern processing model.
[0066] 1. Before any processing, it is first necessary to load and preprocess the image. This usually includes steps such as grayscale conversion (if the image is color), size adjustment, and necessary noise elimination.
[0067] Grayscale conversion: Convert the color image to a grayscale image to reduce the complexity of subsequent processing.
[0068] Size adjustment: Adjust the image size to ensure it is divisible by a 3x3 partition, or set a unified processing size as needed.
[0069] Noise elimination: Use a filter (such as a Gaussian filter) to eliminate noise in the image, which helps improve the accuracy of subsequent texture feature extraction.
[0070] 2. Evenly divide the preprocessed image into multiple independent 3x3 small blocks.
[0071] Calculate the adaptation of the image size to the 3x3 blocks to determine how to best divide the image.
[0072] For pixels on the edges that are not divisible by 3, strategies such as padding (e.g., copying edge pixels) or cropping (discarding extra pixels) can be adopted.
[0073] 3. For each 3x3 small block, calculate its grayscale histogram. The histogram can reflect the number of pixels with different grayscale levels within the small block.
[0074] Set the grayscale range (usually from 0 to 255 for 8-bit grayscale images).
[0075] Traverse each pixel in the small block and count at the corresponding position in the histogram according to its grayscale value.
[0076] Generate a grayscale histogram for each small block, with a length equal to the number of grayscale levels (e.g., 256).
[0077] 4. To reduce the impact of lighting changes on texture features, the histogram of each small block can be normalized.
[0078] Divide each value in the histogram by the total number of pixels in the small block to obtain the normalized frequency value.
[0079] This makes the texture features more robust and less affected by changes in image brightness and contrast.
[0080] 5. Convert the normalized histogram data of each small block into a vector form. These vectors are the local texture features of the image.
[0081] If the grayscale level is 256, the length of the texture feature vector for each small block is 256.
[0082] Extract the texture feature vectors of each small block in sequence and arrange them in a certain order to form a global texture feature matrix.
[0083] 6. To further reduce the size of the feature vectors and improve their discriminative ability, dimensionality reduction or feature selection can be performed on the texture features.
[0084] Dimensionality reduction: Use methods such as principal component analysis (PCA) and linear discriminant analysis (LDA) to reduce the dimensionality of the feature space.
[0085] Feature selection: Select the most discriminative feature subset to reduce noise and redundant information.
[0086] In one embodiment of the present invention, based on step S3, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.
[0087] S301. Extract the first audio feature from the audio stream using Mel - Frequency Cepstral Coefficients.
[0088] Pre - process the extracted audio data, including framing, windowing, pre - emphasis, etc.
[0089] Perform a fast Fourier transform (FFT) on the pre - processed audio data to convert the audio signal from the time domain to the frequency domain.
[0090] Use a Mel filter bank to filter the spectrum and calculate the logarithmic energy of each filter.
[0091] Perform a discrete cosine transform (DCT) on the logarithmic energy to obtain MFCC features.
[0092] Multiple MFCC coefficients (such as 13) can be extracted as needed, and their first - order or second - order differences may be added to represent the dynamic characteristics of speech.
[0093] S302. Extract the second audio feature from the audio stream using constant Q - transform technology.
[0094] Perform a CQT transform on the extracted audio data to obtain a time - frequency spectrum matrix.
[0095] In one embodiment of the present invention, based on step S4, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.
[0096] S401. Extract multiple frames of images from the video and extract local texture features for each frame. Subsequently, sort these local texture features according to their respective timestamps to generate an ordered video feature sequence.
[0097] Use a video processing library (such as OpenCV) to extract images frame by frame from the video. Apply a texture feature extraction algorithm (such as LBP, HOG, etc.) to each frame of the image to obtain a series of local texture features. Record the timestamp corresponding to each local texture feature, which is usually achieved through frame numbers or the time code of the video. Sort the local texture features according to the timestamps to form a video feature sequence.
[0098] S402. Match the timestamps of the audio features with the standard timestamps of the video features. This is typically achieved by calculating the time difference between the audio feature timestamps and the standard timestamps.
[0099] Extract features (such as MFCC, CQT, etc.) from the audio signal and record the timestamps corresponding to each feature. Traverse the audio features and calculate the time difference between each audio feature timestamp and each standard timestamp in the video feature sequence. Find the minimum time difference corresponding to each audio feature, which represents the time offset between the audio feature and the closest frame in the video feature sequence.
[0100] S403. Determine the validity of the audio features based on the calculated minimum time difference. If the minimum time difference does not exceed a preset time threshold, the audio feature is considered valid; otherwise, it is determined to be invalid.
[0101] Set a reasonable time threshold, which depends on the accuracy requirements of video - audio synchronization. For each audio feature, check whether its minimum time difference is less than or equal to the time threshold. According to the determination result, classify the audio features into two categories: valid and invalid.
[0102] S404. Filter out all valid audio features and sort them according to their timestamps to generate an ordered audio feature sequence.
[0103] Remove all features determined to be invalid from the audio feature set. Sort the remaining valid audio features according to their timestamps. Form an ordered audio feature sequence, where each feature in the sequence corresponds to a certain frame or set of frames in the video feature sequence.
[0104] In an embodiment of the present invention, based on step S5, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation scheme.
[0105] S501. Calculate the weighted sum of these two feature sequences to generate a unified feature sequence.
[0106] When calculating the weighted sum, assign different weights to each feature sequence according to actual needs. The selection of weights can be based on the importance of the features, the contribution to classification, or the results of experimental verification. The calculation formula for the weighted sum can be expressed as: fused feature sequence = α * video feature sequence+β * audio feature sequence, where α and β are the weights of the video and audio feature sequences respectively, and α + β = 1.
[0107] S502. Before inputting the feature sequence into the classification model, some pre - processing steps may be required to ensure that the format, dimension, and range of the feature sequence match the input requirements of the classification model.
[0108] If the dimension of the feature sequence is high, dimensionality reduction may be required to reduce the computational complexity and the risk of overfitting. If the range of the feature sequence is large, normalization may be needed to scale the feature values to a suitable range. If the classification model requires the input data to be in a specific format (such as vectors, matrices, etc.), then the feature sequence needs to be transformed accordingly.
[0109] S503. After preprocessing, the feature sequence is input into a pre-trained classification model, which has been trained with a large amount of labeled data and has learned how to map the input features to the target classes.
[0110] The classification model used is the Support Vector Machine (SVM). SVM is a commonly used classification algorithm, especially suitable for classification problems of high-dimensional data. When the feature sequence is input into the SVM model, the model calculates a decision function value based on its feature values, which represents the probability or confidence of the feature sequence belonging to each class. According to the decision function value, the SVM model selects the class with the highest probability or confidence as the prediction result.
[0111] S504. Based on the prediction result of the classification model, determine the category of the short video.
[0112] Read the predicted class label from the output of the SVM model.
[0113] Match the predicted class label with the preset set of class labels to determine the category of the short video.
[0114] The training method of the Support Vector Machine (SVM) is a key process, which involves multiple steps and techniques. The following is a detailed explanation of the SVM training method: I. Training Preparation Data collection: Use various methods to collect a large amount of video and audio data, which should include short videos of multiple categories. Ensure that the data is representative and can comprehensively reflect the characteristics of each category.
[0115] Data preprocessing: Preprocess the video and audio data, extract video features such as color histograms, motion vectors, edge detection, and audio features such as Mel Frequency Cepstral Coefficients (MFCC), spectral energy, pitch, etc. Serialize these features and possibly perform normalization, dimensionality reduction, etc. to ensure the quality and consistency of the data.
[0116] II. Training Process Model selection: According to the complexity of the problem and the characteristics of the data, select a suitable SVM model, including linear SVM and non-linear SVM.
[0117] Linear SVM is applicable to linearly separable datasets, while non-linear SVM maps the data to a high-dimensional space through kernel functions to handle non-linear problems.
[0118] Parameter Tuning: The training process of SVM mainly realizes the tuning of two parameters: the penalty parameter C and the kernel function parameter (such as the order of the polynomial kernel, the bandwidth of the radial basis kernel, etc.).
[0119] The penalty parameter C is used to control the degree of penalty for misclassification. The larger C is, the heavier the penalty for misclassification, which may lead to overfitting; the smaller C is, the lighter the penalty for misclassification, which may lead to underfitting.
[0120] The kernel function parameter affects the mapping method of the data and the classification effect.
[0121] Training Algorithm: Use an optimization algorithm (such as the Sequential Minimal Optimization algorithm SMO) to train the SVM model. During the training process, the algorithm will continuously iterate and update the weights and biases of the model to minimize the classification error.
[0122] Model Evaluation: Use methods such as cross-validation to evaluate the performance of the model, including metrics such as accuracy, recall, and F1-score. Adjust the model parameters according to the evaluation results to improve the generalization ability of the model.
[0123] III. Training Techniques Feature Selection: Select features with a large contribution to classification for training, which can improve the performance of the model. Feature selection algorithms (such as Recursive Feature Elimination RFE) can be used to automatically select features.
[0124] Data Augmentation: Perform augmentation processing on the data, such as rotation, scaling, cropping, etc., which can increase the diversity of the data and improve the robustness of the model.
[0125] Regularization: Prevent model overfitting by introducing regularization terms. Commonly used regularization methods include L1 regularization and L2 regularization.
[0126] Kernel Function Selection: Select a suitable kernel function according to the characteristics of the data. For linearly separable datasets, a linear kernel can be selected; for non-linear datasets, a polynomial kernel, a radial basis kernel, etc. can be selected.
[0127] In some embodiments, the short video classification system may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the short video classification system can be stored in the memory of the computer device and executed by at least one processor to perform the functions of short video classification (see Figure 1 description).
[0128] In this embodiment, the short video classification system can be divided into multiple functional modules according to the functions it performs, such as Figure 2 shown. The functional modules of system 200 may include: a data decomposition module 210, an image processing module 220, an audio processing module 230, a synchronization alignment module 240, and a classification processing module 250. The modules referred to in the present invention refer to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0129] The data decomposition module is used to decompose the short video into a video stream and an audio stream; The image processing module is used to extract key frames from the video stream and extract local texture features from the key frames; The audio processing module is used to extract audio features from the audio stream; The synchronization alignment module is used to align the local texture features with the audio features based on the synchronization of the time stamps of the local texture features and the audio features; The classification processing module is used to identify the category of the short video based on the aligned local texture features and the audio features by using a pre-trained classification model.
[0130] Optionally, as an embodiment of the present invention, extracting key frames from the video stream and extracting local texture features from the key frames includes: Decoding the video stream, identifying and extracting I frames in the video stream, and the I frame is an independent frame in video coding; Selecting the image frame with the earliest timestamp and the image frame with the latest timestamp from all I frames, and randomly extracting three image frames from the remaining I frames; Saving all the selected image frames as images in a specified format; Using a local binary pattern processing model to extract texture features from the images.
[0131] Optionally, as an embodiment of the present invention, saving all the selected image frames as images in a specified format includes: Performing a forward discrete cosine transform on the image to obtain the discrete cosine transform coefficients of the image; Using a weighting function to quantize the discrete cosine transform coefficients, truncating the discrete cosine transform coefficients, and only retaining the coefficient values within a certain range to reduce the data volume; Converting the quantized DCT coefficients into binary form, dividing the binary data into codeword blocks of size 7, and performing BCH coding processing on each codeword block; Removing the parity bits generated during the BCH coding process; Recombine all data blocks after BCH encoding and removing the parity bits into complete image data.
[0132] Optionally, as an embodiment of the present invention, use a local binary pattern processing model to extract texture features from an image, including: Divide the image into 3x3 independent blocks, count the histogram data of each small block, and arrange it into a vector, and the vector is the texture feature.
[0133] Optionally, as an embodiment of the present invention, extract audio features from the audio stream, including: Extract the first audio feature from the audio stream using Mel frequency cepstral coefficients; Extract the second audio feature from the audio stream using constant Q transform technology.
[0134] Optionally, as an embodiment of the present invention, align the local texture features with the audio features based on the synchronization of the timestamps of the local texture features and the audio features, including: Sort the local texture features in the order of their corresponding timestamps to obtain a video feature sequence; Use the timestamps of the local texture features as standard timestamps, calculate the time difference between the timestamps of the audio features and the standard timestamps, and filter out the minimum time difference; If the minimum time difference does not exceed a preset time threshold, determine that the corresponding audio feature is valid; if the minimum time difference exceeds the time threshold, determine that the corresponding audio feature is invalid; Filter out the valid audio features and arrange the valid audio features in chronological order as an audio feature sequence.
[0135] Optionally, as an embodiment of the present invention, use a pre-trained classification model to identify the category of a short video based on the aligned local texture features and the audio features, including: Calculate the weighted sum of the video feature sequence and the audio feature sequence to obtain a feature sequence; Input the feature sequence into the pre-trained classification model to obtain the category of the short video.
[0136] Figure 3The short video classification method provided by the embodiments of this application can be applied to a device. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of this application described herein and / or claimed.
[0137] Among them, the device 300 may include: a processor 310, a memory 320, and a communication unit 330. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It can be a bus structure, a star structure, or may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0138] Among them, the memory 320 can be used to store the execution instructions of the processor 310. The memory 320 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the device 300 can be enabled to execute some or all of the steps in the above method embodiments.
[0139] The processor 310 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 320, and by calling the data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor can be composed of an integrated circuit (Integrated Circuit, abbreviated as IC). For example, it can be composed of a single packaged IC, or can be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 310 may only include a central processing unit (Central Processing Unit, abbreviated as CPU). In the embodiments of the present invention, the CPU can be a single arithmetic core or can include multiple arithmetic cores.
[0140] A communication unit 330 is configured to establish a communication channel, enabling the storage device to communicate with other devices. It receives user data sent by other devices or sends user data to other devices.
[0141] The present invention also provides a computer storage medium. The computer storage medium can store a program which, when executed, may include some or all of the steps in the various embodiments provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.
[0142] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes, and includes several instructions to enable a computer device (which may be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0143] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the descriptions in the method embodiments for the relevant parts.
[0144] In several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the system or module can be in an electrical, mechanical, or other form.
[0145] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module, that is, it may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0146] In addition, in each embodiment of the present invention, each functional module can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0147] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.
Claims
1. A short video classification method, characterized in that: include: Decomposing the short video into a video stream and an audio stream; Extracting key frames from the video stream, and extracting local texture features from the key frames; extracting audio features from the audio stream; Aligning the local texture feature with the audio feature based on synchronization of timestamps of the local texture feature and the audio feature; The pre-trained classification model is used to identify the category of the short video based on the aligned local texture features and the audio features.
2. The method according to claim 1, characterized in that Extracting key frames from the video stream, and extracting local texture features from the key frames, including: Decoding the video stream, identifying and extracting an I frame in the video stream, wherein the I frame is an independent frame in video encoding; Select the image frame with the earliest timestamp and the image frame with the latest timestamp from all I frames, and randomly select three image frames from the remaining I frames; Save all selected image frames as images in the specified format; The local binary pattern processing model is used to extract texture features from images.
3. The method according to claim 2, characterized in that Save all selected image frames as images of the specified format, including: Perform forward discrete cosine transform on the image to obtain discrete cosine transform coefficients of the image; Use a weighting function to quantize the discrete cosine transform coefficients and truncate the discrete cosine transform coefficients to keep only the coefficient values within a certain range to reduce the amount of data; Convert the quantized DCT coefficients into binary form, divide the binary data into code blocks of size 7, and perform BCH encoding on each code block; Remove the check digit generated during the BCH encoding process; All the data blocks after BCH encoding and check bit removal are reassembled into complete image data.
4. The method according to claim 2, characterized in that: The local binary pattern processing model is used to extract texture features from images, including: The image is divided into 3x3 independent blocks, the histogram data of each small block is counted and arranged into a vector, which is a texture feature.
5. The method according to claim 1, characterized in that Extracting audio features from the audio stream includes: Extracting a first audio feature from the audio stream using Mel-frequency cepstral coefficients; A second audio feature is extracted from the audio stream using a constant-Q transform technique.
6. The method according to claim 1, characterized in that Based on the synchronization of the timestamps of the local texture features and the audio features, aligning the local texture features with the audio features, comprising: Sort the local texture features in order according to their corresponding timestamps to obtain a video feature sequence; The timestamp of the local texture feature is used as the standard timestamp, the time difference between the timestamp of the audio feature and the standard timestamp is calculated, and the minimum time difference is selected; If the minimum time difference does not exceed the preset time threshold, the corresponding audio feature is determined to be valid; if the minimum time difference exceeds the time threshold, the corresponding audio feature is determined to be invalid; Valid audio features are screened out, and the valid audio features are arranged in chronological order into an audio feature sequence.
7. The method according to claim 6, characterized in that Using a pre-trained classification model to identify the category of the short video based on the aligned local texture features and the audio features, including: Calculate the weighted sum of the video feature sequence and the audio feature sequence to obtain a feature sequence; The feature sequence is input into a pre-trained classification model to obtain the category of the short video.
8. A short video classification system, characterized in that: include: A data decomposition module, used for decomposing the short video into a video stream and an audio stream; An image processing module, used for extracting key frames from the video stream and extracting local texture features from the key frames; An audio processing module, used for extracting audio features from the audio stream; A synchronization alignment module, configured to align the local texture feature with the audio feature based on the synchronization of the timestamps of the local texture feature and the audio feature; The classification processing module is used to use a pre-trained classification model to identify the category of the short video based on the aligned local texture features and the audio features.
9. A device, characterized in that: include: A memory, used for storing a short video classification program; A processor, used to implement the steps of the short video classification method as described in any one of claims 1-7 when executing the short video classification program.
10. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a short video classification program, which, when executed by a processor, implements the steps of the short video classification method according to any one of claims 1 to 7.